SRE Capacity Intelligence

FI Stack Right-Sizing and Capacity Risk Management — HQ, Ardent, MemCache, Krayt. 📐 Methodology

Data freshness
Overview

Executive Summary

Capacity risk, availability impact, and rightsizing outcomes at a glance.
Loading summary...
Status: checking...
Loading recommendation distribution...

Capacity Trend & Risk ▶ collapsed

What breaks if we do nothing, what it costs, what we gain by fixing it, and execution pipeline progress.
FI Deep Dive

FI Deep Dive — Stack, Changes, Availability & Demand ▼ expanded

Look up any FI for a full picture in one place: its complete deployed stack and everything that changed (with dates, who applied it, and the change record — so a version bump the night before an incident is one search away, not a CR/PCC hunt), plus hosting availability, the incident ledger, and demand & responsiveness. Sources: Orca change-sets (who / when / CR-case, whole stack) + Thanos build_info (running versions, dated) + Snowflake hosting-SLA availability & incidents + console-sessions demand & Valkyrie responsiveness.
Enter an FI to see its complete stack and change history.
Capacity & Rightsizing

Customer Priority Queue ▶ collapsed

Capacity recommendations ranked by joint infrastructure urgency × customer pain. Source: SRE Capacity Intelligence + Support Customer Metrics Grid (Wk??).

Right-Sizing Recommendations ▶ collapsed

Evidence-backed CPU/memory sizing per deployed service — review, approve, and file tickets. (Nomad/envstack-confirmed only; candidate rows excluded.)
Capacity Forecast — About to Breach FIs projected to cross threshold within 30 days
Loading capacity forecast…
Slow Boil — Trajectory to Breach FIs trending toward breach within 90 days
Loading slow-boil projections…
Customer Impact — Before vs. After CRITICAL & HIGH priority FIs: baseline captured 2026-06-17. Fill in outcomes as right-sizings are implemented.
Loading impact data...
Reliability & Customer Signals

Service Reliability Health ▶ collapsed

Fleet-level SLI/SLO rollup by service. Error signal: Nomad alloc restart rate (24h window) — authoritative app-level health. Latency signal: Traefik mean (2h). Incidents: PagerDuty 30d.
Computed live on load · Nomad restart-rate (24h) + Traefik (2h) + PagerDuty (30d) + Grafana SLOs

Hosting Availability (SLA) ▶ collapsed

Authoritative fleet hosting availability from Snowflake (HOSTING_SLA_XLA_DATA). True available-minutes ratio, split into Q2-attributable vs other (vendor/FI) downtime. Distinct from the SLO "% meeting target" above.
Latest complete month · Snowflake SLA views (DB-first, monthly collector)

Login Availability (Valkyrie) ▶ collapsed

Live Valkyrie login alerts on capacity-risk FIs (Undersized + Near Breach) — 7-day history of capacity-correlated outage signals.
Live on every page load · Valkyrie synthetic login checks

PagerDuty Incidents (30d) ▶ collapsed

Live incident counts by capacity app. High incident rates block downsize recommendations.
Planning

Pre-Go-Live Capacity Estimator ▶ collapsed

New FI onboarding — select services, enter FI profile, and get starting allocation estimates from production analogs.

Acquisition / Integration Sizing ▶ collapsed

A live FI is acquiring/integrating a book of business — size the full stack for the go-live user step-up (measured floor + Carbon cross-check + guardrails + go-live buffer). For new/greenfield FIs use the Estimator above instead.

Regional Failover Survivability ▶ collapsed

Q2 runs every bank in two data-center regions at once. If one region goes down, can the other carry all the traffic? Each card is one direction; the table below lists every stack (one bank’s app in one region). Read-only.
System & Audit

Live Integrations ▶ collapsed

Integration status for live data sources. Green = configured, amber = not configured, red = unhealthy.

Review Decision History

Persistent audit trail — every Approve / Reject / Implement / Verify, grouped by session date. Survives server restarts.