FI Stack Right-Sizing and Capacity Risk Management — HQ, Ardent, MemCache, Krayt.
📐 Methodology
Data freshness
Overview
Executive Summary
Capacity risk, availability impact, and rightsizing outcomes at a glance.
Loading summary...
Status: checking...
Loading recommendation distribution...
Capacity Trend & Risk
▶ collapsed
What breaks if we do nothing, what it costs, what we gain by fixing it, and execution pipeline progress.
Loading...
FI Deep Dive
FI Deep Dive — Stack, Changes, Availability & Demand
▼ expanded
Look up any FI for a full picture in one place: its complete deployed stack and everything
that changed (with dates, who applied it, and the change record — so a version bump the night
before an incident is one search away, not a CR/PCC hunt), plus hosting availability, the
incident ledger, and demand & responsiveness.
Sources: Orca change-sets (who / when / CR-case, whole stack) + Thanos build_info (running versions, dated) + Snowflake hosting-SLA availability & incidents + console-sessions demand & Valkyrie responsiveness.
Change window:to
Enter an FI to see its complete stack and change history.
Capacity & Rightsizing
Customer Priority Queue
▶ collapsed
Capacity recommendations ranked by joint infrastructure urgency × customer pain. Source: SRE Capacity Intelligence + Support Customer Metrics Grid (Wk??).
Right-Sizing Recommendations
▶ collapsed
Evidence-backed CPU/memory sizing per deployed service — review, approve, and file tickets. (Nomad/envstack-confirmed only; candidate rows excluded.)
Legend:▲Undersized▼Overallocated●Right-sized
Loading confirmed deployment recommendations...
▶Capacity Forecast — About to BreachFIs projected to cross threshold within 30 days
Loading capacity forecast…
▶Slow Boil — Trajectory to BreachFIs trending toward breach within 90 days
Loading slow-boil projections…
FI Tier × Capacity Risk Matrix
▶ collapsed
Undersized, Watch, and Overallocated FIs by tier — with MRR at risk. Top-right = highest severity.
Loading…
Customer Impact — Before vs. AfterCRITICAL & HIGH priority FIs: baseline captured 2026-06-17. Fill in outcomes as right-sizings are implemented.
Computed live on load · Nomad restart-rate (24h) + Traefik (2h) + PagerDuty (30d) + Grafana SLOs
Loading...
Hosting Availability (SLA)
▶ collapsed
Authoritative fleet hosting availability from Snowflake (HOSTING_SLA_XLA_DATA). True available-minutes ratio, split into Q2-attributable vs other (vendor/FI) downtime. Distinct from the SLO "% meeting target" above.
Live Valkyrie login alerts on capacity-risk FIs (Undersized + Near Breach) — 7-day history of capacity-correlated outage signals.
Live on every page load · Valkyrie synthetic login checks
Loading customer experience data…
PagerDuty Incidents (30d)
▶ collapsed
Live incident counts by capacity app. High incident rates block downsize recommendations.
Click "Sync PD Incidents" to fetch live data (first sync takes ~30-60s).
👻 Ghost Services — Zero-Traffic FIsFIs provisioned in Beacon with no detectable infrastructure traffic. Candidates for deprovisioning outreach.
▼
Loading…
SDK Ghost Detection — Zero-Traffic JobsLive scan: running SDK Nomad jobs with no Traefik traffic in the past 30 days. Candidates for decommission.
▼
Click Scan Now to run a live cross-reference of Nomad SDK jobs against Traefik traffic data.
Planning
Pre-Go-Live Capacity Estimator
▶ collapsed
New FI onboarding — select services, enter FI profile, and get starting allocation estimates from production analogs.
▶ How to use this tool
Step 1 — Load the service catalog
Click Load Catalog. Core Platform services pre-select automatically. If the catalog is stale, click Sync Beacon first (requires server running with Beacon access).
Step 2 — Fill in the FI profile
Type an account number or FI name in the search box — it auto-completes from the FI Directory. Confirm or adjust Tier, Member Count, Region, and Core Provider. These drive analog selection and P50/P75 scaling.
Step 3 — Select services
Check or uncheck services in the catalog. Core Platform services are pre-selected. Add any optional services (Bill Pay, Transfers, etc.) that this FI will launch with.
Step 4 — Run Capacity Estimate
Click Run Capacity Estimate. The tool finds production analog FIs with similar tier/member count and returns P50 and P75 starting allocations for each selected service — instances, CPU MHz, and RAM per app.
Reading the results
P50 = median analog — safe baseline for most FIs. P75 = upper quartile — use for large-tier or high-transaction FIs, or when the core provider has known latency (Jack Henry/Fiserv mainframe cores). Check the analog count — fewer than 3 analogs means low confidence.
Sending the estimate
Click Copy as Email on the results panel to generate a formatted email for the implementation team. The estimate is also saved to History automatically — open the drawer below to review past runs.
Note: These are starting allocations, not permanent right-sizes. The Capacity Intelligence Dashboard will score the FI after 28+ days in production and flag any adjustments needed.
1
FI Identity
2
FI Profile
3
Go-Live Parameters
4
Service Selection
Core Platform services are pre-selected. Check any additional services this FI will launch with.
No catalog loaded
Click Load Catalog above to show available services.
▾ Estimation History
Open to load history.
▾ Estimation Accuracy — Post-Live Feedback
Open to compare estimates vs. actual capacity scoring (FIs live 28+ days).
Acquisition / Integration Sizing
▶ collapsed
A live FI is acquiring/integrating a book of business — size the full stack for the go-live user step-up (measured floor + Carbon cross-check + guardrails + go-live buffer). For new/greenfield FIs use the Estimator above instead.
Loads the FI's measured stack + right-sizing metrics from the app's own store. The recommendation never falls below current production; measured utilization drives Ardent vertical / MCHQ hold.
Enter the acquiring FI and total users after go-live, then Size the acquisition.
Regional Failover Survivability
▶ collapsed
Q2 runs every bank in two data-center regions at once. If one region goes down, can the other carry all the traffic? Each card is one direction; the table below lists every stack (one bank’s app in one region). Read-only.
▶
Service Maturity Scores
SRE maturity 1–5 per service. Low-maturity undersized stacks score higher in the priority queue — less likely to self-heal when degraded.
System & Audit
Live Integrations
▶ collapsed
Integration status for live data sources. Green = configured, amber = not configured, red = unhealthy.
Loading integration status...
▶
Review Decision History
Persistent audit trail — every Approve / Reject / Implement / Verify, grouped by session date. Survives server restarts.