Operational Resilience
Regulated environments where trust, compliance, and operational resilience are non-negotiable.
This interactive experience is the shipped product itself — the same application code customers run in production, mounted read-only in your browser over a real sample journey. Not a video, not a mockup: because the demo and the product are one codebase, it can never drift from the real thing.
Inside this journey
-
Pre-Sales
Qualify and diagnose before investing in a full evaluation cycle.
-
Qualification
Confirm budget runway, decision-makers, regulatory deadlines, and timeline constraints before investing in detailed discovery.
Qualification Questions
Program fit and scope
- So we make the best use of your time — which resilience outcomes should we prioritize in discovery?
- Which critical business services must be included in the initial discovery (one or two key services is fine)?
Budget runway
- Is there an allocated budget or runway for evaluation and a pilot? Please select the range that best fits your current plan.
Decision authority and stakeholders
- Who will sign off on moving from discovery to a pilot or paid engagement? Please list role or team.
- Who else will materially influence the buy decision? Select all that apply.
Timeline and regulatory deadlines
- Is there a regulatory or board deadline that requires a documented dependency map or quarterly scenario testing?
- If you selected a dated deadline or have a target window, please state the compliance date or target go-live quarter (e.g., Q3 2026). Otherwise, leave a short note on timing.
- Are there data access, residency, or third-party contractual constraints that will affect evidence collection during discovery?
-
Enterprise Discovery
Map stakeholders, critical business services, recent outage impact, regulatory obligations, and success criteria across the buying group.
Discovery Questions
How this landed on your desk
- How did the nine-hour payment processor outage change the priorities your leadership has set for operational resilience?
- Tell the story of the board request for a full dependency map, including who will sign the annual resilience statement and which deadlines are fixed
- When did the board require quarterly scenario testing to begin, and are there fixed submission dates they expect?
- Who on your team will own day to day coordination with an external discovery team during the pilot?
- If the platform cannot access inventories for core payment systems within 14 days, would you pause the pilot?
Where your services really rely on fragile links
- Walk me through the single business service that would cause the largest regulator concern if it failed for four hours, and why
- How many distinct systems, third party vendors, and cloud regions support that service in a typical month?
- Describe the last incident where that service degraded, including which dependency surprised you during the postmortem
- Who first notices when that service breaks and which alert channels or reports reach them?
- If your most critical vendor refuses to share architecture diagrams, would that prevent you from delivering a regulator ready dependency map?
What regulators will ask for and how you would prove it
- Point to the single compliance test or examiner question that would fail your institution today, and explain what would need to change to pass
- Which regulatory frameworks matter most for this program and which types of evidence do examiners expect each quarter?
- When examiners review your resilience statement, which documentation gaps are they most likely to flag?
- Can your team produce a full dependency map and scenario log within 90 days if the board requires regulator ready evidence for the next reporting cycle?
- List the internal approvers required before evidence is shared with examiners and the legal gates that must clear
Where the current documentation is brittle
- Name the top two blind spots in your current continuity binders that would show up during a live scenario test
- Tell me which assets in your inventory have not been reconciled to owners in the last 12 months
- Which configuration and evidence sources are the hardest to access for automated pulls?
- Is any critical business service still documented only in paper or spreadsheets, and would that block deployment if left unchanged?
- Estimate the number of calendar days required to reconcile missing ownership records for your top 20 services
The other paths you are weighing
- What would have to be true about your current continuity approach for you to keep it instead of adopting a dedicated platform?
- List the external vendors, incumbent solutions, and internal projects you are actively evaluating for this capability
- For each option you are weighing, what single criterion would make you rule it out immediately?
- Has anyone proposed solving this entirely in house, and if so who would lead that effort and how would it change budget expectations?
- Rate the other vendors' demonstrations on realism versus showmanship
Readiness to share the data we need
- Can your team provide API access to your CMDB, service catalog, and vendor contract store within 21 days and who authorizes that access?
- Identify the systems that must be integrated to build a live dependency map and indicate whether public APIs or service accounts are available
- Name the team that owns API credentials and indicate whether they will issue scoped read only keys for discovery
- Estimate how quickly your legal team can approve a standard data protection addendum
- Provide the number of full time equivalents available to support integrations during the pilot
How success will be judged and who signs
- Identify the decision owners who would sign off to move to production if the pilot demonstrates traceable owners and tested recovery steps for your top services
- Specify measurable acceptance criteria that would convince your CRO to approve a full rollout, including tolerances and required evidence
- Provide the name and role of the person who will present scenario testing results to the board risk committee and indicate whether a presentation is ready
- Will an operational budget already cover scaling after a successful pilot, or will additional approvals be required?
- Indicate the expected frequency for regulator facing scenario tests after deployment
Practical next steps and likely blockers
- What single decision or resource would accelerate a pilot start to within three weeks?
- Describe the top three calendar conflicts, approval gates, or legal actions that commonly delay pilots in your organization
- Indicate the names and roles of the individuals who should receive the initial statement of work and identify who has final contracting authority
- Are there any contractual prohibitions or regulator imposed restrictions that would prevent data sharing for discovery?
- Choose the preferred week to begin pilot discovery calls
-
-
Solution Evaluation
Prove the platform's ability to build live dependency maps, set impact tolerances, and run scenario simulations against the buyer's real-world failure scenarios.
- decision_readiness
- stakeholders
- current_state
- desired_state
- success_criteria
- gaps
- desired_state
- current_state
- stakeholders
- success_criteria
- decision_readiness
- gaps
- stakeholders
- current_state
- gaps
- desired_state
- success_criteria
- decision_readiness
- decision_readiness
- decision_readiness
- decision_readiness
- decision_readiness
-
Solution Scope
Define which business services, technology and third-party dependencies, scenario types, and governance modules are included and how success will be measured.
Scope Configuration
- Import and Reconcile BCP and DR Runbooks
- Map Critical Business Services
- Map Service-to-Technology Dependencies
- Map People, Facilities, and Data Dependencies
- Ingest CMDB and Asset Inventories
- Create Impact Tolerance Framework
- Build Scenario Simulations
- Execute Live Scenario Test
- Simulate Cloud-Region and Vendor Outage
- Generate Board-Ready Resilience Report
- Prepare Regulatory Evidence Package
- Train Governance and Board Presenters
- Configure Continuous Dependency Monitoring
- Deploy Single-Point-of-Failure Alerts
Scope Questions
Import and Reconcile BCP and DR Runbooks
- Do you have a central library of business continuity plans (BCPs) and disaster recovery (DR) runbooks for payment processing and retail banking operations?
- Which document formats contain your existing runbooks for lending, payments, and claims workflows?
- How many distinct BCP or DR runbooks cover services that directly touch customer-facing payments or claims?
- When were the runbooks for the retail payments and core ledger last updated and approved by the resilience owner?
- Where are runbook owners recorded for each critical service (for example an owner field in the runbook or an approval list used at the board level)?
- Provide the biggest known reconciliation issue between runbook steps and live operational procedures for the payment clearing workflow.
Map Critical Business Services
- Which business services must be in scope for the dependency map to satisfy board reporting after the recent third-party payment processor outage (examples: retail payments, corporate payments, claims adjudication)?
- How many service lines do you classify as critical under your operational resilience policy for PRA, DORA, or similar regimes?
- Which measurable business outcome defines unacceptable impact for each critical service (for example number of failed transactions, customer-facing downtime minutes, or lost settlement batches)?
- Identify the service owners who will validate the mapped service boundaries for payment processing and lending workflows.
- Are there regulatory deadlines or exam cycles (for example PRA or ECB submission dates) that require the dependency map to be delivered by a fixed calendar date?
- Describe any recent outage impact metrics (for example nine hours of customer outages, number of failed payments) that must be reflected in the service risk score.
Map Service-to-Technology Dependencies
- Which technology components must be mapped for each service (examples: core ledger, payment gateway API endpoints, batch settlement jobs)?
- How many external API endpoints or vendor integrations are actively called by your retail payments service during peak processing windows?
- What is the required level of fidelity for mapping: logical dependencies only, logical plus environmental (region/zone), or down to specific host and container identifiers for the claims processing path?
- Which authentication or integration surfaces (for example OAuth token endpoints, TLS mutual auth, MQ channels) must be captured for vendor and inter-service calls?
- Where do you expect mapping data to be sourced from for technology components (for example CMDB, network diagrams, cloud inventory)?
- Provide any known undocumented single points of failure in your payments or lending technology stack that should be prioritized in the mapping.
Map People, Facilities, and Data Dependencies
- Which key human roles must be included in the dependency map for end-to-end payments and claims workflows (examples: payment ops, reconciliation DBA, incident commander)?
- How many physical sites (data centers, recovery sites, regional offices that perform settlements) support your critical services and must be modelled?
- Which categories of data stores must be captured and classified (for example customer transaction ledger, payment credentials, claims documents) for regulator-ready evidence?
- Who in your team maintains the runbook access lists and evidence for custody of critical data during tests?
- Are there site-specific regulatory reporting requirements (for example local supervisor exam data residency rules) that change how facilities are scoped?
- Describe any people or facility access constraints (for example key staff on rotating shifts, secure data rooms) that could limit live scenario execution.
Ingest CMDB and Asset Inventories
- Which authoritative inventory sources will you allow us to ingest (examples: your CMDB, cloud tagging, asset management database)?
- How many configuration management database (CMDB) records are expected for in-scope services (estimate order of magnitude)?
- What minimum data quality threshold will you accept after reconciliation between CMDB and discovered runtime topology (for example 95 percent asset match rate)?
- Which connector types must be enabled for ingestion (examples: cloud provider APIs, ServiceNow CMDB, network discovery tool API)?
- Identify any sensitive inventory fields that must be redacted or masked on import (examples: encryption keys, admin credentials).
- Are there scheduled maintenance windows when we should run initial inventory syncs to avoid production impact?
Create Impact Tolerance Framework
- Which regulatory framework should the impact tolerance definitions align to for this engagement (examples: PRA operational resilience policy statements, DORA requirements, OCC guidance)?
- How should unacceptable impact be measured for retail payments: transaction count failed, total customer minutes of outage, monetary value at risk, or a combination?
- Who will define final impact tolerance thresholds for lending and claims (for example chief risk officer, resilience owner)?
- Specify the cadence for reviewing and approving impact tolerances ahead of board sign-off (for example quarterly review ahead of resilience statement).
- Are recovery time and recovery point objectives already defined for settlement and batch jobs that touch third-party processors?
- What acceptance criteria will confirm the impact tolerance framework is approved for regulatory reporting (for example documented thresholds accepted by CRO and evidence in the governance repository)?
Build Scenario Simulations
- Which scenario types should be modelled first to satisfy supervisory concerns after the payment processor outage (examples: vendor outage during peak settlement, simultaneous cloud region failure and core ledger latency)?
- How realistic should injected failure behavior be: synthetic latency/failure tags at API level, simulated DB transaction locks, or full traffic replay of payment streams?
- Which data sets must be used for scenario runs to meet realism requirements (examples: last six months of payment batches, live anonymized transactions)?
- Identify any legal or privacy constraints that prevent using live customer transaction data in simulations.
- Are there internal approvals required to run simulations against production-adjacent environments for claims adjudication?
- Provide the primary success metric for scenario realism (for example percent of fault paths exercised or match to historical outage behavior).
Execute Live Scenario Test
- Which live scenario should be executed first as a proof point (for example payment gateway outage during peak hours or regional cloud disruption affecting settlement batches)?
- When do you have a validated maintenance window and incident commander roster available to perform a live scenario for payments or claims?
- Which evidence streams must be captured during the test to satisfy examiners (examples: timestamps of failed transactions, operator logs, customer impact dashboards)?
- Who will be the accountable owner for test remediation items arising from the live scenario for the settlement pipeline?
- Are there external parties (for example critical third-party vendors) who must be notified and coordinated with before executing the live scenario?
- What acceptance criteria will confirm the live scenario test met objectives for the board and regulators (for example defined impact tolerances not breached, required evidence captured and stored)?
Simulate Cloud-Region and Vendor Outage
- Which cloud regions and vendor endpoints should be included when simulating a simultaneous cloud-region and vendor outage for payment flows?
- How should failover behavior be modelled for cross-region settlement queues (for example queue replay delays, DNS failover times, or instance rebuild latency)?
- Which vendor SLAs must be tested for tolerance (for example vendor MTTR, support responsiveness, queued transactions limits)?
- Are your disaster recovery environments configured to accept injected traffic for end-to-end validation of claims or payments processing?
- Indicate any contractual obligations or escalation clauses with payment processors that affect simulation scope.
- Estimate the maximum transaction volume we should simulate to reflect your peak processing window.
Generate Board-Ready Resilience Report
- Which elements must appear in the board-ready resilience report for the upcoming risk committee (examples: service dependency map, impact tolerance table, scenario results summary)?
- How long should the executive summary be for board consumption (for example one page, two pages)?
- Who will be the named presenter on the board report for operational resilience outcomes?
- Are there specific regulator expectations we must reference in the report (for example PRA thematic letters or DORA provisions)?
- Specify the format required for board evidence packs (for example PDF binder plus CSV attachments for raw test logs).
- Describe any board-level KPIs you track that should be included (examples: quarterly scenario pass rate, number of remediation items open).
-
Mutual Commit
Finalize commercial and data-access terms, acceptance criteria for assessments and tests, and confirm program governance and timelines.
Agreement Modules
- Master Services Agreement (MSA)
- Statement of Work (SOW)
- Subscription Agreement and Order Form
- Service Level Agreement (SLA)
- Data Processing Agreement (DPA)
- Data Access and Security Addendum
- Acceptance Test Plan and Acceptance Certificate
- Program Governance and Timeline Agreement
- Change Order Agreement
- Termination and Exit Plan
- Financial Services Compliance Addendum
-
Deployment
Lock readiness facts and configuration values before execution begins.
-
Pre-Deployment Readiness
Lock owners, access permissions, environments, evidence sources, and regulatory reporting requirements the deployment depends on.
Pre-Deployment Questions
Environment and access
- Which environments will the deployment touch? List environment names (e.g., production, pre-prod, staging, DR) — this defines scope and cutover order.
- Are the named environments accessible to the deployment team today, or will access still be provisioned?
- If access is partial or not available, what is the expected date full access will be granted? (so we can schedule the kickoff)
Data and configuration
- Which authoritative evidence and dependency source categories will feed the mapping (select all that apply)? We use these to size integrations and identify automation paths.
- For the primary sources you selected, what is the access method available to the platform (choose the closest match) — this determines whether evidence collection can be automated.
- Who owns the source-of-truth records for the primary sources above? Provide name and role for each source category so we can route approvals and test access.
People and ownership
- Who is the program sponsor/overall owner for this deployment? Provide name, role, and best contact (so governance decisions are rapid).
- Please provide day-to-day owners (name and role) for these workstreams: environments & ops; integrations & APIs; security & access; regulator reporting & compliance.
- Has the buyer confirmed who will approve regulator-ready test evidence and acceptance criteria?
Timing and constraints
- Are there any blackout windows or high-risk dates we must avoid (e.g., quarter close, major releases, regulatory reporting days)? List date ranges so we can sequence work.
- Are there immovable regulatory or board deadlines tied to this deployment (e.g., upcoming resilience statement, examiner review) that set a fixed delivery date?
- Do any environments or evidence sources require third-party coordination for access (cloud provider, payment processor, vendor)? Who will obtain that coordination?
-
Configuration Details
Capture exact configuration and integration values the implementation will use — connectors, API endpoints, data mappings, and test scenarios.
Configuration Details
Deployment & Access Overview
- Primary environment name for this deployment (enter the exact instance name used in URLs and UI — Default: prod).
- Primary API base URL for production (format: https://your-subdomain.example.com — enter exact base URL the platform will call).
- Identity provider type for admin access (select the single method you will use; Default: SAML-based IdP).
- Admin SSO entity ID or OIDC client ID (enter the non-secret identifier exactly as configured in your IdP).
Connectors & Data Mappings
- Connector types to ingest topology and third-party registry data (select all that apply).
- Exact endpoint URL or file path for the primary connector you will use (format examples: https://api.example.com/endpoint or s3://bucket/path or \\server\share\file.csv). Enter one value only.
- Canonical identifier field name used for business-service records in your source system (enter the exact field name the connector will map to; Default: service_id).
Testing & Acceptance
- Primary initial test scenario to run first (select one — this guides the initial scenario build and simulation schedules).
- If you selected 'Other' above, enter the exact scenario name to use for the initial run (leave blank if not applicable).
- Enable scenario simulation with time-series data (Default: Yes).
- Maximum number of business services to include in the initial deployment (numeric — Default: 50).
-
Deployment
Execute rollout with sequenced tasks, scenario testing schedules, regulator-ready evidence capture, and clear owners for remediation items.
-
-
Success
Validate outcomes against impact tolerances, capture governance-ready evidence for examiners, and maintain a shared channel for issues and enhancement requests.
Success Reviews
- Go-live Health Check (weeks 1-4)
- First Measurement Review (weeks 4-10)
- Acceptance Gate and Formal Ratification (around day 90)
- Operational Burn-down Meeting (monthly, first 6 months)
- Quarterly Business Review, Resilience Operations (quarterly ongoing)
- Annual Resilience Validation and Governance Pack (annual)
Issues & Enhancements
- Prioritize the backlog of persistent issues and commit to a remediation plan for the quarter.
- Reduce the count of P1/P2 remediation items and shorten mean time to remediate toward the targets recorded in Solution Scope.
- Resolve or escalate any blockers that threaten upcoming scenario tests or regulator reporting deadlines.
- Confirm the evidence package readiness for the next scheduled scenario test.
- Close or re-scope remediation items older than the agreed SLA and publish updated resolution dates.
- Prepare the evidence extracts needed for the upcoming scenario test and post them to the shared governance channel.
- Escalate any unresolved cross-team blockers to the executive escalation channel for intervention.
- Quarterly outcomes vs targets
- Confirm the quarterly scenario testing pass rate and average evidence assembly time relative to the targets recorded in Solution Scope.
- Reconfirm success criteria and ownership
- Produce the regulator-ready evidence package and agree the distribution process for examiner access.
- Publish the quarterly resilience report including scenario results and the evidence package for examiner access.
- Schedule remediation work for the top 3 persistent issues and record verification criteria for closure.
- Finalize and publish the next quarter's scenario test calendar and required pre-test evidence checkpoints.
- Consolidated annual outcomes
- Finalize and approve the governance-ready evidence set for examiner review and board reporting.
- Confirm whether the percentage of critical services meeting impact tolerances and the number of evidence packages meet the targets recorded in Solution Scope.
- Agree a timebound plan to remediate any high-severity gaps before the next annual validation.
- Publish the annual governance pack and upload the evidence artifacts to the regulated evidence repository.
- Produce a prioritized remediation roadmap for any high-severity gaps with target close dates.
- Confirm the next year's scenario themes and the anticipated evidence collection adjustments.
- Confirm deployment completed and critical connectors and evidence sources are operational.
- Produce an initial list of open issues with severity and target remediation dates.
- Establish the shared channel for operational issues and enhancement requests.
- Publish the deployment validation checklist and operational runbook updates to the shared workspace.
- Resolve P1 connectivity or permission issues and confirm evidence capture is producing expected artifacts.
- Schedule the First Measurement meeting and circulate required data extracts in advance.
- Present first measurements against targets
- Determine whether the percentage of critical business services with live dependency maps and the scenario test pass rate are trending to the targets recorded in Solution Scope.
- Agree a prioritized remediation plan with timelines that closes the largest gaps before the Acceptance Gate.
- Confirm the data sources and evidence that will be used at the Acceptance Gate.
- Produce a remediation plan listing tasks, expected outcomes, and resolution dates for each metric shortfall.
- Deliver the evidence package extract that will be used at the Acceptance Gate for each evaluated service.
- Update the scenario test cadence and assign owners for the next test window.
- Restate acceptance criteria and numeric targets
- Produce a documented acceptance decision with a named signatory for the enterprise engagement.
- Confirm which acceptance criteria passed and which require remediation with committed resolution dates.
- Complete the incumbent wind-down checklist and record the decommissioning plan status.
- Publish the signed acceptance record and archive it to the governance folder.
- Execute the incumbent decommissioning tasks including contract/renewal handling and data archive confirmation.
- Schedule remediation validation sessions and list the evidence required to close each conditional acceptance item.
- Operational ticket and remediation burn-down
- Deployment and integration validation
- Present outcome data for each criterion
- Diagnose root causes for any metric shortfalls
- Open issue and enhancement backlog review
- Board and examiner evidence package
- Metric trend review
- Document pass or fail per criterion and record the acceptance decision
- Outstanding high-severity remediation plan
- Agree corrective actions and timelines
- Early adoption and usage signals
- Escalations and blocker resolution
- Regulatory reporting and evidence readiness
- Blockers and open issues triage
- Confirm timeline to the Acceptance Gate
- Continuous improvement and next year test calendar
- Incumbent system wind-down confirmation
- Update evidence capture for imminent scenario tests
- Next quarter test calendar and owners