Security Orchestration
High scrutiny and high blast radius; proof and governance matter.
This interactive experience is the shipped product itself — the same application code customers run in production, mounted read-only in your browser over a real sample journey. Not a video, not a mockup: because the demo and the product are one codebase, it can never drift from the real thing.
Inside this journey
-
Outcome Discovery
Align on current SOC workflows, high-volume alert types, staffing constraints, stakeholders, and measurable success criteria for automation.
Discovery Questions
Opening: Where your team spends the clock
- To start, how many full-time analysts staff your SOC and how are shifts organized?
- Tell me about your typical daily alert volume and the three alert categories that appear most often.
- Walk me through the most recent incident that required the longest containment time, including the manual steps analysts performed to investigate.
- On an average shift, roughly what share of analyst time is spent copying indicators between consoles versus making containment decisions?
- Which categories of consoles or tools are you most often copying data between during triage?
- Estimate how many recent incidents from the last 30 days would be suitable for building a playbook for testing.
- If automation could reduce average triage time from your current baseline to under two minutes for those incidents, provide your best estimate of reclaimed analyst hours per week.
Where delays translate into measurable risk
- If a delayed containment event cost you eight hours last quarter, quantify the business impacts that stemmed directly from that delay.
- Describe the most common failure that occurs when manual triage is rushed, and the downstream work it creates.
- How often do duplicated investigations or missed indicators occur because analysts lack a single source of shared context, measured in a typical month?
- Name the one operational failure related to triage that, if it happened during a pilot, would immediately stop the program.
- When containment is delayed, which compliance or SLA exposure is most likely to be triggered within 30 days?
The manual steps that eat your time
- Assuming current manual processes stay in place, how long until alert backlog or overtime becomes a sustained cost for the team?
- List the exact manual actions a triage analyst performs on a high-volume phishing report, in order.
- On average for a single case, count the number of distinct consoles or endpoints an analyst must query.
- Do any of those systems require credentialed access behind a VPN, jump host, or on-prem access that complicates automated connections?
- Name the role that is responsible for escalating when manual isolation or containment is delayed, and capture the escalation SLA.
- Outline the typical cleanup work or operational cost when an IOC is copied incorrectly during manual triage.
What automation must prove to earn trust
- If a single automated action led to an incorrect isolation during the pilot, would that outcome stop you from advancing to active automation?
- List the safety constraints you require before a playbook can take an action without analyst approval.
- How do you measure acceptable timing, false positive tolerance, and rollback windows for a successful pilot?
- Specify the acceptance criteria you would use to judge functional correctness and timing, including any numeric thresholds.
- Are there mandatory boards, legal reviews, or policy signoffs that must approve automated response actions before they can be promoted from advisory to active?
Who must buy in and who can stop this
- Name the single approver whose veto would pause rollout, include their role and the reason their approval is decisive.
- Please identify the budget owner and the approval process for purchasing platform licenses and ongoing costs.
- Estimate the number of SOC users and distinct roles that will need access during the pilot and ongoing operations.
- If the pilot meets acceptance criteria, describe the expected approval timeline and the next administrative step you would take.
- Are network, endpoint, and identity teams able to commit on-call time or access windows needed for playbook execution during the pilot?
Operational readiness and integration gates
- If a required integration lacks a stable API or available credentials, will that tool block the pilot or can you test a reduced scope?
- Select which system categories must integrate for the pilot, pick all that apply.
- For the categories you selected, identify who owns the API credentials and can grant access during the pilot, by role.
- Do any systems require on-prem agents, air-gapped access, or third-party vendor approvals to connect for testing?
- Are there contractual, data residency, or compliance constraints that would prevent the export, enrichment, or automated action on relevant telemetry?
- State your target timeline to have test environments, credentials, allowlists, and incident samples ready for an initial build.
- If those readiness items cannot be met within 4 weeks, will the organization pause the pilot or proceed with an adjusted plan?
The alternatives you are weighing
- Describe the specific outcome that would make you keep your current approach rather than switching to an external automation platform.
- Select the alternatives you are actively evaluating or considering for automation and orchestration.
- Has anyone proposed building equivalent playbooks internally, and if yes, estimate the headcount and months required to reach parity.
- If you stayed with your current vendor, identify the single capability they would need to add to keep your business.
- Identify the one option most likely to win this procurement and the specific reason it would be chosen.
Decision rules, acceptance criteria, and financial thresholds
- If the two pilot playbooks meet your timing and correctness goals on two live incidents, what is the single action you will take that week?
- Pick up to three acceptance criteria that are most important for pilot success.
- Provide the minimum quantitative target for reclaimed analyst hours per week that would justify purchase in your view.
- State the single financial metric your budget owner will use to approve spend, for example annual cost saved or hours recovered.
- If the pilot misses the top acceptance criterion, do you want to iterate with adjusted scope or stop and revisit later?
- Select the level of access playbooks will require during the pilot.
Next practical steps and timeline
- If committing a named pilot owner would accelerate decision by one week, would you assign someone now?
- Provide target dates for pilot start, selection of two incidents to test, and a two-week evaluation window.
- Name the day-to-day contact for the seller during build and test, include role and preferred contact method.
- List any blockers that must be resolved in the next two weeks to keep the proposed schedule.
- If the team cannot commit the required pilot resources within the timeline provided, should we pause scheduling or proceed with a reduced scope?
-
Solution Scope
Define the specific alert types and the two-to-three playbooks to build, required integrations, responsibilities, safety constraints, and acceptance criteria.
Scope Configuration
- Provision platform tenant and admin access
- Integrate SIEM via API connector
- Integrate endpoint detection & response (EDR)
- Integrate email security and phishing report ingest
- Integrate ticketing and ITSM systems
- Connect legacy on‑prem systems via bridge connector
- Develop phishing triage playbook
- Develop endpoint isolation playbook
- Automate indicator enrichment and reputation lookups
- Configure parallel API execution and throttling
- Enable advisory mode with response safeguards
- Deploy playbook test harness and incident replay
- Instrument execution timing and ROI metrics
Scope Questions
Provision platform tenant and admin access
- Do you require a dedicated tenant or a shared evaluation tenant for this engagement?
- Name the role that will be the primary admin for the tenant.
- Provide the list of admin onboarding documents you can supply (SOPs, change control, network diagrams).
- Which authentication method will you use for admin console access (single sign-on, local credentials, OAuth2)?
- Are there IP allowlist constraints or bastion requirements for administrative console access?
- How many administrator accounts need provisioning for initial rollout?
Integrate SIEM via API connector
- Which SIEM export surface will we use for alert ingestion (event API, search API, syslog export)?
- What is the SIEM alert retention window available for replay and validation in days?
- Share an example SIEM alert ID and the schema field you use for priority mapping.
- Does the SIEM require certificate pinning or mutual TLS for API access?
- Describe the SIEM API rate-limit, pagination behavior, or webhook retry semantics we should plan for.
- Identify the team that will supply SIEM API credentials and approve access.
Integrate endpoint detection & response (EDR)
- Which endpoint detection and response (EDR) actions must be callable from playbooks (isolate host, kill process, collect artifacts)?
- List the EDR agent versions and operating system images in scope for automated actions.
- Do you permit automated host isolation from playbooks or require human approval per runbook?
- Specify the acceptable false-positive threshold for isolation actions used to gate automation.
- Name any network segments, VLANs, or OT host groups that must be exempt from isolation actions.
- Name the owner responsible for rotating and maintaining EDR API keys.
Integrate email security and phishing report ingest
- Identify the inbound source that carries user-reported phishing in your environment (shared mailbox, phishing portal, ticketing).
- Which email headers or fields do you rely on to classify phishing incidents (message-id, Received chain, SPF/DKIM verdict)?
- Will you allow playbooks to perform automated sender disable or mailbox quarantine as part of triage?
- What was the average daily volume of user-reported phishing items in the last 30 days?
- Which attachment handling policy should the playbook follow for detonation, quarantine, or analyst review?
- Identify the owner of email security API credentials and mailbox access approvals.
Integrate ticketing and ITSM systems
- Which IT service management (ITSM) surface will be used for case creation and updates?
- Which ticket fields should map to SOC triage status such as severity, owner, and SLA?
- Do you require bi-directional synchronization between playbook state and the ticketing system?
- What ticket SLA thresholds must playbooks respect before auto-closure or escalation?
- Name the approver role authorized to approve automatic ticket updates from playbooks.
- Provide the ticketing API authentication method available (API key, OAuth2, basic auth).
Connect legacy on‑prem systems via bridge connector
- Which legacy on‑prem endpoints require bridging (firewall console, legacy SIEM, air-gapped EDR, management servers)?
- Describe the network path available for the bridge connector (outbound-only, VPN, jump host).
- Are there approved change-control or maintenance windows required to install a bridge connector appliance?
- List the protocols the bridge must support to reach legacy consoles (SSH, RPC, SMB, JDBC).
- Does the legacy system require agent-based access, API-only access, or database-level connectivity?
- Which team will provide firewall rule approvals and network changes for the bridge connector?
Develop phishing triage playbook
- Which phishing triage steps from your runbook should the playbook automate (URL reputation, attachment detonation, mailbox quarantine, user notification)?
- Supply an example phishing incident ID from the last 30 days that we should use for test replay and validation.
- What performance target do you require for automation versus manual triage in seconds per case?
- List the data enrichment sources the playbook must call (threat intelligence feed, URL scanner, DNS logs, internal reputation).
- Which safety checkpoints should require human approval in advisory mode (account disable, host isolate, auto-delete messages)?
- Define the acceptance criteria for the phishing triage playbook including required functional tests, allowable false positive thresholds, and maximum completion time.
Develop endpoint isolation playbook
- Which initial indicators should trigger the endpoint isolation playbook (EDR alert patterns, SIEM correlation rule ID, IOC match)?
- Specify the exact isolation action required from the playbook (network block at host, disable user account, revoke VPN token).
- Which escalation path and on-call rotation should the playbook notify when isolation actions are taken?
- What minimum corroborating evidence threshold must be met before isolation (EDR confidence score, SIEM correlation, multi-sensor match)?
- Identify any whitelisted hosts or host groups (OT controllers, management servers) that must be exempt from isolation.
- How will you validate acceptance for endpoint isolation (test incidents, rollback window, acceptable false isolation rate)?
Automate indicator enrichment and reputation lookups
- Which indicator types should be enriched by automation (IP, domain, URL, file hash)?
- List the external feeds or internal threat intelligence sources you require for enrichment.
- What maximum enrichment latency per call is acceptable to meet your SLA for triage workflows?
- How should the system reconcile conflicting reputation results (most recent, highest severity, require analyst review)?
- Does enrichment require rate-limiting, batching, or caching to avoid upstream API throttling?
- Name the owner responsible for mapping feed fields to your SIEM or endpoint detection and response (EDR) schema.
Configure parallel API execution and throttling
- What parallel execution limits does your environment permit per integration (concurrent API calls)?
- List any integrations that have strict API quotas or burst limits requiring special throttling.
- Which playbook steps must remain serialized for safety even when other steps run in parallel (account disable, host isolate)?
- Which backoff strategy should be used on API 429 rate-limit responses (exponential, fixed delay, queue)?
- What monitoring metrics should alert for throttling events (error rate, latency, retry count)?
- Name the team responsible for updating integration throttle settings if APIs change.
-
Hands-On Evaluation
Build and execute two playbooks on recent real incidents to validate technical fit and acceptance criteria (functional correctness, time-to-complete, advisory-mode safety).
- decision_readiness
- current_state
- gaps
- stakeholders
- desired_state
- success_criteria
- desired_state
- success_criteria
- stakeholders
- gaps
- current_state
- decision_readiness
- stakeholders
- decision_readiness
- current_state
- decision_readiness
- decision_readiness
- decision_readiness
- decision_readiness
-
Mutual Commit
Finalize commercial and access terms, confirm responsibilities, and record the acceptance criteria and go/no-go conditions to proceed to rollout.
Agreement Modules
- Subscription Agreement (Order Form)
- Master Services Agreement (MSA)
- Statement of Work (SOW)
- Service Level Agreement (SLA)
- Data Processing Addendum (DPA)
- Access & Credentials Authorization
- Acceptance Criteria and Go/No-Go Signoff
- Change Order Agreement
-
Deployment
Lock readiness facts and configuration values before execution begins.
-
Pre-Deployment Readiness
Capture concrete readiness facts — environments, owners, access windows, allowlists, and rollback constraints — the deployment depends on.
Pre-Deployment Questions
Environment and site access
- Which environments will the deployment touch? (select all that apply)
- Is production access for the platform connectors and orchestration APIs ready?
- If access is not yet ready, what is the target date when production access and any required allowlist entries will be available? (so we can schedule the cutover window)
Data and configuration scope
- Which alert types and initial playbooks are in scope for this rollout? (select primary types; we will request the playbook list in DeploymentConfig)
- Who owns source-of-truth mappings and playbook acceptance (per system category — e.g., SIEM, EDR, Ticketing)? Select the ownership model
- Provide the named owner (name and role) responsible for configuration approvals and the primary contact for each system category above (so we have a single point to validate mappings and changes)
People, approvals, and change control
- Have the deployment approvers and emergency rollback approvers been identified?
- Provide the approved maintenance windows and any blackout periods for production work (dates/times + timezone). Please include why the window is required (so we can schedule the deployment and testing)
Timing, network, and safety constraints
- Are connector allowlists, IP exceptions, or firewall changes required before deployment?
- What rollback constraints or escalation steps must the deployment team follow if an automated action must be reversed? (select the model that matches your process)
- Has the advisory-to-automated gating and the approver who will authorize moving a playbook to full automation been defined?
-
Configuration Details
Lock connector credentials, API endpoints, integration mappings, playbook parameters, and any on-prem/legacy endpoint details the deployment team will use.
Configuration Details
Deployment Snapshot
- Enter the target deployment environment name (single value; exact string used in the platform environment selector — e.g., 'prod', 'staging', 'pilot')
Locking Down Connectors & Secret Handling
- Select the authentication method the connector(s) will use (Default: API key). This choice is consumed by the connector settings page and determines which non-secret identifiers we will request.
- Enter the connector instance name to create in the platform (single value; short alphanumeric, used verbatim in the connector settings page — e.g., 'siem-prod-1')
- Provide your secrets store name and secret identifier for the connector credentials (non-secret: format '<vault-name> / <secret-key-name>' — the secret value will be loaded via your vault at deployment)
- Who is the credential owner responsible for granting platform access to the stored secret? (enter team or role email; e.g., 'platform-ops' or '[email protected]')
Endpoint Reachability & Network
- Enter the primary integration API endpoint URL the connector will call (format: https://api.example.com or https://host.example.local — this value is used verbatim in connector endpoint configuration)
- Is the primary integration API endpoint publicly reachable from the platform? (This toggles whether we configure public API calls or require private connectivity options)
- If the endpoint is NOT publicly reachable, enter the private access hostname or jump-host FQDN or CIDR the platform must use to reach it (enter 'N/A' if publicly reachable)
Mapping the Signals (Field-level)
- Provide the exact source field name your SIEM/alert source uses as the unique alert identifier (single value; exact label used for de-duplication and case linking — e.g., 'event_id' or 'case_number')
- Provide the exact source field name that contains the primary indicator value your playbooks should act on (single value; exact label — e.g., 'url', 'ip', or 'file_hash')
Playbook Runtime & Safety Parameters
- Default playbook execution timeout in seconds (numeric — Default is 120 seconds; entered value is used in playbook runtime settings)
- Default advisory-mode duration for newly deployed playbooks (select one — Default: 'Advisory-only for 2 weeks')
- Maximum parallel API calls allowed per playbook (numeric — Default: 8; used to configure concurrency caps in the execution engine)
Handoff & Deployment Ownership
- Enter the deployment owner role or team responsible for executing the rollout (single value; exact team/role name used for assignment and notifications — e.g., 'SOC Ops' or 'Platform Engineering')
-
Deployment
Execute rollout of prioritized playbooks, validate parallel execution across integrations, and transition approved playbooks from advisory to automated response.
-
-
Success
Measure reclaimed analyst hours, confirm playbook reliability and safety, and track issues and enhancement requests for continuous improvement.
Success Reviews
- Go-live Health Check (Week 1-4)
- First Measurement Review (Week 4-10)
- Acceptance Gate Review (Day ~90)
- Quarterly Success Review (Ongoing)
Issues & Enhancements
- Initiate a safety audit for any playbook that caused reopened cases and schedule remediation work.
- Produce a clear pass or fail verdict for each numeric acceptance criterion recorded in Hands-On Evaluation.
- Capture a documented go or no-go decision by the buying owner and publish it to the shared workspace.
- Define remediation tasks and timelines for any failed criteria so the loop is closed within the agreed remediation window.
- Publish the acceptance decision record showing pass/fail per criterion and the buying owner's documented decision.
- Log remediation tasks for failed criteria with acceptance tests and completion dates.
- Schedule a verification checkpoint after remediation completion to re-measure the affected metrics.
- Trend review for reclaimed hours and playbook success rate
- Confirm the solution continues to meet weekly reclaimed hours and playbook success rate expectations relative to targets in Hands-On Evaluation, or document corrective plans.
- Close or move forward the top enhancement requests and safety remediations for the next quarter.
- Determine whether to keep the quarterly cadence or add monthly operational check-ins for faster blocker burn-down.
- Publish the quarterly realization dashboard showing reclaimed hours and playbook success rate trends with annotated incidents.
- Prioritize and schedule the top 3 enhancement requests or fixes for the upcoming quarter with acceptance tests.
- Reconfirm acceptance criteria and owners
- Confirm the prioritized playbooks are executable in advisory mode with integrations reporting telemetry.
- Create a visible list of open blockers with owners and resolution target dates.
- Confirm who will publish the first measurement data feed for the next review.
- Publish a deployment verification checklist showing connector status, environment access, and advisory-mode telemetry availability.
- Log and assign remediation tasks for critical blockers with expected resolution dates.
- Enable or verify monitoring that will capture playbook run counts and basic success/failure flags for the first measurement.
- Present first measurement data
- Determine whether hours reclaimed per week and median playbook time-to-complete meet the numeric targets recorded in Hands-On Evaluation.
- Document root causes for any metric shortfalls and agree concrete remediation tasks and dates.
- Confirm the data sources and report schedule that will feed the Acceptance Gate meeting.
- Export and publish the playbook run dataset that produces the hours-reclaimed and time-to-complete metrics for audit.
- Implement agreed playbook changes or connector fixes and schedule verification runs in advisory mode.
- Produce an updated timeline to the acceptance gate showing remediation completion and validation steps.
- Restate acceptance criteria and numeric targets
- Deployment and connector validation
- Present outcome data against each criterion
- Safety incidents and reopened case review
- Diagnose gaps and root causes
- Enhancement requests triage and prioritization
- Agree corrective actions and owners
- Early adoption signals and usage patterns
- Document pass or fail per criterion and record decision
- Confirm readiness timeline to acceptance gate
- Agree remediation plan for any failed criteria
- Open issues and next quarter objectives
- Blockers, defects, and open items
- Agree immediate remediation actions