Network Automation
Complex platform, content, and network decisions where revenue, rights, and customer experience intersect.
This interactive experience is the shipped product itself — the same application code customers run in production, mounted read-only in your browser over a real sample journey. Not a video, not a mockup: because the demo and the product are one codebase, it can never drift from the real thing.
Inside this journey
-
Outcome Discovery
Align on outage drivers, reliability goals, stakeholder roles, and measurable success criteria for automating network configuration and remediation.
Discovery Questions
Start Here: a quick network snapshot
- How many network devices are in scope for this initial evaluation, broken down by routers, switches, and firewalls?
- When was your most recent severity-one outage caused by a configuration change, and what was the primary failure mode?
- Tell me about the teams who make production configuration changes, their size, and typical shift patterns
- Estimate the weekly hours your engineers spend on manual CLI changes versus time spent troubleshooting post-change incidents
- Select the primary outcome your leadership will judge this project on
When a change breaks the network: incident realities and accountability
- If a single misconfiguration caused 50,000 subscribers to be down for four hours, who in your organization would be held accountable and what immediate actions would they take?
- Give a concrete incident example from the last 18 months where an intended configuration change produced unexpected forwarding-plane impact, include the change, scope, and remediation timeline
- On average, how long from detection to full restoration after a configuration-induced outage?
- Who first detects these incidents today, NOC, field engineers, monitoring, or customer reports, and how are they escalated?
- Which metrics does leadership track for these events, for example customer minutes impacted, SLA penalties, or incident cost
Where configuration risk concentrates and why it matters
- Which hidden configuration gaps in your estate are most likely to cause the next large outage, by region, vendor, or device role
- How many devices are currently flagged as drifting from approved baselines, provide percentage or a device count if possible
- List the specific artifacts you use as the approved baseline, for example canonical templates, vendor images, or golden running configs
- Which automated compliance checks do you run now, and which gaps generate the most false positives
- If a pilot found 30 percent of devices non-compliant, what would have to change internally for you to proceed with a phased automation rollout
How you validate changes today, and where testing breaks down
- Why do you still rely on manual smoke testing for broad config changes instead of automated pre-checks, staged rollouts, and rollback automation
- Do you have a lab that mirrors production for core routing and edge aggregation, including multi-vendor interoperability tests
- Who owns the validation test plan and is that owner enabled to sign off on rollouts
- Which of the following validation checks do you perform after a change, BGP neighbor establishment, route convergence, ACL verification, or forwarding-plane validation
- If pre-checks could prevent the top two failure modes you listed earlier, what internal approval or budget decision would be unlocked
Obstacles that would derail automated rollouts
- What single operational constraint today would force you to pause any automated change rollout immediately
- Which maintenance windows are acceptable for staged rollouts, nights, weekends, or limited windows during business hours
- Who must authorize emergency rollback in production, and how quickly can they be reached
- Is there any regulatory or contractual prohibition on automated changes or on granting vendor or third-party access to device configurations
- If you discovered a vendor device lacked API access and required manual intervention, would that stop the pilot or require a custom integration plan
Competitive landscape, what options are you weighing
- Why would you choose to keep doing manual changes or extend a vendor-specific automation tool instead of standardizing on a vendor-neutral platform
- Which alternatives are you actively evaluating, including internal tooling, vendor-specific automation, or broader IT automation platforms
- For any incumbent or vendor-specific option you consider, what must be true about its multi-vendor coverage or validation to keep you from switching
- Has anyone proposed an internal build to solve automation, and if so, what is the projected time to a working pilot
- What evaluation result or metric would make you decide to stay with your current approach and not run an external pilot
Operational readiness and integration constraints
- If APIs, credentials, or lab parity are not available within 6 weeks, can your team still run the evaluation on schedule
- Which systems must we integrate with, NMS, ticketing, CMDB, device inventory, CI CD, or others
- Who owns API tokens and device credentials today, and are those stored in automated vaults or handled manually
- What is the headcount you can commit in full-time equivalent engineers for the pilot
- Are there compliance approvals, legal reviews, or change board gates that must be satisfied before device access or production testing
- If a required integration points to a third-party system that has no API, would you be willing to grant temporary SSH access for the pilot or is that a deal blocker
Success metrics, acceptance, and next steps
- If the pilot proves a measurable reduction in configuration-caused outages, what would realistically stop you from signing a production deployment agreement within 30 days
- Select the primary acceptance metrics we should measure in the lab pilot, mean time to detect, mean time to restore, percent of successful rollbacks, compliance delta
- Who will be the decision maker with authority to sign off on pilot success and commit budget for rollout
- What budget or procurement constraints determine the latest date you can commit to a purchase decision
- If the lab pilot achieves the agreed acceptance criteria, when could you begin a staged production rollout, immediately, within 1 to 3 months, or later
- Which single missing approval or condition would prevent you from moving to production even if the pilot met all technical criteria
-
Solution Walkthrough
Walk through how intent-based automation, pre-checks, staged rollouts, and automatic rollback will address multi-vendor support, validation, and safety concerns using the buyer's scenarios and lab test plan.
Solution Experience
- Solution Walkthrough: Intent-Based Automation
- Confirm the current state and its cost to your team
- You confirm the demonstrated workflow prevents the specific manual-change outage scenario from Discovery.
- Seller prepares a mapped lab runbook that ties each buyer scenario to the exact pre-checks, staged rollout steps, and validation checks to be executed in the demo.
- You agree that the pre-checks and validation checks in the demo would catch the edge cases you worry about in production.
- Map your scenarios to the intent model
- Buyer provide the lab test plan and a list of representative devices with OS versions and access method for the demo.
- Execute the lab pre-check and staged rollout for the scenario
- Buyer define the acceptance criteria metrics for the lab runs, including rollback timeout, acceptable convergence time, and pass/fail conditions.
- You accept a concrete set of lab acceptance criteria and the next evaluation runs required to prove multi-vendor coverage.
- Demonstrate automatic rollback and remediation
- You identify any vendor feature gaps that must be resolved before production rollout.
- Seller produce a short gap log after the demo listing any vendor integrations or feature gaps requiring custom work and estimated effort.
- Confirm multi-vendor coverage and integration gaps
- Validate that this maps to your needs
- Agree remaining evidence and next evaluation steps
- Solution Walkthrough: Intent-Based Automation
- Solution Experience Deck
- Solution Brief - Intent-Based Automation Walkthrough
- meeting
- slides
- document
-
Solution Scope
Define modules, device coverage, responsibilities, lab validation tests, acceptance criteria, and explicit out-of-scope items for the engagement.
Scope Configuration
- Vendor-Neutral Intent Model Templates
- Convert CLI Runbooks to Intent Templates
- Automated Provisioning Workflow Deployment
- Pre-Change Validation and Pre-Check Rulesets
- Staged Rollout and Canary Deployment Orchestration
- Automatic Rollback and Snapshot Configuration
- Closed-Loop Change Validation Tests (Lab)
- Multi-Vendor Device Adapter Development
- Compliance Baseline Enforcement and Remediation
- Continuous Compliance Scanning and Reporting
- Forwarding-Plane Convergence Verification Scripts
- Auto-Remediation Policy Tuning and Thresholds
- Engineer Hands-On Automation Training
Scope Questions
Vendor-Neutral Intent Model Templates
- Which device families in your control plane and core edge inventory should be modeled as intent templates (for example: core routers, PE routers, aggregation switches)?
- How many distinct service intents (for example: BGP PE-to-PE connectivity, segment-routing L3VPN, access VLAN provisioning) do you need templated for the evaluation?
- Who is the owner of the canonical single-line diagram (SLD) or network topology document we should use to derive the intent templates?
- When an intent template changes, what configuration elements must be treated as immutable versus parameterized (examples: ASN, loopback addresses, VRF names)?
- Provide the expected template input parameters you require for each service intent (for example: peer ASN, neighbor IP, route-targets, interface list).
Convert CLI Runbooks to Intent Templates
- List the existing CLI runbook procedures (by name and last-updated date) you want converted into intent-driven workflows.
- Specify typical runbook preconditions we must encode as pre-checks (examples: no active BGP flaps, interface admin-up, CPU < 70%).
- Identify any runbook steps that require manual confirmation from your network operations center (NOC) before automated execution.
- Describe the canonical CLI snippets or configuration fragments that must be parameterized versus those that should remain verbatim in the template.
- Confirm whether archived CLI runbooks are stored in a versioned repository (for example a git repo, change ticket system) and provide access method.
Automated Provisioning Workflow Deployment
- Which provisioning scenarios do you prioritize for the initial deployment (examples: new VRF creation, subscriber access VLAN assignment, MPLS label distribution)?
- How many devices will participate in the first staged provisioning run (provide a count split by device role: core, aggregation, access)?
- Specify the integration endpoints for orchestration actions (examples: network management system API, ticketing system webhook, CMDB read endpoint).
- Identify any credentials or vault method you require for provisioning (examples: vault path, role-based service account, SSH key rotation cadence).
- Provide your desired rollout schedule characteristics for provisioning workflows (for example: batch size, quiet hours, inter-batch delay).
Pre-Change Validation and Pre-Check Rulesets
- Specify the exact pre-change checks you require for production-affecting changes (examples: BGP neighbor state is Established, route count within expected range, interface counters below threshold).
- Identify the authoritative sources for pre-change state (examples: live routing table, telemetry feed, syslog events, netflow samples).
- List the acceptable device health thresholds that must be met before a change proceeds (examples: CPU < 70%, memory < 80%, control-plane process up for 24h).
- Confirm whether you require a dry-run validation that renders the post-change desired config and a diff against the running-config before commit.
- What failure thresholds or evidence will define a pre-check rejection that stops execution (for example: BGP not Established, more than 5% route delta)?
Staged Rollout and Canary Deployment Orchestration
- Provide the staging strategy you prefer for canaries (examples: single device, single site, percentage of population, per-population by role).
- How many incremental stages do you want for a typical rollout (examples: canary -> small batch -> regional -> full)?
- Identify metrics to watch during a canary (examples: BGP session count, forwarding-plane drop rate, traffic loss percentage) and the telemetry source for each.
- Specify the automatic pause or manual approval points you require between stages (examples: automatic pause for 30 minutes, manual sign-off after canary).
- List the outage impact tolerance for canary failures (examples: rollback on any BGP flap, allow transient flap < 30s, allow 0.1% traffic loss).
Automatic Rollback and Snapshot Configuration
- Which snapshot cadence do you require for config backups prior to changes (options: immediate pre-change, hourly, daily, both running and startup configs)?
- Describe your preferred rollback triggers and their thresholds (examples: persistent BGP down > 30s, forwarding-plane loss > 0.5%, control-plane PID crash).
- Identify whether rollback should be automatic, automatic with notification, or manual with operator confirmation for production changes.
- Specify how configuration snapshots must be stored and retained (examples: encrypted object store with 90-day retention, versioned git repo with signed commits).
- Confirm the maximum acceptable time-to-rollback from trigger detection to completed rollback in production.
Closed-Loop Change Validation Tests (Lab)
- Provide a lab topology reference or SLD for the validation tests including device roles, link capacities, and test traffic generators.
- Which test cases must be executed in the lab to mirror your production scenarios (examples: BGP session bring-up, LDP/segment-routing label distribution, multicast join and leave)?
- Who on your engineering team will observe and sign off on lab test runs and where will you capture test evidence (examples: pcap, route table dumps, telemetry graphs)?
- List the test traffic profiles and volumes the lab should generate to validate forwarding-plane behavior (examples: 1 Gbps IMIX, ramp to 10 Gbps, BFD session churn).
- What measurable acceptance criteria will confirm the closed-loop validation passed in the lab (examples: BGP convergence < 60s, 0 packet loss for control-plane management traffic, forwarding-table delta < 0.1%)?
Multi-Vendor Device Adapter Development
- Identify the device OS families and firmware ranges that currently lack adapter support and require development.
- Specify access methods for each device family (examples: NETCONF, gNMI, SSH CLI, RESTCONF) and whether out-of-band management exists.
- Describe any custom configuration constructs or local templates on devices that an adapter must translate (examples: proprietary ACL macros, custom TCL scripts).
- Confirm whether you can provide lab devices or device images for adapter development and functional testing.
- List the telemetry and state artifacts an adapter must surface to the platform (examples: RIB/adj-RIB-ins/out, interface counters, BGP neighbors).
Compliance Baseline Enforcement and Remediation
- Which compliance frameworks or standards govern your configs (examples: internal approved baseline, NERC CIP control list, ISO network configuration guidelines)?
- How many baseline templates do you maintain and where are they stored (examples: central S3 bucket, git repo, ticketing attachments)?
- Specify the remediation actions you permit for noncompliant devices (examples: auto-apply baseline, create a remediation ticket, quarantine device).
- Identify audit frequency and reporting cadence you require for baseline compliance (examples: real-time alerting, daily report, weekly executive summary).
- Describe exceptions process for devices that intentionally diverge from baseline (examples: documented exception in CMDB, temporary change ticket ID).
Continuous Compliance Scanning and Reporting
- Which telemetry streams will be available for continuous scanning (examples: syslog, streaming telemetry, periodic config snapshots)?
- Provide the set of reports you require out of the box (examples: noncompliant devices list, drift trend by device group, remediation time-to-fix).
- Specify role-based access to compliance reports and whether exports to your ticketing system are required.
- Identify acceptable false-positive tolerance for automated compliance alerts (for example: allow up to 5% false positives before requiring manual review).
- Describe your retention and archival requirements for compliance audit artifacts (examples: keep snapshots for 1 year, signed change logs for 7 years).
Forwarding-Plane Convergence Verification Scripts
- Provide the convergence metrics you need verified after a change (examples: route propagation time to all peers, next-hop reachability, traffic microburst recovery time).
- Which test endpoints or traffic flows should be used to validate forwarding-plane convergence (examples: probe IPs, service prefixes, telemetry collectors)?
- List the scriptable artifacts we can use in the lab and production for verification (examples: traceroute scripts, BFD session monitors, flow collector queries).
- What evidence will validate forwarding-plane convergence to your standards (examples: traceroute showing expected path within X ms, no packet loss across N probes)?
- Confirm acceptable measurement windows after a change for declaring convergence (examples: 30s, 60s, 5 minutes).
Auto-Remediation Policy Tuning and Thresholds
- Specify the categories of incidents you allow auto-remediation to act on (examples: config drift, interface admin-down, flapping BGP neighbor).
- Which threshold values should trigger auto-remediation for transient faults versus persistent faults (examples: flap count, error-rate percentage, packet-loss duration)?
- Identify notification routes and escalation windows when auto-remediation runs (examples: Slack channel, paging list, 5-minute escalation).
- Describe the rollback-to snapshot policy when an auto-remediation action fails to restore desired state.
- List any change categories that must be excluded from auto-remediation for compliance or safety reasons (examples: firmware upgrade, license changes).
-
Mutual Commit
Finalize commercial and legal terms, access and data-sharing agreements, SLAs, and the sign-off criteria that permit evaluation and deployment to proceed.
Agreement Modules
- Non-Disclosure Agreement (NDA)
- Subscription & Order Form
- Master Services Agreement (MSA)
- Statement of Work (SOW)
- Service Level Agreement (SLA)
- Data Processing Agreement (DPA)
- Access & Data-Sharing Agreement
- Acceptance Criteria & Sign-Off Certificate
- Change Order Agreement
- Regulatory Compliance Addendum
- Commercial Summary & Payment Terms
-
Deployment
Lock readiness facts and configuration values before execution begins.
-
Pre-Deployment Readiness
Capture concrete readiness facts — lab/production environments, device access, maintenance windows, named owners, and rollback authorization — required before execution.
Pre-Deployment Questions
Environment and site access
- Which environments will the deployment touch? (select all that apply; this determines scope and rollback boundaries)
- Are device-level access credentials and read/write authorization provisioned for devices in scope in both lab and production? (we need this to run pre-checks and push staged changes)
- From the deployment host (the platform), are management-plane integration endpoints reachable for all devices/systems in scope? (so we can validate connectivity before execution)
Data and configuration
- Is a single source-of-truth identified for device inventory and approved configuration baselines, and does the deployment team have access?
- Has the initial device list for the rollout (per-site device count and identifiers) been finalized and shared with the deployment team?
- Have the lab validation test cases and acceptance criteria for this evaluation been agreed and documented (so go/no‑go decisions are unambiguous)?
People and ownership
- Please provide the named technical execution owner, deployment approver, and emergency network contact (name, role, and preferred contact method). These owners will receive change notifications and have authority during rollout.
- Has rollback authorization been delegated (automatic rollback allowed, and who holds manual rollback authority)? Select the current state.
Timing and constraints
- Are approved maintenance windows for each affected production environment confirmed (if yes, please provide the date/time per environment so we can schedule the cutover)?
- Are there any regulatory or change-management freeze windows, blackout dates, or dependent projects that will block the rollout? (select all that apply; if Other, you will describe it)
-
Configuration Details
Lock the exact configuration values, integration endpoints, vendor support gaps, credentials, staged rollout parameters, and validation checks the deployment will use.
Configuration Details
Environments & Integration Endpoints
- Enter the production environment name (exact inventory name used by your systems). This value is used by the deployment to target live devices (example format: prod-core-1).
- Enter the production management-plane endpoint the connector will use (format: https://your-mgmt-host.example or IP address). This endpoint is consumed by the connector setup step.
- Source inventory type for device population (select one) — this determines how the deployment reads your device list.
Authentication & Credential Handling (NO SECRETS)
- Name of the credential object in your secrets manager (enter the exact secret name/identifier; do NOT paste the secret value). The deployment will request this secret at kickoff.
- Secrets manager category (select one) — indicates where deployment will retrieve the credential.
- Credential owner (enter the person or group name responsible for approving retrieval of the secret; exact directory/group name preferred).
Deployment Parameters & Staged Rollout
- Maximum concurrent device changes during staged rollout (numeric). Default is 5.
- Rollout stage sizes (enter a single comma-separated list of ascending numeric device counts consumed by the rollout engine; default: 2,10,100). Example: '2,10,100'.
- Per-device validation timeout in seconds (numeric) — maximum time to wait for post-change forwarding-plane convergence per device. Default is 30.
Validation & Safety Checks
- Enable automatic rollback on validation failure? Default: Yes.
- Select the post-change validation checks the deployment should run (these are consumed directly by the validation engine). Choose all that apply.
- Failure threshold per rollout stage as a percentage (numeric). If the percent of failed devices in a stage ≥ this value, the rollout pauses. Default is 10.
Vendor Support & Integration Gaps
- List any vendor + model combinations NOT supported by your inventory (single comma-separated string, format: 'Vendor Model, Vendor Model'). Enter 'None' if all devices are supported.
- Will you require custom driver/adapter development for unsupported devices? (This drives an integration work item.)
-
Deployment
Execute the staged rollout with pre-checks, forwarding-plane convergence validation, monitored rollback behavior, and clear owners and escalation paths.
-
-
Success
Confirm outcomes against acceptance criteria, document operational lessons, and maintain a shared channel for issues, enhancements, and continuous improvement.
Success Reviews
- Go-live Health Check
- First Measurement Review
- Acceptance Gate and Incumbent Wind-down
- Operational Retrospective and Runbook Handoff
- Quarterly Operational Review
Issues & Enhancements
- Confirm training plan and a shared channel plus escalation path for ongoing issues.
- Restate acceptance criteria and numeric targets
- Produce a documented pass/fail outcome for each acceptance criterion recorded in Solution Scope.
- Capture a formal acceptance decision with the buyer's named signatory and date, or record the conditional acceptance plan.
- Agree the incumbent wind-down plan or retention terms and schedule any required data migration or archiving tasks.
- Publish the formal acceptance record with per-criterion pass/fail results and the buyer signatory information.
- Create a remediation tracker for any failed criteria with tasks and hard completion dates.
- Document the incumbent wind-down steps and schedule the final decommission or retention review.
- Document lessons learned from deployment and early operations
- Capture a prioritized list of operational lessons and concrete runbook updates.
- Finalize and publish operator runbooks and validation checklists for steady-state operations.
- Re-confirm acceptance criteria and owners
- Publish the operational retrospective document including prioritized lessons and runbook edits.
- Deliver updated runbooks and validation checklists for day-to-day operations and rollback procedures.
- Distribute the training plan and schedule hands-on sessions to validate operator proficiency.
- Present the quarterly metric dashboard
- Confirm whether quarterly metrics meet the Solution Scope targets or require additional remediation.
- Ensure persistent incidents are tracked with remediation owners and deadlines until closure.
- Maintain a prioritized backlog of enhancements and vendor gaps with agreed delivery timelines.
- Publish the quarterly performance report with trend charts and the remediation tracker status.
- Update the enhancement backlog with committed timelines for the next quarter.
- Prepare compliance artifacts required for the next audit window and distribute to auditors or internal stakeholders.
- Confirm all acceptance criteria and owners recorded in Solution Scope are accurate and acknowledged.
- Validate that the initial rollout shows no critical forwarding-plane failures and identify any urgent technical fixes.
- Create a short remediation list with target dates for all high-priority blockers.
- Publish the go-live health summary including open issues and timelines for remediation.
- Collect and share logs and validation artifacts for any failed staged rollout attempts.
- Update the escalation path document and circulate it to on-call teams.
- Present first measurement data
- Establish whether each named metric is trending toward the Solution Scope targets or requires corrective action.
- Document a concrete remediation plan with resolution dates that will bring off-target metrics back in line.
- Confirm whether the onboarding timeline to the acceptance gate remains realistic.
- Publish the measurement report including raw metric calculations and supporting validation artifacts.
- List specific pre-check or test additions required to reduce change failure rate and assign resolution dates.
- Schedule any necessary vendor integration work to close device support gaps identified during diagnosis.
- Deployment and rollout validation
- Runbook and playbook review and sign-off
- Present outcome data against each criterion
- Root-cause diagnosis for gaps
- Persistent incidents and root-cause updates
- Training and competency gaps
- Agree corrective actions and timelines
- Pass/fail determination and signatory capture
- Enhancement and bug backlog review
- Early adoption signals and usage patterns
- Confirm timeline to acceptance gate
- Remediation plan for any failed criteria
- Compliance and audit readiness check
- Monitoring, alerts, and validation checks tuning
- Blockers and open technical issues
- Agree immediate remediation actions
- Agree ongoing shared channel and escalation path
- Short meeting option when stable
- Incumbent system and process wind-down check
- Update risk and issue register
- Close and next steps