Predictive Maintenance
Complex deployments where integration, safety, and operational handoff determine production success.
This interactive experience is the shipped product itself — the same application code customers run in production, mounted read-only in your browser over a real sample journey. Not a video, not a mockup: because the demo and the product are one codebase, it can never drift from the real thing.
Inside this journey
-
Reliability Discovery
Align on critical assets, failure modes, available sensor and historian data, stakeholders, and measurable success signals for predictive maintenance.
Discovery Questions
Plant Snapshot: The assets and signals you actually care about
- How often do unplanned failures on your critical assets occur in a typical year?
- Which asset classes create the largest unplanned cost for your site?
- Tell me about the most recent failure on your site that cost over $100,000, including the asset, failure mode, and impact.
- Do you currently have vibration, temperature, pressure, current, or other sensors on those assets?
- Which historian or data archive categories hold the longest records for those assets at your plant?
Where maintenance tools fall short for you
- If one reliability metric on your site had to improve this quarter to prove value, which would it be and why?
- How many false alerts per month does your maintenance team tolerate before trust erodes?
- Who loses credibility internally on your team when an alert turns out to be a false positive?
- Describe your current workflow from anomaly detection to planning a corrective work order.
- If alerts came with clear failure mode context and a recommended work scope, what would stop your planners from acting within 48 hours?
Failures you keep seeing and the data behind them
- Name the recurring failure mode on your assets that surprises you most despite inspections and maintenance.
- Estimate the number of distinct failure events of that type your team has logged in the last 24 months.
- Tell me which sensor trends or historian signals on your assets you think showed the earliest signs of that failure.
- When was the last time your team could not find enough labeled failure examples to train a model for an asset?
- Would having access to 10 past failure incidents from your plant with full sensor traces remove the biggest technical barrier to predicting that failure?
Where integrations and approvals live in your world
- Who signs the go ahead in your organization when a predictive alert would trigger a CMMS work order?
- List the CMMS, SCADA, or historian endpoint categories at your site that must be connected for the pilot to be meaningful.
- Does your CMMS support API based work order creation today, or would the integration rely on file drops or manual entry?
- Name the individual roles in your organization that own those system credentials and will approve external access.
- Would a security or vendor onboarding review lasting longer than 6 weeks stop your pilot timeline?
What's getting in the way — deal killers and showstoppers
- What single site constraint at your facility, if unresolved, would force you to cancel this project now?
- Estimate how long it typically takes your IT or OT team to provision external read only historian access.
- Describe any internal headcount or skills gaps on your team that would slow model training, integration, or validation.
- List the regulatory, safety, or union approvals at your site that could gate a pilot and who owns them.
- Is there a hard budget cap or procurement rule in your organization that would prevent piloting without a purchase order in place?
Other options on the table and what keeps you with the status quo
- Identify the external providers your team is evaluating for predictive maintenance or whether you are weighing an internal build.
- Point to the specific features or proof points those alternatives offer that would make you stay with them instead of switching.
- Does anyone inside your organization believe these problems are solvable without an outside vendor or partner?
- Specify the single measurable improvement your current approach would need to demonstrate for you to keep it as the chosen path.
- Provide the fastest realistic timeline your organization could use to halt evaluations and proceed if a pilot meets your acceptance metrics.
Are you ready to run a pilot? The technical checklist
- Identify the single integration dependency at your site most likely to delay the pilot beyond 8 weeks.
- Do you have a named data owner who can grant historian extracts, sensor mappings, and maintenance logs for the assets in scope?
- Provide the formats and interfaces where your time series data lives, for example raw CSV exports, OPC UA, historian API, or edge gateway streams.
- Are there edge compute nodes in your environment that can run models, or must everything run in the cloud?
- Please give the role and contact method for the day to day person in your organization responsible for connectivity and security during the pilot.
- Assuming historian extracts cannot be provided within two weeks, can your team support a parallel approach using raw sensor exports for the pilot?
When a pilot deserves a yes: metrics that seal the deal
- Specify the minimum average lead time and maximum false positive rate your team requires to accept the pilot.
- Choose whether your team needs diagnostic context on every alert or only on a percentage of high severity alerts.
- Select the pilot duration your team considers statistically meaningful, 4 weeks, 8 weeks, or 12 weeks.
- Give the acceptable false positive rate in the terms your planners use, for example one per 100 assets per month.
- Report the approvers in your organization who must sign deployment approval and the target time frame for signing if the pilot meets acceptance metrics.
The bottom line: cost levers and measurable impact
- Point to the single financial metric your leadership will use to judge this project, reduced reactive spend, avoided downtime, or increased throughput.
- Give your best estimate for the average cost of a critical asset failure on your site in USD.
- State the number of high criticality incidents per year on your site for which a predictive win would materially reduce spend.
- Select the benefits your organization would prioritize in any business case, faster mean time to repair, fewer emergency work orders, or reduced spare parts inventory.
- Report the budget owner role in your organization and the typical approval lead time to release funds if the pilot meets the business case assumptions.
If we agreed today, what's the fastest path to a pilot?
- Assuming a minimal scope pilot, what is the earliest calendar date your team could start given security, data access, and operations constraints?
- Recommend three assets from your site you would pick to maximize signal and minimize integration effort for a first pilot.
- Please identify the roles from your side to include in the kickoff, covering operations, IT, and maintenance.
- Choose whether your team prefers a limited scope pilot focused on prediction quality or a run to failure pilot that also tests end to end CMMS integration.
- What would need to happen in your organization in the next two weeks to get the pilot statement of work signed?
-
Solution Walkthrough
Translate how model types, diagnostics, alert context, and integration patterns will deliver the buyer's reliability outcomes using their asset examples.
Solution Experience
- Solution Walkthrough Session
- You confirm the demonstrated model-to-asset mappings would detect the failure mode you described with the lead time you need.
- Confirm the current state and its cost
- Run a sample model run on the provided historian extracts and deliver accuracy and false-positive estimates prior to the pilot kickoff.
- You confirm the alert context and integration pattern supply the diagnostic detail required to create correct CMMS work orders and preserve planner trust.
- Provide 30-90 days of historian data, recent maintenance logs, and a representative failure incident for each selected asset example.
- Map your asset examples to model types
- Demonstrate diagnostic alerts and integration patterns
- Identify 1-3 pilot assets and the primary site owner contacts for each asset, including CMMS endpoint details.
- You agree on the remaining evidence and data access needed to run a pilot that will validate prediction accuracy and false-positive rate.
- Agree on pilot acceptance metrics including required lead time, acceptable false-positive rate, and diagnostic value criteria.
- Validate that this matches your needs
- Solution Walkthrough Session
- Solution Walkthrough Deck
- Solution Brief — Solution Walkthrough
- meeting
- slides
- document
-
Solution Scope
Define asset coverage, model types, integration endpoints (CMMS, SCADA, historians), deployment topology (edge/cloud), responsibilities, and acceptance metrics.
Scope Configuration
- Ingest Sensor and Historian Data
- Preprocess and Signal-Process Time Series
- Train Equipment-Specific Failure Models
- Deploy Edge Inference Agents
- Cloud Model Hosting and Serving
- Anomaly Detection and Alert Generation
- Failure-Mode Classification and Diagnostics
- Remaining Useful Life Estimation
- CMMS Integration for Work Orders
- Maintenance Record and CMMS Data Ingestion
- Alert Management with Diagnostic Context
- Asset Health Dashboards and Plant Scorecards
- API Connectors for IIoT and SCADA
- Automated Model Retraining and Drift Handling
Scope Questions
Ingest Sensor and Historian Data
- Which asset tags and telemetry streams (example: motor_vibration_axis1, bearing_temp, suction_pressure) should we ingest for the pilot assets?
- How many months of historian archives for each pilot asset are available (specify per asset tag if varying)?
- Do any of your sensors or PLC/RTU tags use nonstandard naming conventions that require a mapping file (for example: site_tag -> asset_tag)?
- Provide the typical sampling rates for each sensor type involved in the pilot (examples: vibration 12 kHz, temperature 1 Hz, pressure 1 Hz).
- Identify the historian endpoints and access method you can provide (examples: historian archive export, ODBC database snapshot, time-series API) and the estimated timeframe to grant read access.
Preprocess and Signal-Process Time Series
- Which preprocessing steps are required to match your operations data (examples: timezone normalization, gap-filling strategy, unit conversions like psi->bar)?
- How should we handle sensor bursts or intermittent sampling on the pilot assets (options: aggregate to 1-min, flag gaps, request raw segment pull)?
- Specify any signal-processing transforms you need preserved for diagnostics (examples: FFT spectra, envelope analysis, cepstrum) for vibration or acoustic sensors.
- Are there known calibration offsets or sensor relocations in the historian that require timestamped correction records?
- Indicate the acceptable latency for preprocessing batches to be available to model training (examples: hourly, daily, weekly).
Train Equipment-Specific Failure Models
- List the specific failure modes and their historical evidence you want models trained for (examples: bearing outer race brinelling, pump cavitation, heat-exchanger fouling with recorded work orders or root-cause reports).
- How many confirmed failure events (with timestamps and corrective work order) exist in your records per targeted failure mode and asset type?
- Specify the minimum model performance thresholds you will accept from the pilot for each failure mode (examples: true positive rate >= 75%, false positive rate <= 5%, lead time >= 7 days).
- Who in your organization will approve model features or derived signals (examples: rotating equipment SME, reliability engineer) and provide subject-matter feedback during training iterations?
- Describe any asset-specific constraints for training (examples: variable-speed pump duty cycles, repeated startup transients, seasonal fouling) that must be encoded in the training labels or windows.
Deploy Edge Inference Agents
- Which sites or asset groups require edge inference rather than cloud-only inference (provide asset IDs or plant areas)?
- Specify the edge hardware profile available at the site (examples: industrial PC with 4 CPU cores, 8 GB RAM; gateway with ARM CPU) and whether OT network access allows agent install.
- Do you have existing OT change-control or security policies we must follow for edge installation (examples: maintenance window, network isolation, code signing requirements)?
- Indicate the acceptable inference latency at the edge for alerts from a given sensor stream (examples: <1s for protective, <1 min for condition monitoring).
- What evidence will validate a successful edge deployment on site (examples: agent connected to historian, inference logs for 72 hours, sample alert with diagnostic payload)?
Cloud Model Hosting and Serving
- Which cloud tenancy model do you require for model hosting (examples: single-tenant VPC, shared tenancy with logical separation)?
- Specify the maximum acceptable model response time for an API call that returns a prediction for an asset (example: 200 ms, 1s, 5s).
- Do you require model versioning and rollback controls tied to a CI/CD pipeline or is manual promotion sufficient?
- Identify any data residency or encryption requirements for hosted model inputs and outputs (examples: at-rest AES-256, region-restricted storage).
- Are there specific monitoring and audit logs you require for model serving requests (examples: per-request latency, user ID, asset ID, prediction score)?
Anomaly Detection and Alert Generation
- Which anomaly thresholds or scoring semantics should be used to trigger an alert for each sensor type (examples: envelope RMS > X, sudden spectral energy increase of Y dB)?
- How should alerts be prioritized and routed (examples: high for immediate planning, medium for next shift, low for trending) and which plant roles receive each priority?
- Do you require a suppression window or cooldown after an alert to avoid repeated notifications for the same condition (example: suppress identical alert for 24 hours)?
- Specify the minimum lead time you require from an alert to when maintenance must be scheduled for the asset class (examples: 3 days, 2 weeks).
- Indicate any regulatory or safety conditions under which alerts must escalate immediately (examples: rotating equipment overspeed events, high-pressure venting).
Failure-Mode Classification and Diagnostics
- For each targeted failure mode, list the diagnostic artifacts you expect with the classification (examples: spectral peak at 1x rpm, imbalance severity, fouling factor estimate).
- How should diagnostic confidence be represented in alerts (examples: probability score, top-3 ranked causes, required human review flag)?
- Do you have recently completed root-cause analysis reports or teardown photos that can be used to validate classifier outputs?
- Specify any domain rules to apply to diagnostics (examples: ignore transient startup spikes, exclude data during known maintenance windows).
- Who on your team will be the diagnostic SME to resolve classifier disagreements and label additional examples during model iteration?
Remaining Useful Life Estimation
- Which asset types require remaining useful life (RUL) estimates rather than binary alerts (examples: heat exchanger fouling, compressor bearing wear)?
- What planning horizon do you require for RUL outputs (examples: weeks, months) and how should uncertainty be expressed (examples: median plus 90% CI)?
- Do you have periodic inspection records or non-destructive test (NDT) results that can be used as RUL ground truth for calibration?
- Specify how RUL estimates should feed planning workflows (examples: convert to recommended lead-time for spare parts procurement, create CMMS preventive tasks).
- Are there hard operational thresholds after which RUL is no longer useful (examples: component must be replaced at 80% wear regardless of RUL)?
CMMS Integration for Work Orders
- Which CMMS endpoints and methods can you provide for work-order creation (examples: REST API, SOAP API, file drop, database direct insert)?
- List the CMMS work-order fields required by your planners that must be populated from an alert (examples: asset ID, failure code, priority, estimated labor hours).
- Will new work orders created by the platform follow an approval workflow in your CMMS, and if so what is the typical approval path and SLAs?
- Specify the acceptance criteria we must meet for successful CMMS integration (examples: work order created with all required fields, attachment of diagnostic PDF, API success response within 5s).
- Who is your CMMS technical contact that will provide sandbox credentials and map asset IDs between historian tags and CMMS asset records?
Maintenance Record and CMMS Data Ingestion
- Which historical maintenance record tables or exports are available for ingestion (examples: work_order_history.csv, labor_hours table, spare_parts usage)?
- How complete are your maintenance records for pilot assets (percentage of events with causal codes and timestamps)?
- Do maintenance records include post-failure root-cause tags or manufacturer repair codes that can be used to label training events?
- Specify any retention or PII constraints for maintenance logs that affect how long we can store and process ingested records.
- Who will provide the mapping between CMMS asset IDs and historian asset tags for join operations during model training?
Alert Management with Diagnostic Context
- Which communication channels should alerts and diagnostic context be delivered to (examples: email to reliability team, SMS for duty engineer, CMMS work order creation)?
- How much diagnostic detail must accompany each alert for the maintenance planner to act (examples: short summary + 1-page diagnostic, full spectral plots, recommended spare parts list)?
- Do you require templated alert text or structured fields to be inserted into your planning workflows (examples: prefilled failure code, estimated downtime)?
- Identify who on your team will be the primary alert recipient for the pilot and who should be CCed for visibility (roles, not names).
- Are there any shift or union rules that influence when alerts may be acted on (examples: no nonessential work after 10pm, permit-to-work requirements)?
Asset Health Dashboards and Plant Scorecards
- Which KPIs do you want on the plant scorecard for pilot assets (examples: percent of assets with high risk, mean remaining useful life, monthly false-positive rate)?
- How frequently should dashboards refresh and who needs access (examples: shift leads view live, managers weekly report)?
- Do you require role-based dashboard views (examples: planner view, reliability engineer view, plant manager view) with different data columns?
- Specify any export or data-slice capabilities you need from dashboards (examples: CSV export of high-risk assets, PDF diagnostic snapshot).
- Are there legacy reporting formats or scorecards we must reproduce to match internal governance reports?
-
Pilot Evaluation
Run a pilot against agreed acceptance criteria to validate prediction accuracy, false-positive rate, diagnostic value, and end-to-end integration with operations.
- decision_readiness
- success_criteria
- stakeholders
- gaps
- desired_state
- current_state
- stakeholders
- decision_readiness
- success_criteria
- current_state
- desired_state
- gaps
- stakeholders
- gaps
- current_state
- desired_state
- success_criteria
- decision_readiness
- decision_readiness
- decision_readiness
- decision_readiness
- decision_readiness
-
Mutual Commit
Finalize commercial, data-access, SLA and governance terms, and confirm acceptance criteria and responsibilities to proceed to deployment.
Agreement Modules
- Master Services Agreement (MSA)
- Statement of Work (SOW)
- Subscription Agreement
- Order Form — Commercial Terms
- Service Level Agreement (SLA)
- Data Processing and Access Agreement (DPA)
- Acceptance Criteria & Handoff Annex
- Governance & Escalation Charter
- Change Order Agreement
- Work Product and IP License Exhibit
- Security and Regulatory Compliance Addendum
- Termination and Transition Services Annex
-
Deployment
Lock readiness facts and configuration values before execution begins.
-
Pre-Deployment Readiness
Confirm concrete readiness facts — data availability, historian/SCADA access, CMMS endpoints, site owners, and go-live timing the deployment depends on.
Pre-Deployment Questions
Environment and site access
- Which site(s) are included in this deployment? List each site name exactly as shown in your operations systems (so we target the correct historian/CMMS records).
- For each listed site, is production historian/SCADA read access already provisioned for the deployment team?
- If access is not yet provisioned for any site, what date will read access be available? (enter a date per site or 'N/A' if already provisioned) — so we can schedule model training.
- Are CMMS work-order integration endpoints available for the sites in scope (API or integration middleware), or is CMMS integration not in scope?
Data and configuration readiness
- Is historical sensor and maintenance data for the pilot asset classes retained and accessible for the agreed model training window?
- Are tagged asset master records and maintenance history available in a single searchable source (asset registry, EAM/CMMS extract, or similar) that the deployment team can query?
- Are there any regulatory, IT, or security constraints that prevent the seller from receiving time-series or event metadata from the historian (e.g., data cannot leave-premises, strict access approvals)? If yes, briefly state the constraint and the approving authority.
People and ownership
- Who is the site/operations owner for cutover approvals and on-site access? Provide name, role, and best contact (email or phone).
- Who is the technical owner for CMMS/EAM integrations (name and role)? This person will approve integration tests and API scheduling.
- Who is the historian/SCADA data owner who can approve data extracts, retention confirmation, and access requests (name and role)?
Timing and operational constraints
- What is the earliest acceptable go-live date for this deployment? (provide a target date so we can align training, integration, and cutover milestones).
- Are there scheduled blackout windows, maintenance freezes, or other site-level constraints that would block deployment activities? If yes, list dates or 'No known constraints'.
- Is remote edge device provisioning from the seller's cloud permitted, or must provisioning and any device configuration be completed on-site by the buyer?
-
Configuration Details
Capture exact configuration values the deployment team will use — API credentials, field mappings, model training windows, thresholds, and integration settings.
Configuration Details
Environments & Endpoints
- Enter the production environment label used in deployment manifests (short, no spaces; e.g., 'prod-plant1')
- Enter your production historian/SCADA read endpoint URL (format: https://host[:port]/path). This value is consumed verbatim by the connector.
- Choose the deployment region for cloud-hosted model training (Default: us-east-1)
Authentication & Credential Handling (DO NOT PASTE SECRETS)
- Select the authentication method the seller will use to connect to your historian/SCADA (choose method only; do NOT paste secrets)
- Enter the non-secret credential identifier that matches the selected method (e.g., service-account-username or client-id). Do NOT paste tokens, passwords, or keys.
- Who will hold the integration secret and how will it be exchanged at deployment kickoff? (select one)
Field Mappings & CMMS Integration
- Enter the exact field name in your historian or CMMS that identifies the asset (case-sensitive). This maps to the seller asset_key.
- Select the CMMS work-order type the seller should request when creating an alert (this value will be written to the CMMS 'type' field)
- Enter the exact CMMS field name that should receive the seller 'diagnostic_code' (case-sensitive).
Model Training, Thresholds & Topology
- Model training window in months (integer historical data range used for initial training). Default is 12 months — enter integer.
- Alert confidence threshold for generating prediction alerts (decimal between 0.0 and 1.0). Default is 0.85 — enter numeric.
- Edge deployment required for this environment? (Default: No)
-
Deployment
Execute model training, data pipelines, CMMS work-order integration, edge provisioning, and operational handover with clear owners and milestones.
-
Go-Live Validation
Verify acceptance criteria, validate prediction performance in production, and confirm operational readiness before declaring the rollout complete.
Checklist items
- Receive written acceptance-criteria sign-off from buyer approver
- Verify production data-feed integrity for model inputs
- Execute shadow-mode prediction run in production environment
- Calculate and deliver production performance report against acceptance metrics
- Perform end-to-end CMMS work-order test using staged alerts
- Validate alert content and diagnostic context for operational use
- Confirm alerting and escalation workflows in production
- Verify monitoring, logging, and model-observability are active
- Execute fail-safe and rollback test
- Store baseline model snapshot and rollback checkpoint in secure repository
- Obtain per-site operational go/no-go sign-off
- Deliver go-live handover package and record go-live declaration
-
-
Success
Review outcomes against success signals, refine models, and maintain a shared channel for issues, enhancements, and continuous improvement.
Success Reviews
- Go-live Health Check
- First Measurement Review
- Acceptance Gate Review (90-day)
- Quarterly Success Review and Continuous Improvement
Issues & Enhancements
- Close or reassign open issues older than 30 days and update the shared tracker with expected resolution dates.
- Schedule any model retraining or threshold adjustments and document the validation protocol to be used before the acceptance gate.
- Restate Pilot Evaluation acceptance criteria and numeric targets
- Each acceptance criterion recorded in Pilot Evaluation marked pass or fail and the buyer's acceptance decision documented.
- Remediation plan with specific resolution items and dates captured for any failed criterion.
- Incumbent system decommissioning status confirmed or a retention/archival plan recorded.
- Publish the acceptance gate decision and criterion-level results to the shared workspace.
- Track and publish remediation items with resolution timelines until all acceptance criteria are satisfied.
- Confirm and document the incumbent system's decommissioning or retention-read-only plan and archive status.
- Production performance trends
- Confirm production prediction performance remains within acceptable bounds versus the numeric targets recorded in Pilot Evaluation or surface required remediations.
- Reduce the count of open issues older than 30 days and lower weekly false-positive alerts.
- Agree a prioritized quarterly improvement backlog with validation windows and deployment dates.
- Publish the quarterly performance dashboard and a short root-cause memo for any regressions.
- Prioritize the model improvement backlog and schedule retraining or A/B test windows.
- Re-confirm success criteria and ownership
- All core integrations (historian, SCADA, CMMS) verified as connected and returning expected data.
- Early operational users have access and have acknowledged at least one platform alert or work order.
- All critical blockers logged with a target resolution window and tracking owner.
- Publish the deployment checklist completion status to the shared channel within 24 hours.
- Provide an export of user accounts, roles, and first-use activity for adoption analysis.
- Resolve critical integration failures and document remediation steps and validation tests.
- Present first-window prediction performance
- Determine whether prediction true positive rate and false positive rate are trending toward the numeric targets recorded in Pilot Evaluation.
- Confirm CMMS work-order automation success rate and whether integration fixes are required prior to the acceptance gate.
- Agree a concrete remediation plan and timeline to reach the Acceptance Gate.
- Deliver the raw event-level dataset and labeling rules used to compute prediction metrics for independent validation.
- Produce an integration error log and proposed fixes for CMMS work-order failures.
- Present outcome data against each acceptance criterion
- Operational KPI review
- Deployment and integration validation
- Review CMMS work-order automation performance
- Document pass/fail per criterion and capture acceptance decision
- Early adoption and usage signals
- Open issues and incident burn-down
- Root-cause diagnosis for metric gaps
- Blockers and open issues triage
- Agree corrective actions and timeline to acceptance gate
- Model refinement and deployment pipeline
- Agree remediation plan for any failed criteria
- Agree immediate remediation actions
- Incumbent system decommissioning and data migration status
- Quarterly action plan and next checkpoints