Defect Analysis
Complex deployments where integration, safety, and operational handoff determine production success.
This interactive experience is the shipped product itself — the same application code customers run in production, mounted read-only in your browser over a real sample journey. Not a video, not a mockup: because the demo and the product are one codebase, it can never drift from the real thing.
Inside this journey
-
Outcome Discovery
Align on defect patterns, target scrap-reduction goals, stakeholders, required data sources, and pilot success signals.
Discovery Questions
A quick line snapshot
- How would you summarize the current scrap problem on the candidate line in one sentence?
- Describe the defect capture workflow on a typical shift, from detection to where the record lands today.
- In the last three months, how many consecutive months has scrap exceeded your target?
- Which role on your team will own coordinating the pilot and granting MES or data access?
Where the scrap risk actually shows up
- What single defect pattern, if unresolved, would force you to stop shipments or trigger a major corrective action?
- Walk me through the recurring defect types on that line and how operators currently classify them during a run.
- How many defects per million opportunities does the line record when the spikes occur, using your best estimate or historical numbers?
- When a defect spike appears today, who is notified and what immediate containment steps are taken?
- Describe the last incident where a material lot, tool change, or environment shift coincided with a defect surge, and what you learned from it.
What stands between you and a faster root-cause
- Who on your floor has authority to stop a line over quality issues, and would they support a pilot that changes a small part of operator workflow?
- Do you have a formal defect taxonomy today, and if so, which of these best describes it?
- Which MES endpoints or exports currently include defect codes, timestamps, cycle IDs, and lot numbers?
- Have past attempts to reduce scrap created alert fatigue or added paperwork that hurt adoption?
- What breaks downstream when operators must spend more time logging defects or following extra steps?
- If we couldn't access one of the required MES endpoints within 4 weeks, would you still proceed with a pilot on that line?
The other options your team is weighing
- If the incumbent vendor or your existing toolchain could prove the same root-cause link in 4 weeks, what would keep you from staying with them?
- List the types of alternatives you are evaluating right now for defect analytics.
- Under what conditions would your team decide to keep the current approach rather than switching to a new platform?
- Has anyone inside your organization proposed building this capability internally instead of purchasing from an outside vendor?
- Would a working integration to your vision systems and SPC data make you less likely to switch away from your incumbent?
- Name the single proof or result that would make you change suppliers within two weeks of pilot close.
Can we actually run the pilot here?
- Do your MES, vision, and SPC systems provide APIs or scheduled extracts we can access under your security policies?
- Identify the owner for each integration point (MES, vision, SPC) and the typical SLA for granting access requests.
- Is 90 days of historical line data, including part IDs, timestamps, defect labels, lot IDs, and key process parameters, stored centrally and permissioned for export?
- Can you commit a subject matter expert to validate field mappings and test API credentials for 4 to 6 hours per week during the pilot?
- Are there regulatory approvals, data-sharing agreements, or plant access permits that typically take longer than 4 weeks to obtain for projects like this?
- Would a missing integration endpoint prevent the pilot from proceeding?
The measure of success we need to see
- At what measurable change in scrap rate would your plant manager commit to a production rollout?
- List the KPIs you need validated during the pilot, for example scrap reduction percentage, false positive rate, operator entry time, or mean time to root cause.
- Specify the maximum operator entry time that keeps cycle time acceptable, in seconds or steps.
- Provide the minimum monthly dollar savings your finance team needs to see to approve the production spend.
- Name the person or role whose signature would start procurement within seven days of pilot success.
Decision steps and next practical moves
- Outline the decision steps and expected dates that would allow procurement to issue an order within 30 days if the pilot succeeds.
- Confirm whether budget exists for a production rollout this quarter and which budget owner controls that spend.
- Provide the roles and backup approvers needed to greenlight deployment so we can plan the stakeholder invites.
- Estimate the typical response time from your IT or integration team to credential requests and API tests.
- Assuming the pilot meets the targets we agree, can procurement complete contract signature within 14 days?
-
Pilot Evaluation
Run a 4–6 week hands-on pilot: load 90 days of line data, configure the defect taxonomy, validate MES integration, and verify the platform surfaces known root-cause correlations.
- desired_state
- current_state
- decision_readiness
- stakeholders
- gaps
- success_criteria
- stakeholders
- decision_readiness
- current_state
- desired_state
- success_criteria
- gaps
- stakeholders
- desired_state
- current_state
- decision_readiness
- success_criteria
- gaps
- decision_readiness
- decision_readiness
- decision_readiness
- decision_readiness
-
Solution Scope
Define scope for production rollout: lines, integrations, taxonomy granularity, training, acceptance criteria, and measurable ROI targets.
Scope Configuration
- Load 90-Day Production and Defect Data
- Ingest Machine-Vision Images and Metadata
- Connect and Sync MES Defect Streams
- Configure Defect Taxonomy and Entry Forms
- Deploy Shop-Floor Defect Capture UI
- Map Material Lot and Traceability Links
- Ingest Process Sensor and Equipment Data
- Configure SPC Rules and Real-Time Alerts
- Run Correlation Engine and Root-Cause Models
- Create Root-Cause Correlation Dashboards
- Automate Pareto and Trend Reporting
- Tune Alert Sensitivity and Reduce False Positives
- Deploy Platform to Additional Lines
Scope Questions
Load 90-Day Production and Defect Data
- Specify the line IDs whose contiguous 90 days of production and defect records we should load (e.g., Line 1, Line A).
- Provide the export file formats you will supply for the 90-day load (choose all that apply).
- Indicate the primary key that links production rows to defect records on each line (e.g., serial number, cycle timestamp, part-id + position).
- Attach the typical per-shift or per-day row volume for the specified line(s) to help size ingestion (parts/hour or rows/day).
- Estimate the proportion of records expected to include defect codes or annotations in your historical export (used for model training).
- Confirm the acceptance criteria for the 90-day data load in terms of record completeness and timestamp coverage (example: >=99% timestamps present, max 5% null critical fields).
Ingest Machine-Vision Images and Metadata
- List the camera systems or image sources on the line and their primary output (e.g., 2D inspection camera, high-speed line-scan, AOI), including any camera model identifiers you can share.
- Name the image file formats and typical resolution your vision system produces (e.g., JPEG 2048x1536, PNG, TIFF, raw frames).
- Identify how each image is linked to a part or cycle in your current data (e.g., cycle timestamp, part serial printed on label, camera timestamp + conveyor encoder).
- Describe your image retention policy for inspection frames on the line (e.g., 30 days on local NAS, 12 months in cold storage).
- Specify whether metadata streams (exposure, inspector verdict, AOI bounding boxes) are exported alongside images and in what schema (CSV, JSON per image, protobuf).
- Select the connection method available for image ingestion from your vision controller (choose all that apply).
Connect and Sync MES Defect Streams
- Which MES endpoints expose defect/event streams for the line (name the endpoint type: REST API, message queue, database table, or flat-file export).
- How many defect code sets (e.g., station-specific defect lists) must be mapped from the MES to the platform for this line?
- Does your MES support push notifications or only scheduled pulls for defect updates on the line?
- Identify the field names in your MES defect payload that link to production (for example: workorder_id, operation_id, lot_id, timestamp).
- Provide the expected latency tolerance for MES-to-platform sync for real-time alerts (e.g., under 30 seconds, under 5 minutes).
- Confirm the acceptance criteria for MES defect stream synchronization (for example: no more than X minutes lag and >=Y% event delivery accuracy).
Configure Defect Taxonomy and Entry Forms
- Specify the target taxonomy depth for defect codes on the line (levels of granularity, e.g., category > family > subtype).
- Provide three representative defect examples and the current label operators use on the shop floor (use case: map to taxonomy).
- Identify whether defect entry on the line must support image attachment or only text/code selection.
- Describe the expected operator selection time target for defect entry on the tablet or HMI (the evaluation target is under 10 seconds).
- State whether you require conditional entry logic on forms (e.g., selecting 'scratch' reveals depth and length fields).
- Select who will own taxonomy governance after deployment.
Deploy Shop-Floor Defect Capture UI
- List the operator devices that will run the capture UI by model or OS (e.g., Android tablet Model X, Windows HMI, iOS).
- Name the authentication method operators will use to log defects on the shop floor (e.g., badge scan, username/password, single sign-on).
- Identify the required offline behavior for the capture UI if network is lost (e.g., cache locally for X minutes, disable entry).
- Describe any accessibility or language requirements for operator screens on the line (languages, icon-only modes, large buttons).
- Choose the target maximum keystrokes or taps per defect entry you need to meet the cycle-time constraint.
- Provide the planned training window for operators on the new UI prior to go-live (hours per operator or scheduled training dates).
Map Material Lot and Traceability Links
- Identify the source of lot/traceability IDs on the line (ERP lot number, MES lot, barcode printed on pouch, supplier lot).
- Provide the scan method available at the point-of-use for lot capture (handheld scanner, fixed barcode reader, operator entry).
- Specify the trace depth required downstream and upstream (for example: link defect back to supplier lot and forward to affected workorders).
- Describe the formatting of lot IDs you use (prefixes, date codes) if parsing rules are required for mapping.
- Select whether batch-to-unit-level traceability is required (per-part serial number linked to lot) or lot-level only.
- Assign the owner responsible for providing lot master-data and batch history during rollout (team or role).
Ingest Process Sensor and Equipment Data
- List the key sensor types and tags on the line to ingest (e.g., torque sensor tag TQ01, temperature probe TMP_A, vibrational accelerometer AX1).
- Provide the sampling frequency for each sensor class you plan to ingest (e.g., 1Hz, 10Hz, 100Hz, event-driven).
- Identify the historian or PLC connection method available for sensor data (OPC-UA, historian export, MQTT, direct database).
- Describe time-synchronization constraints between sensor data and inspection frames (do you provide a common clock or timestamps need alignment?).
- Indicate whether equipment changeover and tool-event logs are available to correlate with sensor shifts (yes/no).
- Specify retention and archival requirements for high-frequency sensor streams after ingestion (e.g., raw for 30 days, downsampled thereafter).
Configure SPC Rules and Real-Time Alerts
- Select the statistical process control (SPC) rule set you want applied initially (for example: basic control limits, Western Electric rules, custom multi-rule).
- Choose the sources for the SPC metrics on the line (inspection measurements, process sensors, tester outputs).
- Indicate alert channels for SPC violations and root-cause flags (choose all that apply).
- Describe the debounce windows or alert suppression you require to avoid repeated alarms for the same event (example: suppress identical alert for 5 minutes).
- Identify the team or role that will approve SPC rule changes after deployment.
- Set the initial sensitivity preference for alerts to balance detection vs noise.
Run Correlation Engine and Root-Cause Models
- Identify three known historical root-cause relationships we should validate during model runs (for example: supplier lot X -> spike in scratch defects; tool #5 -> burr generation).
- Provide the minimum labeled examples per defect type you can supply for model validation (number of defect records with ground-truth labels).
- Describe the expected runtime SLA for model re-scoring on a line's 24-hour dataset (e.g., under 1 hour, under 4 hours).
- Indicate whether you require causal-style outputs (e.g., ranked probable causes with confidence intervals) or only correlation scores.
- Clarify any forbidden data fields for model training due to privacy or supplier restrictions (for example: supplier cost fields, PII).
- Approve the schedule for initial model runs and validation checkpoints during the pilot (give target dates or cadence).
Create Root-Cause Correlation Dashboards
- Select the core KPIs you need on the root-cause dashboard for this line (choose up to five).
- Describe the drill-down path an engineer should take from dashboard signal to evidence (example: KPI -> defect trend -> linked images -> sensor traces).
- Identify dashboard user roles that need access and their primary view (quality engineer: full filter; supervisor: summary).
- State the desired refresh cadence for dashboards during pilot and production (real-time, 5-minute, hourly, nightly).
- Select export formats required for dashboard reports (PDF, CSV of raw evidence links, Excel, scheduled email snapshot).
- Recommend any pre-built widgets or visualizations you consider mandatory (for example: Pareto by cause, heatmap of line positions).
-
Mutual Commit
Finalize commercial terms, data-access authorizations, responsibilities for integrations, and pilot-to-production acceptance gates.
Agreement Modules
- Subscription Agreement
- Order Form & Payment Schedule
- Master Services Agreement (MSA)
- Statement of Work (SOW)
- Data Processing Agreement (DPA)
- Integration & Data Access Authorization
- Pilot Acceptance Certificate
- Service Level Agreement (SLA) Addendum
-
Deployment
Lock readiness facts and configuration values before execution begins.
-
Pre-Deployment Readiness
Confirm concrete readiness facts the rollout depends on — MES endpoints, data ownership, operator access, training windows, and go-live dates.
Pre-Deployment Questions
Environment and site access
- Is the MES production endpoint for the target line available and reachable? (so we can plan the integration cutover)
- Is there a staging/test environment that mirrors production for integration verification? (so we can validate without impacting live production)
- Who is the environment owner responsible for connectivity and firewall exceptions? (name, role, email)
Data and configuration
- Is 90 days of line-level production and defect data extractable and authorized for ingestion? (so we can load the pilot dataset)
- Who owns the source-of-truth for defect taxonomy and master data? (select the team responsible for taxonomy decisions)
- Has the field-mapping approach been decided (who will produce and approve the mappings)? (this determines the mapping workstream owner)
People and ownership
- Who is the named integration owner for this rollout? (name, role, email — responsible for MES/API coordination)
- Who is the named data owner for historical and streaming data access? (name, role, email — responsible for authorizations)
- Who will be the operator training lead on site? (name, role, email — responsible for scheduling operator sessions and sign-off)
Timing and constraints
- What is the target go-live date for this line? (enter a date or 'TBD' — we'll use this to build the rollout timeline)
- Are there production blackout windows or constrained hours that block any deployment or validation work? (list recurring windows or select 'none')
- Are operator accounts, shop-floor tablets, and barcode scanners provisioned for go-live, or will accounts/devices need to be created? (so we can sequence access and training)
-
Configuration Details
Capture exact configuration values the deployment team will use — API credentials, field mappings, image ingestion settings, taxonomy defaults, and SPC thresholds.
Configuration Details
ENVIRONMENTS & ENDPOINTS
- Platform instance base URL (format: https://...). Default is https://platform.example.com — enter the exact URL the deployment will configure.
- Primary MES / integration endpoint URL for this rollout (format: https://...). Enter 'None' if no MES integration for this line. This value is consumed by the MES connector during line rollout.
AUTHENTICATION & CREDENTIALS
- Authentication method for API/integration connections (select one). Secrets themselves are not collected here — specify how the secret will be exchanged: the credential owner will provide the secret via the option you select.
- Integration client identifier or integration user account name (non-secret). Enter the exact client_id or username the deployment should reference (the secret/credential will be handed off via your chosen secrets manager).
IMAGE INGESTION & TAXONOMY
- Image ingestion setting (select one). Default: Enabled with JPEG/PNG. Choose the variant the ingestion pipeline should be configured to accept.
- Default defect taxonomy granularity for operator entry (select one). Default is Medium (category + subcategory). This determines default UI fields and pilot CSV mappings.
FIELD MAPPINGS & LIMITS
- Exact MES field name (or CSV column name) that contains the defect code — enter the exact string the connector will map (e.g., defect_code).
- Exact MES field name (or CSV column name) that contains the production timestamp — enter the exact string the connector will map (timestamps must be ISO 8601 or state timezone in the field value).
- Default SPC control-limit width in sigma (numeric). Default: 3 — enter a numeric value the SPC engine will use as the default control limit.
- Historical data retention window in days (numeric). Default: 90 — enter the number of days of historical production/defect data the platform will retain for analyses.
-
Deployment
Execute the line-by-line rollout with task owners, sequencing, operator training, MES integration steps, and verification checkpoints.
-
-
Success
Confirm outcomes against success signals, capture lessons learned, and maintain a shared channel for issues, enhancements, and operator feedback.
Success Reviews
- Go-live Health Check (weeks 1-4)
- First Measurement Review (weeks 4-10)
- Acceptance Gate Review (around day 90)
- Quarterly Success Review
Issues & Enhancements
- Close resolved incident records and update the root-cause knowledge base with confirmed correlations.
- Tune SPC thresholds to reduce alert noise while preserving meaningful signals.
- Patch MES integration to capture missing fields used in root-cause correlations and confirm data flow.
- Restate acceptance criteria and numeric targets from Solution Scope
- Produce a documented acceptance decision against the Solution Scope targets.
- Agree remediation tasks, resolution dates, and a short-term realization tracking plan for any unmet criteria.
- Publish the acceptance decision document with the recorded outcomes and next steps.
- For any failed criteria, define remediation steps with deadlines and track progress to closure.
- Establish the ongoing review cadence and the dashboard views that will be used for realization tracking.
- Metric trends and variance analysis
- Confirm whether key metrics remain within Solution Scope targets and surface any long-term trends.
- Ensure outstanding blockers are progressing to resolution and that operator feedback is tracked into the enhancement backlog.
- Prioritize enhancement requests and publish a delivery timeline for the next quarter.
- Schedule operator refresher training sessions where logging time or taxonomy use is slipping.
- Re-confirm agreed success criteria and owners
- Confirm deployment health and data pipelines are operational for live use.
- Document incumbent wind-down status and next steps to prevent dual-system usage.
- Agree immediate remediation actions for any live blockers with clear completion dates.
- Archive historical spreadsheet tracking and set read-only access to prevent further edits.
- Resolve any MES endpoint or ingestion errors and confirm successful downstream data flows.
- Schedule a 30-minute operator refresher on defect logging workflows and taxonomy use.
- Present first outcome data against targets
- Establish whether the primary metrics are trending toward Solution Scope targets and document deviations.
- Agree a prioritized list of corrective actions and a clear timeline to the acceptance gate.
- Adjust defect taxonomy granularity to improve operator consistency and reduce misclassification.
- Deployment and data pipeline validation
- Open issues and blocker burn-down
- Present outcome data against each criterion
- Diagnose root causes for any metric gaps
- Operator feedback and training needs
- Document pass/fail per criterion and capture the acceptance decision
- Validate known root-cause correlations
- Early adoption and usage signals
- Enhancement and feedback backlog review
- Agree remediation plan for any failed criteria
- Agree corrective actions and timeline to acceptance gate
- Incumbent system wind-down checkpoint
- Open issues and immediate remediation actions
- Close loop and define realization tracking cadence