Why investigations fail
Three failure modes dominate loss investigations in manufacturing plants:
- Jumping to the first plausible cause. A line stops, an operator says "the sensor is wrong", the sensor is replaced, the line restarts, and the real root cause — a worn-out gear box — is never found. Three weeks later the line stops again.
- No timeline. Nobody wrote down what actually happened in the 30 minutes before, during and after the loss. By the time the investigation starts the next day, witnesses have forgotten and the shift pattern has rotated.
- No evidence scoring. Three hypothetical root causes are written on a whiteboard but nobody scores them against the available data. The one with the loudest advocate wins.
The method below exists to prevent all three.
Step 1 — Trigger and scope (within 1 hour)
Within one hour of the loss being reported, define three things in writing:
- What was lost. Units, hours, euros, batches. A specific number, not "a lot".
- Where. Line, asset, workcenter, lot.
- Why we are investigating. Is this a one-off, or is this the third time this quarter? The investigation depth should scale with recurrence.
Step 2 — Reconstruct the timeline (within 24 hours)
Pull every piece of data available for the window from 2 hours before the loss to 1 hour after. Minimum sources:
- Downtime log entries with timestamps and reason codes
- Production counts per minute for the involved line
- Quality checks performed in the window
- Maintenance work orders on the involved asset in the last 90 days
- Shift handover notes
- Any operator observations captured in free-text comments
The output is a single table with one row per minute, columns for each data source. The Loss Investigator in our platform generates this automatically; without automation, allow 2-3 hours.
Step 3 — Hypothesize root causes (not conclude)
Write down at least three candidate root causes. Not one. Not "the obvious one". Three. For each candidate, write:
- What would we expect to see in the data if this hypothesis is correct?
- What would we expect to see if it is wrong?
- What evidence do we currently have for and against it?
This is the step most teams skip. Skipping it is why plants fix the symptom, not the cause.
Step 4 — Score the evidence (0-10 per hypothesis)
For each hypothesis, score the evidence quality on three dimensions:
- Pattern match — does the timeline match what this cause would produce?
- Historical recurrence — has this cause produced similar losses on this asset in the last 90 days?
- Mechanical plausibility — does the physics / chemistry / information flow make sense?
Sum the three scores. The highest-scoring hypothesis becomes the probable root cause. If no hypothesis scores above 15/30, write "insufficient evidence" and go collect more data before proposing a countermeasure.
Step 5 — Confirm with a focused test
Before implementing a countermeasure, run a focused test that would fail if the probable root cause is wrong. Examples:
- If "worn gear box" is the hypothesis, do a vibration measurement on the gearbox.
- If "operator error on changeover" is the hypothesis, review the last 10 changeovers by that operator on that line.
- If "raw material deviation" is the hypothesis, pull the lot genealogy and compare parameters to the previous 20 lots.
If the test confirms, proceed to Step 6. If the test fails to confirm, return to Step 3 and generate new hypotheses.
Step 6 — Select countermeasures and link to CAPA
Two types of countermeasures should always be proposed:
- Immediate — stops this specific loss from continuing (replace the gearbox, retrain the operator, hold the lot).
- Systemic — stops this type of loss from recurring (preventive-maintenance interval change, SOP update, incoming-material spec tightening).
Both must be linked to a formal CAPA record with an owner, a due date, and an evidence requirement for closure. Immediate countermeasures without systemic ones guarantee the loss will recur.
Step 7 — Close the loop with measured evidence
A CAPA is not closed when the countermeasure is implemented. It is closed when there is measured evidence that the loss type has not recurred for a defined watch period — typically 30, 60 or 90 days depending on the loss severity. The watch period must be defined up-front when the CAPA is opened, not negotiated at closure.
In the platform, this step is automated: the CAPA stays "open" until the loss investigator confirms that no similar event has triggered within the watch window on the same asset.
The automation story
Every step above is time-expensive if done manually. A full investigation on a mid-size loss typically takes 4-8 hours of a quality engineer's time. The Loss Investigator inside the Synergy Axella platform automates Steps 1, 2, 3 and 4 — pulling the data, reconstructing the timeline, generating candidate hypotheses from asset history + similar past events across the plant network, and scoring evidence against them. The quality engineer then spends their time where it matters: on Steps 5, 6 and 7.
Average time-to-close on investigations in plants using the Loss Investigator: 2.3 days. In matched plants without it: 11 days.
Try the method on your next loss
The seven steps work with or without software. Start with your next downtime event of 30 minutes or more and run the method on paper. You will be surprised how often the "obvious" root cause is wrong.
Then, when you want to compress a 4-hour investigation into 20 minutes, take the Plant Assessment to see how your current loss-investigation maturity scores and what the Loss Investigator feature would unlock on your specific lines.
