GuideIndustrials
OEE root cause analysis: from automatic stop capture to fixes that hold
OEE root cause analysis works when three things are true: machine states come straight from the controller rather than from memory, every lost minute lands in one agreed loss category, and the reason behind each stop is confirmed by the person who saw it. Machine learning helps most with the third, by suggesting a reason for each stop that operators accept or correct. The method below runs from signal capture to the maintenance or process change that removes the loss.
On this page
- Definition errors that make OEE numbers impossible to explain
- The six big losses, the OEE factor they reduce and the signal that reveals them
- How a stop travels from the controller to a corrective action
- Building automated loss attribution on one line
- Choosing how much the model decides about stop reasons
- A packaging line where micro-stops hid behind the word other
- How OEE gets gamed once it becomes a target
- Routing findings to maintenance, quality and changeover owners
- Questions and answers
- Sources
Definition errors that make OEE numbers impossible to explain
Overall equipment effectiveness multiplies availability, performance and quality, and ISO 22400-2 lists it among its standard key performance indicators for manufacturing operations management, built from defined time elements1. Most plants that cannot explain their OEE have a definition problem before they have an analytics problem.
Four errors recur. The ideal cycle time is set to the budget rate instead of the fastest rate the machine has demonstrably sustained, which hides speed loss. Changeovers are booked as planned downtime on some lines and as availability loss on others. Reworked parts are counted as good, so first-pass quality looks better than it is. And plant OEE is calculated as a simple average of lines rather than weighted by planned production time.
Write the definitions down, version them and lock changes behind an approval, because every later step inherits them.
The six big losses, the OEE factor they reduce and the signal that reveals them
| Loss | OEE factor | Signal you can capture automatically | Evidence that points to a cause |
|---|---|---|---|
| Equipment failure | Availability | Fault state with controller fault code | Fault history, CMMS work orders, component age |
| Setup and adjustment | Availability | Changeover state between last good part of one SKU and first good part of the next | SKU pair, crew, tooling, first-off inspection results |
| Idling and minor stops | Performance | Short stops below an agreed threshold, often without a fault code | Jams, sensor trips, material splices, upstream gaps |
| Reduced speed | Performance | Actual rate against the ideal rate while running | Speed overrides, product format, material lot |
| Process defects | Quality | Reject counts from vision systems, check-weighers or test stations | Defect class, setpoints, tool wear, ambient conditions |
| Reduced yield at startup | Quality | Rejects in the window after a start or changeover | Warm-up procedure, recipe download, operator |
The six big losses come from total productive maintenance practice. Whatever taxonomy you use, every lost minute should land in exactly one row.
How a stop travels from the controller to a corrective action
- Controller signals
Running, faulted, blocked, starved and changeover bits, plus part and reject counters.
- State engine at the edge
Turns raw tags into timestamped machine states and stop events beside the line.
- Reason suggestion
A model proposes the most likely reason for each stop from codes, context and history.
- Operator confirmation
The operator accepts or corrects the reason on the line terminal; corrections become labels.
- Loss analysis
Pareto by category and duration, then stratified comparison against context data.
- Owned action
A CMMS work request, a quality investigation or a changeover standard, with a check date.
Building automated loss attribution on one line
Start on the line that constrains output, because a minute recovered there is a minute of plant output.
Map controller tags to machine states
Work with the controls engineer to identify the bits that mean running, faulted, blocked by downstream, starved by upstream and in changeover, and the counters for total and rejected parts. Where a state has no bit, derive it from counters and timing rather than asking operators.
Set the minor-stop threshold
Agree a duration below which a stop is classed as a minor stop automatically and needs no reason code. Pick it so operators are only asked about stops they can genuinely recall.
Separate blocked and starved time
A machine waiting on its neighbor is not the cause of the loss. Attribute blocked and starved time to the machine that started the chain, which also shows where the true constraint sits.
Rebuild the reason-code list
Shorten it to reasons that lead to different actions, remove catch-all codes and group the rest by machine area. Long lists produce whichever reason sits first on the screen.
Add suggested reasons with confirmation
Train a classifier on confirmed stops using fault code, preceding signals, SKU, time since changeover and recent history. Show its top suggestion on the line terminal; the operator accepts or picks another. Track how often suggestions are accepted, by machine and shift.
Rank losses, then stratify
Build a Pareto of lost time by category and reason. For the top few, compare loss rates across material lots, suppliers, crews, SKUs, setpoints and days since maintenance, holding other factors steady where you can.
Test the fix and record the outcome
Turn each hypothesis into a change with an owner, then compare the loss before and after over comparable production. Close the action only when the loss category moves.
Choosing how much the model decides about stop reasons
- If
The line has clean fault codes that map one-to-one to causes, such as a single servo fault.
ThenAssign the reason automatically from the code and let operators override.
Asking people to confirm what the controller already knows wastes attention.
- If
Stops share codes but have several possible causes, as with a generic jam sensor.
ThenShow a ranked suggestion and require one-tap confirmation.
Confirmation turns operator knowledge into training labels instead of discarding it.
- If
Suggestion acceptance is low or varies widely between shifts.
ThenReview the reason tree and retrain before trusting any Pareto built on it.
Low acceptance usually means the reasons are ambiguous, not that operators are careless.
- If
Reason data is being used to rate individuals or crews.
ThenStop, and report losses by machine and cause only.
Once codes carry blame, people choose the safest code rather than the true one.
A packaging line where micro-stops hid behind the word other
How OEE gets gamed once it becomes a target
Planned downtime expands
Early signalOEE rises while total output per calendar day does not.
MitigationReport loss hours and output against calendar time alongside OEE.
Ideal rates drift down
Early signalIdeal cycle times change without an engineering reason on file.
MitigationKeep an audit trail of ideal-rate changes with approver and justification.
Rework is counted as good
Early signalQuality stays high while rework stations stay busy.
MitigationCount only first-pass good parts toward the quality factor.
Overproduction to protect the number
Early signalLines keep running products nobody scheduled.
MitigationPair OEE with schedule adherence, and never tie pay to OEE alone.
Routing findings to maintenance, quality and changeover owners
Loss analysis only pays when findings land in the systems where work is planned. Repeating equipment faults become work requests in the CMMS with the stop history attached. Defect clusters open a quality investigation with the lots and setpoints involved. Long or variable changeovers go to a SMED-style review with the recorded sequence of states as evidence.
Review the top losses in the daily tiered meeting, but review the definitions and the reason tree quarterly. Where a repeating fault looks like degradation rather than a one-off, it becomes a candidate for condition monitoring; the predictive maintenance readiness checklist helps decide whether that asset is worth instrumenting. Broader process redesign across sites belongs with operations work.
Questions and answers
Can we automate OEE reason codes without operators confirming them?
Only where a controller fault code maps cleanly to one cause. For stops that share sensors or codes, an unconfirmed model will assign plausible but wrong reasons, and nobody will notice because the Pareto still looks tidy. Operator confirmation of a suggested reason costs a single tap and gives you labeled data to keep improving the model, so it is the better default for ambiguous stops.
Where should machine state logic run: in the PLC, at the edge or in the cloud?
Keep the controller program unchanged where possible and compute states on an edge gateway beside the line, which can buffer events if the network drops and serve the operator terminal with low delay. Aggregated stop logs and the reason model's training can sit centrally. Writing new logic into a validated PLC program usually triggers change control you do not need.
How do we set an honest ideal cycle time for older machines?
Use the fastest rate the machine has sustained for a meaningful run on that product, from its own data, and record it with the date and evidence. Nameplate speeds on older equipment may be unreachable, while budget rates hide speed loss. Review the figure when the machine, product or tooling changes, and require an engineering reason for any reduction.
Does OEE root cause analysis work on batch and process lines?
Yes, with adapted states. Batch plants usually track phases, holds and waits rather than discrete parts, and quality losses show up as off-spec batches or yield variance. The method is the same: agreed definitions, automatic state capture from the control system, confirmed reasons for holds and stratified analysis against materials and recipes.
What should we do when the top loss is starved or blocked time?
Follow the chain to the machine that started it, because the waiting machine is not the cause. Persistent starvation often points to an upstream constraint, buffer sizing or material supply, while persistent blocking points downstream. Fix the originating machine's losses first and check whether the line constraint has moved before investing elsewhere.
Sources
- ISO 22400-2:2014 Automation systems and integration: Key performance indicators (KPIs) for manufacturing operations management, Part 2: Definitions and descriptions — International Organization for Standardization · checked 10 October 2026