GuideIndustrials

OEE root cause analysis: from automatic stop capture to fixes that hold

OEE root cause analysis works when three things are true: machine states come straight from the controller rather than from memory, every lost minute lands in one agreed loss category, and the reason behind each stop is confirmed by the person who saw it. Machine learning helps most with the third, by suggesting a reason for each stop that operators accept or correct. The method below runs from signal capture to the maintenance or process change that removes the loss.

Reviewed 8 min read

On this page
  1. Definition errors that make OEE numbers impossible to explain
  2. The six big losses, the OEE factor they reduce and the signal that reveals them
  3. How a stop travels from the controller to a corrective action
  4. Building automated loss attribution on one line
  5. Choosing how much the model decides about stop reasons
  6. A packaging line where micro-stops hid behind the word other
  7. How OEE gets gamed once it becomes a target
  8. Routing findings to maintenance, quality and changeover owners
  9. Questions and answers
  10. Sources

Definition errors that make OEE numbers impossible to explain

Overall equipment effectiveness multiplies availability, performance and quality, and ISO 22400-2 lists it among its standard key performance indicators for manufacturing operations management, built from defined time elements1. Most plants that cannot explain their OEE have a definition problem before they have an analytics problem.

Four errors recur. The ideal cycle time is set to the budget rate instead of the fastest rate the machine has demonstrably sustained, which hides speed loss. Changeovers are booked as planned downtime on some lines and as availability loss on others. Reworked parts are counted as good, so first-pass quality looks better than it is. And plant OEE is calculated as a simple average of lines rather than weighted by planned production time.

Write the definitions down, version them and lock changes behind an approval, because every later step inherits them.

The six big losses, the OEE factor they reduce and the signal that reveals them

LossOEE factorSignal you can capture automaticallyEvidence that points to a cause
Equipment failureAvailabilityFault state with controller fault codeFault history, CMMS work orders, component age
Setup and adjustmentAvailabilityChangeover state between last good part of one SKU and first good part of the nextSKU pair, crew, tooling, first-off inspection results
Idling and minor stopsPerformanceShort stops below an agreed threshold, often without a fault codeJams, sensor trips, material splices, upstream gaps
Reduced speedPerformanceActual rate against the ideal rate while runningSpeed overrides, product format, material lot
Process defectsQualityReject counts from vision systems, check-weighers or test stationsDefect class, setpoints, tool wear, ambient conditions
Reduced yield at startupQualityRejects in the window after a start or changeoverWarm-up procedure, recipe download, operator

The six big losses come from total productive maintenance practice. Whatever taxonomy you use, every lost minute should land in exactly one row.

How a stop travels from the controller to a corrective action

01Controller signals02State engine at the edge03Reason suggestion04Operator confirmation05Loss analysis06Owned action
  1. Controller signals

    Running, faulted, blocked, starved and changeover bits, plus part and reject counters.

  2. State engine at the edge

    Turns raw tags into timestamped machine states and stop events beside the line.

  3. Reason suggestion

    A model proposes the most likely reason for each stop from codes, context and history.

  4. Operator confirmation

    The operator accepts or corrects the reason on the line terminal; corrections become labels.

  5. Loss analysis

    Pareto by category and duration, then stratified comparison against context data.

  6. Owned action

    A CMMS work request, a quality investigation or a changeover standard, with a check date.

Conceptual flow of automated OEE loss attribution on one line. It shows the order of steps, not a measured system or a specific product.

Building automated loss attribution on one line

Start on the line that constrains output, because a minute recovered there is a minute of plant output.

  1. Map controller tags to machine states

    Work with the controls engineer to identify the bits that mean running, faulted, blocked by downstream, starved by upstream and in changeover, and the counters for total and rejected parts. Where a state has no bit, derive it from counters and timing rather than asking operators.

    Output
    Tag-to-state map per machine
    Owner
    Controls engineer
  2. Set the minor-stop threshold

    Agree a duration below which a stop is classed as a minor stop automatically and needs no reason code. Pick it so operators are only asked about stops they can genuinely recall.

    Output
    Threshold in the definitions document
    Owner
    CI lead
  3. Separate blocked and starved time

    A machine waiting on its neighbor is not the cause of the loss. Attribute blocked and starved time to the machine that started the chain, which also shows where the true constraint sits.

    Output
    Constraint view of the line
    Owner
    Production engineer
  4. Rebuild the reason-code list

    Shorten it to reasons that lead to different actions, remove catch-all codes and group the rest by machine area. Long lists produce whichever reason sits first on the screen.

    Output
    Reason tree per machine
    Owner
    Area supervisor
  5. Add suggested reasons with confirmation

    Train a classifier on confirmed stops using fault code, preceding signals, SKU, time since changeover and recent history. Show its top suggestion on the line terminal; the operator accepts or picks another. Track how often suggestions are accepted, by machine and shift.

    Output
    Confirmed stop log with labels
    Owner
    Data engineer
  6. Rank losses, then stratify

    Build a Pareto of lost time by category and reason. For the top few, compare loss rates across material lots, suppliers, crews, SKUs, setpoints and days since maintenance, holding other factors steady where you can.

    Output
    Ranked loss hypotheses
    Owner
    CI lead
  7. Test the fix and record the outcome

    Turn each hypothesis into a change with an owner, then compare the loss before and after over comparable production. Close the action only when the loss category moves.

    Output
    Verified countermeasure log
    Owner
    Plant manager

Choosing how much the model decides about stop reasons

  • If

    The line has clean fault codes that map one-to-one to causes, such as a single servo fault.

    Then

    Assign the reason automatically from the code and let operators override.

    Asking people to confirm what the controller already knows wastes attention.

  • If

    Stops share codes but have several possible causes, as with a generic jam sensor.

    Then

    Show a ranked suggestion and require one-tap confirmation.

    Confirmation turns operator knowledge into training labels instead of discarding it.

  • If

    Suggestion acceptance is low or varies widely between shifts.

    Then

    Review the reason tree and retrain before trusting any Pareto built on it.

    Low acceptance usually means the reasons are ambiguous, not that operators are careless.

  • If

    Reason data is being used to rate individuals or crews.

    Then

    Stop, and report losses by machine and cause only.

    Once codes carry blame, people choose the safest code rather than the true one.

A packaging line where micro-stops hid behind the word other

How OEE gets gamed once it becomes a target

Planned downtime expands

Early signalOEE rises while total output per calendar day does not.

MitigationReport loss hours and output against calendar time alongside OEE.

Ideal rates drift down

Early signalIdeal cycle times change without an engineering reason on file.

MitigationKeep an audit trail of ideal-rate changes with approver and justification.

Rework is counted as good

Early signalQuality stays high while rework stations stay busy.

MitigationCount only first-pass good parts toward the quality factor.

Overproduction to protect the number

Early signalLines keep running products nobody scheduled.

MitigationPair OEE with schedule adherence, and never tie pay to OEE alone.

Routing findings to maintenance, quality and changeover owners

Loss analysis only pays when findings land in the systems where work is planned. Repeating equipment faults become work requests in the CMMS with the stop history attached. Defect clusters open a quality investigation with the lots and setpoints involved. Long or variable changeovers go to a SMED-style review with the recorded sequence of states as evidence.

Review the top losses in the daily tiered meeting, but review the definitions and the reason tree quarterly. Where a repeating fault looks like degradation rather than a one-off, it becomes a candidate for condition monitoring; the predictive maintenance readiness checklist helps decide whether that asset is worth instrumenting. Broader process redesign across sites belongs with operations work.

Questions and answers

Can we automate OEE reason codes without operators confirming them?

Only where a controller fault code maps cleanly to one cause. For stops that share sensors or codes, an unconfirmed model will assign plausible but wrong reasons, and nobody will notice because the Pareto still looks tidy. Operator confirmation of a suggested reason costs a single tap and gives you labeled data to keep improving the model, so it is the better default for ambiguous stops.

Where should machine state logic run: in the PLC, at the edge or in the cloud?

Keep the controller program unchanged where possible and compute states on an edge gateway beside the line, which can buffer events if the network drops and serve the operator terminal with low delay. Aggregated stop logs and the reason model's training can sit centrally. Writing new logic into a validated PLC program usually triggers change control you do not need.

How do we set an honest ideal cycle time for older machines?

Use the fastest rate the machine has sustained for a meaningful run on that product, from its own data, and record it with the date and evidence. Nameplate speeds on older equipment may be unreachable, while budget rates hide speed loss. Review the figure when the machine, product or tooling changes, and require an engineering reason for any reduction.

Does OEE root cause analysis work on batch and process lines?

Yes, with adapted states. Batch plants usually track phases, holds and waits rather than discrete parts, and quality losses show up as off-spec batches or yield variance. The method is the same: agreed definitions, automatic state capture from the control system, confirmed reasons for holds and stratified analysis against materials and recipes.

What should we do when the top loss is starved or blocked time?

Follow the chain to the machine that started it, because the waiting machine is not the cause. Persistent starvation often points to an upstream constraint, buffer sizing or material supply, while persistent blocking points downstream. Fix the originating machine's losses first and check whether the line constraint has moved before investing elsewhere.

Sources

  1. ISO 22400-2:2014 Automation systems and integration: Key performance indicators (KPIs) for manufacturing operations management, Part 2: Definitions and descriptions — International Organization for Standardization · checked 10 October 2026

More in Industrials

Back to Industrials

Next step

Get a second opinion on your OEE definitions and reason tree

Send your current OEE definitions, a sample week of stop data and your reason-code list. We will point out where the categories conflict and suggest which line to automate first.

Share your OEE data