ProcessOperations

Process mining for automation: finding the steps worth automating

Process mining rebuilds how a process really runs from timestamps your systems already record. Applied to automation, it answers a narrow question: which steps are frequent, rule-bound and stable enough to automate, which need a person, and which should be redesigned before anyone builds a bot or an agent. Here is the applied method, from event log to ranked backlog.

Reviewed 8 min read

On this page
  1. The designed process versus the process that actually runs
  2. Event-log fields every analysis depends on
  3. Where event data hides in ERP, CRM and ticketing systems
  4. From raw system tables to a ranked automation backlog
  5. Running the analysis: what each pass looks for
  6. Matching each candidate to rules, orchestration, an AI agent or redesign
  7. Mining employee activity data: GDPR and works councils
  8. A hypothetical purchase-to-pay log with hundreds of variants
  9. Questions and answers
  10. Sources

The designed process versus the process that actually runs

Every operation has two versions of each process: the one in the procedure manual and ERP configuration, and the one people run when a supplier quotes the wrong price or an approver is away. Automation built on the first breaks on the second, because exceptions are where the effort goes.

Process mining closes that gap with data instead of workshops. Transactional systems write a timestamped record whenever a document changes state; stitched together per case, those records show every path a case took, how long it waited and who handled it. In ColdAI's operations work this belongs to the diagnostic phase, which maps processes and quantifies performance gaps before any automation is designed1. The wiki entry on process mining explains the discipline; this page applies it to automation decisions.

Event-log fields every analysis depends on

An event log is a table of events. Three fields are mandatory, one is strongly recommended, and the exchange format decides how faithfully the log can describe the process.

Case ID
The identifier that ties events into one instance of the process: a purchase order, a claim, a ticket. Choosing it is a design decision, because one invoice can cover several orders.
Activity
A named step, such as Post goods receipt, usually derived from status changes and renamed in the language the team uses.
Timestamp
When the event happened, ideally with time of day and time zone. Date-only timestamps hide waiting time and scramble the order of same-day steps.
Resource
The user, team or system account behind the step. Needed to analyze handoffs and the share of work already done by batch jobs.
XES
The eXtensible Event Stream standard, IEEE 1849-2023, an XML-based format for exchanging event logs that organizes each log around a single case notion2.
Object-centric event log (OCEL 2.0)
A format in which one event can relate to several objects, such as orders, deliveries and invoices3. It avoids forcing a many-to-many process into one case ID, which distorts frequencies and durations.

Where event data hides in ERP, CRM and ticketing systems

Logs rarely arrive ready-made, and each kind of system has a typical blind spot worth checking before analysis starts.

SourceUsual case notionWhere events come fromCommon gaps
ERP purchase-to-payPurchase order line or invoiceDocument tables, change logs and workflow logsApprovals given by email never appear; batch postings share one timestamp
ERP order-to-cashSales order or deliveryOrder, delivery, billing and payment documentsPartial deliveries create many-to-many links; credit-block releases may sit elsewhere
Ticketing and IT service managementTicket or incidentStatus transitions and assignmentsReassignment is visible but its reason is not; work done on calls leaves no event
Shared mailboxes and spreadsheetsRarely a clean oneMessage metadata or file versions, if anythingOften the largest gap: the steps that most need automation leave no structured trace

Use the gaps column as an extraction checklist: a known gap can be labelled, a missed one becomes a false finding.

From raw system tables to a ranked automation backlog

01Extract and jointables02Build the event log03Discover variants04Check conformance05Measure waiting time06Rank candidates07Choose the tool
  1. Extract and join tables

    Pull change logs and document tables for the scope and window.

  2. Build the event log

    Map status changes to activities and fix the case notion.

  3. Discover variants

    Group cases by path to separate the main flow from the long tail.

  4. Check conformance

    Compare actual paths with the intended process and with policy.

  5. Measure waiting time

    Find where cases sit idle between steps and what they are waiting for.

  6. Rank candidates

    Score steps and sub-flows on volume, rule-bound share, exceptions and delay.

  7. Choose the tool

    Assign rules, orchestration, an AI agent or redesign to each candidate.

Conceptual flow of process mining applied to automation planning. The stages overlap in practice and the diagram implies no durations.

Running the analysis: what each pass looks for

  1. Fix the scope and the question

    Pick one end-to-end process and a window that includes month-end and seasonal peaks. Write down the decision the analysis supports, such as which matching steps to automate this year.

    Output
    Scope note and extraction specification
    Owner
    Process owner with the analyst
  2. Discover the variants

    Group cases by their exact sequence of activities; a few variants usually carry most of the volume. Look for rework loops, ping-pong between teams and steps out of order, such as an invoice posted before its goods receipt.

    Output
    Variant map with volume per path
  3. Check conformance

    Compare actual paths with the intended model and with policy: skipped approvals, orders without a purchase order, payments released before matching. These findings decide which controls the automated flow must enforce.

    Output
    Deviation list by type and frequency
    Owner
    Process and control owners
  4. Measure waiting, not just processing

    Split elapsed time into working and waiting. In back-office processes waiting usually dominates and gathers at handoffs. Automating a quick step changes little if the case then sits in the next queue.

    Output
    Waiting-time map by handoff
  5. Score the candidates

    For each step or sub-flow, record volume, the share of cases that follow explicit rules, the exception rate, the handoffs it removes and the waiting attached to it. Then check that its inputs are structured and reachable through an API or a stable screen.

    Output
    Scored candidate list
    Owner
    Analyst with the automation lead
  6. Validate with the people doing the work

    Walk the top candidates through with the team that handles them. They will explain variants the data cannot, such as a supplier that always invoices in two parts.

    Output
    Validated, ranked backlog

Matching each candidate to rules, orchestration, an AI agent or redesign

The ranking says where to act. The shape of the data says how.

  • If

    Inputs are structured and the decision follows written rules with few exceptions.

    Then

    Implement deterministic rules in the system of record or a rules engine.

    Rules are cheaper to test, audit and change.

  • If

    Each step is simple, but the case crosses several systems and teams.

    Then

    Use workflow orchestration to move data and tasks between systems, for example with the tools covered under workflow automation.

    When delay sits in handoffs, removing the handoff matters more than speeding up the step.

  • If

    Inputs are unstructured documents or messages, or the step needs bounded judgment.

    Then

    Consider an AI agent that reads, classifies and proposes, with people approving the exceptions.

    This is where rules break down and where reading every item by hand costs the most.

  • If

    The step sits in a long tail of variants, or conformance shows the process is not followed.

    Then

    Standardize or redesign the process first, then mine it again.

    Automating a process nobody follows encodes the confusion.

Mining employee activity data: GDPR and works councils

Event logs almost always carry user IDs, which makes them personal data about employees. Settle this before extraction.

GDPR, Regulation (EU) 2016/679

European Union and EEA

Applies whenThe log includes user IDs or other fields that identify employees, as most ERP and ticketing extracts do4.

  • Set a lawful basis and a defined purpose for the analysis, and keep to it.
  • Assess whether a data protection impact assessment under Article 35 is required, especially if results could be used to evaluate individuals4.
  • Check national rules on employment-context processing adopted under Article 884.
  • Pseudonymize resource IDs or aggregate to team level where possible.

Works Constitution Act (Betriebsverfassungsgesetz), Section 87

Germany

Applies whenTechnical devices designed to monitor employee behavior or performance are introduced or used, the co-determination right in Section 87(1) item 65. Process mining tools that process user IDs are generally treated as within scope.

  • Involve the works council before introduction; where co-determination applies, the employer cannot proceed alone.
  • Agree a works agreement on purpose, data fields, access, retention and limits on evaluating individuals.
  • Other countries have their own consultation rights for monitoring systems; check every site in scope.

A hypothetical purchase-to-pay log with hundreds of variants

Questions and answers

How much historical data does process mining need?

Enough to include the cycles that change behavior: month-end, quarter-end, seasonal peaks and several complete runs of the slowest cases, which usually means many months. More is not always better, because system upgrades or new approval rules can make older data describe a process that no longer exists.

What if the steps we most want to automate leave no trace in our systems?

That is common for work done in email, spreadsheets and chat. Add light status capture to the workflow, use task mining on a consenting sample of desktops, or run time-boxed manual sampling, all subject to the same privacy checks. Label these steps so a missing step is not mistaken for a fast one.

How is process mining different from task mining?

Process mining uses system event logs to show how cases move across steps, systems and teams. Task mining records desktop interactions, such as clicks and application switches, to show how one step is performed. The first tells you where to look; the second helps design the automation of a single step, and is more intrusive.

How does process mining connect to straight-through processing?

It finds the steps worth automating, and once automation is live the same log shows where items still fall out and why. That makes it the measurement layer for raising straight-through processing, where the work shifts from choosing what to automate to removing the causes of exceptions.

Sources

  1. Operations capability: diagnostic, design, implement and sustain — ColdAI
  2. IEEE 1849-2023 Standard for eXtensible Event Stream (XES) for Achieving Interoperability in Event Logs and Event Streams — IEEE Standards Association · checked 10 October 2026
  3. OCEL 2.0: Object-Centric Event Log standard — OCEL Standard · checked 10 October 2026
  4. Regulation (EU) 2016/679 (General Data Protection Regulation) — EUR-Lex · checked 10 October 2026
  5. Works Constitution Act (Betriebsverfassungsgesetz), English translation — German Federal Ministry of Justice · checked 10 October 2026

More in Operations

Back to Operations

Next step

Send us one process and the systems it runs through

Name the process, the systems that record it and the automation decision you need to make. We will reply with a view on whether your event data supports process mining and what to extract first.

Discuss a process mining diagnostic