Operations

Process mining

Process mining rebuilds how a business process really runs from the event logs of its software systems, then checks it against the intended process and shows where time and rework pile up.

In short

Process mining is a technique that rebuilds how a business process really runs from the event logs of its software systems. Each event records a case ID, an activity and a timestamp. Three tasks follow: discovery builds a process model from the log, conformance checking compares log and model, and enhancement adds timing and bottlenecks to the model.

Origin
Wil van der Aalst and the IEEE Task Force on Process Mining (earlier roots in Cook and Wolf, and Agrawal et al.), 1998 (first log-based discovery papers); 2011 (Process Mining Manifesto)
Level
401 · Expert
Fits
Enterprise
Time to apply
2 to 4 weeks for a first project on one process, most of it spent extracting and cleaning the log
What you need
one process with a clear start and end, such as order to cash or patient referral to first visit · read access to the system that records it, or an export of its history table · someone who knows the process and can say what a normal path looks like

Process mining is a way to see how a business process actually runs by reading the trail it leaves in software. The IEEE Task Force on Process Mining defines the idea as discovering, monitoring and improving real processes, as opposed to assumed ones, by extracting knowledge from event logs that information systems already keep. A flowchart drawn in a workshop shows how people think the work goes. An event log shows what happened, case by case.

Where process mining comes from

Process mining grew out of several lines of work in the late 1990s. Cook and Wolf published a method for discovering models of software processes from event data in 1998. In the same year Agrawal, Gunopulos and Leymann presented a way to mine process models from workflow logs. Wil van der Aalst’s group, first in Eindhoven and later at RWTH Aachen, developed the field into a toolbox, including the alpha algorithm, whose best-known form is a 2004 paper with Weijters and Maruster.

In 2011 the task force published a Manifesto, written by more than 75 people from more than 50 organizations according to the task force’s own page. It sets out six guiding principles and eleven open challenges, and it is still the shortest authoritative description of the field. Van der Aalst’s book, now in its second edition as Process Mining: Data Science in Action, is the standard textbook.

What an event log is

An event log is a table in which every row is one event, and every event names a case and an activity. Fluxicon, a process mining vendor, states the minimum as three elements: a case ID, an activity and a timestamp. The Manifesto adds that techniques also use the resource, meaning the person or device, and data recorded with the event, such as the size of an order.

The case is the unit you follow. In a purchasing process, Fluxicon’s example, one purchase order is one case, and its events (created, approved, paid) share the order’s ID. Microsoft’s documentation gives the same rule for its tool: the case ID must be present for all activities, and may be a patient ID, an order ID or a request ID.

The data must be an event history. Fluxicon’s blog post on data requirements says it should not be aggregated to the case level, since one row per case erases the sequence. Many system tables hold only the current state of a record. Microsoft’s guide notes that not all events you care about are logged, and that you may need to join the history table with other tables to get the names and IDs. The IEEE 1849 XES standard exists so that tools can exchange event logs in one format.

A table of seven events for cases 101 and 102 with columns Case, Activity and Time, and an arrow to a process map of four boxes: Order received, Check, Approve and Pay. A blue arrow goes straight from Order received to Approve, labelled 1, because case 102 skipped the check.
Two cases in a log become one map; the blue arrow is the skipped check that no workshop flowchart would show.

The three types of process mining

The Manifesto names three kinds of process mining, defined by what goes in and what comes out.

Three rows, each with an input box, a blue technique box and an output box: event log to Discovery to Model, event log plus model to Conformance checking to Diagnostics, event log plus model to Enhancement to New model.
Discovery needs only a log; the other two techniques compare or extend a model that already exists.

Discovery takes an event log and produces a model without any prior information. The alpha algorithm builds a Petri net from the order in which activities follow each other. Leemans, Fahland and van der Aalst’s inductive miner, published in 2013, produces block-structured models. PM4Py, an open-source Python library, includes it.

Conformance checking compares an existing model, or a rule, with the log. Rozinat and van der Aalst measured two things in 2008: fitness, whether the recorded behaviour follows the model, and appropriateness, whether the model describes what was observed. The question it answers is whether people follow the approval policy.

Enhancement uses the log to extend or repair the model. Van der Aalst, Adriansyah and van Dongen showed in 2012 that replaying history on a model and combining it with timestamps reveals bottlenecks, which is how most business teams find the slow step.

How process mining differs from nearby methods

Method Starting point What it shows Weak spot
Process mining Event logs from systems The real paths, with counts and timings Blind to work that leaves no log
BPMN mapping Interviews and workshops The designed process, in a standard notation Shows intent, not what people do
Value stream mapping Walking the flow Waiting and handoffs across a physical flow Manual, a snapshot of a few cases
SIPOC A short workshop Scope: suppliers, inputs, process, outputs, customers Too coarse for analysis
Task mining Desktop and web interaction logs How one person does a task Does not follow a case across systems
Data mining A table of records Patterns in the data Not built around process models

On the last row, van der Aalst’s 2012 overview says classical data mining techniques are often used to analyze a specific step, while process mining focuses on end-to-end processes. Leno and colleagues call the research on finding routines in user interaction logs robotic process mining. It feeds business process automation by showing which routines are worth scripting. Process mining and the event tracking plan are close cousins: a good tracking plan is what makes a product’s events usable as a log.

Where it goes wrong

Data is the main risk. Suriadi and colleagues catalogued recurring imperfections in real logs, such as missing, duplicated or wrongly ordered events, and proposed patterns for cleaning them. Process change is another: Bose and colleagues showed that a process can drift while you analyze it, so one map may mix two versions of the process.

Scale and structure also matter. The Manifesto reports that ASML monitors its wafer scanners continuously and that existing tools struggled with the petabytes of data. Classical logs assume one case notion, while real systems often link one event to several objects, such as an order and its items. The OCEL 2.0 standard from RWTH Aachen records events with the objects involved for exactly that reason.

A caution on evidence. A 2007 industrial application by van der Aalst and colleagues noted that few of the advanced techniques had been tested on real-life processes. Vendor pages, such as Celonis’s, claim faster and cheaper results than workshops and cite customers, but offer no independent evidence on the page. Read these as claims to test on your own log.

Reading a process map in a team is a friction audit done with data. If the processes behind your funnel and back office are poorly documented, a marketing operational system review can use an event log as the first input.

How to apply Process mining, step by step

  1. Start from a question. Write the question the analysis must answer: why do payouts take 4 days, where does rework start. The Process Mining Manifesto says log extraction should be driven by questions, because an ERP database holds thousands of tables and nobody can pick the right ones without one. Result: a one-sentence question and a named process.
  2. Choose the case. Decide what one run of the process is: one purchase order, one patient visit, one merchant application. Every event must carry that ID. Result: a case definition that all data sources can be joined on.
  3. Extract one row per event. Pull case ID, activity name and timestamp for every event, plus the resource and any attribute you will filter on, such as amount or channel. Keep the history, do not roll it up to one row per case. Result: a raw event log.
  4. Clean the log. Check for events without timestamps, duplicated rows, activity names that differ only in spelling, and events recorded long after they happened. Fix them or drop them and write down what you dropped. Result: a log you trust enough to show to the process owner.
  5. Discover the process and read the variants. Load the log into a tool, draw the process map, and list the variants, which are the distinct paths cases took. Start with the few variants that cover most cases, then look at the rare ones. Result: a map of what happens, with counts and durations on each step.
  6. Check the map against the rules. Compare the map with the intended process or a policy, such as a four-eyes approval before payment. Mark each deviation as a justified exception or a defect. Result: a short list of deviations with the share of cases affected.
  7. Fix, then re-run on fresh data. Change the process or the system for the biggest defect, then repeat the extraction after a few weeks. The Manifesto lists continuous process mining as a guiding principle. Result: a before and after comparison on the same measures.

Examples

A clinic from referral to first visit

Illustrative. A clinic exports its booking system's history table: one row per status change with the booking ID, the status and the time. Discovery shows most bookings go Referral received, Call made, Visit booked, Visit done. About one in six has a second Call made step before a booking. Conformance checking against the rule 'every referral gets a call within one working day' flags the cases that waited over a weekend. The fix is a Friday afternoon call slot, and the next export shows whether the weekend waits disappeared.

A payments company onboarding merchants

Illustrative. The case is one merchant application, and the events come from the CRM, the compliance tool and the payments back office, joined on the application ID. The map shows applications bouncing between Documents requested and Documents received three or four times. Filtering by country shows the loops are concentrated in two markets whose document templates differ from the rest. The team rewrites those two requests and compares loop counts a month later.

When to use it

Use it when a process runs through software that logs its steps, has enough volume that a few cases are not a sample (hundreds or more), and the real path is disputed, hidden or too slow. It suits order to cash, purchase to pay, claims, onboarding and support. It also helps before automation, to find which steps are worth automating and which variants will break a script.

When not to use it

Skip it when the work happens outside systems (phone calls, paper, chat) and leaves no timestamps, or when only a handful of cases pass through each month. For a process that does not exist yet, draw it with BPMN first. If the main question is where value is lost in a physical flow, a value stream map on the shop floor will teach more.

Common mistakes

  • Starting from the tool, not a question. A map of every table in the ERP is unreadable. Pick one process and one question first.
  • Using a log that is not an event history. Tables that store only the current state of a record cannot show how it got there. Check that every status change is stored with a time.
  • Trusting timestamps blindly. Events that were entered later than they happened, or in batches, distort durations and orders of steps. Ask how each timestamp is created.
  • Treating every deviation as a defect. Some deviations are the exception handling that keeps customers happy. Review them with the process owner before changing anything.
  • Running it once. Processes drift, and a map from last spring may describe a process that no longer exists.

FAQ

What data do you need for process mining?

An event log with at least a case ID, an activity name and a timestamp for every event, one row per event. Resource and attributes such as amount or channel are optional but make the analysis richer. Aggregated data with one row per case is not enough, because it loses the sequence of steps.

What is the difference between process mining and data mining?

Data mining looks for patterns in data and is often applied to one step of a process. Van der Aalst's 2012 overview in ACM Transactions on Management Information Systems says classical data mining techniques do not focus on process models, while process mining analyses end-to-end processes from event data.

What is the difference between process mining and task mining?

Process mining uses event logs from business systems. Task mining, studied in research as robotic process mining, uses logs of user interactions with desktop and web applications to find repetitive routines that could be automated. The first shows how a case moves across systems, the second how one person does a task.

Who invented process mining?

No single inventor. Cook and Wolf published process discovery from software process data in 1998, and Agrawal, Gunopulos and Leymann published mining of workflow logs the same year. Wil van der Aalst's group in Eindhoven, later at RWTH Aachen, developed much of the method and helped write the 2011 Manifesto.

Sources

  1. Wil van der Aalst, Process Mining: Data Science in Action, 2nd edition, Springer, 2016
  2. IEEE Task Force on Process Mining, Process Mining Manifesto, BPM 2011 Workshops, LNBIP 99, Springer, 2012
  3. IEEE Task Force on Process Mining, Process Mining Manifesto (task force page, with the full text)
  4. Wil van der Aalst, Ton Weijters, Laura Maruster, Workflow mining: discovering process models from event logs, IEEE Transactions on Knowledge and Data Engineering 16(9), 2004
  5. Rakesh Agrawal, Dimitrios Gunopulos, Frank Leymann, Mining process models from workflow logs, EDBT 1998, Springer LNCS 1377
  6. Jonathan Cook, Alexander Wolf, Discovering models of software processes from event-based data, ACM Transactions on Software Engineering and Methodology 7(3), 1998
  7. Anne Rozinat, Wil van der Aalst, Conformance checking of processes based on monitoring real behavior, Information Systems 33(1), 2008
  8. Wil van der Aalst, Arya Adriansyah, Boudewijn van Dongen, Replaying history on process models for conformance checking and performance analysis, WIREs Data Mining and Knowledge Discovery 2(2), 2012
  9. Sander Leemans, Dirk Fahland, Wil van der Aalst, Discovering block-structured process models from event logs: a constructive approach, Petri Nets 2013, Springer LNCS 7927
  10. R. P. Jagadeesh Chandra Bose, Wil van der Aalst, Indre Zliobaite, Mykola Pechenizkiy, Dealing with concept drifts in process mining, IEEE Transactions on Neural Networks and Learning Systems 25(1), 2014
  11. Wil van der Aalst, Process mining, ACM Transactions on Management Information Systems 3(2), 2012
  12. Wil van der Aalst, Hajo Reijers, Ton Weijters, Boudewijn van Dongen and others, Business process mining: an industrial application, Information Systems 32(5), 2007
  13. Marlon Dumas, Marcello La Rosa, Jan Mendling, Hajo Reijers, Fundamentals of Business Process Management, 2nd edition, Springer, 2018
  14. Wil van der Aalst, Josep Carmona (eds.), Process Mining Handbook, Springer LNBIP 448, 2022 (open access)
  15. Jochen De Weerdt, Moe Wynn, Foundations of process event data, in Process Mining Handbook, Springer, 2022
  16. Suriadi, Andrews, ter Hofstede, Wynn, Event log imperfection patterns for process mining: towards a systematic approach to cleaning event logs, Information Systems 64, 2017
  17. XES, the IEEE 1849 standard for exchanging event logs
  18. RWTH Aachen, OCEL 2.0: the object-centric event log standard
  19. Fluxicon (vendor), Process mining book: data requirements and the minimum requirements for an event log
  20. Fluxicon (vendor), Data requirements for process mining, 2012
  21. Microsoft Learn (vendor docs), Prepare processes and data for process mining in Power Automate
  22. Celonis (vendor), What is process mining (marketing claims, not independent evidence)
  23. Process Intelligence Solutions (Fraunhofer FIT spin-off), PM4Py open-source library
  24. Volodymyr Leno, Artem Polyvyanyy, Marlon Dumas, Marcello La Rosa, Fabrizio Maggi, Robotic process mining: vision and challenges, Business & Information Systems Engineering 63(3), 2021
  25. Eduardo Rojas, Jorge Munoz-Gama, Marcos Sepulveda, Daniel Capurro, Process mining in healthcare: a literature review, Journal of Biomedical Informatics 61, 2016
  26. Thomas Reinkemeyer (ed.), Process Mining in Action: Principles, Use Cases and Outlook, Springer, 2020
  27. RWTH Aachen Process and Data Science group, processmining.org

Last updated Oct 9, 2026

Ilia PushinFounder, PUSHERS & COO Fintech ServiceIlia builds operating systems for growing companies in fintech and healthcare. Since 2021 he has run cross-border payments at ARBI Exchange, a licensed currency exchange in Thailand, including KYC and AML and the move into new jurisdictions.About the authorLinkedIn
Related frameworks
More frameworks
Want Process mining running inside your company?Request an operations audit