ProcessOil & Gas
Pipeline integrity management with machine learning: from in-line inspection data to a dig list
Machine learning earns a place in pipeline integrity when it sits inside the engineering process rather than beside it. Here is a six-step workflow, from quality-checking in-line inspection (ILI) data to a ranked dig list an integrity engineer signs, with the US rules and standards each step must satisfy and the points where field results flow back into the models.
On this page
- Threats, high consequence areas and what an integrity programme must show
- Rules and standards an ML-assisted integrity process must fit
- The integrity loop from inspection to recalibration
- Six steps from raw ILI data to a ranked dig list
- Deterministic or probabilistic corrosion growth for your line
- What machine learning should not replace in an integrity programme
- Where ML-assisted integrity programmes go wrong
- A hypothetical crude oil line with three inspection runs
- Questions and answers
- Sources
Threats, high consequence areas and what an integrity programme must show
Integrity management asks of every pipeline segment which threats could make it fail, how soon, and with what consequence. ASME B31.8S groups threats into time-dependent ones such as external corrosion, internal corrosion and stress corrosion cracking; stable ones such as manufacturing and construction defects; and time-independent ones such as third-party damage, incorrect operations and outside force. US gas transmission rules require operators to identify threats to each covered segment using those categories and to integrate the relevant data1.
Consequence enters through high consequence areas (HCAs): locations where a release would harm people or, for hazardous liquids, drinking water sources and unusually sensitive environments. Operators must assess pipe that could affect them, remediate conditions within set response times and keep records of how each decision was reached23. Any machine learning output has to fit this frame: a threat-specific estimate with an uncertainty, traceable to inputs an auditor can inspect.
Rules and standards an ML-assisted integrity process must fit
| Instrument | What it governs | What it means for the model |
|---|---|---|
| 49 CFR Part 192, Subpart O | Gas transmission integrity management, including threat identification under § 192.917 and remediation under § 192.93312 | Outputs must map to the rule's threat categories and remediation conditions |
| 49 CFR § 195.452 | Hazardous liquid pipelines that could affect HCAs, with immediate, 60-day and 180-day repair conditions3 | Ranking orders work within and across categories; it never overrides a regulatory category |
| ASME B31.8S | Integrity management of gas pipelines: threat categories, data integration and response | Gives the threat taxonomy that features and labels should follow |
| API RP 1160 | Integrity management of hazardous liquid pipelines, including risk assessment and data integration | Treat the model as part of the documented risk assessment, reviewed when inputs change |
| API Standard 1163 | Qualification of in-line inspection systems and validation of their performance | Dig results that verify the tool are the same data that recalibrate the model |
| CSA Z662 and other national codes | Design, operation and integrity of pipeline systems outside the US, such as in Canada | Confirm local thresholds before reusing a model across jurisdictions |
Simplified summaries; the US rules have been amended several times, so read the current text before relying on any threshold.
The integrity loop from inspection to recalibration
- ILI run
Magnetic flux leakage, ultrasonic or crack-detection tools record features along the line.
- Data QA
Coverage, speed excursions, sensor loss and stated tolerances checked against the specification.
- Run-to-run alignment
Girth welds and anomalies matched across runs so growth can be measured.
- Growth and threat model
Growth estimates combined with protection, pressure cycling and geohazard data.
- Failure-pressure ranking
Projected failure pressure and time to threshold ordered into a dig list.
- Dig and field verify
Excavations measure real dimensions, which are compared with the ILI calls.
Six steps from raw ILI data to a ranked dig list
Quality-check the ILI data against the tool specification
Confirm the run met its specification: coverage, tool speed, sensor dropouts, and the depth and length tolerances the vendor states at a given certainty. Mark sections below specification so later models treat them as higher uncertainty rather than clean pipe.
Align runs girth weld by girth weld
Match girth welds, valves, tees and other fixed features between runs, then match anomalies within each joint. Odometer drift and different tool technologies make distance matching unreliable, so keep unmatched features visible instead of forcing pairs.
Estimate corrosion growth, deterministic or probabilistic
Deterministic approaches assign a growth rate per matched anomaly or segment, often with a conservative floor. Probabilistic approaches treat depth and growth as distributions that combine tool tolerance and matching uncertainty. Machine learning can estimate growth for anomalies with no reliable prior match by learning from similar joints, coatings and soils.
Integrate operating, protection and environmental data
Join cathodic protection surveys, pressure cycling from SCADA history, soil and coating records, crossings and geohazard maps to each joint. Pressure cycling matters for fatigue-sensitive features such as dents; protection gaps often explain why corrosion clusters where it does.
Assess failure pressure and rank
Calculate remaining strength with a recognised method such as original ASME B31G, modified B31G or an effective-area method; for gas transmission, predicted failure pressure is determined under § 192.7122. Project depth forward with the growth distributions, estimate when each anomaly could cross a remediation threshold, then rank by that time and by consequence.
Issue the dig list, verify in the field and recalibrate
An engineer reviews and signs the dig list, including anomalies already in a regulatory repair category. Crews measure each excavated anomaly; unity plots of field depth against ILI depth reveal tool bias, and the results update both the tolerance used in the first step and the growth model.
Deterministic or probabilistic corrosion growth for your line
- If
You hold at least two aligned runs from comparable tools and a modest anomaly population.
ThenStart with matched-pair deterministic growth rates and a documented conservative adjustment.
The method is transparent to auditors and the data supports direct measurement.
- If
Runs used different tool technologies or vendors, or the anomaly population is large.
ThenUse probabilistic growth that carries tool and matching uncertainty into the projection.
Point estimates from mismatched tools produce negative or implausibly high growth rates.
- If
Only one run exists for the line.
ThenBorrow segment-level rates from corrosion surveys, coupons or comparable lines, and bring the next run forward.
A single run gives the model nothing on this line to learn growth from.
- If
Model growth estimates disagree with field measurements from recent digs.
ThenStop using the model for ranking, investigate the cause and revert to the documented deterministic method meanwhile.
Unity plots are the evidence an inspector will ask about.
What machine learning should not replace in an integrity programme
Where ML-assisted integrity programmes go wrong
Forced matches between runs
Early signalThe growth distribution has a long tail of very high or negative rates.
MitigationKeep unmatched anomalies separate and review alignment where odometer offsets are large.
Learning failure from too few failures
Early signalThe training set holds almost no in-service failures or near-failure anomalies.
MitigationTrain on measured depth and growth, which are plentiful, and keep the failure criterion physics-based.
Optimistic validation
Early signalAccuracy drops sharply when the model meets a new inspection year.
MitigationSplit validation by run and by line section, never by randomly sampled anomalies.
Unpiggable segments read as clean pipe
Early signalSections with no ILI coverage show low risk instead of high uncertainty.
MitigationAssess them by direct assessment or pressure testing and mark model outputs as not applicable there.
A hypothetical crude oil line with three inspection runs
Questions and answers
How accurate are in-line inspection tools?
Each tool's performance specification states tolerances for depth, length and width at a given certainty, and real performance varies by technology, feature type and pipe condition, so treat a reported depth as a distribution. Unity plots of dig measurements against ILI calls show how the tool performed on your line; API Standard 1163 describes how operators validate tool performance this way.
Can machine learning replace ASME B31G?
Not as the engineering basis. B31G and its modified and effective-area variants are recognised remaining-strength methods that regulators reference and engineers can check by hand. Machine learning is more useful upstream of that calculation, in classifying features, matching them across runs and estimating growth, which are the inputs carrying most of the uncertainty. A direct failure-pressure model can serve as a screening cross-check, but it should not set repair categories.
What can be done for unpiggable pipelines?
Lines that cannot take conventional tools because of tight bends, diameter changes or low flow can sometimes be inspected with tethered or robotic tools built for them. Otherwise operators rely on direct assessment, such as external corrosion direct assessment, or on pressure testing. Machine learning can help choose excavation sites, but missing ILI data should always appear as higher uncertainty in the ranking.
How do US regulators view machine learning in integrity programmes?
The integrity rules do not name analytical tools such as machine learning. They require operators to identify threats, integrate data, assess pipe and remediate conditions on time, with records to prove it. A model fits that structure when it is documented, validated on your own data, reviewed by a qualified engineer and versioned. Expect an inspector to ask how a model's output changed a decision and what evidence supports it.
Sources
- 49 CFR § 192.917 — How does an operator identify potential threats to pipeline integrity and use the threat identification in its integrity program? — Legal Information Institute, Cornell Law School · checked 10 October 2026
- 49 CFR § 192.933 — What actions must be taken to address integrity issues? — Legal Information Institute, Cornell Law School · checked 10 October 2026
- 49 CFR § 195.452 — Pipeline integrity management in high consequence areas — Legal Information Institute, Cornell Law School · checked 10 October 2026