Deep diveElectric Power & Natural Gas
Transformer condition monitoring: combining DGA, thermal data and machine learning
Machine learning adds most to transformer condition monitoring when it builds on the diagnostics engineers already trust. This deep dive covers dissolved gas analysis and its interpretation standards, thermal aging models, the other signals worth adding, how rule-based health indices compare with ML scores, and how a health score becomes a risk ranking that drives inspection, repair and replacement.
On this page
- Why transformer decisions carry so much weight
- Diagnostic terms used in transformer condition assessment
- How the established DGA interpretation methods differ
- Thermal loading, moisture and insulation aging
- Signals to inventory before building a health model
- From raw transformer signals to an intervention decision
- Rule-based health indices, ML scores and explanations engineers accept
- Turning a health score into a fleet risk ranking
- Data pitfalls in transformer analytics
- A hypothetical distribution utility ranks substation transformers
- Questions and answers
- Sources
Why transformer decisions carry so much weight
Power transformers are long-lived, expensive and hard to replace quickly. Large units are often custom-built, spares are few, and a failure can mean fire, an oil spill and an extended outage for everyone the substation serves. Choosing which unit to work on next is among the most consequential calls an asset manager makes.
Most utilities already hold the evidence: oil samples, loading history, inspections, test results and work orders. The difficulty is turning scattered records into priorities that hold up across a large fleet. Predictive maintenance for transformers is one of the use cases ColdAI lists for power and gas operators1; the approach below keeps the engineering diagnostics and uses machine learning where it improves consistency.
Diagnostic terms used in transformer condition assessment
- Dissolved gas analysis (DGA)
- Laboratory or online measurement of gases dissolved in insulating oil, produced when oil and paper break down under electrical or thermal stress.
- Key gases
- Hydrogen, methane, ethane, ethylene and acetylene from oil decomposition; carbon monoxide and dioxide from cellulose. Oxygen and nitrogen reflect how the oil is sealed.
- Duval triangle and pentagons
- Graphical methods that place relative gas proportions into zones for partial discharge, low- and high-energy discharges and thermal faults of rising temperature.
- Hot-spot temperature
- The hottest point in the winding insulation, estimated from top-oil temperature, load and design data; it sets how fast paper ages.
- Aging acceleration factor
- In IEEE C57.91, the ratio of insulation aging at a given hot-spot temperature to aging at the reference temperature; integrating it over time gives equivalent aging4.
- Furans and degree of polymerization
- Degree of polymerization measures paper strength directly but needs a paper sample; furanic compounds in oil indicate paper degradation indirectly.
How the established DGA interpretation methods differ
Each method answers a slightly different question, so practitioners combine them, and an ML model should treat them as features rather than competitors.
| Method | What it uses | Strength | Limitation |
|---|---|---|---|
| Key gas method | The dominant gas in the mix | Simple to explain to field staff | Coarse; mixed faults confuse it |
| Gas ratio methods | Ratios of gas pairs read against code tables | Long history and familiar to engineers | Some combinations match no code at all |
| Duval triangle | Relative shares of methane, ethylene and acetylene | Always returns a zone; separates thermal from electrical faults well | Ignores hydrogen and ethane; unreliable near detection limits |
| Duval pentagons | Hydrogen plus four hydrocarbon gases | Adds detail on partial discharge and stray gassing | Diagnoses fault type, not how fast it is developing |
| IEEE C57.104-2019 status approach | Gas levels and rates of increase against statistical norms | Separates whether gassing is unusual from what fault it suggests | Population norms may not match your own fleet |
| Supervised ML classifier | All gases, rates, loading and history together | Can learn fleet-specific patterns and interactions | Needs labeled faults that most fleets lack |
A common pattern is to use the status approach to decide whether gassing is abnormal and a graphical method to suggest the fault type2; IEC 60599 sets out the equivalent interpretation guidance in IEC markets3.
Thermal loading, moisture and insulation aging
Paper insulation ages with temperature, moisture and oxygen, and its loss of strength usually ends a transformer's life. The IEEE loading guide estimates hot-spot temperature from top-oil temperature and load, then converts it into an aging acceleration factor4; IEC 60076-7 plays the same role in IEC markets. Applied to real load history, it gives an equivalent age that can differ sharply from calendar age.
Inputs are modest: loading from SCADA or aggregated meter data, ambient temperature, cooling status and nameplate thermal data. Fiber-optic hot-spot probes on newer units replace estimates with measurements and help calibrate sister units without probes.
Moisture speeds aging and lowers dielectric strength. Readings swing with temperature as water moves between oil and paper, so use relative saturation or temperature-normalized values, not raw parts per million.
Signals to inventory before building a health model
From raw transformer signals to an intervention decision
- Collect signals
Lab and online DGA, loading, oil tests, bushing and tap changer data.
- Validate and align
Flag lab changes, drift and gaps; align each unit to one timeline.
- Apply standard methods
Compute status and fault-type results engineers recognize.
- Score health
Combine interpretations, trends and equivalent age, with stated drivers.
- Add consequence
Weigh safety, environmental, network and financial impact at that site.
- Rank and choose action
Choose resampling, monitoring, repair or replacement, and record why.
Rule-based health indices, ML scores and explanations engineers accept
A rule-based health index scores each diagnostic, weights the scores and often lets the worst subsystem set a floor, so a failing bushing cannot be averaged away by good oil. It is transparent and auditable, but weights rest on judgment and fixed thresholds treat every unit alike.
Supervised failure prediction meets a hard limit: failures are rare and causes are recorded inconsistently. Three alternatives fit the data most fleets have. Peer comparison flags units that diverge from similar ones of the same design, age and duty. Learning from expert assessments makes the rankings engineers already produce consistent. Survival analysis estimates failure probability over time from age and condition, using retired units as well as failed ones.
The practical answer is usually a hybrid. Standard interpretations become inputs and guardrails, the model adjusts within them, and every score shows its drivers in engineering terms, such as acetylene rising across consecutive samples while load is flat. A score that cannot be explained that way will not change a maintenance plan.
Turning a health score into a fleet risk ranking
Health describes likelihood of failure; risk adds consequence. Consequence depends on the site: proximity to the public and fire exposure, oil near watercourses, the customers served and whether another transformer can carry the load, and the cost and lead time of replacement. Great Britain's distribution network operators work with a regulator-approved methodology that combines health-based probability of failure with financial, safety, environmental and network consequences to give monetized risk5, a useful reference even outside that regime.
The ranking should end in actions: resampling sooner, fitting an online monitor, oil processing, repair, relocating a spare, derating or replacement. Recording each unit's chosen action and reason turns the ranking into a plan auditors and regulators can follow.
Data pitfalls in transformer analytics
Laboratory variability
Early signalFleet-wide step changes when the lab or extraction method changes.
MitigationRecord laboratory and method with every result, flag changes and compare labs on shared samples.
Scarce, inconsistent failure labels
Early signalFailures logged as unknown, or by outage cause rather than mechanism.
MitigationHarmonize cause codes, capture teardown findings and use expert assessments as labels.
Online sensor drift
Early signalSlow trends across many monitors of one model, or disagreement with lab samples.
MitigationCross-check monitors against periodic lab samples and recalibrate or correct drift.
Sampling and survivorship bias
Early signalSuspect units are sampled more, and early replacements vanish from the data.
MitigationModel sampling frequency explicitly and keep retired units in the training set.
A hypothetical distribution utility ranks substation transformers
Questions and answers
Can machine learning replace the Duval triangle or IEEE C57.104 interpretation?
It should not try to. The standard methods encode long field experience and are what engineers, insurers and regulators recognize. Machine learning is most useful combining their outputs with trends, loading and other diagnostics to rank a large fleet consistently. When a model disagrees with the standard interpretation, treat that as a prompt for engineering review, not as the answer.
Are online DGA monitors worth it compared with laboratory samples?
For critical or suspect units, often yes: online monitors catch fast-developing faults between lab samples and show rates of change in near real time. For the bulk of a fleet, periodic lab analysis usually remains the economic choice. Many utilities use risk rankings to decide where monitors go and keep lab samples as the calibration reference for them.
What data does a transformer health model need to start?
A usable start is several years of lab DGA per unit, nameplate and installation data, and loading history. Oil quality, bushing and tap changer results and failure records improve it. Gaps are normal; first inventory which units have which signals, so the model can state lower confidence where data is thin.
How do you validate a transformer risk model without many failures?
Combine several checks. Back-test against the failures and major repairs you do have, compare rankings with blinded engineering assessments, and inspect a sample of high-ranked and low-ranked units to see whether findings match. Track how rankings change as new samples arrive. Validation is continuous, because each intervention and teardown adds evidence.
Sources
- Electric Power & Natural Gas: predictive asset maintenance use case — ColdAI
- IEEE C57.104-2019: IEEE Guide for the Interpretation of Gases Generated in Mineral Oil-Immersed Transformers — IEEE · checked 10 October 2026
- IEC 60599:2022 Mineral oil-filled electrical equipment in service: guidance on the interpretation of dissolved and free gases analysis — International Electrotechnical Commission · checked 10 October 2026
- IEEE C57.91-2011: IEEE Guide for Loading Mineral-Oil-Immersed Transformers and Step-Voltage Regulators — IEEE · checked 10 October 2026
- DNO Common Network Asset Indices Methodology (CNAIM), version 2.1 — Ofgem · checked 10 October 2026