Deep diveElectric Power & Natural Gas

Transformer condition monitoring: combining DGA, thermal data and machine learning

Machine learning adds most to transformer condition monitoring when it builds on the diagnostics engineers already trust. This deep dive covers dissolved gas analysis and its interpretation standards, thermal aging models, the other signals worth adding, how rule-based health indices compare with ML scores, and how a health score becomes a risk ranking that drives inspection, repair and replacement.

Reviewed 8 min read

On this page
  1. Why transformer decisions carry so much weight
  2. Diagnostic terms used in transformer condition assessment
  3. How the established DGA interpretation methods differ
  4. Thermal loading, moisture and insulation aging
  5. Signals to inventory before building a health model
  6. From raw transformer signals to an intervention decision
  7. Rule-based health indices, ML scores and explanations engineers accept
  8. Turning a health score into a fleet risk ranking
  9. Data pitfalls in transformer analytics
  10. A hypothetical distribution utility ranks substation transformers
  11. Questions and answers
  12. Sources

Why transformer decisions carry so much weight

Power transformers are long-lived, expensive and hard to replace quickly. Large units are often custom-built, spares are few, and a failure can mean fire, an oil spill and an extended outage for everyone the substation serves. Choosing which unit to work on next is among the most consequential calls an asset manager makes.

Most utilities already hold the evidence: oil samples, loading history, inspections, test results and work orders. The difficulty is turning scattered records into priorities that hold up across a large fleet. Predictive maintenance for transformers is one of the use cases ColdAI lists for power and gas operators1; the approach below keeps the engineering diagnostics and uses machine learning where it improves consistency.

Diagnostic terms used in transformer condition assessment

Dissolved gas analysis (DGA)
Laboratory or online measurement of gases dissolved in insulating oil, produced when oil and paper break down under electrical or thermal stress.
Key gases
Hydrogen, methane, ethane, ethylene and acetylene from oil decomposition; carbon monoxide and dioxide from cellulose. Oxygen and nitrogen reflect how the oil is sealed.
Duval triangle and pentagons
Graphical methods that place relative gas proportions into zones for partial discharge, low- and high-energy discharges and thermal faults of rising temperature.
Hot-spot temperature
The hottest point in the winding insulation, estimated from top-oil temperature, load and design data; it sets how fast paper ages.
Aging acceleration factor
In IEEE C57.91, the ratio of insulation aging at a given hot-spot temperature to aging at the reference temperature; integrating it over time gives equivalent aging4.
Furans and degree of polymerization
Degree of polymerization measures paper strength directly but needs a paper sample; furanic compounds in oil indicate paper degradation indirectly.

How the established DGA interpretation methods differ

Each method answers a slightly different question, so practitioners combine them, and an ML model should treat them as features rather than competitors.

MethodWhat it usesStrengthLimitation
Key gas methodThe dominant gas in the mixSimple to explain to field staffCoarse; mixed faults confuse it
Gas ratio methodsRatios of gas pairs read against code tablesLong history and familiar to engineersSome combinations match no code at all
Duval triangleRelative shares of methane, ethylene and acetyleneAlways returns a zone; separates thermal from electrical faults wellIgnores hydrogen and ethane; unreliable near detection limits
Duval pentagonsHydrogen plus four hydrocarbon gasesAdds detail on partial discharge and stray gassingDiagnoses fault type, not how fast it is developing
IEEE C57.104-2019 status approachGas levels and rates of increase against statistical normsSeparates whether gassing is unusual from what fault it suggestsPopulation norms may not match your own fleet
Supervised ML classifierAll gases, rates, loading and history togetherCan learn fleet-specific patterns and interactionsNeeds labeled faults that most fleets lack

A common pattern is to use the status approach to decide whether gassing is abnormal and a graphical method to suggest the fault type2; IEC 60599 sets out the equivalent interpretation guidance in IEC markets3.

Thermal loading, moisture and insulation aging

Paper insulation ages with temperature, moisture and oxygen, and its loss of strength usually ends a transformer's life. The IEEE loading guide estimates hot-spot temperature from top-oil temperature and load, then converts it into an aging acceleration factor4; IEC 60076-7 plays the same role in IEC markets. Applied to real load history, it gives an equivalent age that can differ sharply from calendar age.

Inputs are modest: loading from SCADA or aggregated meter data, ambient temperature, cooling status and nameplate thermal data. Fiber-optic hot-spot probes on newer units replace estimates with measurements and help calibrate sister units without probes.

Moisture speeds aging and lowers dielectric strength. Readings swing with temperature as water moves between oil and paper, so use relative saturation or temperature-normalized values, not raw parts per million.

Signals to inventory before building a health model

0 of 8 checked

From raw transformer signals to an intervention decision

01Collect signals02Validate and align03Apply standard methods04Score health05Add consequence06Rank and choose action
  1. Collect signals

    Lab and online DGA, loading, oil tests, bushing and tap changer data.

  2. Validate and align

    Flag lab changes, drift and gaps; align each unit to one timeline.

  3. Apply standard methods

    Compute status and fault-type results engineers recognize.

  4. Score health

    Combine interpretations, trends and equivalent age, with stated drivers.

  5. Add consequence

    Weigh safety, environmental, network and financial impact at that site.

  6. Rank and choose action

    Choose resampling, monitoring, repair or replacement, and record why.

Conceptual flow of a transformer health and risk workflow. Engineers can review each stage; it is not a measured result or a product diagram.

Rule-based health indices, ML scores and explanations engineers accept

A rule-based health index scores each diagnostic, weights the scores and often lets the worst subsystem set a floor, so a failing bushing cannot be averaged away by good oil. It is transparent and auditable, but weights rest on judgment and fixed thresholds treat every unit alike.

Supervised failure prediction meets a hard limit: failures are rare and causes are recorded inconsistently. Three alternatives fit the data most fleets have. Peer comparison flags units that diverge from similar ones of the same design, age and duty. Learning from expert assessments makes the rankings engineers already produce consistent. Survival analysis estimates failure probability over time from age and condition, using retired units as well as failed ones.

The practical answer is usually a hybrid. Standard interpretations become inputs and guardrails, the model adjusts within them, and every score shows its drivers in engineering terms, such as acetylene rising across consecutive samples while load is flat. A score that cannot be explained that way will not change a maintenance plan.

Turning a health score into a fleet risk ranking

Health describes likelihood of failure; risk adds consequence. Consequence depends on the site: proximity to the public and fire exposure, oil near watercourses, the customers served and whether another transformer can carry the load, and the cost and lead time of replacement. Great Britain's distribution network operators work with a regulator-approved methodology that combines health-based probability of failure with financial, safety, environmental and network consequences to give monetized risk5, a useful reference even outside that regime.

The ranking should end in actions: resampling sooner, fitting an online monitor, oil processing, repair, relocating a spare, derating or replacement. Recording each unit's chosen action and reason turns the ranking into a plan auditors and regulators can follow.

Data pitfalls in transformer analytics

Laboratory variability

Early signalFleet-wide step changes when the lab or extraction method changes.

MitigationRecord laboratory and method with every result, flag changes and compare labs on shared samples.

Scarce, inconsistent failure labels

Early signalFailures logged as unknown, or by outage cause rather than mechanism.

MitigationHarmonize cause codes, capture teardown findings and use expert assessments as labels.

Online sensor drift

Early signalSlow trends across many monitors of one model, or disagreement with lab samples.

MitigationCross-check monitors against periodic lab samples and recalibrate or correct drift.

Sampling and survivorship bias

Early signalSuspect units are sampled more, and early replacements vanish from the data.

MitigationModel sampling frequency explicitly and keep retired units in the training set.

A hypothetical distribution utility ranks substation transformers

Questions and answers

Can machine learning replace the Duval triangle or IEEE C57.104 interpretation?

It should not try to. The standard methods encode long field experience and are what engineers, insurers and regulators recognize. Machine learning is most useful combining their outputs with trends, loading and other diagnostics to rank a large fleet consistently. When a model disagrees with the standard interpretation, treat that as a prompt for engineering review, not as the answer.

Are online DGA monitors worth it compared with laboratory samples?

For critical or suspect units, often yes: online monitors catch fast-developing faults between lab samples and show rates of change in near real time. For the bulk of a fleet, periodic lab analysis usually remains the economic choice. Many utilities use risk rankings to decide where monitors go and keep lab samples as the calibration reference for them.

What data does a transformer health model need to start?

A usable start is several years of lab DGA per unit, nameplate and installation data, and loading history. Oil quality, bushing and tap changer results and failure records improve it. Gaps are normal; first inventory which units have which signals, so the model can state lower confidence where data is thin.

How do you validate a transformer risk model without many failures?

Combine several checks. Back-test against the failures and major repairs you do have, compare rankings with blinded engineering assessments, and inspect a sample of high-ranked and low-ranked units to see whether findings match. Track how rankings change as new samples arrive. Validation is continuous, because each intervention and teardown adds evidence.

Sources

  1. Electric Power & Natural Gas: predictive asset maintenance use case — ColdAI
  2. IEEE C57.104-2019: IEEE Guide for the Interpretation of Gases Generated in Mineral Oil-Immersed Transformers — IEEE · checked 10 October 2026
  3. IEC 60599:2022 Mineral oil-filled electrical equipment in service: guidance on the interpretation of dissolved and free gases analysis — International Electrotechnical Commission · checked 10 October 2026
  4. IEEE C57.91-2011: IEEE Guide for Loading Mineral-Oil-Immersed Transformers and Step-Voltage Regulators — IEEE · checked 10 October 2026
  5. DNO Common Network Asset Indices Methodology (CNAIM), version 2.1 — Ofgem · checked 10 October 2026

More in Electric Power & Natural Gas

Back to Electric Power & Natural Gas

Next step

Send us your transformer fleet data inventory

List the diagnostics you hold, how far back they go and how failures are recorded. We will reply with a view on whether a health and risk ranking is feasible now and what to fix first.

Discuss transformer analytics