In short: There is no universal normal gas level. IEEE C57.104-2019 replaced fixed limits with 90th and 95th percentile values conditioned on transformer age and the oxygen-to-nitrogen ratio, so the same methane reading can be unremarkable in one unit and worth investigating in another. Rate of change carries equal diagnostic weight, and for acetylene any increase counts. Ratios computed below the analytical floor are noise, not evidence.
What is a normal dissolved gas level in a transformer?
The honest answer is that the question does not have the shape you expect. There is no number. There is a distribution, and where your unit sits in that distribution depends on how old it is and whether its oil sees air.
That sounds like an evasion, so let us be concrete about why it is not. Search for acceptable gas levels in transformer oil and you will find pages handing you a tidy list: a figure for hydrogen, one for acetylene, a caution band, an alarm band. Those numbers came from a real place. They were the condition tables of the 1991 and 2008 editions of the IEEE guide, and for thirty years engineers used them because they were the best consensus available. The 2019 revision withdrew that approach. It did not tighten the numbers or loosen them. It concluded that one set of numbers applied to every transformer was the wrong instrument, and replaced them with percentiles computed separately for different populations of machine.
So if you are looking for a normal level, the question you actually need answered is narrower and more useful: normal for what kind of transformer, at what age, with what oxygen exposure, and moving how fast. Ask it that way and it becomes answerable.
People who go looking for fixed limits are not being lazy. They usually have a spreadsheet of lab results, a maintenance window closing, and a need to sort fifty units into ones they can leave alone and ones they cannot. A single threshold column is the fastest tool imaginable for that job. It is also the reason a large fraction of the DGA screening in service today is quietly non-conformant with the standard it claims to follow.
Why the fixed limits you found online are obsolete
The old condition tables had a structural problem that was visible long before 2019. They were built from a population of transformers that no longer resembles the population in service, and they made no distinction between a two-year-old sealed unit and a forty-year-old free-breathing one. A single hydrogen number had to cover both. It could not.
What IEEE C57.104-2019 changed
The 2019 edition rebuilt the screening criteria from a large statistical study of laboratory DGA results. Clause 6.1.3 sets out the norms in four tables. The first gives 90th percentile gas concentrations as a function of the O2/N2 ratio and transformer age. The second gives the 95th percentile of the same. The third and fourth are not about level at all: they give 95th percentile values for how much a gas moved between two consecutive samples, and for the rate fitted across a series of three to six samples.
The output is no longer a condition number. It is one of three DGA statuses: acceptable screening results with routine operation continuing, then incipient or modest recent gas production, then high levels or continuing significant production where mitigative action is warranted. Note the hedging in the standard’s own language. It says a transformer above the statistical norms “is not necessarily faulty,” only that its behaviour is unusual for its population. That is a screening tool talking, not a diagnosis.
The percentile framing also tells you how much work to expect. A 90th percentile norm will, by construction, flag about one result in ten for review. The standard says this openly. If your screening flags nothing across a fleet of a hundred units, your thresholds are wrong, not your fleet.
Age and O2/N2: why the same number means two different things
Two conditioning variables do the heavy lifting. Age enters as bands: the first decade, roughly the second and third, and beyond thirty years, with a separate column for units whose age is unknown. Oxygen exposure enters through the ratio of dissolved oxygen to dissolved nitrogen, split at 0.2. The standard is explicit that this ratio is a proxy rather than a label: a value at or below 0.2 is seen in most nitrogen-blanketed units and in a majority of membrane-sealed ones, while a value above 0.2 is seen in air-breathing units. It is a way of asking the oil whether it has been meeting air, which is a better question than asking the nameplate what kind of preservation system was fitted in 1978.
The spread between those populations is not a rounding difference. For the low-temperature hydrocarbon group, the sealed-unit screening value sits at roughly five times the air-breathing one. For acetylene in the oldest age band the sealed figure runs on the order of ten times the air-breathing figure. Hydrogen is closer, around double. Apply an air-breathing criterion to a sealed unit and you will generate false alarms across your fleet; apply a sealed criterion to an air-breathing one and you will sail past gas production that the standard would have flagged.
There is a second-order trap worth naming. Because oxygen and nitrogen are themselves measured quantities, a unit near the 0.2 boundary can flip categories between samples, and its screening criteria flip with it. The guide notes this. Knowing it happens is the difference between investigating an apparent step change and shrugging at it.
Why the trajectory ranks equal to the level
Half of the 2019 norms describe motion, not position. That is the part most screening implementations drop, and dropping it is not a partial implementation. It removes the more time-sensitive of the two signals.
A concentration is an integral. It tells you how much gas the oil has accumulated since it was last processed, which may be a decade of unremarkable ageing plus three weeks of something new. A rate tells you what is happening now. A transformer whose ethylene has doubled in four months while sitting comfortably below every level criterion is a transformer with something going on, and the level tables will say nothing about it for another year. This is why the rate norms in the 2019 guide are built specifically from samples where all gas levels were below the 90th percentile table. They exist precisely to catch the quiet ones.
The two rate tables answer different questions. One looks at the step between two consecutive lab results, which is the natural thing to compute when a new sample lands. The other fits a rate across a series of three to six points over a period between roughly four months and two years, which suppresses the noise of any single sample and is the honest way to characterise a slow drift. Use both. They disagree often enough to be informative when they do.
The acetylene exception
Acetylene does not get a rate threshold. In both rate tables, the entry for acetylene is not a number. It reads “any increase” in one and “any increasing rate” in the other.
The chemistry justifies the asymmetry. Acetylene requires the triple carbon-carbon bond, which needs far more energy than the bonds broken by thermal faults, and in oil it is associated with arcing at temperatures above roughly a thousand degrees. There is no benign slow accumulation mechanism the way there is for methane or ethane. If acetylene is rising, something is discharging. The guide suggests operators may wish to have a policy of watching for active acetylene production, especially at low levels where it would otherwise be dismissed.
One caution attaches to this. The 2019 guide warns specifically about the case where acetylene is the only gas above its screening value, because that pattern is more often an analytical artefact or a load tap changer contamination path than a fault in the main tank. Rising acetylene is a reason to sample again quickly and check the story, not a reason to trip a unit.
Why ratios stop meaning anything at low concentrations
Everything above concerns screening: is this unit worth a closer look. Fault identification is a different exercise, and it runs on ratios between gases rather than on the gases themselves. Duval triangles, the IEC ratio scheme, the Rogers approach: all of them take proportions and map them onto fault types. Our guide to dissolved gas analysis of power transformers works through how each of those methods reaches its verdict.
Ratios have a failure mode that levels do not. Divide one small uncertain number by another and the uncertainty does not add, it compounds. IEC 60599:2022 puts figures on this in clause 6.2. Above ten times the analytical detection limit, uncertainty on a DGA value is typically around fifteen percent. Below that it climbs fast, reaching roughly thirty percent at five times the limit. Two values each carrying thirty percent error produce a ratio you should not steer maintenance decisions with.
The same standard is directive about it. Clause 6.1 says gas ratios are significant and should be calculated only when at least one gas is above typical values in both concentration and rate of increase, and it says plainly: avoid calculating ratios when gas concentrations are not high enough to be reasonably accurate. That instruction is widely ignored, because software that always returns a fault type looks more capable than software that sometimes declines.
Picture what ignoring it produces. A healthy transformer with barely measurable hydrocarbons still yields a defined ratio, that ratio still lands somewhere on a Duval triangle, and the triangle still names a zone. You now have a thermal fault classification for a machine with no fault.
What a system should do when it cannot tell
It should say so.
That sentence is short because the principle is. When gases sit below the analytical floor, or when a ratio has a zero or an unmeasurable quantity in its denominator, the correct output is not determined. Not a low-confidence guess. Not a default to the most common fault type. Not determined, with the reason attached, so the engineer reading it knows whether to order a fresh sample, request a lower detection limit from the lab, or simply wait.
Abstention is a design decision, not a gap
Abstention has to be designed in rather than left to emerge. The IEC ratio scheme declines by construction when a ratio is not calculable. The secondary Duval triangles need gating so that they only run after the primary one has established that their preconditions hold, because Triangle 4 is defined for cases already identified as partial discharge or low-temperature thermal, and Triangle 5 for cases already in the thermal range. Running them unconditionally would produce zone assignments outside the domain their author defined.
Anyone who has spent a career reading lab reports knows the value of this, because the same discipline governs their own written judgements. Nobody credible writes a fault type into a condition report when the data will not carry it. Software has historically been rewarded for the opposite, since a screen that fills every field reads as more complete in a demo.
Why raw accuracy is the wrong way to judge a diagnostic system
Here is the uncomfortable corollary, and it works against us commercially.
Accuracy on a labelled DGA dataset counts an honest abstention exactly the same as a wrong diagnosis. Both are scored as not correct. A system that guesses on every low-concentration sample will therefore beat a system that declines on them, on the metric, while being worse in the field. The metric rewards the behaviour you do not want.
Published accuracy figures for tabular DGA classification scatter widely, and the scatter is the finding. The same family of method, evaluated on different datasets, lands anywhere from the low seventies to the high eighties, which tells you the number is a property of the evaluation set at least as much as of the method. Figures of 98 to 100 percent appear regularly in vendor material and in the literature, and they almost always indicate a dataset small enough that a handful of samples moves the number several points, or leakage, where samples from the same transformer land on both sides of the train and test split. So ask anyone quoting an accuracy figure two questions: how many distinct transformers were in the evaluation set, and was the split done by transformer or by sample. If the answer to the second is by sample, the figure is not measuring what you think it is.
Better questions to ask of a diagnostic tool are about behaviour rather than score. On what fraction of samples does it decline, and does that fraction correspond to the samples a competent engineer would also call inconclusive. When it is confident, how often is it right. When it is wrong, is it wrong in the safe direction or the other one.
What the published method chain covers, and where it stops
The chain runs from fault identification through to condition, and every link rests on a named source: Duval Triangle 1 for the primary fault type; Triangles 4 and 5 for mineral-oil sub-types under the gating described above; Triangle 3 and its refinements for esters, by liquid family; the Duval pentagons; the IEC 60599 ratio scheme; the C57.104-2019 conditioned percentiles for both level and rate; oil quality following IEC 60422 by voltage class; 2-FAL and degree of polymerisation for the paper; and a deterministic health index following CIGRE TB 761. An implementation that can name its source for each link is one you can audit, which is the question worth putting to any vendor. Ours is described on the RONIN platform page.
The more useful list is where the published sources stop, because that is the point at which an implementation is either honest about it or quietly inventing.
One more limit belongs on that list, about the health index itself. It is a constructed index, not a measurement: nothing physical in a transformer corresponds to a number between zero and a hundred. It is a defensible way to rank a fleet so that limited outage windows go to the units that need them most, and it should always travel with a statement of confidence. Treating it as a verdict on a single machine asks it to do something it was never built to do.
How to use this on your own fleet
If you are screening against a fixed threshold column, the first job is to find out where those numbers came from. Usually they are inherited from a spreadsheet whose author left years ago, and they trace back to the 2008 edition of the guide or a manufacturer’s recommendation from the same era. Neither is fabricated. Both are superseded.
Then check whether you have the two conditioning variables at all. Age is usually somewhere in the asset register. Oxygen and nitrogen are the gap: labs report them, DGA extracts drop them on import, and without the O2/N2 ratio you fall back to the unknown column, which is broader and less sensitive. Getting those two columns into your pipeline is often the single change that improves screening most, and it costs nothing at the lab.
After that, look at whether your process computes rates at all, and over what window. A tool that only compares the latest result against a level threshold is doing half the screening the standard describes. Annual sampling makes the multi-point rate impossible to compute over a short period, which is an argument for sampling more often on the units where it matters.
Last, be suspicious of any tool, ours included, that returns a confident fault type for a sample where every gas sits near the detection limit. Ask what it does when it cannot tell. If the answer is that it always tells you something, you have learned what you needed to know.
Primary sources
Values from these documents are not reproduced here. Clauses are cited so that anyone holding a licensed copy can check the reasoning against the source.


