Bias in transformer health assessment matters because the resulting priorities can influence inspections, maintenance and investment. A McKinsey discussion of bias in AI and human judgement provides background to this question. For an engineering comparison, however, broad discussion must be followed by a defined assessment method and evidence that others can examine.
Project: Human and Ronin was introduced in 2022 to explore differences between expert judgement and AI-based transformer health indexing.
This article sets out the exercise as an exploratory protocol: which information to compare, how to define a reference assessment and what evidence would be needed to evaluate human and model decisions.
Bias risk in the transformer industry
Condition assessment serves different operational decisions. A utility may need to prioritize network assets with high consequences of failure; an industrial operator may also be concerned about production interruption and the availability of a spare unit. Those priorities can differ even when the underlying condition evidence is the same. A study should distinguish differences in operational objectives from errors in interpreting measurements.
Transformer health indexing combines selected indicators into a score. It can help organize a fleet review, but the score depends on the inputs, thresholds and weighting method. Condition should be considered alongside criticality when prioritizing work. A health score alone is not a complete risk assessment, and disagreement with that score does not by itself demonstrate bias.
Project: Human and Ronin
The project invited asset managers and industry advisors to examine transformer-oriented decisions and consider how to identify, evaluate and reduce bias. The exercise described three transformer cases using condition-monitoring information.
The input categories listed for the exercise were:
- Gas-in-oil concentrations in ppm for hydrogen, methane, ethane, ethylene, acetylene, carbon monoxide and carbon dioxide
- Water content in ppm
- Acidity in mg KOH/g
- Breakdown voltage in kV
- Interfacial tension in mN/m
- Tan delta, with the reporting convention specified
- 2-furfuraldehyde in ppm
- Transformer age in years
These categories cover several aspects of oil and insulation condition. Their interpretation also requires context: sampling dates, test methods, operating conditions, treatment history and relevant missing information. A comparison should document which of these were available to each participant and to the model.
Seetalabs’ Ronin AI health index is described on a product-specific scale from 0, very poor, to 100, very good. That scale is an assessment convention. It should not be read as a percentage probability of survival, a failure forecast or a universally prescribed CIGRE scale. Any reference to recommendations must identify the document and the specific method used.
Accuracy and error
A fair comparison first needs a reference. This might be an independently established condition assessment or a clearly defined diagnostic outcome. If the task instead asks experts to assign a preferred health index, agreement with one model is not an independent measure of correctness.
Accuracy and error need explicit definitions, denominators and uncertainty estimates. For a numerical health index, a study might measure deviation from an independent reference; for a fault category, it might measure classification errors. Missing-data performance needs a controlled comparison using the same cases, comparable information, predefined handling of missing values and evaluation on cases not used to develop the model.
Asset-wise review
The three cases provide a structure for exploratory review. A larger evaluation would need representative assets, with a separate evidence trail for each case:
- Asset 1: record the available measurements, the reference assessment and the range of participant responses.
- Asset 2: identify which information was missing and whether that changed the interpretation or the confidence assigned.
- Asset 3: examine whether multiple concerns were represented and whether a composite score concealed an issue requiring separate action.
Participant role, operating environment and experience could be investigated as possible influences on judgement. They should be reported with their sample sizes and assessment conditions.
Variability and error discrepancies
Differences between cases may reflect data completeness, task difficulty, model limitations or differences in the reference assessment. These explanations need to be tested rather than inferred from a few summary percentages. A useful review would preserve the original responses and the reasons given for each decision.
Questions for a documented follow-up include:
- Does performance change when complete and incomplete records are evaluated under the same protocol?
- Which fault patterns are missed by the model, by participants or by both?
- How do participants handle data-quality concerns before assigning a score?
- Do operational priorities change recommended actions without changing the diagnosis?
- Does the model communicate uncertainty and allow reviewers to inspect conflicting indicators?
Mitigating the risk of bias in health indexing
For Seetalabs, the practical value of this exercise is to define a more transparent comparison. A follow-up should document the protocol, participant and asset counts, reference labels, model version and evaluation metrics before drawing conclusions.
The NIST AI Risk Management Framework provides voluntary guidance for incorporating trustworthiness into AI design, use and evaluation. It can help organize a documented evaluation and assign responsibilities for reviewing its findings.
The project raises relevant questions about decision support. Demonstrating reduced bias or better performance will require a reproducible assessment with a broader, clearly described set of cases and appropriate review of both human and model errors.




