Seetalabs
Product IndustriesTools Knowledge Base About Contact Discuss Ronin AI
Transformer failure with flames at substation - predictive maintenance and early fault detection prevent these events

Synthetic data, LLMs and edge AI: what they actually mean for transformer diagnostics

Synthetic data, language models and edge computing address different parts of a diagnostic workflow. Their value depends on the task, the evidence and the conditions of deployment.


Industrial AI is often discussed as if one model could solve every asset-management problem. Transformer diagnostics presents a more specific set of needs: interpreting measurements, finding relevant history, identifying uncertainty and deciding when an engineer should investigate.

Those tasks require different evidence. A benchmark result on text, for example, does not establish fault-detection performance on a transformer fleet.

This article examines selected research and technical literature available through early 2026. The sources include journal studies, reviews, arXiv and SSRN manuscripts, and technical commentary. These publication types should be distinguished when assessing a claim.

The useful question is what each study tested and how closely its conditions match the proposed application. Dataset composition, the reference diagnosis and the evaluation split can change the meaning of an apparently strong result.

The following sections connect those research questions to practical transformer assessment.

The persistent problem: we still do not have enough failure data

Confirmed failure cases can be difficult to assemble into a representative training dataset. Repeated measurements from normal operation do not provide the same information as independently labelled fault events. The number of readings and the number of distinct assets are therefore both relevant.

The Applied Sciences review on DGA fault classification and machine learning provides a starting point for the diagnostic literature. When evaluating this literature, check the distribution of fault classes, how labels were established and whether repeated samples from the same transformer appear in both training and test data.

Balancing methods can change the data seen during training through sampling, filtering or generation. The Energy Reports article linked in the related reading is a reference for further examination. A method should be selected by its performance on an appropriate independent test set, not by a headline accuracy figure without its experimental context.

A balanced test set and an operating fleet with infrequent faults answer different evaluation questions. Overall accuracy can conceal missed faults or a high false-alarm workload. Report class-specific performance and retain a test distribution that reflects the intended use, or explain how the difference is handled.

Synthetic data is one possible response to limited coverage. It can support experiments, but the generated records need their own checks and should remain identifiable throughout development.

Synthetic data: the promise and the pitfalls

A January 2026 review in the Journal of Intelligent Manufacturing examined 86 studies on synthetic data for predictive maintenance. Its taxonomy covers classical augmentation, deep generative methods, physics-informed or hybrid methods, and feature-space transformations. These approaches create different types of variation and make different assumptions.

For transformer diagnostics, physical plausibility is a useful design requirement. Physics-informed generation may help constrain examples, but the presence of a physical model does not itself establish that generated samples represent real faults or improve a deployed classifier.

For synthetic DGA records, review the relationships among gas concentrations, fault mechanisms and operating context. Oil type, temperature, sampling, treatment and mixed faults can complicate interpretation. Matching marginal distributions or reproducing a familiar gas ratio is insufficient to establish a realistic joint record.

A practical evaluation can compare a model trained on field data with one trained on the same data plus synthetic examples. The final test data should remain independent of the generator and represent actual observations. Report where augmentation helps, where it harms and which classes remain poorly covered.

The October 2025 arXiv manuscript on synthetic data in LLM pretraining addresses a different setting: single-round language-model training with natural text, synthetic text and mixtures. It studies how generation method and mixture affect language-model performance. It is not a transformer-diagnostics experiment.

The manuscript does not establish that recursive model collapse occurs only when all real data is replaced. Nor does it show that every mixture improves performance. Findings about a specific language-training setup should not be turned into a universal rule for repeated model training or DGA augmentation.

For diagnostic development, preserve traceable field records and an independent evaluation set. Synthetic scenarios can help explore rare or difficult conditions, while the field observations remain necessary to assess whether the resulting model is useful on real equipment.

Large Language Models: useful, but not for what you might think

Language models can help retrieve and summarize text, but their role in diagnosis needs a clearly defined boundary. A fluent explanation can contain a wrong value, an unsupported interpretation or a reference that does not support the answer. Evaluation must examine those errors directly.

Results from another industrial domain cannot be transferred to transformers without new testing. A compressor-monitoring benchmark, for example, may use different sensors, labels and event definitions. Numerical recall or AUC results are meaningful only with that context and should not be treated as evidence of general diagnostic superiority.

One candidate application is to connect structured measurements with maintenance notes and inspection reports. This requires reliable retrieval, correct asset matching and preservation of dates and units. The system must also distinguish a confirmed finding from an operator hypothesis or a copied historical note.

A useful output could point an engineer to the records relevant to an assessment and show the supporting passages. Whether a language model improves that task should be tested against the existing search and review process, including unsupported-answer rates and the time needed to verify its output.

An ASRC Procedia review published in April 2025 describes hybrid approaches using multiple diagnostic sources. It identifies data heterogeneity, cybersecurity and deployment standardization as challenges. This indicates a research direction; it does not establish that multimodal language models are an adopted industry standard.

RONIN AI brings together transformer condition indicators such as DGA, oil-quality information, paper-condition indicators and age. Extending a workflow to unstructured records requires a separate evaluation: what information is retrieved, how it is checked and whether it changes the assessment in a useful way. Existing product functionality should be distinguished from a proposed extension.

Edge deployment: bringing intelligence to remote substations

Here is a practical problem that anyone managing geographically distributed transformer fleets understands: connectivity is not always reliable. Remote substations may have intermittent network access or none at all. Sending data to cloud servers for analysis introduces latency and creates dependencies on infrastructure you do not control.

Running suitable analysis locally can reduce dependence on a continuous cloud connection. It does not remove every connectivity problem: data synchronization, remote support, model updates and central fleet review may still require communication.

A review published in Frontiers in Computer Science on 14 May 2025 discusses LLMs and edge computing across application areas. It provides architectural context and identifies resource constraints. It is not evidence of a validated transformer installation.

Model size is one constraint, alongside memory bandwidth, energy consumption, latency and the software runtime. The device must support the required workload while sharing resources with other monitoring functions. A cloud-service model name is not a specification for what can run on a substation computer.

Quantization reduces numerical precision and can reduce model storage. The January 2026 survey on deploying LLM Transformers at the edge is related reading. Moving a parameter representation from 16 to 4 bits reduces those raw parameter bits by three quarters; total runtime memory also includes metadata, activations, caches and other overhead. Quality and resource savings must be measured for the actual configuration.

A hardware benchmark helps identify practical constraints, provided its tasks are kept separate from the intended diagnostic task.

The authors’ preprint of the study subsequently published in ACM Transactions on Internet of Things evaluates 28 quantized model variants on a Raspberry Pi 4 with 4 GB of RAM. The benchmark covers five language and reasoning datasets and measures energy, latency and output accuracy. The ACM publication record identifies the journal version.

These measurements do not establish that Q4 or Q8 quantization is inherently more reliable for transformer diagnosis than Q3. Energy variability and diagnostic error are different outcomes. Select a compression level by testing the intended task on the target hardware, including difficult examples, memory pressure and the consequences of incorrect outputs.

Small Language Models: a pragmatic choice for industry

Small language models may fit selected local tasks within tighter resource limits. Their suitability depends on the application rather than a market-growth forecast. A smaller model that meets a specific retrieval or classification requirement can be more practical than a larger model whose additional capability is unused.

Offline operation. A technical article on industrial IoT discusses local small-model applications. This is technical commentary rather than a peer-reviewed field benchmark. For a remote substation, test the complete workflow without connectivity, including access to records and handling of stale information.

Data privacy. Local processing can reduce the transfer of commercially sensitive information to external services. It does not automatically keep every record on-site: telemetry, backups, updates and remote-access arrangements must also be examined. Security responsibilities remain necessary on local hardware.

Latency. Local inference can avoid a network round trip, but its own processing time may still be substantial. Define the required response time for the task and measure it under expected load. Monitoring support should not be confused with a protection function requiring a separately engineered response.

Questions to ask of the research

A review manuscript posted to SSRN in April 2025, dated January 2025, discusses transformer fault diagnosis and identifies interpretability, dataset and deployment challenges. An SSRN posting alone does not establish peer-review status. The practical questions below should be applied to individual studies and proposed deployments.

Can another team inspect the data or reproduce the evaluation? Proprietary records may be necessary, but the study can still describe asset counts, sampling, labels, exclusions and evaluation splits. Open benchmarks are useful when they represent the intended task and are not mistaken for the entire operating population.

Does the evaluation measure the errors that matter? False negatives can leave a developing fault unaddressed. False positives can trigger additional tests, avoidable outages or misplaced maintenance effort. Precision, recall, alert burden and the consequences of each action should be considered alongside overall accuracy.

Does the evidence come from a retrospective test, a controlled pilot or routine operation over time? These stages answer different questions. A held-out test set can measure generalization within a defined dataset; it does not by itself demonstrate sustained performance as assets and operating conditions change.

Can engineers trace a recommendation to its inputs? Feature-attribution methods such as SHAP can show how variables influence a model output, but they do not prove a causal mechanism or validate the diagnosis. The Technologies article listed below is further reading. Any explanation method should be checked for stability and presented alongside the underlying measurements and model limits.

What this means for asset managers

For asset managers, the research supports a practical evaluation agenda:

Define the role of the language model. Retrieval and reporting assistance, fault classification and maintenance decisions are distinct tasks. Compare each proposed function with the relevant diagnostic methods and engineering review process.

Evaluate synthetic data against real observations. Preserve provenance, isolate the test set and measure the incremental effect of augmentation. A realistic-looking sample is not sufficient evidence of a useful training example.

Treat edge deployment as a complete system. Model compression is only one part. Hardware limits, integration, cybersecurity, updates and support all need testing in the intended environment.

Improve data quality and compare baselines. Consistent units, timestamps, labels and maintenance history make evaluation more informative. Whether a simpler or more complex model performs better remains an empirical question.

Where things are headed

Combining these techniques creates possible workflows worth testing, rather than a predetermined route to autonomous diagnosis.

A transformer diagnostic system of the near future could process DGA data in real-time on local hardware, with no cloud dependency. It could automatically integrate structured sensor readings with unstructured maintenance notes and inspection reports. It could generate synthetic scenarios to stress-test predictions and identify edge cases the training data did not cover. And it could provide clear explanations for its recommendations, citing the specific factors that influenced each decision.

Each component needs evidence for its own function and for its interaction with the others. A system can retrieve the correct record yet misinterpret it, or classify a sample correctly while recommending an unsuitable action.

For teams building tools in this field, the next step is a bounded pilot with representative data, an agreed reference and a documented review process. Record failures and uncertainty as carefully as successful demonstrations. This is how a research result can become evidence for a particular deployment.

The value lies in decisions that become more informed and reviewable, with the limits of the available evidence visible to the engineer.


For transformer engineers, asset managers and grid operators. References include journal publications, reviews, preprints and technical commentary; results from other domains require separate validation for transformer diagnostics.


Selected publication records and further reading. Publisher access and the availability of full text vary: