An oil-and-gas maintenance agent can prepare an investigation pack by resolving equipment identities, retrieving applicable records and checking comparisons with calculation tools. The useful output is evidence an engineer can inspect, including unresolved conflicts. It is not permission to change the plant.
The pilot below tests that limited role with synthetic or replayed compressor records. It has no SCADA write access, no production work-order writes and no automatic physical control. Its workflow, test cases and numerical examples are original proposals.
Define the Deliverable Before Choosing an Agent
Set a task the team can judge: assemble the evidence needed to review a recurring compressor vibration alert, using only records available at a specified cutoff. Require equipment identity, applicable document revisions, a dated event sequence, supported calculations and an unresolved-questions list.
Published industrial agent-building documentation describes defining a workflow and evaluation cases before configuring the agent. This is supplier guidance, not measured evidence of diagnostic competence. Test retrieval, calculations and interpretation separately: finding the right manual does not prove the agent understands the fault.
Keep the pilot independent of a vendor demonstration. An engineer should prepare reference answers and identify acceptable uncertainty before the supplier configures the agent. Include cases where asking for missing information is the correct result.
Give Each Tool a Small, Testable Job
For the hypothetical compressor case, use a frozen asset register, approved manual revisions, maintenance exports and timestamped sensor extracts. Each returned item needs a source identifier and access classification. An archived note can inform the history without becoming an approved current procedure.
Freeze document versions, not just event dates. Suppose a January work order was amended in June with the eventual diagnosis. A replay of a March decision may use only the version actually available in March, not June's text attached to January's date. Record when each revision became available to the intended user. This guards against temporal leakage described in the authors' research on evaluation design.
| Workflow step | Permitted tool result | Required check before continuing |
|---|---|---|
| Resolve equipment | Candidate tag, serial number and historical aliases | Stop the assessment if identity remains ambiguous |
| Retrieve records | Authorized record versions available before the cutoff | Confirm equipment, revision and availability timestamp |
| Extract observations | Value, unit, qualifier and source location | Compare critical extracted fields with the original |
| Calculate comparisons | Output from a versioned, tested calculation service | Validate units, inputs, applicability and error state |
| Assemble evidence | Timeline, disagreements and missing attachments | Distinguish reported facts from proposed explanations |
| Handover | Draft investigation pack with named reviewer | Engineer accepts, corrects or requests further evidence |
The LLM selects among approved tools and explains their results. A validated deterministic engine performs arithmetic and any domain calculations. A failed calculation returns an error; the agent must not invent a plausible replacement number.
For example, suppose two confirmed RMS vibration-velocity readings are 2.0 and 2.6 mm/s, taken at the same point and axis, with the same frequency band and comparable operating conditions. A calculator returns a 0.6 mm/s difference and a 30% increase relative to the first reading. Those hypothetical numbers do not establish a fault or an alarm threshold. The engineer still needs the measurement method, operating point, applicable machine guidance and uncertainty. Do not mix RMS velocity with peak velocity, displacement or acceleration.
A transformer supplying the compressor may have a separate laboratory investigation. Its DGA evidence must retain sampling compartment, liquid type, units and qualifiers. A missing gas and a result below detection remain different from zero; shared equipment context does not make compressor and transformer diagnoses interchangeable.
Test Whether the Evidence Pack Is Correct
Build a small acceptance set with clean cases, conflicting revisions, renamed equipment, missing attachments and intentionally misleading retrieved text. Use records excluded from configuration work. Freeze the expected evidence set for each question, and keep later investigation outcomes out of the replay.
Measure retrieval recall against the engineer's required evidence set, citation support against factual claims, identity accuracy, calculation agreement and correct abstention. Report the denominator for each. Also record elapsed time, reviewer correction time and failed tool calls; a quick answer that takes longer to repair is not a useful saving.
In an illustrative test, reviewers define 80 required evidence items across 20 cases. Count each case-item pair once after deduplication. The agent retrieves 72 of those pairs and 30 irrelevant pairs, giving 72/80 = 90% retrieval recall but only 72/102 = 70.6% retrieval precision for that test. Reviewers also find that 45 of 50 factual claims in the generated packs are supported by the cited records, giving 90% claim support. These pooled percentages are not diagnostic accuracy, and they do not make the five unsupported claims acceptable. A citation supporting a claim does not establish that the underlying record is correct.
Investigate which eight evidence items were missed. A missing superseding procedure may matter more than several omitted background notes. Keep such critical misses visible instead of averaging them away. Repeat the same cases to expose inconsistent tool selection and use a fresh holdout set after repairs.
The wider condition-monitoring workflow still needs an accountable reviewer. Pilot acceptance should establish that evidence preparation is useful and controlled, not that failures have been prevented.
Treat Retrieved Text as Data, and Review Changes Separately
A maintenance attachment can contain instructions that try to redirect an agent. OWASP identifies prompt injection as a risk that retrieval-augmented generation alone does not remove. OWASP LLM01:2025.
Insert a synthetic attachment asking the assistant to export records or close a work order. Score three things separately: whether the model requests a forbidden action, whether the executor denies it, and whether any unauthorized disclosure or state change occurs. A blocked request exposes a model failure but may demonstrate that the executor boundary worked. A polite refusal alone proves neither.
In the authorized test harness, call the executor directly with forbidden operations. Also test an allowed retrieval tool against another user's asset, and an export directed to an unapproved recipient. The executor must check identity, permissions, target and destination on every call; valid argument syntax is not authorization. Use read-only source credentials, restricted destinations and an isolated output store. Retain denied-call evidence and test access revocation.
Prefer synthetic records. Real maintenance exports can contain names, signatures and contact details: minimize these, define permitted purposes and recipients, and set access, retention and deletion rules for records and execution logs. Where GDPR applies, internal pilot approval does not replace a lawful processing basis. Have the responsible privacy team check external-provider and transfer arrangements. European Commission GDPR principles.
For US processes covered by OSHA's process safety management rule, 29 CFR 1910.119(l) requires written management-of-change procedures for changes within its scope, except replacements in kind. Have the responsible process-safety team assess whether a proposed change to a real maintenance workflow falls within that scope before proceeding. Calling a trial isolated does not itself establish an exemption. OSHA's published rule.
Start with one recurring evidence-preparation task and its failure cases. To discuss the transformer-assessment component, Talk to an engineer.
Sources and Boundaries
- The linked industrial agent documentation is supplier guidance; no deployment outcome or benchmark result is claimed.
- The linked leakage research supports cutoff discipline; the document-version test is an original application.
- OWASP LLM01:2025 supplies security context. OSHA 29 CFR 1910.119(l) is cited only for covered-process change review.
- The European Commission guidance supports the personal-data boundary, not a finding of legal compliance.
Evidence checked through 30 September 2026. The compressor case, tool contract and evaluation counts are hypothetical, not results of a pilot run. No agentic Ronin AI integration or autonomous expert capability is claimed.




