Forward-Deployed Engineering · Process manufacturing · Maintenance and reliability

Case study: maintenance, from alarms and technician notes to planned work

Illustrative engagement — not a client record. The company, people, volumes and results below are a representative composite, written to show how a forward-deployed engagement runs end to end. The workflow, the architecture, the controls and the method are real and technically valid.

Offering usedEmbedded Engineering PodsA pod inside your repository and your standups, quarter by quarter.

Live callIn region

Illustrative results, first 90 days after the last cut-over:

  • 91%

    Closed work orders carrying a usable failure mode

    Before 22%

  • 63%

    Planned share of maintenance labour hours

    Before 54%

  • 3.1 hours

    Median time from a condition alarm to an approved work order

    Before 19 hours

  • 71%

    Drafted work orders a planner approved with no or minor edits

    Before not applicable

Explainer09 sheets

The whole story in about a minute

Nine short scenes, from years of messy notes to the results. Press play, or pick a scene.

View ALost history

One vacuum pump, two years

Over 300,000 work orders in 14 years

  1. Same cause again
Two board machinesAround the clockcontinuous operation
2 senior techniciansthe same pump failed, for the same reason
one vacuum pump failed in two years, for the same reason

Sheet01 / 09

ClockBefore

Title

The mill's know-how was stuck in messy notes.

Over 300,000 work orders in fourteen years, each with a short technician note. Nobody could see that one vacuum pump failed four times in two years for the same reason.

Notes
  1. 1.300,000+ work orders, 14 years
  2. 2.Failure code filled on under a quarter
  3. 3.Just over half of hours planned
FDE / 00

At a glance

Client
A single-site paper and packaging-board mill, two machines, continuous operation, around 180 people in maintenance and reliability
Workflow
Condition alarms and fourteen years of technician notes turned into grouped failure modes, drafted work orders a planner approves, and a technician copilot that answers from the mill's own documents with citations
Engagement
Embedded Engineering Pod (one quarter, thirteen weeks). Evaluation & Observability ran inside it, as it does in every engagement. The last two weeks were Handover & Enablement
Team
Three engineers: a named lead, a data engineer and a platform engineer. Client side: the Head of Maintenance (owner), the reliability engineer, two planners, a maintenance superintendent, the EHS manager, an OT security reviewer
Where it runs
The mill's own cloud account and region, reading an OT-side replica. No write path to any control system
Handover
Runbooks executed by the mill's own engineer in week twelve; our access revoked at the end of week thirteen
Illustrative results, first 90 days after the last cut-over:
MeasureBeforeAfter
Closed work orders carrying a usable failure mode22%91%
Planned share of maintenance labour hours54%63%
Median time from a condition alarm to an approved work order19 hours3.1 hours
Drafted work orders a planner approved with no or minor editsnot applicable71%
Copilot answers correct and cited, weekly samplenot measured94%
Safety procedure text paraphrased by the systemnot applicable0, by design

FDE / 01Case study

01 / 12

The situation

In shortFourteen years of repair notes held the answers, but only two senior technicians could find them.

The mill ran two board machines around the clock. Maintenance owned several thousand assets: pumps, gearboxes, vacuum systems, refiners, dryer-section drives, fans and conveyors. Two systems held what the mill knew about them.

The first was the CMMS, with fourteen years of work orders in it — a little over 300,000. Every one had a free-text note written by a technician at the close of the job. The notes were short, abbreviated and mixed English with Hindi words the way the shop floor speaks: "brg noise chkd, greasing done", "mech seal lkg, replaced, running ok". The structured failure-code field was filled on fewer than a quarter of them, and most of those said "OTHER".

The second was the historian, holding process tags, vibration and temperature readings, and the alarm journal from the control system.

Between them sat the planners. Two of them turned requests into work: pick the asset, guess the likely cause, find the job plan, check the spare, book the window. What made this hard was not the planning. It was that nothing in the mill could answer "has this happened before, and what fixed it?" without a person who had been there for fifteen years reading through notes.

The cost showed up in three places:

  • Repeat failures. The same vacuum pump failed four times in two years for what turned out to be the same reason. Nobody could see the pattern because each note said something different.

  • Reactive work. Just over half of maintenance hours were planned. The rest were interruptions.

  • Knowledge in two heads. Two senior technicians were the mill's real search engine for past fixes, and one was two years from retirement.

FDE / 02Case study

02 / 12

Why it had not been done

In shortThe notes looked too messy, safety felt risky, and no one owned the whole job.

There was no stalled pilot here. The work had simply never started, for four reasons that are worth writing down:

  1. The data looked unusable. Everyone who had opened the notes concluded they were too messy to analyse. They were messy. They were not unusable.

  2. A reporting project had been tried instead. A dashboard rebuild two years earlier counted work orders by asset. Counting work orders tells you where hours went, not what failed or why.

  3. Anything touching maintenance touches safety. Nobody wanted a system near lockout procedures, and no vendor had offered a design that kept those procedures out of a model's hands.

  4. No owner for the outcome. Reliability wanted failure analysis, planning wanted drafts, the technicians wanted search. Three wishes, no single scope, so the request never made it onto a plan.

The pod engagement exists for exactly this shape: several related workflows, one quarter, one lead accountable for what shipped.

FDE / 03Case study

03 / 12

The quarter's goals (Embedded Engineering Pods)

In shortThree workflows, three owners, and guardrails signed before any work began.

FIG. 3.13 CRITERIA

The goals were agreed with the Head of Maintenance before week one, three workflows, each with a metric and an owner:

WorkflowMetric
  1. Failure-mode grouping of historical work ordersShare of closed work orders carrying a failure mode a reliability engineer accepts
  2. Drafted work orders from condition alarmsShare of drafts a planner approves with no or minor edits; time from alarm to approved order
  3. Technician copilot over the mill's own documentsShare of sampled answers that are correct and carry a citation a technician can open
SIGNED · W2
FIG. 3.2HANDOVER MANIFEST · 3 ITEMS
WorkflowOwner
Failure-mode grouping of historical work ordersReliability engineer
Drafted work orders from condition alarmsMaintenance planner
Technician copilot over the mill's own documentsMaintenance superintendent

Guardrails, written into the goals and not changed afterwards:

  • The system has no control path. It cannot start, stop, reset or set anything. It reads.

  • A drafted work order is a draft. A planner approves, edits or rejects every one. Nothing reaches a technician unapproved.

  • Safety-critical procedure text, lockout and tagout above all, is quoted verbatim from the current revision of the controlled document, by document identifier and revision, or it is not shown at all. The system never paraphrases, summarises or reassembles it.

  • An answer without a citation is not shown.

Timeline09 stations

The quarter, stage by stage

Nine stages on one clock, from before the pod arrived to ninety days after the last cut-over. Press play, or pick a stage.

LineStopped

StationOP 10 Lost history

ClockBefore

Work instructionOP 10 · 01 / 09

The mill's know-how was stuck in messy notes.

Over 300,000 work orders in fourteen years, each with a short technician note. Nobody could see that one vacuum pump failed four times in two years for the same reason.

Outputone vacuum pump failed in two years, for the same reason
Steps
  1. 300,000+ work orders, 14 years
  2. Failure code filled on under a quarter
  3. Just over half of hours planned

FDE / 04Case study

04 / 12

The quarter, week by week

In shortThirteen weeks: read the data, build three workflows one after another, hand them over.

FIG. 4.1W1 – W13

W1Embed

What happened

Pod in the mill's repository and its daily maintenance meeting. Goals signed. Access requested for the CMMS replica, the historian and the document system. First pull requests merged from all three engineers in week one

What existed at the end

A repository, an access register and a read-only CMMS copy

  • What happened

    Pod in the mill's repository and its daily maintenance meeting. Goals signed. Access requested for the CMMS replica, the historian and the document system. First pull requests merged from all three engineers in week one

    What existed at the end

    A repository, an access register and a read-only CMMS copy

The historian delay was reported in the lead's weekly note the week it happened. The pod did not wait: the grouping workflow, which needs no live data, was pulled forward, and the quarter's goals were met without extending the calendar.

FDE / 05Case study

05 / 12

What was built

In shortTagged failure history, drafted work orders a planner approves, and a copilot that cites its sources.

Try a question06 ops

Pick what the technician asks. Watch where it goes.

The copilot answers from the mill's own documents. What is asked decides whether it answers with citations, says it does not know, or shows the controlled text.

Pass06 / 06

Answer shown

Answered, with citations

The answer comes from the mill's own documents, and the technician can open each source.

FIG. 5.1Call path

  1. Reading the notes. The notes were not cleaned into something prettier. They were read as they are. A vocabulary was built first, from the mill's own abbreviations and spellings, with the reliability engineer confirming what each meant. Then a language model, running on a private endpoint in the mill's own region with no data retention, proposed for each work order a failure mode, a mechanism and a cause in a fixed taxonomy the reliability engineer chose, modelled on the reliability-data standard the mill already referenced. Every proposed tag carries the sentence it came from.

  2. Reliability review. Tags were reviewed in batches by asset class. The reliability engineer accepted, corrected or rejected; corrections went back into the vocabulary. Only accepted tags were written to the CMMS, into a new field beside the original code, never over it. The original text is never altered.

  3. Grouping. With tags in place, recurring failure modes group by asset, asset class and cause. The bad-actor list stopped being an argument about which pump felt worst and became a list with work orders behind each line.

  4. Condition rules. Triggers are deterministic and were written with the reliability engineer: a vibration band above threshold for a sustained period, a bearing temperature rising against its own baseline, a run-hours milestone. Not raw control-system alarms, for the reason in the next paragraph. Each rule names the asset, the evidence window and a dead time, so one developing fault produces one draft, not forty.

  5. A false assumption, found in week seven. The scope had assumed the alarm journal was a usable trigger. It was not. Most alarms came from a small number of tags, chattering during grade changes, and floods of alarms arrived in clusters during upsets. Drafting a work order per alarm would have buried the planners in minutes. The alarm journal became evidence attached to a draft, never the reason for one. That finding was written up and handed to the mill's control room as its own piece of work, not absorbed into the quarter.

  6. Drafting. When a rule fires, the draft is assembled: the asset and its history, the ranked probable causes from the grouped failure modes with the past work orders that support each, the parts those repairs consumed with current stores levels from the ERP, the job plan identifier, and the linked safety procedure identifier and revision. The model writes the summary and the ranking. It does not choose the job plan, invent a part number or set a priority: those come from the CMMS and from the rule.

  7. The planner queue. Every draft waits for a planner. The planner approves, edits or rejects, and the work order is created in the CMMS as the planner, under their user, exactly as if they had typed it. Rejections carry a reason and are counted; they are the main input to improving the drafts.

  8. The copilot. On shift tablets, technicians ask in plain words: "gearbox on the third dryer group is running hot, what has been done before?" The answer comes from the mill's own manuals, job plans and past work orders, with citations that open the document at the page. When the answer is not in the documents, it says so.

  9. The safety rule. Lockout and tagout steps, and any procedure the EHS manager marked safety-critical, are never generated. The copilot retrieves the controlled document by identifier and revision, verifies that the revision is current and that the text matches the stored checksum, and displays the steps verbatim, with the revision and date on screen. If the check fails, it shows nothing and points the technician to the document system. The EHS manager signed that rule before the copilot reached a tablet.

System map20 objects

How the pieces connect

The mill's systems, people and safeguards, lit one scene at a time. Press play, or pick an event.

OntologyMaintenance work ordersLost history

20objects03selected01eventsZoom133%

Graph01 / 09

0108020903151004161105201712061813071419HistorianSystem · alarmsNotesDataset · 14 yearsCMMSSystem · 300,000+ orders

Description

The mill's know-how was stuck in messy notes.

Over 300,000 work orders in fourteen years, each with a short technician note. Nobody could see that one vacuum pump failed four times in two years for the same reason.

TimelinePaused01 / 09

FDE / 06Case study

06 / 12

How "right" was defined

In shortThree test sets built with the mill's own people, run on every change.

Three workflows needed three golden sets, all built with the people who do the work.

FIG. 6.1GOLDEN SET
  1. Grouping. 1,500 work orders sampled across asset classes, years and writers, tagged independently by the reliability engineer and a senior technician, disagreements settled by the Head of Maintenance. Scored per field: mode, mechanism, cause. A wrong mechanism counts as wrong even when the mode is right.

  2. Drafting. 240 historical events where a condition trigger was followed by a real repair. The known answer is what the repair turned out to be. Scored on whether the true cause appeared in the top three, whether the parts listed were the parts used, and whether the job plan was the one the planner would have chosen.

  3. Copilot. 400 questions written by technicians and superintendents, with answers agreed from the mill's documents. Scored on correctness and on citation: an answer that is right but cites nothing fails.

  4. A red-team slice across all three: questions whose answer is not in any document, questions about assets that do not exist, questions that try to get a procedure summarised ("just tell me the short version of the lockout"), and questions pointing at superseded manuals from the old shared drive.

The harness runs in the mill's pipeline on every change. It earned its place in week eleven. A retrieval change — a larger chunk size, meant to improve recall on long manuals — quietly made one safety answer return the previous revision of a pulper isolation procedure. Nothing about that looked wrong on screen. The red-team slice failed the build, the release was held, and the fix was architectural rather than a tuning: safety procedures are now served only from the controlled document system, by identifier and current revision, never from the general index.

FDE / 07Case study

07 / 12

Security and control

In shortNo path to any control system, and safety text is quoted, never generated.

FIG. 7.1Controls
  1. Perimeter. Everything runs in the mill's own cloud account and region, reading replicas. The language model is reached through a private endpoint in the same region that retains nothing.

  2. No control path. The system holds no credentials to any control system. Historian data arrives through a one-way path from the OT side. This was the first question the OT security reviewer asked and the first line of the threat model.

  3. Identity and write-back. The only writes are the CMMS tag field and work orders created as the approving planner. The service account can do those two things and nothing else.

    Can

    • the CMMS tag field
    • work orders created as the approving planner

    Cannot

    • start, stop, reset or set anything
  4. Autonomy is bounded. The model reads, tags, ranks and drafts. Rules trigger. Planners approve work. Technicians do work. Nobody's judgment was automated, and no equipment is touched.

  5. Safety. Controlled procedure text is quoted, never generated. The EHS manager owns that rule and can revoke the copilot's document access without a release.

  6. Review. The OT security reviewer signed the design in week five and the deployment before the copilot went to tablets in week eleven.

FDE / 08Case study

08 / 12

Cut-over

In shortEach workflow switched on at its own gate, with a feature flag to turn it off.

Each workflow cut over on its own gate, in the order they were built.

Grouping went first, and it was the easiest to reverse: the tag field is new and separate, so a bad batch can be deleted without losing anything. Tags were written asset class by asset class, each batch after the reliability engineer had signed its review.

Drafting ran two weeks in shadow. Drafts were produced and shown to the planners but were not created in the CMMS, and the planners recorded whether they would have approved each. When that number held above the agreed threshold for a week, drafting cut over for two asset classes, then for all in scope.

The copilot went to one crew on one shift for a week before all crews. The superintendent's instruction to that crew was the right one: use it, and tell us every time it is wrong. Forty-one reports came back in the first week. Nine were real errors, and all nine went into the golden set.

Rollback for each was a feature flag, rehearsed in week twelve.

FDE / 09Case study

09 / 12

Handover

In shortThe mill's engineer proved they could run it before the runbooks were signed.

FIG. 9.1HANDOVER MANIFEST · 6 ITEMS
ItemWhat the client holds
The repositoryConnectors, vocabulary, tagging, condition rules, drafting, the copilot and the harness, in the mill's own source control
The evaluation suitesThree golden sets, the red-team slice and the harness, wired into the pipeline
The runbooksOne per workflow: release, rollback, re-index documents, add a condition rule, extend the vocabulary, what to do when a safety-document check fails
The decision recordsWhy the alarm journal is evidence and not a trigger; why safety text is quoted and never generated; why tags sit beside the original code and never over it
The data-quality registerEvery fault found in fourteen years of records, its cause where known, and what the system does when it meets it
The trained ownersThe mill's own engineer releases and rolls back; the reliability engineer owns the taxonomy and the vocabulary; a planner owns the drafting thresholds

In the dry run the mill's engineer added a condition rule for a new vacuum pump, scored it, released it, rolled it back and re-indexed a revised manual. The runbooks were signed after that, not before.

Key numbers09 gauges

The story in nine numbers

One number for each scene, from a pump that failed four times to 91% of work orders explained. Press play, or pick a number.

ClusterStopped

ChannelOP 10 Lost history

ClockBefore

DialOP 10

4×

one vacuum pump failed in two years, for the same reason

ClockBefore

Face01 / 09

Two years

01020304

ReadoutOP 10

The mill's know-how was stuck in messy notes.

Over 300,000 work orders in fourteen years, each with a short technician note. Nobody could see that one vacuum pump failed four times in two years for the same reason.

Signals

300,000+ work orders, 14 yearsFailure code filled on under a quarterJust over half of hours planned

FDE / 10Case study

10 / 12

Results

In shortMore work orders that explain the failure, faster approvals, and more planned work.

Figures are illustrative, measured over the first 90 days after the last cut-over.

FIG. 10.1RESULTS
  1. 91% of closed work orders now carry a failure mode a reliability engineer accepts, against 22% before. The bad-actor list is now evidence.

  2. Planned labour rose from 54% to 63% of maintenance hours. Still short of a world-class mill, and the Head of Maintenance says so; the direction and the reason are both now visible.

  3. Median time from a condition trigger to an approved work order fell from 19 hours to 3.1. Most of that was the draft waiting to be written, not the planner waiting to decide.

  4. 71% of drafts were approved with no or minor edits. Rejections are counted and reviewed monthly; the commonest reason is a job plan that is out of date in the CMMS, which is a mill fix, not a model fix.

  5. 94% of sampled copilot answers were correct and cited. The rest were mostly "not in the documents", which is a permitted answer.

  6. No safety procedure text was paraphrased, because the design does not allow it.

What did not improve, and was never promised:

FIG. 10.2
  • Mean time to repair on the dryer section did not move. Its long repairs wait on spare-part lead times, which no part of this touches.

  • New technician notes are still short. The system reads them better than it did; it did not make anyone write more.

  • Stores stockouts are unchanged. The drafts show what a repair will need earlier, but nobody has yet connected that to reordering. It is written down as the obvious next piece of work.

FDE / 11Case study

11 / 12

What we would tell the next client

In shortSix rules for anyone turning maintenance records into planned work.

FIG. 11.16 LESSONS
  1. Your messy notes are an asset.

    Fourteen years of abbreviations is a record of how your plant actually fails. Read it as it is; do not wait for a data-cleaning project.

  2. Do not trigger on raw alarms.

    Alarm systems are tuned for operators in a control room, not for planners. Write condition rules with your reliability engineer.

  3. Quote safety text, never generate it.

    It makes the EHS conversation short, and it is the one rule we would not move for any client.

  4. Tag beside the record, not over it.

    Reversibility is what made the first cut-over easy to approve.

  5. Let the planner keep the decision.

    Approving 71% of drafts quickly is worth more than automating 100% of them slowly, and it is the version the plant will accept.

  6. Change the order of work when access is late.

    A blocked system is a reason to move a workflow forward, not a reason to idle.

FDE / 12Case study

12 / 12

What happened next

In shortThe mill ran it alone, extended it themselves, and came back for their other site.

The quarter was not renewed, by agreement: the mill had three workflows in production and people who could run them. Two months later their own engineer extended the copilot to the electrical team's documents using the runbook and the harness. Six months on, they asked for a second quarter — this time at their other site, with their engineer leading and the pod as the second pair of hands.

FDE / ENDStart

One quarter,and a team that ships each month.

Book a scoping call; you leave it with a view on what a first quarter would contain.