Forward-Deployed Engineering · Logistics · Exception management
Case study: shipment exceptions, from first signal to customer update
Illustrative engagement — not a client record. The company, people, volumes and results below are a representative composite, written to show how a forward-deployed engagement runs end to end. The workflow, the architecture, the controls and the method are real and technically valid.
Offering usedEmbedded Engineering PodsA pod inside your repository and your standups, quarter by quarter.
Illustrative results, first 90 days after cut-over, in-scope clients only:
88%
Exceptions linked to the right shipment without a person searching
Before 0%
0.9 hours
Median time from first signal to customer update
Before 4.1 hours
6%
Exceptions breaching a client's notification clause
Before 23%
96%
Damage and shortage claims filed inside the contractual window
Before 71%
09Explainer
The whole story in about a minute
Nine short scenes, from scattered signals to the results. Press play, or pick a scene.
Itinerary09 stops
Exceptions to reportAbout 6,500 consignments a day
Clause breached
Exceptions a day
400 – 500
more at month-end and festival peaks
28 coordinators, two shifts
1 in 14consignments had something go wrong on the way
Stop 01Scattered signals
1 in 14
consignments had something go wrong on the way
Problems arrived as scattered signals.
About 6,500 consignments a day, and one in fourteen went wrong. Coordinators had to find each shipment before telling anyone, so nearly a quarter of exceptions breached a notification clause.
- About 6,500 consignments a day
- 400 to 500 exceptions a day
- 28 coordinators, two shifts
NextMuted dashboard
At a glance
- Client
- A third-party logistics provider running eleven warehouses and a line-haul network for around forty B2B clients, with a control tower of 28 coordinators across two shifts
- Workflow
- Shipment exceptions: detect, classify, link to the shipment, gather evidence, propose the next action from the client's SLA playbook, draft the customer update, prepare damage and shortage claims
- Engagement
- Embedded Engineering Pods (one quarter, three engineers), then Managed Operations (monthly, after handover)
- Team
- A named lead engineer, a platform engineer and a data engineer. Client side: the Head of Control Tower (owner), a TMS product analyst, two platform engineers, the claims manager, a security reviewer
- Where it runs
- The client's own cloud account and region. No data retained outside it
- Handover
- Runbooks executed by the client's engineers in the last two weeks of the quarter; the pod's standing access revoked at the gate, then scoped operations access granted for Managed Operations
| Measure | Before | After |
|---|---|---|
| Exceptions linked to the right shipment without a person searching | 0% | 88% |
| Median time from first signal to customer update | 4.1 hours | 0.9 hours |
| Exceptions breaching a client's notification clause | 23% | 6% |
| Damage and shortage claims filed inside the contractual window | 71% | 96% |
| Claims filed or liability admitted without a person | not applicable | 0, by design |
FDE / 01Case study
01 / 12
In shortExceptions arrived as scattered signals, and coordinators had to find each shipment by hand before telling anyone.
The company moved around 6,500 consignments a day: full and part truckloads between its own warehouses, client plants and distributors, using its own fleet and about 120 contracted carriers. Roughly one consignment in fourteen had something go wrong on the way. That meant 400 to 500 exceptions a day, more at month-end and during festival peaks.
Exceptions did not arrive as exceptions. They arrived as signals, scattered across five places:
Carrier status events, from the larger carriers' feeds and from a GPS aggregator, many of them late, duplicated or missing.
Driver messages, on a messaging app, often a photograph and three words.
Customer emails, to a shared control-tower mailbox, asking where a load was or reporting what had arrived.
Warehouse scans, from the WMS: short receipts, damaged-on-arrival flags, cartons scanned at the wrong dock.
Documents, chiefly proofs of delivery with missing seals or signatures, and e-way bills close to expiry on delayed loads.
Coordinators read all of it, searched the TMS for the shipment, looked up what the client's contract said to do, and wrote to the customer. The cost showed up in three places:
SLA penalties. Most client contracts set a notification clause, such as a delay over four hours reported within two, and service credits below an on-time threshold. Nearly a quarter of exceptions breached a notification clause, not because nobody knew, but because the person who knew had not yet found the shipment.
Lost claims. Claims against carriers need evidence and a written notice inside a fixed period. The evidence sat in driver chats and scan history, and almost three claims in ten were filed late or not at all.
People. The best coordinators were the ones who remembered each client's playbook. Night shift had fewer of them.
FDE / 02Case study
02 / 12
In shortA dashboard and email parser alerted on everything, linked nothing, and was soon muted.
Two years earlier the company had bought a control-tower dashboard and added an email parser built on keyword rules. It did not fail outright. It faded.
It alerted on everything. Every late carrier event raised an alert, including the many that were simply late to arrive. Within a month coordinators had muted it.
It read, but did not link. The parser extracted docket numbers when they were typed cleanly. Customers rarely typed them cleanly. The search for the shipment was still a person's job.
It knew nothing about contracts. The dashboard showed a delay. It did not know that one client wanted a phone call for a delay and another wanted nothing until it passed six hours.
Nobody measured it. There was no set of real exceptions with known answers, so no one could say whether a rule change made things better or worse.
The pieces that were missing were the ones forward-deployed engineering exists for: the link to the system of record, the client's own rules written down, a way to score the result, and an owner who would stand behind it.
FDE / 03Case study
03 / 12
In shortA week on site, then two signed goals: triage for the twelve largest clients, and claims packs.
The work was larger than one workflow on one system. It touched the TMS, the WMS, the mailbox, the driver messaging channel, the carrier feeds and the claims register. The Head of Control Tower chose a pod for one quarter rather than a single sprint, with Managed Operations after it, because her platform team was small and already carried the TMS.
The pod spent its first week on site: two night shifts in the control tower, a day in a cross-dock, a day with the claims manager, and the rest reading systems and a month of exception history.
- two night shifts in the control tower
- a day in a cross-dock
- a day with the claims manager
- the rest reading systems and a month of exception history
What was found:
FIG. 3.2- The TMS exposed an approved API for reading shipments and writing notes and status codes. Write-back could go through the path coordinators already used, with each action attributed.
- About 40% of "delay" alerts from carrier feeds were events that arrived after the next event had already happened. Many delays were really data latency.
- Each client's SLA playbook existed as a PDF annex to the contract, and in practice as the memory of three senior coordinators. The written annexes and the practice disagreed on about one clause in five.
- Claims had a clock nobody watched. The contractual window with carriers was shorter than the statutory notice period, and it was the contractual one that mattered.
The quarter's goals, signed by the Head of Control Tower before week one:
- Goal 1Exception triage for the twelve largest clients: every signal classified, linked to its shipment, evidence gathered, next action proposed from the client's playbook, customer update drafted
- Goal 2Damage and shortage claims packs: evidence assembled, carrier and window identified, claim drafted for the claims team
- Out of scopeVoice notes, customs and cross-border shipments, route re-planning, carrier rate disputes, clients outside the twelve
- MetricsShare of exceptions linked to the correct shipment at or above 98% precision on the golden set; share breaching a notification clause; share of claims filed inside the window
- OwnerHead of Control Tower; the claims manager for Goal 2
- GuardrailsNo customer message is sent without a coordinator until an approved low-risk template has passed the shadow period. No claim is filed, and no liability is admitted or compensation offered, by the system
- EndingHandover in the last two weeks of the quarter, then Managed Operations, with the client's engineers second on the rota from the first month
Before any code was written, the senior coordinators and the account managers rewrote the twelve playbooks as tables: exception type, threshold, who is told, how soon, through which channel, with what wording. That exercise surfaced the disagreements between contract and practice, and each one went to the client account owner to settle.
09Timeline
The engagement, stage by stage
Nine stages on one clock, from before the quarter to ninety days after cut-over. Press play, or pick a stage.
Delivered
Event 09 / 09
Results
Origin
Before
Scattered signals
Destination
+90 days
Results
Event log09 / 09 events
- 01BeforeScattered signals1 in 14
- 02Two years earlierMuted dashboardMuted
- 03W1On site first~40%
- 04W2–W6One signalRules
- 05W2–W10A person0
- 06W2–W9Scored first2,400
- 07W7–W11Cut-over3
- 08W12–W13Handover13th
- 09+90 daysResults0.9 h
Scan · 09+90 days
Faster updates, claims filed on time.
Illustrative, first 90 days: the median time to a customer update fell from 4.1 to 0.9 hours, and claims filed inside the window rose from 71% to 96%.
0.9 hmedian time to a customer update, down from 4.1 hours
Claims filed inside the window
71%96%
- 88% linked with no search
- Clause breaches 23% → 6%
- Claims in window 71% → 96%
FDE / 04Case study
04 / 12
In shortThirteen weeks: connect the signals, link and triage, cut over by client, add claims, hand over.
W1M1
What happened
Pod on site and in the control tower's standups; access requested for TMS API, WMS read replica, mailbox, messaging export and carrier feeds. Quarter goals signed
What existed at the end
First pull requests merged by all three engineers: mailbox and scan-event intake landing in the client's storage
What happened
Pod on site and in the control tower's standups; access requested for TMS API, WMS read replica, mailbox, messaging export and carrier feeds. Quarter goals signed
What existed at the end
First pull requests merged by all three engineers: mailbox and scan-event intake landing in the client's storage
What happened
Linking prototype running on real signals behind a feature flag. Golden set drafted with senior coordinators. The docket-number assumption failed (see below)
What existed at the end
A working linker on a composite key, and a golden set of 2,400 historical exceptions with known answers
What happened
Classifier, evidence gatherer and playbook engine built. The carrier aggregator's feed arrived eight working days late; recorded in the weekly report the day it slipped, and the owner moved carrier-feed classification to W5
What existed at the end
Triage producing proposed actions in parallel with coordinators, without acting
What happened
Carrier feeds joined; draft-update generation; coordinator review screen inside the TMS; audit log; dashboards. Shadow run began in W6
What existed at the end
A system triaging every in-scope exception alongside the team
What happened
Goal 1 cut over by client: two clients, then six, then all twelve, each step gated on shadow-run agreement
What existed at the end
Coordinators working from proposed actions and drafted updates, rollback one switch away
What happened
Goal 2: claims evidence packs, window tracking, draft claim letters; golden set extended with 300 historical claims. Harness caught a regression in W9
What existed at the end
Claims packs in the claims team's queue
What happened
Shadow period for auto-send ended; three low-risk templates approved by the owner for auto-send on four clients who opted in
What existed at the end
Auto-send live for approved templates only, every message logged
What happened
Handover & Enablement: runbooks, dry run by the client's engineers, rota set up for Managed Operations
What existed at the end
Signed runbooks, decision record, quarter report, the pod's access revoked
The false assumption. At scoping, everyone said the docket number was unique. In week two the golden set showed it was not: two line-haul partners reused the same number series, and a third restarted it every financial year. On real data, 3% of links were confidently wrong. The linker was changed to a composite key of docket number, e-way bill number, vehicle number and a booking-date window, and the confident-wrong rate on the golden set fell to under 0.2%. Had the pod trusted the interview rather than the data, the system would have written updates to the wrong customers.
FDE / 05Case study
05 / 12
In shortA model reads and drafts; rules link shipments and apply each client's playbook; people send and file.
05Try a signal
Pick a signal. Watch where the exception goes.
The same triage handles every signal. What arrives decides whether it ends as a customer update or with a person.
Triage log09 / 09
- 01SignalLands in the shared control-tower mailbox
- 02IntakeStored with its source, time received and time of the event
- 03ClassifyThe model fills a strict schema: type, identifiers, and where each came from
- 04LinkRules match one shipment in the TMS on the composite key
- 05EvidenceScan history, GPS trail, booked and revised ETA pulled in
- 06PlaybookThe client's table says who is told, by when, and how; the countdown starts
- 07Draft updateWritten from the client's approved wording, filled with facts from the evidence
- 08CheckedEvery number, date and identifier matches the evidence
- 09Update sentA coordinator reads, edits if needed and sends
Destination01
Update sent
Update sent by a coordinator
No one searched the TMS for the shipment. The coordinator only read the draft and sent it.
Exception status
- Claims filed by the system
- 0
- Runs in
- The client's cloud region
Intake and de-duplication. Every signal lands in the client's storage with its source, time received and time of the event it describes. A carrier event older than the latest known event for that vehicle is kept for the record but does not raise an exception, which removed most of the false delays found at scoping.
Classify. A language model, reached through a private endpoint in the client's region with no data retention, reads an email, a driver message or a document and fills a strict schema: exception type (delay, damage, short shipment, excess, wrong address, refused delivery, POD problem, e-way bill risk, other), severity signals, identifiers mentioned, and the sentence each value came from. Photographs from drivers are read for damage cues and for visible identifiers such as a docket label. The model's output is a classification with evidence, not a decision.
Link. Deterministic matching against the TMS on the composite key. The model's extracted identifiers are candidates; the rules decide. A signal with one match above the threshold is linked. A signal with none, or with two, goes to an unlinked queue where a coordinator finds the shipment and the choice is recorded for the golden set.
Gather evidence. For a linked exception, the system pulls the scan history from the WMS, the GPS trail, the booked and revised ETA, the POD image and its fields, the e-way bill's validity, and the related messages, each with its source and time. A reviewer sees the evidence beside the proposal, not a summary in place of it.
Playbook. A rules engine, not a model, reads the client's playbook table and returns the next action: whom to tell, by when, through which channel, whether a phone call is required, whether an internal escalation opens. It also computes the notification deadline and shows a countdown on the queue. Two checks run on every delayed load:
whether the e-way bill will expire before the revised ETA, in which case the queue shows an extension task for a person to raise on the government portal, which the system does not touch;
whether the POD has the fields the client needs to accept an invoice: seal number, receiver name, signature or stamp, and quantity received.
Draft update. The model drafts the customer message from the client's approved wording for that exception type, filled with facts taken from the evidence, not from the model's own reading. A deterministic check compares every number, date and identifier in the draft with the evidence record, and blocks the draft if any differ. Drafts never contain an admission of liability, an offer of compensation or a promised delivery time that the TMS does not already hold.
Send. By default a coordinator reads, edits if needed and sends. After the shadow period the owner approved three templates for auto-send, for four clients who asked for it in writing: a revised ETA under four hours with no damage, a delivery completed with a clean POD, and a pickup rescheduled inside the same day. Anything else waits for a person.
Claims pack. Where damage or shortage is documented by a scan, a POD remark or a photograph, the system assembles a pack: the evidence, the carrier responsible for that leg, the contractual claim window and the date it closes, the statutory notice date as a backstop, and a draft notice. The claims team decides whether to file, and files it themselves.
Where the decisions sit. A model reads, classifies and drafts. Rules link shipments, apply playbooks and check drafts against evidence. A named coordinator sends customer messages outside the approved templates, and a named person on the claims team files every claim.
Audit and monitoring. Every signal, link, rule result, draft, edit and send is written to the client's log store with the model version. Dashboards show volume by type, linking rate, notification deadlines at risk, edits made to drafts and claims windows closing. Thresholds page the platform engineer when the linking rate drops or the unlinked queue grows past its limit.
20System map
How the pieces connect
Every system, queue and team in the story, lit one scene at a time. Press play, or pick an event.
Hub · 01 / 09Before
Control tower
Site28 coordinators
Problems arrived as scattered signals.
About 6,500 consignments a day, and one in fourteen went wrong. Coordinators had to find each shipment before telling anyone, so nearly a quarter of exceptions breached a notification clause.
1 in 14consignments had something go wrong on the way
Linked objects · 2
- SIGSignalsfive placesSource
- TMSTMSapproved APISystem
FDE / 06Case study
06 / 12
In short2,400 past exceptions with known answers, which every change is scored against.
The golden set was built with the senior coordinators, who were the only people who knew what the right answer had been.
2,400 historical exceptions across the twelve clients, every exception type, both shifts, festival peaks and quiet weeks, emails, driver messages and scans.
Answers taken from what happened: the shipment the exception was finally attached to in the TMS, the action the playbook required, and whether the customer was told in time.
A red-team slice: reused docket numbers, a customer quoting another customer's load, a driver photograph of the wrong truck, a POD image for a different consignment, and emails asking the system to promise a refund.
Separate scores for classification, linking precision, playbook action and draft fidelity. Linking precision carried the highest threshold, because a wrong link sends the right message to the wrong customer.
300 historical claims added for Goal 2, scored on evidence completeness and window dates.
The suite runs in the client's CI. In week nine it earned its place: a change that improved reading of one client's email format lowered the short-shipment versus damage distinction on WMS remarks from another client. The build failed. Nobody on the floor would have noticed for a week, and by then claims would have been drafted under the wrong heading.
Draft quality was scored by coordinators, not by the pod. In the shadow run they marked each draft "sent as is", "edited" or "rewritten". The auto-send templates were approved only for exception types where "sent as is" held above the owner's threshold for two consecutive weeks.
FDE / 07Case study
07 / 12
In shortIt runs in the client's own cloud, can only write notes and status codes, and never files a claim.
Perimeter. Everything runs in the client's cloud account and region. Egress is limited to the private model endpoint, which retains nothing, the carrier feeds already in use and the TMS API.
Personal data. Driver phone numbers, receiver names and POD signatures stay in the client's storage, are masked in dashboards, and follow the client's existing retention schedule.
Identity. The service account can read shipments, write notes and status codes, and nothing else. It cannot change a rate, a consignee address or a delivery instruction. The pod's access went through the client's identity provider and appeared in their audit log.
Can
- read shipments
- write notes and status codes
Cannot
- change a rate
- a consignee address
- a delivery instruction
Autonomy is bounded. The model reads and drafts; rules decide what is linked and what the playbook requires. Customer messages outside the three approved templates are sent by a coordinator. Claims are filed by a person. An address change requested by email is never applied; it becomes a task to confirm with the consignee on a known number, because redirecting a load is a known fraud pattern.
The owner holds the switches. Auto-send can be turned off per client, per template or entirely, in configuration, without a release.
Review. The security team reviewed the design in week three and the deployment in week six, and again before auto-send in week eleven. Their findings and how each was closed are in the decision record.
FDE / 08Case study
08 / 12
In shortSwitched on client by client after a shadow run, with one flag to switch it back.
The shadow run was the gate for triage, and a second, longer shadow period was the gate for auto-send.
For a week in month two the system triaged every in-scope exception while the coordinators worked as before. Disagreements were sorted into three piles: the system was wrong, the coordinator was wrong, and the playbook was unclear. The third pile went back to account owners as decisions, and eleven playbook rows changed.
For a week in month two
the system was wrong
the coordinator was wrong
the playbook was unclear
went back to account owners as decisions
Rollout went by client: two clients chosen for high volume and simple playbooks, then six, then all twelve. Each step needed two clean days on linking precision and no missed notification deadline caused by the system. Rollback was one flag that returned every signal to the old mailbox and dashboard. It was tested in week seven and used once, for forty minutes, when the TMS API was down for maintenance; the queue then drained on its own.
01two clients
02then six
03then all twelve
EACH STEP · two clean days on linking precision and no missed notification deadline caused by the system
the TMS API was down for maintenance
- for forty minutes
used once
the queue then drained on its own
CONTROL ·one flag that returned every signal to the old mailbox and dashboard
Auto-send came four weeks later. For those four weeks every draft that would have been auto-sent was still sent by a coordinator, and the edits were counted.
FDE / 09Case study
09 / 12
In shortThe client's engineers added a client and rolled back a release before the pod left.
The quarter was not renewed as a pod. Its last two weeks were Handover & Enablement, and Managed Operations began after the gate.
In the dry run, a client platform engineer added a thirteenth client's playbook, scored it against a new slice of the golden set, released it, withdrew an auto-send template and rolled the release back, with the pod in the room and off the keyboard. The runbooks were signed after that.
09Key numbers
The story in nine numbers
One number for each scene, from one consignment in fourteen going wrong to 96% of claims filed in time.
Board · 09 readings
+90 days
Platform 09Results
Gate · 09 / 09+90 days
median time to a customer update, down from 4.1 hours
Faster updates, claims filed on time.
Illustrative, first 90 days: the median time to a customer update fell from 4.1 to 0.9 hours, and claims filed inside the window rose from 71% to 96%.
- 88% linked with no search
- Clause breaches 23% → 6%
- Claims in window 71% → 96%
09 / 09
FDE / 10Case study
10 / 12
In shortFaster customer updates, fewer missed notification clauses, and claims filed on time.
Figures are illustrative, measured on the twelve in-scope clients over the first 90 days after cut-over.
88% of exceptions linked to the right shipment with no search. The rest reached the unlinked queue with the candidates shown.
BEFORE0%AFTER88%Median time to customer update fell from 4.1 hours to 0.9. Most of the gain was the end of searching, not faster typing.
Notification-clause breaches fell from 23% to 6%. The countdown on the queue did as much as the drafting.
BEFORE23%AFTER6%96% of damage and shortage claims filed inside the contractual window, up from 71%, because the pack was ready before anyone asked for it.
BEFORE71%AFTER96%About a third of customer updates went out on approved templates without a coordinator, all in the three low-risk types.
No claim filed and no liability admitted by the system, because the design does not allow it.
The control tower was not reduced. Two senior coordinators moved to owning client playbooks and carrier escalations.
What did not improve, and was never promised:
FIG. 10.2Carrier data quality. Roughly one small carrier in six still sends status events late or not at all. The system stops those gaps from raising false alarms. It cannot make the events arrive.
On-time delivery. Exception handling makes the customer better informed and the claims better evidenced. It does not make a truck faster.
Voice notes in regional languages from drivers were out of scope and are still heard by a person.
FDE / 11Case study
11 / 12
In shortSix rules for anyone automating shipment exceptions.
- Write the playbooks as tables first.
A fifth of the contract annexes disagreed with practice. No model could have resolved that.
- Check your keys against the data, not the interview.
A number everyone calls unique may not be.
- Late data looks like a late truck.
Order events by when they happened before you alert on them.
- Earn auto-send one template at a time.
Count coordinator edits in a shadow period and approve only what is sent unchanged.
- Keep the claims clock in the system.
Evidence decays and windows close whether anyone is watching or not.
- Let people file and promise.
Claims, compensation and redirected loads stay with named people; it kept the security review and the client conversations short.
FDE / 12Case study
12 / 12
In shortThe client took triage in-house and asked for a quarter on voice notes and cross-border exceptions.
The client kept Managed Operations for five months. In that time the golden set was extended twice, for a new client and for festival-peak phrasing, and one model upgrade was held back because it lowered linking precision on the regression run. From the fourth month the client's platform engineer was first on the rota for triage. They have since taken triage in-house, kept operations on claims for another quarter, and asked for a pod quarter on voice notes and cross-border exceptions, both out of scope the first time.

