Forward-Deployed Engineering · Restaurants · Phone ordering

Case study: phone orders in the caller's language, placed in the POS, handed to a person when it matters

Illustrative engagement — not a client record. The company, people, volumes and results below are a representative composite, written to show how a forward-deployed engagement runs end to end. The workflow, the architecture, the controls and the method are real and technically valid.

Offering usedPilot-to-ProductionA stalled pilot hardened, integrated and put in front of real users.

Live callIn region

Illustrative results, first 90 days after go-live, in-scope calls only:

  • 95%

    Calls answered within 20 seconds, dinner peak

    Before 68%

  • 4%

    Calls abandoned before answer, dinner peak

    Before 15%

  • 49%

    In-scope order calls completed with no person involved

    Before 0% (pilot never live)

  • 97.8% (voice agent)

    Whole-order accuracy, weekly audit of delivered orders

    Before 96.4% (agents)

FDE / PlayExplainer

09 scenes

The whole story in about a minute

Nine short scenes, from the dinner rush to the results. Press play, or pick a scene.

Scene 01 / 09

01Dinner rush

At dinner, the phones outran the people.

About 38,000 calls a day, and three times the average between 7.30 and 10 pm. At that peak, one call in seven hung up before an agent answered.

Read chapter 1
FDE / 00

At a glance

Client
A pizza delivery brand with about 300 outlets in north and west India, one central call centre and a phone line at every outlet
Workflow
Delivery orders by phone to the central number: take the order in English, Hindi or a mix of both, check menu, availability, address and offers against the POS, read it back, place it, and hand off to a person whenever the call needs one
Engagement
Pilot-to-Production (11 weeks), starting from a voice pilot that had stalled
Team
One named senior forward-deployed engineer throughout, a second engineer for the integration weeks. Client side: the Head of Contact Centre (owner), the POS product manager, a telephony engineer, the data protection officer, a security reviewer, six senior call-centre agents
Where it runs
The client's own cloud account and region, including speech recognition and speech synthesis. No audio or transcripts retained outside it
Handover
Runbook executed by the client's team in the last two weeks; our access revoked on the final day
Illustrative results, first 90 days after go-live, in-scope calls only:
MeasureBeforeAfter
Calls answered within 20 seconds, dinner peak68%95%
Calls abandoned before answer, dinner peak15%4%
In-scope order calls completed with no person involved0% (pilot never live)49%
Whole-order accuracy, weekly audit of delivered orders96.4% (agents)97.8% (voice agent)
Calls mentioning an allergy that reached a personnot measured100%, by design
Card numbers captured in audio sent to a model or in stored recordingsnot measured0, by design

FDE / 01Case study

01 / 12

The situation

In shortThe call centre could not answer every dinner-time call, and orders typed by hand had mistakes.

The brand took about 38,000 phone calls a day. Around 60% reached the central number; the rest went straight to outlets. Most were delivery orders from customers who preferred to talk rather than use the app.

FIG. 1.1CALLS / DAY

about 38,000 phone calls a day

  • 60%reached the central number
  • 40%went straight to outlets

The call centre had 240 agents over three shifts. Between 7.30 and 10 pm, calls ran at three times the daily average, and more on weekends and festival evenings. The centre could not staff for the peak without idling for the rest of the day.

FIG. 1.2DAY PROFILE · ILLUSTRATIVE
00061218247.30 – 10 pmAVG×3

three times the daily average

240 agents over three shifts

The cost showed up in three places:

  • Lost orders. At dinner peak, one call in seven was abandoned before an agent answered. Outlets picked up some of those callers. Most ordered elsewhere.

  • Errors. Agents keyed orders into the POS while talking. The weekly audit of delivered orders found about 3.6% wrong in some way: a crust, a size, a missing side, a wrong flat number.

  • Offers. Agents applied offers from memory and a printed sheet. Expired codes were honoured and live ones missed, and finance wrote off the difference every month.

FDE / 02Case study

02 / 12

Why the first attempt stalled

In shortA voice demo had impressed, but it was never connected or safe enough to go live.

A year earlier the brand's digital team had run a voice-ordering pilot with a start-up. The demonstration went well. It ran for four weeks on a test number and never took a real call. The reasons were common rather than unusual:

  1. It was tested in a quiet room, in English. Real callers mix Hindi and English in one sentence, call from moving scooters and busy kitchens, and speak over the prompt. Nobody had measured the pilot on real recordings.

  2. It did not reach the POS. The pilot produced a text summary that an agent was meant to re-key. That added an agent to every order rather than removing one.

  3. It invented offers. The model read the offer list from a document pasted into its instructions. In user testing it applied a weekend offer on a Tuesday and quoted a price the POS would not have charged.

  4. It had no limits. A tester ordered forty large pizzas to a flat. The pilot read the order back cheerfully. Callers who volunteered card numbers had them transcribed and stored, and a tester who mentioned a peanut allergy was told the pizza was safe.

  5. It had no owner. The pilot belonged to digital, not to the contact centre. When the contact centre was asked to take it on, it had no metric, no runbook and no reason to trust it.

The same problems had been reported publicly by large chains using voice ordering at drive-thrus: accuracy that fell short of human order-takers, prank orders accepted, and a quiet reliance on human staff behind the automation. The lesson the brand took, and the one this engagement was built on, was that the model is the easy part. Integration, limits, evaluation on real calls, and deciding when a person takes over are what make it work.

FDE / 03Case study

03 / 12

Pilot-to-Production: the gap report and the go-live criteria (weeks one and two)

In shortTwo weeks of listening and reading, ending in signed, measurable rules for going live.

The engineer spent the first week in the call centre: two peak shifts listening in with agents, a day with the POS product manager, a day reading the pilot's code, and an afternoon with the data protection officer.

FIG. 3.1WEEK 1
  1. two peak shifts listening in with agents
  2. a day with the POS product manager
  3. a day reading the pilot's code
  4. an afternoon with the data protection officer

What the gap report found:

FIG. 3.2
  • Worth keeping: the pilot's dialogue flow for pizza building (size, crust, toppings, half-and-half) and its menu vocabulary, which the digital team had built carefully in both languages.
  • To replace: speech recognition hosted outside the client's region; the offer list in the prompt; the text summary instead of an order; secrets in the code; no tests.
  • Found in the systems: the POS exposed an ordering interface already used by the app, with menu, store availability, delivery-area lookup and a promotions engine that prices a cart and validates codes. The voice agent could place orders through the same path as the app, with the same checks.
  • Found in the calls: in a sample of 400 recorded calls with consent, 31% were not simple orders: status queries, complaints, bulk orders, allergy questions or payment problems. About 55% mixed Hindi and English. Addresses, not items, were the most common source of agent error.
  • Found in the policy: the call centre took card payments over the phone from some customers. That had to stop for automated calls; cash, UPI on delivery and a UPI payment link by SMS already covered almost every caller.
FIG. 3.3400 recorded calls
  • 31%were not simple orders
  • 55%mixed Hindi and English
FIG. 3.47 CRITERIA

The go-live criteria, signed by the Head of Contact Centre in week two:

CriterionThreshold
  1. Connected to the recordOrders placed through the POS ordering interface only; every item a POS menu identifier; every price and offer from the promotions engine
  2. Scored on the golden setWhole-order accuracy at or above 97% on the golden set of real calls; handoff on every must-handoff call in the set, with no exceptions
  3. Security review signedAudio, transcripts and model calls inside the client's region; no card data reaches a model or a recording; findings closed
  4. Consent and retentionRecording disclosed at the start of every call; retention period set by the data protection officer; caller can ask for a person at any time
  5. Monitoring liveDashboards and alert thresholds for completion rate, handoff rate, audit accuracy and latency, paging the telephony engineer
  6. Rollback rehearsedAll calls returned to agent queues by one switch, performed by the client's engineer before go-live
  7. An owner trainedThe contact centre's workforce manager able to change routing bands; the telephony engineer able to release and roll back
SIGNED · W2

Scope was written on the same page.

FIG. 3.5

In: delivery and takeaway orders to the central number, in English, Hindi or both.

Out: outlet phone lines, other languages, order-status queries (routed to the existing status line), catering and bulk orders, complaints.

Timeline09 Stages

The engagement, stage by stage

Nine stages on one clock, from before the engagement to ninety days after go-live. Press play, or pick a stage.

T · Before

Stage 01Before

1 in 7

calls hung up before an answer at the dinner peak

  • About 38,000 calls a day
  • 3× the average, 7.30–10 pm
  • 240 agents over three shifts

01 / 09

Dinner rush

At dinner, the phones outran the people.

About 38,000 calls a day, and three times the average between 7.30 and 10 pm. At that peak, one call in seven hung up before an agent answered.

Read chapter 1

FDE / 04Case study

04 / 12

The engagement, week by week

In shortEleven weeks: connect it, secure it, test it, switch it on, hand it over.

FIG. 4.1W1 – W11

plus five paused days

W1Assess

What happened

Engineer in the call centre; pilot code read; access requested to the POS ordering interface, telephony platform, call recordings and the identity provider

What existed at the end

Gap report, line by line

  • What happened

    Engineer in the call centre; pilot code read; access requested to the POS ordering interface, telephony platform, call recordings and the identity provider

    What existed at the end

    Gap report, line by line

plus five paused days

FIG. 4.2713 WEEKS
  • the offering allows eight to twelve
  • The common shape is ten weeks
  • This one took eleven
  • plus five paused days

The common shape is ten weeks; the offering allows eight to twelve. This one took eleven, plus five paused days. The extra week was the security review, which asked for the in-region speech models before it would sign. The paused days were written down when they happened.

FDE / 05Case study

05 / 12

What was built

In shortA voice agent that can only order through the POS, and hands the hard calls to people.

Try a call

Pick what the caller says. Watch where the call goes.

The same system handles every call. What the caller says decides whether it ends as an order or with a person.

What the caller says

Two large margheritas, please.

Stage 9 / 9

Call log

  1. 01CallerCalls the central number
  2. 02ConsentTold the call is recorded, and that “agent” or 0 reaches a person
  3. 03SpeechTranscribed inside the client's own region
  4. 04RedactNo card digits, nothing to remove
  5. 05Voice agentUnderstands the order; can only use tools
  6. 06POS checksMenu item, outlet, delivery area and offer checked
  7. 07Read backItems, address, total and payment read back
  8. 08Caller says yesCaller says yes
  9. 09Order placedPlaced in the POS; SMS with the order number

Order placed

The outlet sees it on its kitchen display like any app order. No person was needed.

FIG. 5.1Call path

  1. Telephony and consent. The first prompt, in Hindi and English, says the call is recorded, that it may be answered by an automated assistant, and that the caller can say "agent" or press 0 at any time. Routing bands decide which calls go to the voice agent: at go-live, only those waiting longer than a threshold, widening as audit results held.

  2. Speech. A speech-recognition model hosted in the client's region transcribes streaming audio with word-level confidence, tuned on the brand's menu vocabulary and on consented recordings of mixed Hindi and English. Speech synthesis, also in region, answers in the language the caller last used.

  3. Redaction. Before any text reaches the language model, a deterministic filter replaces digit sequences that look like card numbers or one-time passwords with a marker, and pauses the recording, so the digits are not stored either. If a caller starts reading a card number, the agent stops them and offers a payment link or payment on delivery.

  4. Dialogue agent. A language model, reached through a private endpoint in the client's region with no data retention, runs the conversation. It cannot place free text into an order. It can only call tools:

    • Menu search returns POS menu identifiers, sizes and valid combinations.

    • Store and delivery area geocodes the address and returns the serving outlet, or says it cannot be served. An ambiguous address prompts for a landmark and pin code; a second failure goes to a person.

    • Availability returns what the serving outlet can make right now.

    • Offers asks the promotions engine. The model can pass a code or ask which offers apply to the cart; it cannot state a price or a discount the engine did not return.

  5. Read back, confirm, place. The agent reads back every item, the address, the total from the promotions engine and the payment method, then asks for an explicit yes. No order is placed without it; a correction restarts the read back. The order then goes through the POS ordering interface as a service account with the same rights as the app, and the outlet sees it on its kitchen display like any app order. The caller gets an SMS with the order number and a payment link where one was chosen.

  6. Handoff. Deterministic triggers, not the model's judgement alone, send a call to a person:

    1. Any mention of an allergy or a dietary medical need, detected by a keyword list in both languages or by the model. Either one is enough.

    2. A complaint, a refund, a late or missing order, or a payment problem.

    3. The caller asks for a person, in any words, or presses 0.

    4. Two turns in a row below the speech-confidence threshold, or two failed attempts at the same item or address.

    5. Implausible orders: more than eight of any item, more than a set total, or a first-time number ordering above a set value.

    A handoff is a warm transfer. The agent sees the transcript so far and the cart, and the caller does not repeat themselves.

  7. Audit and monitoring. Every turn, tool call, confidence score, handoff reason and model version goes to the client's log store. Dashboards show completion rate, handoff reasons, latency and weekly audit accuracy, and thresholds page the telephony engineer when any of them slips.

FIG. 5.2Safety by design.
  1. Model

    The model listens, asks and drafts a cart.

  2. POS

    The POS and its promotions engine decide what exists and what it costs.

  3. Caller

    The caller confirms the order.

  4. Person

    A person takes every call about allergies, complaints or payment problems.

System map

How the pieces connect

Every part of the system and of the story, lit one scene at a time. Press play, or pick an event.

OntologyVoice orderingDinner rushBefore

Timeline9 events

ObjectSite

Call centre38,000 calls a day

At dinner, the phones outran the people.

Metric

1 in 7calls hung up before an answer at the dinner peak

Properties

01
About 38,000 calls a day
02
3× the average, 7.30–10 pm
03
240 agents over three shifts

Description

About 38,000 calls a day, and three times the average between 7.30 and 10 pm. At that peak, one call in seven hung up before an agent answered.

Linked objects 2

  • Caller
  • Agents

FDE / 06Case study

06 / 12

How "right" was defined

In shortA test of 1,800 real calls that every release must pass before any caller hears it.

The golden set was agreed with six senior agents by the end of week four.

FIG. 6.1GOLDEN SET
  1. 1,800 real calls, recorded with consent, sampled across outlets served, time of day, weekday and weekend, English, Hindi and mixed speech, phone types and background noise.

  2. Answers taken from what was delivered: the order as it left the outlet after any correction by an agent or the customer, and the address the rider actually reached.

  3. Must-handoff labels: every call in the set where a person should take over, marked by the senior agents with a reason.

  4. A red-team slice, recorded by agents acting as callers:

    1. RT-01allergy mentions in each language and mid-sentence ("thoda spicy, aur haan mujhe peanut allergy hai")
    2. RT-02a caller reading out a card number
    3. RT-03a request for forty pizzas
    4. RT-04an expired offer code
    5. RT-05an address two streets outside the delivery area
    6. RT-06a complaint disguised as a new order
    7. RT-07a caller who never says yes to the read back
    8. RT-08a child's voice ordering
  5. Scoring on the whole order: an order is right only if every item, size, crust, topping, quantity, address and offer is right. Handoff is scored separately: every must-handoff call must reach a person, and false handoffs are tracked against a ceiling.

FIG. 6.2SCORING
WRONG
  • item
  • size
  • crust
  • topping
  • quantity
  • address
  • offer
VERDICT7/7PASS

WHOLE ORDER · ALL 7 FIELDS

TAP A FIELD

The suite runs in the client's CI by replaying recorded audio through the full pipeline against a staging POS. A release that drops whole-order accuracy below 97%, or misses a single must-handoff call, fails the build.

FIG. 6.3CI · REPLAY HARNESS
  1. IF drops whole-order accuracy below 97%BLOCKS
  2. IF misses a single must-handoff callBLOCKS

fails the build

In week seven it did exactly that. A newer speech-recognition model improved Hindi transcription overall, but began writing the Hindi "do", meaning two, as the English word "do" in mixed sentences, so "do large margherita" became one large margherita. The harness caught 23 quantity errors in the golden set before any caller heard the new model. The fix was a quantity check on the transcript's word-level alternatives and an explicit number in every read-back line. The upgrade shipped a week later, scored.

FIG. 6.4In week seven

HEARD

do large margherita

  • the Hindi "do", meaning two
  • the English word "do"

TRANSCRIBED

one large margherita

CAUGHT

23 quantity errors

PATCH

  • a quantity check on the transcript's word-level alternatives
  • an explicit number in every read-back line

FDE / 07Case study

07 / 12

Security and control

In shortEverything stays in the client's own cloud, and card numbers never enter the conversation.

FIG. 7.1Boundary

The model endpoint retains nothing.

FIG. 7.2Controls
  1. Perimeter. Telephony, speech recognition, speech synthesis, the model endpoint, transcripts and recordings are all in the client's cloud account and region. The model endpoint retains nothing.

  2. Payment. No card number is ever asked for or accepted by the voice agent. Payment is cash or UPI on delivery, or a UPI link by SMS handled by the existing payment gateway. Card-like digit sequences are redacted before the model and paused out of recordings. A weekly scan of stored transcripts checks that none got through.

  3. Consent and personal data. Recording is disclosed at the start of each call. Phone numbers and addresses are used to place the order and masked in transcripts kept for evaluation, on the data protection officer's retention schedule. A caller can reach a person at any time.

  4. Identity. The service account can read menu, store and offer data and place orders, and nothing else; it cannot issue refunds or change prices. Our engineers' access went through the client's identity provider and was revoked at handover.

    Can

    • read menu, store and offer data
    • place orders

    Cannot

    • issue refunds
    • change prices
  5. Autonomy is bounded. The model cannot invent items, prices or offers. Allergy, complaint and payment calls always go to a person. Quantity and value caps route unusual orders to a person. The workforce manager can tighten caps and routing bands in configuration without a release.

  6. Review. The security team reviewed the threat model in week three and the deployment in week six. The data protection officer approved the consent prompt and retention. Both are in the decision record.

FDE / 08Case study

08 / 12

Go-live

In shortSwitched on in small steps, with one switch to turn it off.

Shadow calls were the first gate. For four days, agents took calls as usual while the voice agent listened to the same audio and built a cart without speaking. Disagreements went into three piles: the voice agent was wrong, the agent was wrong, and the menu or offer rules were unclear. The third pile went to the POS product manager.

FIG. 8.1SHADOW CALLS

For four days

  1. the voice agent was wrong

  2. the agent was wrong

  3. the menu or offer rules were unclear

    went to the POS product manager

Live calls went by queue and time band, not by date: weekday afternoon overflow first, then weekday dinner overflow, then all in-scope calls. Each step needed three clean days of audit samples and no missed must-handoff call. Rollback was one routing switch back to the agent queues, rehearsed by the client's telephony engineer before the first live call. It was used once, for twenty minutes, when a POS release slowed the menu lookup during a Saturday peak: the latency threshold paged the engineer, who switched and switched back.

FIG. 8.2LIVE CALLS
  1. 01weekday afternoon overflow

  2. 02weekday dinner overflow

  3. 03all in-scope calls

EACH STEP · three clean days of audit samples and no missed must-handoff call

FIG. 8.3ROLLBACK · USED ONCE
  1. a POS release slowed the menu lookup during a Saturday peak

  2. the latency threshold paged the engineer

  3. for twenty minutes

    switched

  4. switched back

CONTROL ·one routing switch back to the agent queues

FDE / 09Case study

09 / 12

Handover

In shortThe client's own team proved they could run it before we left.

The engagement ended on the go-live criteria and the manifest, not on the calendar.

FIG. 9.1HANDOVER MANIFEST · 6 ITEMS
ItemWhat the client holds
The repositoryEvery commit in their source control, reviewed by their reviewers, the pilot's reusable parts included
The evaluation suiteThe 1,800-call golden set, the red-team slice, the replay harness against a staging POS, wired into CI
The runbookRelease, rollback, credential rotation, adding a menu item or offer, adding a handoff phrase, changing caps and routing bands, what to do when latency or audit accuracy drops
The decision recordEach architectural choice, the options considered and why one was taken, dated, with the security and data protection sign-offs
The accounts and keysAll under the client's cloud account and identity provider; no standing access for us
The trained ownersThe workforce manager owns routing bands and caps; the telephony engineer owns releases and on-call; a senior agent owns the weekly audit and the golden set

In the handover dry run, the client's telephony engineer added a new pizza to the menu vocabulary and two Hindi handoff phrases, scored the release against the golden set, released it and rolled it back, with our engineer in the room but not at the keyboard. The runbook was signed after that, not before.

Key numbers

The story in nine numbers

One number for each scene, from one call in seven lost to 95% answered. Press play, or pick a number.

Before01 / 09

1 in 7

calls hung up before an answer at the dinner peak

01Dinner rush

At dinner, the phones outran the people.

About 38,000 calls a day, and three times the average between 7.30 and 10 pm. At that peak, one call in seven hung up before an agent answered.

Read chapter 1
  • About 38,000 calls a day
  • 3× the average, 7.30–10 pm
  • 240 agents over three shifts

FDE / 10Case study

10 / 12

Results

In shortFaster answers at peak, 49% of order calls needing no person, and fewer wrong orders.

Figures are illustrative, measured on in-scope calls over the first 90 days after go-live.

FIG. 10.1RESULTS
  1. 95% of dinner-peak calls answered within 20 seconds, up from 68%. Abandoned calls at peak fell from 15% to 4%.

  2. 49% of in-scope order calls completed with no person involved. The rest reached an agent with the transcript and the cart already on screen.

  3. 97.8% whole-order accuracy on the weekly audit of orders taken by the voice agent, against 96.4% for agent-taken orders measured the same way. Addresses improved most, because every address is checked against the delivery area before the order exists.

  4. Every allergy mention reached a person. 1.1% of in-scope calls were handed off for this reason.

  5. No card number reached a model or a stored recording, confirmed by the weekly transcript scan.

  6. Offer write-offs stopped, because the promotions engine, not a person or a model, decides which offers apply.

  7. No agent was let go. The centre stopped hiring temporary staff for festival peaks and moved twelve agents to an outbound team.

What did not improve, and was never promised:

FIG. 10.2
  • Calls straight to outlet phones. They were out of scope, and outlets still take them by hand at peak.

  • Callers in other languages, and some older callers who ask for a person as soon as they hear the assistant. About 8% of in-scope calls ask for a person in the first ten seconds. That share has not moved.

  • Average handling time for calls handed to agents rose by about 40 seconds: the easy calls now go to the voice agent, so agents take the harder ones.

FDE / 11Case study

11 / 12

What we would tell the next client

In shortSix rules for anyone putting a voice agent on real orders.

FIG. 11.16 LESSONS
  1. Measure on your real calls, not a quiet room.

    Noise, accents and mixed languages are the workload, not the edge case.

  2. Never let the model name a price.

    The POS knows the menu and the offers. The model should ask it every time.

  3. Decide the handoffs before the dialogue.

    Allergies, complaints and payment go to a person. Write the list, test it, and fail the build if one is missed.

  4. Cap what looks like a prank.

    Forty pizzas can be real. A person should decide that, not a read back.

  5. Keep card numbers out of the conversation.

    A payment link or payment on delivery removes a whole class of risk.

  6. Start with overflow.

    Sending the voice agent the calls you would otherwise miss earns trust before it takes the ones you would answer.

FDE / 12Case study

12 / 12

What happened next

In shortThe client took it in-house, added a language themselves, and came back for more.

The client took the system in-house and did not buy managed operations. Two months after handover, their own team added a third language using the same harness and a new golden set recorded by their agents. They came back for a second engagement to extend the voice agent to outlet phone lines, which had been out of scope the first time and need a different routing design.

FDE / ENDStart

Eight to twelve weeks,and the pilot is in production.

Book a scoping call and bring the pilot; you leave knowing what stands between it and real users.