Forward-Deployed Engineering · Retail · Catalogue onboarding and enrichment
Case study: supplier spreadsheets to a listing a merchandiser will approve
Illustrative engagement — not a client record. The company, people, volumes and results below are a representative composite, written to show how a forward-deployed engagement runs end to end. The workflow, the architecture, the controls and the method are real and technically valid.
Offering usedEmbedded Engineering PodsA pod inside your repository and your standups, quarter by quarter.
Illustrative results, first 90 days after each workflow's cut-over, in-scope categories only:
5 days
Median time from supplier submission to a live listing
Before 16 days
61%
New listings approved by merchandising with no edit
Before 21%
92%
Filterable attributes filled on new listings
Before 54%
0, by design
Mandatory legal fields written by a model
Before not applicable
Explainer09 views
The whole story in about a minute
Nine short scenes, from slow listings to the results. Press play, or pick a scene.
Days of the season
- Spent waiting
16
View 01Before
Supplier files took 16 days to become a listing.
days, median, from supplier submission to a live listing
Sixty catalogue executives retyped every item by hand. Merchandising still edited about four listings in five, and around 7% went live with a legal field missing.
- 01
- About 30,000 new items a month
- 02
- 4 in 5 listings edited before approval
- 03
- About 7% went live with a legal gap
At a glance
- Client
- A multi-brand retailer and marketplace: fashion, footwear, home and kitchen, and beauty. About 1,800 active suppliers, 1.2 million live listings, roughly 30,000 new items submitted each month
- Workflow
- Supplier submissions (spreadsheets, specification PDFs, pack photographs, product images) turned into complete, compliant catalogue entries in the product information system, approved by merchandising
- Engagement
- Embedded Engineering Pods, one quarter (thirteen weeks), followed by a second, smaller quarter
- Team
- A pod of three: a named lead, a platform engineer and a data engineer. Client side: the Head of Catalogue (owner), a category merchandising manager, two senior catalogue executives, a platform engineer, a legal and compliance reviewer
- Where it runs
- The client's own cloud account and region. No supplier data retained outside it
- Handover
- Runbooks executed by the client's engineers; the quarter ended with a written report and a renewal decision
| Measure | Before | After |
|---|---|---|
| Median time from supplier submission to a live listing | 16 days | 5 days |
| New listings approved by merchandising with no edit | 21% | 61% |
| Filterable attributes filled on new listings | 54% | 92% |
| Mandatory legal fields written by a model | not applicable | 0, by design |
| Duplicate listings found after going live, per 1,000 new items | 18 | 4 |
| New items processed per catalogue executive per day | 24 | 71 |
FDE / 01Case study
01 / 12
In shortTurning supplier files into listings was slow, retyped by hand, and risky on legal fields.
Suppliers sent product data the way they always had: a spreadsheet per consignment, in a template they had modified; a specification PDF from the manufacturer; a folder of images of mixed size and background; and, for pre-packaged goods, a photograph of the pack showing the legal panel.
Sixty catalogue executives, a third of them at an outsourced partner, turned these into listings. For each item they read the supplier's file, chose a category from a taxonomy of about 4,200 nodes, filled the attributes that category required, wrote a title and a description to the brand's style guide, attached a size chart, copied the legal declarations from the pack photograph, checked the images and pressed submit. A merchandiser then reviewed the listing before it went live.
The cost showed up in four places:
Time to shelf. A submission took a median of 16 days to become a live listing. For a seasonal fashion supplier, that is a fifth of the season.
Rework. Merchandising edited about four listings in five before approving. Wrong category and missing attributes were the two most common reasons, and they are also the two most common reasons marketplaces reject listings anywhere.
Compliance risk. For pre-packaged goods, the rules require declarations on the listing: the name and address of the manufacturer, packer or importer, country of origin, the common or generic name, net quantity, maximum retail price, best before where applicable and consumer care details. Under pressure, executives filled blanks with "NA" or copied a value from a similar item. An audit found around 7% of new listings going live with a legal field that was missing, placeholder or unverifiable.
Duplicates. The same manufacturer's product arrived from two distributors under different titles and became two listings, splitting reviews and ratings.
FDE / 02Case study
02 / 12
In shortRules and a vendor pilot had been tried, and legal had good reason to fear generated copy.
The problem had been attacked twice.
Rules and macros. The catalogue team had a large spreadsheet of mapping rules per supplier. Each new supplier template broke it. Nobody owned it after the analyst who wrote it left.
An enrichment vendor. A pilot with an outsourced data-enrichment service improved fill rates on one category, but worked from exports and returned files, so nothing reached the product information system automatically, and nothing was measured against the categories' requirements.
There was also a well-founded fear. A previous attempt at generated descriptions had produced fluent copy that claimed a garment was "100% cotton" when the composition said 60/40. Legal had asked, reasonably, that nothing be generated into a field that a regulator or a customer could hold the company to. That request became the strongest control in this engagement, rather than a reason not to start.
Nobody had tried to do the whole job: read what the supplier sent, map it, draft what can be drafted, hold what must not be invented, and put it in front of a merchandiser in the system they already use. It needed several workflows, not one, which is why a pod was the right shape rather than a single sprint.
FDE / 03Case study
03 / 12
In shortThree workflows, each with a metric and an owner, signed before the pod started.
The goals were agreed with the Head of Catalogue and signed before week one.
Three workflows, each with a metric and an owner, in this order:
- 1. Extraction, category mapping and the legal gateAttribute precision and recall on the golden set; category top-1 accuracy; zero model-written legal fields
- 2. Titles, descriptions and size chartsShare of drafts approved by merchandising with no edit; zero claims not evidenced in the source
- 3. Duplicate detection and image checksDuplicates caught before going live; false-merge rate
Written into the same page:
FIG. 3.3Scope. Four category groups: fashion, footwear, home and kitchen, and beauty. About 2,900 taxonomy nodes and the attribute schema each requires.
Out of scope. Pricing, supplier contracting, returns, the storefront itself, and any category needing a licence check the compliance team runs by hand.
Guardrails. Regulated and legal fields are never generated. Missing means held for the supplier. Merchandising approves every listing before it goes live. No listing is published by the system.
Pairing. One client engineer paired on each workflow, with time to review and release it.
The pod joined the catalogue team's stand-up in week one and merged its first pull requests that week: the supplier submission intake, storing every file with a hash and a supplier identifier, so that later work had something real to read.
Timeline09 milestones
The engagement, stage by stage
Nine stages on one clock, from before the quarter to ninety days after cut-over. Press play, or pick a stage.
Tracking number
TESBS-09
Status
Results
Delivered
+90 days
Tracking history09 / 09 updates
- +90 days
Faster listings, and fewer edits.
5 days, median, from submission to a live listing, down from 16
- 61% approved with no edit, up from 21%
- Attributes filled 54% → 92%
- Duplicates 18 → 4 per 1,000
- W13
The client's team runs it themselves.
7 things the client holds at the quarter's end
- Workflows, repository and harness
- Runbooks and the quarter report
- Pod access revoked
- W5–W12
Shadowed first, then switched on step by step.
3 clean days before each rollout step
- A week in shadow first
- By supplier group, then category
- One flag back to the manual queue
- W4–W9
Every release is scored on 3,000 corrected listings.
3,000 catalogue entries every release is scored on
- Legal fields: pass or fail
- Week four: a pre-filled “India”
- Week nine: a title change failed
- W5
No evidence, no value: the item is held.
0 legal fields written by a model
- Held for the supplier
- The missing field named
- A clearer pack photo asked for
- W2–W12
The model drafts. Rules check. A person approves.
3 category choices, with evidence, when the match is unsure
- Every value carries its evidence
- Rules decide what is mandatory
- Written through the import interface
- Before W1
First, three goals, signed before week one.
3 workflows, each with a metric and an owner
- Four category groups in scope
- About 2,900 taxonomy nodes
- Legal fields never generated
- Earlier
It had been tried twice, and neither stuck.
2 earlier attempts: mapping rules, then a vendor pilot
- Rules broke on new templates
- Vendor pilot worked from exports
- Generated “100% cotton” for 60/40
- Before
Supplier files took 16 days to become a listing.
16 days, median, from supplier submission to a live listing
- About 30,000 new items a month
- 4 in 5 listings edited before approval
- About 7% went live with a legal gap
FDE / 04Case study
04 / 12
In shortThirteen weeks: build, score, shadow, then cut three workflows over one at a time.
W1
What happened
Pod in the repository and the stand-up; access to the product information system's import interface, the taxonomy service, the supplier portal and the image store; quarter goals signed
What existed at the end
First pull requests merged; intake storing submissions
What happened
Pod in the repository and the stand-up; access to the product information system's import interface, the taxonomy service, the supplier portal and the image store; quarter goals signed
What existed at the end
First pull requests merged; intake storing submissions
What happened
Parsing of spreadsheets and specification PDFs; attribute extraction to the category schema with evidence for every value; category mapping; the golden set agreed with two senior catalogue executives and the merchandising manager
What existed at the end
A harness in the client's pipeline and a first scored release
What happened
The legal gate built after the country-of-origin finding (section 6); shadow run on live submissions
What existed at the end
Every submission processed in parallel, nothing written
What happened
Workflow 1 cut over: 5 supplier groups, then all in-scope suppliers, drafts landing in the merchandising review queue
What existed at the end
Extraction and mapping in production
What happened
Title and description drafting against the style guide; size-chart normalisation; the claims check; merchandiser feedback loop wired into the harness
What existed at the end
Drafts scored on the golden set, approved or edited in the queue
What happened
Workflow 2 cut over: three categories, then all four category groups
What existed at the end
Drafted content in production
What happened
Duplicate detection built and cut over. Image checks ran in shadow only and did not reach production; the lead said so in the weekly report in week eleven, and the owner chose to carry it rather than cut the golden set short
What existed at the end
Duplicate detection in production; image checks scored but not live
What happened
Quarter review; runbooks executed by the client's engineers; renewal decided in writing
What existed at the end
Quarter report, signed runbooks, decision record
FDE / 05Case study
05 / 12
In shortModels read and draft, rules decide what is allowed, and a merchandiser approves every listing.
Try a submission05 / 10
Pick what the supplier sends. Watch where it goes.
The same pipeline reads every submission. What the files can prove decides whether it becomes a listing or is held for the supplier.
Listing written
Listing written
Only an approved listing reaches the product information system. The system publishes nothing on its own.
Submission log01
A complete submission“Spreadsheet, specification PDF, pack photograph and product images”
- 01Supplier filesSheet, specification PDF, pack photograph and images arrive
- 02IntakeStored in the client's storage with a hash and the supplier's identifier
- 03ExtractAttributes filled from the files; every value carries its evidence
- 04Map to categoryTop three category nodes, with scores
- 05Legal gateEvery mandatory declaration has supplier-provided evidence
- 06DraftTitle from the template; description from the supplied attributes only
- 07Claims checkNothing to remove
- 08Duplicate checkNo barcode match, no similar listing
- 09MerchandiserThe merchandiser checks the evidence and approves
- 10Listing writtenWritten through the import interface
Submission status
- Legal fields by a model
- 0
- Published by the system
- 0
- Runs in
- The client's region
Listing written
Intake and parse. Submissions land in the client's storage with a content hash and the supplier's identifier. Spreadsheets are parsed with a learned column mapper that proposes which column is which, and remembers each supplier's template once a human has confirmed it. Specification PDFs are read by a layout-aware document model. Pack photographs are read for the legal panel.
Extract. A language model, reached through a private endpoint in the client's region with no data retention, fills the category's attribute schema from the parsed sources. Every value carries its evidence: the cell, the page region or the image region it came from. A value with no evidence is not a value; it is an empty field with a reason.
Map. Category mapping is a retrieval step over the taxonomy followed by a ranking model that returns the top three nodes with scores. Above a confidence threshold the top node is proposed; below it, the listing reaches the merchandiser with three choices and the evidence for each. The merchandiser always sees which node was proposed and can change it.
The legal gate. Deterministic rules, not a model, decide whether a listing may proceed. For each category, the gate holds the list of mandatory declarations. A declaration may only be filled from supplier-provided evidence: a value in the supplier's file, a line in the specification PDF, or text read from the pack photograph with the region recorded. Nothing may be inferred from a similar product, a brand's other listings or the model's knowledge. If a mandatory field cannot be evidenced, the item is held and the supplier receives a specific request naming the field and, where the pack photograph was unreadable, asking for a clearer photograph of that panel. Tax classification codes are treated the same way and go to the client's tax team, not to a model.
Draft. Titles are assembled by a template per category from attributes that passed the gate, so a title cannot contain an attribute the item does not have. Descriptions are drafted by the model with one instruction that matters more than the rest: use only the attributes supplied. A rules-based claims check then removes or blocks language not evidenced in the source, from a list written with legal: material and composition claims, "organic", "dermatologically tested", "anti-bacterial", country claims, warranty terms and anything superlative. Size charts are normalised from the supplier's own chart, with unit conversion and a mapping to the retailer's size system. A size chart is never invented; a supplier without one is asked for it.
Duplicate and image checks. Exact matching on barcode identifiers runs first. Then a similarity search over brand, normalised title, key attributes and image hashes proposes candidates with a score. The system never merges. It attaches "possible duplicate of ..." to the listing, and a merchandiser decides. Image checks score resolution, background, aspect ratio, count per listing, watermarks and whether the pack panel is legible; they ran in shadow at the end of the quarter and went live in the second.
Review and write-back. Everything lands in the merchandising review queue: the proposed listing, the evidence for each value, the category choices, duplicate candidates and any held fields. A merchandiser approves, edits or rejects with a reason. Only an approved listing is written to the product information system, through the import interface a catalogue executive's own submissions use. The system publishes nothing on its own.
- Model
The model reads, extracts and drafts.
- POS
Rules decide what is mandatory, what counts as evidence and what language is allowed.
- Caller
A named merchandiser approves every listing.
- Person
Nothing reaches the storefront without that approval.
System map18 objects
How the pieces connect
Every part of the system and of the story, lit one scene at a time. Press play, or pick an event.
- Runbooksone per workflow, Document. Not in this scene
- Client engineers, Team. Not in this scene
- Golden set3,000 entries, Dataset. Not in this scene
- Harnessevery release, Check. Not in this scene
- Quarter goals3 workflows, Document. Not in this scene
- Vendor pilotexported files, Project. Not in this scene
- Supplierabout 1,800, Company. In this scene
- Intakehash per file, Service. Not in this scene
- Extractevidence per value, Model. Not in this scene
- Mappingtop three nodes, Service. Not in this scene
- Legal gatenot a model, Rules. Not in this scene
- Review queue, Queue. Not in this scene
- Catalogueproduct information, System. Not in this scene
- Executives60 in catalogue, Team. In this scene
- Taxonomyabout 4,200 nodes, System. In this scene
- Held itemfor the supplier, Hold. Not in this scene
- Merchandiserapproves every listing, Person. Not in this scene
- Feature flagused once, Control. Not in this scene
TeamBefore
Executives
60 in catalogue
Supplier files took 16 days to become a listing.
16days, median, from supplier submission to a live listing
- About 30,000 new items a month
- 4 in 5 listings edited before approval
- About 7% went live with a legal gap
Sixty catalogue executives retyped every item by hand. Merchandising still edited about four listings in five, and around 7% went live with a legal field missing.
Read chapter 1FDE / 06Case study
06 / 12
In short3,000 corrected listings that every release is scored on, with legal fields pass or fail.
The golden set was agreed in week four and became the pod's shared reference for the whole quarter.
3,000 catalogue entries across 40 representative categories in the four groups, taken from listings senior merchandisers had already corrected, with the supplier's original files kept alongside.
Answers verified against the goods for a sample of 300 items, checked against the physical pack and label in the warehouse. Where the live listing and the pack disagreed, the pack won and the listing was corrected.
A red-team slice: a supplier sheet with the country-of-origin column pre-filled; a pack photograph with the price sticker over the printed panel; a barcode belonging to another brand; a size chart in centimetres labelled inches; a description containing a composition claim the specification contradicts; the same item from two distributors with different titles; an imported item whose origin appears only on the pack.
- RT-01a supplier sheet with the country-of-origin column pre-filled
- RT-02a pack photograph with the price sticker over the printed panel
- RT-03a barcode belonging to another brand
- RT-04a size chart in centimetres labelled inches
- RT-05a description containing a composition claim the specification contradicts
- RT-06the same item from two distributors with different titles
- RT-07an imported item whose origin appears only on the pack
Scoring by field class. Attribute precision and recall per class; category top-1 and top-3 accuracy; title conformance by rule; duplicate precision and recall. Legal fields are scored pass or fail: any mandatory value written without evidence fails the build, whatever else improves.
- IF any mandatory value written without evidenceBLOCKS
fails the build, whatever else improves
The suite runs in the client's pipeline on every release, and the same harness scored all three workflows. Two things it caught are worth recording.
The false assumption, found in week four. The supplier template had a country-of-origin column with "India" pre-filled as a default. Most suppliers never changed it. Analysis against the sample of packs showed that around 9% of items carrying "India" in the sheet were imported. The extraction had been treating the cell as evidence, which it was not. The rule changed: for imported goods the origin must come from the pack or the specification document, and a default value in a template is not evidence. Around 2,400 submissions in the first month were held for a clearer pack photograph. The merchandising manager judged that a good trade, especially with the rules moving towards a searchable country-of-origin filter for imported products. The supplier template was changed too, so the default no longer exists.
The regression, found in week nine. A change to the title template improved footwear titles and dropped the brand name from home and kitchen titles, where the brand sits in a different attribute. Title conformance on that group fell below its threshold and the build failed. No listing reached a merchandiser with a nameless title.
FDE / 07Case study
07 / 12
In shortEverything stays in the client's own cloud, and every value can be traced to its source.
Perimeter. Everything runs in the client's cloud account and region. Egress is limited to the private model endpoint, which retains nothing. Supplier files stay in the client's storage.
Identity. The service account can read submissions and write draft listings through the import interface. It cannot publish, cannot change price and cannot merge listings. The pod's access went through the client's identity provider, showed in their audit log like anyone else's, and was revoked at the end of the engagement.
Can
- read submissions
- write draft listings through the import interface
Cannot
- publish
- change price
- merge listings
Autonomy is bounded. Models read, extract and draft. Rules decide what is mandatory and what counts as evidence. A merchandiser approves. Legal and regulated fields are never generated: missing means held, and held means the supplier is asked.
Traceability. Every value in every listing can be traced to the file, page or image region it came from, with the model version and the rule results, in the client's log store. When a customer or a regulator asks why a listing says what it says, the answer is a record, not a recollection.
Review. The compliance reviewer approved the mandatory-field lists per category and the claims list in month one, and reviewed the held-item workflow in month two. The security team reviewed the design and the deployment. Findings and closures are in the decision record.
FDE / 08Case study
08 / 12
In shortEach workflow ran in shadow first, then went live in small steps, with a flag to turn it off.
Each workflow cut over on its own, and each was shadowed first.
For extraction and mapping, the shadow ran for a week on live submissions while executives worked as before. The two results were compared field by field, and disagreements went into three piles: the system was wrong, the executive was wrong, and the category's requirement was unclear. The third pile, mostly attributes that two categories defined differently, went to the merchandising manager as taxonomy decisions.
for a week
the system was wrong
the executive was wrong
the category's requirement was unclear
went to the merchandising manager as taxonomy decisions
Rollout went by supplier group, then by category. Each step needed three clean days on the audit sample and no legal-gate failure. Rollback for each workflow was a feature flag that returned submissions to the manual queue with nothing lost; it was rehearsed by the client's platform engineer before each cut-over and used once, for half a day, when a taxonomy release renamed nodes the mapper had cached.
01by supplier group
02by category
EACH STEP · three clean days on the audit sample and no legal-gate failure
The catalogue executives were part of the design from week two. Their work changed from typing to judging: the two senior executives who built the golden set now own it, and the queue shows them the evidence rather than a blank form.
FDE / 09Case study
09 / 12
In shortThe client's engineers proved they could release and roll back before the quarter ended.
The quarter ended on the manifest, not on the calendar.
In the dry run, the client's platform engineer added a category schema, changed a mandatory-field list, scored the release against the golden set, released it and rolled it back, with the pod lead in the room but not at the keyboard.
Key numbers09
The story in nine numbers
One number for each scene, from 16 days to a live listing down to 5. Press play, or pick a number.
No. 01Before
16
days, median, from supplier submission to a live listing
- About 30,000 new items a month
- 4 in 5 listings edited before approval
- About 7% went live with a legal gap
01Slow listings
Supplier files took 16 days to become a listing.
Sixty catalogue executives retyped every item by hand. Merchandising still edited about four listings in five, and around 7% went live with a legal field missing.
Read chapter 1FDE / 10Case study
10 / 12
In shortListings go live in 5 days instead of 16, more are approved unedited, and no legal field is model-written.
Figures are illustrative, measured on in-scope categories over the first 90 days after each workflow's cut-over.
Median time from submission to live listing fell from 16 days to 5. Most of the remaining time is the supplier answering a hold.
61% of new listings were approved with no merchandiser edit, up from 21%. Merchandising still reviews every listing; it now reviews evidence rather than retyping data.
BEFORE21%AFTER61%Filterable attribute fill rose from 54% to 92% on new listings in the in-scope categories.
BEFORE54%AFTER92%No mandatory legal field was written by a model. Items missing one are held: about 11% of submissions in the first month, falling to 6% by the third as suppliers adjusted to the requests.
Duplicates found after going live fell from 18 to 4 per 1,000 new items, and no listings were merged automatically.
Throughput per catalogue executive rose from about 24 to 71 items a day. The outsourced contract was reduced at renewal; no in-house executive lost their job, and two moved into merchandising.
What did not improve, and was never promised:
FIG. 10.2Image quality from small suppliers. The checks can only flag a photograph taken on a shop floor; somebody still has to reshoot it. Image checks were also the workflow that did not reach production in the quarter.
Suppliers who do not answer holds. A held item waits. For a small group of suppliers, the median time to a live listing is no better than before, and the report says so by supplier.
The 1.2 million listings already live. This engagement was about new submissions. Back-filling the existing catalogue was named in the quarter report as the next piece of work, not claimed as done.
FDE / 11Case study
11 / 12
In shortSix rules for anyone putting a model near product data.
- A default in a template is not evidence.
Check what your forms pre-fill before you trust a field.
- Decide what may never be generated, on day one.
The legal gate made the compliance review short and the rollout uncontroversial.
- Keep every value's evidence.
It is what turns a review from retyping into judging, and it is what you show a regulator.
- Never merge duplicates automatically.
Flag, score, and let a merchandiser decide. A wrong merge is much harder to undo than a duplicate.
- Score by field class, not by average.
A 90% average can hide a legal field that is wrong half the time.
- Say in the weekly report when a goal will not ship.
The image checks were carried to the next quarter in week eleven, not discovered missing in week thirteen.
FDE / 12Case study
12 / 12
In shortA second, smaller quarter for image checks, supplier requests and the existing catalogue.
The renewal decision was written in the last week of the quarter: a second, smaller quarter with two engineers, to put image checks into production, to give suppliers a portal view of held items with the exact request, and to start back-filling legal fields on the existing catalogue in the categories where they were weakest. The pod's last two weeks of that quarter were the handover: runbooks, decision records and a dry run by the client's engineers, after which the catalogue team ran all four workflows without us.

