Physical & operations · 15
Evaluation sets built from your own traffic, red-teaming that runs before launch and keeps running after it, guardrails enforced outside the model, and the evidence your risk, security and compliance teams will actually accept.
For teams about to put a model in front of customers or a regulator, and for teams who already have and want to know, rather than hope, that it behaves.
01Overview
EvaluationA model is evaluated the way it will be used: on the questions your users actually ask, scored by judges checked against people, and attacked by a team whose job is to make it misbehave before anyone else does.
01/ 02
Evaluation
A model is evaluated the way it will be used: on the questions your users actually ask, scored by judges checked against people, and attacked by a team whose job is to make it misbehave before anyone else does.
Evaluation sets are drawn from your own conversations, documents and tickets, labelled with your team, and versioned. A model is scored on what it will meet, not on a public benchmark that says nothing about your domain.
Automated judges score thousands of outputs, but only after they agree with your reviewers on a labelled sample. Judge drift is measured, and the judge is re-calibrated when it moves.
Prompt injection, jailbreaks, data extraction, tool misuse and harmful content are attempted systematically, by people and by automated attackers, before launch and on a schedule afterwards. Findings become tests that run on every release.
An agent with tools has an attack surface: what it reads, what it can call and what it can be talked into. We write the threat model with your security team and design the controls against it, not against a generic checklist.
02/ 02
Security
Safety does not come from asking the model nicely. It comes from guardrails enforced around it, monitoring that watches what it does, a response plan for when it does the wrong thing, and a paper trail that a risk committee or a regulator can read.
Inputs are screened for injection and out-of-scope requests, tool calls are checked against permissions and limits, and outputs are checked for leakage, harm and policy before they are shown. Each layer is a deterministic control, tested on its own.
Evaluation results, red-team reports, threat models, control descriptions and change logs are produced in the form your risk, security and compliance teams need, mapped to the frameworks they work to.
A harmful output, a successful injection or a tool called out of bounds is an incident with a runbook: contain, roll back, notify, analyse and add the test that would have caught it. We rehearse it before launch.
Uptime says the model answered. Behaviour monitoring says what it answered: refusal rates, judge scores on live samples, tool-call patterns, topic drift and guardrail triggers, with alerts when any of them move.
At a glance
8 figures, one per capability. Open any to read it in full.
How it runs
Every engagement runs the same five steps, whatever the service.
We sit with the people who do the work today and write down every step, exception and hand-off before anything is built.
A held-out set of real cases, agreed with you, is the bar each build has to clear before it goes anywhere near production.
The system runs in parallel with the team for as long as it takes, and every disagreement between them is reviewed together.
The code, the prompts, the evaluation set and the runbooks are handed over in your accounts, under your keys.
We watch the runs, retrain and repair as the inputs drift, or train your own team to do the same.
Outcomes
Scores on your own golden set, a red-team report and a threat model in front of the people who sign off.
Every release runs the same evaluation and attack suite, and a drop fails the pipeline.
Controls, tests, results and incidents documented in the form your risk and compliance teams file.
Details
Questions we are asked about evaluation, safety and security
Yes. Evaluation, red-teaming and guardrails are applied to any model or vendor product you run, and are often the first engagement before any building starts.
From your own traffic and documents, sampled to cover the topics and the edge cases, labelled with your reviewers, and versioned so a score today can be compared with a score next quarter.
Only once it agrees with your reviewers on a labelled sample, and only while that agreement holds. We measure it and re-calibrate when it drifts.
The ones your risk and compliance teams work to. We produce the evidence in their format rather than asking them to learn ours.
Behaviour monitoring, scheduled red-teaming and an incident runbook that we rehearse with your team. Findings become tests on every subsequent release.
Start