Physical & operations · 15

Proof that the model does what you think it does.

Evaluation sets built from your own traffic, red-teaming that runs before launch and keeps running after it, guardrails enforced outside the model, and the evidence your risk, security and compliance teams will actually accept.

For teams about to put a model in front of customers or a regulator, and for teams who already have and want to know, rather than hope, that it behaves.

01Overview

Evaluation

A model is evaluated the way it will be used: on the questions your users actually ask, scored by judges checked against people, and attacked by a team whose job is to make it misbehave before anyone else does.

01/ 02

Evaluation

Tested on your traffic, judged by calibrated judges, attacked on purpose.

A model is evaluated the way it will be used: on the questions your users actually ask, scored by judges checked against people, and attacked by a team whose job is to make it misbehave before anyone else does.

  1. 01

    Golden sets from your real traffic

    Evaluation sets are drawn from your own conversations, documents and tickets, labelled with your team, and versioned. A model is scored on what it will meet, not on a public benchmark that says nothing about your domain.

  2. 02

    LLM-as-judge, calibrated by people

    Automated judges score thousands of outputs, but only after they agree with your reviewers on a labelled sample. Judge drift is measured, and the judge is re-calibrated when it moves.

  3. 03

    Red-teaming before launch and after

    Prompt injection, jailbreaks, data extraction, tool misuse and harmful content are attempted systematically, by people and by automated attackers, before launch and on a schedule afterwards. Findings become tests that run on every release.

  4. 04

    Threat models for agents and prompts

    An agent with tools has an attack surface: what it reads, what it can call and what it can be talked into. We write the threat model with your security team and design the controls against it, not against a generic checklist.

02/ 02

Security

Controls outside the model, and evidence you can file.

Safety does not come from asking the model nicely. It comes from guardrails enforced around it, monitoring that watches what it does, a response plan for when it does the wrong thing, and a paper trail that a risk committee or a regulator can read.

  1. 05

    Guardrails at input, tool and output

    Inputs are screened for injection and out-of-scope requests, tool calls are checked against permissions and limits, and outputs are checked for leakage, harm and policy before they are shown. Each layer is a deterministic control, tested on its own.

  2. 06

    Evidence your risk team can file

    Evaluation results, red-team reports, threat models, control descriptions and change logs are produced in the form your risk, security and compliance teams need, mapped to the frameworks they work to.

  3. 07

    Incident response for model behaviour

    A harmful output, a successful injection or a tool called out of bounds is an incident with a runbook: contain, roll back, notify, analyse and add the test that would have caught it. We rehearse it before launch.

  4. 08

    Monitoring that watches behaviour, not uptime

    Uptime says the model answered. Behaviour monitoring says what it answered: refusal rates, judge scores on live samples, tool-call patterns, topic drift and guardrail triggers, with alerts when any of them move.

At a glance

Every capability, at a glance

8 figures, one per capability. Open any to read it in full.

How it runs

From first call to running unattended

Every engagement runs the same five steps, whatever the service.

  1. 01

    Map the process

    We sit with the people who do the work today and write down every step, exception and hand-off before anything is built.

  2. 02

    Build against a golden set

    A held-out set of real cases, agreed with you, is the bar each build has to clear before it goes anywhere near production.

  3. 03

    Run beside the team

    The system runs in parallel with the team for as long as it takes, and every disagreement between them is reviewed together.

  4. 04

    Hand over the keys

    The code, the prompts, the evaluation set and the runbooks are handed over in your accounts, under your keys.

  5. 05

    Keep it running

    We watch the runs, retrain and repair as the inputs drift, or train your own team to do the same.

Outcomes

What this looks like when it lands.

  • 01

    A launch decision made on evidence

    Scores on your own golden set, a red-team report and a threat model in front of the people who sign off.

  • 02

    Regressions caught before customers see them

    Every release runs the same evaluation and attack suite, and a drop fails the pipeline.

  • 03

    A risk file that survives an audit

    Controls, tests, results and incidents documented in the form your risk and compliance teams file.

Details

On the spec sheet.

Service
AI Evaluation, Safety & Security
Group
Physical & operations
Capabilities
8
Engagement
Map, build, run beside the team, hand over, keep running
Ownership
Your accounts, your keys, your region

Questions we are asked about evaluation, safety and security

Asked before signing.

Do you evaluate models you did not build?

Yes. Evaluation, red-teaming and guardrails are applied to any model or vendor product you run, and are often the first engagement before any building starts.

How is a golden set built?

From your own traffic and documents, sampled to cover the topics and the edge cases, labelled with your reviewers, and versioned so a score today can be compared with a score next quarter.

Can an automated judge be trusted?

Only once it agrees with your reviewers on a labelled sample, and only while that agreement holds. We measure it and re-calibrate when it drifts.

What frameworks do you map evidence to?

The ones your risk and compliance teams work to. We produce the evidence in their format rather than asking them to learn ours.

What happens after launch?

Behaviour monitoring, scheduled red-teaming and an incident runbook that we rehearse with your team. Findings become tests on every subsequent release.

Start

Tell us what you are about to ship and we will tell you what we would test first.