Independent outcome acceptance · Switzerland

Your eval scores the agent.
We verify the outcome.

Most tests grade what an AI agent says. We check what actually happened. Was the refund created? Is the booking right? Does the record match the promise? We add those checks to your existing test suite, and your team keeps them.

For vendors and teams putting AI agents into real work: customer service, claims, bookings, internal operations. Whether they face customers or run behind the scenes.

Acceptance gap 1 open

gap #04: the refund recorded does not match the amount the agent confirmed.

reproducible · 7/10 runs · ready for retest


The gap

A high eval score can hide a wrong outcome.

Once an agent can act (issue a refund, change a booking, update a record), a fluent answer is no longer proof. The agent can say “done” while your system did something else, or nothing at all. That gap is what stalls a launch or an internal go-live.

What we actually checkone journey, end to end
the customerasked
the agentdid
the systemrecorded

We reconcile all three. If they disagree, we show you exactly where, with a reproducible case your team can rerun. Not a screenshot or a score.

Why it matters

These failures already happened, in public.

None of them were exotic attacks. Each one is the kind of case an outcome test is built to catch before launch. And each one cost real money or trust.

Airline · ruled Feb 2024

The chatbot invented a policy.

Air Canada’s chatbot told a customer a bereavement fare could be claimed after the flight. The real policy said otherwise. A tribunal held the airline liable for its chatbot’s answer.

To the customer, the agent’s answer is the company speaking.

Moffatt v. Air Canada, 2024 BCCRT 149 ↗
Government · audit Dec 2025

Many answers, failed job.

An official audit found New York City’s MyCity chatbot could not give accurate, consistent information. The wider programme had already spent over $100 million.

A service can answer many questions and still fail the job it was funded for.

NYC Comptroller audit ↗
Retail · ended Jun 2024

Right words, wrong order.

McDonald’s ended its AI drive-through trial after orders kept going wrong. Accents, corrections and background noise defeated the system, however fluent it sounded.

The unit of quality is a completed order, not a clean transcript.

Associated Press, Jun 2024 ↗
How it works

Evidence, not theatre.

Three moves, inside your own tools. We use your existing eval stack and add only what it cannot prove.

Three moves
01

Define the outcome

Agree with your domain owner what a correct result and a prohibited result actually are, for one journey that matters.

02

Challenge the path

Encode the difficult, ambiguous and adversarial cases your suite is missing, and inspect the tool activity behind each answer.

03

Verify the state

Assert what the system of record finally did, then hand the executable tests back to your team.

See the full method
Two questions, one engagement

Is it correct, and does it get there efficiently?

Outcome

Did the journey end correctly?

We verify the business result across conversation, tool activity and final system state. Every case we accept becomes a reusable release test.

Operation

The same outcome, run leaner.

Once the outcome is provable, we can test whether a smaller model and less data reach the same accepted result. That lets you cut run-cost without gambling on quality. We sell no model, so the answer is honest.

Early access

Join the waitlist.

We’re onboarding a small number of teams at a time. Leave your email and we’ll get in touch when a slot opens.