Your eval scores the agent.
We verify the outcome.
Most tests grade what an AI agent says. We check what actually happened. Was the refund created? Is the booking right? Does the record match the promise? We add those checks to your existing test suite, and your team keeps them.
For vendors and teams putting AI agents into real work: customer service, claims, bookings, internal operations. Whether they face customers or run behind the scenes.
gap #04: the refund recorded does not match the amount the agent confirmed.
reproducible · 7/10 runs · ready for retest
A high eval score can hide a wrong outcome.
Once an agent can act (issue a refund, change a booking, update a record), a fluent answer is no longer proof. The agent can say “done” while your system did something else, or nothing at all. That gap is what stalls a launch or an internal go-live.
We reconcile all three. If they disagree, we show you exactly where, with a reproducible case your team can rerun. Not a screenshot or a score.
These failures already happened, in public.
None of them were exotic attacks. Each one is the kind of case an outcome test is built to catch before launch. And each one cost real money or trust.
The chatbot invented a policy.
Air Canada’s chatbot told a customer a bereavement fare could be claimed after the flight. The real policy said otherwise. A tribunal held the airline liable for its chatbot’s answer.
To the customer, the agent’s answer is the company speaking.
Moffatt v. Air Canada, 2024 BCCRT 149 ↗Many answers, failed job.
An official audit found New York City’s MyCity chatbot could not give accurate, consistent information. The wider programme had already spent over $100 million.
A service can answer many questions and still fail the job it was funded for.
NYC Comptroller audit ↗Right words, wrong order.
McDonald’s ended its AI drive-through trial after orders kept going wrong. Accents, corrections and background noise defeated the system, however fluent it sounded.
The unit of quality is a completed order, not a clean transcript.
Associated Press, Jun 2024 ↗Evidence, not theatre.
Three moves, inside your own tools. We use your existing eval stack and add only what it cannot prove.
Define the outcome
Agree with your domain owner what a correct result and a prohibited result actually are, for one journey that matters.
Challenge the path
Encode the difficult, ambiguous and adversarial cases your suite is missing, and inspect the tool activity behind each answer.
Verify the state
Assert what the system of record finally did, then hand the executable tests back to your team.
Start small. Add cover only when a change demands it.
Fixed scope, agreed up front. Every engagement leaves reusable tests with your team. No platform, no lock-in.
Gap Review
We review your existing tests against one journey and name, in writing, what stays unproven.
Core · two weeksOutcome Acceptance Sprint
One agent, one journey. We close the acceptance gap and leave the executable tests with your team.
Recurring · on changeRelease Challenge
We independently re-run the high-consequence suite whenever something material changes.
Is it correct, and does it get there efficiently?
Did the journey end correctly?
We verify the business result across conversation, tool activity and final system state. Every case we accept becomes a reusable release test.
The same outcome, run leaner.
Once the outcome is provable, we can test whether a smaller model and less data reach the same accepted result. That lets you cut run-cost without gambling on quality. We sell no model, so the answer is honest.
Join the waitlist.
We’re onboarding a small number of teams at a time. Leave your email and we’ll get in touch when a slot opens.
We use your address only to contact you about working with Arctiq. No newsletter spam. See the privacy notice.