Services

Start small. Add cover only when a real change demands it.

Fixed scope, inside your own tools, for agents that face customers or run inside your operation. Every engagement leaves reusable tests with your team.

Entry

Eval-to-Outcome Gap Review

a few days · fixed scope

We review your existing test set against one journey and the release decision, and name what stays unproven.

  • a written acceptance-gap statement
  • a go / no-go on a full sprint
  • best when the gap isn’t yet scoped
Core

Outcome Acceptance Sprint

two weeks · fixed scope

One agent, one journey. We close the acceptance gap and leave the tests with your team.

  • eval-gap map & outcome contract
  • executable cases + final-state assertions
  • reproducible evidence, ranked by impact
  • customer-owned tests + one retest
Recurring

Managed Release Challenge

ongoing · change-triggered

We independently re-run the high-consequence suite when something material changes: a model, a prompt, a policy, a tool.

  • change-triggered reruns
  • production failures added as tests
  • a release delta each cycle

Fixed scope, agreed up front. You keep the tests either way. No lock-in, no automatic renewal. Pricing is set per engagement after a short scoping call.

Who it is for

We are honest about when to call us.

A fit when
  • an agent, customer-facing or internal, goes live or changes within 90 days
  • it retrieves account data or triggers a business workflow
  • an enterprise customer needs acceptance evidence
  • a staging path and an accountable owner exist
Probably not yet when
  • it only answers low-consequence FAQs
  • a human approves every material action and QA is enough
  • there is no named release decision
  • you need certification or a generic pentest
What you keep

Every engagement leaves an asset behind, not just a report.

Whatever the scope, you walk away with executable tests your own team can rerun on the next release. Nothing is locked to us.

Acceptance packhanded to your teamyours to keep
Outcome contract
correct & prohibited states
agreed
Executable test cases
in your existing stack
rerunnable
Final-state assertions
system-of-record checks
deterministic
Reproducible evidence
transcripts + traces
replayable
Prioritised findings
ranked by business impact
with owners
One remediation retest
after your fixes
confirmed

Illustrative. ○ observed · ● verified · ■ final state · caveat.

Fair questions

What teams ask before they call us.

“We already have evaluation tools.”

Good. We expect that, and we use them. We add value only where an important business outcome, exception or final-state check remains unproven. If nothing does, we say so.

“The vendor should test its own agent.”

It should. We come in where supplier and enterprise need independent evidence, or share an acceptance question neither side can settle alone.

“We can’t share production data.”

We default to staging, test identities, synthetic or minimised data, and customer-controlled access. If the required evidence can’t be reached safely, we narrow the scope and state that limit in the report.

“Can you certify the agent?”

No. We report what was tested and what was observed, with the remaining uncertainty stated. Legal classification and certification are outside our scope, on purpose.

Early access

We onboard a small number of teams at a time.

Join the waitlist and we’ll get in touch when a slot opens.