Method

We connect what the agent said to what the tools did, and what the system finally recorded.

One journey, inside your own eval stack. We add the missing business truth and final-state checks, challenge them independently, and leave the tests behind. Ten working days, fixed scope.

Outcome contractclaim intake · v3accepted with conditions
customer journey
journey_7f2a9c2b
defined
tool trace
trace_9b3204e7
observed
final state
state_0c8f1a6d
asserted
edge case: ambiguous date
case_e1…
not yet reproduced

Illustrative. ○ observed · ● verified · ■ final state · caveat / edge case.

The engagement

Five moves, ten working days.

01

Inspect the current suite

We establish what your eval already proves and what stays manual, assumed or disputed. If it already covers everything, we say so and stop there.

02

Contract the outcome

With your domain owner we map one journey: the business rules, the acceptable variations, and the prohibited states.

03

Add the missing tests

We encode the difficult, ambiguous, multilingual and adversarial cases, plus deterministic assertions on the final system state, in your existing tooling.

04

Challenge and retest

We independently review the whole chain, rank findings by business consequence, and verify one remediation cycle.

05

Hand over the coverage

The outcome contract, state assertions and accepted failures stay in your suite, ready for the next release.

What we look through

Six lenses on one journey.

Lenses one to five ask is it correct? The sixth asks does it get there well?

01 · truth

Truth

Does the answer match your actual policy, prices and remedies?

  • exceptions, exclusions and effective dates
  • stale or contradictory knowledge
  • invented remedies, deadlines or entitlements
  • does it qualify or escalate when unsure?
02 · understanding

Understanding

Does it grasp the real situation and hand off at the right moment?

  • ambiguous, emotional or changing requests
  • accents, dialects and the Swiss languages
  • vulnerable or exceptional cases
  • handoff with the context already gathered
03 · action

Action & state

Does the system record exactly what the agent promised? Our core check.

  • right tool, right parameters
  • duplicate actions, retries, partial failures
  • “success” messages when the tool failed
  • confirmation before irreversible actions
04 · boundaries

Boundaries

Can untrusted content or manipulation push it past its authority?

  • direct & indirect prompt injection
  • cross-customer / cross-tenant data access
  • poisoned documents, tickets or tools
  • data leaked through links or replies
05 · change

Change

Does the journey still pass after something in the system changes?

  • model, prompt, temperature or policy
  • updated knowledge, tools or permissions
  • a new language, channel or segment
  • a fix that quietly breaks another case
06 · efficiency

Efficiency & footprint

Does it reach the accepted outcome the efficient way?

  • could a smaller / cheaper model still pass?
  • data & context it pulls but doesn’t need
  • steps that could be cached or deterministic
  • does it leave a clean, auditable trace?
Honest scoping

Three evidence levels.

We only claim what we can observe. The deeper the access, the stronger the proof. We always say which level applies.

Level 1

Conversation

We evaluate the observable interaction only.

Level 2

+ Trace

We verify tool choices and intermediate events behind each answer.

Level 3

+ Final state

We assert what the system of record actually did, whether that is a CRM, claim, booking, payment or internal system.

What this is

A scoped test that produces reproducible evidence, executable tests you keep, and a defensible acceptance decision for one journey.

What this is not
  • legal certification or EU AI Act classification
  • a general infrastructure penetration test
  • another dashboard to adopt
  • a guarantee that the agent is “safe”