We connect what the agent said to what the tools did, and what the system finally recorded.
One journey, inside your own eval stack. We add the missing business truth and final-state checks, challenge them independently, and leave the tests behind. Ten working days, fixed scope.
Illustrative. ○ observed · ● verified · ■ final state · • caveat / edge case.
Five moves, ten working days.
Inspect the current suite
We establish what your eval already proves and what stays manual, assumed or disputed. If it already covers everything, we say so and stop there.
Contract the outcome
With your domain owner we map one journey: the business rules, the acceptable variations, and the prohibited states.
Add the missing tests
We encode the difficult, ambiguous, multilingual and adversarial cases, plus deterministic assertions on the final system state, in your existing tooling.
Challenge and retest
We independently review the whole chain, rank findings by business consequence, and verify one remediation cycle.
Hand over the coverage
The outcome contract, state assertions and accepted failures stay in your suite, ready for the next release.
Six lenses on one journey.
Lenses one to five ask is it correct? The sixth asks does it get there well?
Truth
Does the answer match your actual policy, prices and remedies?
- exceptions, exclusions and effective dates
- stale or contradictory knowledge
- invented remedies, deadlines or entitlements
- does it qualify or escalate when unsure?
Understanding
Does it grasp the real situation and hand off at the right moment?
- ambiguous, emotional or changing requests
- accents, dialects and the Swiss languages
- vulnerable or exceptional cases
- handoff with the context already gathered
Action & state
Does the system record exactly what the agent promised? Our core check.
- right tool, right parameters
- duplicate actions, retries, partial failures
- “success” messages when the tool failed
- confirmation before irreversible actions
Boundaries
Can untrusted content or manipulation push it past its authority?
- direct & indirect prompt injection
- cross-customer / cross-tenant data access
- poisoned documents, tickets or tools
- data leaked through links or replies
Change
Does the journey still pass after something in the system changes?
- model, prompt, temperature or policy
- updated knowledge, tools or permissions
- a new language, channel or segment
- a fix that quietly breaks another case
Efficiency & footprint
Does it reach the accepted outcome the efficient way?
- could a smaller / cheaper model still pass?
- data & context it pulls but doesn’t need
- steps that could be cached or deterministic
- does it leave a clean, auditable trace?
Three evidence levels.
We only claim what we can observe. The deeper the access, the stronger the proof. We always say which level applies.
Conversation
We evaluate the observable interaction only.
+ Trace
We verify tool choices and intermediate events behind each answer.
+ Final state
We assert what the system of record actually did, whether that is a CRM, claim, booking, payment or internal system.
A scoped test that produces reproducible evidence, executable tests you keep, and a defensible acceptance decision for one journey.
- legal certification or EU AI Act classification
- a general infrastructure penetration test
- another dashboard to adopt
- a guarantee that the agent is “safe”