03 / evidence over impressions
Private Evaluation
Would the result survive another run?
Define success as a predicate, not a persuasive explanation. Evaluate the final state, recovery behavior and tool constraints separately.
A private engagement begins with your evaluation question, environment interface and permitted data. We agree held-out tasks, run conditions and reporting boundaries before testing.
We do not publish benchmark scores without an executed protocol and auditable evidence. A dataset QA pass is not an agent performance result.
WHAT AN ENGAGEMENT LOOKS LIKE
- Problem
- A convincing answer may hide an incorrect environment outcome.
- What we produce
- Held-out tasks and a reproducible evaluation protocol with deterministic checks where applicable.
- What you provide
- The agent interface, evaluation question, constraints and authorized evaluation material.
- What you receive
- An executed evaluation report, per-task evidence and explicit limitations; no unrun scores.