WINDWAY CONTROLLED BENCHMARK
blind_retry
Local policy · WFRB v0
Experimental · Not a standard · Not peer reviewed
- Agent
- blind_retry
- Public interface tested
- Not applicable: a local Windway-controlled environment.
- Evaluation date
- 2026-10-04
- Benchmark version
- wfrb-0.1.0
- Sample size
- 300 cells per policy, across 50 layouts and 150 scenario designs. Two seeds; repeated once.
- Scenario categories
- blocked_path, nonadjacent_action, premature_completion
WFRB v0 / LOCAL CONTROL EXPERIMENT
Experimental · Not a standard · Not peer reviewed
These charts describe local control trials. They do not measure public agents, LLMs or customer systems.
Loading verified aggregate data…
Recovery by category
Pass / fail / inconclusive
Injected failure class distribution
Evaluations over time
Deterministic replay verification
Target comparison
Customer target comparison is unavailable. These local controls do not support a cross-model ranking.
Inspect local benchmark evidence ↗Scenario heatmap
Tap or expand a cell to inspect its exact denominator. Each layout/category contains two seeded trials per local policy in one invocation.
PASS means exact-cost recovery. FAIL means that recovery criterion was not met. Replay verification does not establish reproduction of a customer-agent failure.
Read the methodology ↗ · Inspect aggregate data ↗Observed behaviors
Exact-cost recovery passed in the retained local invocation:
0/300Reproducible failures
Local result streams repeat byte-for-byte. No public-agent failure reproduction is published.
Conversation replay
Awaiting verified public runs
No public conversation is available. Local navigation actions are not presented as User/Agent dialogue.
Results describe observed behavior of the publicly accessible interface tested at the stated date and version. They do not represent the vendor's complete production system.
This disclosure will apply to public-agent runs. This page concerns a local controlled experiment.
Read the methodology ↗