WINDWAY DATA

EVALUATION EVIDENCE

Observe behavior.
Check the evidence.

LIVE EVALUATION NETWORK

Public Agent Evaluations

Awaiting verified public runs

No public-agent metrics are published yet. Local research results are listed separately below.

Targets evaluated
0
Scenarios executed
0
Agent turns
0
Failures observed
0
Reproducible failures
0

LOCAL RESEARCH / WFRB v0

50Layouts
150Scenario designs
1800Executed local trials

900 experimental cells repeated once. Local control policies only; no LLM or customer agent tested.

Explore evaluations ↗

Windway controlled benchmark

Local policy

Planner

300/300

Exact-cost recovery in one local invocation. No third-party agent score.

View evaluation ↗

Local policy

Random

4/300

Exact-cost recovery in one local invocation. No third-party agent score.

View evaluation ↗

Local policy

Blind repetition

0/300

Exact-cost recovery in one local invocation. No third-party agent score.

View evaluation ↗

WFRB v0 / LOCAL CONTROL EXPERIMENT

Experimental · Not a standard · Not peer reviewed

These charts describe local control trials. They do not measure public agents, LLMs or customer systems.

Loading verified aggregate data…

Recovery by category

Pass / fail / inconclusive

Injected failure class distribution

Evaluations over time

Deterministic replay verification

Target comparison

Customer target comparison is unavailable. These local controls do not support a cross-model ranking.

Inspect local benchmark evidence ↗

Scenario heatmap

Tap or expand a cell to inspect its exact denominator. Each layout/category contains two seeded trials per local policy in one invocation.

PASS means exact-cost recovery. FAIL means that recovery criterion was not met. Replay verification does not establish reproduction of a customer-agent failure.

Read the methodology ↗ · Inspect aggregate data ↗

Evaluate your agent ↗ · Become a design partner ↗