WINDWAY DATA / LOCAL RESEARCH
A narrow experiment. A checkable protocol.
Experimental · Not a standard · Not peer reviewed
Protocol & denominator
50 pinned layouts × 3 failure categories × 3 policies × 2 seeds = 900 cells per invocation. Repeated twice: 1,800 executed trials. Each policy has 300 cells per invocation; there are 150 scenario designs, not 1,800 independent problems.
- Reset the pinned fixture and execute its reference prefix where required.
- Inject a typed, semantically invalid action and require the expected rejection.
- Run the policy until target completion or budget exhaustion.
- Verify transitions and the original exact-cost goal oracle; replay the complete sequence.
Action budget: min(128, max(16, remaining shortest-path length + 8)). Seeds 0 and 1 affect random walk only. Repeated seeds for deterministic policies are controls.
Evidence retained
Inputs, expected target and cost, actual states, actions, observations, budgets, metrics and trace hashes. A separate audit recomputes costs with Bellman–Ford rather than the runner’s Dijkstra planner.
900/900 traces replayed, eight metrics recomputed and 900 corrupted traces rejected by the audit. The runner separately rejects 100 wrong-cost or forged-transition controls.
Reproduction boundary
Measured on Windows 11 with CPython 3.13.5 using the standard library. Cross-platform reproduction has not been measured. This page exposes a sanitized summary; the full local research package is not a public download.
Declared source hashes match. An extra unmanifested cache file in the historical source snapshot prevents a clean full-inventory claim; that snapshot was not repaired.
Limits on interpretation
Failure rejection does not modify the map. Full visibility and exact-cost acceptance make the planner a positive control. No confidence interval, cross-model ranking or commercial improvement is claimed.
Inspect the sanitized evidence summary ↗