Windway Data · CC BY 4.0 · Исследовательский выпуск 001
Сбой не бинарен
Локальный контролируемый парный эксперимент
Исследовательский выпуск 001 · Публикация одобрена · Windway Data
| Условие | Проверенное завершение | Ложные заявления о завершении |
|---|---|---|
| Контроль | 18/18 | 0/18 |
| Условие со сбоем | 15/18 | 3/18 |
false_success__01: Ложное завершение в условии со сбоем 3/3; Ложное завершение в сопоставленных контролях 0/3.
Исследовательский интервал неопределённости включает ноль. Общий причинный эффект или рыночная частота сбоев не установлены.
Уровень доказательств: E3 по внутренней системе Windway, только для этой модели, конфигурации и выбранных задач. Исходные наблюдения остаются E1/E2; внешнее воспроизведение E4 не установлено.
Методология и верификатор
Шесть выбранных задач из пяти семейств сбоев, по три новых повтора на задачу и условие. В каждой паре совпадали модель, цели, инструменты, начальное состояние, верификатор и конфигурация. Менялось только внедрение сбоя. Порядок условий был сбалансирован.
Верификатор проверяет устойчивое состояние и требуемую процедуру независимо от заявления агента о завершении. Воспроизведение сравнивает записанные операции и вердикты; неуспешная задача может воспроизводиться корректно.
Поддержка зерна не проверена. Завершение считается вызовом. Бюджет составляет 60 вызовов для обычных задач и 180 для длинных.
Ограничения
Небольшая целевая синтетическая выборка, шесть кластеров задач и одна локальная конфигурация модели. Исследовательский кластерный интервал [0, 0.5] включает ноль. Универсальная причинность, производственная или общая модельная частота сбоев и статистическое доказательство для всего рынка не заявляются.
Инжектор частичного успеха использовал версионированное прерывание процесса после двух из трёх устойчивых записей. Это отличается от исторического возмущения только пакетной операции. Результаты разных экспериментов не объединяются.
Воспроизводимость и доступность
Windway проверила согласованность агрегатов. Необработанные трассы, промпты и закрытые доказательства не публикуются. Полное независимое воспроизведение и повторная генерация моделью не установлены. Загружаемый архив сохраняет исходные метаданные кандидата; текущая веб-публикация одобрена владельцем.
Для проверки агрегатов нужен манифест сопровождающего, не включённый в этот веб-экспорт. Независимый исполняемый комплект воспроизведения не заявляется.
Одобренный снимок агрегатов · Область воспроизведения · Документ о методологии · Документ об ограниченияхЛицензии и атрибуция
Windway Data владеет выбранным собственным контентом. Статья, иллюстрации и документация: CC BY 4.0. Выбранный код: Apache-2.0. Исключённые сторонние веса, зависимости и необработанные сгенерированные материалы не перелицензируются.
CC BY 4.0 · Apache-2.0Cite this research note
Windway Data. (2026). Failure Is Not Binary: Evaluating Recovery, False Completion and Loops in Tool-Using Agents. Research Release 001. https://windwaydata.com/research/failure-is-not-binary/
CITATION.cff · BibTeX · Plain citation · Metadata JSON
Collective author: Windway Data. First web publication: . Article language: English. CC BY 4.0. This technical research note is not peer reviewed.
Исходная статья на английском
Ниже приведена полная одобренная статья на английском. Резюме, методология, ограничения и навигация локализованы; переводы не прошли профессиональную проверку.
Failure Is Not Binary: Evaluating Recovery, False Completion and Loops in Tool-Using Agents
Windway Data research note · 2026-10-05 · Published on windwaydata.com on 2026-10-06; not peer reviewed
Abstract
An agent’s decision to finish is an observable behavior. Whether the required work is complete is a separate question.
In one local Reliability v1 experiment, a Qwen2.5-Coder-14B-Instruct Q5_K_M configuration attempted 30 synthetic tool-use scenarios and explicitly declared completion in 27. An independent state-and-procedure verifier accepted 18 of those declarations and rejected nine. Five rejected declarations had incorrect durable state. Four had the requested record counts but omitted required event processing.
These observations concern this configuration and these inspected tasks. They do not estimate the reliability of Qwen in general, another model, a vendor or a customer system. They suggest a useful evaluation distinction: finish declarations, durable effects, required procedure, fault exposure and termination deserve separate records.
Why binary success is insufficient
AGENT CLAIMED SUCCESS != VERIFIED SUCCESS. A finish signal records the model’s assertion. Windway separately checks final environment state and any required mutation/procedure history. A task can fail after an assertion, fail without an assertion, or end with correct record counts but invalid required procedure.
Experimental systems
This note presents Reliability v1 and a separate ComputerOps AI baseline. Their samples, adapters, tasks and verifier predicates differ. Neither is a customer evaluation; their counts are not pooled into a model-wide score.
Reliability v1
Reliability v1 is a Windway-owned ledger and response-queue simulator. Tasks request exactly-once deliveries, with event dependencies or checkpoints where specified. The model receives public task instructions and documented tool contracts. The independent verifier checks the resulting state and required process invariants against private ground truth.
The inspected run contains 30 scenarios across 15 configured failure classes, with two variants per class and two synthetic task families. Each scenario has one model trial. The model generated 455 responses that produced 428 tool actions; a finish declaration is not an environment tool action.
Configuration: Qwen2.5-Coder-14B-Instruct, Q5_K_M, a local LM Studio/OpenAI-compatible endpoint, context length 8192, temperature 0, requested seed 42, maximum response length 220 tokens. Scenario action budgets are 60 or 180. Requested seed support was not verified. The captured model artifact SHA-256 and adapter/verifier fingerprints belong in the accompanying evidence metadata; the inference-engine build is unavailable. These fields should not be filled with guesses.
The setup is a fully observable synthetic laboratory. It does not reproduce live distributed transport, production permission infrastructure or asynchronous service timing. Fault labels and tool vocabulary are disclosed. The original suite has no matched clean-model arm or repeated trials; the later six-task replication is a distinct controlled experiment.
ComputerOps
A separate internal ComputerOps AI baseline evaluated ten predeclared tasks using the same named Qwen2.5-Coder-14B-Instruct Q5_K_M resource through a bounded text-DOM browser adapter. The agent could choose visible control interactions; it was not a vision or desktop model evaluation. The subset was reserved before model inference from a catalog already exercised by scripted controls, not a newly generated task holdout.
| Outcome | Count |
|---|---|
| Tasks attempted | 10 |
| Verified success | 5 |
| Self-reported success | 8 |
| False completions | 3 |
| Recorded environment-operation replay PASS | 10 |
All five verified successes had a success claim; three other success claims failed verification. Two tasks had no success claim and failed. The three false completions include missing requested effects; one also violated the allowed-mutation predicate. We do not retrofit Reliability’s exact state/procedure subtype taxonomy onto a different verifier.
The final-state and mutation-history checks run without trusting the model’s done success=true response. The ten saved environment-operation replays passed, including failed tasks. This is agreement with recorded operations and verdicts, not repeated fresh model sampling or complete browser/UI replay.
Configuration: context 8192, temperature 0, requested seed 42 unverified, response cap 180 tokens, action budget 32, one admitted complete trial per task. Partial interrupted inference is retained privately but excluded from admitted metrics. The model consumed rendered text and control metadata, not screenshots or native validation bubbles. Runtime discovery established a Bionic service during the resumed phase; the original phase’s executable owner was not recorded. API compatibility must not be confused with proof of standalone LM Studio ownership.
This small, separate benchmark supports the question of self-report versus verified state. It does not establish a market-wide false-completion rate, a ranking, or performance on a prospect’s system.
Controlled replication: a scoped E3 experiment
A prospective paired study followed the exploratory observations. It registered six selected tasks across five fault families, three fresh repetitions per task and condition, for 36 new episodes and 18 matched pairs. Each invocation reset the environment and conversation. Control disabled injection only; goals, tools, initial state, verifier and configuration were matched within each pair. Arm order was counterbalanced across repetitions. These purposive tasks are not 36 independent task types.
| Condition | Episodes | Verified success | False completions |
|---|---|---|---|
| Control | 18 | 18 | 0 |
| Failure treatment | 18 | 15 | 3 |
All 18 treatments encountered their configured failure; all 36 saved-action replays passed. On false_success__01, false completion occurred in 3/3 treatment repetitions and 0/3 matched controls. This is a task- and configuration-specific replicated observation, not a market-wide rate.
The injector was fixed and versioned before inference: a partial-success variant now interrupts the workflow at two unique durable records out of three, whether the agent uses scalar or batch writes. This closes a historical scalar-write bypass. It is not an identical historical batch-only perturbation, so pilot and replication estimates must not be pooled. Regression traces and intermediate oracle checks establish actual partial exposure; they are distinct from model outcomes.
The declared E3 quality gate passed: complete captures, matched configuration and first prompts, fixed model/source hashes, actual treatment exposure, recorded-action replay and disk/prompt-lineage checks. E3 is the proposed Windway evidence designation for this local controlled experiment, limited to these six tasks, the tested model and exact configuration. It is not external certification or E4 reproduction.
The equally weighted task-level treatment-minus-control false-completion estimate was 1/6. The exploratory six-task-cluster bootstrap 95% interval was [0, 0.5], using 6000 replicates. It includes zero. The tiny purposive cluster sample and fixed-temperature repetitions do not support statistical proof of a general causal effect, generalization to another model, production agents or market-wide failure rates. Requested seed behavior remains unverified. Controlled design improves the declared evidence without resolving those inferential limits.
Configuration remains Qwen2.5-Coder-14B-Instruct Q5_K_M, context 8192, temperature 0, requested seed 42 and response cap 220. Ordinary model-call budget is 60, long-horizon 180, counting finish calls. Model artifact and backend/source hashes were pinned; uncontrolled shared-host load limits latency comparisons. No outcome-driven sample expansion occurred.
False completion taxonomy
| Explicit finish declaration | Full verifier passes | Full verifier fails |
|---|---|---|
| Yes | 18 | 9 |
| No | 0 | 3 |
The nine rejected finishes are 9/30 attempted episodes, or 30%. Among the 27 finish declarations, they are 9/27, approximately 33.3%. The first denominator asks how often an attempted episode ended with a false declaration; the second asks how often a declaration was rejected. Neither is a universal model error rate.
State mismatch
Five declarations failed the durable-state predicate. The two tasks configured for partial success made no durable writes: their versionless batch requests failed before the configured partial-success injection was reached. One uncommitted ambiguous-response variant omitted a requested record. Two configured false-success variants each omitted a requested record.
The partial-success cases therefore show protocol competence failures, not demonstrated inability to reconcile an actually executed partial batch. The configured fault label alone cannot establish what the agent encountered.
Protocol violation
Four other declarations failed required event procedure. The duplicate-delivery and out-of-order-event variants had the requested record counts, but the agent bypassed process_event. Their event-processing invariant failed. The configured event faults did not trigger because the response bus was not queried.
Calling all nine cases “wrong final data” would erase this difference. Record counts can be correct while required provenance or procedure is missing. Conversely, a procedure checker is only as comprehensive as its encoded requirements; deterministic verification is not an infallibility claim.
For this analysis we propose three operational measures, each tied to a verifier version:
- Full-verifier false completion: explicit finish and full-verifier failure, divided by eligible attempts—9/30 here.
- Durable-state false completion: explicit finish and durable-state failure—5/30 here.
- Procedure-only false completion: explicit finish, durable-state pass and required-procedure failure—4/30 here.
These are proposed reporting definitions, not adopted industry standards. Report the conditional declaration denominator separately, and keep no-finish outcomes visible.
Recovery needs an exposure denominator
All 30 scenarios had configured failures, but only 22 episodes actually triggered an injection. Sixteen of those reached verified completion. The descriptive recovery count is therefore 16/22 exposed episodes. Overall task completion remains 18/30; two successful episodes did not trigger their configured fault.
The recovery funnel must show attempted tasks, actual injection and verified completion after injection. Explicit fault diagnosis is a different measure: the source records three diagnoses among the 22 exposed episodes. Do not place diagnosis as a mandatory step in a nested recovery funnel, since completion can occur without an explicit diagnosis.
Because some faults never triggered, class-level outcomes cannot be interpreted as a causal estimate of each fault’s effect. The later six-task replication predeclared matched conditions and exposure accounting; it does not retrofit controls onto this original suite. The present analysis does not establish degradation relative to the original Arena baseline, which used a different protocol and task set.
Recovery loops
The current write-focused detector identified five loop episodes. It counts at least three equal write, process or batch intent fingerprints since the last durable delivery, ignoring version and request identifiers. Intervening reads and logical-clock changes do not reset that count. It does not cover every planning or read-only loop.
| Configured class | Tool actions | State reads | Largest repeated intent count in final no-effect segment | Stop |
|---|---|---|---|---|
| partial_success, variant A | 11 | 5 | 5 | false finish |
| partial_success, variant B | 11 | 5 | 5 | false finish |
| tool_schema_drift, variant B | 60 | 20 | 19 | budget exhausted |
| recovery_loop, variant A | 60 | 1 | 14 | budget exhausted |
| recovery_loop, variant B | 60 | 1 | 14 | budget exhausted |
The partial-success tasks repeatedly made versionless batches despite state reads. The schema-drift task reused an obsolete key contract. The endpoint-loop tasks confused a fallback endpoint with resource identity and alternated endpoint-unavailable and missing-resource errors.
Some metadata or resources can change while requested deliveries do not progress. A diagnostic should retain those changes as context without treating every change as successful recovery. The five-loop count belongs to the existing detector; a broader cycle detector would require a versioned reanalysis, not a silent rewrite.
The metric classified 99/428 tool actions as rejected under its enumerated error definition. The remaining 329 are not automatically successful task progress. A rejected call is also not necessarily an unsafe action: these are protocol and execution outcomes under this simulator’s rules.
A separate security experiment exposes a capability confound
Agent Security Bench v0 is a different experiment and must not be pooled with Reliability v1. Its local model run had 19/20 policy-compliant episodes and 0/20 completed legitimate procedures. All episodes were incomplete; one was noncompliant.
Three forbidden send requests in one episode were blocked by an enforcing broker. That establishes the observed enforcement outcome for those requests. It does not establish benign agent intent. Eight episodes encountered an attack payload and none reached the scored injection objective, but important trusted revocation, memory and destructive paths were not exercised by this model.
With no procedural completions, policy compliance conditional on demonstrated completion is undefined. A near-total lack of capability can suppress opportunities to violate policy. A useful next study should first establish clean procedural capability, preserve identical grants, and pair clean and attacked tasks prequalified independently. Selecting only successful attacked episodes afterward would introduce selection bias.
Report capability and compliance jointly. High observed compliance alongside absent task capability cannot be read as robust agent security.
Verifier methodology
The verifier checks environment state and allowed transition/procedure history independently of the agent’s completion signal. Replay re-executes saved operations and compares their results; it can reproduce a failure correctly. Encoded predicates and negative controls constrain the claim, but cannot certify requirements absent from the verifier.
Evidence levels
Under Windway’s proposed E0–E4 terminology, the original Reliability v1 and ComputerOps model outcomes are E1: observed executions with raw provenance, without a fresh-generation reproducibility claim. Pinned recorded-action replay and raw-to-aggregate re-derivation are E2 for those declared evidence objects. This publication preparation reran the Reliability recorded-action replay and checked source hashes. No new model generation was performed. The later controlled replication has scoped local E3 status; the original observations remain E1 for generation and E2 for replay. External E4 reproduction is not established.
Replay can faithfully reproduce a failed outcome. A replay PASS establishes replay agreement, not task success. Likewise, a GOLD trajectory can be a verified negative example. The original Arena dataset is separate and has a known historical prompt-capture defect; it must not be presented as exact prompt/completion training data.
The candidate figures and tables are generated from pinned raw traces, scenario definitions, canonical metrics and synthesis findings. Each has a source/verifier fingerprint, denominator and definition. The private review package includes the raw-to-figure derivation and a recorded-action replay receipt. Windway Data owns the first-party authored code, fixtures, synthetic environments, experiment recordings, reports and datasets supported by repository provenance. Selected licensing and publication approval remain separate fields. Selected article, figures and documentation in this package are licensed CC BY 4.0; selected authored aggregate-checking code is Apache-2.0. No license is applied to excluded third-party weights or raw model text. Third-party model weights, dependencies and model-generated content are tracked separately; ownership is not inferred for those components. Raw private evidence is not included in Release 001. Public full-experiment reproduction remains pending an explicitly approved evidence bundle. Do not treat unavailable evidence as an empty or successful result.
Limitations
Both experiments have small synthetic samples and one admitted trial per task. They use different interfaces and predicates and must not be pooled. Reliability lacks matched clean-model controls; configured faults sometimes never trigger. ComputerOps is text-DOM, lacks image/native-validation inputs and includes a resumed execution. Verifier coverage is finite. Seed support and fresh-generation determinism are unestablished. Only the later six-task controlled replication has scoped E3 status; external E4 evidence is absent. The original Arena prompt-capture defect belongs to that historical corpus and is not silently attributed to these later captures.
Reproducibility
PUBLIC_METRICS.json and PUBLIC_FIGURES contain aggregate observations with source fingerprints. REPRODUCE.md separates public aggregate consistency checks, private source replay and future fresh model generation. Raw prompts, private browser traces and canaries are excluded. Full independent public replay is unavailable until the corresponding evidence is explicitly approved and released. This limitation does not change first-party ownership.
Future work
Beyond the completed local replication, our next question is narrow: under matched clean/fault conditions and a fixed interface, how often does an agent reconcile an uncertain effect, exit repeated failed intent, and finish with both correct durable state and required procedure? The present observations justify measuring those outcomes separately. They do not settle that question across agents.
License: CC BY 4.0. Owner: Windway Data. Publication authorization: approved for windwaydata.com only on 2026-10-06.