# Limitations

- Separate internal benchmarks; no customer or market-wide evaluation. Samples must not be pooled into an accuracy or false-completion rate.
- One admitted model realization per task, small synthetic task families and a disclosed interface. The original observations lacked matched repetition;later replication has scoped local E3,no E4.
- Reliability configured faults sometimes never trigger. Partial/event failures can reflect contract failure or path bypass. Observed associations do not isolate causal fault effects.
- Current loop detection is write-focused. State reads and metadata changes do not necessarily constitute meaningful recovery.
- ComputerOps uses text-DOM metadata, not vision/desktop perception, and excludes native validation bubbles/screenshots from model input. A catalog prevalidated by scripted controls is not a task-generation holdout.
- ComputerOps continued after interrupted transport; complete admitted episodes and excluded partial responses are distinguished. Runtime executable ownership was measured only during continuation, not the initial phase.
- Verifiers cover encoded state/procedure predicates. Broader task semantics, useful language understanding and real asynchronous/production behavior remain outside these claims.
- Recorded-action/environment replay is not fresh model generation or complete browser/UI replay. A failure can replay perfectly.
- Requested seed support is unverified. Missing engine/build fields are not reconstructed.
- Public aggregates and figures do not supply full experiment reproducibility; raw evidence stays private without explicit approval. This is a release-scope choice, not an assertion of unclear first-party ownership.
- Security Bench's high compliance with zero procedural completions cannot establish robust capable security. Zero-completion conditional compliance is undefined.

Future work: matched clean/fault arms, repeated executions, independently prequalified capability and broader task families with versioned predicates/exposure analysis.


## Controlled replication update — scoped E3

36 fresh episodes across six tasks,three repetitions per condition:control18/18 verified and0/18 false completion;treatment15/18 verified and3/18 false completion;36/36 recorded-action replay PASS. `false_success__01` reproduced false completion3/3 in treatment and0/3 in matched control.

Local E3 applies only to this prospectively controlled experiment,not retroactively to observational Reliability or ComputerOps. The exploratory task-cluster95% interval [0,0.5] includes zero;no statistical proof of a general causal effect. No pooling,other-model,production-agent or market-wide generalization. The partial injector was versioned from historical batch-only behavior to workflow interruption after two of three durable records.

Selected authored content license:CC BY4.0. Publication approval remains pending. Raw private evidence stays excluded.
