# Methodology

AGENT CLAIMED SUCCESS != VERIFIED SUCCESS.

Preserve task/configuration, model output, tool action, observation, committed state, required procedure, finish signal and independent verifier verdict as distinct objects. Verify requested final state and allowed mutation/procedure history against an owned fixture oracle; do not use model prose as the success predicate. Document what the verifier cannot check.

Reliability v1 checks exact deliveries and event/checkpoint requirements. Its explicit finish signal is independent of full verification. Configured fault and actual trigger are distinct: only exposed episodes enter the descriptive recovery denominator. Five state-mismatch and four procedure-only false finishes are separate subtypes; correct record counts alone cannot establish required event provenance. The existing write-intent loop detector counts repeated intent since durable progress and does not cover all read-only loops.

ComputerOps independently checks requested/strict state and allowed mutation history. Its model uses current text-DOM controls; recorded environment operations are replayed without new model sampling. Three false success claims are verified failures under ComputerOps predicates, not automatically assigned Reliability subtypes.

Evidence terminology: E0 unexecuted design; E1 observed execution; E2 reproduced declared evidence object with mode named; E3 prospectively controlled matched repeated experiment; E4 external independent reproduction. These outcomes are E1 for model generation and E2 for internal saved-action/state audit. No broad project upgrade,no pooled model score;controlled replication is separately scoped E3.

Original evidence is preserved. Aggregates reference source/verifier hashes. Missing runtime metadata remains missing. Exact capture and model weight fingerprints improve provenance without proving universal model identity or trustworthy timestamps. First-party ownership, selected license and publication approval are separate records.


## Controlled replication update — scoped E3

36 fresh episodes across six tasks,three repetitions per condition:control18/18 verified and0/18 false completion;treatment15/18 verified and3/18 false completion;36/36 recorded-action replay PASS. `false_success__01` reproduced false completion3/3 in treatment and0/3 in matched control.

Local E3 applies only to this prospectively controlled experiment,not retroactively to observational Reliability or ComputerOps. The exploratory task-cluster95% interval [0,0.5] includes zero;no statistical proof of a general causal effect. No pooling,other-model,production-agent or market-wide generalization. The partial injector was versioned from historical batch-only behavior to workflow interruption after two of three durable records.

Selected authored content license:CC BY4.0. Publication approval remains pending. Raw private evidence stays excluded.
