1. Benchmark tools are part of the claim

Do Agent Benchmarks Do What They Say?, submitted 29 September, audits the executable contracts behind mutating tools in four agent benchmarks. The authors confirm seven tool defects and one evaluator property across 34 audited tools. On AgentDojo's complete 25-tool mutating surface, at least five tools diverge from the advertised surface as interpreted by the audit.

The finding matters because an evaluator can reward the expected request payload even when the represented state never changed. The checker itself is not a universal solution: it produced no false positives in 25 injected flags, but its dynamic probes missed most covered defects. The study is stronger as an existence proof and audit method than as a cross-benchmark defect-rate estimate.

Harness implication — new pattern and substantive support: treat benchmark tools and evaluators as executable contracts. Preserve the state transition, implementation version and score-at-risk provenance in shared evidence. A benchmark verdict should not cross the qualification gate when the tool or evaluator below it is unqualified.

2. Added harness machinery must earn its place

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?, submitted 30 September, compares several open-source MLE harnesses with a minimal coding-agent session under the same frontier backbone and time budget. In the reported MLE-bench experiments, added planning, search and multi-agent scaffolding did not outperform the minimal session; access to the runtime, shell and filesystem mattered more.

The result is deliberately narrow. It is concentrated on public Kaggle-style ML tasks, some comparisons have wide confidence intervals, and most interventions are single-worker. It therefore does not show that evidence, authorization, isolation or assurance controls are unnecessary.

Harness implication — contradiction/concern: this contradicts blanket expansion of performance scaffolding, not governed execution. Every orchestration layer should show measurable outcome or control value. That supports Least Sufficient Governance: centralized authority, shared evidence and gates remain, while redundant cognitive machinery is removed.

3. Failure reporting needs a structured evidence contract

Failure-Transparent Agents, submitted 28 September, fixes a failed tool observation before response generation so the reporting behavior can be audited directly. Across 100 tasks, six models and 3,600 human-annotated responses, false-success reports fall from 22.8% under the baseline policy to 0.8% with a structured evidence contract. Useful responses increase from 74.9% to 98.8%.

The intervention bundles several requirements, so the experiment does not isolate which component caused the improvement. It is also synthetic, English-only and centered on one-step failed prerequisites rather than partial success, contradictory evidence or long multi-agent trajectories.

Harness implication — new control pattern: persist failed observations in the shared evidence state and bind user-facing completion claims to a typed schema. Synthesis may propose recovery, but it should not convert missing evidence into success; the reporting claim passes through the same qualification logic as the action.

4. Tool errors are model-facing control inputs

MCP Error Messages Written for Developers Hurt the Most Capable Agents Most, submitted 28 September and revised 29 September, finds 949 next-step instructions in 3,001 error messages from 150 widely used MCP servers. Half depend on caller state that the server cannot observe. In tool-only BFCL tasks, terminal or configuration instructions perform poorly, while naming an available recovery tool or the exact call to retry materially improves recovery.

The evaluation covers five OpenAI models in benchmarked tool-only conditions, not every MCP client, model family or hostile server. Even so, it exposes a control boundary that is usually treated as text formatting.

Harness implication — substantive support with concern: retain the raw error as evidence, then normalize the recovery instruction against the authorized tool surface. An error message must not silently widen the model's authority, introduce an unavailable action or bypass the execution gate.

5. External observations compete for governed capacity

When Should Agents Check External State?, submitted 29 September and revised 30 September, models the checks needed by stored intentions as consumers of a shared episode budget. BudgetPM retains 99.9–100% of unconstrained PM-Bench quality with 42–54% fewer observations. Under severe scarcity, its sequential policy matches quality and on-time recall using 16–33% fewer observations than the strongest tested natural monitoring schedule.

The evidence comes from two prospective-memory benchmarks, three backbones and call-count cost models. It does not establish safety under adversarial, rapidly changing or legally time-bounded production conditions.

Harness implication — new memory/context scheduling pattern: make observation capacity an explicit governed resource and enforce it outside the model. Optimization may choose among discretionary checks, but safety, legal, freshness and high-consequence verification requirements need non-optimizable minimum floors.

Release view

Harness Architecture v1.3.0 was finalized this week. Its MAIN additions govern effective deployed capability, workflow progression, inherited policy, skill evolution, external-server trust, cross-layer evidence and outcome-based assurance.

No additional semantic-version increment is justified by the findings above. They sharpen evaluation, reporting, tool-recovery and observation controls, but do not incompatibly change the v1.3.0 authority model.