1. Transect: retain observability across long runs

Transect, revised 8 October, aligns events, token use, sub-agent activity and model-generated behavior labels on a shared turn timeline. In its 14.5-million-token case study, repeated judge rolls were unanimous for 81.7% and 82.1% of two label sets.

Limitation: agreement is not validity. The case study is one descriptive epoch; causal claims require repeated runs, independent expert labels and explicit statistical analysis.

Architecture implication: retain evaluator configuration, source-turn pointers, judge disagreement and capture gaps in shared evidence. A synthesis layer should expose uncertainty instead of turning one label into fact.

Classification: new observability pattern and substantive support for evidence-qualified judging.

2. WorkflowOps: orchestration can learn from history

WorkflowOps, 6 October, uses first-order handoff frequencies as soft DAG guidance, embeddings as a deterministic capability-matching fast path and an LLM only for ambiguous routing. The paper reports more than 80% fewer LLM routing calls.

The cost profile is important: the reference setup averages 14.8 LLM calls, 43.6k tokens and 605 seconds per query. Agent-pool expansion consumes 1.9× calls and 2.7× tokens for a 1.7-point gain over the cheapest ablation.

Limitation: no long-term test shows whether generated agents remain useful, safe, curated or pruned; the transition matrix is first-order and the evaluated domains are code, math and QA.

Architecture implication: collaboration history can guide coordination as a versioned, reversible and costed prior. It must not become admission, promotion or consequential authority.

Classification: new orchestration-memory pattern with a cost and self-expansion concern.

3. EvalResearchBench: agents can design evaluators, but not own promotion

EvalResearchBench, 3 October, tests nine researcher agents that build frozen evaluators for 13 candidate models against 14 target benchmarks. The best evaluator reproduces about 75% of candidate-pair ordering, below a 91% ceiling set by disagreements among the target benchmarks.

Limitation: a human-designed public-task sample remains strong; development-set winners do not consistently win on sealed targets, and generated evaluators can truncate answers, exhaust budgets or over-weight a few tasks.

Architecture implication: agent-generated evaluation belongs in a candidate branch with sealed validation, coverage and cost receipts. The maker cannot be the final judge or promotion authority.

Classification: substantive support for independent qualification, with a contradiction/concern for recursive self-evaluation.

4. Harness-Aware Distillation: learn what the fixed harness cannot supply

Harness-Aware Distillation, 2 October, contrasts teacher actions with and without harness information and discards preferred targets that contradict harness-observed state. On the reported ALFWorld summary, HAD reaches 57.4% performance and 81.0% harness utilization versus 43.5% and 73.1% for ordinary harness-equipped distillation.

Limitation: the experiments use text environments, students up to 2B parameters and a fixed hand-designed harness. A harness-consistent action can still be wrong.

Architecture implication: distill residual competence needed to use stable harness information, validate teacher targets against observed state and keep deterministic runtime controls outside the learned model.

Classification: substantive support for the existing control-retaining distillation pattern.

Operational changes worth carrying into tests

The Monday watch also found three implementation-level signals. A LangGraph branch-reload fix shows why live and reloaded state must be tested for equivalence: 276 of 464 new fork cases failed before the fix, and legacy-thread and exit-durability limits remain. Inspect AI now retains provider request/response IDs, HTTP status, retries and batch attribution. NVIDIA SkillSpector provides a useful scan/evaluate/sign admission reference for agent skills.

These are implementation signals rather than new constitutional rules. They reinforce branch lineage, provider-call provenance and artifact-bound admission evidence.

Bottom line

The strongest thread is not “more autonomy”. It is better separation: learned coordination from authority, behavior labels from validated observations, model-written evaluators from promotion decisions, and harness-aware learning from runtime enforcement.