1. Transect: retain observability across long runs
Transect, revised 8 October, aligns events, token use, sub-agent activity and model-generated behavior labels on a shared turn timeline. In its 14.5-million-token case study, repeated judge rolls were unanimous for 81.7% and 82.1% of two label sets.
Limitation: agreement is not validity. The case study is one descriptive epoch; causal claims require repeated runs, independent expert labels and explicit statistical analysis.
Architecture implication: retain evaluator configuration, source-turn pointers, judge disagreement and capture gaps in shared evidence. A synthesis layer should expose uncertainty instead of turning one label into fact.
Classification: new observability pattern and substantive support for evidence-qualified judging.
2. WorkflowOps: orchestration can learn from history
WorkflowOps, 6 October, uses first-order handoff frequencies as soft DAG guidance, embeddings as a deterministic capability-matching fast path and an LLM only for ambiguous routing. The paper reports more than 80% fewer LLM routing calls.
The cost profile is important: the reference setup averages 14.8 LLM calls, 43.6k tokens and 605 seconds per query. Agent-pool expansion consumes 1.9× calls and 2.7× tokens for a 1.7-point gain over the cheapest ablation.
Limitation: no long-term test shows whether generated agents remain useful, safe, curated or pruned; the transition matrix is first-order and the evaluated domains are code, math and QA.
Architecture implication: collaboration history can guide coordination as a versioned, reversible and costed prior. It must not become admission, promotion or consequential authority.
Classification: new orchestration-memory pattern with a cost and self-expansion concern.
3. EvalResearchBench: agents can design evaluators, but not own promotion
EvalResearchBench, 3 October, tests nine researcher agents that build frozen evaluators for 13 candidate models against 14 target benchmarks. The best evaluator reproduces about 75% of candidate-pair ordering, below a 91% ceiling set by disagreements among the target benchmarks.
Limitation: a human-designed public-task sample remains strong; development-set winners do not consistently win on sealed targets, and generated evaluators can truncate answers, exhaust budgets or over-weight a few tasks.
Architecture implication: agent-generated evaluation belongs in a candidate branch with sealed validation, coverage and cost receipts. The maker cannot be the final judge or promotion authority.
Classification: substantive support for independent qualification, with a contradiction/concern for recursive self-evaluation.
4. Harness-Aware Distillation: learn what the fixed harness cannot supply
Harness-Aware Distillation, 2 October, contrasts teacher actions with and without harness information and discards preferred targets that contradict harness-observed state. On the reported ALFWorld summary, HAD reaches 57.4% performance and 81.0% harness utilization versus 43.5% and 73.1% for ordinary harness-equipped distillation.
Limitation: the experiments use text environments, students up to 2B parameters and a fixed hand-designed harness. A harness-consistent action can still be wrong.
Architecture implication: distill residual competence needed to use stable harness information, validate teacher targets against observed state and keep deterministic runtime controls outside the learned model.
Classification: substantive support for the existing control-retaining distillation pattern.
Operational changes worth carrying into tests
The Monday watch also found three implementation-level signals. A LangGraph branch-reload fix shows why live and reloaded state must be tested for equivalence: 276 of 464 new fork cases failed before the fix, and legacy-thread and exit-durability limits remain. Inspect AI now retains provider request/response IDs, HTTP status, retries and batch attribution. NVIDIA SkillSpector provides a useful scan/evaluate/sign admission reference for agent skills.
These are implementation signals rather than new constitutional rules. They reinforce branch lineage, provider-call provenance and artifact-bound admission evidence.
Bottom line
The strongest thread is not “more autonomy”. It is better separation: learned coordination from authority, behavior labels from validated observations, model-written evaluators from promotion decisions, and harness-aware learning from runtime enforcement.