The v1.4.0 control path
The release connects software delivery and runtime control instead of treating them as separate governance problems:
- Intent and specification: define the outcome, constraints, acceptance criteria and non-goals.
- Contracted execution unit: bind the model to the harness, tools, state, artifacts, evidence and failure behavior it will actually use.
- Distributed coordination: specialists may discover, self-assign and work in parallel inside bounded capability and budget envelopes.
- Executable verification: tests and independent checks attempt to falsify the completion claim and compare the result with the active specification.
- Progression authority: an external decision decides whether the verified work may move toward effect.
- Pre-effect enforcement: current identity, delegated authority, policy, target, data, environment, time and budget are checked outside agent-writable state.
- Observed-state reconciliation: runtime and provider evidence show what changed; drift or missing evidence reopens qualification.
1. Compose model–harness pairs, not model names
Two agents are not interchangeable because they can both return text. Their effective behavior depends on the model, prompt and context policy, tool surface, state semantics, artifact contract, evidence output, retry policy and execution environment.
v1.4.0 therefore treats the model–harness pair as the execution unit. A planner may propose a worker, but admission requires contract compatibility. A valid DAG does not prove that artifacts, state transitions or failure semantics line up.
Raven provides research evidence for composing heterogeneous harnesses. It supports the contract problem; it does not establish universal reliability or transfer authority to a host router.
2. Decentralize coordination without decentralizing authority
Central consequential authority does not require one scheduler to mediate every collaboration event. Agents can discover tasks, self-assign work, coordinate through a shared workspace and parallelize execution.
The authority topology remains different: proposed results enter durable shared evidence, pass independent qualification and reach an explicit gate before an external effect. Agensh is useful evidence for large-scale organizational coordination, but scale and consensus do not create authorization.
3. Let the model edit working context, not governance state
Models can increasingly decide what to summarize, compact, retain or discard. That can improve long-running work, but it also creates a dangerous category error if working summaries become the source of truth.
v1.4.0 separates model-editable working context from immutable source evidence, provenance, policy and authority state. The same rule applies to structural code memory: derived repository maps must remain bound to the exact source revision and freshness state from which they were built.
Context Language Models and AutoCompact support adaptive context mechanics. They do not justify allowing a model to rewrite the evidence or permission basis for a high-consequence decision.
4. Make the specification durable and convergence explicit
Spec-Driven Development is now a built-in, vendor-neutral process profile for software, configuration, policy and infrastructure changes. The minimum path preserves intent, specification, plan, tasks, implementation evidence, verification, convergence and promotion.
GitHub Spec Kit is primary implementation evidence for carrying a specification into planning, tasks and implementation. It is not mandatory. The architecture requires semantic traceability, versioned mutation and risk-adaptive depth rather than one tool or document format.
5. Verification must be able to prove the agent wrong
An agent’s own completion statement is not evidence that the work is correct. The executable verification harness selects checks capable of falsifying the claim: schema and type checks, contracts, properties, architecture and policy tests, integration and end-to-end evidence, differential checks, independent verifiers and, where necessary, accountable human review.
The depth is proportional to consequence and uncertainty. More checks are not automatically better, and a check written by the maker is not independent ground truth. Verification produces evidence for a release decision; it does not issue the decision itself.
6. Enforce before effect, then verify the effect
For consequential actions, policy must be evaluated at an enforcement point outside prompts, agent-editable files and model-generated code. Tool availability, a valid credential or a low semantic-risk score is not authorization.
After an allowed action, v1.4.0 binds the authorized intent to execution identity, provider/runtime events, observed resource state and drift reconciliation. This closes the gap between an approved plan, a successful API call and the state that actually exists.
Where the credible consequence window is shorter than a human-only response path, preventive or containment controls must already be positioned at the enforcement point. Vendor security advisories can also trigger qualification before CVE and scanner ecosystems converge.
7. Learn residual competence; keep controls external
Harness-Aware Distillation adds a useful training pattern: when the harness stays fixed, train the smaller model on what the teacher contributes beyond that harness and reject targets that contradict harness-observed state.
The boundary matters. Better harness utilization is not control effectiveness, and learned response to harness information is not independent enforcement. Runtime controls remain in place until separate qualification justifies any change.
What did not change
The constitutional architecture remains stable. Consequential decision authority is explicit; distributed cognition and coordination do not inherit it; source evidence and authority state stay durable; models, tools and frameworks remain replaceable; and promotion or execution still depends on current policy, qualification and accountable decision rights.
That makes v1.4.0 a MINOR release: it adds backward-compatible contracts and control branches without breaking the v1.3.0 authority and evidence model.
Claim boundaries
- Architecture decisions: contracted model–harness units, authority-preserving coordination, governed editable context, built-in SDD, executable verification and pre-effect enforcement are decisions in this reference architecture.
- Research evidence: Raven, Agensh, Context Language Models, AutoCompact and Harness-Aware Distillation report results in their evaluated settings; they are not independent production assurance.
- Implementation evidence: GitHub Spec Kit demonstrates one SDD mechanism; it is not a mandatory dependency.
- Assurance boundary: specifications, tests, signatures, telemetry and receipts contribute evidence but do not themselves authorize deployment or prove safe effect.