An additive architecture release, not a platform certification.
Central decision authority, distributed cognition, shared evidence, synthesis and appropriate independent judging remain intact. The change adds proportional control profiles and stronger end-to-end evidence.
1. Control the trajectory and justify its depth
Individually permitted actions can compose into a prohibited outcome. Step authorization therefore needs trajectory state: identities, parent steps, destination changes, accumulated privileges, retries, compensation, external effects and budgets. Operational, policy, outcome and engineering observability answer different questions. A successful task and zero logged violations do not prove that the user's intent was safely satisfied.
The seven-rule Constitution compresses MAIN: centrally governable authority; effect-bounded grants; durable state outside model context; evidence-based qualification; enforceable boundaries; promotion-controlled learning; economically proportional governance. A source that reinforces a rule normally updates its evidence, not the list of primitives. This is a Harness design decision, not a claim of scientific optimality.
H0–H3 profiles select reasoning, verification and human-review depth after binding obligations and hard consequence floors. Classify decision latitude first: V (Verbatim), E (Equivalent), RP (Risk-Proportional), A (Appetite-Driven). Mandatory controls cannot be optimized away. Track total lifecycle cost, human attention, latency, retry waste, Governance Tax and cost per accepted outcome.
Todd Tucker's FAIR Institute analysis, 16 September, supports treating risk appetite as a business constraint rather than the sole objective. Its sponsored customer-benefit example is not quantified universal risk-reduction evidence. Harness infers a practical decision record separating risk, assurance and operational value, with accountable owners and no double counting.
- Central authority
- Bounded parallel cognition
- Task-scoped context and tools
- Durable evidence
- Memory provenance
- Synthesis and appropriate judge
- Execution gate
- Step and trajectory policy
- Allow, deny, pause, escalate, terminate
- Outcome assurance
- Observe effects and side effects
- Verify, reconcile, promote qualified learning
2. Admit tools by relevance and close remediation by evidence
Authorization answers whether a tool may execute. Relevance answers whether its definition should shape the current model context. Keep the authorized registry outside resident context where feasible; admit a small task-relevant candidate set just in time. Regression tests must include matched no-tool cases and necessary-tool utility, not selection accuracy alone. This refinement is supported by the new unnecessary-tool study in the reading list below.
The Continuous Verified Remediation Loop distinguishes validated, disproven and inconclusive findings. Failed environment setup is not disproval. Separate candidate generation from verification; obtain accountable ownership, independent fix checks and execution authority; retest the deployed state before closure. OpenAI's Defense Factory provides a concrete control-plane/data-plane case study with isolated reproducible environments and post-deployment checks. It is first-party implementation evidence, not independent assurance or a mandatory vendor choice.
The normalized detection-response branch adds normalize → enrich → detect → prioritize → investigate → authorize → respond → verify → tune. Google SecOps architecture is one implementation reference. Normalized telemetry is not automatically decision-ready evidence; provenance, confidence and action authority still matter.
ENISA's 11 September SRP launch operationalizes an existing regulatory reporting path. Preserve awareness/deadline state, applicable reporter role, guidance versions, approval, submission and closure evidence. Manufacturers' reporting applicability and later steward applicability must not be conflated. Submission remains a consequential authorized action; this is not new legislation.
The adopted conformance direction is invariant → contract → enforcement point → telemetry → positive/negative test → evidence. Architecture, Conformance Suite and future Platform Contract versions are separate. A deterministic depth-selector prototype and eight unit-test functions now exist; the complete conformance suite, reference stack and production operating-effectiveness proof remain work to do. Open PR #190 is outside this release.
3. Reading shortlist: new primary material, 11–18 September
Monday's Unified Weekly Watch informed discovery; primary artifacts and subsequent changes were checked. Two items meet the practical-relevance and date requirements. No unchanged recommendations from the previous release are repeated.
When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering — 12 September, v1. Across 500 query pairs, ten single-tool domains and six models, pooled answer rate falls 98.2%→63.5% with unnecessary tools, even without invocation. One trial per condition, no confidence intervals and model judging limit generalization. Substantive support with concern. Inferred delta: relevance-scope exposure and separately test correctness and tool utility. Already reflected in Harness context admission; useful because it gives a direct regression target.
OrcaReplay: harness-native agent structure capture — 14 September, following schema additions on 12 September. Paired proxy/SDK-span capture exposes missing handoff-origin identity and non-network guardrails; schema 0.2.0 and capture/tests address this. Project-local tests and bounded SDK coverage do not prove universal completeness. Substantive support. Inferred delta: correlate native agent/policy events with tool, proxy and effect evidence, explicitly retaining capture gaps. Worth reading to avoid treating network replay as complete shared evidence or independent judging.
Release notes and source map · Architecture changelog · Previous v1.0.0 release. Repository access may be required for the detailed evidence records.