1. Model completion is not workflow completion

An agent saying “done” should not, by itself, advance a consequential business process. v1.3.0 separates cognition from progression authority: a model may propose or interpret, but a governed state transition occurs only after typed output and relevant evidence pass an external qualification gate.

This is the same architectural idea behind deterministic approval and execution controls, but applied one step earlier. The question is no longer only whether an effect is authorized. It is also whether the procedure is actually entitled to move into the next state.

2. Govern the effective deployed system

A model can remain unchanged while the effective system changes materially. Tool access, context, credentials, agent composition, reasoning budget, runtime configuration or supplier dependencies can alter what the system is capable of doing.

v1.3.0 therefore treats material non-model capability changes as reassessment triggers. Capability declaration, authorization, risk discovery, assurance evidence and production approval remain separate decisions.

This matters for skills too. Microsoft SkillOpt provides a concrete research/reference implementation in which natural-language skill documents are optimized while the base model stays frozen. The useful architecture lesson is not “self-edit skills automatically”; it is the opposite: candidate generation, held-out validation, adoption and production authority should remain separate. The linked research preprint is evidence for the mechanism and reported experiments, not independent proof of production effectiveness.

3. Effective policy matters more than configured policy

Central policy administration can reduce duplication, but a root policy is not evidence that every resource is correctly covered. The architecture now distinguishes policy definition, inheritance, resolved/effective policy, enforcement and evidence.

Databricks Unity Catalog metastore-level ABAC is a useful implementation reference because it exposes hierarchical inheritance and effective-policy inspection. It remains vendor documentation for a beta feature, not independent assurance and not a universal ABAC requirement.

The general rule is portable: configured policy ≠ effective policy ≠ control effectiveness. Classification quality, resource coverage, identity context and actual enforcement still need evidence.

4. Interoperability does not establish trust

Tool protocols solve discovery and invocation problems. They do not automatically establish that the remote server, operator, deployed artifact, domain or infrastructure is an approved enterprise dependency.

v1.3.0 adds an explicit external-server governance boundary and a second rule: discovery metadata cannot be the sole authority for establishing the identity that will receive credentials or privileged material. Fallback and compatibility paths should preserve the same trust property as the primary path rather than silently weakening it.

5. Security testing has to follow the deployed path

Testing only the model endpoint misses the system that creates real effects. The AI Exchange security-testing guidance supports a broader test surface: reasoning, tools, infrastructure, orchestration, session state, retrieval and deterministic enforcement boundaries.

v1.3.0 turns that into a coverage discipline: declare what was tested, preserve production-path parity, test deterministic boundaries directly, repeat probabilistic tests where appropriate, report untested surfaces explicitly, and retest after remediation. The source describes itself as discussion/community guidance, so it informs engineering methodology rather than creating a legal or certification requirement.

6. Evidence has to cross layers

Valid credentials are evidence, not a verdict. Network telemetry is evidence, not a verdict. Application logs are evidence, not a verdict. The architecture now requires material signals from identity, application, workload and network layers to be correlatable in a shared evidence state before higher-confidence decisions are made.

The ANSSI publication on the DGFiP attacks is a useful incident-level illustration: stolen legitimate credentials, exposed applications, telemetry gaps and uncorrelated signals can coexist. The architectural conclusion is not that every environment needs every possible log, but that a material missing evidence source must be represented as missing rather than interpreted as a clean signal.

7. Assurance maps are not assurance conclusions

The release also strengthens the connection from obligations and organization-level requirements to local risks, controls, assurance work and decisions. The key is to preserve source authority, applicability, evidence quality, independence and gaps rather than flattening everything into one score.

HM Treasury's Orange Book provides authoritative risk-management and assurance-mapping guidance within its UK central-government scope. Its broader lesson is useful elsewhere only as contextual guidance: assurance coverage should expose gaps, overlaps, source quality and independence; a completed map is not itself proof that controls are effective.

Outcome-based resilience follows the same logic. An organization should be able to show that critical functions achieve the required outcome under the relevant risk context, not merely that a framework has been mapped or a control exists on paper.

What did not change

The core architecture remains stable: consequential decision authority stays explicit; distributed cognition does not imply distributed authority; evidence is durable and provenance-aware; tool and model choices remain replaceable; and promotion or execution still depends on current policy, qualification and accountable decision rights.

Version 1.3.0 is therefore an additive release. It closes control gaps around system change, state progression, policy resolution, supplier trust and assurance without turning every reference implementation into a mandatory dependency.

Claim boundaries

  • Architecture decisions: progression authority, effective-system reassessment, effective-policy introspection, trust-anchor-before-discovery and cross-layer evidence correlation are design decisions in this reference architecture.
  • Implementation evidence: SkillOpt and Databricks demonstrate concrete mechanisms; they do not prove the architecture is universally optimal.
  • Practitioner/community guidance: AI Exchange testing guidance informs coverage design but is not a statutory or ISO requirement.
  • Government evidence: ANSSI and HM Treasury sources retain their incident, jurisdiction and scope limitations.