Research area: compositional alignment

Established July 2026 · program statement · oversight manuscript in preparation

The prevailing plan for safe advanced AI is, implicitly, monolithic: train one very capable system and arrange for alignment to live inside it — in the weights, in the training objective, in the constitution it was tuned on. This research area develops the alternative this program's systems work keeps arriving at:

Alignment is a property of composition, not of components. Trust is enforced at the joints — the gates, verifications, and acceptance authorities between parts — and is never assumed to reside inside any single part. Not the models, and not the humans either.

On this view a trustworthy autonomous system is assembled, the way a safe bridge is assembled from members none of which is individually a bridge: the safety case lives in the verified connections. The consequences are what make it a research program rather than a slogan.

Four consequences

  1. Model-agnostic by construction. If safety was never a property of the model, the protocol composes with any model — including untrusted ones, including future ones. The systems built under this program treat every component, model or human, as a party whose work is accepted by an authority external to it.
  2. Trust is earned per-domain, witnessed, reversible — and never transferred. An agent's autonomy is a portfolio of graded, per-domain permissions, each earned by a record of externally verified work in that domain, each revocable. Competence demonstrated in one domain confers nothing in another. Self-assessment is structurally capped: a component's own verdict on its own work can at most escalate for review — it can never accept.
  3. Failures are preserved, not hidden. Every acceptance and rejection is recorded as a first-class artifact. The record of failures — kept, examined, and folded into the next configuration — is the foundation the system's reliability stands on, and the corpus from which acceptance authority can later be safely delegated.
  4. Composition is itself gated. Parts join into larger systems only through explicit acceptance events; a composite built from unverified parts is treated as a type error, not as a risk to be managed downstream.

The mechanism inventory

Each element of the protocol is stated with its honest status — running, implemented, or in preparation. Nothing below is claimed beyond its receipt.

protocol elementstatus
Graded oversight as handler substitution — under what conditions the external acceptance authority over an agent's work can be progressively and reversibly delegated while remaining accountablethe oversight manuscript, in preparation (papers)
Acceptance records with self-assessment capped at escalation — verdict schema and recorders under which no component accepts its own workimplemented with a passing test suite; standalone release forthcoming
External verification of generated claims — logic-engine certification of produced content, with a public worked example run end to endrunning; the worked example is published (the film, proved like a theorem)
Self-enforcing harness boundaries — the harness track: boundaries the agent holds while working, with the environment enforcing what the framing describesdesign complete, experiment pending (contemplative systems)
Substrate structure — the context an agent runs on, shaped so that navigation and disclosure are inspectablepublished; running code (Paper 1, tools)

Prior art

The oldest sustained treatment of the problem this area studies — how a powerful intelligence comes to be trusted — is not in the machine-learning literature. Contemplative traditions maintained working protocols for exactly this: capability conferred in witnessed, graded stages; attainment certified externally, never self-certified; transmission carrying provenance; the whole apparatus revocable. The contemplative-systems research area studies those structures as engineering material; this area is where their load-bearing pattern — trust as a composition protocol — is stated for autonomous systems generally.

Scope note

This page is a program statement, and its claims ladder is deliberate: the joint-level mechanisms above are running or implemented as stated; the composition of verified parts into larger verified systems is design work in progress; and the limit case — highly capable systems assembled under this protocol — is stated as the program's direction, not as a result. The program's discipline is that the second and third rungs are never described in the grammar of the first.