Established July 2026 · program statement · oversight manuscript in preparation
The prevailing plan for safe advanced AI is, implicitly, monolithic: train one very capable system and arrange for alignment to live inside it — in the weights, in the training objective, in the constitution it was tuned on. This research area develops the alternative this program's systems work keeps arriving at:
Alignment is a property of composition, not of components. Trust is enforced at the joints — the gates, verifications, and acceptance authorities between parts — and is never assumed to reside inside any single part. Not the models, and not the humans either.
On this view a trustworthy autonomous system is assembled, the way a safe bridge is assembled from members none of which is individually a bridge: the safety case lives in the verified connections. The consequences are what make it a research program rather than a slogan.
Each element of the protocol is stated with its honest status — running, implemented, or in preparation. Nothing below is claimed beyond its receipt.
| protocol element | status |
|---|---|
| Graded oversight as handler substitution — under what conditions the external acceptance authority over an agent's work can be progressively and reversibly delegated while remaining accountable | the oversight manuscript, in preparation (papers) |
| Acceptance records with self-assessment capped at escalation — verdict schema and recorders under which no component accepts its own work | implemented with a passing test suite; standalone release forthcoming |
| External verification of generated claims — logic-engine certification of produced content, with a public worked example run end to end | running; the worked example is published (the film, proved like a theorem) |
| Self-enforcing harness boundaries — the harness track: boundaries the agent holds while working, with the environment enforcing what the framing describes | design complete, experiment pending (contemplative systems) |
| Substrate structure — the context an agent runs on, shaped so that navigation and disclosure are inspectable | published; running code (Paper 1, tools) |
The oldest sustained treatment of the problem this area studies — how a powerful intelligence comes to be trusted — is not in the machine-learning literature. Contemplative traditions maintained working protocols for exactly this: capability conferred in witnessed, graded stages; attainment certified externally, never self-certified; transmission carrying provenance; the whole apparatus revocable. The contemplative-systems research area studies those structures as engineering material; this area is where their load-bearing pattern — trust as a composition protocol — is stated for autonomous systems generally.
This page is a program statement, and its claims ladder is deliberate: the joint-level mechanisms above are running or implemented as stated; the composition of verified parts into larger verified systems is design work in progress; and the limit case — highly capable systems assembled under this protocol — is stated as the program's direction, not as a result. The program's discipline is that the second and third rungs are never described in the grammar of the first.