Status: working paper — not peer-reviewed. Code: github.com/sancovp/skilltree (MIT).
Practitioners building tool-using agents report a counterintuitive failure: connecting an agent to more tools can make its selection worse, not better.6 A prominent response has been to migrate from MCP toward "skills."5 But skills are shipped flat—a directory of self-contained packages, each disclosing only its own files—while skills are used relationally: to apply one correctly a model must place it among the others (general/specific, part/whole, dependency). That relational structure is latent in the library—a fact about the skills, not about how they are stored—and a flat list discards it, forcing the model to reconstruct it on every task. skilltree makes that structure explicit: a coordinate-addressed tree (in general, a graph) the agent navigates instead of reassembling, demonstrated with tools that install in one command. The claim is substrate-correction, not agent-performance: we supply the structure the disciplined folder approach already reaches for, so a reader should expect results of that kind—delivered by the structure rather than by hand—not a new performance record. The corresponding agent-task benchmark is future work this paper specifies, not a result it claims.
The field's guidance moved twice. First, away from programming agents: the most successful builds were found to use "not… complex frameworks or specialized libraries, but instead… simple, composable patterns."14 Then, away from building agents at all — rather than a new agent per use case, give one general agent domain-specific skills: folders of procedural knowledge loaded on demand,720 a shift popularly summarized as "don't build agents, build skills."
That second move is the one this paper examines, because as stated it is incomplete. A skill can carry an agent's knowledge, but not its behavior across a skill library. The moment there are enough skills to require navigation, the platform's own progressive disclosure does not extend across them (§6): a skill cannot cause a deeper skill to be injected. Exactly two things restore that traversal. One is a harness—orchestration code wrapped around the skills—which is the programming the field had just spent a release escaping. The other is a graph: a structure in which progressive disclosure applies to itself, so the library is navigable without a program wrapped around it. That is the whole point: the skills move only completes once progressive disclosure closes on itself, and that closure is structural, not programmatic. Everything below is the cheapest such structure—a tree on the filesystem (§3–§8)—and the substrate it approximates—a graph (§9).
To state the object precisely. A skill package is one local patch of procedural knowledge; a skill library is the set of available packages. The library has a latent semantic structure: its packages stand in typed relations—general/specific, part/whole, dependency, sequence—to one another whether or not those relations are encoded anywhere; the structure is a fact about the skills, not about how they happen to be stored. That structure is in general a graph (a skill can belong under more than one parent), and a tree is a runnable projection of it—one chart, selected for a purpose (§9). To use a skill correctly a model must reconstruct this structure from whatever it is handed. A flat list hands it the leaves and discards the structure, so the model re-derives it on every task; a tree (one projection, made explicit and navigable) or the graph (the structure itself) supplies it in advance. The whole contribution follows in one line: the structure the model needs already exists in the library, the platform ships it flat, and skilltree is the program that makes a runnable projection of it explicit and enforceable—so the model navigates it instead of reassembling it.
Most builders treat the arrangement of context—what to load, in what order, how to chunk it—as ergonomics: a matter of taste and token budget. It is not. Consider what an agent is while it runs: a trajectory through a sequence of context-states, each transition chosen from what the current state makes available. Modeled this way, the agent is a flow, and the space it flows over is precisely the structure of its context—nodes are states, edges are admissible transitions. The arrangement of context is therefore not decoration on top of the agent; it is the space the agent moves in. A flat substrate offers only an unstructured state space—every sibling available, none declared as following another—so at every step the agent must reconstruct which move can follow which. A tree or graph supplies those paths in advance: what can follow from what, what excludes what, where the next descent goes. At the moment of use the structure the agent commits to is acyclic—an unresolved cycle is not a decision procedure—which is why a tree, or a DAG, is the operative shape and not an arbitrary graph.
This is the consequence the field has been paying for without naming. Structure supplies the flow with its admissible moves; flatness withholds them—and withholds them silently, because the model still emits fluent output, merely without the structure it would otherwise exploit. The claim is not metaphorical: flatten the context far enough and the space the agent was supposed to move through is gone.
Why is there structure to reconstruct in the first place? Because meaning is not constituted as a list. A list merely orders items; meaning comes from the typed relations among them—containment, specialization, contrast, dependency, sequence, inheritance, transformation—and those relations are tree- and graph-shaped. They are present in a skill library whether or not the directory encodes them: meaning can be serialized as a list, but it is not constituted as one. A model, moreover, carries those relations only softly—attention-mediated and implicit—and cannot make them durable and explicit unless the substrate does. A flat list therefore hands the model exactly the serialization that has dropped the relations that gave the items their meaning, and asks it to recover them on every pass.
The structural claim of §1 stands on its own; this subsection offers two reasons to expect it, neither of which the argument requires. First, a plausible lower-level account: a leading mechanistic study of in-context learning identifies induction heads—attention circuits that, having attended to [A][B] earlier in a sequence, complete a later [A] with [B].1 The evidence is strongest in small attention-only models and more indirect in large ones; we take from it only the shape of the operation, A→B, a directed edge—and edge-shaped machinery has little to lock onto in an edgeless substrate. Second, and independent of any mechanism, models are documented to degrade when relevant content is buried among irrelevant content2: the same loss seen from the outside. Treat both as corroboration of the structural claim, not as its foundation.
This would be a private matter of taste if these tools were products. They are infrastructure: a population of builders stands on the platform's defaults, so a default that mis-shapes context is not one team's cost but a levy paid across every project built on top, continuously.3 And this is no longer one vendor's default: the skill format has been released as an open standard and adopted across multiple major agent platforms.15 We do not claim the adopting runtimes share identical semantics—this paper examines the tested Claude Code mechanics specifically (Appendix A.2)—but the gap it leaves is shared, and the reach of that gap is the reach of the standard. The scope of that standard is itself the diagnosis—it successfully standardizes the unit of procedural capability, the skill package, but not the graph over units, the relations between skills. The node is standardized; the edge is not. And the missing edge is a specific thing: not a hyperlink but a state transition. In a state machine a transition does not merely name the next state—it enters it; a skill-to-skill relation (A specializes into B, delegates to C, must be followed by D) has to inject the next skill into the working context, or the relation is only prose. The standard fixes the package and withholds this transition operator over packages, which is why builders end up rebuilding the missing state machine as an external harness—the very programming the skills move was meant to retire (Background). The missing structure over packages (§1) is therefore a shared gap, multiplied across every platform that adopted the unit without it. No intent is claimed, and none is required—the shape of the failure is identical whether it was chosen or merely shipped quickly. What the reader pays is concrete: building on the substrate correctly is the capability that was not provided, and the capabilities that were provided stand in the way of building it oneself.
Stated generally, this is the principle the rest of the paper instantiates: when a system exposes an object in a shape that violates the object's natural structure, downstream consumers must spend inference, tokens, tool calls, harness code, and human supervision to recover the structure the interface destroyed. For skills the object is the library's relational structure, the violating shape is the flat list, and those five costs are exactly the ones the following sections document—inference re-derived per step (§1), tokens (§4), tool calls and harness code to navigate by hand (§3, Background), human supervision to supply the missing edges (§7). We establish the instance for skill libraries; the general statement is the frame it fits, not a law claimed across all interfaces.
The frustrations practitioners report look unrelated; they are one bug. In each case the platform ships the nodes and omits the first-class edge that would connect them. Conversations arrive as a flat list—grouping exists, but groups cannot reference one another, so cross-group review is assembled by hand; a single tagging primitive would have made a graph. Agents can be defined but cannot be composed as first-class, observable subcalls without dropping to lower-level orchestration code, at which point the visible product surface no longer carries the nesting relationship.4 Skills carry a description but cannot cause a deeper skill to be injected (§6). The complaints differ; the missing edge is the same—and §1 has already identified agent operation as exactly the process that edges structure. Seen this way, the list stops being a backlog of feature requests and becomes a single diagnosis.
Much of the ecosystem has rearchitected on one piece of guidance: prefer skills (or code execution) to direct MCP19 tool-calling, because loading all tool definitions wastes tokens. The platform states the cost in its own words—agents connected to many tools "need to process hundreds of thousands of tokens before reading a request," and loading on demand cut usage "from 150,000 tokens to 2,000 tokens—a … saving of 98.7%."5 The number is real; the framing is not. It compares two representations of tools and asks which is cheaper, which is the wrong axis. Per unit the two are roughly even—a JSON schema is a docstring, and a wrapping skill exposes code of comparable size. The difference is the loading model: every schema upfront versus a tool tree descended on demand,16 O(#tools) versus O(depth) (Figure 1). The selection failures that appear past a threshold of available tools6 are the symptom of the flat loading model, not of MCP. So the migration does not restore the library's latent structure—it moves a flat representation to a different flat representation. The resolution on either side is the same: impose a tree. Figure 1 is an asymptotic model, not a measured benchmark: it visualizes the structural prediction implied by the sources above—flat exposure grows with the number of tools, descended exposure with the number of levels traversed—while the empirical crossover remains to be measured. The O(depth) bound assumes a bounded branching factor, branch summaries informative enough to choose the right child, traversal errors that are rare or recoverable, and tasks that usually need one branch rather than many siblings at once; where those assumptions fail, descent costs more than the bound suggests.
O(#tools) versus tree O(depth). The flat curve loads every tool definition before the first request; the tree curve descends a branching menu, reading only one node's children per level. The two square/circle markers are Anthropic's own measured points—150,000 tokens loaded upfront and 2,000 tokens on-demand, a 98.7% reduction5; the model is calibrated so the flat line passes exactly through the former. The robust content is the asymptotics—flat is linear in the number of tools, the tree depth-bounded under a fixed branching assumption—and the crossover: below a handful of tools the flat load is trivially cheaper, which is why small toolsets feel fine and the tax goes unnoticed; past it the linear cost explodes while the tree stays nearly flat. (Model: flat(N)=150·N; tree(N)=200+30·5·⌈log₅N⌉. The crossover location depends on these constants; the divergence does not. Generated by figures/figure1_crossover.py; vector master at figures/figure1_crossover.svg.)Interactive: the two loading models of this section — every definition upfront versus a tree descended on demand (§6's flat forest and its correction) — can be experienced directly in the skilltree demo, with both context meters counted live.
The cost of flatness can be stated more generally as representational distance. A task does not begin from a neutral list of options; it begins inside a structured option-space. Before a run selects a route, the library is a hypergraph of possible acyclic projections: a skill can specialize several parents, participate in several workflows, or serve as a shared primitive across branches. Once a task, root, or relation policy is selected, that option-space collapses to the DAG the model must actually traverse, and a tree is the cheapest runnable projection of that DAG.
If the interface exposes the hypergraph as a flat list, the natural geometry does not disappear—it becomes an implicit constraint the model must recover before acting. Each time the model is handed the wrong shape it must perform an extra meta-level step: infer the higher-order structure, choose a projection, then return to the object-level task. The farther the exposed interface is from the required projection, the more inference is spent merely recovering the frame. Geometric debt is paid as inference.
That inference cost then surfaces as binding cost: extra always-on instructions, system-prompt rules, documentation indexes, naming conventions, and harness logic whose purpose is to make the model bind the right context at the right time.18 Vercel's result is a concrete instance of the shift: a small always-on AGENTS.md index outperformed skills, while skills were often not invoked at all.12 In this paper's terms the always-on index partially restores the missing traversal scaffold—it binds the route earlier, because the skill surface itself provides no first-class transition over skills. The better correction is not to keep moving structure into ever-larger prompts but to represent the transition directly: front-load the graph, not the corpus.
The sharpest case is also where the platform's own terminology makes the missing recursion most visible, and it must be stated precisely because the casual version is false. "Progressive disclosure" is the platform's own term—a section heading in the same post, where the remedy is to "read tool definitions on-demand, rather than reading them all up-front."5 And skills do disclose progressively within a skill: a three-level descent from frontmatter to body to referenced files.7 The failure is at the boundary—a skill cannot cause another skill to be injected into the harness. Progressive disclosure is recursive by nature; a tree discloses at every node, not only the root. Progressive disclosure that does not apply to the skill library itself is not closed under composition. A depth-one implementation therefore cannot apply the principle to itself: one discloses into a skill but never from one skill into a more specific one. The feature was named for the tree and shipped as a star—one center, depth-one leaves, no descent—and a set of stars with no connecting edges is a flat forest, precisely at the scale (a library of hundreds of skills) where disclosure was the point. This is not an omission to be patched by more prose inside a skill: the public skill primitive provides no first-class skill-to-skill edge—a skill can disclose its own files but cannot recursively disclose the skill library as a navigable tree—so the recursion is structurally foreclosed. What the reader needs is therefore not a richer skill but the missing level itself (§8).
Rather than argue from first principles, we ran the most disciplined instance of the flat-folder approach—the Interpretable Context Methodology, "folders, not agents"8—with a standard model (deliberately not an enhanced one, so we measure the system and not the agent) on a course-deck-production task: two staged passes, extraction then curriculum, over a short source, run straight and then debriefed; the run record—the inputs, the staged contracts, and the points where the output diverged from a correct deck—is included as supplementary material so the reader need not take the account on trust. This is an engineering case study, not a controlled benchmark: the claim is about observed failure modes in a standard non-interactive run, not population-level performance across all folder workflows. It holds for repeatable, human-reviewed workflows, and its value is locatable: an explicit inputs table, an audit step that forces verification before output, and context scoping. Tellingly, the lift came from the structure—the tagged schema, the staged contracts—not the surrounding prose, exactly as §1 predicts. And it breaks predictably the moment the human steps out: checkpoints stall a non-interactive agent, the inter-stage hand-off is fragile (shared state copied forward with no canonical owner; an under-specified naming convention that silently mismatches), and the audit is shallow enough to pass plausible-but-incorrect output. The lesson is not that folders are wrong but that a flat substrate carries structure only while a human supplies the missing edges by hand—which does not scale, and is not autonomy. ICM is the empirical clue: folders help because they approximate a context graph. The corrective is therefore not a better agent than ICM's; it is that graph made explicit and self-maintaining—supplied by the tree instead of by a human's hand. A reader should expect results of ICM's kind, carried by the structure rather than reconstructed each run, not a new performance record.
The correction restores structure one axis at a time, as tools that install in a single command and were verified at the level of CLI execution, filesystem mutation, generated output, and observed runtime loading behavior—not by assertion. skilltree map renders a flat skill folder as a coordinate-addressed, navigable tree—the missing level of §6. skilltree search provides zero-dependency retrieval over any folder, with a coordinate --scope that restricts it to a subtree. The same tool's discover/cohere/emit/unemit loop (Appendix A) keeps the materialized tree coherent as the library changes. Each closes one of the loops §3 left open. The contribution is not algorithmic—most of the machinery is commodity—it is the supply of structure, at zero configuration, to a reader (and a model) that otherwise cannot construct it. A temporal axis—a state machine over the tree, restoring cross-turn memory—is in progress and is not claimed here as complete.
Stated as a demand rather than a description: a single tree that self-manages, and an agent that can navigate it. The tools above approximate this on the filesystem because the filesystem is the weakest substrate that supports it—lacking first-class edges, it must emulate them, with coordinates standing in for identity and breadcrumbs for edges. A substrate with first-class, typed relationships—a graph—is the native form of that latent structure: it holds the relations directly rather than emulating them.9 More precisely, there is no single canonical tree: the latent structure is a hypergraph, and any one tree is a chart selected from it for a purpose—the same skills organize into different trees under different relations. A tree is then the chart a given task takes; the general substrate is the hypergraph that holds them all. The clearest case is a shared operational primitive—one skill used by several workflows: a single-parent tree must either duplicate it under each parent or appoint one artificial owner and cross-reference the rest, whereas a typed hyperedge holds it once and makes it reachable from all. The tree is a deployable projection of the graph; what a projection drops is exactly that shared reality. Systems such as OpenCog Hyperon illustrate the graph-native end of the design space: graph-shaped cognition treated as an AGI architecture.10 But Hyperon sits above the bridge layer this paper targets—it asks the user to enter a formal metagraph programming environment before any structure is gained. The missing layer is the same insight made runnable for the terminal-capable operator who can run an agent and let it modify local files: folders, coordinates, commands, scoped loading, and descent. The binding constraint there is runnability, not power. A minimal, dependency-free graph-native form of the mechanism is provided as a small experimental tool17—skills as nodes, typed relations as edges where an edge is a state transition that injects context, with tree-charts derived from the hypergraph on demand—the runnable floor to Hyperon's ceiling. This paper does not specify the maximal substrate. It establishes only what the reader can act on now: the operation being built is graph-shaped (§1), the prevailing defaults are flat, and the distance between them is inexpensive to close.
This is also where the practitioner's intuition that "the folders become the agent" becomes precise rather than mystical. If the agent is a flow and the graph is the space it flows over (§1), then constructing the context-graph is not configuring a tool the agent uses—it is building the space the agent is. Reified far enough, the structure is not consulted by the agent from outside; it is what the agent runs as. We state this as the organizing conjecture, not a proven result: the filesystem tree is the runnable first step toward that identity, a graph is its native form, and the empirical content—how far a constructed structure can stand in for a hand-built agent—is exactly what the benchmark this paper specifies would measure. (The soundness layer required to operate such a substrate autonomously is out of scope.)
The required ingredients were not exotic; the difficulty was recognizing the missing layer. Hierarchical agents that select a minimal toolkit are no longer exotic to build. The decisive change is not that trees became possible but that tool and skill corpora grew large enough for flat exposure to fail visibly. Two conditions delayed it, and both have lifted. Models were not previously reliable enough at instruction-following and tool use to drive their own navigation of a tree;11 that constraint has weakened enough to make tree navigation practical in ordinary agent loops. And the tools were built flat—which is a default, not a law.13 The reader who felt the problem of §3 now has both the reason it persisted and the evidence that the cost of removing it is small.
A candid limitation bounds the contribution. This paper does not prove that trees outperform flat tool lists on arbitrary agent tasks; it argues that flat exposure creates a structural scaling failure, demonstrates a filesystem-level correction, and specifies the benchmark needed to measure the crossover. The corresponding task benchmark—flat skills versus retrieved skills versus tree or graph traversal across increasing library sizes—is future work by design, not an omitted result: including it here would change the claim from substrate-correction to agent-performance. The measurement in scope here is substrate-level, not task-level: Figure 1 models the context-cost crossover, while Appendix A verifies the runtime injection behavior required for explicit descent. The tools of §8 are separate single-purpose programs, and the graph-native substrate of §9 is demonstrated in minimal, runnable form (the experimental tool above). What remains open is not whether such transitions can be represented, but how to package one structure across development, plugin, template, and served forms. We call this open problem the packaging geometry, and this paper does not solve it. The goal is one structure that is, simultaneously, four things: (i) a development tree one builds and reasons inside; (ii) an installable plugin, distributed and versioned through the platform's marketplace; (iii) a forkable template any agent can adopt wholesale as its own starting structure; and (iv) a thing served—as an MCP endpoint or a skill-pack—to consumers who never see its internals. Each packaging pulls the layout a different way, and the ways conflict: the development tree wants nested directories with edges emulated by breadcrumbs; the plugin wants a flat, served skills/ directory under a manifest; the forkable template wants self-containment with no external references; the served surface wants a stable external namespace independent of the internal one. The same skill must, under these four readings, sit in four different places without being copied four times and drifting. Reconciling them into a single self-consistent geometry is unfinished work, and we decline to present a design we do not have.
This is deliberately the last level of the program, not the first. Above it sits a further layer still—the composition of trees into reusable higher-order structures (a "framework" layer; the working name, nomicon or otherwise, is unsettled), and the means by which such structures are authored, named, and shared—which is deferred in full to subsequent papers. What the present paper rests on is the level below the packaging: that agent operation is graph-shaped (§1), that the prevailing defaults are flat, and that a tree closes the gap cheaply (§8, Appendix A). That result does not depend on the unified artifact existing; it depends only on the base code, which runs today. Accordingly we ship the base code as a public repository and claim nothing more for it—not that it is yet a plugin, a template, or a service, only that it demonstrates the correction and that the packaging is later, tractable engineering rather than a precondition. To conflate "the fix exists" with "the fix is packaged for every consumer at once" would be exactly the overclaim §2's argument forbids us; the honest statement is that the structural correction is implemented and the distribution geometry is open.
Program note. This paper is the first layer of a larger program; it addresses only the skill library's latent structure—making a runnable projection of it explicit and navigable. Later layers treat notation construction, skill composition, retrieval steering, and exchange/federation across trees. Each is deferred to its own work; none is claimed here.
heaven-tree-repl (the present author's own; MIT; github.com/sancovp/heaven-tree-repl), first published August 2025, roughly two months before skills shipped (October 2025) and before on-demand loading was advocated over upfront schemas (November 2025, note 5). Neither folder navigation nor an executable tree is a new idea; the skill mechanism is itself, in abstract form, a navigation CLI over such a tree. The point is precedence, not credit—the construction was demonstrated and went unremarked. It was there. ↩--help is documented in current engineering writing (e.g. Bradley Walters, "The best code is no code: Composing APIs and CLIs in the era of LLMs," Jan. 19, 2026). We cite this as industry-practice corroboration—the same descent shape at the tool granularity—not as a proven result. ↩graph-skills-experimental (an artifact by the present author; MIT; github.com/sancovp/graph-skills-experimental): a pure-stdlib demonstration—skills as nodes; typed edges (specializes/delegates_to/must_follow/contains); a state machine in which firing an edge injects the target node's context into an accumulating working set; and chart(), which derives a tree from the hypergraph (different roots/relations yield different trees, the "tree-options"). It lifts a skilltree tree as the degenerate case. Provided as a runnable existence proof of the graph-native form, not a production substrate. ↩The argument above is structural; this appendix makes it concrete. The listings are reproduced from the running tools of §8, not paraphrased. They show three things: what a tree is on a flat substrate (A.1), that the substrate's load behavior is exactly as the argument requires (A.2, verified against the live runtime), and that the structure the platform will not maintain can be made to maintain itself (A.3).
Code availability. The library is released by the present author as
skilltree(MIT) atgithub.com/sancovp/skilltree(v0.2.0, commit883cd11, 62 tests passing). References below arepath:lineagainst that repository; listings are excerpts of the released source. The single-file map rendering of §6's missing level is theskilltree mapsubcommand; the retrieval of §8 isskilltree search(with a coordinate--scope).
Lacking first-class edges (§9), the filesystem emulates them. Every node is a plain directory carrying its own one-skill .claude/; only the root auto-loads, and each node's menu carries the edges to its children as explicit instructions to read them.
domain/.claude/skills/0-domain/SKILL.md # root menu — the ONLY node that auto-loads
domain/A/.claude/skills/0.1-A/SKILL.md # child A — loaded only when Read
domain/A/A1/.claude/skills/0.1.1-A1/SKILL.md # grandchild — one level deeper
domain/B/.claude/skills/0.2-B/SKILL.md # child B
The root menu's body, generated verbatim:
## Index summary
[0] domain — a sc index node; opens to 2 branch(es): A (cor), B (ac). Reachable below: A, A1, A2, B.
## Descend — the next layer (2)
Only this layer is loaded now. To descend, use the Read tool on a child below — that injects
the child's layer. A shell `cat` reads the bytes but loads nothing; you must use the Read tool:
- 0.1-A (cor): Read `…/domain/A/.claude/skills/0.1-A/SKILL.md`
- 0.2-B (ac): Read `…/domain/B/.claude/skills/0.2-B/SKILL.md`
Two parts do the work. The index summary is composed from the subtree's own vocabulary, so a query about a descendant (A1) still retrieves the branch that leads to it—internal nodes carry a summary of what lies below, recovering retrievability that a flat list of leaves loses. The descend block is the set of out-edges, expressed as the single operation that loads them. A coordinate (0.1.1) stands in for node identity; a breadcrumb stands in for an edge. Both are emulations of what a graph would hold natively (§9). (Source: the breadcrumb emitter src/skilltree/materialize.py:36 (_write_node); the Read breadcrumb format at :25; the gate that requires a resolvable Read breadcrumb for every child, src/skilltree/validate.py:60.)
The argument needs three facts about the substrate to be true. We verified each directly—observed as a state-change in the live skill manifest, not asserted:
| Observation | Result | What it grounds |
|---|---|---|
A .claude/skills nested below the directory in scope | not auto-loaded | §6: depth is invisible until entered, so a depth-one library is a flat forest |
| Reading a file in a directory (the Read tool) | injects that directory's CLAUDE.md + rules + skills, one level; ancestors load, descendants do not | §8–§9: descent is real—a Read into the next node, accumulating the path while excluding the unentered rest |
A shell cat of the same file | injects nothing | the trigger is the Read tool, not the byte-read; the descend instruction must say "Read," or descent silently fails |
These are observed properties of the tested Claude Code runtime, not claims about filesystems in general; in that runtime the Read tool is a context-loading operation while a shell cat is only a byte read. They are the mechanism behind both the disease and the cure. The first row is why "progressive disclosure" shipped as a star (§6): the tested runtime enforces depth invisibility unless descent is performed through Read. The second is why a tree of plain directories is traversable at all: entering a node loads exactly its layer and no more, which is progressive disclosure performed by the act of reading. The third is a correctness condition on the emulation—an edge the reader cannot follow is not an edge—found only by running it. (The descend instruction is fixed to the Read tool at src/skilltree/materialize.py:25; the observations were reproduced by the marker-probe method described under Provenance.)
A structure the platform will not maintain must maintain itself, or it decays to the flat forest it replaced. Four operations close that loop. discover reconstructs the on-disk tree from reality, independent of any manifest. cohere reports the drift between that reality and the engineered shape—a bare (unrooted) forest, a stale breadcrumb, a coordinate that no longer matches, a skill dropped in flat. emit re-coheres in place; given a flat forest it tree-ifies it, relocating each skill—with all of its files, and resolving any symlink to a served source—into a nested node and writing the edges, journaling every move so that unemit restores the prior state byte-for-byte. A scheduled check writes its verdict into a single managed rule, so a decohered tree announces itself in every subsequent session. All four operations are implemented in one MIT-licensed tool (skilltree), whose test suite passes at this writing. The algorithms are commodity; the contribution is that the maintenance exists at all, on a substrate that provides none. (Source: src/skilltree/cohere.py — discover:117, cohere:152, emit:261, unemit:339; the scheduled check write_notifications:430 / watch:448.)
Provenance. All listings and the A.2 observations are reproduced from tools that install in one command and were exercised on a real filesystem and the live runtime; the substrate behavior in A.2 was confirmed by planting a marker skill in a nested directory and observing whether it entered the model's tool manifest under each access pattern. No claim in this appendix rests on a self-certified "verified" label; each rests on an observed state-change.