Thomas C. Ricks · September 2026
1. What Treebeard Is
Treebeard is a single-operator, self-hosted multi-agent AI system. In this project, production means the live instance I use for real work, not a customer deployment. That distinction matters because I do not want portfolio language to imply scale or maturity the system has not earned.
The live system already has a working router, model/provider selection, typed provider fallthrough, telemetry, cost controls, agent execution paths, and regression infrastructure. The captured source also includes agent-scoped memory components. Router 2 is therefore a migration, not a greenfield rewrite.
The routing problem is broader than model selection. A turn can require mission selection, context retrieval, authorization, execution-shape selection, provider choice, fallthrough, and evidence capture. If those decisions are implicit or coupled, simple work becomes expensive, ambiguous work is hard to diagnose, and failures become difficult to reproduce.
The enterprise version of the same problem has higher consequences: proprietary data needs declared boundaries, authorization must be independent of model confidence, provider policy must actually be enforceable, and fallback behavior cannot silently broaden authority.
2. The Product Goal
The product promise is simple: give each request the cheapest competent path, the correct owner, the minimum necessary reasoning, and exactly the authority it has been granted. Keep enough evidence to reconstruct the decision later.
- Resolve obvious work with deterministic or inexpensive routing before paying for semantic inference.
- Separate routing, authorization, execution, and evidence into explicit contracts.
- Preserve operator direction and correction as stronger signals than inferred preference.
- Keep provider/model selection cost-aware and capability-aware.
- Make fallthrough finite, typed, and observable.
- Capture disagreements and negative results instead of rewriting the story after the fact.
- Promote capability in stages rather than treating implementation as proof of readiness.
3. Evidence States
I use four evidence states throughout the project because source code, runtime behavior, and future architecture are not the same thing.
| State | Meaning |
|---|---|
| Runtime | Observed working in the live system with retained evidence. |
| Source | Present in captured source with self-tests; full production integration is not implied. |
| Designed | Specified and reviewed, but not yet integrated as the production default. |
| Proposed | A research or product hypothesis with no implementation claim. |
Current runtime baseline
The working router remains the executable reference implementation during Stage 1. Characterization is performed against a named, frozen revision rather than against a moving target. Active Router 1 work can continue, but every Router 2 comparison must identify the exact baseline revision it is being compared against.
Current source-state memory work
The captured source includes a salience-based experience-memory store, a daily consolidation path, and long-term memory with provenance and per-agent write boundaries. An explicitly configured supervisory role can read across agent stores without receiving cross-agent write authority. The modules also include self-test seams, failure contracts, overlap protection, and atomic-write handling.
Per-agent read isolation has point-in-time runtime evidence, including distinct per-agent markers and a denied cross-agent read. Retrieval is conditional rather than universal, so Router 2 still has to define exactly when memory signals are eligible to influence routing.
That last boundary is important: memory can help answer where this work probably belongs. It cannot answer what this agent is authorized to do.
4. Development Method
The development method is incremental and bounded: one change small enough to understand, one explicit contract, then regression before the next change.
Each work packet defines inputs, outputs, invariants, acceptance criteria, and responsibilities that are deliberately out of scope. A component is not promoted because it was implemented. It is promoted only after its behavior is characterized and tested against the current system.
Inventory current behavior → define contract → characterize baseline → implement one bounded component → component test → contract test → full regression → replay known cases → shadow comparison → disagreement review → limited authority → expanded authority → retire the legacy dependency.
Continuity across session boundaries depends on versioned artifacts, hashes, regression evidence, and explicit next-step gates rather than developer memory.
Evidence practices already in use
Clean-room qualification. A release candidate was rejected when a clean installation exposed a configuration-path defect. The corrected candidate subsequently passed clean-room qualification. The rule adopted from that failure is simple: declare the expected result and the measurement method before execution, then accept or reject what the run actually shows.
The test harness is part of the product. A suite that cannot execute is not green. A skipped hard gate is not a pass. A silent check needs a positive control when silence could also mean the check never ran. Retained logs show registered safe-tier coverage growing from 9 suites in mid-August to 30 by August 31. The suite definitions changed during that period, so that number describes coverage surface, not quality.
Negative results stay negative. One bounded autonomous-development exercise failed five of eight tests on a locked acceptance suite, with all five failures in a single functional area. I recorded it as a capability ceiling rather than a partial pass. That is the starting prior for Stage 2: automated code modification has to earn broader authority through better evidence.
Authority is explicit state, not narration. A model can describe evidence, but generated text is not proof that an action occurred. Execution-status claims must come from deterministic system state such as dispatch and receipt records.
5. Three-Stage Roadmap
| Stage | Goal | Authority posture |
|---|---|---|
| 1. Governed Router 2 MVP | Replace routing internals incrementally while preserving a working reference implementation. | Committed product. Router 2 earns authority by request class through qualification gates. |
| 2. Adaptive Mirror | Let Treebeard modify and test a copy of Treebeard inside a separate controlled environment. | Future gated experiment. Broad code authority inside the mirror; no control over the enforcement or promotion plane. |
| 3. Treebeard OS | Explore an AI-oriented operating environment with broad delegated administration and a separate root of trust. | Research thesis, not a committed build. |
6. Stage 1: Governed Router 2 MVP
Stage 1 is the actual product commitment.
Operating loop
Core requirements
- Create one canonical routing record per attempt containing request identity, candidate missions, evidence, routing-cognition level, authority state, provider target, fallthrough state, provenance, latency, and final disposition.
- Run explicit operator direction and deterministic evidence before model-assisted inference.
- Support control turns, explicit mission references, lexical routing, thread continuity, genuine no-match, bounded tie resolution, and model-assisted fallback without conflating them.
- Keep mission inference and authorization structurally separate. Memory, history, reasoning, and provider selection cannot create authority.
- Separate routing cognition from execution cognition. A difficult work item does not automatically require an expensive routing decision, and a difficult routing decision does not automatically require an expensive executor.
- Select provider/model from declared capability, health, cost policy, privacy constraints, and required execution behavior.
- Use finite, typed provider fallthrough. A required capability cannot be silently downgraded to rescue a call.
- Persist reconstructable provenance for every mission-selecting or execution-shaping production route.
- Capture operator corrections as feedback linked to the originating decision without directly mutating active policy.
- Support replay, shadow evaluation, rollback, and a Router 1 fallback only before external dispatch begins.
Routing precedence
Precedence is explicit because hidden precedence is where authority mistakes begin.
| Priority | Evidence source | Rule |
|---|---|---|
| 1 | Policy/store invariant | Hard boundary. Lower signals cannot override it. |
| Control | Trusted system/administrative provenance | Bypasses mission selection, carries no mission-scoped tool authority, and cannot be created merely by request content. |
| 2 | Explicit operator assignment | Binding when policy-compliant. |
| 3 | Persisted operator routing binding | Applies until superseded or expired by declared policy. |
| 4 | Explicit mission reference in the turn | Deterministic when exactly one eligible mission is named. |
| 5 | Deterministic lexical winner | Uses a configured score floor and margin; no model call. |
| 6 | Thread continuation | Deterministic when no stronger current-turn evidence is present; never overrides priorities 1-5. |
| 7 | Bounded history/memory tie resolution | May choose only among already eligible candidates. |
| 8 | Model-assisted fallback | Cold path. May rank eligible candidates but cannot add authority. |
| Terminal | No-match or clarification | Used only when priorities 1-8 do not resolve an eligible route. Clarification is reserved for cases involving an external side effect, irreversible write, or spend above a configured threshold. |
Routing-cognition levels
- Fast: one of the deterministic priority rules resolves the route without semantic escalation.
- Deliberative: two or more eligible candidates remain within the configured ambiguity band, or deterministic signals materially disagree.
- Deep: Deliberative routing remains unresolved, or the routing decision gates a consequential external side effect and the configured policy requires a deeper check.
These are routing levels, not execution-model tiers. The separation is deliberate.
Authority closure
The governing invariant is:
Candidate discovery is constrained to missions the request is eligible to enter. Governance is the only layer that can write authority state. Execution shaping may choose only among capabilities already inside the governance-approved envelope. If shaping requires a capability outside that envelope, governance is re-run before dispatch.
Provider selection happens after authority is established. A cheaper or more capable model does not gain permission merely because it can perform the task.
Provider failure handling
| Failure | Handling | Evidence state |
|---|---|---|
| Reasoning exhaustion / empty answer at budget ceiling | Bounded in-rung retry, then permitted fallthrough. | Observed in runtime and retained as a named failure mode. |
| Backend exception | Walk the permitted ladder. | Covered by component/contract evidence. |
| Successful call with empty output | Treat as a failed attempt; continue only if policy permits. | Deterministic test evidence. |
| Rate limit within ladder | Move to the next permitted provider. | Requires complete production qualification before Stage 1 exit. |
| Ceiling reached | Fail or retry within policy. Never silently lift the ceiling. | Policy behavior, with further Stage 1 qualification required. |
Stage 1 acceptance gates
| Gate | Required evidence | Current status |
|---|---|---|
| G0 - Characterize | Frozen Router 1 revision; routing behavior, authority outcomes, provider/fallthrough behavior, latency/cost baseline, historical defects. | Partially complete; rerun against the frozen Stage 1 baseline. |
| G1 - Shadow | Canonical request/evidence/decision records; no side effects; replay set; deterministic fast path. | Not started. |
| G2 - Limited authority | Green hard-gate suite; disagreement review; rollback proven; no governance regression. | Not started. |
| G3 - Production default | Stable operation; fault injection; audit completeness; Router 1 fallback tracked and limited to exceptional pre-dispatch failures. | Not started. |
| G4 - Router 1 retirement | Standalone Router 2; legacy adapters removed; archived reference reproducible. | Not started. |
7. Stage 2: Adaptive Mirror
Stage 2 is a gated experiment, not an MVP requirement.
The prior autonomous-development benchmark failed five of eight tests on a locked acceptance suite. That is evidence of a current ceiling, not evidence that automated code improvement is already reliable.
The Adaptive Mirror would contain a runnable copy of Treebeard and its development surface. Treebeard could inspect source, modify source, create candidate configurations, run tests, benchmark changes, rebuild designated services, and prepare a promotion proposal.
It would not control the layer that defines those permissions.
| Allowed inside the mirror | Outside Treebeard's authority |
|---|---|
| Modify Treebeard application and router code | Change host-level enforcement policy |
| Run unit, regression, replay, benchmark, and adversarial tests | Read production secrets or unrestricted operator credentials |
| Build dependencies inside declared sandbox policy | Introduce unreviewed external dependencies into a promotion candidate |
| Create candidate configurations and learned artifacts | Promote itself to production |
| Restart and rebuild mirror services | Disable audit, rollback, snapshot, or resource limits |
| Generate a promotion proposal with evidence | Modify the trust boundary or use production external-action authority during self-test |
Each improvement cycle must declare its target metric and pass threshold before the change is made. A cycle counts only against that declared metric. This prevents post-hoc selection of whichever result happens to look favorable.
Promotion remains external: observe → identify defect/opportunity → propose → modify mirror → component test → full regression → replay/benchmark → adversarial review → external promotion gate → promote or reject → retain evidence.
8. Stage 3: Treebeard OS
Stage 3 is a research thesis. The point is not to claim that Treebeard needs a custom operating system today. The point is to make the privilege model explicit.
The hypothesis is that an AI-centric operating environment can give Treebeard broad delegated administrative capability while keeping the authority that defines those capabilities in a separate privilege domain.
Treebeard's operating plane could administer assigned services, packages, development environments, containers or VMs, storage, schedulers, observability, and candidate system images.
The superior trust plane would retain exclusive authority over boot trust, key material, audit integrity, recovery, promotion, and the rules that define Treebeard's capability ceiling.
Privilege decomposition alone is not enough on a conventional kernel. Several individual capabilities are practical paths back to full root. The ceiling therefore has to be enforced from a separate privilege domain - for example a hypervisor, mandatory access control policy controlled outside the delegated plane, or an equivalent mechanism - rather than by ordinary file permissions inside the environment Treebeard administers.
Audit records likewise need to live outside Treebeard's write authority. The system cannot claim independent oversight if the subject being audited can rewrite the audit trail.
9. Threat Model
| Failure mode | Control |
|---|---|
| Inference becomes authority | Only governance writes authority state. |
| Automated code changes disable controls | Enforcement and promotion remain outside the writable mirror/OS plane. |
| Provider outage or degradation | Finite typed fallthrough; required capability is never silently downgraded. |
| Prompt or repository content attempts policy injection | Untrusted content remains input/evidence; policy and authorization are separate control channels. |
| Cross-agent memory leakage | Per-agent scope; supervisory cross-read is explicit; cross-agent write remains prohibited. |
| Runaway cost | Provider tiers, budget ceilings, telemetry, and escalation evidence. |
| Silent regression | Corrected defects become regression evidence; replay/shadow comparison precedes authority promotion. |
| Model narration presented as proof | Execution claims derive from deterministic dispatch/receipt state. |
| Control exists but is ineffective | Critical gates require behavioral false/deny tests; configuration presence alone is not evidence that the control works. |
| Fallback silently changes authority state | Fallback may reduce optional context but may not silently alter mission, authorization, provider ceiling, or required execution capability. |
| Background lifecycle service does not run | Any safety property that depends on scheduled work requires an independently monitored schedule/health signal and a fail-safe state. |
10. Metrics
The metric table is intentionally honest about what is not yet instrumented. The first product gate is to produce the baseline, not to pretend it already exists.
| Metric | Current evidence | Next decision |
|---|---|---|
| Fast-path rate | Not measured as a production rate. | Establish at G0, then set a target. |
| Mission correction rate | Not captured. | Instrument, baseline, then reduce. |
| Authorization integrity | No recorded incident in which inference, memory, or narration granted capability that policy did not already permit. This is an absence of recorded incidents, not a measured rate, and authority-boundary test coverage has itself required correction. | Target 0 once instrumented at G1. |
| Provider recovery rate | Partial evidence exists for named provider failures, including reasoning exhaustion. | Qualify every named mode before G3. |
| Routing latency | No stage-level p50/p95/p99 production distribution yet. | Baseline and optimize without changing authority semantics. |
| Cost per successful work call | Not yet established as a stable clean production rate. | Separate production traffic from synthetic probes before publishing a rate. |
| Replay disagreement rate | Not applicable until Router 2 shadow/replay exists. | Categorize every disagreement. |
| Regression escape rate | Not tracked as a rate. | Instrument and trend. |
| Automated code-improvement yield | One attempt, n=1: failed five of eight acceptance tests and was retained as a capability ceiling. | Define Stage 2 promotion threshold before Mirror authority is granted. |
| Recovery confidence | A clean-room candidate failed, exposed a release defect, and the corrected frozen candidate subsequently passed qualification. | Repeat qualification and recovery drills as routine release evidence. |
11. Defects as Product Evidence
The project has been most useful when a defect changed the product rules rather than merely producing a patch.
| Defect | How it was found | What changed |
|---|---|---|
| Operator provider kill-switch was silently discarded | Code inspection plus a live false-gate probe | An explicit false value now survives into child execution; critical controls require behavioral false/deny tests. |
| Reasoning exhaustion produced no usable answer | Telemetry/provider response evidence | Reasoning exhaustion became a named provider failure handled by bounded retry/fallthrough. |
| Execution-status annotation contradicted actual multi-agent messaging | Live multi-agent transcript | Status text was corrected to derive from system state rather than generated narration, and the corrected behavior was added to regression coverage. |
| Clean-room candidate failed installation | Clean-room qualification | The candidate was rejected; the configuration path was corrected before the next qualification. |
| Autonomous-development benchmark failed five of eight tests | Locked acceptance suite | The result stayed a failed qualification and became the prior for Stage 2 gating. |
Each of these became a requirement, a regression case, or a qualification rule in the current product plan.
12. What Stage 1 Does Not Include
- The Adaptive Mirror.
- Treebeard OS.
- Unrestricted autonomous modification of the production system.
- Treebeard control of the superior root of trust.
- Automatic promotion of learned routing policy without versioning and rollback.
- Automatic permission expansion because a model, memory signal, or historical pattern is confident.
- Optimization for benchmark performance at the expense of provenance, recoverability, or authority integrity.
13. Open Validation Questions
- What percentage of real turns can Router 2 resolve deterministically once stage-level instrumentation is live?
- Which current Router 1 behaviors are intentional product contracts and which are implementation artifacts?
- What are the real p50/p95/p99 costs of lexical, thread, memory, and reasoning stages?
- How often do operator corrections identify candidate-discovery failure versus final-selection failure?
- Which classes of software changes can Treebeard implement reliably under locked acceptance criteria?
- What is the minimum external enforcement surface required to give an Adaptive Mirror meaningful development freedom without giving it a viable path to promotion or policy control?
- For a future Treebeard OS, which administrative operations can be delegated safely, and which must remain in a superior privilege domain?
- Which lifecycle properties depend on scheduled services, and how is failure of those services independently detected?
- How should synthetic probe traffic be separated from real work so cost and recovery metrics remain defensible?
14. Product Takeaway
The differentiator in Treebeard is not that it uses multiple models or multiple agents. Those are implementation details.
The useful product idea is the separation between routing, authorization, execution, evidence, and promotion. A routing signal does not become permission. A model's narration does not become proof that something ran. A new capability does not become production authority because it passed once.
Stage 1 turns those ideas into a qualified Router 2 migration. If that works, Stage 2 asks whether Treebeard can improve a copy of itself inside a boundary it does not control. Only repeated success there would justify the Stage 3 research thesis.
The immediate job remains straightforward: characterize the working router, build Router 2 one bounded component at a time, measure the differences, and promote only what the evidence supports.
No comments:
Post a Comment