Friday, September 18, 2026

Project Treebeard: Building a Governed Adaptive AI Development Environment


Thomas C. Ricks · September 2026

1. What Treebeard Is

Treebeard is a single-operator, self-hosted multi-agent AI system. In this project, production means the live instance I use for real work, not a customer deployment. That distinction matters because I do not want portfolio language to imply scale or maturity the system has not earned.

The live system already has a working router, model/provider selection, typed provider fallthrough, telemetry, cost controls, agent execution paths, and regression infrastructure. The captured source also includes agent-scoped memory components. Router 2 is therefore a migration, not a greenfield rewrite.

The routing problem is broader than model selection. A turn can require mission selection, context retrieval, authorization, execution-shape selection, provider choice, fallthrough, and evidence capture. If those decisions are implicit or coupled, simple work becomes expensive, ambiguous work is hard to diagnose, and failures become difficult to reproduce.

The enterprise version of the same problem has higher consequences: proprietary data needs declared boundaries, authorization must be independent of model confidence, provider policy must actually be enforceable, and fallback behavior cannot silently broaden authority.

2. The Product Goal

The product promise is simple: give each request the cheapest competent path, the correct owner, the minimum necessary reasoning, and exactly the authority it has been granted. Keep enough evidence to reconstruct the decision later.

  • Resolve obvious work with deterministic or inexpensive routing before paying for semantic inference.
  • Separate routing, authorization, execution, and evidence into explicit contracts.
  • Preserve operator direction and correction as stronger signals than inferred preference.
  • Keep provider/model selection cost-aware and capability-aware.
  • Make fallthrough finite, typed, and observable.
  • Capture disagreements and negative results instead of rewriting the story after the fact.
  • Promote capability in stages rather than treating implementation as proof of readiness.

3. Evidence States

I use four evidence states throughout the project because source code, runtime behavior, and future architecture are not the same thing.

State Meaning
RuntimeObserved working in the live system with retained evidence.
SourcePresent in captured source with self-tests; full production integration is not implied.
DesignedSpecified and reviewed, but not yet integrated as the production default.
ProposedA research or product hypothesis with no implementation claim.

Current runtime baseline

The working router remains the executable reference implementation during Stage 1. Characterization is performed against a named, frozen revision rather than against a moving target. Active Router 1 work can continue, but every Router 2 comparison must identify the exact baseline revision it is being compared against.

Current source-state memory work

The captured source includes a salience-based experience-memory store, a daily consolidation path, and long-term memory with provenance and per-agent write boundaries. An explicitly configured supervisory role can read across agent stores without receiving cross-agent write authority. The modules also include self-test seams, failure contracts, overlap protection, and atomic-write handling.

Per-agent read isolation has point-in-time runtime evidence, including distinct per-agent markers and a denied cross-agent read. Retrieval is conditional rather than universal, so Router 2 still has to define exactly when memory signals are eligible to influence routing.

That last boundary is important: memory can help answer where this work probably belongs. It cannot answer what this agent is authorized to do.

4. Development Method

The development method is incremental and bounded: one change small enough to understand, one explicit contract, then regression before the next change.

Each work packet defines inputs, outputs, invariants, acceptance criteria, and responsibilities that are deliberately out of scope. A component is not promoted because it was implemented. It is promoted only after its behavior is characterized and tested against the current system.

Migration unit:
Inventory current behavior → define contract → characterize baseline → implement one bounded component → component test → contract test → full regression → replay known cases → shadow comparison → disagreement review → limited authority → expanded authority → retire the legacy dependency.

Continuity across session boundaries depends on versioned artifacts, hashes, regression evidence, and explicit next-step gates rather than developer memory.

Evidence practices already in use

Clean-room qualification. A release candidate was rejected when a clean installation exposed a configuration-path defect. The corrected candidate subsequently passed clean-room qualification. The rule adopted from that failure is simple: declare the expected result and the measurement method before execution, then accept or reject what the run actually shows.

The test harness is part of the product. A suite that cannot execute is not green. A skipped hard gate is not a pass. A silent check needs a positive control when silence could also mean the check never ran. Retained logs show registered safe-tier coverage growing from 9 suites in mid-August to 30 by August 31. The suite definitions changed during that period, so that number describes coverage surface, not quality.

Negative results stay negative. One bounded autonomous-development exercise failed five of eight tests on a locked acceptance suite, with all five failures in a single functional area. I recorded it as a capability ceiling rather than a partial pass. That is the starting prior for Stage 2: automated code modification has to earn broader authority through better evidence.

Authority is explicit state, not narration. A model can describe evidence, but generated text is not proof that an action occurred. Execution-status claims must come from deterministic system state such as dispatch and receipt records.

5. Three-Stage Roadmap

Stage Goal Authority posture
1. Governed Router 2 MVP Replace routing internals incrementally while preserving a working reference implementation. Committed product. Router 2 earns authority by request class through qualification gates.
2. Adaptive Mirror Let Treebeard modify and test a copy of Treebeard inside a separate controlled environment. Future gated experiment. Broad code authority inside the mirror; no control over the enforcement or promotion plane.
3. Treebeard OS Explore an AI-oriented operating environment with broad delegated administration and a separate root of trust. Research thesis, not a committed build.

6. Stage 1: Governed Router 2 MVP

Stage 1 is the actual product commitment.

Operating loop

Incoming turn → normalize → control/directive checks → deterministic and fast-path evidence → ambiguity gate → bounded history/reasoning escalation → mission commit → governance validation → execution shaping → provider/model selection → dispatch/fallthrough → feedback and telemetry.

Core requirements

  1. Create one canonical routing record per attempt containing request identity, candidate missions, evidence, routing-cognition level, authority state, provider target, fallthrough state, provenance, latency, and final disposition.
  2. Run explicit operator direction and deterministic evidence before model-assisted inference.
  3. Support control turns, explicit mission references, lexical routing, thread continuity, genuine no-match, bounded tie resolution, and model-assisted fallback without conflating them.
  4. Keep mission inference and authorization structurally separate. Memory, history, reasoning, and provider selection cannot create authority.
  5. Separate routing cognition from execution cognition. A difficult work item does not automatically require an expensive routing decision, and a difficult routing decision does not automatically require an expensive executor.
  6. Select provider/model from declared capability, health, cost policy, privacy constraints, and required execution behavior.
  7. Use finite, typed provider fallthrough. A required capability cannot be silently downgraded to rescue a call.
  8. Persist reconstructable provenance for every mission-selecting or execution-shaping production route.
  9. Capture operator corrections as feedback linked to the originating decision without directly mutating active policy.
  10. Support replay, shadow evaluation, rollback, and a Router 1 fallback only before external dispatch begins.

Routing precedence

Precedence is explicit because hidden precedence is where authority mistakes begin.

Priority Evidence source Rule
1Policy/store invariantHard boundary. Lower signals cannot override it.
ControlTrusted system/administrative provenanceBypasses mission selection, carries no mission-scoped tool authority, and cannot be created merely by request content.
2Explicit operator assignmentBinding when policy-compliant.
3Persisted operator routing bindingApplies until superseded or expired by declared policy.
4Explicit mission reference in the turnDeterministic when exactly one eligible mission is named.
5Deterministic lexical winnerUses a configured score floor and margin; no model call.
6Thread continuationDeterministic when no stronger current-turn evidence is present; never overrides priorities 1-5.
7Bounded history/memory tie resolutionMay choose only among already eligible candidates.
8Model-assisted fallbackCold path. May rank eligible candidates but cannot add authority.
TerminalNo-match or clarificationUsed only when priorities 1-8 do not resolve an eligible route. Clarification is reserved for cases involving an external side effect, irreversible write, or spend above a configured threshold.

Routing-cognition levels

  • Fast: one of the deterministic priority rules resolves the route without semantic escalation.
  • Deliberative: two or more eligible candidates remain within the configured ambiguity band, or deterministic signals materially disagree.
  • Deep: Deliberative routing remains unresolved, or the routing decision gates a consequential external side effect and the configured policy requires a deeper check.

These are routing levels, not execution-model tiers. The separation is deliberate.

Authority closure

The governing invariant is:

Routing can decide where work belongs. Routing cannot create authority.

Candidate discovery is constrained to missions the request is eligible to enter. Governance is the only layer that can write authority state. Execution shaping may choose only among capabilities already inside the governance-approved envelope. If shaping requires a capability outside that envelope, governance is re-run before dispatch.

Provider selection happens after authority is established. A cheaper or more capable model does not gain permission merely because it can perform the task.

Provider failure handling

Failure Handling Evidence state
Reasoning exhaustion / empty answer at budget ceilingBounded in-rung retry, then permitted fallthrough.Observed in runtime and retained as a named failure mode.
Backend exceptionWalk the permitted ladder.Covered by component/contract evidence.
Successful call with empty outputTreat as a failed attempt; continue only if policy permits.Deterministic test evidence.
Rate limit within ladderMove to the next permitted provider.Requires complete production qualification before Stage 1 exit.
Ceiling reachedFail or retry within policy. Never silently lift the ceiling.Policy behavior, with further Stage 1 qualification required.

Stage 1 acceptance gates

Gate Required evidence Current status
G0 - CharacterizeFrozen Router 1 revision; routing behavior, authority outcomes, provider/fallthrough behavior, latency/cost baseline, historical defects.Partially complete; rerun against the frozen Stage 1 baseline.
G1 - ShadowCanonical request/evidence/decision records; no side effects; replay set; deterministic fast path.Not started.
G2 - Limited authorityGreen hard-gate suite; disagreement review; rollback proven; no governance regression.Not started.
G3 - Production defaultStable operation; fault injection; audit completeness; Router 1 fallback tracked and limited to exceptional pre-dispatch failures.Not started.
G4 - Router 1 retirementStandalone Router 2; legacy adapters removed; archived reference reproducible.Not started.
STAGE 1 ENDS HERE.

7. Stage 2: Adaptive Mirror

Stage 2 is a gated experiment, not an MVP requirement.

The prior autonomous-development benchmark failed five of eight tests on a locked acceptance suite. That is evidence of a current ceiling, not evidence that automated code improvement is already reliable.

The Adaptive Mirror would contain a runnable copy of Treebeard and its development surface. Treebeard could inspect source, modify source, create candidate configurations, run tests, benchmark changes, rebuild designated services, and prepare a promotion proposal.

It would not control the layer that defines those permissions.

Allowed inside the mirror Outside Treebeard's authority
Modify Treebeard application and router codeChange host-level enforcement policy
Run unit, regression, replay, benchmark, and adversarial testsRead production secrets or unrestricted operator credentials
Build dependencies inside declared sandbox policyIntroduce unreviewed external dependencies into a promotion candidate
Create candidate configurations and learned artifactsPromote itself to production
Restart and rebuild mirror servicesDisable audit, rollback, snapshot, or resource limits
Generate a promotion proposal with evidenceModify the trust boundary or use production external-action authority during self-test

Each improvement cycle must declare its target metric and pass threshold before the change is made. A cycle counts only against that declared metric. This prevents post-hoc selection of whichever result happens to look favorable.

Promotion remains external: observe → identify defect/opportunity → propose → modify mirror → component test → full regression → replay/benchmark → adversarial review → external promotion gate → promote or reject → retain evidence.

8. Stage 3: Treebeard OS

Stage 3 is a research thesis. The point is not to claim that Treebeard needs a custom operating system today. The point is to make the privilege model explicit.

The hypothesis is that an AI-centric operating environment can give Treebeard broad delegated administrative capability while keeping the authority that defines those capabilities in a separate privilege domain.

Treebeard's operating plane could administer assigned services, packages, development environments, containers or VMs, storage, schedulers, observability, and candidate system images.

The superior trust plane would retain exclusive authority over boot trust, key material, audit integrity, recovery, promotion, and the rules that define Treebeard's capability ceiling.

Privilege decomposition alone is not enough on a conventional kernel. Several individual capabilities are practical paths back to full root. The ceiling therefore has to be enforced from a separate privilege domain - for example a hypervisor, mandatory access control policy controlled outside the delegated plane, or an equivalent mechanism - rather than by ordinary file permissions inside the environment Treebeard administers.

Audit records likewise need to live outside Treebeard's write authority. The system cannot claim independent oversight if the subject being audited can rewrite the audit trail.

9. Threat Model

Failure mode Control
Inference becomes authorityOnly governance writes authority state.
Automated code changes disable controlsEnforcement and promotion remain outside the writable mirror/OS plane.
Provider outage or degradationFinite typed fallthrough; required capability is never silently downgraded.
Prompt or repository content attempts policy injectionUntrusted content remains input/evidence; policy and authorization are separate control channels.
Cross-agent memory leakagePer-agent scope; supervisory cross-read is explicit; cross-agent write remains prohibited.
Runaway costProvider tiers, budget ceilings, telemetry, and escalation evidence.
Silent regressionCorrected defects become regression evidence; replay/shadow comparison precedes authority promotion.
Model narration presented as proofExecution claims derive from deterministic dispatch/receipt state.
Control exists but is ineffectiveCritical gates require behavioral false/deny tests; configuration presence alone is not evidence that the control works.
Fallback silently changes authority stateFallback may reduce optional context but may not silently alter mission, authorization, provider ceiling, or required execution capability.
Background lifecycle service does not runAny safety property that depends on scheduled work requires an independently monitored schedule/health signal and a fail-safe state.

10. Metrics

The metric table is intentionally honest about what is not yet instrumented. The first product gate is to produce the baseline, not to pretend it already exists.

Metric Current evidence Next decision
Fast-path rateNot measured as a production rate.Establish at G0, then set a target.
Mission correction rateNot captured.Instrument, baseline, then reduce.
Authorization integrityNo recorded incident in which inference, memory, or narration granted capability that policy did not already permit. This is an absence of recorded incidents, not a measured rate, and authority-boundary test coverage has itself required correction.Target 0 once instrumented at G1.
Provider recovery ratePartial evidence exists for named provider failures, including reasoning exhaustion.Qualify every named mode before G3.
Routing latencyNo stage-level p50/p95/p99 production distribution yet.Baseline and optimize without changing authority semantics.
Cost per successful work callNot yet established as a stable clean production rate.Separate production traffic from synthetic probes before publishing a rate.
Replay disagreement rateNot applicable until Router 2 shadow/replay exists.Categorize every disagreement.
Regression escape rateNot tracked as a rate.Instrument and trend.
Automated code-improvement yieldOne attempt, n=1: failed five of eight acceptance tests and was retained as a capability ceiling.Define Stage 2 promotion threshold before Mirror authority is granted.
Recovery confidenceA clean-room candidate failed, exposed a release defect, and the corrected frozen candidate subsequently passed qualification.Repeat qualification and recovery drills as routine release evidence.

11. Defects as Product Evidence

The project has been most useful when a defect changed the product rules rather than merely producing a patch.

Defect How it was found What changed
Operator provider kill-switch was silently discardedCode inspection plus a live false-gate probeAn explicit false value now survives into child execution; critical controls require behavioral false/deny tests.
Reasoning exhaustion produced no usable answerTelemetry/provider response evidenceReasoning exhaustion became a named provider failure handled by bounded retry/fallthrough.
Execution-status annotation contradicted actual multi-agent messagingLive multi-agent transcriptStatus text was corrected to derive from system state rather than generated narration, and the corrected behavior was added to regression coverage.
Clean-room candidate failed installationClean-room qualificationThe candidate was rejected; the configuration path was corrected before the next qualification.
Autonomous-development benchmark failed five of eight testsLocked acceptance suiteThe result stayed a failed qualification and became the prior for Stage 2 gating.

Each of these became a requirement, a regression case, or a qualification rule in the current product plan.

12. What Stage 1 Does Not Include

  • The Adaptive Mirror.
  • Treebeard OS.
  • Unrestricted autonomous modification of the production system.
  • Treebeard control of the superior root of trust.
  • Automatic promotion of learned routing policy without versioning and rollback.
  • Automatic permission expansion because a model, memory signal, or historical pattern is confident.
  • Optimization for benchmark performance at the expense of provenance, recoverability, or authority integrity.

13. Open Validation Questions

  1. What percentage of real turns can Router 2 resolve deterministically once stage-level instrumentation is live?
  2. Which current Router 1 behaviors are intentional product contracts and which are implementation artifacts?
  3. What are the real p50/p95/p99 costs of lexical, thread, memory, and reasoning stages?
  4. How often do operator corrections identify candidate-discovery failure versus final-selection failure?
  5. Which classes of software changes can Treebeard implement reliably under locked acceptance criteria?
  6. What is the minimum external enforcement surface required to give an Adaptive Mirror meaningful development freedom without giving it a viable path to promotion or policy control?
  7. For a future Treebeard OS, which administrative operations can be delegated safely, and which must remain in a superior privilege domain?
  8. Which lifecycle properties depend on scheduled services, and how is failure of those services independently detected?
  9. How should synthetic probe traffic be separated from real work so cost and recovery metrics remain defensible?

14. Product Takeaway

The differentiator in Treebeard is not that it uses multiple models or multiple agents. Those are implementation details.

The useful product idea is the separation between routing, authorization, execution, evidence, and promotion. A routing signal does not become permission. A model's narration does not become proof that something ran. A new capability does not become production authority because it passed once.

Stage 1 turns those ideas into a qualified Router 2 migration. If that works, Stage 2 asks whether Treebeard can improve a copy of itself inside a boundary it does not control. Only repeated success there would justify the Stage 3 research thesis.

The immediate job remains straightforward: characterize the working router, build Router 2 one bounded component at a time, measure the differences, and promote only what the evidence supports.

No comments:

Post a Comment