Wednesday, September 23, 2026

Treebeard Router 2 — Technical Architecture Case Study


Project Treebeard

Router 2

Governed Cognitive Routing for Agentic Systems

Technical Architecture Case Study

Thomas C. Ricks

Limited portfolio excerpt — architecture, not implementation specification.


Limited Portfolio Excerpt This document describes selected architectural principles from Project Treebeard. Implementation details, operational configuration, proprietary evaluation material, and reproduction-critical specifications have been intentionally omitted.

Project Treebeard Router 2: governed cognitive routing for agentic systems.

1. Executive Summary

The problem

Project Treebeard is a self-hosted, multi-agent AI system. Work arrives as conversational turns from an operator and has to be handled by the right mission, specialist, tool, or model. Treebeard already has a working router ("Router 1"). It supports operator-directed routing, explicit mission references, deterministic lexical selection, thread continuity, a memory-assisted tie-break, model-assisted fallback, provider fall through, and routing telemetry.

Router 1 works, but it has reached the limits of its shape. Its decisions are hard to explain after the fact. Its confidence is implicit. The cost of thinking about a decision is not separated from the cost of doing the work. Authority checks, candidate selection, and model choice are entangled in ways that make each harder to change safely. Router 1 has no mechanism by which a better router could prove it is better before being given authority over real work.

The solution

Router 2 is a proposed replacement, designed as cognitive middleware rather than a model selector. For every incoming turn it determines:

  • what kind of work the turn represents;
  • which mission or capability owns it;
  • whether that conclusion is obvious, tied, ambiguous, or genuinely absent;
  • how much reasoning the routing decision itself deserves;
  • whether the result is permitted;
  • what execution capability the work requires;
  • what happens if the preferred path fails;
  • what evidence survives so the decision can be explained, replayed, corrected, and learned from.

The architecture is organized around a Tri-Cortex model:

  • a fast deterministic Intuition layer;
  • a bounded semantic Reasoning layer;
  • a historical Subconscious/Memory layer.

These three are surrounded by single-owner components for mission arbitration, actionability, cognitive budgeting, governance, execution shaping, and provider delivery.

The design principles that matter most

Cheap authoritative evidence before expensive inference. The router consumes evidence in a fixed order, from explicit operator intent down to bounded semantic reasoning, and stops as early as correctness allows. The common case involves no model call at all.

One owner per consequential decision. Three different components answer three different questions: which candidate currently wins, whether that result is resolved enough to act on, and how much more routing cognition may be spent. No component answers another's question.

Authority is structurally separate from inference. Only the governance layer can write authority. Confidence, memory, model reasoning, learning, and provider selection have no contractual path to create or broaden permission.

Autonomy is earned, not declared. Every component, including components generated by Treebeard's own agents, passes the same mechanical gates: contract tests, invariant tests, replay against historical decisions, shadow operation, and evidence packets that cannot hide a failed hard gate.

Migration is incremental. Router 2 enters by controlled replacement of a live system. Router 1 is characterized first, then wrapped. Router 2 runs in shadow and earns authority one responsibility at a time. Router 1 is retired only when evidence justifies it, and retirement must reduce complexity.

Status of the work

Router 2 exists as a frozen architecture. It was developed through independently designed section studies, three unified drafts, an adversarial review of the third draft, and a final consolidation. The consolidation resolved every identified ownership conflict and moved every implementation-dependent fact into a characterization question.

Router 2 is not implemented. The architecture explicitly defers all numeric thresholds, latency budgets, and promotion criteria until the existing router has been characterized. The next step is that characterization: 36 questions across ten areas, to be answered by inspecting Router 1 rather than by assumption.

This document describes a design, its reasoning, and its validation strategy. It does not describe a running system.


2. The Product Problem

Why agentic systems need routing at all

A single-model chatbot has one decision to make: what to say next. An agentic system has a prior decision that is easy to overlook: who should handle this, with what, under whose authority?

In Treebeard, work is organized into missions: bounded, long-running objectives with their own scope, context, specialists, and permissions. A turn from the operator might:

  • continue the current mission;
  • start work in a different one;
  • issue a control command to the system itself;
  • ask a general question that belongs to no mission;
  • or explicitly redirect a previous decision.

Getting this wrong is not a cosmetic failure. A turn routed to the wrong mission operates on the wrong context, may use the wrong tools, and in the worst case acts under the wrong authority.

As the number of missions, specialists, tools, and model providers grows, so does the routing surface. Each new capability adds another place a request could plausibly go and another way for it to be mis-routed.

Why "pick the best model" is not the answer

Much of the industry conversation about routing is really about model selection: send easy prompts to cheap models and hard ones to expensive models. That is a legitimate optimization. It is also the last decision in a correct routing pipeline, not the first.

Model selection cannot know which mission owns a request, whether the operator explicitly assigned it, whether the route is permitted, or whether the current conversation thread already settled the question. If those questions are answered implicitly, inside a prompt or as a side effect of which model happened to be chosen, the system becomes fast, cheap, and unaccountable.

Router 2 therefore answers product-semantic questions first (ownership, authority, execution shape) and chooses a model only after those are fixed. Provider choice happens late enough that product intent is already established.

Why autonomy turns routing into a governance and integration problem

Autonomy changes the risk profile of routing. When a human is in the loop for every action, a mis-route is annoying. When agents act on their own, a mis-route can trigger an external action, consume budget, or cross a permission boundary before anyone notices.

Three failure patterns motivated the architecture more than any others.

Inference quietly becoming authority. A model infers that a request "probably" belongs to a mission. Memory recalls that similar requests went there before. Neither is permission, but if the system has no structural separation between where something belongs and whether it may act, a confident inference can behave like an authorization.

"The agent said it succeeded." Agentic systems are prone to accepting self-reported success. A test suite exits zero because it discovered zero tests. A component claims completion without evidence. Treebeard's own history includes this failure mode, and the architecture treats it as a first-class design constraint rather than a process reminder.

Uncontrolled self-modification. A system that learns from its own outcomes can drift, oscillate, or learn from its own hypotheticals. If learned adjustments rewrite live behavior directly, there is no way to evaluate a change before it takes effect and no clean way to undo it.

Router 2 is designed so that each of these failures has no structural path to occur, rather than being guarded against by policy prose.


3. Design Principles

Eight principles govern the architecture. They are stated here once and applied throughout rather than repeated in each section.

3.1 Cheap authoritative evidence before expensive inference

The router consumes evidence in a fixed order:

explicit operator / invariant evidence
  → deterministic evidence
  → fresh contextual evidence
  → bounded historical evidence
  → bounded semantic reasoning
  → genuine no-match or clarification

Each step is more expensive and less authoritative than the one before it. The pipeline stops as soon as correctness allows. This single ordering improves four things at once:

  • Latency, because most turns resolve on local deterministic evidence.
  • Cost, because model calls are reserved for real ambiguity.
  • Predictability, because deterministic decisions stay deterministic.
  • Governance, because the most authoritative signal, the operator's explicit instruction, is examined first and can never be outvoted by inference.

3.2 One owner per consequential decision

Every consequential decision in the architecture has exactly one owner, recorded in an ownership matrix, and no implementation may create a second owner implicitly. This was the single most productive discipline in the design process. Most contradictions found during review were ultimately two components believing they owned the same decision.

3.3 Authority is separate from inference

Routing answers where does this belong? Governance answers is this permitted? They never share a function, and only governance can write authority state.

3.4 Bounded cognition

Every expensive operation is behind a gate. Reasoning requires an escalation verdict, an explicit budget, and new evidence since the last pass. There is no path by which the system can think in a loop.

3.5 Evidence and provenance over assertion

Every routing claim is an evidence item with a source, a class, and a producer. Every decision records the exact effective configuration it ran under. Success is established by recorded evidence, never by a component's report about itself.

3.6 Graceful degradation, hard boundaries

Optional intelligence (memory, reasoning, learning, detailed telemetry) degrades when it fails. Mandatory correctness and authority boundaries fail closed. An optimization failure must never become a governance failure.

3.7 Reversible learning

Learning produces versioned, evaluated, atomically activated artifacts that can be rolled back. It never mutates the decision currently being made, and it never rewrites raw history.

3.8 Incremental migration

The architecture is designed for an existing working system. Existing behavior is characterized before it is replaced. Authority is transferred one responsibility at a time. Temporary migration machinery carries an explicit deletion condition from the day it is created.


4. Architecture at a Glance

4.1 The overall shape



Router 2 treats the cortices as evidence producers, separates mission selection from confidence and cognitive budgeting, and places governance before execution and delivery.

The cortices produce evidence. The central band turns evidence into a decision. The boundary turns the decision into a permitted commitment. The lower band turns the commitment into work. Telemetry and learning observe everything but decide nothing in-line.

4.2 Component responsibilities

Each consequential component has a narrow job and an explicit non-job:

  • Router Core answers: In what order do things happen, and what is the one committed result? It never scores, ranks, reasons, or authorizes.
  • Intuition (Cortex 1) answers: What does cheap, local evidence say about this turn? It never calls a model, performs a store read, or commits a mission.
  • Reasoning (Cortex 2) answers: What does bounded semantic analysis add? It never names missions outside its supplied candidate set, authorizes, or dispatches.
  • Subconscious / Memory (Cortex 3) answers: What does history say about these candidates? It never creates candidates or authority and never rewrites raw history.
  • Mission Selection answers: Which candidate currently wins under precedence? It does not decide whether that result is good enough to act on.
  • Confidence & Escalation answers: Is the current result actionable, and would more evidence help? It does not call components, choose missions, or grant permission.
  • Cognitive-Level Controller answers: How much routing cognition may be spent next? It does not judge confidence or select a model by name.
  • Consequence Classifier answers: How costly would a wrong routing decision be? It does not use model inference.
  • Governance answers: Is this proposed route permitted? It does not consider routing confidence as a permission signal.
  • Execution Shaping answers: What kind of execution does the committed work require? It never reopens the routing decision.
  • Model / Provider Selection answers: Which provider/model can realize the required capability under policy? It never decides what capability is required.
  • Telemetry answers: What happened, exactly, under what effective state? It does not influence the live decision.
  • Learning answers: What should future behavior be? It never mutates the live decision or raw history.

4.3 The canonical pipeline

The architecture defines a single canonical stage order that every section uses:

  1. Intake. Normalize the turn once and capture one coherent set of immutable state snapshots.
  2. Control classification. Distinguish work turns from control commands.
  3. Intuition. Produce fast evidence: operator directives, deterministic lexical evidence, thread continuity, and pre-assembled historical hints.
  4. Mission Selection. Construct and rank candidates under precedence.
  5. Consequence classification. Assign a risk class.
  6. Confidence and escalation. Accept, resolve a tie, escalate to reasoning, conclude no-match, or ask.
  7. Governance validation and commit. One atomic step.
  8. Execution Shaping.
  9. Model/provider selection.
  10. Pre-dispatch receipt gate. Conditional; applies only to a narrow, policy-defined class.
  11. Delivery, with bounded fall through.
  12. Outcome.
  13. Telemetry seal and feedback.

Two properties of this pipeline are worth drawing out.

First, short-circuiting skips cognition, never boundaries. A turn with an explicit operator assignment may skip historical retrieval and semantic reasoning entirely. It does not skip mission selection, the confidence check, governance, execution shaping, provider policy, or essential telemetry. The fast path is fast because it avoids expensive thinking, not because it avoids checks.

Second, re-entry never restarts the turn. When the memory layer or the reasoning layer produces new evidence, that evidence is appended to the same per-turn working record, and only the stages whose inputs changed rerun. Nothing computed earlier is discarded or silently overwritten.

4.4 The per-turn working record

All state for a routing attempt lives in one structured working record owned by Router Core. Every stage reads from it. No stage writes to it directly: stages return typed results, which Core validates and applies. Evidence within the record is append-only.

This design choice exists to prevent the most common failure of pipeline architectures, where stages pass loosely structured state to each other and gradually acquire hidden co-ownership of fields. With one writer and one record, every piece of state has an unambiguous origin, and the record doubles as the decision's provenance.

Three artifacts are deliberately kept distinct:

  • the working record, mutable until commit;
  • the routing decision, immutable and created only at the governance/commit boundary;
  • the telemetry decision record, a persistent projection that references the other two but is never the authoritative runtime object.

Conflating the runtime decision with its telemetry record is a subtle but common error. It lets the monitoring system become an unintended part of the decision path.


5. The Tri-Cortex Model

The three cortices are architectural responsibility domains, not three processes and not three model calls. In the common case, only Intuition does any work, and it involves no model at all. The metaphor is useful because it maps to three genuinely different kinds of evidence with different costs, re-liabilities, and failure behaviors.


The Tri-Cortex model separates fast deterministic evidence, bounded semantic reasoning, and historical memory while keeping governance and execution responsibilities outside the cortices.


5.1 Intuition — the fast path

Intuition produces cheap evidence about the current turn from information already in hand. It is deterministic over its declared inputs, model-free, performs no store reads beyond the turn's immutable snapshots, and is exactly re-playable.

It evaluates, in order:

  • Operator directives. Current-turn assignments, standing assignments, explicit mission references, exclusions, and corrections.
  • Deterministic lexical evidence from the existing selector.
  • Thread continuity. Is this a follow-up to the mission the conversation already belongs to?
  • Pre-assembled historical hints, such as a recent operator correction or a known pair of commonly confused missions.
  • Negative evidence.

The most consequential design decision inside Intuition was creating a single, sole producer of directive evidence. Earlier drafts allowed operator-direction signals to be recognized in several places. The final architecture routes all interpretation of operator intent through one normalizer that distinguishes an assignment ("do this under Mission A") from an incidental mention ("like we discussed in Mission A").

Treating every mention as an assignment is a surprisingly common routing bug. It is also a quiet way for user content to steer a system.

Intuition emits evidence, not decisions. It never constructs candidate objects and never commits a mission. It may offer a fast recommendation, but that recommendation carries no special standing downstream.

Historical hints available to Intuition are bounded. They come from a snapshot that the memory layer assembles off the routing path, so Intuition never waits on retrieval. These hints can veto fast acceptance ("the operator corrected this exact choice two turns ago") but can never nominate a candidate, reorder deterministic results, or promote weak evidence into strong evidence.

5.2 Reasoning — the deliberative cortex

Reasoning is a bounded semantic adjudicator, invoked only after the confidence controller escalates and the cognitive-level controller issues a budget. It never receives an open question like "figure out where this belongs." It receives a typed routing question over a fixed candidate set whose permissible answers are defined before the model runs. The question types include:

  • disambiguate between named candidates;
  • continue the current mission or switch;
  • check whether a lexical match is a false positive;
  • resolve conflicting evidence.

Normally the candidate set is the top three. When the escalation concerns candidate recall, the mission selection layer, not the model, may extend the set by a small number of additional, already-eligible, low-prior candidates. An earlier draft allowed the reasoning layer itself to "restore" candidates from a deferred pool. That capability was removed in the final architecture. Any operation that lets a model decide which options are on the table is a step toward a model manufacturing scope.

Reasoning's output is structured evidence, not prose. Every semantic claim must cite a source reference to something actually present in its input:

  • what the operator intends;
  • what the request is about;
  • what a pronoun refers to;
  • how well the request fits a mission's declared purpose;
  • whether the topic has shifted.

Uncited claims are discarded by deterministic validation before the result is used. The model's self-reported confidence is never an input to routing confidence. Model output is data until validation passes.

Context is fed through a ladder. The minimal context comes first. Recent thread summaries or candidate objectives are added only when the model identifies a named information gap. Targeted historical evidence is added only when a specific historical gap remains. There is no general "load the conversation" path. A single authorized invocation is limited to a small, fixed number of model passes, and a reasoning failure yields zero new evidence while leaving the fast-path result intact.

Reasoning does not call providers directly. Its model requests go through the same provider selection and fall-through subsystem as all other model work, inheriting capability enforcement, cost policy, health awareness, and failure typing. An earlier section design had its own model-calling contract. Consolidating onto one provider stack eliminated what would have become a second, divergent delivery system.

5.3 Subconscious / Memory — historical experience

The memory cortex provides historical evidence in two planes and several latency tiers.

The raw routing ledger is an immutable record of what happened: candidates, decisions, bases, corrections, outcomes, provider failures, specialist rejections. It is never rewritten, including by learning.

Derived routing memory holds versioned, re-buildable artifacts computed from the ledger:

  • thread continuity state;
  • commonly confused mission pairs;
  • mission affinities;
  • correction summaries;
  • reviewed durable routing rules.

Anything derived can be thrown away and regenerated from raw history.

This separation matters most when something goes wrong. If a bad period of data is identified, such as a failed experiment, a telemetry bug, or a temporary mission definition, its events can be quarantined by range and the derived artifacts rebuilt without them. Recovery from contaminated learning is only possible if facts and interpretations are physically distinct.

The latency tiers keep memory off the hot path:

  • Tier 0: a small pre-assembled snapshot containing the current thread owner, last decision, recent correction, bounded cached priors, and staleness flags. It is assembled off-path and read by Intuition at effectively zero marginal retrieval cost.
  • Tier 1: bounded, question-shaped retrieval scoped only to named candidates. It runs only after ambiguity exists and only under an explicit budget.
  • Tie-break: a bounded mechanism for choosing among mechanically attested genuine ties. It is never a general ambiguity resolver.

There is no "hydrate all memory" operation anywhere on the routing path.

Label quality is modeled explicitly. An explicit operator correction is strong evidence. A route that was never complained about is weak evidence, since silence is not approval. Every derived artifact records the composition of the evidence it was built from, so aggregation can never launder many weak labels into something that looks like a strong one.

The tie-break. Treebeard has an existing memory-assisted tie-resolution mechanism called Dreams. Router 2 preserves it but narrows it sharply. Dreams may run only when the confidence controller certifies, through a mechanical predicate, that a candidate set is a genuine tie. The input is only the tied candidate IDs and a neutral summary of the turn. The output is either one of those candidates or an abstention. It runs at most once per tie cycle, and its result can never create authority.

The distinction between a genuine tie and an unresolved conflict is one of the more important ideas in the design. It is covered in the next section.

5.4 When each cortex runs

The invocation pattern is intentionally asymmetric:

  • Explicit operator assignment: Intuition resolves it; memory and reasoning are not invoked.
  • Clear deterministic winner: Intuition resolves it; memory and reasoning are not invoked.
  • Obvious follow-up in a fresh thread: Intuition resolves it with the Tier-0 memory snapshot only.
  • Two genuinely symmetric candidates: Intuition identifies the state; bounded historical evidence and the tie-break may run; semantic reasoning is used only if the tie-break abstains.
  • Deterministic evidence conflicts with fresh thread continuity: Intuition records both; bounded memory may contribute; Reasoning may adjudicate.
  • No candidate has meaningful evidence: Intuition reports the gap; bounded memory may contribute; Reasoning gets one bounded attempt before the system concludes no-match.

The expected operating distribution, which the architecture requires be measured rather than assumed, is that the large majority of turns resolve on the fast path.


6. Mission Selection, Confidence, and Cognitive Budgeting

This separation is one of the strongest ideas in the architecture. Three components answer three different questions, and conflating any two of them produces predictable failures.

The separation is easiest to remember as three questions:

  • Mission Selection: Which candidate currently wins under the evidence and precedence rules? This is pure arbitration.
  • Confidence & Escalation: Is that result sufficiently resolved to act on, and would more evidence help? This is pure policy evaluation.
  • Cognitive-Level Controller: If more routing cognition is justified, how much may we spend? This is pure budget allocation.

None of the three performs I/O, calls a model, or touches authority. They are pure functions over the working record and versioned policy. That makes them fast, exactly re-playable, and testable without any model or store present.

6.1 Mission Selection: arbitration under precedence

The mission selection engine alone constructs candidates, admits them, ranks them, applies precedence, and declares the provisional result: a sole winner, a tied set, an unresolved conflict, or no match.

Precedence is strict across tiers:

  1. structural invariants and control classification;
  2. current-turn operator assignment;
  3. standing operator assignment;
  4. explicit mission reference;
  5. deterministic lexical or structural evidence;
  6. fresh thread continuity;
  7. valid historical and tie-break evidence;
  8. validated semantic evidence;
  9. genuine no-match;
  10. clarification.

A lower tier never silently overwrites a higher one. A model verdict never overrides stronger deterministic evidence without a recorded, policy-defined reason.

Two design choices here are worth noting for anyone building similar systems.

Precedence tiers, not a unified score. The architecture rejects adding evidence types into a single weighted score. Evidence classes are not commensurable. A weighted sum means that enough historical affinity can, in principle, outvote an operator's explicit instruction, and a summed score cannot be explained after the fact. Within a tier, comparison is allowed. Across tiers, the higher tier dominates.

Directed targets are preserved, even when blocked. If the operator explicitly directs work to Mission A, and Mission A turns out to require authorization, the result is authorization required for Mission A. The router never quietly routes to the runner-up because the requested target is inconvenient. Substituting a permitted alternative for a blocked requested one looks helpful. In fact it treats a governance outcome as routing evidence, which is precisely the boundary the architecture exists to keep clean.

6.2 Confidence: structured, not a single number

Confidence is represented as structured evidence rather than a single opaque probability. It records several distinct things:

  • a grade;
  • the class of the strongest supporting evidence;
  • how separated the leader is from the runner-up;
  • whether the evidence agrees or conflicts;
  • how fresh it is;
  • whether any evidence sources were degraded or missing.

The reason is concrete. "The operator named this mission" and "a lexical scorer produced a very high score" might both look like 0.99 on a scalar scale, but they must behave differently. The first survives contradicting lexical evidence. The second does not survive a recent operator correction. A scalar also hides why confidence is low, which is exactly what the next stage needs to know.

Evidence classes are ordered from authoritative through deterministic, contextual, and historical to probabilistic. Only validated authoritative evidence can reach the top confidence grade. Multiple weak signals do not accumulate into a strong one. Historical evidence can refine or cap confidence but cannot independently produce a strong route.

All confidence behavior lives in one immutable, versioned policy made of declarative tables. It is not scattered across conditional logic in multiple modules. Every decision records the policy version it ran under, so historical decisions can be re-evaluated under alternative policies during calibration.

6.3 Genuine ties versus unresolved conflict

A candidate set is eligible for the tie-break only if a mechanical predicate holds. In substance, the candidates must be:

  • supported by the same highest evidence class;
  • within the configured tie band;
  • free of open conflict or semantic-divergence markers;
  • equivalent in governance eligibility and execution surface;
  • undifferentiated by thread continuity or any higher-precedence directive.

If any required field is missing, the default is not a tie.

If two missions are equally plausible and would do materially different things, the situation is not a tie. It is unresolved ambiguity, and it goes to bounded reasoning. Only genuinely symmetric choices, where the difference does not matter much, go to the tie-break. An earlier draft left "genuine tie" as a semantic judgment. The final architecture made it mechanical and conservative, because a tie-breaker applied to real semantic conflict produces confident-looking coin flips.

6.4 Cognitive budgeting: how much thinking is worth it

When the confidence controller decides more evidence is justified, the cognitive-level controller decides how much. It issues a budget from a small set of classes:

  • Fast. No model reasoning.
  • Tie resolution. Bounded historical retrieval and at most one tie-break; no model.
  • Deliberative. One bounded semantic pass, optionally one context-enriched re-evaluation.
  • Deep. Expanded but still bounded.

A budget specifies capability class, context allowance, candidate limit, retrieval allowance, pass count, deadline, and cost ceiling. It never names a model; that belongs to provider selection.

Every turn starts at the fast level. Escalation is monotonic and bounded, and three mechanisms make looping structurally impossible:

  1. Each escalation step has a hard upper limit on depth.
  2. The no-new-evidence rule. If the previous cognitive pass appended no new evidence to the working record, another equivalent escalation is forbidden. This is checked independently by both the confidence controller and the cognitive-level controller.
  3. The deepest level is terminal.

Escalation also considers consequence. A small declarative classifier assigns each decision a risk class based on:

  • action class;
  • external side effects;
  • reversibility;
  • the impact of every plausible candidate, not just the leader.

The same uncertainty may be accepted with an acknowledged-uncertainty flag on a low-consequence route and escalated on a high-consequence one. A marginal route whose mistake would cost one follow-up message is not worth a model call.

6.5 Routing cognition is not execution cognition

The architecture permanently separates how hard the router thought about where work belongs from how capable the worker that does the work must be. A difficult task may be trivially routed: "redesign the billing service" under an explicit mission assignment needs no routing cognition at all. A trivial task may be hard to route: a five-word follow-up that could belong to either of two active missions.

Conflating the two leads to deep routing for every hard task, and to cheap workers for every easily routed one.

6.6 A simplified decision sequence


Cheap, authoritative evidence is evaluated first; more expensive historical or semantic inference is invoked only when the route remains materially unresolved.

Clarification is a typed last resort. It is permitted only when:

  • internal evidence has been exhausted or is unavailable;
  • the ambiguity materially changes what the system would do;
  • the missing information exists only in the operator's head.

A clarification must state the candidate interpretations, the material distinction between them, and why internal evidence could not resolve it. The router should not ask the operator to repeat context that Treebeard already possesses merely because a cheap component failed to retrieve it.


7. Governance and Bounded Autonomy

7.1 Routing is not permission

The architecture draws one hard line. Routing decides where work belongs. Governance decides whether that route may act. Confidence is never an input to permission. A route the router is certain about has exactly as much authority as one it is barely sure of, which is to say none until governance says otherwise.

This is enforced structurally rather than by instruction. Only the governance layer can produce authority state, and no other component's output contract contains a writable authority field. As a result, none of the following can create, broaden, or substitute for authority: mission inference, historical memory, the tie-break mechanism, model reasoning, learning, or provider selection. This is not because each one is told not to. Their outputs simply have no place to put a permission. One structural rule replaces a long list of prose prohibitions, and one test verifies it.

Governance interacts with routing at two points:

  1. A lightweight eligibility snapshot, used during candidate construction. Missions that are structurally unavailable for inferred routing are never ranked in the first place, so inference cannot even see them.
  2. Commit-time validation of the actual proposed target. Stale eligibility information can never become unsafe authority.

7.2 The commit boundary

An intermediate draft of the architecture committed the routing decision first and validated authority afterward. Review identified the flaw. Between those two steps there existed an "immutable" decision that might not be permitted. Every downstream component would need to remember to check.

The final architecture makes proposal, governance validation, and commit one atomic boundary:



Routing proposes where work belongs. Governance decides whether that route may act. Only authorized executing dispositions cross the commit boundary.

The dispositions split into two classes. Executing dispositions dispatch work and require authorization at commit. Non-executing dispositions dispatch nothing and claim no authority. These include:

  • authorization required;
  • blocked;
  • clarification required;
  • router failure.

They remain safe to commit even when governance is unavailable. A clarification response is non-executing: it does not consume the proposed work target’s execution authority, and it can therefore be committed safely while the system waits for operator input.

Governance unavailability is terminal-safe. If authority cannot be established at commit, no production dispatch occurs, and no fallback path can be used to bypass that. The legacy router may take over only if it independently performs valid governance validation through a working authority path.

Governance uncertainty never degrades into permissive behavior, and an ambiguous authorization result is treated as an absent one.

7.3 Bounded autonomy: where humans belong

The goal of governance in Router 2 is not to require human approval for everything. A system that asks permission for every routine action can reduce oversight to habitual rubber-stamping, which defeats the purpose of human review.

The goal is to locate human judgment at consequential boundaries and let safe, deterministic, reversible work proceed on its own. The architecture ties this together through three mechanisms.

Consequence classification identifies which decisions are high-impact: externally visible, mutating, governance-affecting, or hard to reverse. High consequence raises the evidence required to act and justifies deeper routing cognition. It never manufactures authority.

The established authorization path. When a route requires authorization it does not currently have, the router surfaces it to the operator through Treebeard's existing authorization path, preserving the operator's original target. It does not silently refuse and does not silently substitute. A grantable target remains distinguishable from a hard denial.

Governance does not manufacture resistance. The architecture explicitly states that ordinary operator instructions that change priorities or routing are not treated as attempts to bypass governance. A control system that interprets legitimate direction as a threat is as broken as one that grants anything asked. The authority model enforces the boundaries actually specified, in both directions.

The result is that human attention is spent on what humans are uniquely positioned to decide:

  • granting authority that does not yet exist;
  • resolving ambiguity that only the operator can resolve;
  • approving the promotion of new behavior into production.

Everything else runs autonomously and leaves evidence behind.

7.4 Provider selection cannot bypass authority

Governance supplies provider selection with declarative constraints, such as allowed and disallowed providers, local-only requirements, data classes, and tool restrictions. Provider selection enforces them by eliminating candidates that violate them. It does not interpret policy, and no cost or latency consideration can re-admit an eliminated option. Because provider selection runs after the commit boundary, authority is already fixed by the time any model is chosen.


8. Execution Shaping and Capability Routing

8.1 Routing decision versus execution decision

Knowing where work belongs does not determine how it should be done. A request committed to a mission might be handled in several ways:

  • directly;
  • by the mission's general worker;
  • by one of its specialists;
  • by a deterministic tool;
  • by a combination of these.

Each of those implies a different capability requirement, side-effect profile, and cost.

In early drafts, this translation lived in unnamed glue between the router and the executor. The adversarial review flagged it as an ownerless decision. The final architecture makes Execution Shaping a first-class subsystem with a single owner.

8.2 What Execution Shaping owns

Execution Shaping receives the committed routing decision plus the mission's declared execution metadata: specialists, tool requirements, side-effect declarations, governance capability restrictions, and the consequence class. From these it produces an execution shape, a complete contract the executor consumes without re-deriving anything. The shape covers:

  • the handling mode and target;
  • the worker capability required;
  • the external side-effect class and idempotency requirement;
  • the data-handling class that constrains provider choice;
  • the policy for what happens if the required capability cannot be delivered;
  • whether a durable decision receipt is required before dispatch.

That last item is a narrow, policy-defined class, discussed in Section 10.

8.3 When the specialist is unclear

Deterministic shaping from declared metadata comes first. If several specialists remain materially plausible, the shaper can request a single, bounded semantic adjudication. It uses the reasoning cortex's machinery, under a budget issued by Execution Shaping, not by the routing cognitive controller. The result comes back as evidence only, and the shaper makes the final decision.

This detail was one of three corrections applied at the design freeze. Earlier, specialist selection was budgeted by the routing cognition controller, which quietly merged routing cognition with execution adjudication. Separating the budgets keeps a post-commit execution question from ever reopening the routing decision.

8.4 Degraded capability is a declared policy

For each shape, the shaper declares in advance what happens if the required capability cannot be delivered:

  • no downgrade permitted;
  • one tier down permitted;
  • defer the work;
  • fail.

Provider delivery reports what was actually delivered, and the router applies the declared policy. No component makes an ad hoc judgment about whether "close enough" is acceptable at the moment of failure, which is exactly when such judgments are worst.


9. Models, Providers, and Failure-Aware Delivery

9.1 Provider independence

The provider layer realizes a capability requirement. It never decides what capability is required, and it never names or depends on a specific mission.

Provider-specific behavior is confined to adapters. Nothing above the adapter boundary sees a provider SDK type, a provider error code, or a provider-specific finish reason. The test for leakage is mechanical: if any code outside an adapter names a provider, the boundary is broken.

The product consequence is significant. Changing providers, adding a local model, or dropping a vendor does not require redesigning the product. Product intent is defined in capability terms, and the provider layer is a replaceable realization of those terms.

9.2 Deterministic, snapshot-based selection

Selection is a pure function over four immutable inputs:

  • the capability requirement;
  • the capability registry;
  • the provider health snapshot;
  • provider and cost policy.

No network probes happen during selection. Given identical inputs, selection is bit-identical, and every decision records the snapshot versions it used.

Hard constraints eliminate. They include capability floor, structured-output reliability, context size, data boundary, and policy. Soft preferences rank the survivors. The ranking is deliberately lexicographic rather than a weighted score:

  1. health;
  2. then coarse reliability;
  3. then the configured cost or locality strategy;
  4. then a stable tie-break.

This choice was flagged for challenge in review and retained. A lexicographic ordering can be explained in one sentence per decision. It cannot be quietly tuned into letting a large cost advantage buy back a reliability deficit. It is directly testable stage by stage. Weighted trade-offs may be introduced later if, and only if, evaluation shows they outperform the simpler policy.

After hard capability, policy, health, and reliability constraints are satisfied, the selection policy may prefer the lower-cost viable option. Intelligence is not equated with the most expensive model. A stronger model serving a lesser requirement is permitted but flagged as over-provisioned, so that the pattern "we are paying frontier prices because cheap tiers are unhealthy" is visible as a fleet-health signal.

9.3 Bounded, typed fall-through

The provider layer builds a finite, ordered fallback chain once, at selection time. When an attempt fails, the failure is classified into a named failure type, and that classification determines how the remaining chain is used. The failure types include:

  • timeout;
  • quota exhaustion;
  • provider outage;
  • empty output;
  • structurally invalid output;
  • reasoning exhaustion;
  • context overflow;
  • unexpected refusal.

A quota failure skips the rest of that provider rather than marching into the same quota. Reasoning exhaustion prefers stronger capability. Invalid structure prefers models with stronger structured-output reliability. An authorization failure is never retried.

The chain is never extended at runtime, previously failed entries are never re-entered, and retries share one budget across the whole plan. Termination is guaranteed by construction: there is no loop, only a finite list walked once.

The central semantic rule is that transport success is not semantic success. A provider that returns HTTP 200 with an empty body, a truncated structured output, or output that fails its schema has failed. Distinguishing these cases matters because each suggests a different remedy. Collapsing them into "no usable answer" would throw away exactly the signal fall-through needs.

9.4 No silent downgrade

If the required capability cannot be supplied, the provider layer returns an explicit degraded-capability result stating what was required and what was delivered. It is never reported as success. The execution shape's declared policy decides what happens next.

This rule is structural, not procedural. The flag is set by the same logic that permits a lesser option, so there is no code path that produces a downgrade without producing the record.


10. Evidence, Telemetry, and Replay

10.1 Why "the agent said it succeeded" is not verification

Agentic systems generate a great deal of self-description. Components report success, tests report green, models report confidence. None of these is evidence. Router 2 is built on the principle that every consequential claim must be reconstructable from recorded facts.

That principle drives four design features.

One canonical evidence contract. Every routing claim is an evidence item:

  • Intuition's lexical result;
  • a historical correction;
  • a reasoning conclusion;
  • a tie-break result.

Each item records its source, its producing component and version, its evidence class, what it is about, whether it supports or opposes, how fresh it is, and the configuration epoch it was produced under. Earlier drafts allowed the memory layer its own evidence language. Consolidating to one contract meant every consumer could reason about every producer uniformly.

Append-only evidence. Within a routing attempt, evidence is only ever added. Deterministic results are recorded verbatim and never rewritten by later stages. When reasoning disagrees with a lexical result, both remain in the record, along with which prevailed and why.

Effective-state provenance. Recording "router version 2.3" is insufficient once behavior depends on configuration, policy tables, and learned artifacts. Every decision references the exact effective state it ran under:

  • component manifest;
  • configuration;
  • confidence policy;
  • learned artifacts;
  • mission roster;
  • governance policy;
  • provider registry and health snapshot.

Frequently reused states are content-addressed, so a decision stores a reference rather than a copy. Historical behavior is never inferred from dates or deployment logs.

Immutable decisions, append-only outcomes. The original decision record is never edited. Later facts attach to it by reference:

  • execution success or failure;
  • an operator correction;
  • a specialist rejection;
  • an actual billed cost.

Explicit operator corrections are captured at the point they happen. The operator is not left to hope the system infers them from silence. The system can always answer both "what did the router believe at the time?" and "what turned out to be true?"

10.2 The flight recorder

Telemetry is designed as a flight recorder with a strict split:

  • In-path, a bounded in-memory record builder. No filesystem writes, no network calls, no heavyweight serialization.
  • Off-path, an asynchronous emitter and a local-first store supporting queries, aggregation, and replay.

Telemetry has its own measured overhead budget. If it becomes material, telemetry is optimized; routing is not distorted to accommodate it.

Ordinary telemetry fails open: a metrics outage never stops routing. The one exception is a narrow, governance-defined class of actions that must have a durable decision receipt before any external action. That class passes through an explicit pre-dispatch receipt gate that fails closed for that class only. It is declared in the execution shape, measured as hot-path cost, and structurally unable to spread to ordinary routing.

A recurring design risk in audit-heavy systems is that "must be logged" slowly becomes the default for everything. This design is built to prevent that.

10.3 What the flight recorder must answer

For any decision, the recorded evidence must be able to answer the following:

  • what came in and what evidence was available;
  • what the top three candidates were at each stage and which stage changed the ranking;
  • why escalation happened, or why it did not;
  • what configuration was actually active;
  • what governance decided;
  • what execution shape and provider plan were chosen;
  • what failed, and how fall-through proceeded;
  • how long each stage took and what it cost;
  • whether the operator later corrected the decision.

The top-three structure is instrumented separately. This lets two different failures be distinguished:

  • Candidate generation failed: the correct answer was never among the finalists.
  • Selection failed: the correct answer was present but not chosen.

Those two failures have different owners and different fixes. A single accuracy number cannot tell them apart.

10.4 Replay

Because deterministic stages are pure over recorded state, a sealed decision record plus its referenced snapshots is sufficient to exactly replay every deterministic routing stage. This is the foundation of regression testing and of learning evaluation. "Does today's code make yesterday's decision from yesterday's inputs?" becomes a mechanical question.

Model-assisted stages are handled honestly. They can be replayed from recorded responses, which tests everything around the model deterministically, or re-queried live. Live re-query is an explicitly labeled approximate experiment. It is never presented as a reproduction, because it conflates model drift with router change.

10.5 Privacy

Telemetry stores structured routing facts and references, not duplicated conversation. Operator input appears as fingerprints and references. Nothing from missions uninvolved in a decision is recorded. Credentials and raw model reasoning are never recorded.

Access to content-bearing diagnostic data is a separate capability from reading metrics, so routine engineering diagnosis never requires content access.


11. Learning Without Self-Corruption

11.1 The central rule

Learning may propose future behavior. It never silently rewrites the live decision currently being made, and it never rewrites the history it learned from.

A self-improving router is valuable and dangerous in equal measure. Uncontrolled, it can learn from its own mistakes, over-react to a single bad day, oscillate between configurations, or quietly trade away a previously fixed behavior. Router 2 keeps the value and removes the danger by treating learning as a controlled artifact lifecycle.

11.2 Raw experience, derived artifacts, and corrections

Three kinds of data are kept strictly apart.

  • Raw experience is the immutable routing ledger: what happened.
  • Derived artifacts are versioned interpretations of it, such as lexical profile adjustments, confidence-policy proposals, confusion pairs, priors, and reviewed durable rules.
  • Operator corrections are the highest-trust labels. They are captured explicitly at the moment of correction and linked to the decision they correct.

Several discipline rules follow from this separation:

  • Label strength is carried through every derived artifact.
  • A single correction, however valuable, cannot move a parameter beyond a bounded per-event step.
  • A repeated correction pattern does not automatically become a standing rule; promotion is a separate, gated act.
  • Shadow predictions, decisions that were computed but never dispatched, are never treated as outcomes. A decision that was never executed has no result to learn from.

11.3 The lifecycle

The learning lifecycle is deliberately staged:

OBSERVE → PROPOSE → BUILD CANDIDATE ARTIFACT → REPLAY → TEST → BENCHMARK → ACCEPT/REJECT → ATOMIC ACTIVATE → MONITOR → ROLL BACK IF NEEDED

Every candidate artifact identifies:

  • the evidence it was built from;
  • its version;
  • what built it;
  • the evaluation packet that justified it;
  • its activation state.

Activation is atomic: a request runs entirely under the old artifact set or entirely under the new one, never a mixture. Rollback is a pointer change, cheap enough to be the default response to doubt. A factory baseline is always retained and never modified; it is the floor beneath every rollback.

Promotion requires that a candidate not regress the fixed historical evaluation set on protected case classes, operator-corrected cases in particular. New learning therefore cannot silently undo old fixes. If consecutive promotions keep reversing each other's direction, adaptation freezes pending review, because oscillation is evidence the signal is noise.

11.4 Increasing autonomy is earned by artifact class

This is still self-learning. The difference is where autonomy lives. The architecture defines a staged path:

  1. Automatically generate candidate artifacts.
  2. Automatically run replay, test, and benchmark.
  3. A human approves activation.
  4. After a demonstrated record of success, the lowest-risk artifact classes may activate automatically, with automatic rollback triggers retained.
  5. Only later, and only on evidence, are more consequential surfaces considered.

Precedence rules and governance policy remain outside ordinary learned authority entirely. Learning can make the router better at choosing among permitted options. It can never change what is permitted.


12. Failure Containment

12.1 Why one generic fallback is dangerous

The easiest way to make a routing system "reliable" is a catch-all: if anything goes wrong, route somewhere reasonable. That approach is dangerous because it erases the distinction between very different situations. Some of them demand opposite responses.

The architecture keeps several superficially similar outcomes mechanically distinct:

  • Uncertainty: components worked, but the evidence is insufficient. This is not a failure; the router may escalate within budget, accept with acknowledged uncertainty, or clarify.
  • Optional-component failure: memory, reasoning, or detailed telemetry is unavailable. The router degrades to a named mode and continues using the evidence that remains.
  • Provider/model failure: a delivery attempt failed in a named way. The system uses typed fall-through within the pre-built chain and reports degraded capability explicitly if that chain is exhausted.
  • Authorization failure: the proposed route is not permitted, or permission cannot be established. The system fails closed, surfaces the authorization path, and never substitutes another route merely because it is permitted.
  • Router failure: a mandatory component failed or an internal invariant was violated. The attempt fails closed; legacy fallback is legal only before any external dispatch.
  • Genuine no-match: the router worked correctly and determined that no mission owns the request. That is a legitimate non-mission outcome, not an error.

A generic fallback would turn uncertainty into a guess, a governance failure into permissive routing, and a router defect into an invisible wrong answer. The architecture keeps these outcome classes mechanically distinct. No-match, governance denial, provider failure, and router failure are never conflated in outcomes, telemetry, or learning.

12.2 The preservation hierarchy

When failure forces the system to sacrifice something, it sacrifices in reverse order of this list:

  1. governance and authorization integrity;
  2. routing correctness;
  3. availability;
  4. latency;
  5. cost optimization;
  6. diagnostic richness.

Two corollaries follow. The system degrades capability before correctness: losing memory or reasoning yields a less clever router, never a less honest one. And it degrades optimization before governance: no optimization subsystem sits on the authorization path, so its failure cannot touch authority.

12.3 Degradation is named, not emergent

Degradation happens through a small set of named modes, such as no-learning, reduced-context, no-routing-reasoning, provider-degraded, telemetry-reduced, and legacy-router-fallback. Each mode declares:

  • its trigger;
  • what capabilities remain;
  • what is disabled;
  • how it recovers.

The alternative is a large combination space of quietly half-broken features, which no one can reason about and no test suite can cover. Recovery requires sustained health, not one lucky success, so a marginal dependency does not make the system flicker between modes.

12.4 A decision spine

Failure handling is centralized in the router's orchestration spine rather than scattered through components. Components report typed failures. The spine classifies them, applies policy, and writes the failure record. This concentrates risk in one place, which is a deliberate trade-off. The spine is kept small, deterministic, invariant-checked, and heavily fault-tested, and the legacy router backstops it during migration. Scattered, improvised failure handling is itself treated as a defect.

12.5 The dispatch boundary

The most important boundary in failure handling is the first external dispatch attempt.

Before it, routing is pure computation. It can be safely rerun, and falling back to the legacy router is legitimate. After it, the request may have caused an external effect, so recovery stays within the delivery domain. The request is never restarted through the legacy router, which could duplicate the action.

Dispatch carries an idempotency key tied to the routing decision. Its lifecycle distinguishes planned, attempted, confirmed, failed, and attempted but confirmation unknown. The last state is surfaced explicitly rather than guessed in either direction, so a crash cannot silently repeat an external side effect.

Router 1 fallback, when it happens, receives the clean original request, not Router 2's partially computed state, because that state is suspect by definition. The fallback opens a new, correlated routing attempt. It is one-way: the legacy router never routes back into Router 2.

12.6 Resource pressure

Under resource pressure, the system sheds work in a fixed order:

  1. shadow evaluation;
  2. replay and benchmark jobs;
  3. learning and consolidation;
  4. verbose telemetry;
  5. optional enrichment;
  6. semantic escalation where policy permits;
  7. authoritative routing, last.

Governance checks are never shed into permissive behavior.


13. Testing, Benchmarking, and Earned Autonomy

13.1 The premise: authority is earned, not declared

Nothing in Router 2 is trusted because it works once, because a respected author wrote it, or because a model generated it confidently. That includes the router's own components, its learned artifacts, and the router as a whole.

The architecture divides validation into two separate planes:

  • Mechanical verification asks whether the machine obeys its contracts.
  • Comparative acceptance asks whether the candidate is actually better than what already runs, at what cost, and with what new risks.

Passing the first plane is a precondition for the second, never a substitute for it.

13.2 Characterize before replacing

The first test artifact is not a Router 2 test at all. It is a behavioral baseline of Router 1:

  • its existing regression suite, run and recorded;
  • a corpus of known-good routing cases;
  • known failures, explicitly labeled as failures so they are never mistaken for targets;
  • historical operator corrections;
  • decisions across every level of the precedence ladder;
  • provider fall-through traces;
  • measured latency distributions.

The baseline exists for change detection, not correctness. Every behavioral difference Router 2 produces must be classified as either an intentional improvement, traceable to a specific design decision and pre-registered as such, or an unexplained regression, which stops progress until explained. "It changed and it's probably fine" is not an allowed classification.

13.3 A test hierarchy with a deliberate cost gradient

Tests are organized in layers ordered by cost:

  1. pure component tests;
  2. contract tests for every interchangeable component;
  3. property and invariant tests;
  4. pipeline integration tests;
  5. historical regression;
  6. adversarial and fault-injection tests, including compound failures;
  7. replay against historical decisions;
  8. shadow-isolation tests;
  9. load and soak tests;
  10. release qualification.

A component cannot enter integration without passing its component and contract layers. Qualification tiers re-execute everything they depend on; they never cite earlier results.

Some representative mechanisms:

Contract testing. Every replaceable component, whether an alternative lexical scorer, confidence policy, reasoning implementation, or provider selector, is tested by binding it into its interface slot and running the slot's contract suite unchanged. The harness does not know which implementation it is testing. This is what makes "swap this component" a safe operation.

Invariant testing. Architectural invariants are exercised with generated inputs. Examples:

  • narrowing never emits a candidate not in its input;
  • confidence never leaves its valid domain;
  • an unauthorized route can never become dispatchable regardless of confidence, memory, or model output;
  • fall-through always terminates;
  • concurrent requests never share ephemeral state.

Failing seeds are shrunk into minimal fixtures and preserved.

Fault injection. Every failure the architecture claims to handle must be deliberately induced in test, through injection points at every dependency the spine calls. Each architectural invariant must have at least one test that actively tries to violate it. A curated set of compound failures covers plausible combinations, for example a primary provider down while memory is slow, or a restart in the middle of dispatch.

Targeted mutation. Deliberate small corruptions are applied to decision-critical code, such as threshold comparisons, precedence order, and governance checks. A surviving mutant in a governance check is a test defect that blocks qualification.

13.4 False-green prevention

Treebeard has direct experience with a suite that reported success while not actually testing anything. The architecture treats that as a design constraint.

A suite result is evidence only if it reports the following:

  • expected test count, discovered count, executed count, and skipped count;
  • failures and known failures;
  • fixture, component, and configuration versions;
  • durations.

Several rules are mechanical:

  • Zero discovered tests is a failure.
  • A required test that was skipped blocks qualification.
  • An exception inside the test harness is recorded as a harness error, never as a pass.
  • A test marked as a known failure that unexpectedly passes is itself a failure, because the expectation is stale.
  • Human-readable summaries are generated from structured results and cannot exist without them.
  • Production invariant checks are ordinary explicit code, not language assertions that an optimization flag could silently remove.

Every meaningful defect closes through the same lifecycle: reproduce, minimize, write a failing test, verify it fails, fix, verify it passes, run the related suites, and commit the regression. A code change plus a manual observation is not closure.

13.5 Comparative acceptance

Candidates are scored on a multidimensional scorecard, never a single headline number. The dimensions include:

  • top-1 accuracy and top-2 and top-3 recall;
  • final routing correctness per severity class;
  • confidence calibration;
  • escalation precision and recall;
  • reasoning rescues versus reasoning regressions;
  • tie-resolution outcomes;
  • execution-shaping and provider-selection accuracy;
  • fall-through success;
  • latency percentiles;
  • cost;
  • operator correction rate.

Hard gates are non-compensatory. A governance violation, a shadow component that dispatched, a duplicate-side-effect risk, a failed critical invariant, or a rollback failure cannot be offset by being faster or cheaper. Among candidates that pass every hard gate, quality, latency, and cost are compared as a frontier: a candidate dominated on every dimension has no case for promotion. Choosing among non-dominated candidates is an explicit product decision, recorded with the evidence, rather than a formula buried in the harness.

Confidence calibration is measured directly, as accuracy and correction rate by confidence grade. The aim is confidence that means what it says, not confidence that is as high as possible.

Ablation makes every component justify its complexity: baseline, baseline plus the component, full system without the component, and alternatives. Components whose contribution is indistinguishable from noise are flagged for removal. The removal decision is still a product decision, not an automatic one.

13.6 Protecting the benchmark from the system being benchmarked

A self-improving router will eventually over fit to its benchmark unless the benchmark is structurally protected. The evaluation corpus is partitioned into development, regression, challenge, operator-derived, sealed holdout, and production-shadow sets.

  • The sealed holdout is never readable by the candidate or its development process.
  • Holdout results are reported only in aggregate.
  • Repeated exposure to the holdout is tracked and reported.
  • Candidates cannot write to their own labels, corpus, or gate configuration.
  • An improvement that appears only on development data is reported as fitted, not generalized.

Shadow evidence is handled with similar discipline. A shadow router saying it would have chosen differently is not evidence it would have done better. Every counterfactual is classified as directly applied scores, replay-executable, partially observable, or unobservable. Only the first two may contribute to gate metrics.

13.7 Evidence packets

Every acceptance run produces an evidence packet containing:

  • candidate and baseline identities and configurations;
  • corpus hashes;
  • the gate table;
  • metrics by workload and severity;
  • regressions and improvements with case-level pointers;
  • uncertainty intervals;
  • contamination and exposure reporting;
  • everything needed to reproduce the run.

The packet has no top-level pass field. Its top level is the gate table, and any failed hard gate fails the run regardless of anything else in it. Promotion claims can be made only against a packet.

13.8 Earned autonomy for generated components

Router 2 is designed so that Treebeard's own agents can help build it. That creates a second use for the validation machinery. Components may be produced by Treebeard, by the integration engineer, by implementation agents, or by humans. All face identical gates. Origin is recorded but never consulted by any gate.

Each generated component moves through a fixed ladder:

design contract → contract validation → component test → fault test → replay → shadow → limited authority → production authority

Each rung is gated by recorded evidence. Demotion is cheap. No component can promote itself, and presence in the codebase confers no authority.

Recording the outcome for every component builds an empirical picture of what the development system can and cannot yet do reliably. Outcomes are recorded as accepted, repaired, replaced, or rejected, along with the class of defect found. That picture is analyzed separately and never feeds back into the artifact being measured.


14. Migration from Router 1

14.1 Philosophy

Router 1 works. The migration rejects both failure modes available to it.

  • The big-bang rewrite is unprovable before cut over and cannot be rolled back after.
  • Perpetual layering puts new concepts on top of old code, and the legacy never dies.

Router 2 enters by controlled replacement of an existing working system. Every responsibility eventually has one owner, and every obsolete legacy path is eventually deleted.

14.2 Characterize, then wrap

Phase 0: characterization. This is an inventory of the router Treebeard actually has, produced by reading code rather than projecting the architecture onto it. It covers:

  • entry points and callers;
  • where control commands are detected;
  • what the deterministic selector actually returns and what it costs;
  • how thread state, mission metadata, governance states, the tie-break mechanism, provider fall-through, and receipts actually work;
  • current latency;
  • existing mis-routes and corrections.

Every item gets one disposition: keep, extract, adapt, replace, retire, or requires investigation. Nothing on the critical path migrates while it is still unknown.

Phase 1: wrapping. Router 1 is then wrapped in canonical request and result envelopes with zero semantic change. Facade transparency is proven by showing that routing behavior is unchanged with the wrapper in place. The facade translates; it does not understand. Directive classification, scoring, confidence, shaping, provider logic, and governance are all forbidden inside it. If the facade needs to understand the turn, the seam is in the wrong place.

Missing provenance stays missing. When Router 1's output is translated into the canonical result envelope, fields Router 1 never computed are marked unknown, not synthesized. Fabricated provenance would corrupt every subsequent comparison.

14.3 Build order

The migration order is driven by isolation, determinism, and diagnostic leverage rather than by the architecture's conceptual order. Authority-sensitive responsibilities come last, after the tooling to prove them exists.

  1. Characterization.
  2. Compatibility envelopes around Router 1.
  3. Telemetry spine: canonical decision records produced from Router 1, which bootstraps replay and the evaluation corpus. Migration that cannot be measured is migration on faith, so measurement comes before any decision logic moves.
  4. Explicit correction capture.
  5. The Router 2 working-state skeleton, running in shadow with no dispatch authority.
  6. Intuition wrappers exposing the existing deterministic behavior as canonical evidence.
  7. Mission selection and confidence as pure shadow components.
  8. The atomic governance/commit boundary, wrapping current authority behavior.
  9. Execution shaping extraction.
  10. Provider selection and fall-through extraction.
  11. Limited authority for selected low-risk, deterministic routing classes.
  12. The cold path: historical retrieval, the tie-break flow, cognitive budgeting, and reasoning.
  13. Learning, as candidate artifacts evaluated offline before any activation.
  14. Progressive promotion to production default.
  15. Router 1 retirement.

Reordering requires a concrete dependency, not preference.

14.4 Authority promotion

Deployment and authority are separate states. A component can be installed, test-only, replay-only, shadow, advisory, limited-authority, production-authority, or disabled. Code present does not mean code authoritative. Authority is explicit, versioned state held by a release and activation controller, and it is recorded in every decision's provenance.



Router 2 earns authority in stages: characterize, observe, shadow, grant limited authority, expand it with evidence, and retire Router 1 only after the new system proves itself.

Each migrated responsibility follows the same unit lifecycle:

inventory → contract → characterization tests → new component → component and contract tests → replay → shadow → disagreement review → limited authority → full authority → legacy disabled → legacy deleted

"New implementation exists" is step four of thirteen, not the finish line.

Advancement is evidence-gated at every stage:

  • Component → shadow: component and contract suites are green; deterministic behavior is verified; no governance-class failures remain.
  • Shadow → limited authority: minimum shadow evidence exists across the relevant workload classes; every bad-outcome disagreement has been classified; governance divergences are zero; isolation has been audited.
  • Limited authority → authoritative default: the sealed-holdout acceptance run passes every hard gate; the evidence packet is complete and reproducible; Product Owner approval is recorded against that evidence.
  • Authoritative default → Router 1 retirement: sustained operating evidence is at least as strong as the Router 1 baseline; fall-through and governance have been proven with Router 1 unavailable; rollback has been proven by test.

14.5 Disagreement is classified, not resolved by seniority

When the two routers disagree, the record captures both candidate sets, both winners, confidence, escalation, provider choices, cost, latency, and the eventual outcome where it is observable. Each disagreement is classified as one of:

  • Router 2 improvement;
  • Router 1 correct;
  • both acceptable;
  • both wrong;
  • unscoreable;
  • requires a product decision.

Router 1 is a baseline, not ground truth. Router 2 is never automatically tuned toward Router 1's answers. The label "Router 1 was correct" must be earned by outcome evidence, such as an operator correction, not assigned by default. Mission-selection disagreements are additionally screened for authorization regression, and a single such case stops migration regardless of aggregate metrics.

14.6 Rollback, receipts, and stop conditions

Rollback restores a coherent combination of code, component manifest, configuration, learned artifacts, authority assignment, and compatible persistent state. It never means restoring old code on top of new state. A rollback checkpoint that has not been exercised in test is not a valid checkpoint.

Receipts. Every authority transition emits a migration receipt recording:

  • the previous and new owner;
  • versions and configuration;
  • test, benchmark, and shadow evidence;
  • the rollback point;
  • known issues accepted at activation.

Stop conditions. Migration halts on any of:

  • governance regression;
  • unexplained severe disagreement;
  • state corruption;
  • a shadow side effect;
  • a fallback loop;
  • duplicate-dispatch risk;
  • rollback failure;
  • cross-mission leakage;
  • inability to reproduce the effective state behind a decision.

Resuming requires a recorded root-cause analysis, not a passing retry. Sunk cost is explicitly not a reason to proceed.

14.7 Retirement must remove complexity

Temporary machinery is tracked in a single debt register, including adapters, dual-run flags, compatibility schemas, duplicate telemetry, and fallback scaffolding. Every entry carries an owner and a concrete deletion condition. An adapter without a deletion condition fails review.

Router 1 retires only after sustained standalone Router 2 operation. Retirement then:

  • removes the runtime dependency;
  • deletes adapters;
  • archives a reproducible reference implementation;
  • keeps useful regression fixtures;
  • simplifies configuration and telemetry;
  • reruns full qualification.

Retirement is judged by measured reduction in complexity. The finished system should read as Router 2, not as Router 1 buried under replacement machinery.


15. Generalization to Multi-Product AI Platforms

The architecture was designed for one self-hosted multi-agent system, but its central problem is common. Many companies now own several AI-enabled capabilities, some built and some acquired, that customers increasingly expect to behave like a single platform. A representative platform might contain:

  • a market or data intelligence product;
  • a workflow and automation product;
  • an agentic execution layer;
  • specialized tools;
  • several model providers;
  • legacy customer workflows that cannot be broken.

What a routing layer does not solve

It is important to be precise here. A routing and orchestration layer does not solve product integration by itself. It does not migrate data, reconcile schemas, consolidate customer identity, unify permission models, or merge user experiences. Those are substantial programs in their own right. An architecture that claims otherwise is selling routing as a substitute for integration work it cannot do. If two acquired products model the same customer differently, no router fixes that.

What it does provide

What a governed routing layer can provide is a coherent decision layer across capabilities that remain separately built. It gives each request one accountable answer to these questions:

  • Which capability should handle this request? Evidence ordered by authority, with explicit user direction first, deterministic signals second, and inference last.
  • What context follows it? Only the owning capability's context, plus explicitly referenced context the governance layer permits to cross.
  • What authority accompanies it? Established by a governance layer that no inference path can override.
  • What execution capability is required? Expressed as a capability requirement, independent of which vendor or model realizes it.
  • When should multiple systems collaborate? One primary owner, with other capabilities referenced explicitly rather than sharing ownership ambiguously.
  • Where should human review sit? At consequential boundaries identified by a declared consequence classification, not everywhere.
  • What happens when a preferred capability fails? Typed, bounded fall-through with no silent capability downgrade.
  • How is the decision observed and audited? An immutable decision record with the exact effective configuration, joined later to outcomes and corrections.

Why the migration pattern matters as much as the routing pattern

For a multi-product platform, the migration architecture may be the more transferable half. Wrapping each existing product behind canonical envelopes without changing behavior, instrumenting before modifying, running new orchestration in shadow, and transferring authority one responsibility and one workload class at a time is how a platform can converge without a flag-day cut over that puts existing customer workflows at risk. The same evidence packets that justify promoting a Router 2 component justify routing a customer segment's requests through a new unified path.

The same caution applies at the model layer. Separating product intent from provider realization lets a platform change model vendors, add local or private models for data-sensitive customers, or respond to a provider outage without redesigning products that depend on them.


16. Architectural Tradeoffs and Lessons Learned

The design record includes independent section studies, three unified drafts, an adversarial review, and a freeze with targeted corrections. The most instructive lessons are the places where an earlier formulation was wrong in a way that looked reasonable at the time.

16.1 "Commit, then check authority" creates an unsafe intermediate state

An intermediate draft committed the routing decision immutably and then asked governance whether it was permitted. That looks like clean separation of concerns. In practice it creates a window in which a final-looking decision exists that may not be allowed, and it relies on every downstream component to remember that it might be invalid.

Making validation and commit a single atomic boundary removed the window. The lesson: separation of concerns does not mean separation in time. Some checks belong inside the transition, not after it.

16.2 Unowned glue is where defects hide

Deciding whether committed work is handled directly, by a specialist, or by a tool initially lived in no one's section. It was implied as plumbing between the router and the executor. The adversarial review found that this ownerless zone held consequential decisions: capability requirements, side-effect classification, and what happens under degraded capability.

Elevating it to a first-class component with one owner resolved three separate ambiguities. When a design has a gap between two well-specified components, the gap usually contains a decision, and that decision needs an owner.

16.3 "Genuine tie" had to become mechanical

Letting the confidence controller judge semantically whether candidates were truly tied made the tie-break an escape hatch for confusion. Replacing the judgment with a conservative mechanical predicate, whose default when anything is missing is "not a tie," separated symmetric choices from semantic conflict.

Ambiguity that matters goes to reasoning. Only ambiguity that does not matter goes to the tie-break.

16.4 Never let a model decide what options exist

An earlier reasoning design could "restore" candidates from a deferred pool. It was narrow and well-intentioned, and it was still a model deciding which options were on the table. The final design moved candidate supply entirely to mission selection.

The general principle is that models may rank within a whitelist defined before they run; they may not write the whitelist.

16.5 One interpreter of operator intent

Early drafts recognized operator direction in several places, which produced subtle conflicts about whether a mention of a mission was an instruction. Consolidating into one normalizer that explicitly distinguishes assignment, standing assignment, reference, exclusion, correction, and incidental mention made operator intent both unambiguous and testable.

In a system where the operator's word outranks inference, the parser of that word is the most security-relevant component in the router.

16.6 Fallback is only safe before side effects

Falling back to the legacy router on failure seems obviously safe until the failure happens after an external action was attempted. At that point, a second router starting fresh can repeat the action. Fixing the fallback boundary at the first dispatch attempt, and adding idempotency-aware dispatch, turned an intuitive safety net into a correct one.

16.7 Two evidence languages will drift

Allowing the memory subsystem its own evidence representation seemed reasonable, since historical evidence carries extra attributes. It would have required every consumer to reason about two schemas that inevitably diverge. One canonical evidence contract, with source-specific detail in metadata, was simpler and made uniform replay possible.

16.8 Routing thought and worker thought are different budgets

The last correction applied at the freeze moved the budget for specialist adjudication from the routing cognition controller to Execution Shaping. It was a small change that protected a large principle. If routing and execution share a budget, execution questions can reopen routing decisions, and cognition costs become impossible to attribute.

16.9 Inspectable beats clever, until evidence says otherwise

Several choices deliberately favor explainability over theoretical optimality:

  • lexicographic provider ranking instead of weighted scoring;
  • precedence tiers instead of a unified score;
  • declarative policy tables instead of distributed conditionals.

Each is paired with an explicit condition for revisiting it: if evaluation shows the more complex approach wins. The architecture does not reject sophistication. It requires sophistication to prove itself against a simpler baseline.

16.10 Design documents must decide

Several early drafts exposed a recurring failure mode: restating design questions without committing to an owner, contract, or trade-off. A design document earns its place by removing ambiguity for the implementer. Every final section states a recommendation, the alternatives considered, and the trade-off accepted. Every fact that could not be known without the code is isolated as a characterization question rather than invented.

Accepted costs. The architecture knowingly accepts several costs:

  • Centralized failure handling concentrates risk in the orchestration spine.
  • Fail-closed governance costs availability during authorization outages.
  • Validation adds work to the hot path.
  • Conservative learning adapts more slowly.
  • Keeping Router 1 as fallback doubles some maintenance during migration.

Each cost is bounded and documented, and each was chosen over a failure mode judged worse.


17. Current Status and Next Steps

What exists

Router 1 is Treebeard's working router in the existing Treebeard environment. It already embodies the routing concepts Router 2 preserves: operator-directed routing, explicit mission references, deterministic lexical selection, thread continuity, memory-assisted tie resolution, model-assisted fallback, genuine no-match, control-turn handling, cost-aware model tiers, provider fall-through, and routing telemetry.

The Router 2 architecture exists as a frozen design:

  • seventeen owned architecture sections;
  • a single ownership matrix;
  • canonical contracts and stage order;
  • a failure-ownership table;
  • roughly forty mechanically testable global invariants;
  • an end-to-end migration and retirement plan;
  • a fifteen-step dependency-ordered build sequence.

The freeze criteria state that no unresolved ownership conflicts remain. Further architectural change requires an identified defect, such as a contradiction, an unimplementable contract, an ambiguous boundary, or a safety defect, rather than preference.

What does not yet exist

Router 2 has not been implemented:

  • No Router 2 component is running in shadow or production.
  • No Router 2 benchmark results exist, and no latency, accuracy, or cost improvement is claimed.
  • No numeric thresholds, confidence bands, latency budgets, or promotion criteria have been set. The architecture deliberately defers every one of them until Router 1 has been measured.

What must happen next

Phase 0 characterization comes first. It answers 36 questions across ten areas by inspecting the existing implementation:

  • the router's entry signature;
  • what the deterministic selector actually returns;
  • where control detection lives;
  • whether the existing tie-break can be bounded to an explicit candidate set;
  • which governance states exist and whether "grantable" is distinguishable from "denied";
  • which provider failure modes are currently detectable;
  • current fast-path latency percentiles;
  • existing correction rates.

The architecture treats these answers as binding parameters. They configure the design; they do not reopen it.

After characterization, the build order begins with changes that alter no behavior: compatibility envelopes, a telemetry spine that produces canonical decision records from Router 1, and explicit correction capture. These establish the replay corpus and measurement surface that every later promotion decision depends on. Router 2 decision components then enter in shadow, earn limited authority on low-risk deterministic routing classes, and expand only as evidence accumulates.


Appendix A — Glossary

Ablation. An evaluation technique that removes or disables one component at a time to determine whether that component actually improves the system.

Atomic operation. A change that either completes as one indivisible action or does not occur at all. Router 2 uses an atomic governance-and-commit boundary so that a routing decision cannot become final before its authority is established.

Bounded cognition. Reasoning that operates within explicit limits on time, cost, context, passes, or depth rather than continuing indefinitely.

Calibration. Measurement of whether a system’s stated confidence corresponds to how often its decisions are actually correct.

Candidate. A mission, capability, specialist, model, or other option being considered for selection.

Cognitive middleware. A decision layer that determines ownership, authority, execution shape, and cognitive expenditure before and around model execution, rather than merely selecting a model.

Consequence class. A declarative risk classification—such as low, normal, or high—derived from action type, side effects, reversibility, and the impact of plausible routing choices.

Content-addressed. Stored or referenced according to a fingerprint derived from the content itself, allowing an exact version of configuration or state to be identified later.

Control turn. An operator command directed at the system itself rather than at mission work. Control turns bypass mission inference but still cross consequence classification and the governance/commit boundary.

Deterministic. Producing the same output whenever the same inputs and configuration are supplied. Deterministic Router 2 stages can therefore be reproduced exactly.

Directed target. A mission explicitly named or assigned by the operator. It is preserved through governance evaluation even when authorization is required.

Dreams. Treebeard's existing memory-assisted tie-break. In Router 2 it is limited to mechanically established genuine ties, runs at most once per cycle, and can only choose within the tied set or abstain.

Evidence. A canonical, append-only record of a routing claim carrying its source, class, producer, polarity, freshness, and configuration epoch.

Evidence packet. The structured output of an acceptance run containing the evidence needed to support or reject promotion of a component or configuration.

Execution shape. The complete execution contract for committed work: handling mode, target, worker capability, side-effect and idempotency requirements, data class, degradation policy, and receipt mode.

Fall-through. Trying the next eligible provider or capability after a prior attempt fails. Router 2 uses a finite, predetermined fall-through chain rather than an open-ended retry loop.

Genuine tie. A candidate set satisfying a conservative mechanical symmetry test. It is distinct from unresolved conflict, which proceeds to bounded reasoning.

Governance. The subsystem that determines whether a proposed action is permitted. Routing determines where work belongs; governance determines whether it may act.

Holdout set. Evaluation data deliberately withheld from development so that a candidate can be tested against cases it has not been tuned against.

Idempotency. The property that repeating the same operation does not create an additional external effect after the first successful application.

Idempotency key. A unique identifier attached to an operation so that retries can be recognized as the same action rather than executed as duplicates.

Immutable. Not modifiable after creation. Router 2 keeps final routing decisions and raw routing history immutable.

Invariant. A rule that must remain true regardless of input, configuration, failure, or execution path. Router 2 expresses important architectural boundaries as mechanically testable invariants.

Latency. The elapsed time required for an operation or stage to complete.

Lexical evidence. Evidence derived from words, phrases, structural patterns, or other directly detectable features of the request rather than from model reasoning.

Lexicographic ranking. Ranking by a fixed sequence of priorities: compare the first criterion, then use the second only when needed, then the third, and so on. This differs from combining all criteria into one weighted score.

Mission. A bounded, long-running objective in Treebeard with its own scope, context, specialists, and permissions. Routing primarily determines which mission owns a turn.

Mutation testing. Deliberately introducing small defects into important code to verify that the test suite detects them.

No-match. A legitimate routing result indicating that no available mission or capability appropriately owns the request. It is not the same as a router failure or governance denial.

Provenance. The recorded origin and history of a decision or piece of evidence: what produced it, from which inputs, and under which configuration.

Provider. The service or runtime that supplies a model or other execution capability. Router 2 separates the required capability from the provider that ultimately delivers it.

Replay. Re-running a decision from recorded inputs and state to determine what the system would produce. Deterministic stages can be replayed exactly.

Rollback. Returning from a newly activated component, configuration, or learned artifact to a previously known-good state.

Semantic reasoning. Interpretation based on meaning rather than simple word or pattern matching. Router 2 invokes it only when cheaper evidence does not adequately resolve a decision.

Shadow operation. Running a candidate system on production-equivalent inputs without allowing it to dispatch work, mutate production state, or influence the authoritative result.

Side effect. A change outside the routing calculation itself, such as writing data, sending a message, invoking a tool, submitting a transaction, or otherwise changing external state.

Telemetry. Structured operational evidence about what the system did, why it did it, how long it took, what configuration was active, and what happened afterward.

Tie-break. A bounded mechanism used to select among candidates that have already been established as genuinely equivalent under the routing rules.

Tri-Cortex. The three evidence-producing responsibility domains: Intuition (deterministic and fast), Reasoning (bounded and semantic), and Subconscious/Memory (historical). They are conceptual owners, not separate processes or model calls.

Working record. The structured per-turn state owned by Router Core. Components return typed results to Core rather than modifying the record directly.


Appendix B — Selected Invariants

A representative subset of the architecture's global invariants, each intended to be mechanically testable or objectively reviewable:

  1. Every consequential decision has exactly one owner. No implementation may create a second owner implicitly.
  2. Governance is the sole writer of authority. No other component's output contract contains a writable authority field.
  3. Governance validation and routing commit form one atomic boundary. Exactly one final disposition is committed per routing attempt.
  4. A directed operator target is never silently replaced. Absent authority yields "authorization required" on that target.
  5. Memory, the tie-break, reasoning, learning, and provider selection cannot create, broaden, or substitute authority.
  6. Governance uncertainty or unavailability never degrades into permissive execution.
  7. Short circuits skip optional cognition only, never mandatory boundaries.
  8. Escalation is monotonic and bounded, and is forbidden when no new evidence has appended since the previous cognitive pass.
  9. Models rank only within a candidate set fixed before they run.
  10. Required delivery capability is never silently downgraded. Provider fallthrough is finite, loop-free, and failure-typed.
  11. Legacy-router fallback is legal only before the first external dispatch attempt and cannot re-enter Router 2.
  12. Evidence is append-only. Raw routing history is immutable. Learned state is versioned, reversible, and separate from raw history.
  13. Every routing attempt runs under one coherent, immutable effective state, and every decision records it.
  14. Deterministic stages are exactly replayable from a sealed decision record. Approximate replay is always labeled as such.
  15. Shadow execution structurally cannot dispatch. Experimental components cannot self-promote. Deployment does not imply authority.
  16. Hard governance and safety gates are non-compensatory. Zero discovered tests is not a successful suite.
  17. Components from any origin face identical qualification gates.
  18. Migration adapters contain no routing intelligence, and all temporary migration infrastructure carries an explicit retirement condition.

No comments:

Post a Comment