Tuesday, August 25, 2026

Gungabricks - A finite attention layer for GitHub pull-request triage: suppress routine work, surface uncertainty, preserve human authority.

Gungabricks — Product Requirements Document

A finite attention layer for GitHub pull-request triage: suppress routine work, surface uncertainty, preserve human authority.

Author: Rhombus Ticks

Technical Consultation: TC Ricks, IT Consultant

Research Assistance: Redwin Tursor

Document Product Requirements Document (PRD)
Version v0.7 — Draft
Status Partner circulation draft
Primary Author Rhombus Ticks
Technical Consultation TC Ricks, IT Consultant
Research Assistance Redwin Tursor
Imprint Red Anvil Productions
Designation Project 20 Gem
Scope MVP, controlled pilot, Phase 2 / Phase 3 roadmap
Date August 25, 2026
PRODUCT THESIS
Gungabricks reduces the cost of senior-engineering attention. It does not replace GitHub, and it does not automate judgment away. It removes routine triage only when policy and evidence justify silence; everything uncertain goes to a human.

Revision Notes

v0.7 — this revision:

  • Prepared a partner-circulation draft by removing working-draft meta-references and normalizing the document’s external-facing register.
  • Tightened normative requirement language, including evaluator regression criteria, data-handling fallback, Inspection lifecycle, GitHub authority, notification defaults, and retention exceptions.
  • Resolved the EASY WIN suppression contradiction and clarified RETURN, stale-action, concurrency, and GitHub review wording.
  • Preserved the product thesis, suppression rule, authority-boundary model, kill criteria, pilot runbook, and standing weekly operating agenda.

Document Conventions

This PRD uses MUST, MUST NOT, SHOULD, SHOULD NOT, and MAY in the ordinary requirements sense. MUST and MUST NOT are binding for the stated phase. SHOULD indicates a strong design default that may be overridden only with a documented reason. MAY is optional. Business prices, staffing, and dates are hypotheses until validated by an actual build and pilot partner.

1. Executive Summary

Gungabricks is an attention-management product for software delivery. It sits beside GitHub rather than replacing it as a system of record. It receives pull-request activity, builds canonical state, applies deterministic policy, uses AI-assisted evaluation to synthesize evidence, and decides one thing: does this pull request require a maintainer’s attention now?

The problem is not visibility. Engineering teams are drowning in visibility. The problem is that every notification asks a senior engineer to spend a little judgment. Those fragments are expensive. Gungabricks is valuable only if it returns that judgment capacity without buying efficiency by hiding risk.

The MVP therefore makes a narrow bargain. ACT NOW, DECIDE, and EASY WIN remain human work. IGNORE is the only classification eligible for suppression, and even IGNORE stays human-visible whenever confidence is not high, a deterministic rule vetoes suppression, evidence is incomplete, or an anomaly is present. The model cannot approve, merge, hold credentials, or grant itself authority. Uncertainty fails toward a human.

Human actions are equally constrained. APPROVE and RETURN are explicit maintainer decisions, bound to the exact evaluated commit SHA, held behind a server-enforced undo window, executed using on-behalf-of GitHub identity, and revalidated before execution and every retry. A stale SHA is refusal, not “best effort.”

The commercial thesis is equally testable: if Gungabricks cannot create measurable reduction in active triage time, it is another inbox and does not warrant investment. The controlled pilot therefore measures suppression safety, workflow adoption, realized time returned, operational reliability, and willingness to pay. One pilot can prove the product loop. It cannot prove the market. A second independent design partner is required before broad commercialization.

THE RULE
A classifier that cannot be trusted to stay silent is not an attention product. It is an interruption engine with extra steps.

2. Problem Statement

2.1 Attention, not visibility

GitHub, CI systems, issue trackers, monitoring tools, and chat already expose pull requests, checks, approvals, failures, comments, and status changes. Nothing important is hidden for lack of a dashboard. The unresolved cost is the repeated human decision about whether any of it matters.

Senior-engineering attention arrives in small pieces: open the notification, recover context, check the SHA, scan CI, remember the ownership rule, read the review state, decide whether the change is routine, then return to whatever work was interrupted. The individual interruption looks cheap. The aggregate is not.

Gungabricks attacks that cost directly. It does not promise faster code review. It promises fewer unnecessary reviews of whether a review is necessary.

2.2 The product opportunity

The desired state is a short, credible queue. Routine activity disappears from active workload only when hard policy and current evidence permit it. Ambiguous, risky, stale, contradictory, or partially understood work is escalated. The user should be able to open the product, see what needs them, decide, and leave.

2.3 Existing products solve adjacent problems

Platform / product What it already does Implication for Gungabricks
GitHub Native PR inbox, rulesets, merge queues, and Copilot-assisted review inside the system of record. Highest platform risk. Win on selective attention suppression, evidence, reversibility, and independent audit — not by cloning GitHub review UI.
Graphite PR inbox, full review surface, AI reviews, merge queue, stacking, notifications, developer workflow. Stay narrower. Decide what deserves human attention rather than becoming another PR client.
CodeRabbit AI code review focused on understanding and validating changes. Treat review intelligence as adjacent. Differentiate on routing, trust policy, suppression safety, and human decision governance.
Mergify Rules-based PR workflow automation, merge protections, merge queues, conditional actions. Differentiate by allocating human attention before execution, combining deterministic rules with explainable evaluation.

These competitive references were checked on August 25, 2026 and should be re-validated before launch; they are context, not diligence.

THE WEDGE
Gungabricks governs what does not need to become a human task. The moat, if it earns one, is calibrated policy plus audited evidence plus workflow trust — not access to a generic model.

3. Goals and Non-Goals

3.1 Goals

  • Reduce active senior-engineering triage time by keeping policy-eligible routine PR activity out of the human attention queue.
  • Make every surfaced or suppressed decision reconstructible from immutable PR version, rules, model/prompt version, evidence, routing policy, and human action.
  • Keep repository authority human. AI may classify and explain; it may not approve, return, merge, hold credentials, or issue external mutations.
  • Bind every external review action to the exact SHA the maintainer examined and refuse stale execution.
  • Provide a finite user experience: Today, a secondary Board, a Decision Panel, and Done — not an engagement dashboard.
  • Prove safety and business value through a controlled baseline → shadow → live pilot, with independent review of suppressed predictions.
  • Keep the architecture evolvable toward multi-tenant and multi-SCM operation without building that complexity before the first value proof.

3.2 Non-Goals

  • Not a replacement GitHub client or generalized code-review IDE.
  • Not an autonomous merge or code-modification agent.
  • Not a learned personal productivity score, gamification system, or engagement product.
  • Not a general workflow/ticketing engine.
  • Not multi-SCM in the MVP; GitLab and Bitbucket are Phase 3 candidates.
  • Not shared multi-tenant SaaS in the first pilot; the MVP is dedicated single-tenant.
  • Not broad enterprise identity or custom organization-wide policy authoring in MVP.
  • Not a claim that one design partner proves market generalization.

4. Product Overview

The operating loop is deliberately linear. Every stage has one authority boundary and one failure direction. The system may become more sophisticated later; the authority graph should not.

Stage Input Output / authority rule
1. Intake Signed GitHub webhook + reconciliation Persist delivery before async processing; dedupe by delivery ID.
2. Canonical state Repository + PR activity Immutable PR version bound to exact commit SHA.
3. Deterministic policy Ruleset + PR version Hard vetoes and required HUMAN_ATTENTION conditions.
4. AI-assisted evaluation Bounded, redacted context Bucket, confidence, explanation, evidence, optional alternative/anomaly suggestion. No authority.
5. Routing Rules + validated evaluation + anomalies HUMAN_ATTENTION or NO_ATTENTION. Only policy-eligible high-confidence IGNORE may be suppressed.
6. Human surface Today / Board / Decision Panel Maintainer sees evidence and chooses APPROVE, RETURN, INSPECT, or DEFER.
7. Pending action Human decision + exact evaluation/SHA Server-side undo window; authorization and current state remain authoritative.
8. Controlled executor Valid decision + current auth/SHA Submit GitHub review on behalf of maintainer; retry idempotently or fail visibly.
9. Audit All prior state transitions Append-only reconstruction of why the item was hidden, surfaced, cancelled, retried, or acted on.

The model appears in one stage. Credentials appear in another. The design is safer because those two facts never meet.

5. Component Specifications

5.1 GitHub Intake and Canonical PR State

Purpose. Receive PR activity without treating webhook order as truth, and bind every later decision to a stable version of the repository state.

Behavior. The GitHub App persists signed webhook deliveries before asynchronous processing, deduplicates by GitHub delivery identifier, and runs reconciliation to repair missed or out-of-order events. Each evaluation points to an immutable PR version and exact commit SHA.

REQ-INTAKE-01 — Valid signed webhooks MUST be durably persisted before downstream processing is scheduled.

REQ-INTAKE-02 — Duplicate delivery identifiers MUST NOT create duplicate transitions, evaluations, or external actions.

REQ-INTAKE-03 — Periodic reconciliation MUST compare local state with GitHub and repair missed/out-of-order activity.

REQ-STATE-01 — Every evaluation and human decision MUST reference an immutable PR version and exact commit SHA.

REQ-STATE-02 — The SHA MUST be revalidated before external execution and before every retry. A mismatch MUST refuse the action and trigger re-evaluation.

5.2 Deterministic Policy and AI-Assisted Evaluation

Purpose. Separate hard repository policy from probabilistic classification so the model can contribute judgment without becoming authority.

Behavior. Deterministic rules run first. The evaluator then returns a schema-valid bucket, qualitative confidence, short explanation, structured evidence, and optional alternative bucket or anomaly suggestion. Routing consumes both. A human-required rule vetoes model suppression.

REQ-EVAL-01 — Rules that require human attention MUST have veto power over model output.

REQ-EVAL-02 — Malformed, unavailable, timed-out, or contract-exceeding model output MUST route to HUMAN_ATTENTION; the system MUST NOT heuristically repair it into a suppressible result.

REQ-EVAL-03 — Production evaluation MUST pin model identifier, prompt version, context-extraction version, output schema, and ruleset version.

REQ-EVAL-04 — Repository content MUST be treated as untrusted data. Instruction-like content in repository data MUST NOT alter authority, credentials, or routing policy.

REQ-EVAL-05 — A production evaluator change MUST NOT ship if it regresses on the maintained golden and adversarial evaluation corpora against the currently deployed evaluator.

5.3 Starter Rule Catalog, Evidence, and Anomalies

Purpose. Start the pilot with conservative policy rather than an empty rules engine, and make evidence traceable to sources rather than to model prose.

Starter rule Default effect
Required CI/check failure Force HUMAN_ATTENTION regardless of confidence; not eligible for IGNORE suppression.
Security, authentication, infrastructure, deployment, database migration, or policy-controlled paths changed Force HUMAN_ATTENTION.
First-time/external contributor under configured policy Force HUMAN_ATTENTION.
Required CODEOWNERS review not satisfied Force HUMAN_ATTENTION.
>1,500 changed lines or >50 files, or binary/generated content prevents full evaluation Record PARTIAL_CONTEXT and force HUMAN_ATTENTION.
Model/rule material contradiction Record CONTRADICTION and force HUMAN_ATTENTION.
Evaluation unavailable or schema-invalid Record failure and force HUMAN_ATTENTION.

Behavior. MVP anomalies are STALE, CONTRADICTION, and PARTIAL_CONTEXT. STALE defaults to more than seven calendar days without material activity; new commits, human reviews/comments, or CI/check results count as material. Evidence objects carry a claim, a strengthens/weakens direction, and references to source objects.

REQ-EVID-01 — Every evidence item MUST contain a claim, direction, and one or more source references.

REQ-EVID-02 — The UI MUST allow a user to move from classification → claim → underlying GitHub source without treating the generated explanation as proof.

REQ-ANOM-01 — CONTRADICTION and PARTIAL_CONTEXT MUST force HUMAN_ATTENTION.

5.4 Classification, Routing, and Suppression

Purpose. Keep classification separate from the higher-stakes decision to remove something from the human queue.

Behavior. ACT NOW and DECIDE always route to HUMAN_ATTENTION. EASY WIN also remains human in the MVP. IGNORE is the only bucket eligible for NO_ATTENTION, and eligibility still requires high confidence, complete context, no disqualifying anomaly, no deterministic veto, and an explicit versioned routing policy.

MVP SUPPRESSION RULE
Only high-confidence, policy-eligible IGNORE items can be suppressed. “IGNORE” is a classification; “NO_ATTENTION” is a routing decision. They are not synonyms.

REQ-ROUTE-01 — Only policy-eligible high-confidence IGNORE MAY route to NO_ATTENTION in MVP.

REQ-ROUTE-02 — Any evaluation failure, veto rule, contradiction, partial context, or insufficient confidence MUST route to HUMAN_ATTENTION.

REQ-ROUTE-03 — Routing policy MUST be versioned and recorded with each evaluation.

5.5 Today, Board, Decision Panel, and Done

Purpose. Create a finite work surface that makes the next judgment legible and then gets out of the way.

Behavior. Today is the default screen. It shows an authoritative server-provided count and cards with severity, title, repository, author, age, confidence, and provenance. A secondary Board exposes ACT NOW, DECIDE, EASY WIN, IGNORE, and Inspection context with sorting/filters. The Decision Panel shows classification first, then a short explanation, strengthening/weakening evidence, anomalies, and source links. Detailed code review stays in GitHub. When the active queue reaches zero, Done is terminal.

REQ-UX-01 — Today MUST be the default work surface and MUST derive counts from server state, not client inference.

REQ-UX-02 — The Board MUST NOT provide drag-and-drop reclassification or workflow management in MVP.

REQ-UX-03 — The Decision Panel MUST expose bucket, confidence, provenance, explanation, evidence, anomalies, and an Open in GitHub path.

REQ-UX-04 — Done MUST NOT add streaks, badges, productivity scores, tomorrow previews, or suggestions for more work.

5.6 Disagreement Signal

Purpose. Capture where maintainers reject the system’s classification without silently teaching production behavior from unreviewed feedback.

Behavior. A maintainer may select “Disagree with classification” and optionally provide a short reason. The signal is linked to the evaluation, ruleset, model/prompt version, repository, bucket, evidence types, and reason code. Product and Evaluation review clusters weekly. Production behavior changes only through a versioned proposal and regression process.

REQ-FEEDBACK-01 — A disagreement signal MUST NOT change the current routing outcome or automatically retrain/reconfigure production behavior.

REQ-FEEDBACK-02 — Material concentration around one rule/model/evidence/repository pattern SHOULD trigger manual class review.

5.7 Human Decisions, GitHub Identity, and Execution

Purpose. Let maintainers act through Gungabricks without allowing Gungabricks to become an unaccountable bot identity.

Behavior. APPROVE submits a GitHub review approval only; it never merges. RETURN submits REQUEST_CHANGES with a required structured reason or typed comment. Both execute on behalf of the authenticated maintainer using a GitHub App user access token, subject to the App’s permissions, the user’s permissions, and GitHub branch rules. GitHub remains authoritative.

Decision Assertion System effect
APPROVE “I authorize an approval review on GitHub for this exact evaluated commit.” After undo and validation, submit APPROVE on behalf of maintainer. Never merge.
RETURN “I am returning this PR with a reason.” After undo and validation, submit REQUEST_CHANGES; require reason chip or typed comment.
INSPECT “I need deeper human review.” No GitHub mutation; move to INSPECTION and allow later resolution.
DEFER “I do not need to decide now.” No GitHub mutation; reappear next configured working day at 09:00 local.

REQ-ACT-01 — APPROVE and RETURN MUST use on-behalf-of maintainer identity, not an installation token presented as the human.

REQ-ACT-02 — APPROVE MUST be disabled for a PR author attempting self-approval, and GitHub rejections MUST be treated as authoritative.

REQ-ACT-03 — RETURN MUST require one structured reason or typed comment. Default reason set: Missing tests; CI failure unresolved; Needs architecture review; Insufficient evidence/context; Other. Selecting Other MUST require a non-empty typed comment before RETURN can be committed.

REQ-ACT-04 — No model output MAY directly invoke the GitHub API.

REQ-ACT-05 — GitHub authorization, repository permissions, branch protection, and CODEOWNERS decisions MUST be treated as authoritative; the executor MUST NOT retry a non-retryable GitHub refusal.

5.8 Undo, Authorization Failure, Retry, INSPECT, and DEFER

Purpose. Make human actions reversible before mutation and honest when external execution is blocked.

Behavior. APPROVE and RETURN enter PENDING_ACTION for a server-enforced default five-second window configurable from five to thirty seconds. Undo atomically cancels the pending action. INSPECT and DEFER get the same short undo interaction for consistency, though they create no external mutation. Token expiry is a credential lifecycle, not a reason to lose the decision.

REQ-UNDO-01 — APPROVE/RETURN MUST NOT reach external execution before the configured server-side undo deadline.

REQ-UNDO-02 — Database state MUST make cancellation and execution mutually exclusive.

REQ-AUTH-01 — If user authorization cannot be refreshed after a decision, preserve the accepted decision in AUTHORIZATION_REQUIRED against its original evaluation/SHA and preserve the original undo deadline.

REQ-AUTH-02 — After re-authorization, re-check identity, repository permission, App installation, current SHA, and policy before execution.

REQ-RETRY-01 — Retryable GitHub failures MUST create visible PENDING_RETRY, use backoff/idempotency, and revalidate SHA on every attempt.

REQ-INSPECT-01 — INSPECTION MUST be non-terminal; inspected items MUST be re-openable via the same Decision Panel. A new commit invalidates the inspection evaluation and re-enters evaluation.

REQ-DEFER-01 — DEFER MUST wake at 09:00 on the next configured working day. A new commit MUST cancel the deferral and trigger re-evaluation.

5.9 Shared Queue and Multi-Maintainer Concurrency

Purpose. Allow multiple maintainers to observe the same repository queue without permitting conflicting decisions.

Behavior. Viewing is non-exclusive. The first valid decision atomically moves the item out of ACTIVE and records the acting maintainer as decision owner for that lifecycle. A racing command is rejected atomically and receives the current item state in the response.

REQ-CONC-01 — Exactly one of two racing valid decisions MAY transition the same ACTIVE item.

REQ-CONC-02 — A decision claim MUST be lifecycle-scoped; it MUST NOT silently become Phase 2 team assignment or ownership.

5.10 Onboarding, Roles, and Notifications

Purpose. Make basic operation possible without undocumented database edits or continuous engineering intervention.

Role Allowed in MVP Not allowed
Viewer View Today, Board, Decision Panel evidence, and permitted audit history. Submit decisions, change workspace settings, execute actions.
Maintainer APPROVE, RETURN, INSPECT, DEFER, and disagreement on authorized repositories. Change repository access, security settings, rule history, audit history.
Administrator Install/configure App, select repositories, set timezone/workdays/undo, assign roles, manage starter policy configuration. Bypass GitHub permissions, stale-SHA checks, or executor policy; rewrite audit history.

Behavior. The Administrator selects covered repositories, workspace timezone and workdays, user roles, undo duration, and starter policy. Notifications remain restrained: one daily digest at 09:00 local and immediate outbound notification only when a new ACT NOW item enters outside an active session. Email is sufficient for the pilot; Slack and Teams integration are out of scope for MVP.

REQ-ADMIN-01 — Every active ruleset MUST have a version, owner, and audit record.

REQ-NOTIFY-01 — DECIDE, EASY WIN, and IGNORE MUST NOT generate per-item outbound notifications in MVP.

REQ-NOTIFY-02 — The MVP MUST send one daily digest at 09:00 in the configured local timezone; immediate per-item outbound notification MUST be limited to a new ACT NOW item entering outside an active session.

5.11 Audit, Security, and Data Handling

Purpose. Make every consequential decision reconstructible while minimizing the amount of customer code leaving the dedicated environment.

Behavior. Every evaluation stores PR version/SHA, timestamp, ruleset version, model identifier/configuration, prompt version, output schema version, context hash, normalized output, evidence, routing result, and provenance. Human decisions reference the exact evaluation. Model context is bounded and redacted; credentials and unrelated repository content are excluded.

REQ-AUDIT-01 — The system MUST be able to reconstruct why a PR was hidden, surfaced, cancelled, retried, or acted on at a specific point in time.

REQ-DATA-01 — Model context MUST NOT contain GitHub tokens, private keys, webhook secrets, raw credentials, or unrelated repository content.

REQ-DATA-02 — Secret detection/redaction MUST run before repository content is transmitted to the model provider. If safe context is insufficient, the evaluation MUST record PARTIAL_CONTEXT and route to HUMAN_ATTENTION.

REQ-CONTRACT-01 — Live pilot use MUST be blocked until the selected provider’s DPA/no-training terms, retention/deletion, subprocessors, incident notice, and processing region are approved by the pilot parties.

REQ-RET-01 — At pilot exit, revoke App/user authorizations, cancel outstanding actions, export agreed audit/configuration, and delete source-derived evaluation content within 30 days unless contractually shorter or preserved under a documented legal-hold or regulatory-retention obligation.

REQ-PERM-01 — The GitHub App MUST use least privilege and limit installation to explicitly selected repositories. Repository Administration permission is not required for normal MVP operation.

REQ-INCIDENT-01 — A confirmed dangerous suppression MUST immediately stop live suppression for the affected policy scope and require remediation plus a fresh qualifying shadow sample before resumption.

6. User Personas

The Maintainer. A repository maintainer, technical lead, or senior engineer accountable for PR review and repository health. They do not need Gungabricks to teach them code review. They need it to stop wasting their attention before review begins.

The Administrator. The person responsible for installing the GitHub App, selecting repositories, setting workspace behavior, assigning roles, and maintaining the starter policy without bypassing repository or audit controls.

The Pull-Request Author. Not an active Gungabricks operator in MVP, but directly affected by RETURN. Feedback must be attributable to the maintainer and structured enough to be actionable rather than feeling like an anonymous bot rejection.

The Pilot Champion. The customer-side engineering champion who owns cohort participation, repository selection, workflow change, weekly feedback, and fast escalation of suspected misses. This is a pilot role, not a permanent product permission class.

The Security / Data Reviewer. A pilot stakeholder validating model-provider terms, data categories, retention, permissions, incident behavior, secret handling, and exit/deletion obligations.

The DBA / Data Reliability Reviewer. The reviewer who treats state transitions, concurrency, audit immutability, partitions, backup/PITR, migrations, and restore as part of the product safety case rather than as post-launch housekeeping.

7. User Stories (Illustrative)

User stories use an abbreviated “As a role, I action” form; the outcome clause is captured in the goals, requirements, and pilot metrics.

  • As a maintainer, I open Today and see only pull requests that genuinely require my judgment, so my first task is a decision rather than inbox triage.
  • As a maintainer, I open a borderline item and see evidence that strengthens and weakens the classification, with direct links back to GitHub sources.
  • As a maintainer, I approve an exact SHA and can undo before the server submits the review, even if I immediately navigate to the next item.
  • As a maintainer, I return a PR with a structured reason so the author receives an attributable GitHub REQUEST_CHANGES review instead of a bot-shaped rejection.
  • As a maintainer, I disagree with a classification and can record that signal without the product silently changing today’s behavior.
  • As an administrator, I can install the App, choose repositories, define workdays/timezone, assign roles, and review the active starter policy without requesting a database edit.
  • As a security reviewer, I can verify exactly which repository data may be sent to the model provider, which credentials are excluded, and what happens when context cannot be safely extracted.
  • As a pilot champion, I can stop live suppression immediately if a dangerous miss is confirmed and can audit what the system would have hidden during shadow mode.
  • As a DBA, I can demonstrate that two racing decisions cannot both execute, an undo cannot coexist with commit, audit history cannot be rewritten by normal application roles, and recovery works from tested backups.

8. Business Model and Why-Now

8.1 Why now

AI-assisted review is becoming ordinary, but review intelligence does not remove the attention tax of deciding what to inspect. At the same time, engineering organizations are under pressure to get leverage from senior staff without granting opaque automation broader repository authority. Gungabricks occupies that seam: use probabilistic analysis to reduce unnecessary triage, but keep policy, identity, and mutation control outside the model.

This is also a defensible moment to test the category because GitHub is the largest platform risk. If native tooling absorbs the same value proposition, the window narrows. That argues for a small, explicit proof now — not for premature moat claims.

8.2 Initial customer profile

The first design partner should have roughly 15–30 maintainers, 15–50 active repositories, 100–500 daily pull requests, a visible triage burden, and enough operational discipline to independently review suppressed predictions. The organization must accept dedicated vendor-hosted deployment, baseline measurement, and the right to halt live suppression immediately when a trust gate fails.

8.3 Value model

The economic equation is intentionally simple: maintainers in scope × active triage hours per week × loaded hourly cost × working weeks × realized time reduction. Attention Reduction (M1) is not ROI by itself. If maintainers continue reading the same GitHub notifications and then also process Gungabricks, the product can “suppress” work on paper while returning no time at all. That is why adoption and realized time reduction are product metrics, not post-hoc customer-success narrative.

Scenario Maintainers Triage hrs/wk Loaded cost/hr Realized reduction Annual value returned
Conservative 12 4 $100 40% $96,000
Base planning case 20 5 $120 60% $360,000
Upside 30 6 $140 60% $756,000

These are planning scenarios, not customer claims. The baseline replaces every input with observed or customer-approved values.

8.4 Pricing and unit-economics hypothesis

The working commercial model is value-led enterprise subscription with volume guardrails: approximately $8,000/month up to ~5,000 evaluated PRs/month and $12,000–$15,000/month up to ~15,000. This is a testable hypothesis, not published pricing.

Dedicated single-tenant deployment is inherently more expensive per customer than shared infrastructure; that cost is accepted for the first pilot. Planning assumptions put model spend at roughly $750/month at 15,000 PRs if M7 holds, with dedicated infrastructure/observability/email/backups/storage adding roughly $3,000–$7,000/month before human support. The environment must be right-sized to observed load rather than permanently provisioned for the ceiling. Commercial scale requires either willingness to pay near the upper range or a credible shared-infrastructure cost-down path. Target steady-state software gross margin after that path: at least 70%.

8.5 Controlled pilot

The pilot has three measurement phases: two working weeks of baseline, four working weeks of shadow evaluation, and four working weeks of live suppression. Shadow and live results are reported separately because “what we would have hidden” is not the same business result as “work humans actually stopped doing.”

Baseline uses a small browser instrument that records only active time spent on GitHub PR pages, count of PR pages viewed, and timestamps needed to distinguish active use from idle tabs. It does not capture source code, diff content, comments, keystrokes, or page text. Team-level reporting is supplemented by a short structured end-of-day self-report; the pilot champion approves the privacy notice and telemetry fields.

Shadow mode suppresses nothing. It exists to create a safe counterfactual and an independently reviewable safety sample. Graduation requires zero confirmed dangerous suppressions in at least 1,000 independently reviewed NO_ATTENTION predictions, giving an approximate 95% rule-of-three upper bound of ~0.3%. At least 10% of the sample is double-reviewed. One confirmed dangerous suppression fails the gate and resets the qualifying sample after remediation.

Live suppression begins only after the shadow gate passes. The system removes only high-confidence policy-eligible IGNORE items from the Gungabricks human queue, audits at least 100 suppressed items per week plus every reported incident, and measures realized behavior change. A confirmed dangerous suppression stops live suppression immediately for the affected scope.

9. Risks and Open Questions

Risk Why it matters Required response
Dangerous suppression A critical PR could be hidden from the maintainer who should have seen it. Shadow gate; independent sampling; IGNORE-only suppression; immediate conservative mode after any confirmed event.
GitHub absorbs the category GitHub owns the system of record and can expand native inbox, AI review, rules, and merge flow. Differentiate on attention suppression, evidence, reversibility, governance, and audit; validate quickly; preserve multi-SCM option value.
Adoption / second inbox Users may keep old notification habits and add Gungabricks on top. Pilot champion; workflow training; measured adoption; reversible notification changes; H2/H4 gates.
Model-provider data exposure Customer code leaving the environment can block procurement or create IP risk. DPA/no-training terms; approved region; secret redaction; minimal context; retention/deletion commitments.
Execution identity mismatch Bot attribution or wrong identity can fail branch protections and accountability. User-to-server token; branch/CODEOWNERS integration tests; GitHub remains authoritative.
Prompt injection / adversarial repository content Repository text may try to manipulate model classification. Treat content as data; separate authority; no credentials; adversarial regression corpus; fail HUMAN_ATTENTION.
Single-tenant cost Dedicated environments may not support scalable unit economics. Track per-customer run cost; second-partner trigger for multi-tenant architecture; price/cost-down validation.
Pilot overfitting Rules tuned to one partner may not generalize. Treat first pilot as product proof, not market proof; require second independent design partner.
GitHub rate limits / outages Valid decisions may be delayed; reconciliation can fall behind. PENDING_RETRY; backoff; idempotency; SHA revalidation; visible degraded state.
Approval habituation Fast queue actions can become rubber stamps. Monitor M4 rather than optimize it; keep evidence visible; review disagreement and undo rates; high-risk policy stays human.

Open questions are narrower than the risks: which model provider meets the required contract; which design partner will accept the measurement method; what starter thresholds fit its repositories; whether early willingness to pay supports dedicated economics; and how quickly a second customer justifies multi-tenant investment. Each should be answered with evidence from the pilot, not extrapolation.

10. Roadmap (Phased)

Phase 0 — Specification and partner readiness. Lock the GitHub permission manifest, model/data contract, starter rules, pilot measurement method, repository cohort, incident contacts, and data-exit terms. No live customer code is sent to a model until those boundaries are agreed.

Phase 1 — MVP operating loop. Build signed intake and reconciliation; immutable PR/SHA state; deterministic rules; bounded model evaluation; evidence/anomalies; conservative routing; Today/Board/Decision Panel/Done; roles/settings; decisions; server-side undo; on-behalf-of execution; retry/authorization handling; append-only audit; observability; backup/restore; DBA review; acceptance and adversarial tests.

Phase 1A — Baseline and shadow. Measure current triage behavior, then evaluate every in-scope PR without suppression. Independently review the NO_ATTENTION sample. Do not graduate on average accuracy; graduate on the dangerous-suppression gate.

Phase 1B — Live suppression. Suppress only eligible IGNORE items, audit suppressed samples continuously, make Gungabricks the starting point for covered triage, and measure adoption plus realized time returned.

Phase 2 — Controlled operational autonomy. If the MVP proves safe attention reduction and workflow adoption, candidates include deterministic low-risk automation, richer anomalies, cross-PR reasoning, deeper inspection workflow, additional defer triggers, explicit human reclassification, policy-management UI, multi-repository workspaces, team assignment/ownership, and administrative analytics. The model still does not receive independent action authority.

Phase 2A — Multi-tenant commercialization. A second independent design partner is the trigger to begin architecture/funding work for shared tenant management, billing/entitlements, automated provisioning, tenant-isolation tests, migrations, and fleet operations. Planning assumption: 8–12 engineering weeks and ~$250k–$400k additional product/platform investment, subject to separate validation.

Phase 3 — Adaptive intelligence and ecosystem breadth. Candidate capabilities include learned preferences, historical risk modeling, cross-repository dependency graphs, simulation/model-evaluation tooling, advanced enterprise governance, and additional SCM providers such as GitLab and Bitbucket.

11. Open Decisions

Decision Accountable parties Required before
Name/contract pilot partner; confirm maintainer/repo volume, sponsor/champion, hosting region, and baseline privacy plan. Product + pilot partner Shadow start
Select model provider; approve DPA/no-training terms, retention, region, subprocessors, incident posture. Security / Legal / Product Any live customer-code evaluation
Approve GitHub App permission manifest, expiring user-token refresh design, and branch/CODEOWNERS integration test plan. Tech Lead / Security MVP pilot installation
Confirm starter rule thresholds, sensitive paths, large-diff limits, and adversarial-content policy with repository owners. Product / Tech Lead / pilot champion Shadow start
Confirm commercial pilot terms and willingness-to-pay test. Product + pilot commercial sponsor Pilot contract
Replace planning labor/cloud costs with actual build assumptions. Product / Engineering / Finance or equivalent Investment decision
Set retention/archive period for source-derived content versus immutable operational audit. Security / Legal / DBA Data-processing approval
Confirm final interactive latency and WCAG 2.1 AA pilot criteria. Product / Engineering / Design Pilot readiness
Define second-design-partner threshold that triggers multi-tenant architecture work. Product / Engineering Post-pilot commercialization decision

Appendix A. Success Measures and Pilot Gates

Metric Definition Target / interpretation
M1 — Attention Reduction Evaluated PRs legitimately routed to NO_ATTENTION during live suppression. Provisional >=60%; confirm after baseline/shadow.
M2 — Human Disagreement Rate Surfaced items where maintainer selects “Disagree with classification.” Provisional <=10%; >15% blocks expansion pending root-cause analysis.
M3 — Dangerous Suppression NO_ATTENTION prediction later judged, using information available at evaluation time, to have required human attention for material correctness, security, data integrity, availability, or release risk. Shadow: 0 confirmed in >=1,000 independently reviewed predictions; ~95% upper bound <=0.3% rule of three. Any confirmed event fails gate.
M4 — Decision Time Active time from opening HUMAN_ATTENTION item to submitting decision. Monitor only; investigate median <20 sec or >5 min.
M5 — Stale Action Protection Block APPROVE/RETURN against changed SHA. 100% of induced stale-action test cases refused; zero production incidents.
M6 — Evaluation Latency p95 Canonical PR version ready → routing persisted. <=15 sec within operating envelope.
M7 — Model Cost per PR Average third-party model cost per evaluated PR. Provisional <=$0.05 under bounded-context policy.
M8 — Webhook Processing Success Valid signed webhooks persisted/normalized without manual intervention. >=99.9%; reconciliation measured separately.
Indicator Purpose Pilot expectation
H1 — Time to Surface Webhook receipt → eligible item visible in Today. p95 <=30 sec.
H2 — Workflow Adoption Eligible triage decisions performed through Gungabricks instead of parallel manual triage. >=70% by second half of live; <50% no-go signal.
H3 — Undo Rate Submitted decisions reversed during undo window. Monitor; sustained >5% triggers UI/error review.
H4 — Realized Triage-Time Reduction Baseline active triage time vs. live, adjusted for adoption. Target >=30% realized reduction.

Pilot decision rules

Outcome Required evidence
GO — proceed to second-customer / Phase 2 validation Safety gate green; no confirmed dangerous suppression in live audit; adoption >=70%; realized triage-time reduction >=30%; reliable operations; sponsor wants to continue; pricing supports credible path to >70% steady-state gross margin after cost-down.
HOLD — one remediation extension Safety remains green, but one specific remediable adoption/value/operations issue remains. One two-working-week extension with named owner and unchanged safety gate.
NO-GO — stop or materially redesign Unresolved dangerous suppression; inability to satisfy data/identity requirements; live adoption <50% after intervention; realized time reduction <20%; sponsor unwilling to continue; or no credible willingness-to-pay/cost structure.
KILL immediately Unauthorized repository mutation, material customer-code exposure outside approved terms, or repeated dangerous suppression after remediation.

Appendix B. Operational and Technical Guardrails

Requirement Product meaning
Correctness Every external action applies only to the exact PR version the maintainer examined.
Auditability Every automated and human decision can be reconstructed from retained, immutable history.
Idempotency Duplicate events and retries cannot create duplicate external reviews or state transitions.
Security Repository content cannot grant authority, expose credentials, or widen model access.
Recoverability Backup/restore, including audit partitions and key dependencies, is tested before live suppression.
Availability Model-provider outage keeps repository state available and causes conservative routing, not silent loss.
Performance Today and Decision Panel remain interactive; H1/M6 are processing gates.
Data minimization Only evaluation-relevant customer content is transmitted or retained.
Accessibility Core triage/decision controls support keyboard operation and pass agreed WCAG 2.1 AA review before pilot.

The MVP should remain a modular monolith: frontend, backend/API, workers, scheduler, PostgreSQL, durable task mechanism, object storage, GitHub App, outbound notification service, and model-provider gateway. Internal modules separate identity, providers, ingestion, repositories, policy/rules, evaluation, evidence, routing, attention, decisions, actions, and audit. These are code boundaries and MUST NOT be decomposed into microservices in MVP.

PostgreSQL is the system of record for repositories, pull requests and immutable versions, webhook events, rulesets, evaluations, evidence, anomalies, attention items, decisions, actions, policies, authorizations, and audits. Large raw payloads or evaluation artifacts may live in object storage while hashes and metadata remain relational. Audit history is append-only and time-partitioned.

Durable jobs include webhook processing, reconciliation, PR normalization, evaluation, routing, pending-decision validation, deferred wake-up, notifications, GitHub retry, and audit-sample scheduling. Kafka is out of MVP scope. A PostgreSQL-backed durable queue is acceptable if it satisfies concurrency, recovery, and observability requirements.

Minimum telemetry includes webhook volume/failure, processing delay, queue depth, reconciliation errors, evaluation latency/failure, model cost, GitHub API errors/rate limits, attention-queue size, PENDING_RETRY and AUTHORIZATION_REQUIRED volume/age, stale-action refusals, commit failures, suppression volume, and notification delivery. Structured logs must exclude raw secrets and source diffs.

Appendix C. Model Output Contract and Input Boundary

A valid model-assisted evaluation returns exactly one bucket — ACT NOW, DECIDE, EASY WIN, or IGNORE — plus qualitative confidence (high/medium/low), optional alternative bucket, short human-readable reason, structured evidence, and optional anomaly suggestions. Each evidence item includes a claim, strengthens/weakens direction, and references. Routing is not part of model authority.

Structural validation is mandatory. Typed validation, JSON Schema, or provider-supported structured outputs are acceptable. Malformed output fails fast. No heuristic JSON repair may turn invalid output into NO_ATTENTION.

The input contract is bounded. The default full-evaluation limit is 1,500 changed lines and 50 files, subject to pilot tuning. Larger or binary-heavy changes, or cases where trustworthy context cannot be safely extracted within the cost envelope, may receive metadata-level analysis but must set PARTIAL_CONTEXT and force HUMAN_ATTENTION. Silent truncation followed by high-confidence IGNORE is forbidden.

Appendix D. Database and Data Integrity Review

The database review is a production gate because state is part of the safety model. A system that cannot prove the order and exclusivity of its own transitions cannot credibly claim that actions are reversible, stale-safe, or auditable.

  • Relational schema, primary/foreign keys, uniqueness, nullability, deletion semantics, and intentional JSON use.
  • Indexes and query plans at the stated 50-repository / 500-PR-per-day operating envelope.
  • Isolation levels and locking semantics for attention claims, decision races, undo-versus-commit, retries, and authorization recovery.
  • Append-only audit schema permissions; ordinary application roles cannot update or delete historical rows.
  • Time-based partitioning, archival policy, retention boundaries between source-derived content and operational audit.
  • Backup, point-in-time recovery, restore drills, encryption-key dependencies, and documented recovery objectives.
  • Migration/rollback procedure, including behavior when a migration is interrupted or must be reversed.
  • Database roles, credential separation, and least privilege for API/workers/audit/operations.

DBA sign-off MUST cover the partition key, archival policy, restore path, migration plan, and concurrency semantics before live suppression.

Appendix E. MVP Acceptance Matrix

Area Go-live acceptance test
Intake Signed webhook persisted before async processing; duplicate delivery IDs cannot create duplicate work.
State Canonical PR state bound to exact SHA; intentionally missed webhook repaired by reconciliation.
Evaluation Starter rules execute; model output structurally validates; evidence persists; partial/invalid output falls back to HUMAN_ATTENTION.
Routing Only policy-eligible high-confidence IGNORE can become NO_ATTENTION; all vetoes/anomalies route as specified.
Queue concurrency Two maintainers racing on one ACTIVE item cannot produce two valid decisions.
Identity APPROVE/RETURN execute on behalf of authenticated maintainer; self-approval blocked; GitHub branch rules remain authoritative.
Decisions APPROVE, RETURN, INSPECT, DEFER, disagreement signal work; RETURN reasons and Inspection lifecycle behave as specified.
Undo APPROVE/RETURN cannot execute before server-side deadline; cancellation and execution mutually exclusive.
Stale safety SHA change before execution/retry refuses action and triggers re-evaluation.
Retry / authorization Transient failure → visible PENDING_RETRY; auth failure → visible AUTHORIZATION_REQUIRED preserving decision/deadline and revalidating before resume.
Security Secret redaction, provider contract boundary, least privilege, credential rotation, uninstall invalidation, adversarial injection, and deletion flow tested.
Audit Rules/models/prompts/evidence/decisions/cancellations/retries/identity/outcomes reconstructible and immutable.
Operations Restore test, partition strategy, observability, support runbook, adversarial suite, and DBA review signed off.
Shadow gate 0 confirmed dangerous suppressions in >=1,000 independently reviewed NO_ATTENTION predictions; ~95% rule-of-three upper bound <=0.3%.
Live gate M1/H2/H4 and reliability meet pilot criteria with no confirmed dangerous suppression in live audit.

Appendix F. Delivery Plan, Team, and Planning Cost

The schedule below is an illustrative build sequence if work begins in September 2026. It is not an executive commitment. Dates assume the pilot partner and model provider are contracted by kickoff; slippage in either MUST move the schedule. Safety-critical dependencies are allowed to move the date.

Milestone Illustrative window Primary outcome Planning cost
A — Foundations Sep 7–Sep 25, 2026 GitHub App, signed intake, durable storage, canonical PR state, SHA/version model, reconciliation. $60k–$90k
B — Evaluation Sep 21–Oct 16 Starter rules, bounded model gateway, schema validation, evidence/anomalies/versioning, golden/adversarial dataset. $80k–$120k
C — Routing + Queue Oct 5–Oct 30 Versioned routing policy, shared queue, multi-maintainer concurrency, authoritative counts/live updates. $55k–$85k
D — Product Experience Oct 19–Nov 13 Settings, Today, Board, Decision Panel, disagreement, Inspection, Done, notifications. $80k–$120k
E — Actions + Security Nov 2–Nov 27 On-behalf-of identity, undo, SHA guard, executor, RETURN, retry/auth, permissions/data controls. $85k–$130k
F — Hardening + Pilot Readiness Nov 23–Dec 11 Observability, backup/restore, DBA review, adversarial tests, acceptance suite, support/pilot runbooks. $60k–$95k
Baseline + Shadow + Live ~10 working weeks plus holiday pause 2-week baseline, 4-week shadow, 4-week live suppression. $30k–$60k incremental external/run cost + retained team support
Role Typical allocation Responsibility
Product lead 0.5–1.0 FTE Pilot requirements, customer alignment, metrics, decision policy, product governance.
Tech lead / backend 1.0 FTE Architecture, state model, GitHub integration, concurrency, executor.
Backend / platform engineers 2.0 FTE Ingestion, evaluation, routing, data, jobs, audit, observability.
Product / frontend engineer 1.0 FTE Today, Board, Decision Panel, Settings, notifications, keyboard workflow.
Evaluation / applied AI engineer 0.75–1.0 FTE Model gateway, context extraction, schemas, evaluation harness, golden dataset.
Product designer 0.25–0.5 FTE Interaction design, trust communication, pilot usability.
Security / SRE 0.25–0.5 FTE Secrets, data handling, threat model, deployment, incident readiness.
DBA / data reliability 0.1–0.2 FTE Schema, concurrency, partitions, restore, migration review.

Planning envelope: approximately $450k–$700k total MVP build plus controlled pilot. Replace this range with organization-specific labor and cloud costs before any investment decision. The low end assumes platform reuse and internal security/DBA support; the upper end allows integration complexity, external review, and pilot hardening.

Appendix G. Pilot Runbook

Phase Entry criteria Required activities Exit artifact
Preparation Model/data and GitHub permission boundaries agreed; design partner contracting. Confirm champion; repo/maintainer inventory; working days/timezone; baseline privacy notice; starter-rule workshop; incident contacts. Signed pilot scope, permission manifest, data terms, rule catalog, measurement plan.
Baseline — 2 weeks Cohort enrolled; measurement approved/installed. Measure active PR-page time/PRs viewed; self-report; record notification workflow and historical volume; no suppression. Baseline report with participation, error/missing-data range, loaded-cost inputs, provisional targets.
Shadow — 4 weeks MVP acceptance complete; partner scope active. Evaluate all in-scope PRs; suppress nothing; independently review stratified NO_ATTENTION sample; review M2 clusters; run adversarial/operational drills. Shadow report; >=1,000 reviewed predictions; M3 bound; signed live-start decision.
Live — 4 weeks Shadow gate passed; notification rollback tested. Enable eligible IGNORE suppression; make Gungabricks the triage start; audit >=100 suppressed/week; track H2/H4, retries, authorization blocks, incidents, and cost. Live results pack and completed decision scorecard.
Decision / exit Live complete or kill criterion triggered. GO/HOLD/NO-GO; if HOLD, one two-week remediation plan; export audit; restore notifications if stopping; uninstall/revoke/delete per policy. Decision record, exit/continuation record, next-investment recommendation.

Standing weekly agenda

Order Agenda item Owner / output
1 Safety: dangerous-suppression audits, incidents, conservative-mode status. Engineering + Product; written safety status.
2 Data/identity: provider, redaction, authorization failures, App permission changes. Security / Engineering; exceptions/actions.
3 Adoption/value: H2, H4, baseline quality, notification behavior. Product + pilot champion; interventions.
4 Evaluation quality: M1/M2 clusters, anomalies, golden/adversarial regressions. Evaluation lead; review queue and versioned proposals.
5 Operations/economics: latency, retries, run cost, model cost, support load. Engineering + commercial/finance owner; actions.
6 Gate readiness / decisions. Pilot sponsor/delegate; decisions and owners logged.

Appendix H. Glossary and Revision History

Term Meaning
Attention item A pull request routed to the durable human attention queue.
Attention-suppressed item A policy-eligible PR routed to NO_ATTENTION and therefore not placed in the human queue.
Bucket ACT NOW, DECIDE, EASY WIN, or IGNORE.
Confidence Qualitative evaluation confidence: high, medium, or low.
Provenance Origin of a conclusion or change: deterministic rule, model, or human.
Alternative bucket A plausible secondary classification for a borderline case.
STALE No configured material activity for longer than the stale threshold.
CONTRADICTION Rules/evidence materially disagree in a way that forces HUMAN_ATTENTION.
PARTIAL_CONTEXT Evaluator could not safely inspect enough of the change; forces HUMAN_ATTENTION.
Evidence Structured claim, strengthens/weakens direction, and source reference(s).
PR version Immutable snapshot of a pull request at an exact commit SHA.
Dangerous suppression A NO_ATTENTION prediction later judged to have required human attention using information available at evaluation time.
PENDING_ACTION Server-held decision inside the undo window.
PENDING_RETRY Valid human decision awaiting external execution after retryable GitHub failure.
AUTHORIZATION_REQUIRED Accepted decision blocked because on-behalf-of GitHub authorization cannot currently be refreshed/used.
Decision owner Maintainer whose atomic decision command moved an ACTIVE item into its decision lifecycle; not long-lived team assignment.
Disagreement signal Lightweight feedback that a maintainer would classify/route differently; not operational reclassification.
On-behalf-of identity A GitHub App user-access-token pattern in which an external action is submitted under the authenticated maintainer’s identity and permissions rather than as an installation bot.
Version Date Status Material change
0.1 Aug 25, 2026 Draft Initial detailed MVP requirements and roadmap.
0.2 Aug 25, 2026 Business rewrite Reframed technical outline as business-facing PRD.
0.3 Aug 25, 2026 Cross-functional candidate Added trust model, initial metrics, operating envelope, action/retry semantics.
0.4 Aug 25, 2026 Pilot/investment candidate Added pilot design, competitive risk, data handling, GitHub identity, M2/M3, business case, pricing, change management.
0.5 Aug 25, 2026 Gated sign-off candidate Added funding gates, token-refresh UX, Inspection lifecycle, adversarial harness, baseline instrumentation, economics, scorecard, runbook.
0.6 Aug 25, 2026 Project 20 draft Rebuilt in Project 20 PRD template and Redwin voice; normalized roles; removed assumed corporate approval scaffolding; retained safety/DBA/pilot substance.
0.7 Aug 25, 2026 Partner circulation draft Removed working-draft meta-references; tightened normative requirements and partner-facing register; clarified GitHub authority, notification defaults, retention, and suppression wording.

External Product and Platform References

These references were checked on August 25, 2026 and should be re-validated before implementation; GitHub identity, token lifecycle, and platform behavior can change.

  • GitHub Apps acting on behalf of a user
  • GitHub App user access-token refresh and expiry
  • GitHub REST API — pull request reviews
  • GitHub reviewing proposed changes / self-approval rules
  • GitHub code review product surface
  • GitHub Copilot code review behavior
  • Graphite product / PR inbox / AI review
  • CodeRabbit AI code review
  • Mergify workflow automation

End of document — Gungabricks PRD v0.7 (Partner circulation draft). Project 20 Gem. Red Anvil Productions.

No comments:

Post a Comment