Autonomous bug pipeline — workflow definition¶
Status: partly built. This document is the shared vision, written before any of it existed. Phases 1–5 of §12 have since shipped — the bug spec side (
dor-bug-agent), the repro contract, and theDOR_AUTOBUILDcarve-out that lets a certified bug skip the value gate are all live. Later phases are still a proposal, and anything here that contradicts dor-state-machine.md or operationalization.md is out of date — those describe what runs today. The feature pipeline is unaffected — see §10.
A reported bug that the DoR probe has certified — reproduced against real code, root cause pinned, regression test drafted — goes from report to a merge-ready PR without a human in the loop. The only human action is the merge review, and it is a review of evidence, not of a promise.
That review has three answers, not one: approve, "let me see it running myself", or "the evidence isn't enough" — and the last two are one command each, not manual work (§5).
1. Why a bug needs no value gate¶
The value gate exists because "should we build this?" is a genuinely contestable question — for a feature. Someone has to weigh desirability, product coherence, and opportunity cost, and no amount of AI diligence answers it. That gate stays.
For a bug, that question is already answered, and not by us: the product made a promise and broke it. Once the probe has certified that the break is real and reproducible, "is it worth fixing?" has no interesting answer left. Approving it is ceremony.
But today's single gate quietly bundles three different jobs. Splitting them is what makes it safe to drop:
| Function of the gate | The question it answers | For a certified bug |
|---|---|---|
| Value | Is this worth building at all? | Answered by certification → drop |
| Spend | What will this cost in runner time and model quota? | Real, but a policy question — bound it with concurrency + budget caps, not a per-issue click (§8) |
| Blast radius | Should the AI change this part of the codebase unsupervised? | Real, and the one worth keeping — but as an automatic rule, not a human judgement call (§6) |
The trade this makes: the gate moves from before the work to after it. Before, a human guesses
whether the work is worth doing. After, a human reads what was actually done, with proof. The second
is a better decision made on better information — and it already exists as a required step, because
main is branch-protected and needs an approving review.
That only holds if the machine can prove its work. That is the rest of this document.
2. The three proof obligations¶
Autonomy is bought with falsifiability. Each obligation below exists because without it a specific failure mode passes silently.
2.1 A falsifiable certification — the repro contract¶
Today the probe emits prose plus a route label. Prose cannot be checked. If certification is the only thing standing between a bug report and an autonomous code change, it must be a contract the build is then measured against.
The probe additionally emits .dor/out/contract.json, validated by the deterministic post step the
same way route.txt already is (schema check → reject → no action). Fields:
| Field | Purpose |
|---|---|
symptom |
The observable defect, in the reporter's terms |
assertion |
The specific assertion that must FAIL before the fix and PASS after |
root_cause |
Layer + file(s) that produce the wrong value — the "fix at the source" target |
blast_radius |
Globs the fix is predicted to touch |
test_tier |
unit | api | e2e — where the regression test belongs |
repro_path |
The user-visible route to the symptom, for the live-env replay |
confidence |
certain | likely — likely routes to a human, never to an autonomous build |
duplicate_of |
Issue number, if the backlog scan found one |
The contract is what makes the later checks mechanical instead of narrative.
2.2 Red before green¶
A test that was never red proves nothing. A green suite after a fix is equally consistent with "bug fixed" and "test doesn't touch the bug" — and an AI that writes both the fix and its test has every opportunity to produce the second by accident.
So the build commits in a fixed order, and the flow — not the model — checks it:
- Test-only commit. Run it. It must fail, and fail on
contract.assertion. If it passes, the test does not reproduce the bug → Exceptions. This is cheap: same checkout, no Docker. - Fix commit. Re-run. It must pass.
Both runs are captured verbatim into the evidence bundle. The red output is the single most valuable artifact the pipeline produces, because it is the only one that cannot be faked by optimism.
2.3 The symptom is gone on a real deployment¶
Unit-tier green is not "the bug is fixed" — it is "one assertion changed state". The reporter's symptom must be replayed against the actual running product.
The build deploys the branch to its sidekick, seeds demo data, runs the context plugins, then replays
contract.repro_path as an e2e against https://N.build.identityatlas.io.
The current smoke fallback must die for bugs. run_feature_e2e() today runs only the e2e specs
the branch touched, and falls back to "does the app serve?" when the branch touched none — which a
unit-only bug fix would sail straight through, and be reported as verified. For a bug: no e2e means
no pass.
3. The workflow¶
flowchart TD
A[Bug reported via Bug Form] --> B[Triage: board + Requested-by]
B --> C[Probe: reproduce, root cause, draft test, emit contract]
C -->|not reproducible / needs info| D[Awaiting reporter]
C -->|confidence: likely, or blast radius restricted| E[Human review queue]
C -->|certified| F[Queue for build]
F --> G[Claim a pool sidekick + reserve it]
G --> H[Commit test only -> MUST FAIL on the contract assertion]
H -->|passes: does not reproduce| X[Exceptions]
H --> I[Commit fix -> MUST PASS]
I --> J[Deploy + seed on the sidekick]
J --> K[Replay reporter symptom as e2e on the live env]
K -->|still reproduces| X
K --> L[Open PR incl. regression test, changelog, docs]
L --> M[Drive CI to green, auto-fix up to N]
M -->|red after N| X
M --> N[Post evidence bundle to the PR]
N --> O[Release the sidekick back to the pool]
O --> P[Human merge review — the only gate]
P -->|approved| Q[Merged -> issue closed, board Done]
P -->|I want hands on it| R[Validation session: re-claim a box, deploy, hand over the URL]
P -->|evidence insufficient| S[Proof-gap loop: strengthen the test, re-prove red vs main]
R --> P
S --> J
| # | Step | Actor | Trigger | Artifact | On failure |
|---|---|---|---|---|---|
| 1 | Intake | Reporter | Bug Form | issue | — |
| 2 | Triage | deterministic | issues.opened |
Bug board item, Requested-by | non-member → notice + assign |
| 3 | Probe / certify | AI (Fable 5) | issues, issue_comment |
verdict comment + contract.json + route |
route to reporter / design / duplicate / out |
| 4 | Queue | deterministic | route = certified |
queue position | over concurrency cap → wait |
| 5 | Claim runner | deterministic | free dor-build runner |
reservation written at claim, not at PR-create | no free runner → stay queued |
| 6 | Red proof | flow | — | failing test output | test passes → Exceptions |
| 7 | Fix | AI (Opus 5) | — | fix commit | no changes produced → Exceptions |
| 8 | Green proof | flow | — | passing test output | still red after N → Exceptions |
| 9 | Deploy + seed | flow | — | live env URL | infra failure → Exceptions (not a fix loop) |
| 10 | Live replay | flow | — | e2e run + trace | symptom persists → Exceptions |
| 11 | PR + CI | flow | — | PR, CI checks | red after N auto-fixes → Exceptions |
| 12 | Evidence + release | deterministic | CI green | evidence bundle comment; sidekick reset | — |
| 13 | Merge review | human | PR ready | approval | → step 14 or 15 |
| 14 | Validation session (optional) | deterministic | /validate on the PR |
re-claimed box, deployed branch, live URL held for the human | TTL expiry → release the box, PR untouched |
| 15 | Proof-gap loop (optional) | AI + flow | reviewer objection | strengthened test, red re-proved against main, appended bundle |
> N iterations → Exceptions |
Steps 6–11 are the Definition of Done. All must hold; any failure routes to Exceptions with the evidence of why, and never silently degrades to "probably fine". Steps 14–15 are the reviewer's two ways of saying "not yet" — see §5.
4. The evidence bundle — what "prove it to me" means¶
Posted as a single structured PR comment when CI goes green. Every row is a claim with a link to the line in a run log that substantiates it. A reviewer reads a checklist, not a story.
| Claim | Evidence |
|---|---|
| The bug was real and reproducible | Certification comment + contract.json |
| The test reproduces this bug | Red run output — the failing assertion, pre-fix |
| The fix resolves it | Green run output, same test, post-fix |
| The reporter's symptom is gone | Live-env e2e result + trace/screenshot + the N.build URL |
| It cannot come back | The regression test now in main's CI — named, linked |
| The fix is where the diagnosis said | Files touched vs contract.blast_radius, diffed |
| Nothing else broke | Full CI status; coverage delta (line + branch, per-file ratchet) |
| It was fixed at the source | The root-cause layer, quoted from the contract, vs the files actually changed |
| What it did not do | Explicit: scope not expanded, no schema/migration, no .github/, no dependency bumps |
The live env stays up until the PR closes. Available for a look — not something anyone has to block on.
5. The review loop — when the reviewer is not satisfied¶
The merge review is now the only gate, so it is only a real gate if "no" is as cheap and as actionable as "yes". A reviewer who has to choose between rubber-stamping and doing the work by hand will rubber-stamp. Two distinct kinds of "not yet", each with its own machinery:
5.1 "I want to see it for myself" — a validation session¶
The evidence may be complete and you still want your hands on the running product. That is a legitimate, permanent need, not a failure of the pipeline.
- Trigger:
/validateas a PR comment, by any org member (same membership gate as everywhere). - Effect: claim a pool sidekick, deploy this PR's branch, seed demo data + run the context plugins, and post the URL together with the reproduction path from the contract and the exact steps the automated replay performed — so you can check the same thing by hand, or deliberately check something else.
- It is a session, not a deployment. The box is held for you: TTL, a warning comment before it
expires, one word to extend, and immediate release on merge, close, or
/release. - The URL is per session, not per PR. The build's box was released back to the pool at CI-green
(nobody's sidekick sits idle overnight), so a validation session usually lands on a different
box with a different
N.buildURL. Never trust an older URL from earlier in the thread. - Priority: a validation claim preempts queued autonomous builds. A waiting human is more expensive than a waiting bot.
- Board: stays Awaiting merge, plus a
manual-validationlabel so the reconcile sweep does not read a human-held env as a stuck build.
5.2 "The evidence is insufficient" — a proof-gap loop¶
This is the more important half, and the one that keeps the whole bargain honest. If the bundle does not convince you, that is a defect in the proof, and the pipeline — not you — should close it.
- Trigger: the objection in your own words on the PR: "the e2e only covers the on-screen path, not the export", "the red run proves the unit case, not the reported one", "I don't see the owner-role variant covered".
- Effect: the objection is treated as an amendment to the Definition of Done, not as chat. It goes back to the build agent — extend or replace the test, widen the assertion, replay a different path — and then the entire chain re-runs: deploy, live replay, CI, new bundle.
- The subtlety that makes this hard to get right: red-first is free on the first pass because
commit ordering supplies it. On a loop iteration it is not — the fix is already committed, so a
newly added test cannot be proven red by ordering. It must be proven red against
origin/main: run the new test on a clean pre-fix tree, where it must fail, then on the branch, where it must pass. Skip that and loop iterations quietly degrade into unproven tests — precisely the failure mode §2.2 exists to prevent. - The bundle is appended, never replaced — "Evidence v2 — what changed since your objection" — so the trail shows what your "no" actually bought.
- Bounded: capped iterations, then Exceptions and a human takes the wheel. Each iteration counts against the budget guard.
Both paths converge on the same gate: CI green, evidence posted, human approval. Nothing auto-merges, ever.
6. Stop conditions — when the machine must escalate¶
An autonomous pipeline is only trustworthy if it is eager to stop. Every one of these routes to Exceptions with a maintainer @-mention, and none of them are recoverable by retrying harder:
| Condition | Why it stops |
|---|---|
confidence: likely in the contract |
Uncertain diagnosis is exactly where autonomy is worst |
| Test passes before the fix | Does not reproduce the bug → the whole proof chain is void |
Fix touches files outside blast_radius |
The diagnosis was wrong, or scope crept — either way a human decides |
| Touches a schema migration, auth/security path, or crawler credential handling | Blast radius a review cannot cheaply undo |
| Diff exceeds size limit (files / lines — value in §13) | "Bug fix" that is really a refactor |
| No test added | Violates the DoD and the coverage ratchet |
| Coverage down, or diff-coverage gate red | Repo hard rule |
| Live replay still shows the symptom | The thing we set out to prove failed |
| CI red after N auto-fix attempts | Flailing; a human reads it faster |
| Review loop past N iterations | The objection is not something more AI passes will close |
| Sidekick died mid-flight | The flow dies with it, so it can never route itself here. The hourly reconcile sweep detects "active phase, no workflow run alive behind it" and flags + comments once (dor-stuck). This is the only stop condition that cannot be self-reported — everything else in this table assumes the pipeline is alive to report it |
| Model usage limit | Not a failure — pause, save the branch, resume (existing dor-resume) |
| Validation-session TTL expired | Not a failure — release the box, leave the PR exactly as it was; /validate again any time |
7. Runner lifecycle¶
Unchanged in shape, three fixes the loss of the human gate makes load-bearing:
- Reserve at claim, not at PR-create. Today
~/.dor-reservationis written after the PR opens; a build that dies before that leaves a box that looks free but has a stack on it. With no human pacing the queue, that collides. - Sweep stale reservations. (built) The reconcile sweep — now hourly — flags an issue that
still claims an
sk:*sidekick with no open PR, and a closed issue that never released one. Without a human gate, a stranded runner silently shrinks the pool until it starves. - A box can be claimed twice in a PR's life — once by the build, later by a validation session
(§5.1). The reservation file therefore
records why it is held (
buildvsvalidation) and, for a session, its expiry. A human-held box must never be swept as a stuck build, and a session must be released on merge/close even if the human never says so.
Release on PR close is already correct (dor-reset.yml: stack down, volumes + images pruned,
reservation cleared, edge placeholder restored). One known gap to close first: sk7–sk10 are
registered runners but absent from the hostname→URL map, so a build landing there fails at step 1
(PR #944).
8. Concurrency, budget, and kill switches¶
The human gate was also, accidentally, the rate limiter. Replace it explicitly:
- Concurrency cap — at most N autonomous builds in flight (N ≤ pool size − 1, so a human can always grab a box). Excess queues; queue order by severity then age. Validation sessions count against the pool but jump the queue — see §5.1.
- Budget guard — a weekly quota ceiling; on breach the pipeline queues instead of building and says so on the issue. Bugs share the Max subscription with the spec side, which must never starve: triage and certification are cheap and always run.
DOR_AUTOBUILD— a new repo variable, separate fromDOR_ENABLED, so autonomy can be turned off without turning off triage and certification. Defaultfalseuntil we have watched it work.- Per-issue opt-out — a
no-autobuildlabel any maintainer can apply, honoured at step 4.
9. What the human still does¶
- Approves the merge. The single gate. Already required by the ruleset (1 approval + CODEOWNERS
- required checks), so this is not new machinery — it is the machinery we stop duplicating.
- Says "not yet" cheaply.
/validateto get the thing running under your own hands, or state the gap and let the pipeline close it (§5). Neither costs you manual work, which is the point: an expensive "no" is not a real gate. - Reads Exceptions. The pipeline's job is to be honest about what it could not prove.
- Objects, if the reporter disagrees. The reporter is notified when the PR opens, with the live URL. Objection before merge routes into the existing feedback flow. Their voice is preserved as an opt-out, not as a blocking opt-in.
10. What this does not change¶
- Features keep the value gate. "Is this worth building?" stays a human question. This document is only about bugs — the distinction is the whole argument.
- Merge stays human. Permanently. No auto-merge, no self-approval; "Actions can approve PRs" stays off.
- The security model is untouched. Untrusted issue text still reaches the model only through the sandboxed reason step; the deterministic post step is still the only thing that writes to GitHub; the build still runs on an isolated, reset-between-uses sidekick and pushes with no merge rights.
- External reporters still cannot trigger anything. Org-member gate unchanged.
11. Board and state model¶
This table is the proposed model, not the current one
The columns as they behave today — all 13, on both boards, with the actor who owns each
transition — are in dor-state-machine.md. Nothing below has retired:
Awaiting approval and Awaiting functional acceptance are both still live and still used by
bugs. Paused is missing from the table entirely.
The Bug board (org project #3) already carries every Status option needed. Under this design:
| Status | Under autonomy |
|---|---|
| Ready for AI probe · Awaiting requestor · Awaiting design · Decompose · Blocked (external) · Out of pipeline | unchanged |
| Awaiting approval | retires — rename the column to Queued for build (certified, waiting on a runner) |
| Building | set at claim |
| Awaiting functional acceptance | retires — replaced by the live replay in step 10 |
| Awaiting merge | set when the evidence bundle posts; the human queue. Stays put during a validation session or a proof-gap loop — the manual-validation / reworking label carries the detail, so the column never lies about where the item is |
| Done · Exceptions | unchanged |
~~Prerequisite, and the reason this cannot ship today: the build side is Feature-board-only~~ —
resolved. dor_set_status.sh now resolves the board from the issue's own labels (bug → Bug
Pipeline, everything else → Feature Pipeline), at the source, so no call site needs an override. The
board follows the issue.
12. Build order¶
Each phase is independently useful and independently revertible.
| Phase | Contents | Value on its own |
|---|---|---|
| 0 | Board resolution moved into dor_set_status.sh; sk7–sk10 pool map (#944) |
Unblocks any bug reaching the build side |
| 1 | (built) Red-first sequencing; mandatory live replay (smoke fallback removed for bugs); "no test → Exceptions" | The proof chain — valuable even with the gate still in place |
| 2 | (built) contract.json from the probe + conformance check against it |
Makes certification falsifiable |
| 3 | (built) Evidence bundle on the PR | Makes the merge review a review of proof |
| 4 | (built) The review loop: proof-gap re-runs from either side, red-proved-against-main. (/validate sessions deferred — the box is held until PR close, so a reviewer's env is already live) |
Makes "not yet" cheap — useful on any bot PR, gate or no gate |
| 5 | (built) Flip the gate: DOR_AUTOBUILD, threshold policy, concurrency cap, size ceiling. (Budget guard NOT built — no spend API; the concurrency cap bounds it instead) |
The autonomy itself — last, on top of everything that proves it |
Phases 1–4 are worth shipping regardless of whether we ever flip phase 5. That ordering is deliberate: build the proof, and the way to reject it, before removing the gate they replace.
13. Decisions I need from you¶
- Scope of autonomy — every certified bug, or only those under a blast-radius/size threshold (my recommendation: threshold, with schema/auth/crawler-credential paths always escalating)?
- Diff size ceiling for "this is a bug fix, not a refactor" — I suggest 10 files / 400 changed lines, escalate above.
- Concurrency cap — my recommendation: 3 concurrent autonomous builds against a 6-box pool.
- Severity filter — do cosmetic/low bugs auto-build too, or only
priority:medium and up? - Reporter veto window — merge as soon as CI is green and a maintainer approves, or hold a fixed window (e.g. 24h) for the reporter to object first?
- Validation-session TTL — how long is a box held for you? I suggest 8h idle with a warning at
7h and
/extendto keep it, always released on merge or close. - Review-loop cap — I suggest 3 iterations, then Exceptions.
- Column renames on the Bug board (§11) — these are manual, one-time, and mine to do only if you want them.