fix(shiftengine,store): a failed writing Run re-opens its Round #36
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "ryangr0/failed-writer-advances"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
shiftengine.evaluatefroze the plan on astuckOutcome and had no case forfailed. A Run the sweeper reclaimed is, as far asRoundCompleteis concerned, simply finished — so the Round completed and the next one opened over work that was never done.The reviewer reviewed a branch that had never been written, approved it, and the Shift closed with the most reassuring reason there is, having produced nothing. The tracker comment told the truth — "Ploeg stopped working this item without opening a pull request" — and the two disagreed.
The distinction the spec was missing
The shift-orchestration spec says "a swept Run does not block its Round forever … the Round can complete and the Shift advances." That was reasoned about readers, and for a reader it is right: a dead reader costs an opinion, and stalling an item over a missing opinion is worse. For a writer the same rule means every later Round reasons about a branch that does not exist.
failedis the sweeper's verdict on a pod that stopped renewing — never an agent's report. That is what makes it retryable, and what separates it fromstuck(R4).The fix
A failed writing Role re-opens the round the Shift is already on.
In place, because
shifts.rounddoubles as the index into the plan (tp.Rounds[si.Round]). Opening a fresh Round to retry would consume the slot belonging to the next planned Round and silently skip it — and that bug, a reviewer that never runs, is harder to see than the one being fixed.store.ReopenRoundinserts a pending Run at the current round number without incrementing it.RoundCompletekeys on the Shift's current round, so the Round correctly becomes incomplete again.(shift, round, role), never stored — the discipline ADR-0012 sets forreserved.store.MaxRunAttemptsthe Shift closes atneeds_humanwithclose_reason = writing_run_failed_repeatedly.Planned Shifts only. Under uniform dispatch a synthesized one-writer Shift already settles by its run's own Outcome, and
failedmaps toqueued— the attempt-capped requeue R5 requires. There is no later Round there to step over, which is whyTestUniform_FailedRequeuesAndRespectsTheAttemptCappasses unchanged.MaxRunAttemptsis deliberately separate fromMaxAttempts:work_items.attemptsincrements per role claim, so it stopped meaning "attempts at this work" the day Shifts landed, and reusing it would let a three-role plan exhaust its budget in one clean pass.Verification
Unit —
go test ./pkg/shiftengine/, all three proven to fail against the unfixed engine first:End to end —
ploeg-bench'shangscenario SIGKILLs the worker mid-run so the lease genuinely lapses (sleeping does not: the worker renews on its own goroutine while the agent works). Against this build:All seven bench scenarios green, including
fix-cap,budget-floorandreader-push.Also here
architecture.md§4 and §9.19 corrected; backlog #116 closed.Numbered 0019 because #35 already claims 0018.
🤖 Generated with Claude Code
docs: a failed writing Run does not stop the planto fix(shiftengine,store): a failed writing Run re-opens its Round11963fb269107d5f760f