fix(worker,shiftengine): a killed run reports its own death, and does not spend the agent budget #39
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "fix/infra-kills-dont-park-the-ticket"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What
Shift 73 parked work item 98 at
needs_humanon 2026-08-13 after three pods were killed mid-run. The agents were working correctly in all three — LiteLLM's ledger records 19, 17 and 1 completed model calls, context growing 18k → 53k tokens, each pod dying within seconds of a call that had just succeeded, for $0.0237 of an $8.00 pool. Nothing about the work was ever tried, and the close reason sent a person to read agent logs.Two defects, one behind the other.
1. A killed worker reported nothing
ploeg-workerinstalled no signal handler at all, so SIGTERM killed it on the default disposition: no outcome, no revoked credential, and the Lease left to the sweeper a full 15 minutes later, attributed to nothing.RunContextnow takes a cancellable parent;cmd/ploeg-workerhands it SIGINT/SIGTERM. The run context is deliberately not derived from that parent — the harness must die with it, but revoke, settle and report have to survive it, or the shutdown reports nothing and we are back where we started.context.WithCancelCausecarries why:errTerminatedmaps tofailed/infra_node,errLeaseLostkeeps the existingstuck/lease_lost.terminationGracePeriodSeconds(new value, default 90) is load-bearing: the 30s default SIGKILLed the pod mid-report and left the handler inert.2. Infrastructure spent the agent's retry budget
store.ExpireLeaseshas always counted infrastructure apart on the pre-Shift path — refunding the attempt, trackinginfra_failures, capping atMaxInfraFailures. The Shift path inherited none of it, and for a Round whose only producer offailedis the sweeper, the ticket's three attempts were being spent entirely on the cluster.FailedRunsInRoundnow returnsInfraAttemptsbesideAttempts, partitioned in SQL bywork.InfraFailureReasons()so query and engine cannot drift.retryFailedWriterbounds them separately and closes withwriting_run_killed_repeatedlywhen infrastructure is at fault — a close reason that points at evictions and node pressure rather than at the ticket.Applied to item 98: three kills would have left all three agent attempts intact.
Also
settleSpendran on the run context, so a cancelled run lost its cost settlement — which is why shift 73 readsspent=0.0000against real gateway spend. It runs detached now.Not fixed here
No backoff between infra retries on this path, where
ExpireLeaseshas 1/5/15/60-minute steps. Gating a reopened Run on a time would changeClaimRoleand the KEDA trigger predicate together, and those must stay byte-identical. Recorded as an accepted cost and a re-evaluation trigger in ADR-0021.Open question for review: the infra budget reuses
MaxInfraFailures(10). Without backoff that is 10 rapid pods on a sick cluster, where it used to stop at 3. A smaller dedicated constant is a two-line change if you'd rather.Records
openspec/specs/shift-orchestration— requirement and scenarios updated. Edited directly rather than via a change directory; say so if you would rather it went through propose to apply to archive.Verification
go build,go vet, full suite and helm goldens green. New tests:TestFailedWriter_InfraKillsDoNotSpendTheAgentBudget,TestFailedWriter_ParksAtTheInfraCapAndSaysSo,TestAbortOnTermination_*, andTestInfraFailureReasons_MatchesIsInfra(fails if the SQL list and the Go predicate ever drift).🤖 Generated with Claude Code