chore(release): promote development to main #234
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "ryangr0/promote-development-site-v0.1.0"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Promotes
developmenttomainso the site gets its first stable release onunfoldhq.dev(ADR-0012, 2026-10-04 entry).mainwas last updated on 2026-09-12, so this carries 1338 commits.What happens on merge:
on_source_change.ymlrunschecksandploeg-pinonmain.site-releasecutsunfold-site-v0.1.0. Unfold's ownreleasejob is skipped onmain, so Unfold does not release.on_release_publishedsees a stable site tag and deploys theunfold-siteWorker onhttps://unfoldhq.dev, with live checks on/,/nl,/robots.txt, the sitemap,/demo/and/privacy.Already in place:
AAAA 100::, applied by the homelab reconciler.unfold-site-signups; staging has its own database since #233.Merge with a merge commit, not squash. After the release, merge
mainback intodevelopmentso the next promotion does not conflict onapps/site/CHANGELOG.md.Merge only when: the checks on this PR are green.
🤖 Generated with Claude Code
- .releaserc.js: conventionalcommits, tagFormat v${version} (Go modules), CHANGELOG commit-back [skip ci], Forgejo-only publish via @saithodev/semantic-release-gitea - ops/docker/ploegd/Dockerfile: cross-compile -> distroless static nonroot, -X main.version - on_source_change.yml: checks -> semantic-release -> explicit dispatch of on_release_published (Forgejo emits no release event for CI-cut releases) - on_release_published.yml: tag parse -> Harbor + GHCR multi-arch publish - on_pull_request.yml replaces ci.yaml for PR checks Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>The worker now mints a per-run key against the LiteLLM admin API before starting the agent subprocess and revokes it via a deferred call that runs on EVERY return path (success, failure, context cancel). The minted key is passed as LLM_API_KEY to the entrypoint, which skips its own mint/revoke when receiving it (agent-runner >= 1.0.1). Key fixes: - Revoke() sends {"keys": [key]} (LiteLLM /key/delete schema), not bare "key" - modelList() strips litellm_proxy/ and openai/ prefixes from model names - Fake LiteLLM servers in tests validate real request schemas (reject bare "key" field with 422; reject model scopes containing "/") - Regression tests: agent failure revokes, agent success revokes, mint failure does NOT revoke, key alias format preserved, model prefixes stripped VIK-585 Agent-Trace-Id: ploeg-b880357c84b3resolveOutcome let the forge PR poll win before runErr was examined, and findPR matches on agent/vik-<externalID> — the branch every retry, review round and (soon) persona turn reuses. A run that crashed instantly found its predecessor's still-open PR and reported pr_opened, marking the work item done on work it never did. Today that mis-closes re-assigned tickets (architecture.md §1's review loop); with a reviewer persona it would ship unreviewed code while the audit log credits the reviewer. execute() now snapshots the branch's PR state before the harness starts and passes prExisted into resolveOutcome: - pr_opened only for a PR that did not exist before this run - clean run on a pre-existing PR -> pr_updated (already in the enum) - failed run keeps stuck/agent_error, carrying the PR in Links for the reviewer's convenience but never as this run's success Also guards the VIK-586 heuristic, which keyed "LLM adapter produced no spend" on CostUSD == 0 alone. That (a) overwrote a FailureReason an adapter had already set, and (b) misreads any harness reporting tokens without a cost — ACP makes UsageUpdate.cost optional, so that is the common case for the adapters coming next. It now requires an unset FailureReason and zero token counts: no evidence of LLM traffic, rather than no evidence of billing. Finally, failureReason joins outcomereport.v1.schema.json. The worker has posted it and the store has persisted it since VIK-597, but the published schema is additionalProperties:false with no such property — the contract was behind the code. Additive-optional, legal under the v1 rules. Gates (docker golang:1.25 + alpine/helm, this branch): gofmt -l . .................... clean go vet ./... .................. ok go build ./... ................ ok go test ./... ................. ok (pkg/store needs a non-root uid; green when run unprivileged) helm lint ops/helm/ploeg ...... 1 chart linted, 0 failed helm template x3 .............. ok (default, keda, cronjob) The six new table rows were run against the unfixed resolveOutcome and all six fail there: crashed_run_does_NOT_inherit_a_pre-existing_PR pr_opened, want stuck clean_run_on_a_pre-existing_PR pr_opened, want pr_updated lost_lease_on_a_pre-existing_PR pr_opened, want stuck tokens_without_cost_is_NOT_infra_llm infra_llm, want "" adapter-set_failure_reason_survives infra_llm, want agent_error exit_0_with_tokens_but_no_cost failed, want no_change_neededDecisions lived in three shapes at once — verdict rows in design.md §8/§9, dossiers in docs/research/, and *[research]*-tagged backlog items. No single place answered "what is decided and is it still in force", and nothing stopped two of them disagreeing. homelab-cluster ADR-0050 exists because exactly that happened there for two days. docs/adrs/ is now the single ledger, MADR 4.0 with two local extensions: - Supersession is append-only. An accepted record's own status is never flipped; the superseding record carries `supersedes: NNNN` and the index shows "superseded by NNNN". The file is the artefact, the index is the current view. - A decision that can change carries `review-by: YYYY-MM-DD` plus a "Re-evaluation triggers" section. This is the one thing §8 did that vanilla MADR has no slot for, and migrating without it would have been a downgrade. Nine records. The seven short market-survey rows collapse into 0005 with their individual verdicts preserved as Pros and Cons — that is what MADR's Considered Options section is for, and it reads better than seven stubs. The four substantial protocol/product verdicts (AHP, A2A, OmniRoute, Paperclip) each keep their own record, triggers, and dossier link. §9's foundation decisions (Go, Apache-2.0, Forgejo-leading home) are migrated too, though the plan only called for §8: leaving them behind would have recreated the dual-ledger problem this commit exists to remove. Evidence does not move. docs/research/*.md stays exactly as it is; each ADR links its dossier under More Information. The ADR carries the verdict, the dossier carries the working. scripts/check_adr_consistency.py is ported from erfbeeld and extended with two checks it asks for in prose but never enforced: every record must carry a "### Confirmation" subsection, and a dated review-by must be backed by a triggers section. Note the adr-writer skill's bundled validator must NOT be used here — it assumes status-flip supersession and rejects this corpus. Verified positively and negatively: 9 record(s), 9 Records row(s), 0 supersession(s), 5 dated review(s) — consistent flipped file status ............ caught (2 violations) dated review, no triggers ...... caught missing Confirmation ........... caught index date drift ............... caught record with no Records row ..... caught supersedes a missing number .... caught CI wiring, the OpenSpec half, and the design.md/AGENTS.md pointer rewrite follow on this branch.Changes now run proposal → specs → design → adr → tasks, with the adr step gating tasks so a change that makes a durable architectural commitment cannot reach implementation without recording it (ADR 0001). The schema is derived from the UPSTREAM `spec-driven` shipped with @fission-ai/openspec 1.6.0, not from erfbeeld's fork of it — erfbeeld's adr instruction carries its own house rules (X1 Dutch/English, H1 fiscal amounts, conformance vectors) that mean nothing here. Two kinds of change on top of upstream: the adr artifact itself, and repo rules folded into each instruction (seam vocabulary, append-only migrations, never-restate-a- published-contract, the gate set, regression-test-must-fail-first). openspec/config.yaml deliberately POINTS at docs/architecture.md, docs/design.md, docs/adrs/, docs/contracts/, docs/domain/ and docs/backlog.md rather than copying them. The contracts are pinned to the Go types by pkg/harness/contract_test.go and the domain docs are generated from model.yaml — a prose mirror in config.yaml would be drift with no gate on it. Keep it a map, not a mirror. The ~3,100 lines of skills and slash commands are NOT committed: they are pure `openspec update` output, identical across repos, and already present user-globally at ~/.claude/commands/opsx/. Run `openspec update` locally to generate them for whichever tools you use. Verified against the real CLI (node 24 + openspec 1.6.0): openspec schemas spec-driven-with-adr (project) Artifacts: proposal → specs → design → adr → tasks openspec doctor OpenSpec root: ok openspec validate --all No items found to validate (no changes yet) openspec status [-] adr (blocked by: design) [-] tasks (blocked by: adr, specs) openspec instructions adr --change <scratch> renders <project_context> and <rules> from config.yaml That last one is the integration point worth knowing about: `openspec instructions` is how config.yaml actually reaches a model, so it is the way to check any edit to the schema, the config or a template before relying on it. Two YAML traps hit while writing config.yaml, both from plain scalars: a colon-space inside a list item ("a requirement: what happens") parses as an implicit key, and a value starting with a backtick is a reserved character. Neither is caught by eye — `openspec doctor` reports them.Closes the migration: design.md §8/§9 become indexes into docs/adrs/, AGENTS.md points at the new homes, and the consistency gate moves from a Python script into the existing `go test ./...` step. The gate moved languages on purpose. The Forgejo runner is guaranteed a Go toolchain and is NOT guaranteed python3, and on_pull_request.yml already carries two comments about network-fragile setup actions — adding a third to run a stdlib script was the wrong trade. internal/ledger reads the corpus as data, needs no new CI step, and runs locally for anyone who can already build the repo. One implementation, not two in different languages. `openspec validate --all` is deliberately NOT wired into CI yet. It needs node + the openspec CLI in the runner, and today it validates zero changes ("No items found to validate"). Wire it with the first real OpenSpec change, when it has something to check. AGENTS.md now states the three homes explicitly — evidence in docs/research/, verdict in docs/adrs/, action in docs/backlog.md — plus one line the plan asked for: .openhands/ and .opencode/ in this repo are dogfooding, not product spec. Ploeg's harness support is pkg/harness and docs/contracts/, never whichever agent config sits in this tree. Gates (docker golang:1.25 + alpine/helm): gofmt -l . .................... clean go vet ./... .................. ok go build ./... ................ ok go test ./... ................. ok (internal/ledger 0.029s; pkg/store green when run with a mapped unprivileged uid) helm lint ..................... 1 chart(s) linted, 0 failed helm template x3 .............. ok internal/ledger was run against seven deliberately broken corpora; every check fails when it should: flipped file status ................ illegal status + index mismatch dated review-by, no triggers ....... caught missing ### Confirmation ........... caught file date drifts from index ........ caught record absent from the index ....... caught supersedes a nonexistent number .... caught real supersession, stale index ..... caught <- the append-only mechanismThe process half of the ACP adapter. Three failure modes it exists to prevent, each covered by a test that re-execs the test binary as a fake agent (the os/exec idiom) — a real child process, no `go build`, no network, no agent binary. 1. Orphaned grandchildren. Deliberately NOT exec.CommandContext: node-based ACP agents (opencode, the npm adapter processes) fork workers, and killing only the direct child leaves a grandchild holding the DinD socket and the per-run LiteLLM key until its TTL expires. Setpgid plus signalling the negative pgid reaps the group. The test forks a real grandchild and asserts it dies. 2. Banners on the protocol channel. Agents print version notices, spinners and stray console.log on stdout. Without a filter the first such line is a JSON-RPC parse error and the session dies for a cosmetic reason. Lines whose first non-space byte is '{' go to the dispatcher; everything else is diverted to the pod log. So a wrong subcommand (`opencode` instead of `opencode acp`) surfaces as "initialize never completed, and here is the text it printed instead" — diagnosable infra, not a mystery. A JSON ARRAY is noise too: JSON-RPC frames are objects. 3. Stdio deadlock. Every client-side response goes through a 64-slot async writer, so a child that stops reading stdin stalls one goroutine instead of the dispatcher. A full queue returns errStdinBacklog rather than blocking or silently dropping — dropping would desync the protocol, and the watchdog can escalate to SIGTERM on the error. Also: Wait is sync.Once-guarded (a second Wait is "wait: no child processes" without it), and Kill is idempotent. One test bug found and fixed while writing this, worth recording because it would have been a flake rather than a failure: the JSONL classification test asserted on the noise buffer after reading a fixed line COUNT, so the filter goroutine had not necessarily classified the trailing line yet. Draining to EOF is the only thing that proves it has. gofmt clean, go vet ok go test ./pkg/harness/adapters/acp/ -count=1 ok 0.341s go test ./pkg/harness/adapters/acp/ -race ok 3.614ssession/request_permission exists so an editor can ask a human. There is no human in a worker pod, so the policy must answer every request immediately and deterministically — it runs on the protocol read loop and must never block or prompt. Default is allow-all, matching the claudecode adapter's bypassPermissions and for the reason recorded there: the pod is a disposable, credential-scoped sandbox whose blast radius is bounded by the per-run LiteLLM key, the repo-scoped forge token, and being destroyed at exit. The permission prompt is not the security boundary; the pod is. What the policy DOES protect against is a runaway agent. An agent asking the same thing 200 times is not progressing, and the storm cap (200 total, 60/min, sliding window) turns a silent 45-minute burn into a stuck reason naming the worst offenders. allow_read_only exists for the reviewer persona in Phase 3: a judge that cannot edit is a stronger guarantee than a judge instructed not to. It maps read/search/fetch/think to allow and everything else to reject. Two traps the tests pin, both of which would silently grant a mutation: - never fall back to "the first option" when nothing matches. A deny decision with only allow-shaped options answers cancelled instead. - "Disallow this tool" contains the substring "allow". The heuristic tier excludes it explicitly. Also: allow_once is preferred over allow_always so a grant never widens beyond the call that asked for it, and an unrecognised mode string is reported as invalid rather than defaulting to allow-all — the caller fails startup instead of running wide open by accident. The clock is injectable, so the sliding rate window is tested without sleeping: three requests in one minute storm, and the same asker two minutes later does not. gofmt clean, go vet ok go test ./pkg/harness/adapters/acp/ -run TestPermission 8/8 PASS go test ./... — all green except pkg/store, which needs a non-root uid (embedded-postgres; unrelated to this change)Completes the working adapter: a phase machine (launch → initialize → session/new → session/prompt → shutdown) that records HOW FAR it got, because a failure before the prompt is an infrastructure problem with the pod rather than a problem with the ticket. That is architecture.md §9.9 / VIK-596 fixed at the source — and it needs no orchestrator change, since resolveOutcome already prefers a valid structured report. The dependency was audited before being taken, not after. `go list -deps` shows the SDK compiles in with ZERO external packages — all 13 modules in its graph are its own test dependencies. Repo goes from 3 direct + 7 indirect to 4 direct, with nothing new linked into the binary. TWO REAL DEFECTS FOUND, both by tests rather than by reading: 1. The SDK drops usage silently, and my first defence was in the wrong place. coder/acp-go-sdk v0.13.5 is generated from schema 0.13.5, whose session/update union has no ACP v1 usage variant ({used,size,cost}) — so an agent sending v1 usage had it discarded by the SDK's decoder before ploeg's tolerant decoder ever ran. A decoder placed AFTER the SDK inherits whatever the SDK failed to parse. State is now folded from the RAW protocol line via a tap in the JSONL filter, BEFORE the SDK sees it; client.SessionUpdate is a deliberate no-op. This is the single most important line of defence in the adapter and it only held once the ordering was right. 2. conn.SetLogger races with the SDK's own read loop. NewClientSideConnection starts that loop before returning, so setting the logger afterwards is a data race — the detector flags it every run. We drop the call; the agent's output already reaches the pod log through the launcher's noise channel. Worth reporting upstream. Also fixed: harness.TailBuffer was not concurrency-safe. RunCommand never noticed because it reads the tail only after cmd.Run() returns, but a session adapter pumps stderr on its own goroutine and reads the tail on early-return paths. Fixed at the shared type (mutex + Bytes returns a copy) rather than locally, so no future adapter has to remember. And a `cp :=` that shadowed the outer checkpointer would have leaked its emitter goroutine — caught while fixing the call sites. Client half is deliberately two methods. fs/* is refused: pkg/worker embeds AGENT_BUILDER_TOKEN in the clone URL so it lands in .git/config, and an fs/read of that file would hand the forge token to the model provider through a protocol-blessed path. terminal/* is refused: tool_call updates already carry the command, its status and its output at the same fidelity. elicitation/* is refused permanently — an agent needing a human IS the stuck state. Profiles: opencode (flagship, per homelab-cluster ADR-0051) and custom. The opencode provider-config key names are the one thing needing verification against a real binary, so PLOEG_ACP_CONFIG_JSON overrides the whole document without an image rebuild. Config is written 0600 and trace-scoped, because ScratchDir is os.TempDir() and shared across concurrent runs. The conformance kernel now drives this adapter with plain shell scripts that do NOT speak ACP — the likeliest real misconfiguration — and every property holds. One kernel assumption had to be corrected: "exit 0 means success" is spawn-and-wait-shaped; for a session adapter a binary that exits 0 without a handshake never spoke the protocol, so an error there is correct. The property now asserts what both shapes actually owe: never fabricate an outcome. gofmt clean · go vet ok · go build ok go test ./pkg/harness/... -race all ok (acp 7.178s) go test ./pkg/store/... ./internal/... ok (unprivileged uid) helm lint + 3 templates okWires the finished adapter into the three places a harness is chosen: - worker.NewAdapter gains `case "acp"`. Profile resolution and permission-mode parsing both happen there, so a misconfigured team fails at worker startup rather than after leasing a ticket it cannot work — the same fail-before-Claim property the other adapters have. ACP is a session adapter, so it implements harness.Adapter directly and takes no RunCommand lift. - cmd/ploeg-worker parses PLOEG_ACP_{PROFILE,ARGV,PERMISSION_MODE, PROMPT_TIMEOUT,IDLE_TIMEOUT,CONFIG_JSON}. The two durations are strict: a typo in a watchdog timeout that silently reverts to its default fires at the wrong moment and reads as an agent bug rather than a config error. - The chart grows harness.acp, overridden per team by the existing field-by-field pattern (explicit hasKey, not sprig merge). A custom profile without argv fails at template time, not at run time. New adapters_test.go pins the registry's one safety property: every misconfiguration is an error at construction. Without it a bad harness value surfaces as a pod that dies holding a lease — indistinguishable from an agent crash, and retried MaxAttempts times before anyone notices. ci/executor-values.yaml gains a third team so helm template exercises the ACP env branch and the argv guard on every PR. Gates: gofmt clean, go vet ok, go build ok, go test ./... ok (store under an unprivileged uid), helm lint 1/0 failed, all three renderings ok, both chart guards negative-tested. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Runs the domain-language skill's generator over model.yaml: python scripts/generate_docs.py docs/domain/model.yaml -o docs/domain/ overview.md gains Shift in the ER diagram and both new open ambiguities at the top; glossary.md gains Shift and Round with the narrowed Lease; entities.md gains the Shift attribute table and Lease's new shift_id/run_id shape; events.md gains ShiftOpened, RoundStarted, BudgetAuthorized and BudgetSettled. rules.md is unchanged — no rule was edited. Corrects the previous commit message, which said no generator exists in this repo. True but beside the point: the generator is the domain-language skill, and the views were left stale for one commit describing a Lease as unique per Work Item, which ADR-0010 had already changed. Health check: 9.8/10. The only deduction is four open ambiguities, two of which this branch added on purpose — the `leased` state name and the Shift close rule. Per the skill's own guidance, honest open questions beat a higher score bought by deleting them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Found by running the local stack rather than by reading the code: POST /api/v1/runs/<token>/outcome {"outcome":"failed","failureReason":"vibes"} -> HTTP 204 and `vibes` was stored verbatim in agent_runs.failure_reason. This is a gap in my own Phase 0 change. It added failureReason to docs/contracts/outcomereport.v1.schema.json as a five-value enum and pinned the schema to the Go types in pkg/harness/contract_test.go — but nothing evaluates that schema at request time. handleOutcome checked Outcome.Valid() and the R4 stuck-reason rule, then passed failureReason straight through. It matters beyond untidiness: pkg/worker's classification DEFERS to an adapter-set failureReason, so `infra-llm` would not error — it would silently override the orchestrator's own judgement about whether a failure is retryable. pkg/harness/harnesstest holds adapters to this rule, but adapters are not the only callers; ops/local/demo.sh posts with curl, and so will anything implementing docs/contracts/executor.md. Validation is extracted to a pure validateOutcomeReport so the rules are testable without a database — pkg/httpapi had no test file at all. Two tests cover the closed enums, R4, and that every published taxonomy value survives the boundary (a conformant adapter must never be rejected for using a documented value). Adds ops/local/probe.sh, the negative half of demo.sh: demo.sh shows the happy path works, probe.sh shows the guards hold. It exercises R4, the closed taxonomy and R2 crash-safety over the real wire, and exits non-zero so it can gate CI later. The general lesson, worth watching for elsewhere: a published schema that nothing evaluates at runtime is documentation, not a gate. Verified: probe.sh green from a clean database (lease reclaimed after ~30s, a swept run's report refused 404); gofmt clean, go vet ok, go test ./... ok across all 16 packages. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>The foundation for several personas working one Work Item, some of them at once. Implements ADR-0010 (Shift owns the item, Lease owns the branch) and ADR-0012 (two-level budgets, authorized and settled). The load-bearing choice, which improved on the ADRs while implementing them: a Round MATERIALISES its Runs. Opening a Round inserts one pending agent_runs row per Role; claiming flips a row to running. That buys three things at once: * The claim predicate and the KEDA scaler query become the SAME statement — pending runs for (team, role), oldest first. ADR-0010 flagged the mirrored scaler query as the design's highest-risk drift point, where undershoot stalls items with no error anywhere, forever. Materialising the round dissolves the hazard rather than guarding it. * Reserved budget is SUM(authorized) over running rows, so it cannot drift from what is actually in flight the way a counter can. ADR-0012 specified a `reserved` column; a derived sum is strictly better and one column less. * The roster is explicit and queryable instead of reconstructed. Concurrency safety is the existing discipline, not a new one: ClaimRole locks the Shift row before the pool arithmetic, so five readers starting together cannot each see the full pool. The unique partial index shifts_one_live_per_item means two Teams racing to open a Shift is a database error, not a race one of them silently wins. Readers take no Lease — that is what lets a fan-out run at once — and OpenRound refuses a Round that mixes a writer with readers, or carries two writers, rather than trusting callers with the one rule the whole concurrency model rests on. Tests are the acceptance criteria the ADRs named for themselves, plus two the implementation turned up: a zero budget means unmetered (every team today), and a refused claim must not consume its slot. One migration fix while writing them: agent_runs.started_at has been NOT NULL since 0001, correct when a row only existed once a pod was running. Pending runs predate their pod, so the NOT NULL is dropped; the DEFAULT stays, leaving the pre-Shift Claim path untouched. Not yet wired: nothing opens a Shift or a Round, settlement does not run, and the sweeper does not release holds. Store-level only. Gates: gofmt clean, go vet ok, go build ok, 9 new store tests pass, all 8 pre-existing store tests still pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Two records claimed things the code disproved. Correcting the ledger rather than letting it drift is the whole point of ADR-0001, and the drift was one commit old. ADR-0012 specified `UPDATE shifts SET reserved = ...`. The implementation derives reserved as a SUM over running Runs instead — one column fewer, and it cannot disagree with what is actually in flight. More importantly it removes the release statement entirely: a Run that stops running stops holding money, so a missed release became impossible rather than unlikely. The record now describes what was built and says plainly that it replaces the counter an earlier draft specified. ADR-0010 gains a "What the implementation changed about this record" section covering the two defects only contact with the code exposed: * the advance-once CAS could not stay on the Lease. ReportOutcome opened with DELETE FROM leases ... RETURNING, and readers hold no Lease, so every reader's report would have failed with ErrUnknownRun — the population the record exists to enable could never have reported an outcome; * liveness could not stay there either, for the same reason: an OOM-killed reader has no lease to expire and would hold budget forever. Both Confirmation sections now name the tests that actually exist, and 0012 states which of its promises is still owed by the orchestration change (the needs_human transition on an exhausted pool) rather than implying it is green. architecture.md gains §10, the complete Shift picture, with five new Mermaid diagrams: how the Lease's three jobs split by lifetime, the Round state machine, the full loop as a sequence (readers concurrently, one writer, review round, human pulled in to merge), the money flow, and the data model. §10.6 is a per-component build-status table — the store layer is built and tested, everything that would drive it is not, and the table says so component by component rather than leaving a reader to infer it. §§1-9 still describe only what is deployed today. model.yaml: Lease narrowed to exclusion with liveness moved onto the Run, and the Run entity gains round, writes, state, authorized and expires_at. Domain views regenerated; health 9.8/10, the deduction being four deliberately open ambiguities. Gates: gofmt clean, go vet ok, go build ok, go test ./... ok across 13 packages (13 Shift tests among them), helm lint 1/0 failed, both chart renderings ok, ADR ledger validator green. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>The Shift layer (0008) shipped with no way to close a Shift, no way to find one, and three latent defects its first callers would have hit. This change completes the store surface the shift engine needs, all dead code until the engine lands: - CloseShift: records why, cancels leftover pending runs so the scale signal drops to zero, idempotent (fast-path and sweeper may both close), and releases the shifts_one_live_per_item slot for a later re-mandate. - LiveShifts / LiveShiftForItem: the sweeper's worklist and the idempotency read EnsureShift needs. - OpenRound gains a compare-and-swap fromRound guard: two evaluators can no longer double-advance a round and materialise a duplicate roster. - ReportOutcome returns OutcomeResult{WorkItemID, ShiftID}, persists findings (migration 0009, ADR-0011), and for shift runs leaves the work_items transition to the engine - three readers reporting must not flip the item's state three times. Legacy shift-less runs keep today's behaviour bit for bit. - RoundReports: the blackboard read serving both the PR comment and the next round's prompt. ShiftsBelowFloor: the parking worklist for exhausted pools. Fixed, each with a regression test proven to fail against the unfixed code: - ClaimRole's RETURNING omitted the target_* columns, so a role-claimed run silently lost the resolved Work Target (ADR-0014) and fell back to the env repo. - Renew only touched leases: a reader (no Lease, ADR-0010) got ErrUnknownRun on its first renew and cancelled itself at TTL/3; a writer outliving one TTL kept a live Lease while ExpireRuns reclaimed its run underneath it. - Checkpoint resolved the item via leases, so readers could not checkpoint. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>TestLauncher_KillReapsTheWholeProcessGroup has been failing the pipeline, and the launcher was never at fault. Diagnosed by dumping the process table at the point of failure: 99 (sh) Z 1 93 ... PID PPID PGID STAT COMMAND 1 0 1 Ssl go 99 1 93 Z sh State Z means the group signal DID arrive — the grandchild is a corpse, not a survivor. It lingers because killing the group also killed its parent, so it was reparented to PID 1, and PID 1 reaps orphans only if PID 1 is an init. Under `docker run golang go test` (how AGENTS.md tells contributors without a local toolchain to run the gates) and on the CI runner, PID 1 is the go driver, which reaps nothing. kill(pid, 0) keeps succeeding for a zombie, so the probe reported the corpse as alive and the assertion tested the container's init rather than the launcher. alive() now discounts zombies, reading Linux's process state from /proc and locating the state character from the LAST ')' — the comm field is parenthesised and may itself contain spaces and parens. On darwin, where PID 1 is launchd and does reap, there is nothing to correct for and the absence of /proc reports false. Verified in both directions: green as written, and still red with the group-kill deliberately removed from execLauncher, so the property it protects — a grandchild does not outlive the run holding the DinD socket and the per-run LiteLLM key — is still enforced. No production behaviour changes: the launcher is untouched, and a worker pod runs one run and exits. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Both executors now range over (team, Role) rather than team. A team with no plan yields exactly one Role-less workload whose rendered manifest is unchanged; a planned team yields one per distinct Role, each with its own image, harness, model, dind setting and replica cap. That is what makes "different agents, different harnesses, different images" literally true — pod shape is fixed at render time, so it cannot be a runtime decision. Role workloads scale on the role predicate, quoted beside the query it must match byte-for-byte: pending runs for (team, role), served index-only by the agent_runs_claimable index migration 0008 created verbatim for it. Overshoot wastes a pod that exits 0; undershoot stalls Work Items with no error anywhere, so this third copy of the predicate carries its leash in a comment. SECURITY — the PolicyException stops tracking workload names. The privileged DinD sidecar needs a Kyverno waiver, and the exception in webgrip/homelab-cluster matched `scaledjob.keda.sh/name` with one entry per team. Under role partitioning that meant a security-repo change for every new writer Role, which is how waivers rot. Two things were wrong with that shape, and both are fixed by keying the exception to the HAZARD instead: * exception-governance forbids wildcards in resources.names[] — it says nothing about label selectors, and this exception never used names. The per-team convention was self-imposed, not required. * the Pod-level match was `app.kubernetes.io/name: ploeg-worker`, the label EVERY worker pod carries. Copper (dind: false) is waived today for privileged-containers, run-as-non-root and drop-all-capabilities while running no privileged container at all, and every reader would have inherited the same over-grant. The chart now emits `ploeg.webgrip.dev/privileged-dind: "true"` only where it actually renders the privileged sidecar. The waiver follows the privilege: adding a Role costs no security-repo change, and readers — which run dind: false — fall outside it entirely. Strictly narrower than today. ADR-0013 tier 1 lands as far as this repo can take it: a reading Role draws AGENT_BUILDER_TOKEN from executor.forgejo.readTokenSecret when configured, so the reader/writer split is enforced by the forge and not only by scheduling. Unset is documented as a known gap rather than a safe default — the repos are private, so "no credential" cannot clone, and closing it is one OpenBao entry plus one ExternalSecret, no chart or code change. Also: LITELLM_KEY_BUDGET degrades to the Role's own cap rather than the team's budget, because for a planned team that value is the SHIFT POOL and handing one Run the whole pool would be wrong if the fallback ever applied. Guarded by committed golden renders (scripts/helm-golden.sh, wired into CI). The pod is a security boundary and the waiver is keyed on a label this chart emits, so a manifest change now has to appear in the diff. Verified in both directions: green as committed, and red when the privileged-dind label is widened to pods that take no privilege. Two render-time guards added with tests of their own: a Role defined twice with different settings fails the render (one Role is one workload, so its shape cannot change between rounds), and a workload name over 63 characters fails rather than being truncated into a collision. CI fixture gains a planned team covering the whole new branch — reader/writer credential split, the dind-less reader pods, and a Role recurring across rounds collapsing to one workload. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>provider.ForgeProvider has been declared since the SPI was carved with zero implementations. This gives it its first one and its first caller in the same change, because the two are the same story: the blackboard (ADR-0011) is a reading Run's findings travelling to where a human is already looking. pkg/provider/forgejo does two things and no more. Comment posts to a pull request through the ISSUES endpoint — Forgejo models a PR as an issue with a branch attached, and /pulls/{n}/comments would be a review comment on a diff hunk, which a round's findings are not. ParseWebhook verifies the raw body against X-Forgejo-Signature BEFORE any JSON parsing and normalizes review_submitted / check_failed / merge_state_dirty, so the Follow-Ups of R9 need no further provider work. Events Ploeg does not act on are dropped without error: a forge subscribes wider than the core consumes, and erroring on every unrelated push is how a webhook ends up disabled. The Vikunja provider's stubs become real. FetchItem gives the thin-payload rule its authoritative half; Comment creates with PUT, not POST (the trap recorded in docs/ops/board.md — a POST there silently does something else); SetStatus writes only `done`, because needs_human and stale are NOT done and marking them so would hide the item from the very board that has to act on it. Inventing a label mapping for needs_human would put a Ploeg concept inside the provider (R7); what a human needs — why it stopped and which PR to look at — travels in the comment instead. Both are opt-in by credential. Without a URL and token they keep the prototype's logging no-op, so a deployment that has not been given tracker credentials still finishes runs; it just does not update the board. The engine now publishes each reading Run's findings when its Round completes rather than at close, so a human watching the thread sees the review while the writer is still working from it, and writes back to the tracker when the Shift closes — the PR link plus a request to merge, which is what turns the handoff from something noticed into something announced. Publication is best-effort, everywhere and deliberately. A forge outage, a tracker outage, an unresolved Work Target or a Shift with no pull request yet must not stall the pipeline or lose an Outcome: every failure is logged and none is returned into the lifecycle, and the tracker write-back happens AFTER the state is durable so an outage can never leave a Shift open. Tests drive both outages through a full plan and assert the Shift still closes and the item still reaches needs_human. Round-1 readers routinely run before any PR exists; their findings are not lost, they reach the writer through the briefing on its claim. Duplicate comments are possible and accepted: two evaluators can both observe a completed Round before one wins the advance CAS. A duplicated comment is visible and harmless; a missing one loses a review a human is waiting for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Implements ADR-0017. A reading Run may return approve or request_changes in its OutcomeReport; a request re-opens the plan's OWN last writing Round with the findings attached, then the review Round after it. Each pair is one fix round. Before this, a reviewer that found a real defect had nowhere to send it — the plan's next entry opened regardless, and if the reviewer was last the Shift closed with the defect recorded and unfixed. MERGING THIS RATIFIES ADR-0017, which is still `proposed`. It is the first place an agent's output influences what runs next, so the boundary is worth reading before it lands: a verdict names no Role, authors no Round, raises no cap and extends no budget. It is one bit that may re-run work the operator already configured. The bounds are checked in the record's order, and each closes with its own reason so "why did this item stop" stays a query rather than a reconstruction: 1. the pool — money is the limit that cannot be argued with, so no Run is ever spawned that the Shift cannot pay for; 2. maxFixRounds — the cap, default off, configured per Team; 3. the verdict — the only bound an agent influences, and checked last. A writer's verdict is ignored twice over: the store blanks it on write (CASE WHEN writes THEN ''), and the loop skips writing Runs when it looks. A writer approving its own work would be the loop grading itself, and one guard could be refactored away without the other noticing. The fix-round count is derived from the Shift's round counter against the plan's length, never stored — the same discipline ADR-0012 applies to `reserved`. It cannot drift from what happened, and it survives a restart mid-loop for free, which a test pins directly. pkg/plan refuses maxFixRounds > 0 on a plan with no writing Round at boot, rather than at the moment a reviewer first asks for changes: a plan that cannot fix anything would otherwise look healthy for hours and then quietly ignore its first real verdict. The reviewer prompt now asks for the verdict and says what each value does — including that request_changes sends work back to the writer, so it is for things that must change rather than for thoroughness. Migration 0010 adds the column with a CHECK constraining it to the two values plus empty; the schema enum, the Go type and the boundary validator change together. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>POST /webhooks/forge/{provider} gives ForgeProvider.ParseWebhook the caller it has never had. The signature is verified against the raw body before anything is parsed (backlog #2), the handler does no expensive work — Forgejo's DELIVER_TIMEOUT is 5 seconds and a slow endpoint becomes a disabled one (backlog #3) — and every accepted event lands in the audit log. It acts on nothing, deliberately. Routing a submitted review into a re-mandate needs the branch-to-Work-Item lookup backlog #107 owes it, and there are now TWO paths that mean "keep going" — an agent's verdict (ADR-0017) and a human's review. Reconciling them is a decision, not a merge order, and ADR-0017 names the arrival of this route as the trigger to make it. What lands now is the endpoint, so the events are recorded from the day the network path opens rather than from the day somebody notices it was never wired. Dedup is a table, not a cache (migration 0011). A forge retries what it thinks failed, and a retry that acts twice turns one review into two fix rounds; an in-memory set would forget across exactly the restart a redelivery is most likely to follow. The insert IS the check — ON CONFLICT DO NOTHING — so two concurrent deliveries of one id cannot both conclude they are first. Ids are swept with the leases after 48 hours, well past any forge's retry window. A missing delivery header is treated as fresh rather than as a duplicate: a forge that sends none must not have every event silently dropped. The event BODY is not stored. It is text written outside the factory (backlog #9), and an audit row is read by humans and future prompts alike; the metadata is what routing will need. Tests: verified event recorded, wrong and missing signatures rejected with nothing touched, three deliveries of one id acting once, no-delivery-header not deduped, unknown provider 404, and a push webhook creating neither an audit row nor a work item. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Two changes that were asked for directly, plus the ADR the second one needed. CONFIGURATION MOVES OUT OF ENVIRONMENT VARIABLES. PLOEG_TARGET_MAP, PLOEG_TEAM_MAP and PLOEG_TEAM_PLANS were three hand-rolled DSLs with no schema, no comments and no diff worth reading. They are replaced by one YAML file, rendered from Helm values into a ConfigMap and mounted at /etc/ploeg/ploeg.yaml. A typo'd key now fails the boot (KnownFields) instead of silently taking a default. The sharper half is the magic numbers. "11/bronze=webgrip/ploeg@development" put a Vikunja project ID into cluster config, where a bare 11 says nothing about which board it is, cannot be reviewed, and silently routes work to the wrong repository the day the project is rebuilt with a new id. Projects are NAMED now: trackers: vikunja: projects: - name: "Ploeg Test" repo: webgrip/ploeg branch: development ploegd asks the tracker which id that name has at boot, logs what it resolved, and refuses to start if the name matches nothing — with the available names in the error, so the operator can fix it. `id:` remains as an escape hatch for a board whose names are not unique. The env vars still work when the file is absent, so this migrates one deployment at a time. It renders to the same wire format pkg/target already parses, so there is one routing resolver in the codebase, not two that drift. PUSH RIGHTS ARE MINTED PER RUN (ADR-0013 tier 2, now accepted). Tier 1 gave readers a weaker static token, which closes the hole that matters. This closes the other one: a writer pod partitioned from ploegd keeps running after its Lease expires, and with a shared static credential it can still push to a branch another Run has since taken over. Now a writing Run gets a write:repository token minted for it alone, named ploeg-run-<12hex>-<repo> so a token in the forge UI traces to a Run, a Shift and a ticket the way the LiteLLM alias does. Revoked on report, on lease expiry by the sweeper, and by a boot sweep for whatever a crashed ploegd left behind — the same three-layer shape ADR-0008 uses for spend, because that pattern is proven here and a second novel one would be a second thing to get wrong. If the mint fails, nothing runs: the Run is finished as a retryable infra failure and the pod exits empty-handed, rather than proceeding with a credential we did not intend to hand out. Readers get nothing minted at all — no Lease, no business pushing. The admin credential lives only in ploegd, never in a worker pod (R6). That is a real escalation and ADR-0013 accepts it explicitly, on ADR-0008's reasoning. Unset, everything falls back to the shared token and the worker path is identical, so this ships dark. Honest limitation, recorded in the code: Forgejo's token API scopes by permission, not by repository, so the repo in the token name is audit rather than enforcement — the bot's own repository access is still what bounds it. ADR-0017 flipped proposed -> accepted: its implementation merged in #26. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>All three broke, or hid, real dispatch on 2026-07-30. 1. The worker ServiceAccount guard inferred create from name. `if and .Values.executor.enabled (not .Values.executor.serviceAccountName)` read a name as "an external account exists", so naming the chart's OWN default suppressed the account the pod template then referenced and every worker Job died with `serviceaccount "ploeg-worker" not found`. Naming the default and saying nothing must be equivalent; they were opposite. Split into executor.serviceAccount.{create,name}, and route the object and the reference through one helper so they cannot disagree again. 2. The DinD hazard label never reached the Job. It was stamped on the pod TEMPLATE only. Kyverno autogens a Job rule for pod-security-baseline-enforce, and a Job selector matches the JOB's own labels — so the PolicyException could not admit the Job and every DinD team was rejected at admission. KEDA copies ScaledJob labels onto its Jobs, so the label now goes there too, gated by the same role->team->global dind resolution the pod template uses (extracted to ploeg.roleUsesDind). Consumers can then key their exception on the hazard instead of on a list of workload names. 3. maxFixRounds was never serialised. PLOEG_TEAM_PLANS was built as {pool, rounds}, dropping the field entirely, so ADR-0017's request_changes loop was unreachable from Helm values — a reviewer verdict could never re-open the writing Round. Verified by rendering: naming the default now both creates and references the account; create:false skips creation and keeps the external reference; the hazard label appears on a dind role's ScaledJob and pod template and on neither for a dind:false role; and maxFixRounds:2 reaches PLOEG_TEAM_PLANS. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>ploeg is the trial for webgrip/workflows ADR-0005: the release toolchain becomes an image instead of an install. Run 122 is the before picture — 413 packages in 2m, 15 more in 21s, an npm audit summary nobody reads, and a yq download, all to spend ~25s cutting v0.2.0-rc.9. None of it locked, so that release and the next one were not guaranteed to run the same plugin versions. harbor.webgrip.dev/webgrip/semantic-release:0.1.2 carries node, git, yq, semantic-release 25 and @webgrip/semantic-release-config from a committed lockfile, and exports SEMREL_PREBAKED — which is all the composite needs to skip installing entirely. Built in webgrip/infrastructure. Two deliberate limits on this commit: The container is on the release job ONLY. `checks` needs go and helm, which this image does not carry, and gluing them in would make the release toolchain a build farm — the thing rust-releaser already is and this family exists not to be. The composite is pinned to a COMMIT on the semrel branch, not @main and not the branch name. @main still installs at release time and would ignore the image; pinning the branch would let a later push change what this release runs. Flip it back to @main once that PR merges — this pin is not meant to outlive the trial. This is the first time the container path runs anywhere. If the image is missing something ploeg's config reaches for, this is where it surfaces: the config is makeConfig({manifest:'helm'}) with the dependency-free Chart.yaml bump, which needs nothing outside the image, so the honest expectation is a clean release with no npm output at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Run 143 was half a test. It ran the composite in harbor.webgrip.dev/webgrip/semantic-release:0.1.2 and showed exactly what the toolchain image is for — zero npm output, 44s for the whole composite against 3m45s of installing — but it analyzed two commits, found neither releasable, and stopped. Everything after analyzeCommits is still unproven on this path: prepare, the makeConfig({manifest:'helm'}) Chart.yaml version/appVersion bump, @semantic-release/git's push-back through the new public-host extraheader, and the Gitea publish. This commit is a `fix:` on purpose. It releases, so the rest of the path runs. The pin moves to 2c5c68c, which changes the composite's prebake guard from "is there a file at $SEMREL_PREBAKED/.bin/semantic-release" to "is that tree actually the ADR-0005 toolchain". ploeg is not what that fixes — in the container the answer is yes either way — but ploeg is what proves the probe does not break the case that already worked. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>The release job has failed on every push to development since 2026-07-31 — seven runs, no rc cut, so the fixes merged on 2026-08-08 have no image. The job never reached its first step. It pinned webgrip/workflows' semantic-release composite at 2c5c68c9, which was amended off fix/semrel-parity-and-hardening (head is now ee06ca5, same parent, same tree but for reworded comments). The commit lives on in local clones, so the pin looked fine from a workstation; on the server it is unreachable from any ref, and the Forgejo runner resolves an action by cloning branch refs over anonymous HTTPS. Hence, every run: could not determine the commit ID of 2c5c68c9...^{commit}: fatal: ambiguous argument ... unknown revision or path not in the working tree Re-pinned to ee06ca5. The comment now records that a SHA pin is only as immutable as the branch it hangs off, and corrects the stale claim that publish had never run on the container path — run 144 cut v0.2.0-rc.15 on it, Chart.yaml bump and git push-back included. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Harbor tags are immutable, so re-pushing an existing version is a hard failure rather than a no-op. The rc.21 backfill (run 173) died on failed to push harbor.webgrip.dev/webgrip/ploegd:0.2.0-rc.21: 'ploegd:0.2.0-rc.21' configured as immutable and took the whole chain with it — signing skipped, both mirrors dead, and the two GitHub jobs never even instantiated, since every one of them needs this job. An already-published version is the desired end state, not an error: downstream jobs only ever copy the image out of Harbor by digest, and that digest is already there. So probe the registry first and skip the build if the manifest resolves. Fail-open — anything other than a clean 200 falls through to the build, so a fresh tag 404s and builds exactly as before. This is inert on the normal release path and only bites on a re-publish. The label verification is gated with the build for the same reason: it asserts image.created against the timestamp stamped in THIS run, and an image that was already published legitimately carries the one from when it was built. Not fixed here: the Harbor CHART push has the identical flaw (412, immutable). That job is independent, so a re-publish now gets the image, the signature and both mirrors, and only the chart step fails. Noted in the dispatch comment, which until now advertised a backfill path that could never work. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>rc.22 (run 175) built and pushed a perfectly correct image and then failed verification on it, taking the signing job and both mirrors down with it. Two independent bugs in the check, both the same class of mistake — matching text instead of data: grep -qF '"org.opencontainers.image.created":"<value>"' looks for `"key":"value"`, while `imagetools inspect --format '{{ json .Image }}'` pretty-prints `"key": "value"` with a space. Every one of the three assertions failed on labels that were present and correct. The 1970 guard was wrong for the same reason: it searched the whole config blob for the epoch, and BuildKit stamps every layer's history entry with exactly that for reproducibility — 12 of them in the rc.22 image. A good build always contains the string, so the guard could only ever fire on a false positive. Both are replaced by one jq assertion over the parsed document, comparing the labels themselves against the expected values — which covers the ARG-default case precisely, since the default is never the expected timestamp. `.Image` is normalised so the single-platform object and the multi-platform map assert identically, and a failure now prints the actual per-platform labels instead of dumping the entire config. Verified by running the step body against the real rc.22 image JSON recovered from run 175 (passes), a single-platform variant (passes), and a copy with one platform's label corrupted to the 1970 default (fails, and names the platform). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Bumps docker-build-push-registry-fast and cosign-sign-attest from v1.0.0 to v1.10.0. The only drift on those two files across that range is the ::error::/:⚠️: sweep — verified with git diff, 15 changed lines, all of them the marker prefix. No behaviour change. That completes the sweep for ploeg: nothing on this repo's release path emits a GitHub annotation command any more, which on Forgejo only ever produced an uglier log line. Also corrects the workflow_dispatch comment. It claimed the Harbor chart push still fails on a re-publish; v1.10.0 probes the registry and skips an already-published version, and run 184 was green end to end on exactly that path. And drops the temporary GHCR_TOKEN fingerprint workflow — it did its job: the stored secret was a reconciled value from OpenBao, not the UI edit we kept making. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Ploeg shipped one provider per seam — Vikunja and Forgejo — which is enough to prove the SPI and not enough to run anywhere else. The second estate's staging cluster is GitLab and ClickUp, so both seams needed a second implementation before Ploeg could be exercised there at all. Both are pure SPI implementations: no core change, nothing vendor-shaped escapes either package (R7), and each is registered under its dialect name so the existing /webhooks/{tracker,forge}/{provider} routes reach them unchanged. GitLab differs from Forgejo in three ways a copy-paste gets silently wrong, and each is commented where it bites: - GitLab does NOT sign webhooks. It echoes a shared secret in X-Gitlab-Token, which authenticates the sender but not the payload. Compared in constant time, since a bare == leaks a secret a byte at a time. - A merge request has two numbers; only `iid` works in an API path. - Projects live at arbitrary subgroup depth, so the path is URL-encoded whole rather than split into owner/name. Review outcomes also arrive as merge_request actions rather than a review object, and a branch pipeline has no MR — reported as PR 0, which the core reads as "nothing to route this to" rather than as merge request zero. ClickUp differs from Vikunja in three ways, likewise commented: - Auth is the raw token, no "Bearer " prefix (ClickUp 401s the prefixed form). - Priority is inverted — id 1 is urgent — and is flipped on the way in rather than leaking backwards ordering into scheduling. - There is no global "done": status is a per-List custom string, so SetStatus needs DoneStatus configured and skips loudly without it instead of guessing a name that would 400 or move the task somewhere nobody chose. ClickUp's webhook is thin and carries no List, so Scope is left empty at parse time and resolved in FetchItem — the thin-payload rule doing exactly what it is for. Wiring: trackers become a registry built once rather than two inline literals, and the forge instance id (ADR-0016) is bound explicitly. With one forge configured it takes the id; with two and PLOEG_TARGET_FORGE naming neither, nothing is bound and it says so — publishing findings to the wrong forge is worse than not publishing. New env: PLOEG_CLICKUP_{SECRET,TOKEN,URL,DONE_STATUS}, PLOEG_GITLAB_{URL,TOKEN,SECRET}. Absent, behaviour is byte-identical to before. Gates: go build, go vet, go test ./... (all green, including cmd/ploegd's existing wiring tests), gofmt clean, helm lint and all three chart renderings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>ploegd has spoken GitLab since rc.31 — pkg/provider/gitlab comments on a merge request and verifies inbound webhooks. ploeg-worker never learned: it polled /api/v1/repos/{owner}/{name}/pulls and its briefing named the Forgejo API as the way to open one. On a GitLab target no change request was ever created, so the Shift had none for a reviewer to comment on, publishRound logged "no pull request on this shift yet", and the review loop could not close. Silent all the way to a human. Less was missing than it looks. Three things were already forge-agnostic and are untouched here: git.go's authURL/plainURL build owner/name into a URL that is already correct for a GitLab subgroup; shiftengine's prPathRe already matches /merge_requests/(\d+); and the Shift takes its change-request URL from the OutcomeReport links, not from the poll. The gap was two places that name a forge, so that is what this changes. RepoRef gains Forge, the API DIALECT. Empty means forgejo, so every stored target, every taskspec and every deployment that never set it keeps its exact current meaning. The dialect travels on the work item — pkg/work.Target has carried Forge all along — and falls back to the worker's configured default, matching the "empty = the default forge" promise ploegd's own registry makes. - findPR dispatches. GitLab filters source_branch server-side, so unlike the Forgejo call it cannot be defeated by a repo with 50+ open requests. The project is addressed by URL-ENCODED full path: acme/internal/widgets is three segments and the slashes must survive as %2F. - The briefing dispatches, in vocabulary as well as endpoint. An agent told to open a "pull request" on GitLab looks for an endpoint that is not there, and the noun is what it searches its tools and the repo's docs for. Noun and endpoint come out of the same switch so they cannot drift. - An unknown dialect fails loudly. Falling back to Forgejo would poll a real endpoint shape against the wrong host and report "no change request" forever — indistinguishable from an agent that never opened one. Chart: executor.forge selects one active forge and executor.gitlab configures it. The new ploeg.forge helper resolves whichever is active into one shape, so no template touches .forgejo or .gitlab directly. That also removes a trap: the worker template used to dereference .Values.executor.forgejo.url unconditionally, which made `forgejo: null` — the documented way to empty an unused block — a nil pointer the moment executor.enabled flipped true. The new GitLab fixture renders with forgejo null precisely so that cannot come back. FORGE_URL is the name; FORGEJO_URL is emitted alongside it and still accepted, so a ScaledJob starting a pod from the previous image mid-upgrade still finds a forge. A fourth golden, executor-gitlab, because selecting a forge changes the worker pod — a different credential Secret and a different API — and the worker pod is the boundary the goldens exist to police. It pins the ADR-0013 tier-1 split holding on GitLab: readers draw agent-reader-token, the writer draws agent-builder-token. Tier 2 is deliberately absent. ploegd mints per-run push credentials through /api/v1/admin/users/forge, a Forgejo admin endpoint with no GitLab equivalent; the nearest analogue is a project access token, a different escalation that deserves its own ADR rather than an implied one. Unset means the shared token, which is the documented pre-tier-2 behaviour. Gates run locally with the pinned toolchain (helm v4.2.3 as CI pins; 4.2.4 disagrees about the blank line before a document separator, as scripts/ helm-golden.sh warns): gofmt -l . clean go vet ./... clean go build ./... ok go test ./... ok, all packages incl. pkg/store helm lint 1 chart linted, 0 failed helm-golden.sh check ok, 4 renders No VIK trailer: this did not come from a board ticket.RepoRef.Forge landed without its schema edit, which docs/contracts/README.md does not allow: v1 changes additively, and the Go type and the published schema change together. `repo` is additionalProperties:false, so a Task Spec carrying the field would have been rejected by any consumer validating against the published contract. The gate did not catch it because fullTaskSpec — the fixture whose whole job is to carry every field — did not set Forge, and omitempty dropped it. Fixed in the order the tasks rules ask for: the fixture first, observed to fail with at '/repo': additional properties 'forge' not allowed then the schema. The enum is constrained to the dialects the worker actually implements, so a Task Spec naming a third forge fails at the contract rather than at run time.Review feedback. Every comment this change added is gone; the diff now adds none. What the prose carried is carried by names, types and the ADR instead. prMatches -> isRunChangeRequest(changeRequest, runBranch, base) findPR -> findOpenChangeRequest listForgejoPRs -> listForgejoPullRequests listGitLabMRs -> listGitLabMergeRequests forgeGet -> getJSON openChangeRequest -> openChangeRequestInstruction Config.Forge -> Config.DefaultForge changeRequest{URL,Head,Base} -> {URL,HeadBranch,BaseBranch} The "no silent fallback" comment is now errUnsupportedForge, a named sentinel the test asserts with errors.Is rather than a substring. The subgroup comment is RepoRef.ProjectPath, which joins and never splits. The chart's helper comment is the helper's own shape. TWO NAMES FOR ONE VALUE, REMOVED. requireEnvOneOf("FORGE_URL", "FORGEJO_URL") was a shim for a rolling upgrade that cannot happen: the release train keeps chart and appVersion in lockstep, so the chart and the image it configures move together. The worker now requires FORGE_URL and nothing else, and the chart emits only that. requireEnvOneOf is deleted. The dialect had the same problem in a worse form: PLOEG_FORGE on the worker was a second spelling of PLOEG_TARGET_FORGE, which ploegd has read since the forge registry landed — the same concept, the same default, two names, one of them invented here. The worker now reads PLOEG_TARGET_FORGE too, and the chart renders it once for both binaries from executor.forge, so they cannot disagree about which forge is the default. Both renames are breaking for anyone setting these by hand and neither is for a chart-driven deployment. That trade is the point: two names for one value is a worse thing to own than a rename under a version bump. values.yaml loses its comment blocks; the meaning moved into values.schema.json descriptions, which is where a Helm chart keeps structured intent — validated, machine-readable, and shown by tooling rather than only to whoever opens the file. Goldens regenerated (helm v4.2.3, as CI pins). Gates: gofmt clean, go vet clean, go build ok, go test ./... ok, helm lint ok, 4 renders ok, helm-golden.sh check ok, go test ./internal/ledger/ ok, openspec validate --all 6 passed.ploeg.workerPodTemplate hardcoded nodeSelector: node.webgrip.io/pool: worker with no values hook. That label exists in the homelab and on no other cluster, so every worker Job on any other estate is Pending forever. Nothing reports it: the ScaledJob is created, KEDA scales it, the pod never schedules, and the Work Item sits queued with a healthy-looking release above it. It is the same trap as the ploegd nodeSelector default, one layer down and worse, because ploegd's could at least be overridden. The default is unchanged, so the homelab renders what it rendered before — the goldens move by exactly one thing, the inline comment that toYaml does not carry. An estate with different labels sets its own; one with none clears the block with null. {} does not clear it, because Helm deep-merges maps, which is the same footgun the existing nodeSelector and forge blocks already carry and is now said once in the schema. ADR-0002's constraint is intact: the reason for a selector is keeping DinD off control-plane nodes. That reason is a deployment's to enforce with its own labels, not the chart's to assume with someone else's. Gates: gofmt clean, go vet clean, go build ok, go test ./... ok, helm lint ok, 4 renders ok, helm-golden.sh check ok.The Project type has carried this comment since routing moved into the file: Team routes this project's work to one team. Empty means the assignee decides, via the team's `assignees` list below. Only the second sentence was ever implemented. `team:` was consumed solely as a routing-table qualifier — "<scope>/<team>=repo" disambiguates WHICH REPO a team's work lands in — while the team itself always came from the assignee mapping, falling through to PLOEG_DEFAULT_TEAM for anyone unmapped. A config that pinned a project to `team: app` routed nothing to app unless the assignee independently mapped there, which makes the pin a comment about an intention rather than a control. The gap has a concrete cost on a real board. On ClickUp, assignees are licensed workspace members — there are no bot users to invent, and the validator's one-assignee-one-team rule means a PERSON can only ever trigger a single tier. The natural gesture — three Lists, one per tier, drop a task in the right one — was unexpressible. Now the pin decides. After mirror() has fetched the item (a thin clickup webhook carries no List, so the scope is only known post-fetch, which is why this cannot live in a provider), ingest overrides the assignee's team with the container's pinned team, then resolves the target — so the team-qualified routing rules see the team the work will actually run as. Unpinned containers are byte-for-byte unchanged: the assignee still decides. Config grows ScopeTeams(): every project with both a pinned id and a team. Name-resolved projects are deliberately absent until a resolver hands their ids back — the shape that needs pinning is the shape that pins ids. Gates: gofmt clean, go vet clean, go build ok, go test ./... ok, goldens untouched (no chart change).Four gateways enforce spend caps, not three: LangSmith's LLM Gateway does it at organisation, workspace, API-key and user level with per-trace attribution. The pattern is unchanged and reinforced, since every one of them is a gateway. Record AG2 v1.0's Network as BAND's closest architectural analogue: a hub with a registry, audit log and typed channels over a pluggable WebSocket transport, Apache-2.0. Also note LangGraph Platform serves A2A at /a2a/{assistant_id}, so A2A adoption is wider than the two GA surfaces the Linux Foundation named. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>The adapter is lifted through RunCommand against a fake claude on PATH that records its argv and working directory, and asserts that the process receives --settings '{"disableAllHooks":true}' as one argument and --strict-mcp-config without an --mcp-config. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>Unassigning a Vikunja or ClickUp item now withdraws its Work Item: the live Shift closes with reason withdrawn_unassigned, pending Runs are cancelled, running Runs are finished and their model keys blocked, the Lease and its push credential are released, and the item moves to the new withdrawn state. No sweep retries it; a new assignment re-queues it with attempts reset. POST /api/v1/operator/work-items/{id}/cancel does the same for an execute-permitted consumer, closing with withdrawn_by_operator. Items bound to an Operator Execution are left to that execution's cancel. Migration 0017 admits the withdrawn state; the operator, task spec and tracker execution schemas list it, and SettleItem no longer settles a withdrawn item back into the queue. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>An outcome report may now carry createdWorkItems (title, description, ready, kind and an optional team). ploegd stores each accepted entry as a follow_up Work Item that exists only in Ploeg, names its source Work Item and Run, sits one level deeper than its source and inherits its Work Target. Entries over a Team's limits are rejected with a reason recorded in the audit log. Each Team gets conservative defaults, overridable under teams.<name>.createdWork: maxCreatedPerRun 5, maxDepth 2, maxOpen 20, an item budget of 2.00 USD that caps the created Work Item's Shift pool, and a 10.00 USD pool per root Work Item. Created work waits as proposed until an operator approves or rejects it through the new /api/v1/operator/work-items/{id}/approve and /reject routes, unless the Team sets autoDispatch. Work that is not Ready goes to a configured refinement team or planner Role, and otherwise stays proposed. A plan Role marked planner gets a planner prompt that asks for createdWorkItems instead of code. VIK-1122 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>Vloer had no guard: one breaking commit would have computed 1.0.0-rc.1, which the homelab Renovate rule (/^0\./) never offers. Vloer now mirrors Ploeg's policy: breaking changes raise the minor version and verifyRelease refuses anything but 0.x.y-rc.N from development. The release-policy action runs the new suite; the isolation test copies both applications' policy scripts. The weekly schedule gains a release-notes job that fails when Forgejo no longer holds every imported semantic-release note, after the second loss on 2026-09-23/24 went unnoticed until a release preview. Also fixes ADR-0028's stale v${version} tag wording. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>The Forgejo broker called /api/v1/admin/users/{bot}/tokens with an admin token. Forgejo 15 serves no such route (404) and refuses token authentication on /api/v1/users/{bot}/tokens ("auth method not allowed"), so per-run push credentials could never be enabled: every writing claim would have failed to mint. Checked against a local Forgejo 15.0.9. ploegd now signs in as the bot with its password (PLOEG_FORGEJO_BOT_PASSWORD, chart executor.forgejo.botPasswordSecret) and mints each token with write:repository and a repository list holding only the Run's repository. Git and the repository API answer 403 for any other repository the bot can see. The chart refuses the old adminTokenSecret, ploegd warns and ignores PLOEG_FORGEJO_ADMIN_TOKEN, and a worker refuses to start with the bot password in its environment. VIK-733 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>rc.3's final distribution still failed for both applications: Forgejo answers the unlink call with HTTP 500 ("no permission to unlink package ... from its repository") because the CI bot has no rights on the archived webgrip/de-vloer and webgrip/ploeg repositories. The link only decides which repository page lists a package, so a link that cannot be moved now logs a warning and publication continues to GHCR, the Go module export and the GitHub release. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>d63e7cblanded 3e4d3f6efdd63e7cblanded' (#4) from agent/vik-1289 into development a309c269d9Vloer opened on Sessions; the owner had to walk four Ploeg tabs and one team lane at a time to see what needed them. Now is the start view: for every team the user may read it shows waiting Work Items, running Runs and recently finished Runs, and stays useful when one Ploeg call fails. - PloegClient.now(user) reuses items/proposed/runs with per-group error capture; GET /api/ploeg/now returns {waiting, running, recent, errors, fetchedAt} and keeps the existing 403 ploeg_scope for scoped callers. - runRow parses the optional observedUsd/reservedModels fields; missing model or spend renders as "not reported", never 0. - waiting rows carry team, state, title, age, cost (2 decimals) and tracker/pull-request/Grafana links; awaiting_review detail is fetched for at most 10 items, same cap as the Proposed lane. - public/now.js renders the three groups with inline retry per group; registered in the static asset map, route/nav and j/k keyboard handling added in app.js, styles in styles.css, demo data extended. - Tests cover the scope 403, per-group failure and "not reported", plus a browser case asserting no horizontal scroll on #now at 390 px. VIK-1389 Agent-Trace-Id: ploeg-b73fda03b059A Request is a Client's ask and its thread; Refinement turns one Request into one or more Work Items, each with its own Quote. Dispute is a billing disagreement between Glide and an Agency only; a Client objecting to a preview is feedback the Agency decides. Reversal refunds a Delivery Fee after a defect-caused revert within 14 days. Markup Tier is the only use of "tier"; a Size maps to a Team. Edition names the plan an Agency is on, and the middle one is renamed from Agency to Studio so the two never share a sentence ("Agency Edition" is retired). How a Request that arrives as a tracker task relates to Ploeg's Tracker Item mirroring stays open as an ambiguity with a recommendation. ADRs 0005-0007 and the Who Glide is for page use the new words. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>The child wrote its pid with writeFileSync, which creates the file before writing to it, and the test polled only for the file to exist. On a busy runner it could read an empty file: Number('') is 0, and kill(0, 0) signals the test's own process group, so the "child is gone" loop ran its full window and the ESRCH assertion failed after about 11 s (PR 16, run 161). The child now writes to a temporary name and renames it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>The browser check is now a runner that starts the demo and live servers once and runs one flow module per area from scripts/browser/ (now, tasks, sessions, shell, settings, feeds, work, login). Each flow exports run({ page, app, live, password, assert, screenshot }), so parallel work on different screens edits different files. Assertions are unchanged; the Sessions heading check moved from before the Tasks steps to the start of the sessions flow. The Now page tests from PR #16 move unchanged from ploeg-view.test.mjs to now-view.test.mjs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>core/format.js is the one formatter: US dollars in nl-NL with two decimals ("US$ 1.234,50"), "< US$ 0,01" for a positive amount under a cent with the exact value in a title, and "Not reported" for an unknown amount, never zero. Dates read 30-09-2026 21:30 in the browser time zone, relative time is compact English ("5 min ago", "in 5 min") and turns absolute after seven days, durations read "1 h 05 min", and plural() counts nouns. A 'browser' locale preference is available. The earlier money, clock and ago names delegate to the new formatters, so every view that still calls them now shows nl-NL amounts and times. core/states.js holds one vocabulary with a label, a tone and a glyph for Work Item states, Run states, outcomes, agent verdicts, failure reasons and session statuses. Done is never "Merged" and an agent verdict never reads as human review. Session labels now come from it, which fixes two labels: a completed session without a review reads "Ready for your review", a failed one "Failed", and the session list shows "Accepted" or "Rejected" once a review is recorded. core/reasons.js derives why a Work Item waits on a person from its close reason, with the corrections from the research critique: plan_exhausted in needs_human names both causes, the prefixed "budget exhausted", "run stuck" and "plan removed" reasons are handled, a writer killed by the cluster is not the Work Item's fault, and an unresolved repository is a secondary "Not routed" warning. The detail reason adds Ploeg's own needs-human sentence and the stuck or failing Run. Every reason ends with re-assigning the task in its tracker, the only way to start again today. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>The sidebar stays dark (Hal) in both themes. It carries the outlined De Vloer lockup inlined with currentColor and groups the navigation: Now; Ploeg with Work, Proposed, Runs, Activity and Insights; Workbench with Tasks and Sessions (only in demo mode, with shared execution or when sessions exist); Settings at the bottom. Counts come from state.counts: what waits on you on Now, proposed work on Proposed, sessions that need you on Sessions. Below 1100 px the sidebar becomes an icon rail with tooltips. The top bar holds breadcrumbs, the search trigger (Ctrl or ⌘ K), a status strip with the Ploeg connection, "Updated … ago", a Live or Paused switch and the DEMO or LIVE badge, and an account menu with the System, Light and Dark theme switch, Preferences, keyboard shortcuts and Sign out. The working-agreement block, the demo ribbon and the hard-coded version footer are gone; the sidebar shows the version only when the bootstrap reports one. Under 720 px a compact top bar opens the navigation in a <dialog> drawer and a bottom bar offers Now, Work, Runs and More. shell(content, { title, subtitle, overline, actions, breadcrumbs, wide }) is the new signature; shell(content, title, subtitle) keeps working. document.title leads with the waiting count, "(4) Work · De Vloer". core/counts.js reads GET /api/ploeg/now every minute through the live scheduler, and the Now page and proposal decisions update it directly. The skip link now moves focus to the main content instead of routing to #main. A 401 in the middle of a session closes open dialogs and the sign-in page says the session expired; the hash survives, and single sign-on keeps the deep link through sessionStorage. styles/shell.css is self-contained with token fallbacks and is imported from styles.css until the layered stylesheet lands. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>Runs - The table fits from a 60rem container up (a 1280 px laptop): fixed columns Status, Work Item, Started, Spend and Model, with the Role in the Work Item's meta line. Model names stack one per line. - A running Run reads "Running for 20 min" with the start time in the title; a pending Run reads "Waiting for a worker" and "Not authorized yet". A failure shows its next step on screen, not only in a title. - The spend meter says "observed so far" to assistive technology while a Run has not settled, drops the percentage in the table, and reads "Not reported" alone when no budget is known either. - The demo shows a muted dash with "Demo · no model calls" in the title instead of repeating "No model calls" on every row. - Phone cards are about half as tall: status row, title, one meta line, one timing line, model and tokens, then the meter. Proposed - One compact footer per card ("Spends from the delivery Team's budget" with Reject and Approve); the budget caveat is said once per page and viewers get one "who decides" note. "Readiness not reported" is gone and the count reads "3 proposals". - The reject dialog's cancel reads "Keep it proposed" and its reason field is described by the hint alone until an error shows. Activity - A dot separates the event label from the Work Item title on wide screens; on phones the title takes its own line. - A long Team name is cut short with its full name in the title on wide screens and wraps on phones instead of overflowing. Insights - The description reads "Last 7 days · 3 Teams"; the settlement note moves under the table. With no Teams only the empty state shows. - Work Item tiles never leave holes (5, 3 + 2, 2 + 2 + 1 full width), phone cards keep amounts on one line ("Reserved"), the demo tile reads "No model calls", and tables are no tab stop. All four - A failed refresh reads "Could not refresh. Showing data from 11:48." in the attention tone; on phones Refresh is an icon beside the title and the filters sit in a two-column grid, with the disabled Outcome filter explained on screen. - Loading skeletons match the loaded layouts. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>The sentence about what search covers moves from the Open Work row into the no-match note, where it explains the empty result, and the row says only where Work leads ("Every lane of every Team"), so it no longer truncates on a phone. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>- Work Item links without a lane open beside the lane the item is in and write it into the hash; the phone top bar goes back to that lane, Tasks and sessions use the same shell back option, and the in-page back links are gone. - The Work Item page takes its reason chip from the list's code ("Every Round ran" replaces the hedging plan_exhausted chip); a unit test checks list and page chips for every demo item. - core/states.js is the one table for audit labels, close reasons, actors ("Agent · delivery", "You"), verdict short labels ("Agent approved", "Agent asked for changes"), tile details and Runs cancelled before they started; Activity, Runs, Now, the Work Item page and sessions all read from it. - A rejected proposal reads as a neutral "Rejected" (circle-slash) in lists and on its page; its state stays done. - Needs you follows one rule on Now and Work: flat rows with their chip, and a sunken overline band only for a reason two or more Work Items share. Now's why line leads with the fix. - Decision boxes never show a disabled primary: without a link they name the task key with a Copy button. - The top bar keeps a fixed search field, folds the session stream and "Updated" into the Ploeg dot's tooltip, makes auto-refresh an icon button (Pause/Resume auto-refresh), uses the shared avatar, dims the Ploeg group while unconfigured, and shows the phone title only once the h1 scrolls away. - Every page header has a subtitle and one icon Refresh; Work shrinks its h1 on detail pages and the item title takes the page-title scale under a "Work Item" overline; Cancel Work Item is a neutral ghost. - One demo disclaimer wording and layout, hidden on phones; one unconfigured-Ploeg empty state; one budget field (US$ prefix, 5,00, comma or full stop); Settings cards use the card component; #settings opens on Preferences. - Meters draw settled spend in neutral ink; the Round ladder's Shift budget spans the card; Running now sizes to its content; the Runs demo drops the Spend and Model columns; meta lines never start a wrapped line with a dot; engine jargon (lease deadlines, Shift ids, "Worker lost") is gone; overlays close with the dialog ×; in-app headings drop their full stops; avatar text is at least 12 px. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>Release run 358 left glide-v0.4.0-rc.15 unpublished: de-vloer-agent had three HIGH findings against a zero budget, all in dependencies npm bundles, advisories published 2026-09-29 (brace-expansion 5.0.9, fixed in 5.0.11/5.0.12; undici 6.28.0, fixed in 6.28.1). No npm release ships the fixes yet, including 12.1.0. The fetch stage packs the fixed releases, pinned as annotated NPM_*_VERSION args so Renovate moves them, and the final stage swaps them into npm's own node_modules. patch-bundled-npm.mjs replaces only an older copy of the same major, logs when npm has caught up, and fails the build unless npm still expands `{a,b}.js` correctly. Verified locally: npm 12.0.2 with brace-expansion 5.0.12 and undici 6.28.1, opencode 1.18.30, and grype reports neither package. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>db49e45committed 45db4712ebA writer writes {"problem","solution"} to PLOEG_OUTCOME_FILE and opens its pull request with the same two sections. The drop box keeps both whatever the adapter concluded, execbin accepts a file without an outcome, and ploegd stores them on agent_runs for writing Runs only (migration 0022). The operator API returns both on every Run. They decide nothing. ADR-0042 (proposed). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>GET /api/v1/operator/work-items/{id}/card returns one card per Work Item with its pull requests as plays, the crew, roster, steward fallback, usage totals and a provenance timeline. Webhooks and the review poller now capture diff stats and the combined CI status for Ploeg pull requests (migration 0024), and Work Targets accept a cardStyle with a skin and theme. ADR-0046 (proposed) records the decision. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>A Work Item page opened on a small "Ready for review" chip above the problem/solution card, with the Run card as a third-width section beside empty space, an overflowing cost ring and a truncated Tokens tile. The Run card now heads the detail across the full content width, states what happened in one headline (state, verdict, pull request, Rounds) and carries the primary action. The separate "Run card" section and caption are gone; the review box keeps only the checks and forge outcomes, folds the neutral checks into one muted line ("2 not reported: CI, tracker link") and closes "On the forge". The vloer-native skin holds the ring amount and "of" amount on two lines and reads "Not reported yet" in full. An empty lane's list column collapses when an item is open. VIK-1697 Agent-Trace-Id: ploeg-964232a0f888Vikunja and ClickUp skipped signature verification when their secret was empty, so anyone who could reach /webhooks/tracker/{provider} could assign or withdraw work. They now reject every delivery without a configured secret, the same as the Forgejo and GitLab forge providers, and ploegd warns at startup when a tracker secret is missing. VIK-1716 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>/healthz and /readyz called store.listSessions(), which parsed every session and decrypted its workspace secret on each probe. Probe cost grew with history, blocked the event loop, and one corrupt old session made liveness fail. /healthz now answers without touching storage and /readyz runs a single SELECT 1 through Store.ping(); both keep their {status, version} response and GET-only method check. The live event stream read every event after the client's cursor in one unbounded query, so a client reconnecting from an old cursor loaded the whole remainder before backpressure applied. It now reads through Store.eventPage() in batches of eventReplayBatch (200) rows, continuing from the last sent id until a short page or a buffered-output limit, and resumes on the next tick. Store.events() keeps returning the full remainder for the history endpoint, the engine and the AHP host, so no caller is truncated. VIK-1724 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>The migration runner checked schema_migrations outside each migration's transaction and held no lock across the loop. Two ploegd processes starting together could both decide a migration was unapplied, and one failed startup (in the regression test, already on the concurrent CREATE TABLE IF NOT EXISTS schema_migrations). Migrate now acquires one pool connection, takes a session pg_advisory_lock keyed on hashtextextended('ploeg.schema-migrations', 0) and runs the whole create-check-apply loop on that connection. The lock is released in a deferred unlock that also runs on error, with a context detached from cancellation; if the unlock fails the connection is hijacked and closed so the pool never hands out a connection that still holds the lock. A test runs two migrators concurrently against a fresh database in the embedded PostgreSQL: both return without error, each migration is recorded once and no advisory lock remains. VIK-1726 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>Signing in through an editor's link no longer completes the ticket. The callback records who signed in and lands on an approval page that names the request ("A VS Code editor is asking to sign in to <workbench> as you"), shows the short code the editor also shows, and offers Approve and Deny. Polling stays pending until approval; a denied, expired, consumed or mismatched ticket yields nothing. Approval issues a vle_ bearer token, stored only as a hash in editor_credentials, scoped to keep it from approving or managing editors, valid for thirty days, and listed under Settings > Signed-in editors where the person can sign it out together with the Agent Host tokens it issued. Ticket starts are limited per address and polls answer slow_down when they come too fast. The extension shows the code, honours slow_down and sends the credential as Authorization: Bearer. VIK-1718 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>PloegClient.announce sent PUT /api/v1/operator/consumer {"vloerUrl"} with backoff until Ploeg answered. No Ploeg release, branch or commit ever had that route, so every live workbench logged consumer_unsupported once and the call was dead. It was also the last place Unfold used the former name on the wire. Ploeg now links reports through deployment-set URL templates (ploeg-hq/ploeg#62), so nothing needs the announcement. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>The Work Item page gets a Context card: an "Add context" dialog sends a file and an optional note to POST /api/ploeg/work-items/{id}/context, which forwards the bytes to Ploeg's operator API as the signed-in person and keeps no copy. The list shows each item's size, file count, who added it and when, marks items added while steering, and says which Runs receive them. The application's engine never reads context, and the demo refuses uploads. Needs Ploeg 0.2.0-rc.3 or later. VIK-1858 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>