fix(chart): the worker node selector is a value, not a constant #47

Merged
ryangr0 merged 1 commit from worker-node-selector into development 2026-09-02 08:13:23 +00:00
Owner

fix(chart): the worker node selector is a value, not a constant

ploeg.workerPodTemplate hardcoded

nodeSelector:
  node.webgrip.io/pool: worker

with no values hook. That label exists in the homelab and on no other cluster,
so every worker Job on any other estate is Pending forever. Nothing reports it:
the ScaledJob is created, KEDA scales it, the pod never schedules, and the Work
Item sits queued with a healthy-looking release above it.

It is the same trap as the ploegd nodeSelector default, one layer down and
worse, because ploegd's could at least be overridden.

The default is unchanged, so the homelab renders what it rendered before — the
goldens move by exactly one thing, the inline comment that toYaml does not
carry. An estate with different labels sets its own; one with none clears the
block with null. {} does not clear it, because Helm deep-merges maps, which is
the same footgun the existing nodeSelector and forge blocks already carry and
is now said once in the schema.

ADR-0002's constraint is intact: the reason for a selector is keeping DinD off
control-plane nodes. That reason is a deployment's to enforce with its own
labels, not the chart's to assume with someone else's.

Gates: gofmt clean, go vet clean, go build ok, go test ./... ok, helm lint ok,
4 renders ok, helm-golden.sh check ok.

Found while wiring code14's staging cluster to v0.3.0-rc.1: its nodes are labelled code14.nl/pool, so every agent Job would have sat Pending with nothing reporting it. Blocks RFC-0013 phase 4 there.

Behaviour for the homelab is byte-identical — the only golden movement is the inline comment, which toYaml does not carry.

fix(chart): the worker node selector is a value, not a constant ploeg.workerPodTemplate hardcoded nodeSelector: node.webgrip.io/pool: worker with no values hook. That label exists in the homelab and on no other cluster, so every worker Job on any other estate is Pending forever. Nothing reports it: the ScaledJob is created, KEDA scales it, the pod never schedules, and the Work Item sits queued with a healthy-looking release above it. It is the same trap as the ploegd nodeSelector default, one layer down and worse, because ploegd's could at least be overridden. The default is unchanged, so the homelab renders what it rendered before — the goldens move by exactly one thing, the inline comment that toYaml does not carry. An estate with different labels sets its own; one with none clears the block with null. {} does not clear it, because Helm deep-merges maps, which is the same footgun the existing nodeSelector and forge blocks already carry and is now said once in the schema. ADR-0002's constraint is intact: the reason for a selector is keeping DinD off control-plane nodes. That reason is a deployment's to enforce with its own labels, not the chart's to assume with someone else's. Gates: gofmt clean, go vet clean, go build ok, go test ./... ok, helm lint ok, 4 renders ok, helm-golden.sh check ok. Found while wiring code14's staging cluster to v0.3.0-rc.1: its nodes are labelled `code14.nl/pool`, so every agent Job would have sat Pending with nothing reporting it. Blocks RFC-0013 phase 4 there. Behaviour for the homelab is byte-identical — the only golden movement is the inline comment, which `toYaml` does not carry.
fix(chart): the worker node selector is a value, not a constant
All checks were successful
On Pull Request / checks (pull_request) Successful in 53s
7cd54a768a
ploeg.workerPodTemplate hardcoded

    nodeSelector:
      node.webgrip.io/pool: worker

with no values hook. That label exists in the homelab and on no other cluster,
so every worker Job on any other estate is Pending forever. Nothing reports it:
the ScaledJob is created, KEDA scales it, the pod never schedules, and the Work
Item sits queued with a healthy-looking release above it.

It is the same trap as the ploegd nodeSelector default, one layer down and
worse, because ploegd's could at least be overridden.

The default is unchanged, so the homelab renders what it rendered before — the
goldens move by exactly one thing, the inline comment that toYaml does not
carry. An estate with different labels sets its own; one with none clears the
block with null. {} does not clear it, because Helm deep-merges maps, which is
the same footgun the existing nodeSelector and forge blocks already carry and
is now said once in the schema.

ADR-0002's constraint is intact: the reason for a selector is keeping DinD off
control-plane nodes. That reason is a deployment's to enforce with its own
labels, not the chart's to assume with someone else's.

Gates: gofmt clean, go vet clean, go build ok, go test ./... ok, helm lint ok,
4 renders ok, helm-golden.sh check ok.
ryangr0 merged commit e5df0021f2 into development 2026-09-02 08:13:23 +00:00
Commenting is not possible because the repository is archived.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
webgrip/ploeg!47
No description provided.