fix(chart): the worker node selector is a value, not a constant #47
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "worker-node-selector"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
fix(chart): the worker node selector is a value, not a constant
ploeg.workerPodTemplate hardcoded
with no values hook. That label exists in the homelab and on no other cluster,
so every worker Job on any other estate is Pending forever. Nothing reports it:
the ScaledJob is created, KEDA scales it, the pod never schedules, and the Work
Item sits queued with a healthy-looking release above it.
It is the same trap as the ploegd nodeSelector default, one layer down and
worse, because ploegd's could at least be overridden.
The default is unchanged, so the homelab renders what it rendered before — the
goldens move by exactly one thing, the inline comment that toYaml does not
carry. An estate with different labels sets its own; one with none clears the
block with null. {} does not clear it, because Helm deep-merges maps, which is
the same footgun the existing nodeSelector and forge blocks already carry and
is now said once in the schema.
ADR-0002's constraint is intact: the reason for a selector is keeping DinD off
control-plane nodes. That reason is a deployment's to enforce with its own
labels, not the chart's to assume with someone else's.
Gates: gofmt clean, go vet clean, go build ok, go test ./... ok, helm lint ok,
4 renders ok, helm-golden.sh check ok.
Found while wiring code14's staging cluster to v0.3.0-rc.1: its nodes are labelled
code14.nl/pool, so every agent Job would have sat Pending with nothing reporting it. Blocks RFC-0013 phase 4 there.Behaviour for the homelab is byte-identical — the only golden movement is the inline comment, which
toYamldoes not carry.ploeg.workerPodTemplate hardcoded nodeSelector: node.webgrip.io/pool: worker with no values hook. That label exists in the homelab and on no other cluster, so every worker Job on any other estate is Pending forever. Nothing reports it: the ScaledJob is created, KEDA scales it, the pod never schedules, and the Work Item sits queued with a healthy-looking release above it. It is the same trap as the ploegd nodeSelector default, one layer down and worse, because ploegd's could at least be overridden. The default is unchanged, so the homelab renders what it rendered before — the goldens move by exactly one thing, the inline comment that toYaml does not carry. An estate with different labels sets its own; one with none clears the block with null. {} does not clear it, because Helm deep-merges maps, which is the same footgun the existing nodeSelector and forge blocks already carry and is now said once in the schema. ADR-0002's constraint is intact: the reason for a selector is keeping DinD off control-plane nodes. That reason is a deployment's to enforce with its own labels, not the chart's to assume with someone else's. Gates: gofmt clean, go vet clean, go build ok, go test ./... ok, helm lint ok, 4 renders ok, helm-golden.sh check ok.