Runbook for restoring or reality-checking the on-prem tcfs namespace when
honey (or another local cluster) is intended to be the active authority.
There are currently three adjacent deployment paths in this repo, and they are not resource-name compatible with each other:
-
infra/k8s/charts/tcfs-backend- Direct Helm chart for the backend worker deployment and namespace scaffolding.
- Default live release name:
tcfs-backend - Expected resource names:
- service account:
tcfs-backend-tcfs-backend - deployment:
tcfs-backend-tcfs-backend-worker - config map:
tcfs-backend-tcfs-backend-config
- service account:
- If Helm is the authority, release state should exist in the namespace as
sh.helm.release.v1.tcfs-backend.vNsecrets.
-
infra/k8s/charts/tcfs-stack- Umbrella chart for blank-cluster bootstrap.
- Useful when you want Helm to create the broader stack, not when you are
reconciling an already-existing direct
tcfs-backendrelease.
-
infra/tofu/modules/tcfs-backend- OpenTofu-managed backend worker deployment.
- Uses different object names (
tcfsd,tcfs-sync-worker) and should not be treated as an in-place reconciler for a Helm-managedtcfs-backend-tcfs-backend-*namespace.
The practical consequence is simple: if the live namespace already contains
Helm-managed tcfs-backend-tcfs-backend-* objects, the direct Helm chart is
the source of truth for restoring that path.
Live readback from the Tinyland honey cluster on 2026-04-28 changed the
operator call from "repair a broken worker" to "decide authority before moving
placement":
nats-0,seaweedfs-0, andtcfs-backend-tcfs-backend-workerare Running.- NATS health returns OK, SeaweedFS reports a leader, and the worker logs show a live NATS connection.
helm list -n tcfsreports no release state.tcfs/natsandtcfs/seaweedfshave live Tailscale exposure annotations, but their last-applied Service configuration had empty annotations.data-nats-0anddata-seaweedfs-0arelocal-pathPVCs whose backing PVs have node affinity tohoney.
That means a ProxyClass-only patch would be cosmetic. It could move the generated Tailscale proxy pods, but it would not source-own the exposure, create release state, or make NATS/SeaweedFS data drain-mobile.
Treat the direct tcfs-backend Helm chart as a recovery authority for the
backend worker objects only. Do not use it as evidence that the whole live
namespace is Helm-owned.
The durable on-prem path is OpenTofu migration, not blind Helm adoption. The current source-owned on-prem environment records retained target PVCs, non-canonical candidate workloads, candidate tailnet Services, and render-only cutover/rollback commands. Do not switch canonical tailnet hostnames or move data outside an approved downtime window.
As of 2026-05-08, no downtime window is open. Keep the on-prem cutover deferred
from the current usage-reality sprint unless #327 is explicitly scheduled
with preflight, rollback, and post-cut smoke owners. The lazy hydration and PZM
FileProvider proof lanes should continue against disposable/live smoke
endpoints rather than depending on this migration.
Before any operator follows the rendered migration commands, render the
source-owned cutover packet and attach it to #327/TIN-720:
TCFS_DOWNTIME_WINDOW='YYYY-MM-DD HH:MM-HH:MM TZ' \
TCFS_PREFLIGHT_OWNER='name' \
TCFS_ROLLBACK_OWNER='name' \
TCFS_POSTCUT_SMOKE_OWNER='name' \
TCFS_CONTEXT=honey \
just onprem-cutover-packetThe packet is non-mutating. Its job is to fail fast when the maintenance window or owner assignments are still placeholders, then print the ordered preflight, render, cutover, smoke, and rollback commands for review.
Reasons:
- The backend worker objects are Helm-shaped, but NATS and SeaweedFS are not.
- The OpenTofu modules already model NATS, SeaweedFS, backend workers, and tailnet exposure as separate source-owned concerns.
- The live backing state is honey-local
local-path, so honey/sting mobility requires storage/data planning either way. - A migration plan can preserve the current healthy singleton while building new source-owned objects with explicit storage class, Tailscale exposure, and smoke gates.
Minimum migration gates:
- capture live NATS JetStream and SeaweedFS data inventory;
- use the retained target storage classes for NATS and SeaweedFS instead of
inheriting
local-path; - render or apply retained target PVCs and non-canonical candidate workloads
only through
infra/tofu/environments/onprem; - smoke candidate Tailscale Services with selectors that do not point at the
live
app=natsorapp=seaweedfspods; - run a dry-run/plan that does not collide with live object names;
- cut over only after smoke tests prove NATS, SeaweedFS, and worker connectivity on the new path;
- remove the old live annotations and retained objects through the same source-controlled transition.
Use the read-only preflight before changing either path:
TCFS_CONTEXT=honey just onprem-preflightConfirm which cluster you are talking to before making changes:
kubectl config current-context
kubectl get ns tcfs
helm list -n tcfs
kubectl get sa,deploy,secret -n tcfs | \
rg 'tcfs-backend|sh.helm.release'Expected direct-chart signs:
- Helm release
tcfs-backendexists in namespacetcfs - service account
tcfs-backend-tcfs-backendis present - deployment
tcfs-backend-tcfs-backend-workerexists - Helm release secrets
sh.helm.release.v1.tcfs-backend.vNexist
If instead the namespace contains tcfsd or tcfs-sync-worker, you are on
the OpenTofu-managed path and should not force the Helm recovery flow onto it.
Use the direct backend chart reconciler from this repo:
bash scripts/tcfs-backend-deploy.shDefaults:
- release:
tcfs-backend - namespace:
tcfs - chart:
infra/k8s/charts/tcfs-backend
This path is intentionally narrow: it restores the backend worker deployment, service account, role binding, config map, and Helm release state for the existing direct chart authority. It does not recreate SeaweedFS or NATS.
Useful flags:
bash scripts/tcfs-backend-deploy.sh --dry-run
TCFS_NAMESPACE=tcfs bash scripts/tcfs-backend-deploy.sh \
--set image.tag=v0.12.12Live recovery note, 2026-04-27: the on-prem namespace can contain
Helm-shaped tcfs-backend-tcfs-backend-* objects without Helm release
secrets. In that state a full helm upgrade --install cannot immediately adopt
the existing ConfigMap / Deployment, and it can also fail before adoption if
optional CRDs such as KEDA ScaledObject or Prometheus ServiceMonitor are not
installed.
If the Deployment exists but pod creation is blocked because the service account is missing, restore only the chart-owned RBAC scaffold first:
bash scripts/tcfs-backend-deploy.sh --rbac-only --dry-run
bash scripts/tcfs-backend-deploy.sh --rbac-only
kubectl rollout restart deployment/tcfs-backend-tcfs-backend-worker -n tcfs
kubectl rollout status deployment/tcfs-backend-tcfs-backend-worker -n tcfsThis is a repair path, not a complete Helm adoption. After the worker is healthy, follow the downtime-gated OpenTofu migration path for NATS, SeaweedFS, and canonical tailnet ownership.
helm list -n tcfs
kubectl get sa tcfs-backend-tcfs-backend -n tcfs
kubectl rollout status deployment/tcfs-backend-tcfs-backend-worker -n tcfs
kubectl logs deployment/tcfs-backend-tcfs-backend-worker -n tcfs --tail=50If the service account is missing but the deployment still references
tcfs-backend-tcfs-backend, re-running the direct Helm release should recreate
it because serviceAccount.create defaults to true in the chart values.
Do not retire the preserved Civo PVC tail until the on-prem namespace has all of the following:
- an explicit deployment authority (
tcfs-backendHelm release or an explicit alternative) - live Helm release state in the namespace if Helm owns it
- restored namespace scaffolding for the backend worker path
- an operator decision recorded on whether Civo remains standby state or can be retired
Once those conditions are true, the residual Civo tcfs PVCs stop being a
guess-driven safety blanket and can be evaluated deliberately.