Hermes Agent 883e264c42
Some checks failed
build-and-deploy / test (push) Failing after 2m36s
build-and-deploy / build-deploy (push) Has been skipped
chore: seed gitea-repo-template cookie-cutter
Reusable repo template for homelab services on git.aridgwayweb.com.

Bakes in (all verified live on armistace/wedding-photos):
- Hard commit guard: shared pre-commit hook (git config core.hooksPath
  ~/dev/git-hooks) + master branch protection (push whitelist [armistace],
  merge whitelist [hermes, armistace]).
- Gitea Actions CI (.gitea/workflows/build_push.yml): test + build-deploy,
  persistent remote buildkit cache, registry push, idempotent deploy that
  preserves hand-provisioned Secrets, cluster injection from repo secrets/vars
  via scripts/reconcile-cluster-inject.sh.
- Persistent buildkit cache (ci/buildkit/): single-replica Longhorn backing.
- scripts/reconcile-cluster-inject.sh: reconcile live Secret/ConfigMap from
  Gitea secrets/vars without clobbering hand-provisioned values.
- RUNBOOK.md: handoff-complete ops doc.

Placeholders (<APP> <OWNER> <NS> <KEY_*>) are filled per-service on repo creation.
2026-09-24 11:45:57 +10:00
..

Persistent Buildkit Cache for CI

Dedicated, long-lived buildkit daemon that persists its layer cache on a single-replica Longhorn PVC so repeated CI image builds reuse cache instead of re-building every time. This manifest set is ported verbatim from armistace/wedding-photos — verified live.

Why single-replica Longhorn (not the default SC)

The buildx kubernetes driver's PVC option creates the volume with NO storageClassName — it falls back to the cluster default SC, which here is longhorn at 3 replicas. That would triple every cache byte across disk-constrained nodes. So this uses a dedicated buildkit-single-1r SC (numberOfReplicas: 1, dataLocality: best-effort): the cache is 1× on disk and only replicates if the hosting node actually dies.

Components

Path Purpose
ci/buildkit/01-storageclass.yaml buildkit-single-1r SC (1 replica, best-effort locality)
ci/buildkit/02-statefulset.yaml buildkit StatefulSet, pod schedules off the control plane, listens tcp://0.0.0.0:1234, 10Gi cache PVC
ci/buildkit/03-service.yaml ClusterIP buildkit.gitea-runner.svc:1234
ci/buildkit/04-networkpolicy.yaml restrict buildkit to the gitea-runner namespace only
.gitea/workflows/build_push.yml setup-buildx-action → driver: remote, endpoint=tcp://buildkit.gitea-runner.svc:1234

Ops tooling is NOT in this repo. The disk watchdog + revert/restore scripts live on the control machine (~/.hermes/scripts/buildkit-cache-monitor.sh, buildkit-re-enable.sh) and run from the buildkit-disk-monitor cron. They are homelab ops, not app code.

Apply (one-time, per cluster — NOT per repo)

kubectl apply -f ci/buildkit/01-storageclass.yaml
kubectl apply -f ci/buildkit/02-statefulset.yaml
kubectl apply -f ci/buildkit/03-service.yaml
kubectl apply -f ci/buildkit/04-networkpolicy.yaml
kubectl -n gitea-runner rollout status statefulset/buildkit --timeout=180s

Buildkit is a cluster-wide shared resource — it should be installed once, not once per repo. If your repo doesn't own the cluster, coordinate with whoever does (it lives in the gitea-runner namespace and only needs applying once).