Files
Hermes Agent 883e264c42
build-and-deploy / test (push) Failing after 2m36s
build-and-deploy / build-deploy (push) Has been skipped
chore: seed gitea-repo-template cookie-cutter
Reusable repo template for homelab services on git.aridgwayweb.com.

Bakes in (all verified live on armistace/wedding-photos):
- Hard commit guard: shared pre-commit hook (git config core.hooksPath
  ~/dev/git-hooks) + master branch protection (push whitelist [armistace],
  merge whitelist [hermes, armistace]).
- Gitea Actions CI (.gitea/workflows/build_push.yml): test + build-deploy,
  persistent remote buildkit cache, registry push, idempotent deploy that
  preserves hand-provisioned Secrets, cluster injection from repo secrets/vars
  via scripts/reconcile-cluster-inject.sh.
- Persistent buildkit cache (ci/buildkit/): single-replica Longhorn backing.
- scripts/reconcile-cluster-inject.sh: reconcile live Secret/ConfigMap from
  Gitea secrets/vars without clobbering hand-provisioned values.
- RUNBOOK.md: handoff-complete ops doc.

Placeholders (<APP> <OWNER> <NS> <KEY_*>) are filled per-service on repo creation.
2026-09-24 11:45:57 +10:00

45 lines
2.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Persistent Buildkit Cache for CI
Dedicated, long-lived buildkit daemon that persists its layer cache on a
single-replica Longhorn PVC so repeated CI image builds reuse cache instead of
re-building every time. This manifest set is **ported verbatim from
`armistace/wedding-photos`** — verified live.
## Why single-replica Longhorn (not the default SC)
The buildx **kubernetes driver's PVC option creates the volume with NO
`storageClassName`** — it falls back to the cluster default SC, which here is
`longhorn` at **3 replicas**. That would triple every cache byte across
disk-constrained nodes. So this uses a dedicated `buildkit-single-1r` SC
(`numberOfReplicas: 1`, `dataLocality: best-effort`): the cache is **1× on
disk** and only replicates if the hosting node actually dies.
## Components
| Path | Purpose |
|---|---|
| `ci/buildkit/01-storageclass.yaml` | `buildkit-single-1r` SC (1 replica, best-effort locality) |
| `ci/buildkit/02-statefulset.yaml` | `buildkit` StatefulSet, pod schedules off the control plane, listens `tcp://0.0.0.0:1234`, 10Gi cache PVC |
| `ci/buildkit/03-service.yaml` | ClusterIP `buildkit.gitea-runner.svc:1234` |
| `ci/buildkit/04-networkpolicy.yaml` | restrict buildkit to the gitea-runner namespace only |
| `.gitea/workflows/build_push.yml` | `setup-buildx-action` → `driver: remote`, `endpoint=tcp://buildkit.gitea-runner.svc:1234` |
> **Ops tooling is NOT in this repo.** The disk watchdog + revert/restore
> scripts live on the control machine (`~/.hermes/scripts/buildkit-cache-monitor.sh`,
> `buildkit-re-enable.sh`) and run from the `buildkit-disk-monitor` cron. They are
> homelab ops, not app code.
## Apply (one-time, per cluster — NOT per repo)
```bash
kubectl apply -f ci/buildkit/01-storageclass.yaml
kubectl apply -f ci/buildkit/02-statefulset.yaml
kubectl apply -f ci/buildkit/03-service.yaml
kubectl apply -f ci/buildkit/04-networkpolicy.yaml
kubectl -n gitea-runner rollout status statefulset/buildkit --timeout=180s
```
Buildkit is a **cluster-wide shared resource** — it should be installed once, not
once per repo. If your repo doesn't own the cluster, coordinate with whoever
does (it lives in the `gitea-runner` namespace and only needs applying once).