# Homelab Kubernetes Pipeline This repo bootstraps a hybrid kubeadm cluster and then hands app delivery to Argo CD. ## Architecture The lab is intentionally small but production-shaped: - a Debian amd64 host runs the kubeadm control plane and local deployment tools - a Raspberry Pi arm64 node runs selected workloads - the Debian host also runs the always-on Gitea Docker service outside Kubernetes - the Debian host keeps a bare GitOps mirror under `/home/jv/git-server/my-homelab-configs.git` - a provisioning layer can PXE boot Debian 13 arm64 VMs for Pimox worker templates - OpenTofu owns the bootstrap layers for cluster, platform, apps, and edge - Argo CD continuously reconciles Kubernetes manifests from this repo - a local registry stores the website and demos images built for the worker architecture - the website image keeps Apache as the public server, uses PHP-FPM with OPcache for dynamic endpoints, and pre-renders read-heavy pages to static HTML during the image build - SOPS with age is the committed secret-management path for future encrypted Kubernetes secrets - an OCI jump box provides the public edge path back into the homelab over Tailscale, with nginx caching, gzip, HTTP/2, TLS session reuse, and upstream keepalive ## Inventory The canonical non-secret inventory is [homelab.yml](homelab.yml). Keep LAN IPs, Tailscale IPs, public hostnames, ports, storage paths, and service ownership there before copying values into scripts, OpenTofu variables, runbooks, or service docs. Secrets and tokens do not belong in that file. `./jeannie` reads `homelab.yml` at startup and exports matching OpenTofu variables for the edge, provisioning, cluster, platform, and apps stacks. This keeps values such as the Debian LAN IP, Debian Tailscale IP, RPi LAN IP, OCI IP, Traefik MetalLB IP, and domain name from drifting across scripts and templates. `scripts/render-docs` also renders [docs/service-catalog.md](docs/service-catalog.md) and the static service catalog page at `apps/demos-static/public/homelab-catalog/index.html` from the same inventory. Run `./jeannie up` and `./jeannie nuke` only from the Debian homelab server. The script intentionally refuses to run from non-Debian machines so a laptop cannot accidentally modify the cluster. ## Flow 1. `bootstrap/provisioning` - prepares a Debian server as a PXE and preseed service for arm64 VMs - serves Debian 13 arm64 netboot assets through TFTP and HTTP - exposes a PXE rescue menu and HTTP rescue page for manual repair paths - creates a golden image install path with Kubernetes, containerd, qemu-guest-agent, cloud-init, and storage client packages ready - is driven by `./jeannie up` when Pimox is reachable, without changing Orange Pi host networking 2. `bootstrap/cluster` - creates the kubeadm control plane on the Debian amd64 node - joins worker nodes such as Raspberry Pi and Pimox Debian arm64 nodes - configures Calico-compatible pod CIDR - configures containerd to pull from the in-cluster NodePort registry - creates retained host directories under `/data/openebs/local` 3. `bootstrap/platform` - installs a minimal Calico deployment through the Tigera operator - installs NodeLocal DNSCache for node-local DNS query caching - installs MetalLB for LAN `LoadBalancer` services - installs Traefik as the single Kubernetes ingress entry point - installs OpenEBS - creates `openebs-hostpath-retain` - installs Argo CD - installs Kyverno with audit-first baseline Pod Security policies and homelab image signature/SBOM admission controls - registers the private GitOps repo without storing the SSH private key in Terraform state 4. `bootstrap/apps` - registers Argo CD Applications from the `applications` map - passes the website image produced by the build step into Argo CD as a Kustomize image override - deploys the tuned website runtime: Apache listens on port 8080, PHP-FPM listens on loopback inside the container, OPcache is enabled, and pre-rendered HTML handles the read-heavy pages - default apps are `container-registry`, `website-production`, `demos-static`, `heimdall`, `n8n`, and `supply-chain-policy` 5. `bootstrap/edge` - connects to the OCI jump box - uploads nginx, HAProxy, Varnish, and Squid configs - obtains and renews Let's Encrypt certificates for the configured hostname - runs the edge cache/proxy chain with Docker Compose - enables gzip, HTTP/2, TLS session reuse, client keepalive, upstream keepalive, and cache headers for the public website path ## Prerequisites On the Debian host: - OpenTofu - Docker with Buildx - kubeadm, kubelet, kubectl, and containerd pinned to the target Kubernetes minor, currently `v1.36` - SSH access to worker nodes - SSH access to the OCI edge host - enough persistent storage for `/data/openebs/local` and Docker's data root After a clean Debian reinstall, bootstrap the host package set and desktop configuration first: ```bash sudo apt-get update sudo apt-get install -y git ansible sudo git clone /home/jv/my-homelab-configs cd /home/jv/my-homelab-configs ansible-playbook -K -i bootstrap/host/inventory.ini bootstrap/host/playbook.yml ``` If you kept the private app-config archive generated from the old Debian host, pass its generated `ansible-vars.yml` with `-e @/path/to/ansible-vars.yml`. Details are in `bootstrap/host/README.md`. The default kubeconfig path is `/home/jv/.kube/config`. Override it with `KUBECONFIG_PATH` or `TF_VAR_kubeconfig_path` when needed. ## Version Discipline The lab should converge on one Kubernetes minor across the API server and every kubelet. The Pimox golden-image default is `v1.36` so rebuilt arm64 VM workers match the Debian control plane and Raspberry Pi worker instead of staying on an older template. Avoid treating a three-minor kubelet skew as normal steady state: it is only a temporary compatibility window and blocks the next control plane upgrade. Check node versions from the Debian host: ```bash ./jeannie doctor-versions ``` The check prints the API server, kubelet versions, kubelet minors, and containerd runtimes. It exits nonzero when kubelet minors differ from the API server minor or when containerd major versions are mixed. The Gitea Actions auto path runs this check before app deploys; kubelet minor drift switches the job to `./jeannie rebuild-cluster` so the Pimox template and worker VMs are replaced from the pinned image. For Pimox workers, prefer rebuilding from the corrected golden template and rejoining the worker over ad hoc live package upgrades. ## Deploying From the Debian server: ```bash cd ~/my-homelab-configs ./jeannie up ``` The script first deploys external Gitea to the Debian host with Docker Compose under `/data/homelab-gitea`, backed by the HP laptop NVMe, so Git stays outside the Kubernetes rebuild blast radius. It then detects the Pimox host at `192.168.100.80` in auto mode. When SSH, `qm`, and `vmbr0` are available, it applies `bootstrap/provisioning`, creates or reuses the Debian 13 arm64 template, creates or reuses one worker VM clone, discovers the guest IP through qemu-guest-agent, and passes that worker into the cluster layer. It then applies the remaining OpenTofu stacks, refreshes Argo CD apps, waits for the local registry, builds the website and demos images when their source changed, pushes them to the registry, recreates pods only after a new image is built, and applies the edge stack. Set `LAB_PIMOX_PIPELINE=false` to skip Pimox automation. Set `LAB_PIMOX_WORKER_COUNT=0` to create or refresh only the template. The pipeline keeps the template on its configured `local` storage, creates new worker VM clones on `opi5_ssd` by default when that Pimox storage is available, checks that the Pimox bridge already exists, refuses `local` as worker clone storage, and refuses to edit Orange Pi host networking. Keep durable data on the HP Debian laptop's `data-vg` or external backups; treat Orange Pi VM disks as rebuildable worker capacity. `LAB_PIMOX_SKIP_WORKER_INDEXES` defaults to an empty value, so the pipeline owns worker index `1` and VMID `9010` when `LAB_PIMOX_WORKER_COUNT=1`. Set `LAB_PIMOX_SKIP_WORKER_INDEXES=1` if an existing manually created first worker must be left untouched, or set `LAB_PIMOX_WORKER_COUNT=2` to manage both VMID `9010` and VMID `9011`. OpenWrt firewall VM automation is available as a standalone command because it attaches to both WAN and LAN bridges. Run `./jeannie openwrt` after `vmbr1` already exists on the Orange Pi. The pipeline downloads the OpenWrt ARM SystemReady EFI image, writes basic WAN/LAN/firewall config into the image, imports it as VM `9100`, attaches `vmbr0` as WAN and `vmbr1` as LAN, and stores the VM disk on the selected non-local Pimox storage. It leaves the VM stopped and not enabled for host boot by default. It does not use the Debian Kubernetes golden-node template for OpenWrt. The website and demos images default to `linux/amd64,linux/arm64` because both deployments can schedule on either the Debian host or ARM workers. Override with `WEBSITE_IMAGE_PLATFORMS` or `DEMOS_IMAGE_PLATFORMS` only when intentionally restricting node placement. Build metadata is written under `.lab/` so repeat runs can skip the website or demos image build when the source hash, platform, image reference, and registry manifest still match. The website build is intentionally performed on the Debian homelab path and the Gitea Actions runner, not on a laptop. The current image uses `php:8.5-fpm-alpine`, installs Apache, routes PHP through PHP-FPM, enables OPcache for immutable image deploys, and runs `apps/website/render-static-pages.sh` during the image build. That render step produces static HTML for `/`, `/cv.php`, `/blog.php`, `/demos.php`, and `/homelab-tree.php` in the static languages. Runtime PHP remains for write and translation endpoints such as `save_idea.php`, `visitor_ideas.php`, `translate.php`, and `save_lang.php`. Set `LAB_GITEA_DEPLOY=false` to skip the external Gitea deployment step when the service is already managed manually. The default Gitea target is `jv@192.168.100.73`, install directory `/data/homelab-gitea`, HTTP port `3000`, and SSH port `32222`. Before continuing into Pimox, cluster, apps, or edge changes, `./jeannie up` runs `./jeannie preflight`. The preflight verifies Gitea, Debian Docker root, Debian Tailscale IP, RPi Docker root fallback state, RPi Tailscale IP, Pimox worker storage, and OCI edge SSH. Run it directly when changing inventory or host storage/networking: ```bash ./jeannie preflight ``` If the Debian Docker root check fails after a reinstall, repair it before running the full pipeline: ```bash ./jeannie fix-debian-docker-root ``` When `./jeannie up` sees an existing tracked kubeadm control plane but the Kubernetes API is down, it starts the saved runtime before applying cluster changes. Use `./jeannie rebuild-cluster` only when you intentionally want to destroy and recreate the cluster. ## Validation Useful checks after a rebuild: ```bash export KUBECONFIG=/home/jv/.kube/config kubectl get nodes kubectl -n argocd get applications kubectl -n container-registry get pods kubectl -n website-production get pods -o wide kubectl -n demos-static get pods -o wide kubectl -n heimdall get pods -o wide ssh jv@192.168.100.73 'cd /data/homelab-gitea && sudo docker compose ps' docker info --format '{{.DockerRootDir}}' df -h / /data /data/openebs/local /data/docker ``` The website should be reached through the configured public hostname, not the raw OCI IP address, because the Let's Encrypt certificate is issued for the hostname. Run a read-only health snapshot from the Debian server with: ```bash ./jeannie status ./jeannie capacity ./jeannie recover-plan ./jeannie recover-power --dry-run ./jeannie gitops-status ./jeannie cert-check ./jeannie release-snapshot ./jeannie backup-status ./jeannie synthetic-checks ./jeannie resource-budget ./jeannie artifact-cache status ./jeannie golden-ledger check ./jeannie route-inventory ./jeannie workers list ./jeannie change-journal list ./jeannie map ``` It reports host memory/disk, systemd services, Docker Compose stacks, Kubernetes health when the API is reachable, Pimox worker VM status, RPi services, Tailscale, key local/public HTTP endpoints, and "what broke" signals such as deployment readiness, pod restarts, node pressure, disk pressure, Traefik 5xx/404 log evidence, and recent Gitea errors. High-signal Prometheus alerts live in `apps/homelab-alerts`. They cover node readiness, pod restart storms, unavailable deployments, node/PVC storage pressure, Traefik 5xx spikes, website availability, and Gitea/Pi-hole scrape coverage. Keep this set small so alerts remain actionable. Application egress is controlled with Kubernetes NetworkPolicies where the lab owns the workload. The website namespace defaults to denied egress and allows only DNS, Redis, and the Debian Ollama endpoint. The security lab namespace is isolated from the cluster and LAN while allowing DNS plus outbound HTTP/HTTPS for controlled practice targets. `capacity` is the placement report for the hardware you already have. It summarizes Debian memory/disk/Docker usage, Kubernetes node and PVC usage, Pimox VM/storage allocation, RPi Docker/disk state, and ends with placement guidance for deciding what can safely run next. `recover-plan` prints the disaster recovery order and lightweight prerequisite checks, from Debian and Gitea through DNS, Pimox, Kubernetes, GitOps apps, and edge verification. `recover-power` runs the post-outage recovery sequence in dependency order: Debian runtime, Gitea, RPi DNS, Pimox workers, Kubernetes, GitOps/apps, and edge/public checks. Use `recover-power --dry-run` to preview the sequence. `gitops-status` prints a focused Argo CD view: application sync and health, out-of-sync or degraded apps, repository secret presence, recent Argo CD events, and recent repo-server/application-controller errors. `cert-check` verifies public DNS, TLS certificate expiry, public website/Gitea HTTP status, and DuckDNS-to-OCI edge IP drift from the canonical inventory. `release-snapshot` writes a timestamped pre-change report under the homelab state directory with Git state, Kubernetes/Argo CD/Helm state, workload images, public URL status, and pointers to the latest Gitea/OpenTofu backups. `backup-status` reports freshness for Gitea backups, OpenTofu state backups, restore drill reports, and repo-managed Pi-hole restore inputs. The main `status` cascade includes this as a warning signal. `synthetic-checks` runs end-to-end probes for Gitea HTTP/SSH, registry API, Pi-hole DNS, Traefik LAN HTTP, and temporary Kubernetes DNS/ingress pods. The registry push/pull smoke test is opt-in with `LAB_SYNTHETIC_REGISTRY_PUSH=true`. `resource-budget` checks the report-only policy in `infra/resource-budgets.yml` against Kubernetes pod requests/limits and Debian host disk/memory signals. Use it before adding heavier workloads so the HP laptop remains usable. `artifact-cache` manages an optional Debian-host cache stack in `infra/artifact-cache` with apt-cacher-ng and a Docker Hub registry pull-through cache. Use `artifact-cache instructions` for client configuration snippets. `golden-ledger` shows or validates `infra/pimox/golden-image-ledger.yml`, the reviewable record of template VMID, storage, OS release, Kubernetes pins, runtime versions, and build metadata for Pimox worker golden images. `route-inventory` reports Kubernetes ingress hosts, paths, backend services, TLS coverage, and visible Uptime Kuma monitor coverage so stale routes and missing monitors are easier to spot. `workers` provides named Kubernetes/Pimox worker lifecycle operations: `list`, `start`, `stop`, `drain`, `uncordon`, `recreate-plan`, and `rebalance`. Use `recreate-plan ` before replacing one broken Pimox worker so the cluster drain, VM stop, and `./jeannie up` path stay ordered. `change-journal` lists local journal entries written before risky commands such as `up`, `apps`, `deploy-gitea`, `rpi-services`, `stop-cluster`, `start-cluster`, `rebuild-cluster`, `release-snapshot`, and `nuke`. Entries live under `${HOMELAB_STATE_DIR}/change-journal` and record Git revision, branch, dirty file count, and command details. `map` prints a dependency map from the canonical inventory, covering public entry, GitOps automation, cluster foundation, DNS, observability, and security learning paths. `./jeannie map --dot` emits Graphviz DOT. Focused doctor commands run narrower read-only checks and print the most likely next step: ```bash ./jeannie doctor-edge ./jeannie doctor-gitea ./jeannie doctor-rpi ./jeannie doctor-cluster ``` ## RPi Services The Raspberry Pi runs lightweight always-on services outside Kubernetes from `infra/rpi-services`: - Pi-hole on DNS port `53` and web port `8081` - Unbound as Pi-hole's first upstream resolver - Uptime Kuma on port `3001` Deploy or refresh them from the Debian server with: ```bash ./jeannie rpi-services ``` `./jeannie up` also runs this step unless `LAB_RPI_SERVICES_DEPLOY=false` is set. The bootstrap prefers Docker data under `/nvme-storage/docker` only when that path is an active writable mount point. If the NVMe is missing or resets, it rewrites Docker's `data-root` to `/var/lib/docker` and restarts Docker so DNS can still come up from the SD/root filesystem. Pi-hole uses Unbound first and public fallback upstreams after that, so DNS can continue resolving if the Unbound container stops. Repo-managed Pi-hole state lives in `infra/rpi-services/adlists.txt`, `infra/rpi-services/local-dns-records.txt`, `infra/rpi-services/cname-records.txt`, and `infra/rpi-services/static-dhcp-hosts.txt`. The bootstrap inserts adlists without deleting manually added Pi-hole lists, then renders the DNS/CNAME/DHCP files into a managed dnsmasq config inside the Pi-hole container. If `PIHOLE_WEBPASSWORD` is not provided, the bootstrap generates one and keeps it in `/opt/homelab-rpi-services/.env`. Uptime Kuma monitor seeds live in `infra/rpi-services/uptime-kuma-monitors.json` and are applied after the first Kuma user exists. `./jeannie status` and `./jeannie doctor-rpi` test Pi-hole DNS, direct Unbound DNS, public fallback resolver readiness, Uptime Kuma HTTP, and whether Docker is currently using the NVMe path or `/var/lib/docker`. ## Adding Nodes For Pimox on Orange Pi 5 Plus, `./jeannie up` can create the Debian 13 arm64 template and worker VM clones automatically when the Orange Pi storage path is healthy. Defaults are intentionally tied to the observed host: Pimox SSH host `192.168.100.80`, bridge `vmbr0`, template VMID `9000` on `local` storage, two 4 GiB worker VMs starting at VMID `9010`, worker clone storage `opi5_ssd`, and no CPU affinity because this Pimox is pinned to Debian Bullseye. The NVMe that was previously attached to the Orange Pi now belongs to the HP Debian laptop as LVM volume group `data-vg`; the current Orange Pi SSD-backed Pimox datastore is `opi5_ssd` and is used for rebuildable VM root disks. Details and override variables are in `bootstrap/provisioning/README.md`. Worker indexes are stable. Index `1` maps to VMID `9010`, node name `pimox-worker-01`, and worker key `pimox01`; index `2` maps to VMID `9011`, and so on. `LAB_PIMOX_SKIP_WORKER_INDEXES=1` leaves the first slot unmanaged while allowing higher indexes to be automated. Run a full cluster rebuild from the Debian server with: ```bash ./jeannie rebuild-cluster ``` That path preserves external Gitea, rebuilds the Pimox template with 2 cores and 4 GiB memory, replaces two Pimox worker VMs with 2 cores and 4 GiB memory, and joins those workers to the Kubernetes cluster. CPU affinity is disabled by default because the Bullseye-pinned Pimox `qm` does not support it. The Raspberry Pi is still included as a Kubernetes worker by default; `nuke` does not clean it unless you explicitly add it to `WORKER_SSH_TARGETS`. To pause the Kubernetes cluster without deleting kubeadm files, OpenTofu state, PV data, or VM disks, run: ```bash ./jeannie stop-cluster ``` This stops `kubelet`, stops Kubernetes CRI containers with `crictl` when available, stops any workers listed in `WORKER_SSH_TARGETS`, and gracefully shuts down automated Pimox worker VMs. It intentionally leaves the host `containerd`/Docker stack running so external Docker Compose services such as Gitea stay online. It does not run `kubeadm reset`, delete CNI files, or remove state. Use `LAB_CLUSTER_STOP_CONTAINERD=true` only when you intentionally want to stop the host container runtime too. Use `LAB_CLUSTER_STOP_VM_TIMEOUT_SECONDS` to change the Pimox shutdown timeout, `LAB_CLUSTER_STOP_FORCE=false` to avoid a hard `qm stop` after timeout, and the usual `LAB_PIMOX_WORKER_COUNT`, `LAB_PIMOX_WORKER_BASE_VMID`, and `LAB_PIMOX_SKIP_WORKER_INDEXES` variables to select VM workers. Resume the runtime with: ```bash ./jeannie start-cluster ``` To exclude the Raspberry Pi from the Kubernetes cluster, set `LAB_INCLUDE_RASPBERRY_WORKER=false`. To manage workers manually instead, add entries to `bootstrap/cluster/variables.tf` or a `.tfvars` file: ```hcl worker_nodes = { raspberrypi = { host = "192.168.100.89" user = "jv" node_name = "raspberry" ssh_key_path = "/home/jv/.ssh/id_ed25519" } } ``` Stateful apps currently pin retained local PVs to the `debian` node. Move or duplicate those PV manifests when you want storage on another node. ## Workload Placement `bootstrap/cluster` labels nodes with homelab placement metadata: - `node-role.kubernetes.io/worker=worker` on every worker so `kubectl get nodes` shows `worker` instead of `` in the ROLES column - `homelab.dev/node-role=control-plane`, `homelab.dev/storage=local`, and `homelab.dev/workload-class=control-plane` on the Debian control plane - `homelab.dev/node-role=edge-app`, `homelab.dev/storage=local`, and `homelab.dev/workload-class=edge` on the Raspberry Pi worker - `homelab.dev/node-role=app`, `homelab.dev/storage=ssd`, and `homelab.dev/workload-class=platform` on automated Pimox worker clones when those workers are enabled Override `control_plane_node_labels`, `worker_node_labels`, `LAB_RASPBERRY_NODE_LABELS_JSON`, or `LAB_PIMOX_WORKER_NODE_LABELS_JSON` when the physical layout changes. The current website, demos, and registry manifests are not moved automatically because the public NodePort path and retained OpenEBS hostpath PVs are node-local. Move workloads only after their storage and edge path are ready on the target node. Gitea is outside Kubernetes and is moved by changing the Raspberry Pi Docker install target instead. The stateless platform controllers are pinned to Pimox worker nodes through `homelab.dev/workload-class=platform` and include hostname topology spread plus preferred pod anti-affinity so future Argo CD, Kyverno, Prometheus operator, and kube-state-metrics scheduling does not collapse onto the first worker that joins. PVC-backed monitoring StatefulSets are intentionally treated separately because their retained OpenEBS hostpath volumes are node-local. Run `./jeannie move-prometheus-stack-workers` from the Debian host to label existing worker nodes, destroy only the existing `prometheus-stack` Helm release, delete its retained PVC/PV objects, and recreate the stack on the worker selector when you intentionally accept losing that monitoring data. A planned monitoring data migration should be handled as a separate maintenance task with backup, delete/recreate or storage migration steps, and post-restore checks. The older NodePort path is now reserved for special cases such as the local registry. `bootstrap/cluster` still contains `homelab-tailscale-nodeport` support, but app traffic should normally enter through Traefik's MetalLB address instead of per-app NodePorts. The cluster stack advertises the LAN subnet from the configured Tailscale worker so the OCI edge can route to the Traefik LoadBalancer address: ```hcl metallb = { enabled = true repository = "https://metallb.github.io/metallb" version = "0.16.0" namespace = "metallb-system" address_pool = ["192.168.100.240-192.168.100.240"] l2_advertisement_enabled = true pool_name = "homelab-lan" } traefik = { enabled = true repository = "https://helm.traefik.io/traefik" chart_version = "40.2.0" namespace = "traefik" load_balancer_ip = "192.168.100.240" ingress_class = "traefik" } tailscale_subnet_routes = { enabled = true worker_key = "raspberrypi" routes = ["192.168.100.0/24"] } ``` DuckDNS resolves `*.lab2025.duckdns.org` to the OCI edge, so public requests for the service hostnames land on the same edge host. For direct LAN testing, point LAN DNS, `/etc/hosts`, or a Tailscale DNS override for app hostnames at Traefik's address: ```text 192.168.100.240 lab2025.duckdns.org 192.168.100.240 demos.lab2025.duckdns.org 192.168.100.240 heimdall.lab2025.duckdns.org 192.168.100.240 n8n.lab2025.duckdns.org 192.168.100.240 prowlarr.lab2025.duckdns.org 192.168.100.240 sonarr.lab2025.duckdns.org 192.168.100.240 radarr.lab2025.duckdns.org 192.168.100.240 qbittorrent.lab2025.duckdns.org 192.168.100.240 kapowarr.lab2025.duckdns.org 192.168.100.240 suwayomi.lab2025.duckdns.org 192.168.100.240 maintainerr.lab2025.duckdns.org 192.168.100.240 grafana.lab2025.duckdns.org 192.168.100.240 argocd.lab2025.duckdns.org ``` The edge stack uses Traefik as its backend by default and validates `http://192.168.100.240:80/` before updating the edge containers. If Tailscale subnet-route approval is not automatic in the tailnet policy, the edge deploy will fail clearly instead of silently keeping a broken public path. For `./jeannie nuke`, set `WORKER_SSH_TARGETS` to a space-separated list of remote SSH targets when more worker nodes exist. Set it to an empty string for a single-node rebuild. ## Adding Platform Tools Add Helm releases through `bootstrap/platform`'s `extra_helm_releases` map. ## Policy Guardrails `bootstrap/platform` installs Kyverno and the upstream baseline Pod Security policies in `Audit` mode. This gives the lab policy reports for unsafe workload settings without blocking existing pods during the first rollout. After reports are clean, individual policies can be promoted to `Enforce` in `bootstrap/platform/main.tf`. `apps/supply-chain-policy` adds a separate Kyverno `ImageValidatingPolicy` for the images built by this repo and pushed to the local registry. The policy targets `192.168.100.73:30500/php-website:*` and `192.168.100.73:30500/demos-static:*`, mutates admitted pods to image digests, and audits whether the image has both: - a valid Cosign signature from the homelab signing key - a signed SPDX SBOM attestation The private Cosign key, generated password, and downloaded Cosign binary live in `${XDG_DATA_HOME:-$HOME/.local/share}/homelab/` by default so the signing identity survives temporary checkout cleanup. Repo-local image state files still live under `.lab/`, which is ignored by Git. `./jeannie apps` creates or reuses the persistent Cosign key, publishes the public key into the Kyverno namespace as `homelab-cosign-public-key`, builds images with BuildKit SBOM attestations, signs the pushed images, attaches a signed SPDX SBOM predicate, and verifies both artifacts before refreshing the workload Applications. The policy starts in `Audit` mode so an existing live deployment can recover while the current images are signed; after `./jeannie apps` verifies both images, change `validationActions` in `apps/supply-chain-policy/local-registry-image-policy.yaml` from `Audit` to `Deny` to make it admission-enforcing. Set `COSIGN_PASSWORD` if you want to manage the key password externally; otherwise the Debian runner creates `cosign/cosign.password` under the homelab state directory for non-interactive runs. Set `COSIGN_KEY_PREFIX`, `COSIGN_PASSWORD_FILE`, or `COSIGN_BIN` to override those paths. ## DNS Cache `bootstrap/platform` installs NodeLocal DNSCache in `kube-system` with `registry.k8s.io/dns/k8s-dns-node-cache`. The default listens on `169.254.20.10` and the kube-dns service IP `10.96.0.10`, which keeps the rollout compatible with the current kube-proxy iptables path without rewriting kubelet DNS settings across the nodes. Override `nodelocal_dns` if the service CIDR or upstream DNS servers change. ## MetalLB MetalLB is present in `bootstrap/platform` but disabled by default. Enable it only after reserving a LAN IP range outside DHCP and outside any future OpenWrt LAN pool: ```bash export TF_VAR_metallb='{ enabled = true repository = "https://metallb.github.io/metallb" version = "0.16.0" namespace = "metallb-system" address_pool = ["192.168.100.240-192.168.100.250"] l2_advertisement_enabled = true pool_name = "homelab-lan" }' ``` Traefik uses MetalLB for a LAN `LoadBalancer` address. App services such as the website, demos, and Heimdall should be `ClusterIP` services behind Kubernetes Ingress objects. The local registry remains a `NodePort` because the cluster nodes use it as a pull endpoint. Gitea is not a Kubernetes service; it runs on the Debian Docker host. ## Secrets Use SOPS with age for secrets that need to live in Git. The Debian host bootstrap installs `age` and `sops`; then run: ```bash ./jeannie secrets-init ./jeannie secrets-check ``` `secrets-init` creates `~/.config/sops/age/keys.txt` if needed and renders `.sops.yaml` from `.sops.yaml.example` with the host's public age recipient. Commit `.sops.yaml`, but keep the private age key outside the repo. Operational notes are in `docs/secrets.md`. ## Edge Services The OCI jump box runs the public edge path: ```text nginx -> HAProxy -> Varnish/Squid -> Traefik MetalLB IP ``` Tailnet access policy is tracked in `infra/tailscale/tailnet-policy.hujson`. Validate it before applying with: ```bash ./jeannie tailnet-policy-check ``` The `bootstrap/edge` stack renders configs from `bootstrap/edge/templates` and deploys them to `/opt/homelab-edge` on the OCI host. Defaults are in `bootstrap/edge/variables.tf`; override them through `TF_VAR_*` or a `.tfvars` file when the public host, SSH key, server name, backend Tailscale IP, or Traefik backend address changes. The `/git/` route is intentionally different from the Kubernetes app routes: it proxies to Gitea on the Debian host through Tailscale instead of Traefik. This keeps public read-only source browsing available even when the cluster has been destroyed. For the main website, nginx terminates TLS, serves cached HTML and static assets, compresses text responses with gzip, advertises HTTP/2, reuses TLS sessions, and keeps upstream connections open to the Traefik backend. The application side marks pre-rendered HTML as `Cache-Control: public, max-age=300, s-maxage=3600`; the edge keeps CSS and JS under the stronger immutable asset policy. In local measurements, cached single-resource requests are dominated by the remote TCP/TLS path, while a reused HTTP/2 asset batch completes in roughly 60 ms after the connection is established. Use the configured `server_name` in the browser, for example `https://lab2025.duckdns.org`. A raw OCI IP address will still show a browser certificate warning because the trusted certificate is issued for the hostname. The edge stack uses HTTP-01 validation and requests one certificate covering `server_name` plus `additional_server_names`. DuckDNS resolves sub-subdomains under `lab2025.duckdns.org` to the same edge IP, so inbound TCP 80 and 443 must be open before `./jeannie up` runs. Set `TF_VAR_letsencrypt_email` to receive expiry notices, or leave it empty to register without an email. Set `TF_VAR_enable_letsencrypt=false` to keep using the temporary local certificate. Operational runbooks: - [Edge failures](docs/runbooks/edge-failures.md) - [Gitea failures](docs/runbooks/gitea-failures.md) - [Cluster stop/start failures](docs/runbooks/cluster-stop-start-failures.md) ## Adding Apps Add Kubernetes manifests under `apps/` and register them in `bootstrap/apps`'s `applications` map. Argo CD will own sync, pruning, and self-healing for the app. The `heimdall` app is intentionally waited on at the end of `./jeannie apps`. It runs the LinuxServer.io Heimdall dashboard, persists `/config` on OpenEBS, and seeds tiles for the website, demo apps, Gitea, Grafana, Prometheus, Alertmanager, Argo CD, the local registry, and Heimdall itself. Its manifest only selects Linux nodes; the bound OpenEBS local PV supplies the node affinity. Do not also pin it to a workload class unless the retained PV is rebuilt on a node with that label. Because Heimdall does not support reverse-proxy subfolder hosting cleanly, it is exposed through the dedicated hostname `heimdall.lab2025.duckdns.org` rather than a `/heimdall/` path. The `n8n` app runs in the `n8n` namespace with retained OpenEBS storage and is exposed at `https://n8n.lab2025.duckdns.org`. It receives `OLLAMA_BASE_URL=http://192.168.100.73:11434` so workflows can call the existing Debian-host Ollama API for translation prewarming and batch jobs. The ARR/media stack is intentionally not deployed by the Kubernetes app pipeline. It is better suited to a Docker Compose stack on the Debian host with persistent volumes under `/data`, where large media paths, qBittorrent writes, and service config directories can survive cluster rebuilds without OpenEBS node-affinity or PVC migration concerns. The legacy `apps/arr-stack` manifests are retained only as migration reference. The active host-level Compose definition lives under `infra/arr-stack` and stores persistent data under `/data/arr`. ## Storage OpenEBS provides the platform storage provisioner. Stateful Kubernetes apps use retained local PV paths such as `/data/openebs/local/registry`; these paths are intentionally outside kubeadm reset paths so data can survive cluster destroy/create cycles. Those critical volumes are declared explicitly as retained local PVs so a rebuilt cluster binds back to the same host paths instead of creating fresh directories. For the current lab, the HP Debian laptop's root filesystem is on NVMe, so the standard Docker root `/var/lib/docker` is acceptable. OpenEBS retained hostpath data still lives under `/data/openebs/local`, and larger service data such as Gitea, Ollama models, and app volumes should stay under `/data`. ## Gitea Gitea is external bootstrap infrastructure. It runs on the Debian host as an always-on Docker Compose service from `infra/gitea/docker-compose.yml`, not as a Kubernetes workload. This keeps Git available when the Kubernetes cluster is destroyed and rebuilt. The default data path is `/data/homelab-gitea/data` on the HP laptop NVMe. Docker may use the standard `/var/lib/docker` root because `/` is also on NVMe. Public source browsing stays available through `https://lab2025.duckdns.org/git/`. Registration is disabled and anonymous users can view public repositories, so the blog can link to code read-only while writes still require an authenticated Gitea account. The Debian bare repo remains the GitOps mirror: ```text /home/jv/git-server/my-homelab-configs.git ``` Argo CD consumes that Debian mirror through the default `gitops_repo_url`. Gitea Actions pushes the `main` commit into the mirror before running the selected deploy command. The platform bootstrap registers the Argo CD repository secret and the SSH host key for the Debian GitOps mirror. If Argo CD reports `knownhosts: key is unknown` after the Debian host was rebuilt or its SSH host key changed, refresh `argocd-ssh-known-hosts-cm` in the `argocd` namespace, restart `argocd-repo-server`, and hard-refresh the affected Application. Deploy or refresh the external Gitea container from the Debian host with: ```bash ./jeannie deploy-gitea ``` ## Gitea Backups `./jeannie up` installs a Debian-host systemd timer named `homelab-gitea-backup.timer`. The timer runs daily, SSHes to the configured Gitea host, executes `gitea dump` inside the Gitea Docker container, copies the dump back to Debian, and stores it under `/home/jv/backups/gitea`. The default retention is 30 days. The same install step also creates `homelab-gitea-restore-drill.timer`. The monthly drill is non-destructive: it verifies the latest backup ZIP, extracts it to a temporary directory, records a report under `/home/jv/backups/gitea-restore-drills`, and removes the temporary extract. It does not write into the live Gitea data directory. Run a manual backup from the Debian server with: ```bash ./jeannie backup-gitea ``` Run the restore drill manually with: ```bash ./jeannie drill-gitea-restore ./jeannie drill-restore ./jeannie drill-pihole-restore ``` `drill-restore` validates the latest Gitea backup archive, Pi-hole repo-managed config inputs, the latest OpenTofu state backup archive when present, and the GitOps rebuild path. It does not replace running services or mutate Kubernetes. Useful checks: ```bash ./jeannie backup-status systemctl list-timers homelab-gitea-backup.timer systemctl list-timers homelab-gitea-restore-drill.timer sudo systemctl start homelab-gitea-backup.service ls -lh /home/jv/backups/gitea ls -lh /home/jv/backups/gitea-restore-drills ``` ## Gitea Actions This repo includes a Gitea Actions workflow at `.gitea/workflows/homelab-main.yml`. It runs validation on pushes to `dev` and `main`, and deploys only from `main`. That keeps the promotion path simple: ```text dev -> ./jeannie validate in Gitea Actions -> main -> Argo CD sync ``` The workflow targets a repository-scoped Debian host runner with the label `homelab-debian`. Run the same validation locally before promoting: ```bash ./jeannie validate ``` ## Identity and access audit Run a read-only access inventory from the Debian server with: ```bash ./jeannie access-audit ``` The audit reports configured SSH targets and key file modes, Git/Gitea remotes, Tailscale ACL policy validation, current Kubernetes authorization, broad cluster-admin bindings, automation service accounts, and Gitea runner status. It does not create users, tokens, kubeconfigs, or keys. Use it to move toward separate identities: - admin kubeconfig only on the Debian control host - read-only kubeconfig for dashboards and audits - separate service accounts for automation - explicit SSH inventory for Debian, RPi, Pimox, and OCI edge - repo-managed Tailscale ACLs The read-only Kubernetes identity is managed in `apps/access-control`. After that app syncs, generate a separate read-only kubeconfig on the Debian host: ```bash ./jeannie kubeconfig-readonly ``` The workflow only blocks automatic deploy for external Gitea service changes: files under `infra/gitea/`, or edits inside the `deploy_gitea`, `install_gitea_backup_timer`, `backup_gitea`, or `drill_gitea_restore` functions in `jeannie`. Other changes use `HOMELAB_ACTION_COMMAND=auto` by default: Actions runs `./jeannie doctor-versions` when the Debian runner already has a cluster kubeconfig. If node versions are aligned it runs `./jeannie apps`; if kubelet minor drift is detected it runs `./jeannie rebuild-cluster`; if no kubeconfig exists it also runs `./jeannie rebuild-cluster`. Set `HOMELAB_ACTION_COMMAND=apps` or `HOMELAB_ACTION_COMMAND=rebuild-cluster` on the runner to force one path. `./jeannie bootstrap-gitea-repo` also registers the Debian host SSH public key with the Gitea repository and switches the Debian working copy's `gitea` remote to `ssh://git@192.168.100.73:32222/jv/my-homelab-configs.git`. The default key is `/home/jv/.ssh/id_ed25519.pub`; set `LAB_GITEA_REPO_SSH_KEY_PATH` to use a different Debian-host key, or `LAB_GITEA_REPO_SSH_BOOTSTRAP=false` to leave SSH access unchanged. The Actions deploy job uses the checked-out Actions workspace as the source commit, updates the first available persistent checkout from `HOMELAB_DEPLOY_DIR`, `/home/jv/my-homelab-configs`, or `/home/jv/repos/my-homelab-configs`, and otherwise deploys directly from the Actions workspace. It does not need SSH read access back to Gitea. Enable Actions for the repository in Gitea, then create a repository-level runner token from: ```text https://lab2025.duckdns.org/git/jv/my-homelab-configs/settings/actions/runners ``` Register and start the Debian runner from the Debian server: ```bash cd ~/my-homelab-configs GITEA_RUNNER_REGISTRATION_TOKEN='' ./jeannie install-gitea-runner ``` The runner is installed as `homelab-gitea-runner.service`, runs as user `jv`, and uses a host label instead of a Docker job container because deployment needs the Debian host's Docker, OpenTofu, kubeconfig, SSH keys, and local state. The deployment job is non-interactive. User `jv` must be able to run `sudo -n true` on the Debian host for deployment commands that require sudo. Useful checks: ```bash systemctl status homelab-gitea-runner.service journalctl -u homelab-gitea-runner.service -n 100 --no-pager ``` ## Renovate `renovate.json` defines dependency update rules for Dockerfiles, OpenTofu providers, Helm chart versions, and the pinned tools used by the Gitea Actions workflow. Renovate should open reviewable update branches or PRs only; it must not auto-merge infrastructure changes. Keep app-only dependency updates on the normal Gitea Actions path, and run `./jeannie up` manually on the Debian server for platform or provisioning updates. ## Destructive Rebuilds `./jeannie nuke` resets kubeadm, containerd runtime state, CNI files, Calico links, iptables rules, and local OpenTofu state. It does not delete retained data under `/data/openebs/local`. For multi-node labs, set `WORKER_SSH_TARGETS` to a space-separated list of SSH targets. It defaults to an empty string so worker nodes are not cleaned unless you explicitly include them. ## Website App The website is a PHP app under `apps/website`. It includes a home page, CV page, blog page, and demos page, plus a lightweight translation flow backed by Redis, n8n, and Ollama. Static language files live in `apps/website/lang`; `en.php` and `nah.php` are curated source files, with the Nahuatl home page intentionally biased toward as many Nahuatl words as possible while keeping technical terms understandable. Unsupported browser languages use the same-origin `/translate.php` endpoint, which calls Ollama server-side through `OLLAMA_HOST` and `OLLAMA_MODEL`; the browser never calls the private Ollama IP directly. Redis caches per-string translation results in the `website-production` namespace, while n8n is available for translation prewarming and batch workflow jobs. The default model is the custom `website-translator` Ollama model defined in `apps/website/ollama/Modelfile`. Generated runtime language JSON is saved through `save_lang.php` on the website PVC, and `translate.php` emits structured logs for cache hits, misses, Ollama latency, JSON parse failures, and timeouts. Create or refresh the Ollama model on the Debian server before deploying a website image that points at it: ```bash ./jeannie website-translation-model ``` Ollama must also listen on the Debian host LAN address so Kubernetes pods on other nodes can reach `OLLAMA_HOST=http://192.168.100.73:11434`: ```bash ./jeannie website-ollama-listen ``` The Debian host bootstrap also manages Ollama when `debian_pc_install_ollama` is true: it installs the service, stores models under `/data/ollama/models`, binds the API to the configured LAN listener, and pulls the lightweight `ai_gateway.model`. On an existing host, apply just this setup with: ```bash ./jeannie ollama-setup ``` `jeannie` also has a backstage local AI helper for doctor commands. It is not a separate CLI action; when `ai_gateway.enabled` is true in `homelab.yml`, doctor commands try the configured Ollama model and silently skip the helper if Ollama is down. The default lightweight model is: ```bash ollama pull qwen2.5:0.5b ``` Disable it for low-CPU sessions with: ```bash LAB_BACKSTAGE_BRAIN_ENABLED=false ./jeannie doctor-edge ``` Build the local homelab knowledge index separately from the main deployment: ```bash ./jeannie ai-index ./jeannie ai-check ``` The index is built from the non-secret source list in `infra/ai/knowledge-sources.txt` and defaults to `/data/homelab-ai/index`. Doctor commands use it as extra context when it exists, but infrastructure deployment does not depend on it. The CV page has two client-side presentation modes: - `Elegant`: dark, minimal, terminal-inspired styling with a square profile image and light green console text. - `Fancy`: centered circular profile image, cursive orbit text, and a cursor-following portrait rotation effect. The Demos page is a catalog in the PHP website. The actual demo applications are served from a separate `demos-static` artifact under `apps/demos-static` and are published through the `demos-static` Argo CD application. Public traffic reaches them through the edge path at `/demo-apps/`. `./jeannie up` builds and pushes two independent images: - a content-hash `php-website` tag generated by `jeannie` and passed to Argo CD as a Kustomize image override - `demos-static:latest` from `apps/demos-static` - Cosign signatures and signed SPDX SBOM attestations for both pushed images The website manifest keeps the stable base image name `php-website:bootstrap`. During bootstrap, `jeannie` hashes `apps/website`, builds `/php-website:src-`, exports that exact reference through `TF_VAR_website_image_ref`, and the Argo CD Application applies it through Kustomize. This keeps the GitOps source generic while the deployed image remains immutable. The Kyverno supply-chain policy mutates admitted pods to the verified image digest, so the workload runs the same digest that was signed and attested. After `./jeannie apps`, the live deployment image should be a content-hash tag, for example `192.168.100.73:30500/php-website:src-...`. If it still shows `php-website:latest`, Argo CD has not rendered the current Application source. Check the `website-production` Application source, sync status, and repository access before restarting pods. The first demo, `The Client-Side Media Cruncher (Wasm + TS)`, currently performs private, browser-only image compression and conversion using native Canvas APIs. Heavier video conversion, such as MP4 to WebM, should use a Rust core compiled to WebAssembly with a TypeScript UI so the codec work stays fast and still avoids backend uploads. The demos are designed to be local-first so the current cluster can serve them from any Linux app node without turning either pod into an application server. The website pod serves the portfolio shell and the `demos-static` pod serves static demo bundles; CPU-heavy work runs in the visitor's browser. Because the deployments can run on either Debian or ARM workers, avoid bundling large ML models, server-side WebSocket probes, or backend video transcoders into either image. If those demos become production-grade, lazy load model assets in the browser or move backend workers to a larger node, such as VMs on the Orange Pi 5 Plus. Current demo inventory: - Client-side media cruncher: image conversion/compression with Canvas; future Rust/Wasm codec path for video. - Internet quality visualizer: live Canvas graph for latency, jitter, and stability using same-origin browser probes; a dedicated WebSocket echo endpoint would be the production version. - Local log and JSON toolbelt: JSON formatting, JWT decoding, URL parsing, and local text-log filtering. - Architecture simulator: click-driven load, crash, and auto-scale simulation. - Offline traveler converter: PWA shell with timezone, currency, and GB/GiB conversions. - Privacy-first redactor: local image redaction prototype; future onnxruntime-web plus quantized YOLO or face model path. - Local sentiment sandbox: lightweight local sentiment, keyword, and summary prototype; future Transformers.js/ONNX path. - Model drift simulator: visual MLOps playground for spikes, corrupted inputs, and retraining. The Kubernetes deployment uses `apps/website/web-app.yaml` as a Kustomize base. Keep `TF_VAR_registry_endpoint` aligned with the local registry endpoint used by the app image build and with the image globs in `apps/supply-chain-policy/local-registry-image-policy.yaml`. Keep the `.terraform.lock.hcl` files committed. They pin provider selections and make bootstrap behavior reproducible across nodes and rebuilds.