|
|
||
|---|---|---|
| .gitea/workflows | ||
| apps | ||
| bootstrap | ||
| docs | ||
| infra | ||
| scripts | ||
| security | ||
| .gitignore | ||
| .sops.yaml.example | ||
| README.md | ||
| README.md.tmpl | ||
| homelab.yml | ||
| jeannie | ||
| renovate.json | ||
README.md
Homelab Kubernetes Pipeline
This repo bootstraps a hybrid kubeadm cluster and then hands app delivery to Argo CD.
Architecture
The lab is intentionally small but production-shaped:
- a Debian amd64 host runs the kubeadm control plane and local deployment tools
- a Raspberry Pi arm64 node runs selected workloads
- the Debian host also runs the always-on Gitea Docker service outside Kubernetes
- the Debian host keeps a bare GitOps mirror under
/home/jv/git-server/my-homelab-configs.git - a provisioning layer can PXE boot Debian 13 arm64 VMs for Pimox worker templates
- OpenTofu owns the bootstrap layers for cluster, platform, apps, and edge
- Argo CD continuously reconciles Kubernetes manifests from this repo
- a local registry stores the website and demos images built for the worker architecture
- the website image keeps Apache as the public server, uses PHP-FPM with OPcache for dynamic endpoints, and pre-renders read-heavy pages to static HTML during the image build
- SOPS with age is the committed secret-management path for future encrypted Kubernetes secrets
- an OCI jump box provides the public edge path back into the homelab over Tailscale, with nginx caching, gzip, HTTP/2, TLS session reuse, and upstream keepalive
Inventory
The canonical non-secret inventory is homelab.yml. Keep LAN IPs, Tailscale IPs, public hostnames, ports, storage paths, and service ownership there before copying values into scripts, OpenTofu variables, runbooks, or service docs. Secrets and tokens do not belong in that file.
./jeannie reads homelab.yml at startup and exports matching OpenTofu
variables for the edge, provisioning, cluster, platform, and apps stacks. This
keeps values such as the Debian LAN IP, Debian Tailscale IP, RPi LAN IP, OCI IP,
Traefik MetalLB IP, and domain name from drifting across scripts and templates.
scripts/render-docs also renders docs/service-catalog.md
and the static service catalog page at
apps/demos-static/public/homelab-catalog/index.html from the same inventory.
Run ./jeannie up and ./jeannie nuke only from the Debian homelab server. The
script intentionally refuses to run from non-Debian machines so a laptop cannot
accidentally modify the cluster.
Flow
-
bootstrap/provisioning- prepares a Debian server as a PXE and preseed service for arm64 VMs
- serves Debian 13 arm64 netboot assets through TFTP and HTTP
- exposes a PXE rescue menu and HTTP rescue page for manual repair paths
- creates a golden image install path with Kubernetes, containerd, qemu-guest-agent, cloud-init, and storage client packages ready
- is driven by
./jeannie upwhen Pimox is reachable, without changing Orange Pi host networking
-
bootstrap/cluster- creates the kubeadm control plane on the Debian amd64 node
- joins worker nodes such as Raspberry Pi and Pimox Debian arm64 nodes
- configures Calico-compatible pod CIDR
- configures containerd to pull from the in-cluster NodePort registry
- creates retained host directories under
/data/openebs/local
-
bootstrap/platform- installs a minimal Calico deployment through the Tigera operator
- installs NodeLocal DNSCache for node-local DNS query caching
- installs MetalLB for LAN
LoadBalancerservices - installs Traefik as the single Kubernetes ingress entry point
- installs OpenEBS
- creates
openebs-hostpath-retain - installs Argo CD
- installs Kyverno with audit-first baseline Pod Security policies and homelab image signature/SBOM admission controls
- registers the private GitOps repo without storing the SSH private key in Terraform state
-
bootstrap/apps- registers Argo CD Applications from the
applicationsmap - passes the website image produced by the build step into Argo CD as a Kustomize image override
- deploys the tuned website runtime: Apache listens on port 8080, PHP-FPM listens on loopback inside the container, OPcache is enabled, and pre-rendered HTML handles the read-heavy pages
- default apps are
container-registry,website-production,demos-static,heimdall,n8n, andsupply-chain-policy
- registers Argo CD Applications from the
-
bootstrap/edge- connects to the OCI jump box
- uploads nginx, HAProxy, Varnish, and Squid configs
- obtains and renews Let's Encrypt certificates for the configured hostname
- runs the edge cache/proxy chain with Docker Compose
- enables gzip, HTTP/2, TLS session reuse, client keepalive, upstream keepalive, and cache headers for the public website path
Prerequisites
On the Debian host:
- OpenTofu
- Docker with Buildx
- kubeadm, kubelet, kubectl, and containerd pinned to the target Kubernetes
minor, currently
v1.36 - SSH access to worker nodes
- SSH access to the OCI edge host
- enough persistent storage for
/data/openebs/localand Docker's data root
After a clean Debian reinstall, bootstrap the host package set and desktop configuration first:
sudo apt-get update
sudo apt-get install -y git ansible sudo
git clone <repo-url> /home/jv/my-homelab-configs
cd /home/jv/my-homelab-configs
ansible-playbook -K -i bootstrap/host/inventory.ini bootstrap/host/playbook.yml
If you kept the private app-config archive generated from the old Debian host,
pass its generated ansible-vars.yml with -e @/path/to/ansible-vars.yml.
Details are in bootstrap/host/README.md.
The default kubeconfig path is /home/jv/.kube/config. Override it with
KUBECONFIG_PATH or TF_VAR_kubeconfig_path when needed.
Version Discipline
The lab should converge on one Kubernetes minor across the API server and every
kubelet. The Pimox golden-image default is v1.36 so rebuilt arm64 VM workers
match the Debian control plane and Raspberry Pi worker instead of staying on an
older template. Avoid treating a three-minor kubelet skew as normal steady
state: it is only a temporary compatibility window and blocks the next control
plane upgrade.
Check node versions from the Debian host:
./jeannie doctor-versions
The check prints the API server, kubelet versions, kubelet minors, and containerd
runtimes. It exits nonzero when kubelet minors differ from the API server minor
or when containerd major versions are mixed. The Gitea Actions auto path runs
this check before app deploys; kubelet minor drift switches the job to
./jeannie rebuild-cluster so the Pimox template and worker VMs are replaced
from the pinned image. For Pimox workers, prefer rebuilding from the corrected
golden template and rejoining the worker over ad hoc live package upgrades.
Deploying
From the Debian server:
cd ~/my-homelab-configs
./jeannie up
The script first deploys external Gitea to the Debian host with Docker Compose
under /data/homelab-gitea, backed by the HP laptop NVMe, so Git stays
outside the Kubernetes rebuild blast radius. It then detects the
Pimox host at 192.168.100.80 in auto mode. When SSH, qm, and vmbr0 are
available, it applies bootstrap/provisioning, creates or reuses the Debian 13
arm64 template, creates or reuses one worker VM clone, discovers the guest IP
through qemu-guest-agent, and passes that worker into the cluster layer. It then
applies the remaining OpenTofu stacks, refreshes Argo CD apps, waits for the
local registry, builds the website and demos images when their source changed,
pushes them to the registry, recreates pods only after a new image is built, and
applies the edge stack.
Set LAB_PIMOX_PIPELINE=false to skip Pimox automation. Set
LAB_PIMOX_WORKER_COUNT=0 to create or refresh only the template. The pipeline
keeps the template on its configured local storage, creates new worker VM
clones on opi5_ssd by default when that Pimox storage is available,
checks that the Pimox bridge already exists, refuses local as worker clone
storage, and refuses to edit Orange Pi host networking. Keep durable data on
the HP Debian laptop's data-vg or external backups; treat Orange Pi VM disks
as rebuildable worker capacity.
LAB_PIMOX_SKIP_WORKER_INDEXES defaults to an empty value, so the pipeline owns
worker index 1 and VMID 9010 when LAB_PIMOX_WORKER_COUNT=1. Set
LAB_PIMOX_SKIP_WORKER_INDEXES=1 if an existing manually created first worker
must be left untouched, or set LAB_PIMOX_WORKER_COUNT=2 to manage both VMID
9010 and VMID 9011.
OpenWrt firewall VM automation is available as a standalone command because it
attaches to both WAN and LAN bridges. Run ./jeannie openwrt after vmbr1
already exists on the Orange Pi. The pipeline downloads the OpenWrt ARM
SystemReady EFI image, writes basic WAN/LAN/firewall config into the image,
imports it as VM 9100, attaches vmbr0 as WAN and vmbr1 as LAN, and stores
the VM disk on the selected non-local Pimox storage. It leaves the VM stopped
and not enabled for host boot by default. It does not use the Debian Kubernetes
golden-node template for OpenWrt.
The website and demos images default to linux/amd64,linux/arm64 because both
deployments can schedule on either the Debian host or ARM workers. Override with
WEBSITE_IMAGE_PLATFORMS or DEMOS_IMAGE_PLATFORMS only when intentionally
restricting node placement.
Build metadata is written under .lab/ so repeat runs can skip the website
or demos image build when the source hash, platform, image reference, and
registry manifest still match.
The website build is intentionally performed on the Debian homelab path and the
Gitea Actions runner, not on a laptop. The current image uses
php:8.5-fpm-alpine, installs Apache, routes PHP through PHP-FPM, enables
OPcache for immutable image deploys, and runs
apps/website/render-static-pages.sh during the image build. That render step
produces static HTML for /, /cv.php, /blog.php, /demos.php, and
/homelab-tree.php in the static languages. Runtime PHP remains for write and
translation endpoints such as save_idea.php, visitor_ideas.php,
translate.php, and save_lang.php.
Set LAB_GITEA_DEPLOY=false to skip the external Gitea deployment step when the
service is already managed manually. The default Gitea target is
jv@192.168.100.73, install directory /data/homelab-gitea, HTTP port 3000,
and SSH port 32222.
Before continuing into Pimox, cluster, apps, or edge changes, ./jeannie up
runs ./jeannie preflight. The preflight verifies Gitea, Debian Docker
root, Debian Tailscale IP, RPi Docker root fallback state, RPi Tailscale IP,
Pimox worker storage, and OCI edge SSH. Run it directly when changing inventory
or host storage/networking:
./jeannie preflight
If the Debian Docker root check fails after a reinstall, repair it before running the full pipeline:
./jeannie fix-debian-docker-root
When ./jeannie up sees an existing tracked kubeadm control plane but
the Kubernetes API is down, it starts the saved runtime before applying cluster
changes. Use ./jeannie rebuild-cluster only when you intentionally
want to destroy and recreate the cluster.
Validation
Useful checks after a rebuild:
export KUBECONFIG=/home/jv/.kube/config
kubectl get nodes
kubectl -n argocd get applications
kubectl -n container-registry get pods
kubectl -n website-production get pods -o wide
kubectl -n demos-static get pods -o wide
kubectl -n heimdall get pods -o wide
ssh jv@192.168.100.73 'cd /data/homelab-gitea && sudo docker compose ps'
docker info --format '{{.DockerRootDir}}'
df -h / /data /data/openebs/local /data/docker
The website should be reached through the configured public hostname, not the raw OCI IP address, because the Let's Encrypt certificate is issued for the hostname.
Run a read-only health snapshot from the Debian server with:
./jeannie status
./jeannie capacity
./jeannie recover-plan
./jeannie recover-power --dry-run
./jeannie scorecard
./jeannie gitops-status
./jeannie cert-check
./jeannie release-snapshot
./jeannie backup-status
./jeannie synthetic-checks
./jeannie resource-budget
./jeannie artifact-cache status
./jeannie golden-ledger check
./jeannie route-inventory
./jeannie workers list
./jeannie change-journal list
./jeannie map
It reports host memory/disk, systemd services, Docker Compose stacks, Kubernetes health when the API is reachable, Pimox worker VM status, RPi services, Tailscale, key local/public HTTP endpoints, and "what broke" signals such as deployment readiness, pod restarts, node pressure, disk pressure, Traefik 5xx/404 log evidence, and recent Gitea errors.
High-signal Prometheus alerts live in apps/homelab-alerts. They cover node
readiness, pod restart storms, unavailable deployments, node/PVC storage
pressure, Traefik 5xx spikes, website availability, and Gitea/Pi-hole scrape
coverage. Keep this set small so alerts remain actionable.
Application egress is controlled with Kubernetes NetworkPolicies where the lab owns the workload. The website namespace defaults to denied egress and allows only DNS, Redis, and the Debian Ollama endpoint. The security lab namespace is isolated from the cluster and LAN while allowing DNS plus outbound HTTP/HTTPS for controlled practice targets.
capacity is the placement report for the hardware you already have. It
summarizes Debian memory/disk/Docker usage, Kubernetes node and PVC usage,
Pimox VM/storage allocation, RPi Docker/disk state, and ends with placement
guidance for deciding what can safely run next.
recover-plan prints the disaster recovery order and lightweight prerequisite
checks, from Debian and Gitea through DNS, Pimox, Kubernetes, GitOps apps, and
edge verification.
recover-power runs the post-outage recovery sequence in dependency order:
Debian runtime, Gitea, RPi DNS, Pimox workers, Kubernetes, GitOps/apps, and
edge/public checks. Use recover-power --dry-run to preview the sequence.
scorecard rolls up backup readiness, GitOps health, DNS/RPi health, cluster
health, edge health, capacity pressure, security posture, public certificates,
inventory, and docs freshness into a compact pass/warn/fail report.
Report-oriented commands can share the reusable renderer in
scripts/report-render. It reads status TSV rows and prints compact grouped
summaries by default, with details or JSON available for future TUI/web output.
gitops-status prints a focused Argo CD view: application sync and health,
out-of-sync or degraded apps, repository secret presence, recent Argo CD events,
and recent repo-server/application-controller errors.
cert-check verifies public DNS, TLS certificate expiry, public website/Gitea
HTTP status, and DuckDNS-to-OCI edge IP drift from the canonical inventory.
release-snapshot writes a timestamped pre-change report under the homelab
state directory with Git state, Kubernetes/Argo CD/Helm state, workload images,
public URL status, and pointers to the latest Gitea/OpenTofu backups.
backup-status reports freshness for Gitea backups, OpenTofu state backups,
restore drill reports, and repo-managed Pi-hole restore inputs. The main
status cascade includes this as a warning signal.
synthetic-checks runs end-to-end probes for Gitea HTTP/SSH, registry API,
Pi-hole DNS, Traefik LAN HTTP, and temporary Kubernetes DNS/ingress pods. The
registry push/pull smoke test is opt-in with
LAB_SYNTHETIC_REGISTRY_PUSH=true.
resource-budget checks the report-only policy in infra/resource-budgets.yml
against Kubernetes pod requests/limits and Debian host disk/memory signals. Use
it before adding heavier workloads so the HP laptop remains usable.
artifact-cache manages an optional Debian-host cache stack in
infra/artifact-cache with apt-cacher-ng and a Docker Hub registry pull-through
cache. Use artifact-cache instructions for client configuration snippets.
golden-ledger shows or validates infra/pimox/golden-image-ledger.yml, the
reviewable record of template VMID, storage, OS release, Kubernetes pins,
runtime versions, and build metadata for Pimox worker golden images.
route-inventory reports Kubernetes ingress hosts, paths, backend services,
TLS coverage, and visible Uptime Kuma monitor coverage so stale routes and
missing monitors are easier to spot.
workers provides named Kubernetes/Pimox worker lifecycle operations:
list, start, stop, drain, uncordon, recreate-plan, and rebalance.
Use recreate-plan <index> before replacing one broken Pimox worker so the
cluster drain, VM stop, and ./jeannie up path stay ordered.
change-journal lists local journal entries written before risky commands such
as up, apps, deploy-gitea, rpi-services, stop-cluster,
start-cluster, rebuild-cluster, release-snapshot, and nuke. Entries live
under ${HOMELAB_STATE_DIR}/change-journal and record Git revision, branch,
dirty file count, and command details.
map prints a dependency map from the canonical inventory, covering public
entry, GitOps automation, cluster foundation, DNS, observability, and security
learning paths. ./jeannie map --dot emits Graphviz DOT.
Focused doctor commands run narrower read-only checks and print the most likely next step:
./jeannie doctor-edge
./jeannie doctor-gitea
./jeannie doctor-rpi
./jeannie doctor-cluster
RPi Services
The Raspberry Pi runs lightweight always-on services outside Kubernetes from
infra/rpi-services:
- Pi-hole on DNS port
53and web port8081 - Unbound as Pi-hole's first upstream resolver
- Uptime Kuma on port
3001
Deploy or refresh them from the Debian server with:
./jeannie rpi-services
./jeannie up also runs this step unless
LAB_RPI_SERVICES_DEPLOY=false is set. The bootstrap prefers Docker data under
/nvme-storage/docker only when that path is an active writable mount point. If
the NVMe is missing or resets, it rewrites Docker's data-root to
/var/lib/docker and restarts Docker so DNS can still come up from the
SD/root filesystem.
Pi-hole uses Unbound first and public fallback upstreams after that, so DNS can
continue resolving if the Unbound container stops. Repo-managed Pi-hole state
lives in infra/rpi-services/adlists.txt,
infra/rpi-services/local-dns-records.txt,
infra/rpi-services/cname-records.txt, and
infra/rpi-services/static-dhcp-hosts.txt. The bootstrap inserts adlists without
deleting manually added Pi-hole lists, then renders the DNS/CNAME/DHCP files into
a managed dnsmasq config inside the Pi-hole container. If PIHOLE_WEBPASSWORD
is not provided, the bootstrap generates one and keeps it in
/opt/homelab-rpi-services/.env.
Uptime Kuma monitor seeds live in
infra/rpi-services/uptime-kuma-monitors.json and are applied after the first
Kuma user exists.
./jeannie status and ./jeannie doctor-rpi test Pi-hole DNS,
direct Unbound DNS, public fallback resolver readiness, Uptime Kuma HTTP, and
whether Docker is currently using the NVMe path or /var/lib/docker.
Adding Nodes
For Pimox on Orange Pi 5 Plus, ./jeannie up can create the Debian 13 arm64
template and worker VM clones automatically when the Orange Pi storage path is
healthy. Defaults are intentionally tied to the observed host: Pimox SSH host
192.168.100.80, bridge vmbr0, template VMID 9000 on local storage, two
4 GiB worker VMs starting at VMID 9010, worker clone storage opi5_ssd,
and no CPU affinity because this Pimox is pinned to Debian Bullseye. The NVMe
that was previously attached to the Orange Pi now belongs to the HP Debian
laptop as LVM volume group data-vg; the current Orange Pi SSD-backed Pimox
datastore is opi5_ssd and is used for rebuildable VM root disks. Details and
override variables are in
bootstrap/provisioning/README.md.
Worker indexes are stable. Index 1 maps to VMID 9010, node name
pimox-worker-01, and worker key pimox01; index 2 maps to VMID 9011, and
so on. LAB_PIMOX_SKIP_WORKER_INDEXES=1 leaves the first slot unmanaged while
allowing higher indexes to be automated.
Run a full cluster rebuild from the Debian server with:
./jeannie rebuild-cluster
That path preserves external Gitea, rebuilds the Pimox template
with 2 cores and 4 GiB memory, replaces two Pimox worker VMs with 2 cores and
4 GiB memory, and joins those workers to the Kubernetes cluster. CPU affinity is
disabled by default because the Bullseye-pinned Pimox qm does not support it.
The Raspberry Pi is still included as a Kubernetes worker by default; nuke
does not clean it unless you explicitly add it to WORKER_SSH_TARGETS.
To pause the Kubernetes cluster without deleting kubeadm files, OpenTofu state, PV data, or VM disks, run:
./jeannie stop-cluster
This stops kubelet, stops Kubernetes CRI containers with crictl when
available, stops any workers listed in WORKER_SSH_TARGETS, and gracefully
shuts down automated Pimox worker VMs. It intentionally leaves the host
containerd/Docker stack running so external Docker Compose services such as
Gitea stay online. It does not run kubeadm reset, delete CNI files, or remove
state. Use LAB_CLUSTER_STOP_CONTAINERD=true only when you intentionally want
to stop the host container runtime too. Use LAB_CLUSTER_STOP_VM_TIMEOUT_SECONDS
to change the Pimox shutdown timeout, LAB_CLUSTER_STOP_FORCE=false to avoid a
hard qm stop after timeout, and the usual LAB_PIMOX_WORKER_COUNT,
LAB_PIMOX_WORKER_BASE_VMID, and LAB_PIMOX_SKIP_WORKER_INDEXES variables to
select VM workers. Resume the runtime with:
./jeannie start-cluster
To exclude the Raspberry Pi from the Kubernetes cluster, set
LAB_INCLUDE_RASPBERRY_WORKER=false. To manage workers manually instead, add
entries to
bootstrap/cluster/variables.tf or a .tfvars file:
worker_nodes = {
raspberrypi = {
host = "192.168.100.89"
user = "jv"
node_name = "raspberry"
ssh_key_path = "/home/jv/.ssh/id_ed25519"
}
}
Stateful apps currently pin retained local PVs to the debian node. Move or
duplicate those PV manifests when you want storage on another node.
Workload Placement
bootstrap/cluster labels nodes with homelab placement metadata:
node-role.kubernetes.io/worker=workeron every worker sokubectl get nodesshowsworkerinstead of<none>in the ROLES columnhomelab.dev/node-role=control-plane,homelab.dev/storage=local, andhomelab.dev/workload-class=control-planeon the Debian control planehomelab.dev/node-role=edge-app,homelab.dev/storage=local, andhomelab.dev/workload-class=edgeon the Raspberry Pi workerhomelab.dev/node-role=app,homelab.dev/storage=ssd, andhomelab.dev/workload-class=platformon automated Pimox worker clones when those workers are enabled
Override control_plane_node_labels, worker_node_labels,
LAB_RASPBERRY_NODE_LABELS_JSON, or LAB_PIMOX_WORKER_NODE_LABELS_JSON when
the physical layout changes. The current website, demos, and registry manifests
are not moved automatically because the public NodePort path and retained
OpenEBS hostpath PVs are node-local. Move workloads only after their storage and
edge path are ready on the target node. Gitea is outside Kubernetes and is moved
by changing the Raspberry Pi Docker install target instead.
The stateless platform controllers are pinned to Pimox worker nodes through
homelab.dev/workload-class=platform and include hostname topology spread plus
preferred pod anti-affinity so future Argo CD, Kyverno, Prometheus operator, and
kube-state-metrics scheduling does not collapse onto the first worker that joins.
PVC-backed monitoring StatefulSets are intentionally treated separately because
their retained OpenEBS hostpath volumes are node-local. Run
./jeannie move-prometheus-stack-workers from the Debian host to label existing
worker nodes, destroy only the existing prometheus-stack Helm release, delete
its retained PVC/PV objects, and recreate the stack on the worker selector when
you intentionally accept losing that monitoring data. A planned monitoring data
migration should be handled as a separate maintenance task with backup,
delete/recreate or storage migration steps, and post-restore checks.
The older NodePort path is now reserved for special cases such as the local
registry. bootstrap/cluster still contains homelab-tailscale-nodeport
support, but app traffic should normally enter through Traefik's MetalLB
address instead of per-app NodePorts. The cluster stack advertises the LAN
subnet from the configured Tailscale worker so the OCI edge can route to the
Traefik LoadBalancer address:
metallb = {
enabled = true
repository = "https://metallb.github.io/metallb"
version = "0.16.0"
namespace = "metallb-system"
address_pool = ["192.168.100.240-192.168.100.240"]
l2_advertisement_enabled = true
pool_name = "homelab-lan"
}
traefik = {
enabled = true
repository = "https://helm.traefik.io/traefik"
chart_version = "40.2.0"
namespace = "traefik"
load_balancer_ip = "192.168.100.240"
ingress_class = "traefik"
}
tailscale_subnet_routes = {
enabled = true
worker_key = "raspberrypi"
routes = ["192.168.100.0/24"]
}
DuckDNS resolves *.lab2025.duckdns.org to the OCI edge, so public requests for
the service hostnames land on the same edge host. For direct LAN testing, point
LAN DNS, /etc/hosts, or a Tailscale DNS override for app hostnames at
Traefik's address:
192.168.100.240 lab2025.duckdns.org
192.168.100.240 demos.lab2025.duckdns.org
192.168.100.240 heimdall.lab2025.duckdns.org
192.168.100.240 n8n.lab2025.duckdns.org
192.168.100.240 prowlarr.lab2025.duckdns.org
192.168.100.240 sonarr.lab2025.duckdns.org
192.168.100.240 radarr.lab2025.duckdns.org
192.168.100.240 qbittorrent.lab2025.duckdns.org
192.168.100.240 kapowarr.lab2025.duckdns.org
192.168.100.240 suwayomi.lab2025.duckdns.org
192.168.100.240 maintainerr.lab2025.duckdns.org
192.168.100.240 grafana.lab2025.duckdns.org
192.168.100.240 argocd.lab2025.duckdns.org
The edge stack uses Traefik as its backend by default and validates
http://192.168.100.240:80/ before updating the edge containers. If Tailscale
subnet-route approval is not automatic in the tailnet policy, the edge deploy
will fail clearly instead of silently keeping a broken public path.
For ./jeannie nuke, set WORKER_SSH_TARGETS to a space-separated list of
remote SSH targets when more worker nodes exist. Set it to an empty string for a
single-node rebuild.
Adding Platform Tools
Add Helm releases through bootstrap/platform's extra_helm_releases map.
Policy Guardrails
bootstrap/platform installs Kyverno and the upstream baseline Pod Security
policies in Audit mode. This gives the lab policy reports for unsafe workload
settings without blocking existing pods during the first rollout. After reports
are clean, individual policies can be promoted to Enforce in
bootstrap/platform/main.tf.
apps/supply-chain-policy adds a separate Kyverno ImageValidatingPolicy for
the images built by this repo and pushed to the local registry. The policy
targets 192.168.100.73:30500/php-website:* and
192.168.100.73:30500/demos-static:*, mutates admitted pods to image digests,
and audits whether the image has both:
- a valid Cosign signature from the homelab signing key
- a signed SPDX SBOM attestation
The private Cosign key, generated password, and downloaded Cosign binary live in
${XDG_DATA_HOME:-$HOME/.local/share}/homelab/ by default so the signing
identity survives temporary checkout cleanup. Repo-local image state files still
live under .lab/, which is ignored by Git. ./jeannie apps creates or reuses
the persistent Cosign key, publishes the public key into the Kyverno namespace as
homelab-cosign-public-key, builds images with BuildKit SBOM attestations, signs
the pushed images, attaches a signed SPDX SBOM predicate, and verifies both
artifacts before refreshing the workload Applications. The policy starts in
Audit mode so an existing live deployment can recover while the current images
are signed; after ./jeannie apps verifies both images, change validationActions in
apps/supply-chain-policy/local-registry-image-policy.yaml from Audit to
Deny to make it admission-enforcing. Set COSIGN_PASSWORD if you want to
manage the key password externally; otherwise the Debian runner creates
cosign/cosign.password under the homelab state directory for non-interactive
runs. Set COSIGN_KEY_PREFIX, COSIGN_PASSWORD_FILE, or COSIGN_BIN to override
those paths.
DNS Cache
bootstrap/platform installs NodeLocal DNSCache in kube-system with
registry.k8s.io/dns/k8s-dns-node-cache. The default listens on
169.254.20.10 and the kube-dns service IP 10.96.0.10, which keeps the
rollout compatible with the current kube-proxy iptables path without rewriting
kubelet DNS settings across the nodes. Override nodelocal_dns if the service
CIDR or upstream DNS servers change.
MetalLB
MetalLB is present in bootstrap/platform but disabled by default. Enable it
only after reserving a LAN IP range outside DHCP and outside any future OpenWrt
LAN pool:
export TF_VAR_metallb='{
enabled = true
repository = "https://metallb.github.io/metallb"
version = "0.16.0"
namespace = "metallb-system"
address_pool = ["192.168.100.240-192.168.100.250"]
l2_advertisement_enabled = true
pool_name = "homelab-lan"
}'
Traefik uses MetalLB for a LAN LoadBalancer address. App services such as the
website, demos, and Heimdall should be ClusterIP services behind Kubernetes
Ingress objects. The local registry remains a NodePort because the cluster
nodes use it as a pull endpoint. Gitea is not a Kubernetes service; it runs on
the Debian Docker host.
Secrets
Use SOPS with age for secrets that need to live in Git. The Debian host
bootstrap installs age and sops; then run:
./jeannie secrets-init
./jeannie secrets-check
secrets-init creates ~/.config/sops/age/keys.txt if needed and renders
.sops.yaml from .sops.yaml.example with the host's public age recipient.
Commit .sops.yaml, but keep the private age key outside the repo. Operational
notes are in docs/secrets.md.
Edge Services
The OCI jump box runs the public edge path:
nginx -> HAProxy -> Varnish/Squid -> Traefik MetalLB IP
Tailnet access policy is tracked in infra/tailscale/tailnet-policy.hujson.
Validate it before applying with:
./jeannie tailnet-policy-check
The bootstrap/edge stack renders configs from bootstrap/edge/templates and
deploys them to /opt/homelab-edge on the OCI host. Defaults are in
bootstrap/edge/variables.tf; override them through TF_VAR_* or a .tfvars
file when the public host, SSH key, server name, backend Tailscale IP, or
Traefik backend address changes.
The /git/ route is intentionally different from the Kubernetes app routes: it
proxies to Gitea on the Debian host through Tailscale instead of Traefik. This
keeps public read-only source browsing available even when the cluster has been
destroyed.
For the main website, nginx terminates TLS, serves cached HTML and static
assets, compresses text responses with gzip, advertises HTTP/2, reuses TLS
sessions, and keeps upstream connections open to the Traefik backend. The
application side marks pre-rendered HTML as
Cache-Control: public, max-age=300, s-maxage=3600; the edge keeps CSS and JS
under the stronger immutable asset policy. In local measurements, cached
single-resource requests are dominated by the remote TCP/TLS path, while a
reused HTTP/2 asset batch completes in roughly 60 ms after the connection is
established.
Use the configured server_name in the browser, for example
https://lab2025.duckdns.org. A raw OCI IP address will still show a browser
certificate warning because the trusted certificate is issued for the hostname.
The edge stack uses HTTP-01 validation and requests one certificate covering
server_name plus additional_server_names. DuckDNS resolves sub-subdomains
under lab2025.duckdns.org to the same edge IP, so inbound TCP 80 and 443 must
be open before ./jeannie up runs. Set TF_VAR_letsencrypt_email to receive
expiry notices, or leave it empty to register without an email. Set
TF_VAR_enable_letsencrypt=false to keep using the temporary local certificate.
Operational runbooks:
Adding Apps
Add Kubernetes manifests under apps/<name> and register them in
bootstrap/apps's applications map. Argo CD will own sync, pruning, and
self-healing for the app.
The heimdall app is intentionally waited on at the end of ./jeannie apps.
It runs the LinuxServer.io Heimdall dashboard, persists /config on
OpenEBS, and seeds tiles for the website, demo apps, Gitea, Grafana,
Prometheus, Alertmanager, Argo CD, the local registry, and Heimdall itself.
Its manifest only selects Linux nodes; the bound OpenEBS local PV supplies the
node affinity. Do not also pin it to a workload class unless the retained PV is
rebuilt on a node with that label.
Because Heimdall does not support reverse-proxy subfolder hosting cleanly, it
is exposed through the dedicated hostname heimdall.lab2025.duckdns.org rather
than a /heimdall/ path.
The n8n app runs in the n8n namespace with retained OpenEBS storage and is
exposed at https://n8n.lab2025.duckdns.org. It receives
OLLAMA_BASE_URL=http://192.168.100.73:11434 so workflows can call the existing
Debian-host Ollama API for translation prewarming and batch jobs.
The ARR/media stack is intentionally not deployed by the Kubernetes app
pipeline. It is better suited to a Docker Compose stack on the Debian host with
persistent volumes under /data, where large media paths, qBittorrent writes,
and service config directories can survive cluster rebuilds without OpenEBS
node-affinity or PVC migration concerns. The legacy apps/arr-stack manifests
are retained only as migration reference. The active host-level Compose
definition lives under infra/arr-stack and stores persistent data under
/data/arr.
Storage
OpenEBS provides the platform storage provisioner. Stateful Kubernetes apps use
retained local PV paths such as /data/openebs/local/registry; these paths are
intentionally outside kubeadm reset paths so data can survive cluster
destroy/create cycles. Those critical volumes are declared explicitly as
retained local PVs so a rebuilt cluster binds back to the same host paths
instead of creating fresh directories.
For the current lab, the HP Debian laptop's root filesystem is on NVMe, so the
standard Docker root /var/lib/docker is acceptable. OpenEBS retained hostpath
data still lives under /data/openebs/local, and larger service data such as
Gitea, Ollama models, and app volumes should stay under /data.
Gitea
Gitea is external bootstrap infrastructure. It runs on the Debian host as an
always-on Docker Compose service from infra/gitea/docker-compose.yml, not as a
Kubernetes workload. This keeps Git available when the Kubernetes cluster is
destroyed and rebuilt.
The default data path is /data/homelab-gitea/data on the HP laptop NVMe.
Docker may use the standard /var/lib/docker root because / is also on NVMe.
Public source browsing stays available through
https://lab2025.duckdns.org/git/. Registration is disabled and anonymous users
can view public repositories, so the blog can link to code read-only while
writes still require an authenticated Gitea account.
The Debian bare repo remains the GitOps mirror:
/home/jv/git-server/my-homelab-configs.git
Argo CD consumes that Debian mirror through the default gitops_repo_url.
Gitea Actions pushes the main commit into the mirror before running the
selected deploy command.
The platform bootstrap registers the Argo CD repository secret and the SSH host
key for the Debian GitOps mirror. If Argo CD reports
knownhosts: key is unknown after the Debian host was rebuilt or its SSH host
key changed, refresh argocd-ssh-known-hosts-cm in the argocd namespace,
restart argocd-repo-server, and hard-refresh the affected Application.
Deploy or refresh the external Gitea container from the Debian host with:
./jeannie deploy-gitea
Gitea Backups
./jeannie up installs a Debian-host systemd timer named
homelab-gitea-backup.timer. The timer runs daily, SSHes to the configured
Gitea host, executes gitea dump inside the Gitea Docker container, copies the
dump back to Debian, and stores it under /home/jv/backups/gitea. The default
retention is 30 days.
The same install step also creates homelab-gitea-restore-drill.timer. The
monthly drill is non-destructive: it verifies the latest backup ZIP, extracts it
to a temporary directory, records a report under
/home/jv/backups/gitea-restore-drills, and removes the temporary extract. It
does not write into the live Gitea data directory.
Run a manual backup from the Debian server with:
./jeannie backup-gitea
Run the restore drill manually with:
./jeannie drill-gitea-restore
./jeannie drill-restore
./jeannie drill-pihole-restore
drill-restore validates the latest Gitea backup archive, Pi-hole repo-managed
config inputs, the latest OpenTofu state backup archive when present, and the
GitOps rebuild path. It does not replace running services or mutate Kubernetes.
Useful checks:
./jeannie backup-status
systemctl list-timers homelab-gitea-backup.timer
systemctl list-timers homelab-gitea-restore-drill.timer
sudo systemctl start homelab-gitea-backup.service
ls -lh /home/jv/backups/gitea
ls -lh /home/jv/backups/gitea-restore-drills
Gitea Actions
This repo includes a Gitea Actions workflow at
.gitea/workflows/homelab-main.yml. It runs validation on pushes to dev and
main, and deploys only from main. That keeps the promotion path simple:
dev -> ./jeannie validate in Gitea Actions -> main -> Argo CD sync
The workflow targets a repository-scoped Debian host runner with the label
homelab-debian.
Run the same validation locally before promoting:
./jeannie validate
Identity and access audit
Run a read-only access inventory from the Debian server with:
./jeannie access-audit
The audit reports configured SSH targets and key file modes, Git/Gitea remotes, Tailscale ACL policy validation, current Kubernetes authorization, broad cluster-admin bindings, automation service accounts, and Gitea runner status. It does not create users, tokens, kubeconfigs, or keys. Use it to move toward separate identities:
- admin kubeconfig only on the Debian control host
- read-only kubeconfig for dashboards and audits
- separate service accounts for automation
- explicit SSH inventory for Debian, RPi, Pimox, and OCI edge
- repo-managed Tailscale ACLs
The read-only Kubernetes identity is managed in apps/access-control. After
that app syncs, generate a separate read-only kubeconfig on the Debian host:
./jeannie kubeconfig-readonly
The workflow only blocks automatic deploy for external Gitea service
changes: files under infra/gitea/, or edits inside the deploy_gitea,
install_gitea_backup_timer, backup_gitea, or drill_gitea_restore
functions in jeannie. Other changes use HOMELAB_ACTION_COMMAND=auto by
default: Actions runs ./jeannie doctor-versions when the Debian runner already
has a cluster kubeconfig. If node versions are aligned it runs ./jeannie apps;
if kubelet minor drift is detected it runs ./jeannie rebuild-cluster; if no
kubeconfig exists it also runs ./jeannie rebuild-cluster.
Set HOMELAB_ACTION_COMMAND=apps or HOMELAB_ACTION_COMMAND=rebuild-cluster
on the runner to force one path.
./jeannie bootstrap-gitea-repo also registers the Debian host SSH public key
with the Gitea repository and switches the Debian working copy's gitea remote
to ssh://git@192.168.100.73:32222/jv/my-homelab-configs.git. The default key
is /home/jv/.ssh/id_ed25519.pub; set LAB_GITEA_REPO_SSH_KEY_PATH to use a
different Debian-host key, or LAB_GITEA_REPO_SSH_BOOTSTRAP=false to leave SSH
access unchanged. The Actions deploy job uses the checked-out Actions workspace
as the source commit, updates the first available persistent checkout from
HOMELAB_DEPLOY_DIR, /home/jv/my-homelab-configs, or
/home/jv/repos/my-homelab-configs, and otherwise deploys directly from the
Actions workspace. It does not need SSH read access back to Gitea.
Enable Actions for the repository in Gitea, then create a repository-level runner token from:
https://lab2025.duckdns.org/git/jv/my-homelab-configs/settings/actions/runners
Register and start the Debian runner from the Debian server:
cd ~/my-homelab-configs
GITEA_RUNNER_REGISTRATION_TOKEN='<repo-runner-token>' ./jeannie install-gitea-runner
The runner is installed as homelab-gitea-runner.service, runs as user jv, and
uses a host label instead of a Docker job container because deployment needs the
Debian host's Docker, OpenTofu, kubeconfig, SSH keys, and local state.
The deployment job is non-interactive. User jv must be able to run sudo -n true on the Debian host for deployment commands that require sudo.
Useful checks:
systemctl status homelab-gitea-runner.service
journalctl -u homelab-gitea-runner.service -n 100 --no-pager
Renovate
renovate.json defines dependency update rules for Dockerfiles, OpenTofu
providers, Helm chart versions, and the pinned tools used by the Gitea Actions
workflow. Renovate should open reviewable update branches or PRs only; it must
not auto-merge infrastructure changes. Keep app-only dependency updates on the
normal Gitea Actions path, and run ./jeannie up manually on the Debian server
for platform or provisioning updates.
Destructive Rebuilds
./jeannie nuke resets kubeadm, containerd runtime state, CNI files, Calico
links, iptables rules, and local OpenTofu state. It does not delete retained data
under /data/openebs/local.
For multi-node labs, set WORKER_SSH_TARGETS to a space-separated list of SSH
targets. It defaults to an empty string so worker nodes are not cleaned unless
you explicitly include them.
Website App
The website is a PHP app under apps/website. It includes a home page, CV page,
blog page, and demos page, plus a lightweight translation flow backed by Redis,
n8n, and Ollama.
Static language files live in apps/website/lang; en.php and nah.php are
curated source files, with the Nahuatl home page intentionally biased toward as
many Nahuatl words as possible while keeping technical terms understandable.
Unsupported browser languages use the same-origin /translate.php endpoint,
which calls Ollama server-side through OLLAMA_HOST and OLLAMA_MODEL; the
browser never calls the private Ollama IP directly. Redis caches per-string
translation results in the website-production namespace, while n8n is
available for translation prewarming and batch workflow jobs. The default model
is the custom website-translator Ollama model defined in
apps/website/ollama/Modelfile. Generated runtime language JSON is saved
through save_lang.php on the website PVC, and translate.php emits structured
logs for cache hits, misses, Ollama latency, JSON parse failures, and timeouts.
Create or refresh the Ollama model on the Debian server before deploying a website image that points at it:
./jeannie website-translation-model
Ollama must also listen on the Debian host LAN address so Kubernetes pods on
other nodes can reach OLLAMA_HOST=http://192.168.100.73:11434:
./jeannie website-ollama-listen
The Debian host bootstrap also manages Ollama when debian_pc_install_ollama
is true: it installs the service, stores models under /data/ollama/models,
binds the API to the configured LAN listener, and pulls the lightweight
ai_gateway.model. On an existing host, apply just this setup with:
./jeannie ollama-setup
jeannie also has a backstage local AI helper for doctor commands.
It is not a separate CLI action; when ai_gateway.enabled is true in
homelab.yml, doctor commands try the configured Ollama model and silently
skip the helper if Ollama is down. The default lightweight model is:
ollama pull qwen2.5:0.5b
Disable it for low-CPU sessions with:
LAB_BACKSTAGE_BRAIN_ENABLED=false ./jeannie doctor-edge
Build the local homelab knowledge index separately from the main deployment:
./jeannie ai-index
./jeannie ai-check
The index is built from the non-secret source list in
infra/ai/knowledge-sources.txt and defaults to /data/homelab-ai/index.
Doctor commands use it as extra context when it exists, but infrastructure
deployment does not depend on it.
The CV page has two client-side presentation modes:
Elegant: dark, minimal, terminal-inspired styling with a square profile image and light green console text.Fancy: centered circular profile image, cursive orbit text, and a cursor-following portrait rotation effect.
The Demos page is a catalog in the PHP website. The actual demo applications are
served from a separate demos-static artifact under apps/demos-static and are
published through the demos-static Argo CD application. Public traffic reaches
them through the edge path at /demo-apps/.
./jeannie up builds and pushes two independent images:
- a content-hash
php-websitetag generated byjeannieand passed to Argo CD as a Kustomize image override demos-static:latestfromapps/demos-static- Cosign signatures and signed SPDX SBOM attestations for both pushed images
The website manifest keeps the stable base image name php-website:bootstrap.
During bootstrap, jeannie hashes apps/website, builds
<registry>/php-website:src-<hash>, exports that exact reference through
TF_VAR_website_image_ref, and the Argo CD Application applies it through
Kustomize. This keeps the GitOps source generic while the deployed image remains
immutable. The Kyverno supply-chain policy mutates admitted pods to the verified
image digest, so the workload runs the same digest that was signed and attested.
After ./jeannie apps, the live deployment image should be a content-hash tag,
for example 192.168.100.73:30500/php-website:src-.... If it still shows
php-website:latest, Argo CD has not rendered the current Application source.
Check the website-production Application source, sync status, and repository
access before restarting pods.
The first demo, The Client-Side Media Cruncher (Wasm + TS), currently performs
private, browser-only image compression and conversion using native Canvas APIs.
Heavier video conversion, such as MP4 to WebM, should use a Rust core compiled
to WebAssembly with a TypeScript UI so the codec work stays fast and still
avoids backend uploads.
The demos are designed to be local-first so the current cluster can serve them
from any Linux app node without turning either pod into an application server.
The website pod serves the portfolio shell and the demos-static pod serves
static demo bundles; CPU-heavy work runs in the visitor's browser. Because the
deployments can run on either Debian or ARM workers, avoid bundling large ML
models, server-side WebSocket probes, or backend video transcoders into either
image. If those demos become production-grade, lazy load model assets in the
browser or move backend workers to a larger node, such as VMs on the Orange Pi 5
Plus.
Current demo inventory:
- Client-side media cruncher: image conversion/compression with Canvas; future Rust/Wasm codec path for video.
- Internet quality visualizer: live Canvas graph for latency, jitter, and stability using same-origin browser probes; a dedicated WebSocket echo endpoint would be the production version.
- Local log and JSON toolbelt: JSON formatting, JWT decoding, URL parsing, and local text-log filtering.
- Architecture simulator: click-driven load, crash, and auto-scale simulation.
- Offline traveler converter: PWA shell with timezone, currency, and GB/GiB conversions.
- Privacy-first redactor: local image redaction prototype; future onnxruntime-web plus quantized YOLO or face model path.
- Local sentiment sandbox: lightweight local sentiment, keyword, and summary prototype; future Transformers.js/ONNX path.
- Model drift simulator: visual MLOps playground for spikes, corrupted inputs, and retraining.
The Kubernetes deployment uses apps/website/web-app.yaml as a Kustomize base.
Keep TF_VAR_registry_endpoint aligned with the local registry endpoint used by
the app image build and with the image globs in
apps/supply-chain-policy/local-registry-image-policy.yaml.
Keep the .terraform.lock.hcl files committed. They pin provider selections and
make bootstrap behavior reproducible across nodes and rebuilds.