Skip to content

Deployment — staging & production

Why this page exists

Every service's exact deployment mechanics live in a per-service README.md inside playtelly-iac (e.g. services/application/coreapi/README.md, which is excellent and worth reading directly for CoreAPI specifically). Nothing at the docs-site level explained the shared mechanism every one of those READMEs assumes you already know — the two-cluster split, GitOps flow, and how a public hostname actually reaches a pod. That's what this page is.


Three environments, not two

Environment Where it runs How it's brought up Data
Local dev Docker Compose on your machine Quickstart repo, make <target> ephemeral — seeded fresh (see Auth — bootstrap & seeding), down -v wipes it
Staging staging namespace, on the nuc cluster GitOps via playtelly-iac + ArgoCD persistent, real Postgres — this is the environment referenced throughout this doc site whenever "verified live against staging" appears
Production production namespace, split across nuc and do clusters same GitOps mechanism, different overlay persistent, real customer data

Staging and prod are not just "the same compose file with different env vars" — they're an entirely separate deployment system (Kubernetes + Kustomize + ArgoCD), living in its own repo, playtelly-iac.


The two clusters, and why there are two

  • nuc — an on-prem k3s cluster (physical host nuc-a). This is where CoreAPI, AuthAPI, all the consoles/admin frontends, Zitadel, Postgres, SeaweedFS, NATS, Nexus (the Docker registry), and most of the platform actually runs — for both staging and prod. See argocd/apps/prod/nuc/ for the full prod list.
  • do — a DigitalOcean cluster. Hosts the public edge (apisix-do, the CDN/streaming pipeline — elb, gslb, streamer, cproxy) plus a handful of standalone apps that don't depend on CoreAPI (castis-website, promoweb, the documentation site itself, mediasoup, headscale). See argocd/apps/prod/do/.

Why split at all: nuc is cheap/physical but not reliably internet-facing; do gives a stable public IP/TLS front door. So public traffic always lands on do first, then gets forwarded privately to nuc.

How a public request actually reaches a pod on nuc

This is the one mechanism worth understanding before touching any overlay — every app-specific README (CoreAPI's included) assumes you already know this:

https://stag.api.castis.io
        │
        ▼
apisix-do  (public ApisixRoute, synced by the "apisix-do" ArgoCD Application)
        │  plugin: redirect http→https only — does NOT proxy real traffic itself
        │  ApisixUpstream → externalNodes: 100.64.x.x   (a Headscale/Tailscale tailnet address)
        ▼
   [ tailnet — Headscale-managed private mesh network, see playtelly-iac/headscale/ ]
        ▼
apisix-nuc  (private Gateway API HTTPRoute / ApisixRoute, synced by the "apisix-nuc" Application, runs ON nuc)
        │  backendRef → the app's ClusterIP Service
        ▼
   Deployment → pod

TLS for the public hostname is issued by cert-manager (letsencrypt-prod) and bound via an ApisixTls resource on the do side; the nuc side never terminates public TLS itself. The public route's only job is to redirect and forward across the tailnet — the private route does the actual proxying. This exact pattern is why CoreAPI's own README shows two route files per environment (public/route/core-route.yaml + private/core-route.yaml) instead of one.

nexus.castis.io / registry.nexus.castis.io (the Docker registry every image gets pushed to and pulled from) work identically — see playtelly-iac/nexus/README.md for that specific walkthrough if you need to debug the registry itself.


GitOps: how a change actually rolls out

Everything above is defined declaratively in playtelly-iac, one Kustomize app per service:

services/application/<app>/
  base/                 # Deployment, Service, PVC — shared shape
  overlays/prod/         # prod-only ConfigMap, HPA, PDB, replica count, image tag
  overlays/staging/      # staging-only ConfigMap, image tag

Each app also has an ArgoCD Application resource (argocd/apps/{staging,prod/nuc,prod/do}/*.yml) pointing at one of those overlay paths, with:

syncPolicy:
  automated:
    prune: true      # deletes anything in the cluster that isn't in the repo
    selfHeal: true    # reverts any manual kubectl edit back to what the repo says

Practical consequence: the cluster is a mirror of main in playtelly-iac, continuously enforced. A manual kubectl edit to "quickly fix something in prod" gets silently reverted by the next ArgoCD reconcile (default poll interval, or immediately if you don't disable auto-sync first). If you need a manual one-off (like CoreAPI's Zitadel migration Job, which is deliberately not in the overlay's resource list so ArgoCD never touches it), that's the pattern to follow — omit it from kustomization.yaml's resources and apply it by hand with kubectl apply -f.

Rolling out a new image is a one-line change, not a deploy pipeline invocation:

# overlays/staging/kustomization.yaml
images:
  - name: registry.nexus.castis.io/playtelly/coreapi
    newTag: staging-3cfbd0ed   # bump this, push to main, ArgoCD does the rest

Forcing an immediate sync instead of waiting for the poll: argocd app sync coreapi-stag.


⚠️ Secrets are plaintext, committed, in every environment

Both overlays/prod/configmap.yml and overlays/staging/configmap.yml, for every application that has one, are plain Kubernetes ConfigMaps — not Secrets — checked into playtelly-iac in cleartext. This includes, for CoreAPI alone: DB_PASS, MINIO_ACCESS_KEY/SECRET_KEY, COREAPI_AUTHAPI_PSK, COREAPI_EMAIL_PASSWORD, COREAPI_CLIENT_SECRET, COREAPI_REFRESH_TOKEN, COREAPI_SECRETS_ENCRYPTION_KEY (the key that decrypts every stored email/payment/POS credential — see Auth), and more. Same pattern repeats per-service (e.g. nexus's cleanup CronJob ships a hardcoded non/non credential).

Treat every credential in this repo as already compromised by anyone with read access to playtelly-iac. This is a real, standing hazard (also called out in CLAUDE.md §13), not a one-off oversight — rotating to real k8s Secrets is flagged work, not done yet.


Verifying an environment is actually healthy

The pattern (shown here for CoreAPI, identical shape for every other app):

kubectl -n staging get pods -l app=coreapi
kubectl -n production get hpa coreapi-hpa        # prod only — staging has no HPA
kubectl -n production get pdb coreapi             # prod only
kubectl -n staging get certificate coreapi-cert   # READY should be True
kubectl -n staging get httproute coreapi-route
curl -I https://stag.api.castis.io/health         # expect HTTP/2 200

Read-only access to the nuc cluster for exactly this kind of check (and to the real staging Postgres, for schema/data verification) is available — see whoever set up your SSH/kubectl key before assuming you need to ask for new access.


See also

  • Auth, Identity & RBAC — what actually runs once a pod is up, and how bootstrap/seeding populates it on first boot
  • playtelly-iac/services/application/coreapi/README.md — the fullest single example of everything on this page applied to one real app; use it as the template when reading (or writing) any other service's README
  • Make Targets — the local-dev equivalent of this page