Skip to content

Standardization Goals — Domains, Pkgs, and Auth/IAM Structure

This is a goals/planning doc, not an implementation record. Nothing below is committed or started unless a linked page says otherwise. It exists so the next standardization pass (by anyone, on any account) can start from an agreed target instead of re-deriving it from scratch.

Two independent tracks, each can be worked in any order:

  • Track A — Domain & pkg structure (CoreAPI-native vs. ticketing/johorzoo)
  • Track B — Auth/Zitadel/IAM structure (AuthAPI ↔ CoreAPI ↔ Zitadel ↔ Postgres)

Track A — Domain & pkg standardization

Where we already are (documented, not aspirational)

index.md already establishes the target domain shape (media, distribution, playback as reference; spatial, provisioning as flat/old-pattern; advertising/catalogue "being migrated"). That page is accurate and doesn't need to be redone here — treat it as the source of truth for "what a domain should look like."

What it does not yet cover, because it predates the finding: CoreAPI is really two applications sharing one binary — the CoreAPI-native domains above, and a separately-ported ticketing subsystem (identity, tenancy*, notifications, commerce, analytics, plus catalogue/advertising's ticketing-only content) wired in via setup.go's SetupTicketing(), on its own johorzoo database, still under active development as of the last audit (commits landing within days of that review, not legacy/frozen code).

* Note: tenancy here is CoreAPI-native (orgs/workspaces, documented under domain/tenancy) — the ticketing system's own customer/org concept is separate and lives in identity/ commerce. Don't conflate the two just because "tenancy" sounds generic in both contexts.

Goal 1 — Fix the response-envelope leak (safe, non-breaking, do first)

Problem: two incompatible response shapes exist — pkg/response ({success, data, error}, CoreAPI-native) and pkg/models+pkg/errors ({respCode, respDesc, result}, ticketing — already a live, shipped external contract, confirmed by you, not safe to change the shape itself). The shared Fiber GlobalErrorHandler hardcodes the ticketing shape for any unhandled error app-wide, so a CoreAPI-native endpoint that throws an uncaught error returns the wrong envelope shape to its own clients.

Goal: make the error handler shape-aware of which side of the app the failing route belongs to, without touching either contract. This is the one item on this list that's a pure bug fix, not a structural change — no reason to bundle it with the bigger split below.

Goal 2 — Split the two apps at the Fiber-app boundary ("Option B")

Status (2026-09-07): code done and verified locally, not committed, not pushed, infra not touched. See "Current rollout status" below for exactly what that means and what's still needed to actually ship it.

Problem: main.go's native routes and setup.go's ticketing routes share one fiber.New() instance, so global config (CORS, ErrorHandler, recover middleware) applies to both whether or not that was ever a deliberate choice for the ticketing side.

Goal: two independent fiber.New() instances in one binary/one deploy artifact (not a full service split yet) — each with its own ErrorHandler, each listening on its own port. This directly resolves Goal 1 structurally instead of patching the shared handler, and is a deliberate stepping stone toward a full service split later (ticketing already has its own DB, its own JWT service — it's most of the way there already).

Proposed folder shape for this goal:

CoreAPI/
├── main.go              # thin: infra init → RegisterModels → RunMigrationMain → start both apps
├── models.go             # NEW — RegisterModels(...) call extracted out of main.go
├── coreapi/               # NEW — CoreAPI-native app boundary
│   ├── app.go             #   fiber.New(..., ErrorHandler: NativeErrorHandler)
│   └── routes.go
├── ticketing/             # RENAMED from setup.go — its own app, not a function on main's
│   ├── app.go             #   fiber.New(..., ErrorHandler: TicketingErrorHandler)
│   ├── bootstrap.go
│   └── schedulers.go
├── domain/                # unchanged either way
├── pkg/
│   ├── response/          # scope now matches name — coreapi/ only
│   ├── models/ errors/    # ticketing/ only — keep as-is, still a live contract
│   └── middleware/
│       ├── auth_native.go
│       ├── auth_ticketing.go
│       ├── native_error_handler.go
│       └── ticketing_error_handler.go

Actual result differs slightly from the proposal above — simpler, no new coreapi//ticketing/ folders yet, just the split done in place in main.go: app (native, port from CORE_API_PORT, unchanged) and a new ticketingApp (TICKETING_API_PORT, defaults to 3010 locally — 3001 was tried first but collides with an existing container's host port mapping in this dev stack, so don't reuse that number). Same CORS config shared between both (both are called directly from browsers). setupTicketingService now takes ticketingApp instead of the shared app. middleware.GlobalErrorHandler renamed to TicketingErrorHandler (identical logic, zero shape change — still the live {respCode, respDesc, result} contract); new middleware.NativeErrorHandler emits pkg/response's {success, error} shape for the native side. Verified locally: independent error shapes on each port, /api/v1/* unreachable on the ticketing port and vice versa, and one login cookie authenticating correctly against both apps (JWKS/RS256 validation is shared middleware, not app-specific, so no auth-side changes were needed).

Current rollout status — what's done vs. what ships it

Three separate things have to happen for this goal to actually reach staging/prod, and only the first is done:

  1. CoreAPI code — done, verified locally (see above). Sitting uncommitted in the CoreAPI checkout as of 2026-09-07; not pushed.
  2. Quickstart compose — one line added (3010:3010 port mapping) so local docker compose exposes the new port. Uncommitted.
  3. playtelly-iac (the actual blocker for staging/prod) — not touched at all. This is the part that needs a real decision, not just code.

Why step 3 is the one that matters: right now, playtelly-iac's core-route.yaml (an APISIX/Gateway-API HTTPRoute) sends every path on stag.api.castis.io/api.castis.io to one place — coreapi-service, port 3000. The k8s Service only exposes port 3000, and the Deployment only declares containerPort: 3000. So even once the CoreAPI code above is merged and deployed, nothing changes for staging/prod until the gateway is told about the second port — the ticketing app would be listening on 3010 inside the pod, but nothing outside the pod could ever reach it. That's actually a safety property, not a gap: it means merging and deploying step 1 by itself, on its own, changes nothing observable in staging/prod — the risk is entirely concentrated in step 3, whenever that's done.

What step 3 actually involves, concretely: add containerPort: 3010 to the Deployment, add a second named port to the Service, and add one more HTTPRoute rule so ticketing's paths (/auth, /api/admin, /api/customer, /api/settings, /api/ticketGroups, /api/tags, /payment, and commerce/notifications/analytics' bare /api/* sub-routes) get routed to port 3010 while /api/v1/* (every native domain) and the / fallback (health/static) keep going to port 3000. Gateway API resolves overlapping PathPrefix matches by specificity (longest prefix wins) regardless of rule order, so /api/v1 naturally wins over a broader /api rule without needing careful ordering. No existing client's URL changes either way — TicketAdmin/TicketCMS keep calling the exact same hostnames and paths; only the internal upstream port the gateway forwards to changes.

Staging vs. prod — why they're two separate decisions, not one: staging traffic is disposable — if a path-matching rule is subtly wrong (e.g. a ticketing sub-route that doesn't start with one of the prefixes listed above and falls through to the native port by mistake), the failure mode is a 404/wrong-shape response that's easy to notice and fix, with no real users affected. Prod traffic is live — the same mistake there is a real outage for TicketAdmin/TicketCMS/whatever mobile or web client hits that path. So the sequencing that actually de-risks this is: merge step 1 → apply step 3 to staging only → let it sit and be exercised for real for a while → only then copy the same rule change to prod's core-route.yaml. There's no code difference between the staging and prod versions of this change, just a deliberate time gap and a "did staging actually prove this out" checkpoint in between — this is why the two were treated as separate questions rather than one "ship the split" decision.

Also note: the reverse-proxy/ingress question this goal's plan originally flagged as "confirm before starting" is now answered — playtelly-iac does front CoreAPI (APISIX via Gateway API), confirmed by reading core-route.yaml/service.yml/deployment.yml directly, not assumed.

Goal 3 — Mechanical cleanups (low-risk, independently schedulable)

None of these block or depend on Goal 2 — pick any subset:

Item Action
provisioning's missing config.go 8+ scattered raw os.Getenv calls — consolidate per the documented domain-config convention
catalogue/advertising mixed flat+layered structure Finish the already-documented split into media/content/scheduling; ticketing-only content stays where it is
CProxyClient, ColorbarClient, ELBClient Move from pkg/ into domain/distribution/clients/ — single consumer each, zero risk
StreamerClient Move into domain/distribution/clients/ too, and fix the StreamerNode/distribution.Streamer duplicate-struct problem as part of the move — have playback/scheduling depend on a local interface, not distribution's concrete type
bootstrap/ vs pkg/seed/ Same conceptual job (idempotent seed data), two names/locations — co-locate or rename, don't delete either
internal/db/models.General vs domain/tenancy.General Duplicate struct exists to work around an import-direction problem — fix the direction instead of maintaining the duplicate

Goal 4 — Replace AutoMigrate-in-production with a tracked migration tool

Deliberately out of scope for this pass per your steer — noted here only so it isn't lost. Context if picked up later: AutoMigrate is additive-only (can't drop/rename columns, chokes on new NOT NULL against already-populated rows), and RunMigrationMain() currently log.Fatals the whole process on any migration error — the likely root cause of past staging corruption that got worked around with down -v instead of diagnosed. pkg/db/migrate.go already has two self-heal functions (selfHealMissingNotNullColumns, ensureChannelSIDPartialUniqueIndex) patching around this — evidence of a workaround-on-a-workaround pattern, not a reason to add a third self-heal function next time this class of bug appears. Candidates when this is picked up: golang-migrate/goose (explicit up/down SQL, tracked table) or ariga/atlas (diff GORM structs against live schema, generate reviewable SQL) — see the existing Migrations page for current behavior in detail.


Track B — Auth / Zitadel / IAM structural goals

Based on the full-codebase auth audit done this session (Zitadel Session API v2, RS256/JWKS verification with 5-min cache, HS256 refresh, PSK service-to-service trust, platform_members vs organization_members, one Postgres server / four separate databases with no cross-DB FKs). Existing docs (authapi/index.md, zitadel/zitadel.md, domain/tenancy/organization.md) are accurate for what they cover — the goals below are gaps found during that audit, not corrections to what's already written.

Scope note: when this track is picked back up, it's meant to cover BizConsole and PlatformConsole's real (not documented-assumed) capabilities too, not just AuthAPI/Zitadel/CoreAPI's own code — Goal 5 below (the PlatformConsole doc correction) is one concrete example of what that broader review already turned up; expect more of that kind of correction once BizConsole gets the same treatment.

Goal 1 — Single source of truth for the platform-org sentinel UUID

Problem, confirmed by reading both codebases: the fixed platform-org UUID (00000000-0000-0000-0000-000000000001) is independently hardcoded in both AuthAPI and CoreAPI. Nothing keeps them in sync — a change in one repo silently diverges from the other, with no test or CI check that would catch it. Documented today only as a fact (authapi/index.md:387), not flagged as a divergence risk anywhere.

Goal: pick one owner for the constant (CoreAPI, since it also owns the platform_members table that's the real authorization gate) and have AuthAPI resolve it — either via an env var both services set from the same source, or a lookup call at boot, rather than a second hardcoded literal. Add a comment at both sites cross-referencing the other, minimum viable fix if a full resolve-at-boot mechanism is too heavy.

Goal 2 — Document the AuthAPI ↔ CoreAPI trust boundary explicitly

Problem: two genuinely different trust mechanisms exist side by side — the PSK (COREAPI_AUTHAPI_PSK) for service-to-service calls, and JWKS-based RS256 verification for end-user tokens — and nothing currently documents why there are two, or which one a new integration should use. This is exactly the kind of thing that gets re-discovered by archaeology (like this session did) instead of read off a page.

Goal: one short page (or a section added to authapi/index.md) stating plainly: PSK is for AuthAPI-initiated calls into CoreAPI where there is no end-user session (e.g. org bootstrap/sync); JWKS/RS256 is for verifying a token that originated from an actual user login. Include the CoreDB-sync gap found in this audit (AuthAPI's own org-creation path never calls its UpsertOrgInCoreDB helper — currently masked only because CoreAPI writes its own Organization row directly during CreateOrg, not because the sync path works) as a known-gap callout, not a silent TODO.

Goal 3 — Reconcile or retire dead/mismatched env vars

Problem, confirmed by grep, not assumption: JWT_ISSUER vs AUTHAPI_JWT_ISSUER naming mismatch (one is read, the other looks like it should be but isn't), an unused AUTHAPI_JWT_ACCESS_SECRET, unused Google/Microsoft IDP env vars (dead config for a feature not wired up), and a duplicate AUTHAPI_JWT_ACCESS_EXPIRY line in .env where last-value-wins silently drops the first — easy to lose hours to if someone edits the first occurrence expecting it to take effect.

Goal: one pass through AuthAPI's env-var surface — delete genuinely dead vars, fix the naming mismatch, collapse the duplicate line, and (per Track A's spirit) leave a comment distinguishing "used" from "reserved for a documented future feature" so the next person doesn't have to grep to find out which is which.

Goal 4 — Decide the real status of "future Google/IDP support"

Surfaced during the comprehensive auth walkthrough as an open question, not yet a decision: unused IDP-related env vars exist, but there's no actual OIDC federation code wired to them. Either this is a genuine near-term roadmap item (in which case it's worth a stub doc saying so, so the vars stop looking like dead config) or it's stale (in which case Goal 3 should remove them, not just tidy them).

Goal 5 — Correct PlatformConsole's documented capability vs. reality

Problem: an existing internal doc (referenced during this audit, "Transfer Plan - Mue.xlsx") describes PlatformConsole as an "embedded mirror of OrgConsole" and claims org member/workspace/team/role CRUD is stubbed. Both are confirmed false by reading the actual code: there's no embed/proxy/module-federation mechanism (it's orphaned, unwired copy-paste code), and CoreAPI's tenancy domain has full, real CRUD — the "stubbed" claim is true only of AuthAPI's own separate, legacy, unused-in-practice routes.

Goal: not a code change — a documentation correction. Whoever owns that handover spreadsheet should update it, or it will keep sending the next reader down the wrong path. Flagging here so it isn't lost between sessions.


Suggested order of attack

Superseded by an explicit sequencing decision (2026-09-07) — the order below is the actual plan, not a menu to pick from:

  1. Land the Fiber-app split (Track A Goal 2) cleanly, end to end. Code is done and verified locally (see Goal 2's rollout-status section) — what's left is committing it, then the three-step playtelly-iac rollout (staging first, prod as its own later step) once you're ready to move past local verification.
  2. Then pkg standardization — Track A Goal 3's mechanical cleanups (client-package moves, bootstrap/pkg/seed merge, the General struct duplication) plus actually folding pkg/models/pkg/errors onto pkg/response's shape now that Goal 2 has already separated where each shape is used. Track A Goal 4 (migration tooling) stays explicitly deferred within this phase, not part of it.
  3. Then Track B — the full AuthAPI/Zitadel/CoreAPI structural pass (sentinel UUID, trust-boundary doc, env-var cleanup, IDP-vars decision), broadened to include BizConsole and PlatformConsole's real capabilities per the scope note above, not just the code-level gaps already found.