Standardization Goals — Domains, Pkgs, and Auth/IAM Structure
This is a goals/planning doc, not an implementation record. Nothing below is committed or started unless a linked page says otherwise. It exists so the next standardization pass (by anyone, on any account) can start from an agreed target instead of re-deriving it from scratch.
Two independent tracks, each can be worked in any order:
- Track A — Domain & pkg structure (CoreAPI-native vs. ticketing/johorzoo)
- Track B — Auth/Zitadel/IAM structure (AuthAPI ↔ CoreAPI ↔ Zitadel ↔ Postgres)
Track A — Domain & pkg standardization
Where we already are (documented, not aspirational)
index.md already establishes the target domain shape (media,
distribution, playback as reference; spatial, provisioning as
flat/old-pattern; advertising/catalogue "being migrated"). That page
is accurate and doesn't need to be redone here — treat it as the source
of truth for "what a domain should look like."
What it does not yet cover, because it predates the finding: CoreAPI
is really two applications sharing one binary — the CoreAPI-native
domains above, and a separately-ported ticketing subsystem (identity,
tenancy*, notifications, commerce, analytics, plus
catalogue/advertising's ticketing-only content) wired in via
setup.go's SetupTicketing(), on its own johorzoo database, still
under active development as of the last audit (commits landing within
days of that review, not legacy/frozen code).
* Note: tenancy here is CoreAPI-native (orgs/workspaces, documented
under domain/tenancy) — the ticketing
system's own customer/org concept is separate and lives in identity/
commerce. Don't conflate the two just because "tenancy" sounds generic
in both contexts.
Goal 1 — Fix the response-envelope leak (safe, non-breaking, do first)
Problem: two incompatible response shapes exist —
pkg/response ({success, data, error}, CoreAPI-native) and
pkg/models+pkg/errors ({respCode, respDesc, result}, ticketing —
already a live, shipped external contract, confirmed by you, not
safe to change the shape itself). The shared Fiber GlobalErrorHandler
hardcodes the ticketing shape for any unhandled error app-wide, so a
CoreAPI-native endpoint that throws an uncaught error returns the wrong
envelope shape to its own clients.
Goal: make the error handler shape-aware of which side of the app the failing route belongs to, without touching either contract. This is the one item on this list that's a pure bug fix, not a structural change — no reason to bundle it with the bigger split below.
Goal 2 — Split the two apps at the Fiber-app boundary ("Option B")
Status (2026-09-07): code done and verified locally, not committed, not pushed, infra not touched. See "Current rollout status" below for exactly what that means and what's still needed to actually ship it.
Problem: main.go's native routes and setup.go's ticketing routes
share one fiber.New() instance, so global config (CORS, ErrorHandler,
recover middleware) applies to both whether or not that was ever a
deliberate choice for the ticketing side.
Goal: two independent fiber.New() instances in one binary/one
deploy artifact (not a full service split yet) — each with its own
ErrorHandler, each listening on its own port. This directly resolves
Goal 1 structurally instead of patching the shared handler, and is a
deliberate stepping stone toward a full service split later (ticketing
already has its own DB, its own JWT service — it's most of the way
there already).
Proposed folder shape for this goal:
CoreAPI/
├── main.go # thin: infra init → RegisterModels → RunMigrationMain → start both apps
├── models.go # NEW — RegisterModels(...) call extracted out of main.go
├── coreapi/ # NEW — CoreAPI-native app boundary
│ ├── app.go # fiber.New(..., ErrorHandler: NativeErrorHandler)
│ └── routes.go
├── ticketing/ # RENAMED from setup.go — its own app, not a function on main's
│ ├── app.go # fiber.New(..., ErrorHandler: TicketingErrorHandler)
│ ├── bootstrap.go
│ └── schedulers.go
├── domain/ # unchanged either way
├── pkg/
│ ├── response/ # scope now matches name — coreapi/ only
│ ├── models/ errors/ # ticketing/ only — keep as-is, still a live contract
│ └── middleware/
│ ├── auth_native.go
│ ├── auth_ticketing.go
│ ├── native_error_handler.go
│ └── ticketing_error_handler.go
Actual result differs slightly from the proposal above — simpler, no
new coreapi//ticketing/ folders yet, just the split done in place in
main.go: app (native, port from CORE_API_PORT, unchanged) and a
new ticketingApp (TICKETING_API_PORT, defaults to 3010 locally —
3001 was tried first but collides with an existing container's host
port mapping in this dev stack, so don't reuse that number). Same CORS
config shared between both (both are called directly from browsers).
setupTicketingService now takes ticketingApp instead of the shared
app. middleware.GlobalErrorHandler renamed to TicketingErrorHandler
(identical logic, zero shape change — still the live {respCode,
respDesc, result} contract); new middleware.NativeErrorHandler emits
pkg/response's {success, error} shape for the native side. Verified
locally: independent error shapes on each port, /api/v1/* unreachable
on the ticketing port and vice versa, and one login cookie authenticating
correctly against both apps (JWKS/RS256 validation is shared middleware,
not app-specific, so no auth-side changes were needed).
Current rollout status — what's done vs. what ships it
Three separate things have to happen for this goal to actually reach staging/prod, and only the first is done:
- CoreAPI code — done, verified locally (see above). Sitting
uncommitted in the
CoreAPIcheckout as of 2026-09-07; not pushed. - Quickstart compose — one line added (
3010:3010port mapping) so localdocker composeexposes the new port. Uncommitted. playtelly-iac(the actual blocker for staging/prod) — not touched at all. This is the part that needs a real decision, not just code.
Why step 3 is the one that matters: right now, playtelly-iac's
core-route.yaml (an APISIX/Gateway-API HTTPRoute) sends every path
on stag.api.castis.io/api.castis.io to one place — coreapi-service,
port 3000. The k8s Service only exposes port 3000, and the Deployment
only declares containerPort: 3000. So even once the CoreAPI code above
is merged and deployed, nothing changes for staging/prod until the
gateway is told about the second port — the ticketing app would be
listening on 3010 inside the pod, but nothing outside the pod could ever
reach it. That's actually a safety property, not a gap: it means merging
and deploying step 1 by itself, on its own, changes nothing observable
in staging/prod — the risk is entirely concentrated in step 3, whenever
that's done.
What step 3 actually involves, concretely: add containerPort: 3010
to the Deployment, add a second named port to the Service, and add one
more HTTPRoute rule so ticketing's paths (/auth, /api/admin,
/api/customer, /api/settings, /api/ticketGroups, /api/tags,
/payment, and commerce/notifications/analytics' bare /api/*
sub-routes) get routed to port 3010 while /api/v1/* (every native
domain) and the / fallback (health/static) keep going to port 3000.
Gateway API resolves overlapping PathPrefix matches by specificity
(longest prefix wins) regardless of rule order, so /api/v1 naturally
wins over a broader /api rule without needing careful ordering. No
existing client's URL changes either way — TicketAdmin/TicketCMS keep
calling the exact same hostnames and paths; only the internal upstream
port the gateway forwards to changes.
Staging vs. prod — why they're two separate decisions, not one:
staging traffic is disposable — if a path-matching rule is subtly wrong
(e.g. a ticketing sub-route that doesn't start with one of the prefixes
listed above and falls through to the native port by mistake), the
failure mode is a 404/wrong-shape response that's easy to notice and fix,
with no real users affected. Prod traffic is live — the same mistake
there is a real outage for TicketAdmin/TicketCMS/whatever mobile or web
client hits that path. So the sequencing that actually de-risks this is:
merge step 1 → apply step 3 to staging only → let it sit and be
exercised for real for a while → only then copy the same rule change to
prod's core-route.yaml. There's no code difference between the staging
and prod versions of this change, just a deliberate time gap and a
"did staging actually prove this out" checkpoint in between — this is
why the two were treated as separate questions rather than one "ship
the split" decision.
Also note: the reverse-proxy/ingress question this goal's plan
originally flagged as "confirm before starting" is now answered —
playtelly-iac does front CoreAPI (APISIX via Gateway API), confirmed
by reading core-route.yaml/service.yml/deployment.yml directly,
not assumed.
Goal 3 — Mechanical cleanups (low-risk, independently schedulable)
None of these block or depend on Goal 2 — pick any subset:
| Item | Action |
|---|---|
provisioning's missing config.go |
8+ scattered raw os.Getenv calls — consolidate per the documented domain-config convention |
catalogue/advertising mixed flat+layered structure |
Finish the already-documented split into media/content/scheduling; ticketing-only content stays where it is |
CProxyClient, ColorbarClient, ELBClient |
Move from pkg/ into domain/distribution/clients/ — single consumer each, zero risk |
StreamerClient |
Move into domain/distribution/clients/ too, and fix the StreamerNode/distribution.Streamer duplicate-struct problem as part of the move — have playback/scheduling depend on a local interface, not distribution's concrete type |
bootstrap/ vs pkg/seed/ |
Same conceptual job (idempotent seed data), two names/locations — co-locate or rename, don't delete either |
internal/db/models.General vs domain/tenancy.General |
Duplicate struct exists to work around an import-direction problem — fix the direction instead of maintaining the duplicate |
Goal 4 — Replace AutoMigrate-in-production with a tracked migration tool
Deliberately out of scope for this pass per your steer — noted here
only so it isn't lost. Context if picked up later: AutoMigrate is
additive-only (can't drop/rename columns, chokes on new NOT NULL against
already-populated rows), and RunMigrationMain() currently log.Fatals
the whole process on any migration error — the likely root cause of past
staging corruption that got worked around with down -v instead of
diagnosed. pkg/db/migrate.go already has two self-heal functions
(selfHealMissingNotNullColumns, ensureChannelSIDPartialUniqueIndex)
patching around this — evidence of a workaround-on-a-workaround pattern,
not a reason to add a third self-heal function next time this class of
bug appears. Candidates when this is picked up: golang-migrate/goose
(explicit up/down SQL, tracked table) or ariga/atlas (diff GORM structs
against live schema, generate reviewable SQL) — see the existing
Migrations page for current behavior in detail.
Track B — Auth / Zitadel / IAM structural goals
Based on the full-codebase auth audit done this session (Zitadel Session
API v2, RS256/JWKS verification with 5-min cache, HS256 refresh, PSK
service-to-service trust, platform_members vs organization_members,
one Postgres server / four separate databases with no cross-DB FKs).
Existing docs (authapi/index.md,
zitadel/zitadel.md,
domain/tenancy/organization.md) are
accurate for what they cover — the goals below are gaps found during
that audit, not corrections to what's already written.
Scope note: when this track is picked back up, it's meant to cover BizConsole and PlatformConsole's real (not documented-assumed) capabilities too, not just AuthAPI/Zitadel/CoreAPI's own code — Goal 5 below (the PlatformConsole doc correction) is one concrete example of what that broader review already turned up; expect more of that kind of correction once BizConsole gets the same treatment.
Goal 1 — Single source of truth for the platform-org sentinel UUID
Problem, confirmed by reading both codebases: the fixed platform-org
UUID (00000000-0000-0000-0000-000000000001) is independently hardcoded
in both AuthAPI and CoreAPI. Nothing keeps them in sync — a change in
one repo silently diverges from the other, with no test or CI check that
would catch it. Documented today only as a fact
(authapi/index.md:387), not flagged as a
divergence risk anywhere.
Goal: pick one owner for the constant (CoreAPI, since it also owns
the platform_members table that's the real authorization gate) and have
AuthAPI resolve it — either via an env var both services set from the
same source, or a lookup call at boot, rather than a second hardcoded
literal. Add a comment at both sites cross-referencing the other,
minimum viable fix if a full resolve-at-boot mechanism is too heavy.
Goal 2 — Document the AuthAPI ↔ CoreAPI trust boundary explicitly
Problem: two genuinely different trust mechanisms exist side by side
— the PSK (COREAPI_AUTHAPI_PSK) for service-to-service calls, and
JWKS-based RS256 verification for end-user tokens — and nothing currently
documents why there are two, or which one a new integration should use.
This is exactly the kind of thing that gets re-discovered by archaeology
(like this session did) instead of read off a page.
Goal: one short page (or a section added to
authapi/index.md) stating plainly: PSK is for
AuthAPI-initiated calls into CoreAPI where there is no end-user session
(e.g. org bootstrap/sync); JWKS/RS256 is for verifying a token that
originated from an actual user login. Include the CoreDB-sync gap found
in this audit (AuthAPI's own org-creation path never calls its
UpsertOrgInCoreDB helper — currently masked only because CoreAPI writes
its own Organization row directly during CreateOrg, not because the
sync path works) as a known-gap callout, not a silent TODO.
Goal 3 — Reconcile or retire dead/mismatched env vars
Problem, confirmed by grep, not assumption: JWT_ISSUER vs
AUTHAPI_JWT_ISSUER naming mismatch (one is read, the other looks like
it should be but isn't), an unused AUTHAPI_JWT_ACCESS_SECRET, unused
Google/Microsoft IDP env vars (dead config for a feature not wired up),
and a duplicate AUTHAPI_JWT_ACCESS_EXPIRY line in .env where
last-value-wins silently drops the first — easy to lose hours to if
someone edits the first occurrence expecting it to take effect.
Goal: one pass through AuthAPI's env-var surface — delete genuinely dead vars, fix the naming mismatch, collapse the duplicate line, and (per Track A's spirit) leave a comment distinguishing "used" from "reserved for a documented future feature" so the next person doesn't have to grep to find out which is which.
Goal 4 — Decide the real status of "future Google/IDP support"
Surfaced during the comprehensive auth walkthrough as an open question, not yet a decision: unused IDP-related env vars exist, but there's no actual OIDC federation code wired to them. Either this is a genuine near-term roadmap item (in which case it's worth a stub doc saying so, so the vars stop looking like dead config) or it's stale (in which case Goal 3 should remove them, not just tidy them).
Goal 5 — Correct PlatformConsole's documented capability vs. reality
Problem: an existing internal doc (referenced during this audit, "Transfer Plan - Mue.xlsx") describes PlatformConsole as an "embedded mirror of OrgConsole" and claims org member/workspace/team/role CRUD is stubbed. Both are confirmed false by reading the actual code: there's no embed/proxy/module-federation mechanism (it's orphaned, unwired copy-paste code), and CoreAPI's tenancy domain has full, real CRUD — the "stubbed" claim is true only of AuthAPI's own separate, legacy, unused-in-practice routes.
Goal: not a code change — a documentation correction. Whoever owns that handover spreadsheet should update it, or it will keep sending the next reader down the wrong path. Flagging here so it isn't lost between sessions.
Suggested order of attack
Superseded by an explicit sequencing decision (2026-09-07) — the order below is the actual plan, not a menu to pick from:
- Land the Fiber-app split (Track A Goal 2) cleanly, end to end.
Code is done and verified locally (see Goal 2's rollout-status
section) — what's left is committing it, then the three-step
playtelly-iacrollout (staging first, prod as its own later step) once you're ready to move past local verification. - Then pkg standardization — Track A Goal 3's mechanical cleanups
(client-package moves,
bootstrap/pkg/seedmerge, theGeneralstruct duplication) plus actually foldingpkg/models/pkg/errorsontopkg/response's shape now that Goal 2 has already separated where each shape is used. Track A Goal 4 (migration tooling) stays explicitly deferred within this phase, not part of it. - Then Track B — the full AuthAPI/Zitadel/CoreAPI structural pass (sentinel UUID, trust-boundary doc, env-var cleanup, IDP-vars decision), broadened to include BizConsole and PlatformConsole's real capabilities per the scope note above, not just the code-level gaps already found.