Skip to content

Consolidation progress — 2026-09-08 → 09

Narrative record of what changed and why. The forward-looking checklist lives in consolidation-plan.md; the rationale behind the goals is in standardization-goals.md.

Goal: collapse CoreAPI's two half-separate applications into one app, one database, without disrupting auth/Zitadel.

Status: Phases 0, 1, 1.5, 2a, 2b, 4a, 4b done. Phase 5 (DB merge) next. All commits local — nothing pushed.

commit phase
439f063 1 structural split, CDN grouping, orphan cleanup
6fd44fa 1.5 HTTP smoke suite
ab4222a 2a pkg/email owns its config; internal/ deleted; dead flags dropped
af999a2 4a ticketing models made migratable; cmd/schema-diff
5b6b873 2b env-default alignment, insecure-default warnings, inventory
08a349d 4b 27 model/schema discrepancies reconciled
ecd897a — timestamptz migration drafted, not applied
bab1b41 0 (Quickstart) seed replacement + healthcheck fix

A direction change worth recording

Earlier work in this effort built the opposite thing: a two-fiber.App split putting ticketing on its own port (feature/ticketing-fiber-app-split). The decision reversed to consolidation, so that branch is abandoned. The response-envelope unification from the same period survives and becomes a prerequisite — one app with two response shapes is not one app.


Phase 0 — baseline

The problem: local dev's ticketing seed had silently forked from staging — 25 tables / 57 general columns vs 28 / 65, frozen since 2026-06-16 while staging was re-dumped 2026-08-28. That fork is why ticket_operating_exception and ticket_group_integration_config did not exist locally, which made every ticketProfile call 500 and left the storefront's product page dead.

Done: replaced with byte-identical copies of the staging schema/data files, loaded by a new 05-ticketing-init.sh. They live in a subdirectory (docker-entrypoint-initdb.d only auto-runs its top level) so the runner controls order and target database — ticketing-data.sql has no \connect, so run directly it executes against POSTGRES_DB (coreapi) and fails.

Also fixed: the compose healthcheck probed http://localhost:3000, but CoreAPI binds 0.0.0.0 (IPv4 only) while localhost resolves to ::1 first inside the container. It therefore always failed, CoreAPI was permanently "unhealthy", and TicketCMS's depends_on: condition: service_healthy never started it. Now 127.0.0.1. The k8s probes were never affected — kubelet hits the pod IP.

before after
tables 25 28
FK constraints 0 17
general columns 57 65
ticketProfile?id=5 500 200
ticketAvailability record not found 200, live POS slots
manual SQL needed 2 tables by hand none

Carried forward: ticketing-data.sql is a production snapshot — 2,367 real customers (name, email, phone, identification_no) and 4,867 orders. Committed knowingly; replacing it with synthetic seed remains open.


Phase 1 — structural tidy

main.go 415 → 40 lines, split into root-level files in package main rather than an app/ subpackage: same readability, zero import churn, no exported-symbol surface to design.

models.go      registerModels()        GORM model registration
infra.go       initInfra()             postgres, redis, nats, mqtt, storage, authapi
routes.go      buildApp() *fiber.App   middleware, static, health, routes
ticketing.go   setupTicketingService   (was setup.go, body unchanged)
schedulers.go  the 5 run*Scheduler helpers

setup.go's body was not touched — it is ~200 lines of order-dependent DI (notifications → tenancy → identity → catalogue → advertising → commerce → analytics). Only the call site moved.

A bug avoided: defer natsclient.Close() sat inside the block that became initInfra(). Left there it would have fired on that function's return — i.e. immediately at startup — silently closing NATS. Hoisted into main().

Also: bootstrap/ → pkg/seed/platform.go (Run → Platform, avoiding a collision with the existing seed.Run); internal/routes.go inlined; scratch/ deleted; pkg/{CProxyClient,ColorbarClient,ELBClient,StreamerClient} → pkg/cdn/{cproxy,colorbar,elb,streamer} with packages renamed to match their directories. pkg/ 33 → 30.


Phase 1.5 — smoke tests (promoted from Phase 3)

Moved ahead of config work because the same checks are needed after every later phase — and hand-run curls had already produced one false result: a malformed URL returned 000 across the board and briefly looked like a total outage.

tests/smoke_test.go, 10 tests / 16 with subtests: boot, auth boundaries (unauthenticated native route and PSK-gated internal route must both 401), login issuing a session CoreAPI accepts, the native /api/v1 surface, the root-mounted ticketing surface, both response envelopes, live POS availability, and the org-scoped email-profile endpoints.

Skipped unless SMOKE_BASE_URL is set, so go test ./... stays green with no stack. Environment-dependent cases (missing seeded admin, unreachable POS) skip rather than fail, so the suite stays honest about what it verified.

Asserting both envelopes is deliberate: it pins today's contract, so when Phase 6 flips ticketing to {success, data, error} the failing test is the signal it worked — and any test that does not flip means a surface was missed.

cd Quickstart && make oc
SMOKE_BASE_URL=http://localhost:3000 go test ./tests/... -v

Not in CI — .gitlab-ci.yml has review / build / update_image_tag and no test stage, and these need a stack. Local gate for now.


Phase 2 — config honesty

2a

pkg/email.NewEmailService took *internal/db/models.General — a 57-field struct — to read 9 fields. That single signature forced two things to exist: a duplicate of domain/tenancy.General under internal/ (so pkg would not import domain) and a ~60-line field-by-field mapper in wire.go to feed the call. The duplicate had also already drifted — 57 fields vs tenancy's 63.

The duplicate was not laziness; it solved a real problem (keeping pkg off domain) with the wrong tool. The idiomatic answer is consumer-defined types: pkg/email now declares its own 9-field SMTPConfig, callers map into it, and it imports neither domain nor internal. wire.go 209 → 161 lines; internal/ deleted outright.

Dead flags removed. cfg.AutoMigrate / cfg.CreateConstraints were read from env, then used in exactly two places — both log.Println — under a comment asserting they "control schema migration on the ticketing DB". Nothing in CoreAPI ever passed ticketing models to a migrator. Deleted rather than left logging "Auto-migration is enabled" while doing nothing; that pair cost real debugging time. Bonus: staging's configmap set TICKETING_AUTO_MIGRATE while the code read AUTO_MIGRATE — misnamed and inert.

2b

Auditing all 82 vars / 107 read sites / 38 files found three whose fallback differed by file, so with the var unset components silently disagreed:

var was now
AUTHAPI_BASE_URL config knew http://authapi:3002; the JWKS validator and AuthAPI client got "" aligned
MINIO_BUCKET media domain used playtelly; storage and seeding used "" aligned
CORE_API_PORT main.go had none (addr became ":"); config said 8080, used by nothing both 3000

Two committed signing-secret defaults now announce themselves at boot via config.WarnInsecureDefaults:

  • TICKETING_JWT_SECRET_KEY = "your-default-secret-key" — signs legacy HS256 customer sessions (ticketcms login, /api/customer/*, /api/orderTicketGroups)
  • CONCIERGE_GUEST_JWT_SECRET = "dev-only-insecure-concierge-guest-secret" — signs concierge guest tokens

Warning, not fatal, deliberately: CoreAPI's prod overlay loads only coreapi-configmap, while staging also loads coreapi-ticketing (43 keys) via patch-env.yml. Prod has no equivalent, so every TICKETING_* var except TICKETING_DB_NAME already falls back in production. Failing fast would convert a latent misconfiguration into an outage.

Scope, not overstated: prod ticketing traffic appears to be served by the legacy stack (ticket-iac) — ticketcms's PlayTelly CI defines only a staging API base URL with the prod block commented out, and there is no prod etiket route in playtelly-iac. So this reads as latent rather than exploited. The prod APISIX core-route does forward all paths to CoreAPI, so the endpoints are reachable regardless. Worth confirming with whoever owns that deployment.

Generated coreapi-env-vars.md — all 82 vars, defaults, read sites.


Phase 4 — ticketing schema → GORM

4a — the models could not migrate at all

19 of 26 failed outright against Postgres. Two MySQL-only constructs inherited from when this was a MySQL application:

type:bigint unsigned AUTO_INCREMENT  ->  syntax error at "unsigned"
type:datetime                        ->  type "datetime" does not exist

96 tags fixed across 10 files → 26/26 migrate cleanly. Notes: type:bigserial is not a valid replacement — it is a CREATE TABLE pseudo-type and fails in ALTER COLUMN TYPE; dropping the tag and letting GORM infer from uint+primaryKey is correct. One tag was malformed (bigint unsigned:int), parsed as an unknown key and silently ignored.

These tags only affect migration — GORM never consults them when querying — so runtime is unchanged.

cmd/schema-diff was added and kept: it AutoMigrates one authoritative model per table into a scratch database so the output can be diffed against the seeded schema. Phase 5 needs the same check after the merge.

Duplicate models found: five tables have two competing definitions. Only identity.OrderTicketGroup genuinely differs (lacks program_cd); commerce's matches the DB. config and schema_migrations have no model.

4b — 27 discrepancies closed

Database treated as source of truth throughout — it is a production snapshot; the structs are what the code merely believes.

  • nullability (20) — 16 columns the DB allows NULL on carried not null; 4 the DB marks NOT NULL were missing it
  • int width (4) — advertising_ticket.placement, report.schedule_day / _day_of_week / _month are integer; uint rendered bigint
  • varchar length (1) — report_attachment.attachment_path is 255, tag said 510
  • missing fields (2) — general.jp_api_key_2 / jp_ag_token_2 exist in the schema but had no struct fields, so CoreAPI could not read the second JohorPay key pair at all

A mistake worth recording: the first pass matched column:updated_at file-wide and dropped not null from six structs that should keep it. The diff caught it immediately — which is the argument for gating on a real diff rather than reasoning about tags. Now scoped per struct.

Result: 0 missing columns, 0 extra columns, 0 type/nullability differences, and 6 timestamp cases deliberately left.

The 6 timestamps — drafted, not applied

25 of 28 tables use timestamptz; three do not. That is the signature of hand-written DDL: pg_dump always emits timestamp with time zone, while a person typing TIMESTAMP gets without. Those three are exactly the hand-appended blocks in the seed — so the seed is internally inconsistent, and matching it would cement an accident.

migration/2026-09-09_ticketing_timestamptz.sql converts them. The explicit AT TIME ZONE 'UTC' is the whole point, demonstrated on a scratch DB:

with    AT TIME ZONE 'UTC':      03:37:40.962023 -> 03:37:40.962023+00
without it, from a +08 session:  03:37:40.962023 -> 2026-08-17 19:37:40+00

Omitting it silently shifts every value 8 hours when run from a Malaysian workstation rather than the server. Verified idempotent; post-migration the schema diffs to zero differences against the models.

⚠️ It fixes an existing database only. A fresh down -v re-seeds from ticketing-schema.sql and reintroduces the inconsistency — the durable fix belongs upstream.


Open upstream requests (playtelly-iac)

  1. ticketing-schema.sql:2247 — CREATE TABLE IF NOT EXISTS ticket_group_integration_config is the only 1 of 28 not johorzoo.-qualified. With the pg_dump preamble at :17 blanking search_path, it fails with "no schema has been selected to create in". Worked around locally via sed in the runner. Staging's own init would fail from scratch too.
  2. The three timestamp without time zone tables — fix in the schema file so fresh volumes stop reproducing it.
  3. ~~Data-file COPY ordering~~ — fixed upstream by f73cb84, which moved ticket_group ahead of its dependents. Note down -v alone does not fix that class of problem; ordering does.

Standing risks

  1. Ticketing schedulers run on every replica with no locking. 5 goroutines, opt-out, no lock or leader election; prod HPA is minReplicas: 2, maxReplicas: 5 and no disable var is set in the IaC. The email job does select-then-mark, so 2–5 pods send the same mail. Likely causing duplicate ticket/report emails in production today. Phase 5 makes pg_advisory_lock trivially available.
  2. Production PII in git, three locations — 2,367 citizens incl. NRIC. Malaysian PDPA applies. History purging and reportability are decisions above a dev session.
  3. Secrets in a plain k8s ConfigMap (coreapi-configmap, prod + staging): DB_PASS, COREAPI_EMAIL_PASSWORD, MINIO_SECRET_KEY, COREAPI_AUTHAPI_PSK.
  4. Live credentials committed in Quickstart/.env and .env.shared-dev.
  5. TLS verification disabled on POS and payment clients.
  6. Per-product email fields have no Go consumer — editable in SpatioViewAdmin, change nothing.