Consolidation progress — 2026-09-08 → 09
Narrative record of what changed and why. The forward-looking checklist lives in consolidation-plan.md; the rationale behind the goals is in standardization-goals.md.
Goal: collapse CoreAPI's two half-separate applications into one app, one database, without disrupting auth/Zitadel.
Status: Phases 0, 1, 1.5, 2a, 2b, 4a, 4b done. Phase 5 (DB merge) next. All commits local — nothing pushed.
| commit | phase | |
|---|---|---|
439f063 |
1 | structural split, CDN grouping, orphan cleanup |
6fd44fa |
1.5 | HTTP smoke suite |
ab4222a |
2a | pkg/email owns its config; internal/ deleted; dead flags dropped |
af999a2 |
4a | ticketing models made migratable; cmd/schema-diff |
5b6b873 |
2b | env-default alignment, insecure-default warnings, inventory |
08a349d |
4b | 27 model/schema discrepancies reconciled |
ecd897a |
— | timestamptz migration drafted, not applied |
bab1b41 |
0 | (Quickstart) seed replacement + healthcheck fix |
A direction change worth recording
Earlier work in this effort built the opposite thing: a two-fiber.App split
putting ticketing on its own port (feature/ticketing-fiber-app-split). The
decision reversed to consolidation, so that branch is abandoned. The response-envelope
unification from the same period survives and becomes a prerequisite — one app
with two response shapes is not one app.
Phase 0 — baseline
The problem: local dev's ticketing seed had silently forked from staging —
25 tables / 57 general columns vs 28 / 65, frozen since 2026-06-16 while
staging was re-dumped 2026-08-28. That fork is why ticket_operating_exception
and ticket_group_integration_config did not exist locally, which made every
ticketProfile call 500 and left the storefront's product page dead.
Done: replaced with byte-identical copies of the staging schema/data files,
loaded by a new 05-ticketing-init.sh. They live in a subdirectory
(docker-entrypoint-initdb.d only auto-runs its top level) so the runner
controls order and target database — ticketing-data.sql has no \connect, so
run directly it executes against POSTGRES_DB (coreapi) and fails.
Also fixed: the compose healthcheck probed http://localhost:3000, but
CoreAPI binds 0.0.0.0 (IPv4 only) while localhost resolves to ::1 first
inside the container. It therefore always failed, CoreAPI was permanently
"unhealthy", and TicketCMS's depends_on: condition: service_healthy never
started it. Now 127.0.0.1. The k8s probes were never affected — kubelet hits
the pod IP.
| before | after | |
|---|---|---|
| tables | 25 | 28 |
| FK constraints | 0 | 17 |
general columns |
57 | 65 |
ticketProfile?id=5 |
500 | 200 |
ticketAvailability |
record not found |
200, live POS slots |
| manual SQL needed | 2 tables by hand | none |
Carried forward: ticketing-data.sql is a production snapshot — 2,367 real
customers (name, email, phone, identification_no) and 4,867 orders. Committed
knowingly; replacing it with synthetic seed remains open.
Phase 1 — structural tidy
main.go 415 → 40 lines, split into root-level files in package main
rather than an app/ subpackage: same readability, zero import churn, no
exported-symbol surface to design.
models.go registerModels() GORM model registration
infra.go initInfra() postgres, redis, nats, mqtt, storage, authapi
routes.go buildApp() *fiber.App middleware, static, health, routes
ticketing.go setupTicketingService (was setup.go, body unchanged)
schedulers.go the 5 run*Scheduler helpers
setup.go's body was not touched — it is ~200 lines of order-dependent DI
(notifications → tenancy → identity → catalogue → advertising → commerce →
analytics). Only the call site moved.
A bug avoided: defer natsclient.Close() sat inside the block that became
initInfra(). Left there it would have fired on that function's return — i.e.
immediately at startup — silently closing NATS. Hoisted into main().
Also: bootstrap/ → pkg/seed/platform.go (Run → Platform, avoiding a
collision with the existing seed.Run); internal/routes.go inlined;
scratch/ deleted; pkg/{CProxyClient,ColorbarClient,ELBClient,StreamerClient}
→ pkg/cdn/{cproxy,colorbar,elb,streamer} with packages renamed to match their
directories. pkg/ 33 → 30.
Phase 1.5 — smoke tests (promoted from Phase 3)
Moved ahead of config work because the same checks are needed after every later
phase — and hand-run curls had already produced one false result: a malformed
URL returned 000 across the board and briefly looked like a total outage.
tests/smoke_test.go, 10 tests / 16 with subtests: boot, auth boundaries
(unauthenticated native route and PSK-gated internal route must both 401), login
issuing a session CoreAPI accepts, the native /api/v1 surface, the
root-mounted ticketing surface, both response envelopes, live POS
availability, and the org-scoped email-profile endpoints.
Skipped unless SMOKE_BASE_URL is set, so go test ./... stays green with no
stack. Environment-dependent cases (missing seeded admin, unreachable POS) skip
rather than fail, so the suite stays honest about what it verified.
Asserting both envelopes is deliberate: it pins today's contract, so when
Phase 6 flips ticketing to {success, data, error} the failing test is the
signal it worked — and any test that does not flip means a surface was missed.
cd Quickstart && make oc
SMOKE_BASE_URL=http://localhost:3000 go test ./tests/... -v
Not in CI — .gitlab-ci.yml has review / build / update_image_tag and no test
stage, and these need a stack. Local gate for now.
Phase 2 — config honesty
2a
pkg/email.NewEmailService took *internal/db/models.General — a 57-field
struct — to read 9 fields. That single signature forced two things to exist: a
duplicate of domain/tenancy.General under internal/ (so pkg would not
import domain) and a ~60-line field-by-field mapper in wire.go to feed
the call. The duplicate had also already drifted — 57 fields vs tenancy's 63.
The duplicate was not laziness; it solved a real problem (keeping pkg off
domain) with the wrong tool. The idiomatic answer is consumer-defined types:
pkg/email now declares its own 9-field SMTPConfig, callers map into it, and
it imports neither domain nor internal. wire.go 209 → 161 lines;
internal/ deleted outright.
Dead flags removed. cfg.AutoMigrate / cfg.CreateConstraints were read
from env, then used in exactly two places — both log.Println — under a comment
asserting they "control schema migration on the ticketing DB". Nothing in
CoreAPI ever passed ticketing models to a migrator. Deleted rather than left
logging "Auto-migration is enabled" while doing nothing; that pair cost real
debugging time. Bonus: staging's configmap set TICKETING_AUTO_MIGRATE while
the code read AUTO_MIGRATE — misnamed and inert.
2b
Auditing all 82 vars / 107 read sites / 38 files found three whose fallback differed by file, so with the var unset components silently disagreed:
| var | was | now |
|---|---|---|
AUTHAPI_BASE_URL |
config knew http://authapi:3002; the JWKS validator and AuthAPI client got "" |
aligned |
MINIO_BUCKET |
media domain used playtelly; storage and seeding used "" |
aligned |
CORE_API_PORT |
main.go had none (addr became ":"); config said 8080, used by nothing |
both 3000 |
Two committed signing-secret defaults now announce themselves at boot via
config.WarnInsecureDefaults:
TICKETING_JWT_SECRET_KEY = "your-default-secret-key"— signs legacy HS256 customer sessions (ticketcms login,/api/customer/*,/api/orderTicketGroups)CONCIERGE_GUEST_JWT_SECRET = "dev-only-insecure-concierge-guest-secret"— signs concierge guest tokens
Warning, not fatal, deliberately: CoreAPI's prod overlay loads only
coreapi-configmap, while staging also loads coreapi-ticketing (43 keys)
via patch-env.yml. Prod has no equivalent, so every TICKETING_* var except
TICKETING_DB_NAME already falls back in production. Failing fast would convert
a latent misconfiguration into an outage.
Scope, not overstated: prod ticketing traffic appears to be served by the
legacy stack (ticket-iac) — ticketcms's PlayTelly CI defines only a staging
API base URL with the prod block commented out, and there is no prod etiket
route in playtelly-iac. So this reads as latent rather than exploited. The
prod APISIX core-route does forward all paths to CoreAPI, so the endpoints are
reachable regardless. Worth confirming with whoever owns that deployment.
Generated coreapi-env-vars.md — all 82 vars, defaults, read sites.
Phase 4 — ticketing schema → GORM
4a — the models could not migrate at all
19 of 26 failed outright against Postgres. Two MySQL-only constructs inherited from when this was a MySQL application:
type:bigint unsigned AUTO_INCREMENT -> syntax error at "unsigned"
type:datetime -> type "datetime" does not exist
96 tags fixed across 10 files → 26/26 migrate cleanly. Notes:
type:bigserial is not a valid replacement — it is a CREATE TABLE
pseudo-type and fails in ALTER COLUMN TYPE; dropping the tag and letting GORM
infer from uint+primaryKey is correct. One tag was malformed
(bigint unsigned:int), parsed as an unknown key and silently ignored.
These tags only affect migration — GORM never consults them when querying — so runtime is unchanged.
cmd/schema-diff was added and kept: it AutoMigrates one authoritative model
per table into a scratch database so the output can be diffed against the
seeded schema. Phase 5 needs the same check after the merge.
Duplicate models found: five tables have two competing definitions. Only
identity.OrderTicketGroup genuinely differs (lacks program_cd);
commerce's matches the DB. config and schema_migrations have no model.
4b — 27 discrepancies closed
Database treated as source of truth throughout — it is a production snapshot; the structs are what the code merely believes.
- nullability (20) — 16 columns the DB allows NULL on carried
not null; 4 the DB marks NOT NULL were missing it - int width (4) —
advertising_ticket.placement,report.schedule_day/_day_of_week/_monthareinteger;uintrenderedbigint - varchar length (1) —
report_attachment.attachment_pathis 255, tag said 510 - missing fields (2) —
general.jp_api_key_2/jp_ag_token_2exist in the schema but had no struct fields, so CoreAPI could not read the second JohorPay key pair at all
A mistake worth recording: the first pass matched column:updated_at
file-wide and dropped not null from six structs that should keep it. The diff
caught it immediately — which is the argument for gating on a real diff rather
than reasoning about tags. Now scoped per struct.
Result: 0 missing columns, 0 extra columns, 0 type/nullability differences, and 6 timestamp cases deliberately left.
The 6 timestamps — drafted, not applied
25 of 28 tables use timestamptz; three do not. That is the signature of
hand-written DDL: pg_dump always emits timestamp with time zone, while a
person typing TIMESTAMP gets without. Those three are exactly the
hand-appended blocks in the seed — so the seed is internally inconsistent, and
matching it would cement an accident.
migration/2026-09-09_ticketing_timestamptz.sql converts them. The explicit
AT TIME ZONE 'UTC' is the whole point, demonstrated on a scratch DB:
with AT TIME ZONE 'UTC': 03:37:40.962023 -> 03:37:40.962023+00
without it, from a +08 session: 03:37:40.962023 -> 2026-08-17 19:37:40+00
Omitting it silently shifts every value 8 hours when run from a Malaysian workstation rather than the server. Verified idempotent; post-migration the schema diffs to zero differences against the models.
⚠️ It fixes an existing database only. A fresh down -v re-seeds from
ticketing-schema.sql and reintroduces the inconsistency — the durable fix
belongs upstream.
Open upstream requests (playtelly-iac)
ticketing-schema.sql:2247—CREATE TABLE IF NOT EXISTS ticket_group_integration_configis the only 1 of 28 notjohorzoo.-qualified. With the pg_dump preamble at:17blankingsearch_path, it fails with "no schema has been selected to create in". Worked around locally viasedin the runner. Staging's own init would fail from scratch too.- The three
timestamp without time zonetables — fix in the schema file so fresh volumes stop reproducing it. - ~~Data-file COPY ordering~~ — fixed upstream by
f73cb84, which movedticket_groupahead of its dependents. Notedown -valone does not fix that class of problem; ordering does.
Standing risks
- Ticketing schedulers run on every replica with no locking. 5 goroutines,
opt-out, no lock or leader election; prod HPA is
minReplicas: 2, maxReplicas: 5and no disable var is set in the IaC. The email job does select-then-mark, so 2–5 pods send the same mail. Likely causing duplicate ticket/report emails in production today. Phase 5 makespg_advisory_locktrivially available. - Production PII in git, three locations — 2,367 citizens incl. NRIC. Malaysian PDPA applies. History purging and reportability are decisions above a dev session.
- Secrets in a plain k8s ConfigMap (
coreapi-configmap, prod + staging):DB_PASS,COREAPI_EMAIL_PASSWORD,MINIO_SECRET_KEY,COREAPI_AUTHAPI_PSK. - Live credentials committed in
Quickstart/.envand.env.shared-dev. - TLS verification disabled on POS and payment clients.
- Per-product email fields have no Go consumer — editable in SpatioViewAdmin, change nothing.