Concierge — WebRTC Calling (staff↔guest)
Flat file, not a full domain page — per this site's own convention (see
Domains), a domain gets promoted to its own index.md/folder
once it has a stable API surface. This one doesn't yet: the mediasoup
signaling surface has been stable at 7 endpoints since 2026-09-01, but real
gaps remain open (queue/routing, generic error masking, push signaling — see
below). Maps to the future Engagement domain in the v3.4 ontology, not a
domain of its own.
This page covers calling only. Same domain/concierge package's chat,
translation and AI hotel guide feature is a separate sibling page:
Concierge — Guest Chat, Translation & AI Hotel Guide.
First written 2026-09-20, from a direct code audit of domain/concierge on
current staging (6ae07d0) — not from the older
WEBRTC_POC_HANDOFF.md/WEBRTC_DEMO_RUNBOOK.md at the repo root, which
describe the 2026-08-31 proof-of-concept snapshot and are now stale on their
two headline findings (see What changed since the POC).
Overview
A 1:1 WebRTC call between a hotel guest (in-room guest webapp) and staff, built on a standalone Node mediasoup-service for the actual media plane. CoreAPI never touches RTP/SRTP — it only relays control-plane calls (capabilities, transports, produce/consume, close) and gates them behind real auth.
Architecture
graph LR
Guest["Guest webapp"]
Staff["Staff console"]
CoreAPI["CoreAPI domain/concierge"]
Mediasoup["mediasoup-service (Node)"]
Guest -->|guest HS256| CoreAPI
Staff -->|platform RS256| CoreAPI
CoreAPI -->|control-plane only| Mediasoup
Guest -.->|RTP/SRTP direct| Mediasoup
Staff -.->|RTP/SRTP direct| Mediasoup
Two separate connections per client, not one relayed through the other: signaling (dashed arrows above wouldn't exist if it were a relay) goes guest/staff → CoreAPI only, while actual media goes guest/staff → mediasoup-service directly. CoreAPI is never in the RTP/SRTP path at all, by construction, not just by convention.
Data Model
CallLog (domain/concierge/models.go) — one row per call attempt:
| Field | Type | Notes |
|---|---|---|
StayID |
uint |
Which guest stay placed the call — not the guest's user id directly |
DepartmentID |
uint |
Which department was called |
RoomID |
string |
"call-<id>", generated after insert (needs the DB id first) — the value both clients pass to every /concierge/rtc/* call |
Status |
CallStatus |
ringing | accepted | ended — see state machine below |
StartedAt |
time.Time |
Set on create |
AcceptedAt |
*time.Time |
Nil until a staff member accepts |
EndedAt |
*time.Time |
Nil until either side ends it |
Department |
Department |
Preloaded FK — org membership for auth checks reads through this |
Auth model
Two separate gates, both real (no header-trust shortcuts):
/concierge/staff/calls/*—middleware.ProtectedRS256, CoreAPI's standard AuthAPI JWT. Full staff auth, same as every other staff-facing domain./concierge/rtc/*(the 7 mediasoup signaling endpoints) —RequireGuestOrStaff(domain/concierge/handlers/handler.go): tries concierge's own guest HS256 token first, falls back to the platform RS256 validator. Both a guest (in the call) and staff (answering it) need to reach the same signaling endpoints.
Org scope comes from the validated token, not a header. An earlier
version of staff calling read a client-supplied X-Organization-ID header
outright — fixed (6ae07d0, 2026-09-14): org now comes from
claims.Org.OrgID on the caller's own RS256 token.
Per-room ownership is enforced on 3 of 7 endpoints. authorizeRoom
(handler_rtc.go) looks up the CallLog behind a roomId and checks the
caller is actually a party to it — a guest must own the call's StayID,
staff must belong to the call's Department's OrganizationID. Applied to
GetRoomCapabilities, CreateTransport, ListRoomProducers. Deliberate,
disclosed scope cut: connect/produce/consume/close only take a
transportId, not a roomId, and don't re-check ownership — a
transportId is only ever handed back to an already-authorized caller by
CreateTransport, so re-checking on every subsequent call was judged
unnecessary for this pass.
Authorization is org-scoped, not department-scoped, on purpose. Any authenticated staff member of the org can see and accept any ringing call for that org's departments — real per-department routing/membership is a known, explicit follow-up, not a blocker.
Call lifecycle — real, not a stub
CallLog (domain/concierge/models.go) used to be an unwired stub with a
hardcoded 3-second fake connect. It's now a real state machine:
CallRinging → CallAccepted → CallEnded, driven by:
| Endpoint | Auth | Effect |
|---|---|---|
Guest StartCall |
guest | Creates CallLog, RoomID = "call-<id>", status ringing |
GET /staff/calls/ringing |
staff | Poll for ringing calls in the org |
GET /staff/calls/:id |
staff | Poll one call's status (mirrors the guest side's GetCall) |
POST /staff/calls/:id/accept |
staff | ringing → accepted, only from ringing — prevents two staff double-accepting |
PATCH /staff/calls/:id/end (staff) / guest EndCall |
either | → ended, calls mediasoup.CloseRoom for authoritative server-side cleanup |
The CloseRoom call on end (11dc6db, 2026-09-02) fixes a real
port-exhaustion leak — abandoned calls (crash, force-quit, declined without
notice) used to leave mediasoup rooms open indefinitely.
Still genuinely open (not fixed by the above)
- No queue/routing.
StartCallties a guest directly to oneDepartmentID— no queueing, no "assign to available staff," no reassignment on no-answer. - No push signaling. New/gone producers and call-status changes are all discovered by polling (~1.5s), not pushed — VerneMQ/MQTT (used elsewhere in this domain for device casting commands) has not been adopted for call signaling.
- Generic error masking.
ErrMediasoupUnavailablecollapses every failure mode — real outage, bad producer ID, closed transport — into the same503. MEDIASOUP_ANNOUNCED_IPis manual and network-specific — an accepted limitation of the current single-laptop deployment, not a code gap.
Planned features
Department-scoped call routing — closes the org-scoped-not-department-scoped
gap above by reusing domain/tenancy's existing Team/TeamMember model
(pkg/db/migrate.go:1009-1023) rather than inventing a new membership
table: add a nullable Department.TeamID, then filter
StaffListRingingCalls through team_members instead of org alone. The
engineering itself is small — one migration, one query change — since the
reusable roster infrastructure already exists and is live elsewhere
(ticketing's WorkspaceHasPermission chain). [PLACEHOLDER — target
date not yet committed]. Two things outside pure engineering scope that
this depends on, not yet scoped or owned:
- [PLACEHOLDER — who decides department→team mapping per org, and when]
- [PLACEHOLDER — no admin UI exists yet to assign staff to a department's team; first rollout would need a manual/SQL-driven roster per org unless this is built too]
Running this locally
mediasoup-service is a separate Node/Express container — CoreAPI does not
start it for you. CoreAPI runs fine without it (the /concierge/rtc/*
routes just 503 with ErrMediasoupUnavailable), so it's easy to bring up
the normal stack, test guest sign-in/ordering/casting, and only then
discover calling doesn't work because this piece was never started.
cd Quickstart
make rtc # CoreAPI + AuthAPI + infra + mediasoup-service, together
make rtc-down # tear it back down
(make rtc — not the quickstart.md doc's make dev/make dev-d, which
don't exist as real targets; see the Makefile itself, not that page.)
Two env vars actually matter once it's up — both already listed in the CoreAPI env var reference, but neither explains why on that page:
MEDIASOUP_SERVICE_URL— CoreAPI's URL for the mediasoup-service container (domain/concierge/config.go:45). Unset by default; leave it unset to intentionally run without calling, or point it at the compose service (e.g.http://mediasoup:34430) to enable it.MEDIASOUP_ANNOUNCED_IP— set on the mediasoup-service container itself, not CoreAPI. Must be a real IP reachable by both the guest and staff device's network — notlocalhost/127.0.0.1unless both clients are on the exact same machine. This is the "accepted limitation of the current single-laptop deployment" mentioned above, not somethingmake rtcconfigures for you.
Testing this on staging (no local setup needed)
This doc page itself is only hosted on production (developer.castis.io
— there's no separate staging docs site), but the actual feature only
exists on staging — it hasn't shipped to production at all. Don't
conflate the two: reading this on prod docs, you still need to point at
staging to actually try it.
- Guest side: stg.guest.castis.io — real, deployed
ConciergeGuestAppstaging build - Backend:
https://stag.api.castis.io— real, deployed CoreAPI staging (VITE_API_BASE_URLin its own configmap) - Staff side: no public web build —
SpatioViewMobile'sstagingEAS profile needs to actually be built/installed (eas build --profile staging --platform ios|android) to test the staff side; there's no browser-based staff client
Both staging deployments already point at the same live mediasoup-service
instance this page's earlier smoke-test verified — no separate mediasoup
setup needed to test end-to-end on staging.
Verifying mediasoup-service can actually relay (without placing a call)
mediasoup-service's own HTTP API has no auth of its own (by design — see Auth model), so it can be smoke-tested directly, with no guest/staff session and no CoreAPI involved:
BASE="http://<mediasoup-service host>:34430"
ROOM="smoketest-1"
# 1. Confirms the router boots and negotiates codecs for a brand-new room
curl -X POST "$BASE/routers/$ROOM/capabilities"
# 2. Confirms it can actually hand out a real, connectable transport
curl -X POST "$BASE/routers/$ROOM/transports" \
-H "Content-Type: application/json" \
-d '{"direction":"send"}'
# 3. rooms/transports counts should now read 1/1
curl "$BASE/health"
# 4. Tear the test room back down — same call CoreAPI makes on real hangup
curl -X POST "$BASE/routers/$ROOM/close"
curl "$BASE/health" # back to 0/0
The transport response's iceCandidates[].ip is the field that actually
matters — it must be the real, internet-reachable ANNOUNCED_IP, not
127.0.0.1 or a private LAN address, since that's the literal address a
real client's WebRTC stack is told to send media to. Confirmed 2026-09-20
against the live staging/production instance: it correctly announced the
DigitalOcean droplet's own public IP on real UDP/TCP ports, not a private
or loopback address.
API reference
Full OpenAPI 3.0 spec for the 12 endpoints above (guest calls, staff calls,
and the 7 mediasoup RTC signaling routes), added 2026-09-20 alongside this
page — same convention as tenant.yaml / spatial.yaml.
Paths are relative to /api/v1/concierge. Validated against the OpenAPI 3.0
schema (openapi-spec-validator) and confirmed to render cleanly through
this site's render_swagger plugin.
What changed since the POC
If you're holding the 2026-08-31 handoff doc's mental model: its two
headline "confirmed still broken" gaps —no per-room ownership check, and
guest-only auth with no staff path at all— were both fixed the very next
week (74eca5b, 2026-09-01) and are described accurately above. The 7-
endpoint mediasoup signaling surface itself didn't grow; the staff-calling
layer on top of it (4 endpoints, CallLog wired up for real) is what's new.