Skip to content

Concierge — WebRTC Calling (staff↔guest)

Flat file, not a full domain page — per this site's own convention (see Domains), a domain gets promoted to its own index.md/folder once it has a stable API surface. This one doesn't yet: the mediasoup signaling surface has been stable at 7 endpoints since 2026-09-01, but real gaps remain open (queue/routing, generic error masking, push signaling — see below). Maps to the future Engagement domain in the v3.4 ontology, not a domain of its own.

This page covers calling only. Same domain/concierge package's chat, translation and AI hotel guide feature is a separate sibling page: Concierge — Guest Chat, Translation & AI Hotel Guide.

First written 2026-09-20, from a direct code audit of domain/concierge on current staging (6ae07d0) — not from the older WEBRTC_POC_HANDOFF.md/WEBRTC_DEMO_RUNBOOK.md at the repo root, which describe the 2026-08-31 proof-of-concept snapshot and are now stale on their two headline findings (see What changed since the POC).


Overview

A 1:1 WebRTC call between a hotel guest (in-room guest webapp) and staff, built on a standalone Node mediasoup-service for the actual media plane. CoreAPI never touches RTP/SRTP — it only relays control-plane calls (capabilities, transports, produce/consume, close) and gates them behind real auth.

Architecture

graph LR
    Guest["Guest webapp"]
    Staff["Staff console"]
    CoreAPI["CoreAPI domain/concierge"]
    Mediasoup["mediasoup-service (Node)"]

    Guest -->|guest HS256| CoreAPI
    Staff -->|platform RS256| CoreAPI
    CoreAPI -->|control-plane only| Mediasoup
    Guest -.->|RTP/SRTP direct| Mediasoup
    Staff -.->|RTP/SRTP direct| Mediasoup

Two separate connections per client, not one relayed through the other: signaling (dashed arrows above wouldn't exist if it were a relay) goes guest/staff → CoreAPI only, while actual media goes guest/staff → mediasoup-service directly. CoreAPI is never in the RTP/SRTP path at all, by construction, not just by convention.

Data Model

CallLog (domain/concierge/models.go) — one row per call attempt:

Field Type Notes
StayID uint Which guest stay placed the call — not the guest's user id directly
DepartmentID uint Which department was called
RoomID string "call-<id>", generated after insert (needs the DB id first) — the value both clients pass to every /concierge/rtc/* call
Status CallStatus ringing | accepted | ended — see state machine below
StartedAt time.Time Set on create
AcceptedAt *time.Time Nil until a staff member accepts
EndedAt *time.Time Nil until either side ends it
Department Department Preloaded FK — org membership for auth checks reads through this

Auth model

Two separate gates, both real (no header-trust shortcuts):

  • /concierge/staff/calls/* — middleware.ProtectedRS256, CoreAPI's standard AuthAPI JWT. Full staff auth, same as every other staff-facing domain.
  • /concierge/rtc/* (the 7 mediasoup signaling endpoints) — RequireGuestOrStaff (domain/concierge/handlers/handler.go): tries concierge's own guest HS256 token first, falls back to the platform RS256 validator. Both a guest (in the call) and staff (answering it) need to reach the same signaling endpoints.

Org scope comes from the validated token, not a header. An earlier version of staff calling read a client-supplied X-Organization-ID header outright — fixed (6ae07d0, 2026-09-14): org now comes from claims.Org.OrgID on the caller's own RS256 token.

Per-room ownership is enforced on 3 of 7 endpoints. authorizeRoom (handler_rtc.go) looks up the CallLog behind a roomId and checks the caller is actually a party to it — a guest must own the call's StayID, staff must belong to the call's Department's OrganizationID. Applied to GetRoomCapabilities, CreateTransport, ListRoomProducers. Deliberate, disclosed scope cut: connect/produce/consume/close only take a transportId, not a roomId, and don't re-check ownership — a transportId is only ever handed back to an already-authorized caller by CreateTransport, so re-checking on every subsequent call was judged unnecessary for this pass.

Authorization is org-scoped, not department-scoped, on purpose. Any authenticated staff member of the org can see and accept any ringing call for that org's departments — real per-department routing/membership is a known, explicit follow-up, not a blocker.

Call lifecycle — real, not a stub

CallLog (domain/concierge/models.go) used to be an unwired stub with a hardcoded 3-second fake connect. It's now a real state machine: CallRinging → CallAccepted → CallEnded, driven by:

Endpoint Auth Effect
Guest StartCall guest Creates CallLog, RoomID = "call-<id>", status ringing
GET /staff/calls/ringing staff Poll for ringing calls in the org
GET /staff/calls/:id staff Poll one call's status (mirrors the guest side's GetCall)
POST /staff/calls/:id/accept staff ringing → accepted, only from ringing — prevents two staff double-accepting
PATCH /staff/calls/:id/end (staff) / guest EndCall either → ended, calls mediasoup.CloseRoom for authoritative server-side cleanup

The CloseRoom call on end (11dc6db, 2026-09-02) fixes a real port-exhaustion leak — abandoned calls (crash, force-quit, declined without notice) used to leave mediasoup rooms open indefinitely.

Still genuinely open (not fixed by the above)

  • No queue/routing. StartCall ties a guest directly to one DepartmentID — no queueing, no "assign to available staff," no reassignment on no-answer.
  • No push signaling. New/gone producers and call-status changes are all discovered by polling (~1.5s), not pushed — VerneMQ/MQTT (used elsewhere in this domain for device casting commands) has not been adopted for call signaling.
  • Generic error masking. ErrMediasoupUnavailable collapses every failure mode — real outage, bad producer ID, closed transport — into the same 503.
  • MEDIASOUP_ANNOUNCED_IP is manual and network-specific — an accepted limitation of the current single-laptop deployment, not a code gap.

Planned features

Department-scoped call routing — closes the org-scoped-not-department-scoped gap above by reusing domain/tenancy's existing Team/TeamMember model (pkg/db/migrate.go:1009-1023) rather than inventing a new membership table: add a nullable Department.TeamID, then filter StaffListRingingCalls through team_members instead of org alone. The engineering itself is small — one migration, one query change — since the reusable roster infrastructure already exists and is live elsewhere (ticketing's WorkspaceHasPermission chain). [PLACEHOLDER — target date not yet committed]. Two things outside pure engineering scope that this depends on, not yet scoped or owned:

  • [PLACEHOLDER — who decides department→team mapping per org, and when]
  • [PLACEHOLDER — no admin UI exists yet to assign staff to a department's team; first rollout would need a manual/SQL-driven roster per org unless this is built too]

Running this locally

mediasoup-service is a separate Node/Express container — CoreAPI does not start it for you. CoreAPI runs fine without it (the /concierge/rtc/* routes just 503 with ErrMediasoupUnavailable), so it's easy to bring up the normal stack, test guest sign-in/ordering/casting, and only then discover calling doesn't work because this piece was never started.

cd Quickstart
make rtc       # CoreAPI + AuthAPI + infra + mediasoup-service, together
make rtc-down  # tear it back down

(make rtc — not the quickstart.md doc's make dev/make dev-d, which don't exist as real targets; see the Makefile itself, not that page.)

Two env vars actually matter once it's up — both already listed in the CoreAPI env var reference, but neither explains why on that page:

  • MEDIASOUP_SERVICE_URL — CoreAPI's URL for the mediasoup-service container (domain/concierge/config.go:45). Unset by default; leave it unset to intentionally run without calling, or point it at the compose service (e.g. http://mediasoup:34430) to enable it.
  • MEDIASOUP_ANNOUNCED_IP — set on the mediasoup-service container itself, not CoreAPI. Must be a real IP reachable by both the guest and staff device's network — not localhost/127.0.0.1 unless both clients are on the exact same machine. This is the "accepted limitation of the current single-laptop deployment" mentioned above, not something make rtc configures for you.

Testing this on staging (no local setup needed)

This doc page itself is only hosted on production (developer.castis.io — there's no separate staging docs site), but the actual feature only exists on staging — it hasn't shipped to production at all. Don't conflate the two: reading this on prod docs, you still need to point at staging to actually try it.

  • Guest side: stg.guest.castis.io — real, deployed ConciergeGuestApp staging build
  • Backend: https://stag.api.castis.io — real, deployed CoreAPI staging (VITE_API_BASE_URL in its own configmap)
  • Staff side: no public web build — SpatioViewMobile's staging EAS profile needs to actually be built/installed (eas build --profile staging --platform ios|android) to test the staff side; there's no browser-based staff client

Both staging deployments already point at the same live mediasoup-service instance this page's earlier smoke-test verified — no separate mediasoup setup needed to test end-to-end on staging.

Verifying mediasoup-service can actually relay (without placing a call)

mediasoup-service's own HTTP API has no auth of its own (by design — see Auth model), so it can be smoke-tested directly, with no guest/staff session and no CoreAPI involved:

BASE="http://<mediasoup-service host>:34430"
ROOM="smoketest-1"

# 1. Confirms the router boots and negotiates codecs for a brand-new room
curl -X POST "$BASE/routers/$ROOM/capabilities"

# 2. Confirms it can actually hand out a real, connectable transport
curl -X POST "$BASE/routers/$ROOM/transports" \
  -H "Content-Type: application/json" \
  -d '{"direction":"send"}'

# 3. rooms/transports counts should now read 1/1
curl "$BASE/health"

# 4. Tear the test room back down — same call CoreAPI makes on real hangup
curl -X POST "$BASE/routers/$ROOM/close"
curl "$BASE/health"   # back to 0/0

The transport response's iceCandidates[].ip is the field that actually matters — it must be the real, internet-reachable ANNOUNCED_IP, not 127.0.0.1 or a private LAN address, since that's the literal address a real client's WebRTC stack is told to send media to. Confirmed 2026-09-20 against the live staging/production instance: it correctly announced the DigitalOcean droplet's own public IP on real UDP/TCP ports, not a private or loopback address.

API reference

Full OpenAPI 3.0 spec for the 12 endpoints above (guest calls, staff calls, and the 7 mediasoup RTC signaling routes), added 2026-09-20 alongside this page — same convention as tenant.yaml / spatial.yaml. Paths are relative to /api/v1/concierge. Validated against the OpenAPI 3.0 schema (openapi-spec-validator) and confirmed to render cleanly through this site's render_swagger plugin.

What changed since the POC

If you're holding the 2026-08-31 handoff doc's mental model: its two headline "confirmed still broken" gaps —no per-room ownership check, and guest-only auth with no staff path at all— were both fixed the very next week (74eca5b, 2026-09-01) and are described accurately above. The 7- endpoint mediasoup signaling surface itself didn't grow; the staff-calling layer on top of it (4 endpoints, CallLog wired up for real) is what's new.