Concierge — Guest Chat, Translation & AI Hotel Guide
Flat file, not a full domain page — same convention as
Concierge — WebRTC Calling, which this page is a sibling
to (same domain/concierge package, different feature slice — chat/AI,
not calling). Gets promoted once the surface is stable enough to justify
an OpenAPI spec of its own.
Written 2026-09-24, from a direct code audit of domain/concierge and
domain/ai on staging, plus a live incident worked end-to-end that same
day (see Reliability incident below).
Overview
Three related but separable features live in the same guest↔staff chat surface:
- Department chat — guest messages a department (Housekeeping, Front
desk, etc.); a canned "X will get back to you shortly" fires once per
thread, then a real staff member replies from
SpatioViewMobile. - Best-effort translation — every guest message gets a background English translation for staff to read, and staff replies get translated back into the guest's detected/chosen language. Fire-and-forget: never blocks the guest's send.
- AI hotel guide — a separate
ChatThread.IsAIthread where the guest asks a question and gets a real, grounded LLM answer synchronously — no human in the loop unless the AI can't answer.
All three share one external dependency: cdn-ai's agent (domain/ai
package), a different team's /chat endpoint built and tuned for CDN/
streaming analytics, not general chat. See that package's own comment
(domain/ai/config.go) for the risk this implies, and
AI — CDN-AI system for that system's own architecture.
Translation
ai.Client (domain/ai/client.go) wraps cdn-ai's /chat:
| Method | Direction | Used by |
|---|---|---|
DetectLanguageAndTranslateToEnglish |
guest → English, once per thread | SendMessage, on the guest's first message |
TranslateToEnglish |
guest → English | SendMessage, every later guest message (language already known) |
TranslateFromEnglish |
English → guest's language | StaffReplyToThread, and the canned reply's own translation |
Ask |
arbitrary prompt → answer | the AI guide only (see below) |
Stored in ChatMessageTranslation (one row per translated message, not a
column on ChatMessage itself) and ChatThread.GuestLanguage (detected
once, nil until then — see that field's own doc comment for the dual
auto-detect/explicit-choice semantics it shares with the guide).
Everything runs go func(){...}(context.Background(), ...) — never
awaited by the request that triggered it. A translation failure only logs
([concierge] translation failed for message %d: %v); the guest/staff
send itself already succeeded. This is why a genuinely broken translation
call shows up as messages just never getting a translatedBody, not as a
visible error anywhere in the UI.
AI hotel guide
answerGuestQuestion (domain/concierge/handlers/handler_ai_guide.go) —
unlike translation, this is awaited: it's the actual content the guest
is waiting for, with a typing indicator shown client-side.
Grounding is deliberately full-context, not vector-retrieved. Every
call sends the guide's entire corpus — every active PolicyEntry (see
that model's own doc comment for why "RAG-lite" is the right call at this
scale) plus every active Outlet/MenuItem — regardless of what the
question is actually about. buildGuideContext assembles it fresh on
every request; there is no caching, no chunking, no embedding step.
Language: ChatThread.GuestLanguage is reused here as an explicit
guest choice (the globe-icon selector in ConciergeGuestApp, not
auto-detection) — once set, every answer is pinned to it regardless of
what language the question happens to be written in
("Always answer in %s, regardless of..."). Auto-detect mode (no
selection made) falls back to "Answer in the SAME language the guest's
question is written in" — this is genuinely less reliable for short/
casual input on a small quantized model; see the reliability incident
below for a concrete example. The selector exists specifically because
auto-detect isn't trustworthy enough on its own.
Formatting: the prompt allows exactly two markdown patterns —
**bold** for a key number/price/time, and - bullets only when listing
3+ items — never more, never for a single fact. Rendered client-side by
ConciergeGuestApp/src/utils/guideMarkdown.jsx, a small dependency-free
parser scoped to AI-guide bubbles only; department/staff messages and
translated text render as plain text, unchanged.
Fails open to concierge.AIGuideFallbackMessage ("having trouble
answering... message Front desk directly") on any ai.Client error — an
honest admission, not a canned "will get back to you shortly" (there's no
automatic human follow-up on this thread the way there is for department
chat; see that constant's own doc comment).
Every call this domain makes — translation and the guide alike — sends
use_tools: false and a short minimalSystemPrompt, opting out of
cdn-ai's own CDN-analytics system prompt and its ~10k tokens of tool
schemas (see the reliability incident for why this exists). Nothing here
has ever needed cdn-ai's ClickHouse tools.
Stateless, no memory — by design, with a real UX consequence
Every /chat call cdn-ai's agent handles is single-shot: run_loop_traced
(question) takes one question string, no conversation history. There is
no session, no "remember what I asked a message ago." A guest asking "What
time does it close?" right after asking about the pool gets an answer with
zero awareness that "it" was supposed to mean the pool — the model only
ever sees that one message plus the full static corpus, every time.
Reliability incident (2026-09-23)
Real, live, worked end-to-end in one session — kept here as the concrete "why" behind several of the design points above, not as a postmortem nobody will read again.
Symptom: the AI guide's honest-fallback message fired for questions
that were plainly covered by the seeded corpus (e.g. "can I bring my
dog?", with a real pets policy entry). cdn-ai's own
agent/api/logs/audit.jsonl showed the actual cause: "Context size has
been exceeded", and separately "Request timed out" at 140-180s per
attempt — not a grounding gap.
Root cause, measured, not guessed: every /chat call — translation or
guide, ours or anyone else's — paid for cdn-ai's own SYSTEM_PROMPT
(agent/api/loop.py, measured ~3,268 tokens of ClickHouse/CDN
instructions) plus its full tool schema list (~10k tokens, per that
project's own CLAUDE.md), regardless of relevance. For a guide question,
that's ~13k of a typical ~15k-token prompt — 85-90% pure waste —
against llama-server's configured --ctx-size 40000 --parallel 6
--kv-unified — a shared pool across 6 concurrent slots, confirmed
directly via ps aux on the Mac mini itself (matches this site's own
AI — CDN-AI system reference, n_ctx = 40192; a
different, unrelated system's docs elsewhere still cited a stale 16384
for a different deployment — not this one). A few concurrent heavy
requests could exhaust that shared pool on their own, no unrelated load
required.
Fix, two parts, both live on staging:
cdn-ai'sChatRequestgaineduse_tools/system_promptfields (defaulting to the exact original behavior — every existing caller unaffected) — a caller with nothing to do with CDN analytics can opt out of both entirely.- This domain's
ai.Clientsends that opt-out on every call, plus a shortminimalSystemPromptreplacing cdn-ai's own — cutting a typical guide call from ~15k tokens to ~2.2-2.5k.
A second, independent fix landed alongside it: Ask() (the guide) was
sharing RequestTimeout's 6-second budget with translation, where a short
timeout is the correct tradeoff (never stall a guest's chat send).
Ask() now runs on its own longer GuideTimeout (default 25s, separate
http.Client) — translation's behavior is untouched.
Also fixed the same day, upstream of the above: cdn-ai's
mcp_client.py was opening a brand-new MCP session (fresh connection +
protocol handshake) against every configured tool server on every
single /chat call, even though tool schemas never change at runtime —
a second, independent point of failure alongside the LLM backend itself.
Now cached at module level, 10-minute TTL.
Restart demo
POST /demo/restart (handler_demo.go) — soft-deletes ChatThreads
(and all messages), hard-deletes Order/OrderItem/ServiceRequest
rows, resets RoomControlState, and rebuilds the two original
FolioLines — but never touches the Stay row, so room-number/
password login can't break. Deliberately placed as a small, unobtrusive
text link below Sign Out in the guest app's profile sheet, not a
prominent button — this is a demo-reset tool, not a guest-facing feature.
Known gap, not yet fixed: resetting from the guest side has no signal
to SpatioViewMobile — a staff device still viewing/polling an
already-deleted thread shows stale state until its own next poll cycle (or
a manual restart) picks up the change.
Still genuinely open
- Auto-detect language reliability — see above; the selector is the mitigation, not a fix to detection itself.
- No conversation memory — see above; a real UX ceiling on follow-up questions.
- Suggestion pills are a fixed, curated list (
ChatThread.jsx'sAI_SUGGEST), not server-driven or derived from real usage — there's no usage-stats pipeline yet to derive "popular questions" from. - RAG stays full-context, not retrieved — the right call at today's
corpus size (see
PolicyEntry's own comment), but revisit if the corpus genuinely grows past a few dozen entries; smarter retrieval was deliberately deferred (2026-09-23) as too risky to introduce this close to a demo/submission deadline.
Testing this on staging
Same staging environment as Concierge — WebRTC Calling:
guest side at stg.guest.castis.io, backend
at https://stag.api.castis.io. Log in with room 1201 / veranda1201
(the seed fixture) to reach both department chat and the AI guide
end to end.