Skip to content

Concierge — Guest Chat, Translation & AI Hotel Guide

Flat file, not a full domain page — same convention as Concierge — WebRTC Calling, which this page is a sibling to (same domain/concierge package, different feature slice — chat/AI, not calling). Gets promoted once the surface is stable enough to justify an OpenAPI spec of its own.

Written 2026-09-24, from a direct code audit of domain/concierge and domain/ai on staging, plus a live incident worked end-to-end that same day (see Reliability incident below).


Overview

Three related but separable features live in the same guest↔staff chat surface:

  1. Department chat — guest messages a department (Housekeeping, Front desk, etc.); a canned "X will get back to you shortly" fires once per thread, then a real staff member replies from SpatioViewMobile.
  2. Best-effort translation — every guest message gets a background English translation for staff to read, and staff replies get translated back into the guest's detected/chosen language. Fire-and-forget: never blocks the guest's send.
  3. AI hotel guide — a separate ChatThread.IsAI thread where the guest asks a question and gets a real, grounded LLM answer synchronously — no human in the loop unless the AI can't answer.

All three share one external dependency: cdn-ai's agent (domain/ai package), a different team's /chat endpoint built and tuned for CDN/ streaming analytics, not general chat. See that package's own comment (domain/ai/config.go) for the risk this implies, and AI — CDN-AI system for that system's own architecture.

Translation

ai.Client (domain/ai/client.go) wraps cdn-ai's /chat:

Method Direction Used by
DetectLanguageAndTranslateToEnglish guest → English, once per thread SendMessage, on the guest's first message
TranslateToEnglish guest → English SendMessage, every later guest message (language already known)
TranslateFromEnglish English → guest's language StaffReplyToThread, and the canned reply's own translation
Ask arbitrary prompt → answer the AI guide only (see below)

Stored in ChatMessageTranslation (one row per translated message, not a column on ChatMessage itself) and ChatThread.GuestLanguage (detected once, nil until then — see that field's own doc comment for the dual auto-detect/explicit-choice semantics it shares with the guide).

Everything runs go func(){...}(context.Background(), ...) — never awaited by the request that triggered it. A translation failure only logs ([concierge] translation failed for message %d: %v); the guest/staff send itself already succeeded. This is why a genuinely broken translation call shows up as messages just never getting a translatedBody, not as a visible error anywhere in the UI.

AI hotel guide

answerGuestQuestion (domain/concierge/handlers/handler_ai_guide.go) — unlike translation, this is awaited: it's the actual content the guest is waiting for, with a typing indicator shown client-side.

Grounding is deliberately full-context, not vector-retrieved. Every call sends the guide's entire corpus — every active PolicyEntry (see that model's own doc comment for why "RAG-lite" is the right call at this scale) plus every active Outlet/MenuItem — regardless of what the question is actually about. buildGuideContext assembles it fresh on every request; there is no caching, no chunking, no embedding step.

Language: ChatThread.GuestLanguage is reused here as an explicit guest choice (the globe-icon selector in ConciergeGuestApp, not auto-detection) — once set, every answer is pinned to it regardless of what language the question happens to be written in ("Always answer in %s, regardless of..."). Auto-detect mode (no selection made) falls back to "Answer in the SAME language the guest's question is written in" — this is genuinely less reliable for short/ casual input on a small quantized model; see the reliability incident below for a concrete example. The selector exists specifically because auto-detect isn't trustworthy enough on its own.

Formatting: the prompt allows exactly two markdown patterns — **bold** for a key number/price/time, and - bullets only when listing 3+ items — never more, never for a single fact. Rendered client-side by ConciergeGuestApp/src/utils/guideMarkdown.jsx, a small dependency-free parser scoped to AI-guide bubbles only; department/staff messages and translated text render as plain text, unchanged.

Fails open to concierge.AIGuideFallbackMessage ("having trouble answering... message Front desk directly") on any ai.Client error — an honest admission, not a canned "will get back to you shortly" (there's no automatic human follow-up on this thread the way there is for department chat; see that constant's own doc comment).

Every call this domain makes — translation and the guide alike — sends use_tools: false and a short minimalSystemPrompt, opting out of cdn-ai's own CDN-analytics system prompt and its ~10k tokens of tool schemas (see the reliability incident for why this exists). Nothing here has ever needed cdn-ai's ClickHouse tools.

Stateless, no memory — by design, with a real UX consequence

Every /chat call cdn-ai's agent handles is single-shot: run_loop_traced (question) takes one question string, no conversation history. There is no session, no "remember what I asked a message ago." A guest asking "What time does it close?" right after asking about the pool gets an answer with zero awareness that "it" was supposed to mean the pool — the model only ever sees that one message plus the full static corpus, every time.

Reliability incident (2026-09-23)

Real, live, worked end-to-end in one session — kept here as the concrete "why" behind several of the design points above, not as a postmortem nobody will read again.

Symptom: the AI guide's honest-fallback message fired for questions that were plainly covered by the seeded corpus (e.g. "can I bring my dog?", with a real pets policy entry). cdn-ai's own agent/api/logs/audit.jsonl showed the actual cause: "Context size has been exceeded", and separately "Request timed out" at 140-180s per attempt — not a grounding gap.

Root cause, measured, not guessed: every /chat call — translation or guide, ours or anyone else's — paid for cdn-ai's own SYSTEM_PROMPT (agent/api/loop.py, measured ~3,268 tokens of ClickHouse/CDN instructions) plus its full tool schema list (~10k tokens, per that project's own CLAUDE.md), regardless of relevance. For a guide question, that's ~13k of a typical ~15k-token prompt — 85-90% pure waste — against llama-server's configured --ctx-size 40000 --parallel 6 --kv-unified — a shared pool across 6 concurrent slots, confirmed directly via ps aux on the Mac mini itself (matches this site's own AI — CDN-AI system reference, n_ctx = 40192; a different, unrelated system's docs elsewhere still cited a stale 16384 for a different deployment — not this one). A few concurrent heavy requests could exhaust that shared pool on their own, no unrelated load required.

Fix, two parts, both live on staging:

  1. cdn-ai's ChatRequest gained use_tools/system_prompt fields (defaulting to the exact original behavior — every existing caller unaffected) — a caller with nothing to do with CDN analytics can opt out of both entirely.
  2. This domain's ai.Client sends that opt-out on every call, plus a short minimalSystemPrompt replacing cdn-ai's own — cutting a typical guide call from ~15k tokens to ~2.2-2.5k.

A second, independent fix landed alongside it: Ask() (the guide) was sharing RequestTimeout's 6-second budget with translation, where a short timeout is the correct tradeoff (never stall a guest's chat send). Ask() now runs on its own longer GuideTimeout (default 25s, separate http.Client) — translation's behavior is untouched.

Also fixed the same day, upstream of the above: cdn-ai's mcp_client.py was opening a brand-new MCP session (fresh connection + protocol handshake) against every configured tool server on every single /chat call, even though tool schemas never change at runtime — a second, independent point of failure alongside the LLM backend itself. Now cached at module level, 10-minute TTL.

Restart demo

POST /demo/restart (handler_demo.go) — soft-deletes ChatThreads (and all messages), hard-deletes Order/OrderItem/ServiceRequest rows, resets RoomControlState, and rebuilds the two original FolioLines — but never touches the Stay row, so room-number/ password login can't break. Deliberately placed as a small, unobtrusive text link below Sign Out in the guest app's profile sheet, not a prominent button — this is a demo-reset tool, not a guest-facing feature.

Known gap, not yet fixed: resetting from the guest side has no signal to SpatioViewMobile — a staff device still viewing/polling an already-deleted thread shows stale state until its own next poll cycle (or a manual restart) picks up the change.

Still genuinely open

  • Auto-detect language reliability — see above; the selector is the mitigation, not a fix to detection itself.
  • No conversation memory — see above; a real UX ceiling on follow-up questions.
  • Suggestion pills are a fixed, curated list (ChatThread.jsx's AI_SUGGEST), not server-driven or derived from real usage — there's no usage-stats pipeline yet to derive "popular questions" from.
  • RAG stays full-context, not retrieved — the right call at today's corpus size (see PolicyEntry's own comment), but revisit if the corpus genuinely grows past a few dozen entries; smarter retrieval was deliberately deferred (2026-09-23) as too risky to introduce this close to a demo/submission deadline.

Testing this on staging

Same staging environment as Concierge — WebRTC Calling: guest side at stg.guest.castis.io, backend at https://stag.api.castis.io. Log in with room 1201 / veranda1201 (the seed fixture) to reach both department chat and the AI guide end to end.