2026-09-23 — AI hotel guide reliability: found and fixed a real prompt-bloat bug
Brief mark, not a full changelog — see
Concierge — Guest Chat, Translation & AI Hotel Guide
for the full investigation. Three repos, all pushed to staging:
CoreAPI, cdn-ai (a different team's repo — changes here are additive/
opt-in, existing callers unaffected), ConciergeGuestApp.
Root cause found
- The AI hotel guide's honest-fallback message was firing for questions
the seeded corpus plainly covered. cdn-ai's own audit.jsonl showed
"Context size has been exceeded" and 140-180s timeouts, not a
grounding gap.
- Measured, not guessed: every /chat call (translation and the guide
alike) was paying for cdn-ai's CDN-analytics SYSTEM_PROMPT (~3,268
tokens) plus its full tool schema list (~10k tokens) — 85-90% of a
typical ~15k-token guide prompt, regardless of relevance.
Fixed — cdn-ai
- mcp_client.py's tool-list fetch was opening a brand-new MCP session
(fresh connection + handshake) on every single /chat call, even
though tool schemas never change at runtime. Now cached, 10-minute TTL.
- ChatRequest gained use_tools/system_prompt (both default to
original behavior — no existing caller affected). A caller with nothing
to do with CDN analytics can now skip both entirely.
Fixed — CoreAPI
- domain/ai.Client sends use_tools: false + a short
minimalSystemPrompt on every call — cuts a typical guide prompt from
~15k tokens to ~2.2-2.5k.
- Ask() (the guide's synchronous answer) now runs on its own
GuideTimeout (25s default), separate from translation's intentionally
short RequestTimeout (6s) — the two were incorrectly sharing one
budget tuned for "fail fast, don't stall a chat send," which is wrong
for an answer the guest is already waiting on with a typing indicator.
- Deepened the RAG corpus: 6 new policy topics (wifi, gym/pool, laundry,
airport transfer, loyalty programme, local attractions), 3 existing
ones enriched (pets, parking, cancellation) — token-budget-checked
against the fix above; adds well under 1k tokens.
- AI guide prompt now allows light markdown (**bold** for a key fact,
- bullets only for 3+ items) — was explicitly "no markdown" before.
Fixed — ConciergeGuestApp
- guideMarkdown.jsx renders the two patterns above, scoped to AI-guide
bubbles only — department/staff chat and translated text unaffected.
- Suggestion pills: 2 more (from the new topics), translated into all 11
supported languages; fixed a real bug where tapping a translated pill
sent the English original as the guest's own message (visually broken —
now sends whatever's displayed).
- One real, verified Unsplash photo (hotel lobby) as a hero banner at the
top of the AI guide thread.
Not done, deliberately deferred - Smarter retrieval (only send relevant corpus entries, not the whole thing) — full-context stuffing is still fine token-wise post-fix; swapping retrieval strategy this close to a demo/submission deadline was judged riskier than the benefit. - Auto-detect language reliability itself (as opposed to the already-shipped explicit-selector mitigation) — a small-model limitation, not something to chase this late.
See: Concierge — Guest Chat, Translation & AI Hotel Guide, AI — CDN-AI system