AI — CDN-AI system
The CDN-AI system answers natural-language questions about real CDN telemetry. It is two services plus an LLM:
- a FastAPI agent (
:8082) — the entry point (chat + REST + UI) and the LLM tool-calling loop. See FastAPI agent. - an MCP server (
:8085) — a standalone server exposing the ClickHouse analytics as callable tools. See MCP server. - an LLM backend — a llama.cpp server; the agent is the LLM's client.
Data lives in ClickHouse (cdn_ai database), fed by Vector from the real
cproxy / elb / streamer logs — see the ClickHouse schema for the
tables and the standard. For pilot status and the tools→table reference, see
the CDN-AI pilot status. For known issues and
how to debug, see Logging & known issues.
flowchart LR
U[User] -->|/chat, /ui/*| A[FastAPI agent :8082]
A -->|OpenAI tool-calling| L[llama.cpp LLM<br/>172.25.25.17:8001]
A -->|MCP streamable-http| M[MCP server :8085<br/>27 tools]
M -->|SQL| C[(ClickHouse<br/>cdn_ai)]
A -->|direct SQL, non-LLM REST| C
V[Vector] -->|cproxy / streamer logs| C
The agent talks to the LLM and to the MCP server; the LLM never touches ClickHouse directly — it only asks the agent to call MCP tools, and the agent runs them against ClickHouse.
Model
| Value | |
|---|---|
| Model | typhoon2.5-qwen3-30b-a3b (Qwen3-30B-A3B family, Q4_K_M GGUF, ~18.5 GB) |
| Runtime | llama.cpp (llama-server, OpenAI-compatible API) |
| Endpoint | http://172.25.25.17:8001/v1 (Mac mini) — same backend for local and staging |
| Context window | n_ctx = 40192 tokens (hard 400 exceed_context_size on overflow) |
| Tool-calling | Verified live to honor the OpenAI tools param (a 200 response alone does not prove this — always test with a real tool round-trip) |
The LLM backend is deliberately backend-agnostic — it's selected purely by
LLM_BASE_URL / MODEL_NAME, so the same code has run on vLLM, MLX, and
llama.cpp. Swapping backends only changes those two settings.
History: an earlier pilot ran the model + services bare-metal on
10.4.7.11(2× Quadro RTX 4000, Rocky 9), installed under/opt/cdn-aias bare Python withllama-serveras a systemd unit. The current live system is containerized on the staging k3s cluster; the LLM stayed on the Mac mini.
Settings we use
Agent (FastAPI) — set via ConfigMap, except the LLM URL which is in the Secret:
| Setting | Staging value | Notes |
|---|---|---|
LLM_BASE_URL |
http://172.25.25.17:8001/v1 |
in secret.yaml (it's an endpoint, not really secret) |
MODEL_NAME |
typhoon2.5-qwen3-30b-a3b |
|
MAX_ITERATIONS |
4 |
max tool-loop steps per question (local default 6); lowered from 12 to stop runaway loops |
MCP_STREAMER_URL |
http://ai-mcp:8085/mcp |
where the agent reaches the MCP server |
| LLM client timeout | 60s |
hard timeout on each model call (loop.py) |
| Tool scoping | on | each question is routed to a relevant subset of tools, not all 27 (see FastAPI agent) |
| Per-request opt-out | added 2026-09-23 | ChatRequest.use_tools/system_prompt — both default to the row above (every existing caller unaffected). A caller whose questions have nothing to do with CDN analytics (CoreAPI's domain/ai package — translation + its AI hotel guide, see Concierge — Guest Chat, Translation & AI Hotel Guide) can skip both SYSTEM_PROMPT (~3,268 tokens) and the tool schema list (~10k tokens) entirely — measured to be 85-90% of a typical off-topic prompt, large enough on its own to threaten n_ctx even without concurrent load |
ClickHouse connection (shared by MCP tools and the agent's direct REST):
| Setting | Staging value | Notes |
|---|---|---|
CLICKHOUSE_HOST |
clickhouse.analytic-stg.svc.cluster.local |
in-cluster DNS |
CLICKHOUSE_DATABASE |
cdn_ai |
|
CLICKHOUSE_USER |
ai_readonly |
read-only |
CLICKHOUSE_SECURE |
false (staging in-cluster) / true behind the public Apisix TLS ingress |
clickhouse_connect doesn't infer TLS from the port |
| empty-result retry | 8 × 0.3 s | workaround for a past replication skew; adds up to ~2.4 s per empty query — a candidate to trim |
Ports
| Service | Port | Protocol |
|---|---|---|
| FastAPI agent | 8082 |
HTTP — /chat, /qoe/*, /security/*, /ui/* |
| MCP server | 8085 |
streamable-http — /mcp |
| LLM (llama.cpp) | 8001 |
HTTP — OpenAI-compatible /v1 |