Skip to content

AI — CDN-AI system

The CDN-AI system answers natural-language questions about real CDN telemetry. It is two services plus an LLM:

  • a FastAPI agent (:8082) — the entry point (chat + REST + UI) and the LLM tool-calling loop. See FastAPI agent.
  • an MCP server (:8085) — a standalone server exposing the ClickHouse analytics as callable tools. See MCP server.
  • an LLM backend — a llama.cpp server; the agent is the LLM's client.

Data lives in ClickHouse (cdn_ai database), fed by Vector from the real cproxy / elb / streamer logs — see the ClickHouse schema for the tables and the standard. For pilot status and the tools→table reference, see the CDN-AI pilot status. For known issues and how to debug, see Logging & known issues.

flowchart LR
  U[User] -->|/chat, /ui/*| A[FastAPI agent :8082]
  A -->|OpenAI tool-calling| L[llama.cpp LLM<br/>172.25.25.17:8001]
  A -->|MCP streamable-http| M[MCP server :8085<br/>27 tools]
  M -->|SQL| C[(ClickHouse<br/>cdn_ai)]
  A -->|direct SQL, non-LLM REST| C
  V[Vector] -->|cproxy / streamer logs| C

The agent talks to the LLM and to the MCP server; the LLM never touches ClickHouse directly — it only asks the agent to call MCP tools, and the agent runs them against ClickHouse.

Model

Value
Model typhoon2.5-qwen3-30b-a3b (Qwen3-30B-A3B family, Q4_K_M GGUF, ~18.5 GB)
Runtime llama.cpp (llama-server, OpenAI-compatible API)
Endpoint http://172.25.25.17:8001/v1 (Mac mini) — same backend for local and staging
Context window n_ctx = 40192 tokens (hard 400 exceed_context_size on overflow)
Tool-calling Verified live to honor the OpenAI tools param (a 200 response alone does not prove this — always test with a real tool round-trip)

The LLM backend is deliberately backend-agnostic — it's selected purely by LLM_BASE_URL / MODEL_NAME, so the same code has run on vLLM, MLX, and llama.cpp. Swapping backends only changes those two settings.

History: an earlier pilot ran the model + services bare-metal on 10.4.7.11 (2× Quadro RTX 4000, Rocky 9), installed under /opt/cdn-ai as bare Python with llama-server as a systemd unit. The current live system is containerized on the staging k3s cluster; the LLM stayed on the Mac mini.

Settings we use

Agent (FastAPI) — set via ConfigMap, except the LLM URL which is in the Secret:

Setting Staging value Notes
LLM_BASE_URL http://172.25.25.17:8001/v1 in secret.yaml (it's an endpoint, not really secret)
MODEL_NAME typhoon2.5-qwen3-30b-a3b
MAX_ITERATIONS 4 max tool-loop steps per question (local default 6); lowered from 12 to stop runaway loops
MCP_STREAMER_URL http://ai-mcp:8085/mcp where the agent reaches the MCP server
LLM client timeout 60s hard timeout on each model call (loop.py)
Tool scoping on each question is routed to a relevant subset of tools, not all 27 (see FastAPI agent)
Per-request opt-out added 2026-09-23 ChatRequest.use_tools/system_prompt — both default to the row above (every existing caller unaffected). A caller whose questions have nothing to do with CDN analytics (CoreAPI's domain/ai package — translation + its AI hotel guide, see Concierge — Guest Chat, Translation & AI Hotel Guide) can skip both SYSTEM_PROMPT (~3,268 tokens) and the tool schema list (~10k tokens) entirely — measured to be 85-90% of a typical off-topic prompt, large enough on its own to threaten n_ctx even without concurrent load

ClickHouse connection (shared by MCP tools and the agent's direct REST):

Setting Staging value Notes
CLICKHOUSE_HOST clickhouse.analytic-stg.svc.cluster.local in-cluster DNS
CLICKHOUSE_DATABASE cdn_ai
CLICKHOUSE_USER ai_readonly read-only
CLICKHOUSE_SECURE false (staging in-cluster) / true behind the public Apisix TLS ingress clickhouse_connect doesn't infer TLS from the port
empty-result retry 8 × 0.3 s workaround for a past replication skew; adds up to ~2.4 s per empty query — a candidate to trim

Ports

Service Port Protocol
FastAPI agent 8082 HTTP — /chat, /qoe/*, /security/*, /ui/*
MCP server 8085 streamable-http — /mcp
LLM (llama.cpp) 8001 HTTP — OpenAI-compatible /v1