Skip to content

FastAPI agent

The agent (agent/api/, container image ai-agent, port 8082) is the entry point to the CDN-AI system. It does three jobs:

  1. The LLM tool-calling loop — POST /chat.
  2. Direct REST wrappers over the MCP tools — /qoe/*, /security/* (no LLM).
  3. Serves the UI pages — /ui/*.

It holds no tools of its own — it is the client of the MCP server and of the LLM. The LLM decides which tool to call; the agent executes it via MCP against ClickHouse and feeds the result back to the LLM.

POST /chat — the tool-calling loop

agent/api/loop.py runs an OpenAI-style tool-calling loop:

  1. Take the user question, build the message list with the SYSTEM_PROMPT.
  2. Scope the tools (see below) and hand only that subset to the model.
  3. Call the LLM (LLM_BASE_URL). If it returns tool_calls, execute each via the MCP server and append the results; otherwise return the answer.
  4. Repeat up to MAX_ITERATIONS steps (staging 4). If it never settles, return "reached the maximum number of steps."

Request / response:

curl -X POST https://castai.castis.io/chat \
  -H 'Content-Type: application/json' \
  -d '{"question": "Any suspicious IPs in the last 30 minutes?"}'
# -> {"answer": "...", "trace": [ {tool, arguments, result}, ... ]}

The trace array is the tool calls the model made — useful for demos and debugging (the same steps are also written to the audit log; see Logging & known issues).

Tool scoping

Handing the model all 27 tools every turn made it slow and unreliable (it would pick wrong tools and burn the whole step budget). So scope_tools() routes each question by keyword to a small relevant subset:

Category Example tools Keyword triggers (examples)
security suspicious_ip_activity, device_ip_correlation suspicious, attack, scan, device, sharing, ip
qoe cmcd_qoe_score, qoe_score_trend qoe, quality, buffer, bitrate, playback, cmcd
stream stream_lifecycle_events, find_new_streams stream, channel, lifecycle, viewer, live, ingest
log find_http_errors, log_level_breakdown error, status code, log level, 5xx, latency
query list_tables, describe_table, run_query always included

Result: a security question sees ~12 tools, a QoE question ~8, instead of 27. The chosen categories + tool count are logged as a tool_scope audit event. A question that matches nothing falls back to a broad set.

Small local models get certain question shapes reliably wrong (counting active streams, "newly created" vs "recently active", "most viewers", self-verifying a correct tool answer with hand-written SQL that then errors). These are handled with dedicated tools + explicit SYSTEM_PROMPT rules, not left to the model to reason out — prefer adding a tool/rule over hoping the model self-corrects.

Direct REST endpoints (non-LLM)

These call an MCP tool directly and return its JSON — no model involved. They back the MCP tools dashboard and are the fastest way to sanity-check a tool.

Group Endpoints
QoE /qoe/score, /qoe/trend, /qoe/devices, /qoe/pop-scores, /qoe/ip-ranges
Security /security/suspicious-ips, /security/top-paths, /security/ip-ranges, /security/component-breakdown, /security/pop-breakdown, /security/device-ip-correlation, /security/simulate

Most take ?minutes=N. Because they share the single agent pod, they fail together whenever that pod is rolling or busy (see Logging & known issues).

UI pages (/ui/*)

Served from agent/interface/ as static files on the same port:

Page Purpose
chat.html the chat interface
tools_dashboard.html runs every MCP tool via the REST endpoints above
player.html real dash.js playback → generates real CMCD/QoE
player_dual.html two devices on one page (device-sharing demo)
sim_player.html simulated traffic for test data

Run / deploy

  • Container: agent/Dockerfile.agent → image ai-agent; runs uvicorn-style FastAPI on 8082 (--app-dir api).
  • Staging: k8s Deployment ai-agent (services/ai/agent), one replica, config via ai-agent-config ConfigMap + ai-agent-secret Secret.
  • Config that matters: LLM_BASE_URL, MODEL_NAME, MAX_ITERATIONS, MCP_STREAMER_URL, and the CLICKHOUSE_* set — see the settings table.
  • Local: part of the ai.yml compose stack (make ai), same image.