Skip to content

Logging & known issues

The audit log

Every /chat question is traced step-by-step to audit.jsonl:

  • In the container: /app/api/logs/audit.jsonl (the agent runs with --app-dir api).
  • One JSON object per line. Event types include question, tool_scope (chosen categories + tool count), model_response, tool_call (name + arguments), tool_result / tool_error, and final_answer.

Read this first before guessing why an answer was wrong — it shows exactly which tools the model called, with what arguments, and what came back. The same tool steps are also returned inline in the /chat response as trace.

# staging (needs kubectl to the cluster)
kubectl -n <ns> exec deploy/ai-agent -- tail -n 40 /app/api/logs/audit.jsonl
# local
docker exec ai_agent tail -n 40 /app/api/logs/audit.jsonl

Other places to look when rows/answers are wrong:

  • Vector logs — if new/changed rows aren't landing in ClickHouse, a broken VRL transform fails the whole config, not just the new logic. Check the Vector pod logs for compile errors first.
  • ClickHouse — query cdn_ai directly (or via run_query) to separate "the data isn't there" from "the tool is wrong".

Known issues / bug log

Reliability (hosted agent)

Symptom Cause Status / fix
HTTP 000 / whole endpoint down briefly Single replica (replicas: 1) — any pod roll is an outage window Add a 2nd replica; roll happens on every image/config change
/chat returns 500 / non-JSON under load One LLM backend serializes (~20–40 s/question); concurrent calls queue past the 60 s client timeout Ask one question at a time; consider a faster/bigger backend
Some tool endpoints slow (2–5 s) Empty-result retry (8 × 0.3 s) in common.py on empty queries Trim the retry count; it's a workaround for a since-fixed replication skew
"model backend failed: exceed_context_size" Tool result too big for n_ctx = 40192 (e.g. stream_lifecycle_events over 13 k rows) Cap tool row/char output or aggregate before returning
Pod flips Ready but isn't really working tcpSocket readiness probes on agent/mcp only check the port is open, not that ClickHouse/MCP/LLM are reachable Use an HTTP /health readiness probe
Tools dashboard "failed to fetch / invalid agent" It's non-AI (direct REST → MCP → ClickHouse); its endpoints 200 when the pod is healthy — it fails with the same single pod when that rolls/busies Same fix as the reliability items above

Operational gotchas

Gotcha Detail
ConfigMap changes don't restart pods Bump the castis.io/config-revision pod-template annotation to force a roll (or run a Reloader). A ConfigMap-only change is otherwise invisible until restart.
Backend silently drops tools Some backends (MLX, a bad GGUF template) accept the tools param and ignore it — a 200 proves nothing. Always verify with a real tool round-trip after a backend/model swap.
n_ctx overflow is silent on some backends llama.cpp returns a clear 400; Ollama's default (2048) silently truncates. Garbled output → suspect context overflow before sampling.
ClickHouse TLS not inferred from port clickhouse_connect needs CLICKHOUSE_SECURE=true explicitly for the HTTPS/Apisix path; the port number doesn't imply it.
Small models get specific shapes wrong Counting active streams, "new" vs "recently active", "most viewers", self-verifying with hand-written SQL. Each is handled by a dedicated tool / SYSTEM_PROMPT rule — see FastAPI agent.

Cross-cluster collection gaps (data, not bugs)

Table State Why
cdn_ai.cproxy_events flowing cproxy runs on the same (nuc/staging) cluster as Vector
cdn_ai.streamer_events flowing same cluster
cdn_ai.elb_events 0 rows elb runs on the DigitalOcean cluster; the staging Vector can't read another cluster's pod logs. Needs a Vector on the DO side shipping to the nuc ClickHouse.