Skip to content

CDN-AI — pilot OKR status

Status of the CDN-AI observability pilot (MCP server + LLM agent over the real JASTV/3BB CDN telemetry). Every commit hash/date below is from git log; every column/row count is from a live run_query against the staging cdn_ai cluster via castai.castis.io, verified 2026-09-22. The model backend is the llama.cpp 30B (typhoon2.5-qwen3-30b-a3b) at 172.25.25.17:8001, confirmed live to honor tool-calls the same day.

Already tracked separately (not repeated here): local + Mac-mini CDN setup, initial MCP server development, anomaly tooling and QoE scoring.

Operational runbook material (real node IPs, VPN/network paths, ArgoCD sync snapshots, Ansible workflows, ClickHouse audit scripts) stays in the cdn-ai repo itself — it's live-changing pilot state, not stable documentation, and the source explicitly warns it "will go stale." See cdn-ai/docs/production-reference/real-infra-onboarding.md for that.


Reference — MCP tools and pipelines

The stable part of the pipeline reference (tool→table map, pipeline→table map) — compiled from a live audit of the staging cluster, 2026-09-17.

Tools → table

cdn_ai.cproxy_events (+ elb_events via UNION) — security: suspicious_ip_activity, ip_request_history, top_suspicious_paths, suspicious_ip_range_activity, component_traffic_breakdown, pop_traffic_breakdown, pop_attack_summary, node_attack_summary, device_ip_correlation. QoE: cmcd_qoe_score, qoe_score_trend, ip_range_qoe_scores, device_qoe_scores, pop_qoe_scores.

cdn_ai.streamer_events — stream: find_new_streams, streams_with_recent_traffic, current_viewer_counts, stream_lifecycle_events. log: log_level_breakdown, top_source_functions. (count_active_streams/get_stream_status hit the streamer REST API directly, not ClickHouse.)

cdn_ai.cproxy_events — log: find_http_errors, top_clients.

All 27 tools target cdn_ai — security/QoE always did; stream/log tools were retargeted in 769d78b, 2026-09-20 (stream_bandwidth_trend retired rather than retargeted, since its underlying log lines are dropped by the streamer firehose filter).

Pipelines → table

Pipeline → table Status
ai_cproxy.toml cdn_ai.cproxy_events ✅ live
ai_elb.toml cdn_ai.elb_events ✅ live, 0 rows (collection gap, see OKR 2 above)
ai_streamer_events.toml cdn_ai.streamer_events ✅ live, 13,172 rows
demo pipelines (ai_cdn_access_*, ai_cproxy_stats, ai_security, ai_streamer) streamer_analytics.* legacy/duplicate — retire once cdn_ai fully confirmed

---

Standardized schema (deployed to staging cdn_ai)

The three per-component tables as actually deployed (DDL from 07-cdn-ai-migration.yaml; column counts confirmed live 2026-09-22: 25 / 22 / 12). cproxy_events and elb_events share a common base (timestamp/node/pop/component/client_ip/device_id/transaction_id/…) so the security tools' UNION ALL lines up; elb drops CMCD and adds node-selection columns; streamer_events is origin-side lifecycle only (no client_ip/CMCD).

-- cproxy_events (25 cols) — edge access + response outcome + CMCD
CREATE TABLE cdn_ai.cproxy_events (
  timestamp DateTime, node String, pop_name String, component String,
  client_ip String, device_id String, transaction_id String,
  client_port UInt32, event_type String DEFAULT 'request',
  path String, user_agent String, referer String, host String,
  log_injection_suspected UInt8,
  status_code UInt16, elapsed_ms Float32, sent_bytes UInt64, origin_key String,
  cmcd_session_id String, cmcd_object_type String, cmcd_bitrate_kbps UInt32,
  cmcd_top_bitrate_kbps UInt32, cmcd_buffer_length_ms UInt32,
  cmcd_buffer_starvation UInt8, cmcd_startup UInt8
) ENGINE = MergeTree() PARTITION BY toDate(timestamp)
  ORDER BY (timestamp, client_ip) TTL timestamp + INTERVAL 30 DAY;

-- elb_events (22 cols) — same base, no CMCD-response outcome, + node selection
CREATE TABLE cdn_ai.elb_events (
  timestamp DateTime, node String, pop_name String, component String,
  client_ip String, device_id String, transaction_id String,
  client_port UInt32, path String, user_agent String, referer String,
  host String, log_injection_suspected UInt8,
  elb_selected_node String, elb_node_index UInt16,
  cmcd_session_id String, cmcd_object_type String, cmcd_bitrate_kbps UInt32,
  cmcd_top_bitrate_kbps UInt32, cmcd_buffer_length_ms UInt32,
  cmcd_buffer_starvation UInt8, cmcd_startup UInt8
) ENGINE = MergeTree() PARTITION BY toDate(timestamp)
  ORDER BY (timestamp, client_ip) TTL timestamp + INTERVAL 30 DAY;

-- streamer_events (12 cols) — origin-side lifecycle (ingest/dirwatch)
CREATE TABLE cdn_ai.streamer_events (
  timestamp DateTime, node String, pop_name String, component String,
  module LowCardinality(String), level LowCardinality(String),
  source_file String, source_function String,
  stream_id String, event_type LowCardinality(String),
  state LowCardinality(String), message String
) ENGINE = MergeTree() PARTITION BY toDate(timestamp)
  ORDER BY (timestamp, stream_id) TTL timestamp + INTERVAL 30 DAY;

Live data as of 2026-09-22: cproxy_events 800 rows, streamer_events 13,172 rows (both flowing); elb_events 0 rows (ingestion not yet wired).


Productionize CDN-AI — two images, CI/CD, live on real k3s

Status: Complete — 2026-09-22

The single shared agent image was split into two deployable services (MCP server + agent), wired to GitLab CI that builds both and auto-bumps their tags in the GitOps repo, and is now running live on the real JASTV staging k3s cluster (reachable at castai.castis.io).

Accomplished:

  • Split one image into Dockerfile.mcp + Dockerfile.agent — d6ab8b6, 2026-09-16
  • GitLab CI builds ai-agent + ai-mcp as separate images — 7dd7666, 2026-09-16 (agent job f6a12c1, same day)
  • CI update_image_tag bumps both tags in playtelly-iac per push — cc2579d, 2026-09-16
  • ArgoCD deployed latest build to staging (staging-769d78ba) — b0a7678 (iac), 2026-09-20
  • Verified live: agent /health = ok and MCP run_query executes against real cdn_ai — live check, 2026-09-22

Evidence:

Tools → table reference -> Reference — MCP tools and pipelines


Persistent device identity + device-sharing detection

Status: Complete — 2026-09-22

Added a persistent per-device id (X-Device-Id header, distinct from the per-playback CMCD session id) that survives into ClickHouse, plus an MCP tool that flags one device seen from many IPs — the credential-sharing signal.

Accomplished:

  • X-Device-Id capture end-to-end + device_ip_correlation sharing detector — 654c7e7, 2026-09-16
  • device_id column landed in the standardized cproxy schema (live: present in cproxy_events) — 8c14598 (iac), 2026-09-19
  • Tool live on staging (security category, 27-tool MCP set) — verified 2026-09-22

Note:

  • Signal is only as good as device_id coverage; rows with no X-Device-Id are excluded, and real viewer traffic is still sparse (cproxy 800 rows) until the XFF/real-IP work lands.

Evidence:

Security tools over cproxy/elb -> Reference — MCP tools and pipelines


Strategic data — standardize ClickHouse schema + Vector collection

Status: In progress — 2026-09-22

Converged the ad-hoc single-table demo schema onto a common per-component shape (cproxy/elb/streamer) with response-outcome fields (elapsed_ms/status_code/sent_bytes/origin_key) and streamer lifecycle events. Schema is standardized and deployed for all three tables; collection is live for cproxy + streamer only — elb is not yet shipping.

Accomplished:

  • Response/outcome + device/txn columns + streamer_events table designed — 8c14598 (iac), 2026-09-19
  • DROP+recreate cproxy_events/elb_events on clean schema — ba5e732 (iac), 2026-09-19
  • Migration Job registered + versioned so ArgoCD actually runs it — 3e01078 / a4dee06 (iac), 2026-09-19
  • Streamer lifecycle pipeline → cdn_ai.streamer_events — c038336 (iac), 2026-09-20 (crash-fix d9250da same day)
  • Live-verified: cproxy_events 25 cols / elb_events 22 / streamer_events 12 — run_query on staging, 2026-09-22
  • Live-verified: cproxy 800 rows + streamer 13,172 rows flowing (latest 2026-09-21 17:18) — 2026-09-22

Not started / deferred:

  • elb collection: table exists but 0 rows live — elb runs cross-cluster on DigitalOcean, log shipping not wired — 2026-09-22
  • gslb + cproxy-shield: not instrumented (design gap only)

Stale-doc claims found while verifying — corrected in this pass:

  • COLLECTION-DESIGN.md said "streamer — designed, NOT built (deferred)"; it is built and flowing (13,172 rows). Corrected and the file committed (was previously local-only).
  • AI-PIPELINE-REFERENCE.md still called the stream/log tools "ORPHANED (retarget to cdn_ai)"; the retarget is done (769d78b). Header corrected to reflect it.

Evidence:

Pipelines → table -> Reference — MCP tools and pipelines

Stream/log tools retargeted to cdn_ai (corrected) -> Reference — MCP tools and pipelines


Agent chat reliability — scope tools + retarget to new schema

Status: Complete — 2026-09-22

Fixed the slow / "reached max steps" chat loop: the agent now routes each question to a small relevant tool subset instead of all 27, and every stream/log tool was moved off the retired cdn_access_logs table onto the per-component cdn_ai tables (retiring the dead bandwidth tool).

Accomplished:

  • MAX_ITERATIONS 12→4 to stop runaway loops — 6f6fb80 (iac), 2026-09-20
  • Tool-scoping + retarget stream/log tools to cdn_ai + retire stream_bandwidth_trend — 769d78b, 2026-09-20
  • Deployed live (staging-769d78ba) — b0a7678 (iac), 2026-09-20
  • Verified live: stream_lifecycle_events runs against streamer_events with NO dead-table error (proves the retarget is the running image) — 2026-09-22

Note (real gap found live):

  • On high-volume streamer_events, stream_lifecycle_events can return enough rows to blow the llama.cpp 40k context window (observed 43,695-token overflow on a 30-min query, 2026-09-22) — the tool needs a tighter row/char cap or aggregation before this is demo-safe.

Evidence:

Analytics pipelines → tables (what each parses) -> Reference — MCP tools and pipelines


Real client-IP / XFF survival through the edge

Status: Not started — 2026-09-22

Real viewer IPs are needed for POP/IP-range and sharing analysis, but the client IP does not survive Apisix's SNAT/XFF stripping today. Only researched — no code or config pushed; the underlying ingress/LB config is untouched.

Not started:

  • No externalTrafficPolicy: Local / dedicated L4 LB change made — 2026-09-22
  • Custom X-Real-Viewer-Ip header path only documented, not implemented — research in 84b0659, 2026-09-17

Nothing to migrate here yet — this is "not started," so there's no completion to document. The live network-path investigation (VPN routes, which node can reach which) stays in cdn-ai's own onboarding doc since it's operational state, not stable reference.


Production demo data from a JASTV node

Status: Not started — 2026-09-22 (tentative target 2026-09-30)

Drive real playback (manifests, CMCD, anomalies) at a JASTV production node so the pilot demos against genuine production traffic rather than staging test data. Nothing built yet — this depends on the elb collection + XFF cards above.

Not started:

  • No demo-data generation against a JASTV prod node exists — 2026-09-22
  • Blocked-behind: elb 0 rows (cross-cluster shipping) and real-IP survival, both open above

Same as above — nothing built yet to document. Live blocker status lives in cdn-ai's own STAGING-DEVOPS-TODO.md, which changes day to day.