CDN-AI — pilot OKR status
Status of the CDN-AI observability pilot (MCP server + LLM agent over the
real JASTV/3BB CDN telemetry). Every commit hash/date below is from
git log; every column/row count is from a live run_query against the
staging cdn_ai cluster via castai.castis.io, verified 2026-09-22. The
model backend is the llama.cpp 30B (typhoon2.5-qwen3-30b-a3b) at
172.25.25.17:8001, confirmed live to honor tool-calls the same day.
Already tracked separately (not repeated here): local + Mac-mini CDN setup, initial MCP server development, anomaly tooling and QoE scoring.
Operational runbook material (real node IPs, VPN/network paths, ArgoCD
sync snapshots, Ansible workflows, ClickHouse audit scripts) stays in the
cdn-ai repo itself — it's live-changing pilot state, not stable
documentation, and the source explicitly warns it "will go stale." See
cdn-ai/docs/production-reference/real-infra-onboarding.md for that.
Reference — MCP tools and pipelines
The stable part of the pipeline reference (tool→table map, pipeline→table map) — compiled from a live audit of the staging cluster, 2026-09-17.
Tools → table
cdn_ai.cproxy_events (+ elb_events via UNION) — security:
suspicious_ip_activity, ip_request_history, top_suspicious_paths,
suspicious_ip_range_activity, component_traffic_breakdown,
pop_traffic_breakdown, pop_attack_summary, node_attack_summary,
device_ip_correlation. QoE: cmcd_qoe_score, qoe_score_trend,
ip_range_qoe_scores, device_qoe_scores, pop_qoe_scores.
cdn_ai.streamer_events — stream: find_new_streams,
streams_with_recent_traffic, current_viewer_counts,
stream_lifecycle_events. log: log_level_breakdown,
top_source_functions. (count_active_streams/get_stream_status hit
the streamer REST API directly, not ClickHouse.)
cdn_ai.cproxy_events — log: find_http_errors, top_clients.
All 27 tools target cdn_ai — security/QoE always did; stream/log
tools were retargeted in 769d78b, 2026-09-20 (stream_bandwidth_trend
retired rather than retargeted, since its underlying log lines are
dropped by the streamer firehose filter).
Pipelines → table
| Pipeline | → table | Status |
|---|---|---|
ai_cproxy.toml |
cdn_ai.cproxy_events |
✅ live |
ai_elb.toml |
cdn_ai.elb_events |
✅ live, 0 rows (collection gap, see OKR 2 above) |
ai_streamer_events.toml |
cdn_ai.streamer_events |
✅ live, 13,172 rows |
demo pipelines (ai_cdn_access_*, ai_cproxy_stats, ai_security, ai_streamer) |
streamer_analytics.* |
legacy/duplicate — retire once cdn_ai fully confirmed |
---
Standardized schema (deployed to staging cdn_ai)
The three per-component tables as actually deployed (DDL from
07-cdn-ai-migration.yaml; column counts confirmed live 2026-09-22:
25 / 22 / 12). cproxy_events and elb_events share a common base
(timestamp/node/pop/component/client_ip/device_id/transaction_id/…) so the
security tools' UNION ALL lines up; elb drops CMCD and adds node-selection
columns; streamer_events is origin-side lifecycle only (no client_ip/CMCD).
-- cproxy_events (25 cols) — edge access + response outcome + CMCD
CREATE TABLE cdn_ai.cproxy_events (
timestamp DateTime, node String, pop_name String, component String,
client_ip String, device_id String, transaction_id String,
client_port UInt32, event_type String DEFAULT 'request',
path String, user_agent String, referer String, host String,
log_injection_suspected UInt8,
status_code UInt16, elapsed_ms Float32, sent_bytes UInt64, origin_key String,
cmcd_session_id String, cmcd_object_type String, cmcd_bitrate_kbps UInt32,
cmcd_top_bitrate_kbps UInt32, cmcd_buffer_length_ms UInt32,
cmcd_buffer_starvation UInt8, cmcd_startup UInt8
) ENGINE = MergeTree() PARTITION BY toDate(timestamp)
ORDER BY (timestamp, client_ip) TTL timestamp + INTERVAL 30 DAY;
-- elb_events (22 cols) — same base, no CMCD-response outcome, + node selection
CREATE TABLE cdn_ai.elb_events (
timestamp DateTime, node String, pop_name String, component String,
client_ip String, device_id String, transaction_id String,
client_port UInt32, path String, user_agent String, referer String,
host String, log_injection_suspected UInt8,
elb_selected_node String, elb_node_index UInt16,
cmcd_session_id String, cmcd_object_type String, cmcd_bitrate_kbps UInt32,
cmcd_top_bitrate_kbps UInt32, cmcd_buffer_length_ms UInt32,
cmcd_buffer_starvation UInt8, cmcd_startup UInt8
) ENGINE = MergeTree() PARTITION BY toDate(timestamp)
ORDER BY (timestamp, client_ip) TTL timestamp + INTERVAL 30 DAY;
-- streamer_events (12 cols) — origin-side lifecycle (ingest/dirwatch)
CREATE TABLE cdn_ai.streamer_events (
timestamp DateTime, node String, pop_name String, component String,
module LowCardinality(String), level LowCardinality(String),
source_file String, source_function String,
stream_id String, event_type LowCardinality(String),
state LowCardinality(String), message String
) ENGINE = MergeTree() PARTITION BY toDate(timestamp)
ORDER BY (timestamp, stream_id) TTL timestamp + INTERVAL 30 DAY;
Live data as of 2026-09-22: cproxy_events 800 rows, streamer_events
13,172 rows (both flowing); elb_events 0 rows (ingestion not yet wired).
Productionize CDN-AI — two images, CI/CD, live on real k3s
Status: Complete — 2026-09-22
The single shared agent image was split into two deployable services (MCP
server + agent), wired to GitLab CI that builds both and auto-bumps their
tags in the GitOps repo, and is now running live on the real JASTV staging
k3s cluster (reachable at castai.castis.io).
Accomplished:
- Split one image into
Dockerfile.mcp+Dockerfile.agent—d6ab8b6, 2026-09-16 - GitLab CI builds
ai-agent+ai-mcpas separate images —7dd7666, 2026-09-16 (agent jobf6a12c1, same day) - CI
update_image_tagbumps both tags inplaytelly-iacper push —cc2579d, 2026-09-16 - ArgoCD deployed latest build to staging (
staging-769d78ba) —b0a7678(iac), 2026-09-20 - Verified live: agent
/health= ok and MCPrun_queryexecutes against realcdn_ai— live check, 2026-09-22
Evidence:
Tools → table reference -> Reference — MCP tools and pipelines
Persistent device identity + device-sharing detection
Status: Complete — 2026-09-22
Added a persistent per-device id (X-Device-Id header, distinct from the
per-playback CMCD session id) that survives into ClickHouse, plus an MCP
tool that flags one device seen from many IPs — the credential-sharing
signal.
Accomplished:
X-Device-Idcapture end-to-end +device_ip_correlationsharing detector —654c7e7, 2026-09-16device_idcolumn landed in the standardized cproxy schema (live: present incproxy_events) —8c14598(iac), 2026-09-19- Tool live on staging (security category, 27-tool MCP set) — verified 2026-09-22
Note:
- Signal is only as good as
device_idcoverage; rows with noX-Device-Idare excluded, and real viewer traffic is still sparse (cproxy 800 rows) until the XFF/real-IP work lands.
Evidence:
Security tools over cproxy/elb -> Reference — MCP tools and pipelines
Strategic data — standardize ClickHouse schema + Vector collection
Status: In progress — 2026-09-22
Converged the ad-hoc single-table demo schema onto a common per-component
shape (cproxy/elb/streamer) with response-outcome fields
(elapsed_ms/status_code/sent_bytes/origin_key) and streamer
lifecycle events. Schema is standardized and deployed for all three tables;
collection is live for cproxy + streamer only — elb is not yet shipping.
Accomplished:
- Response/outcome + device/txn columns +
streamer_eventstable designed —8c14598(iac), 2026-09-19 - DROP+recreate
cproxy_events/elb_eventson clean schema —ba5e732(iac), 2026-09-19 - Migration Job registered + versioned so ArgoCD actually runs it —
3e01078/a4dee06(iac), 2026-09-19 - Streamer lifecycle pipeline →
cdn_ai.streamer_events—c038336(iac), 2026-09-20 (crash-fixd9250dasame day) - Live-verified:
cproxy_events25 cols /elb_events22 /streamer_events12 —run_queryon staging, 2026-09-22 - Live-verified: cproxy 800 rows + streamer 13,172 rows flowing (latest 2026-09-21 17:18) — 2026-09-22
Not started / deferred:
- elb collection: table exists but 0 rows live — elb runs cross-cluster on DigitalOcean, log shipping not wired — 2026-09-22
- gslb + cproxy-shield: not instrumented (design gap only)
Stale-doc claims found while verifying — corrected in this pass:
COLLECTION-DESIGN.mdsaid "streamer — designed, NOT built (deferred)"; it is built and flowing (13,172 rows). Corrected and the file committed (was previously local-only).AI-PIPELINE-REFERENCE.mdstill called the stream/log tools "ORPHANED (retarget tocdn_ai)"; the retarget is done (769d78b). Header corrected to reflect it.
Evidence:
Pipelines → table -> Reference — MCP tools and pipelines
Stream/log tools retargeted to cdn_ai (corrected) -> Reference — MCP tools and pipelines
Agent chat reliability — scope tools + retarget to new schema
Status: Complete — 2026-09-22
Fixed the slow / "reached max steps" chat loop: the agent now routes each
question to a small relevant tool subset instead of all 27, and every
stream/log tool was moved off the retired cdn_access_logs table onto the
per-component cdn_ai tables (retiring the dead bandwidth tool).
Accomplished:
MAX_ITERATIONS12→4 to stop runaway loops —6f6fb80(iac), 2026-09-20- Tool-scoping + retarget stream/log tools to
cdn_ai+ retirestream_bandwidth_trend—769d78b, 2026-09-20 - Deployed live (
staging-769d78ba) —b0a7678(iac), 2026-09-20 - Verified live:
stream_lifecycle_eventsruns againststreamer_eventswith NO dead-table error (proves the retarget is the running image) — 2026-09-22
Note (real gap found live):
- On high-volume
streamer_events,stream_lifecycle_eventscan return enough rows to blow the llama.cpp 40k context window (observed 43,695-token overflow on a 30-min query, 2026-09-22) — the tool needs a tighter row/char cap or aggregation before this is demo-safe.
Evidence:
Analytics pipelines → tables (what each parses) -> Reference — MCP tools and pipelines
Real client-IP / XFF survival through the edge
Status: Not started — 2026-09-22
Real viewer IPs are needed for POP/IP-range and sharing analysis, but the client IP does not survive Apisix's SNAT/XFF stripping today. Only researched — no code or config pushed; the underlying ingress/LB config is untouched.
Not started:
- No
externalTrafficPolicy: Local/ dedicated L4 LB change made — 2026-09-22 - Custom
X-Real-Viewer-Ipheader path only documented, not implemented — research in84b0659, 2026-09-17
Nothing to migrate here yet — this is "not started," so there's no
completion to document. The live network-path investigation (VPN routes,
which node can reach which) stays in cdn-ai's own onboarding doc since
it's operational state, not stable reference.
Production demo data from a JASTV node
Status: Not started — 2026-09-22 (tentative target 2026-09-30)
Drive real playback (manifests, CMCD, anomalies) at a JASTV production node so the pilot demos against genuine production traffic rather than staging test data. Nothing built yet — this depends on the elb collection + XFF cards above.
Not started:
- No demo-data generation against a JASTV prod node exists — 2026-09-22
- Blocked-behind: elb 0 rows (cross-cluster shipping) and real-IP survival, both open above
Same as above — nothing built yet to document. Live blocker status lives
in cdn-ai's own STAGING-DEVOPS-TODO.md, which changes day to day.