Raw + derived telemetry storage for CEV. Plain Python (stdlib sqlite3, no ORM)
and plain SQL. Runs on an edge device, reached over Tailscale. Exposes its
read-only tools over MCP (Streamable HTTP) for any client on the tailnet --
tern-llm's own harness, a CEV member's own Claude client, or evil-ui's
backend -- not just one fixed consumer.
Full design rationale (and the alternatives that were rejected) lives in
inference-agent/4.md in this same working directory. This README is the
"how it's built," that doc is the "why."
# 1. Load static reference data ONCE per track (skip if evil.db already has it)
docker compose run --rm mcp-server python -m evil.scripts.seed_reference_data turn "Turn 1" 42.0 -76.0 30
docker compose run --rm mcp-server python -m evil.scripts.seed_reference_data start-finish 42.0 -76.0 30
# 2. Run the MCP server -- what tern-llm, evil-ui, and CEV members' own
# Claude clients all connect to. Binds 0.0.0.0:8765, reachable over
# Tailscale automatically once the device is on the tailnet.
docker compose up --build
# 3. Separately: the live compiler against the real ROS2 topic. Builds from
# Dockerfile.ros2 (ros:humble-ros-base), not this repo's default
# python:3.12-slim Dockerfile -- see the note below.
docker compose --profile live up
# Or, offline / after a race: bulk-load a recorded session instead of step 3.
docker compose run --rm mcp-server python -m evil.scripts.ingest_recording <recording.csv-or-.db3> <run-id>
# And separately: register a NAS recording against a run, so a derived row
# can be joined back to the exact clip that covers it (find_nas_files).
docker compose run --rm mcp-server python -m evil.scripts.register_nas_file <run-id> <path-on-nas> <kind> <start-ts> <end-ts>
# Or upload a recording from a UI: `upload` (also started by the bare
# `docker compose up` above) exposes a write-only HTTP endpoint on :8766,
# kept off the mcp-server port on purpose -- evil-ui's backend proxies to
# this, or POST directly: curl -F run_id=<id> -F file=@recording.db3 http://<host>:8766/uploadAll of the above share the same evil-data volume (see docker-compose.yml),
so the CLI commands and the running mcp-server/upload services see the
same database.
live-compiler builds from Dockerfile.ros2 (ros:humble-ros-base), not
this repo's default Dockerfile -- mcp-server/upload stay on the lighter
python:3.12-slim image since they never import evil.ingestion.ros2_source
(the only module here that needs rclpy). It runs with network_mode: host
since ROS2's default DDS discovery relies on multicast, which isn't reliable
across an isolated docker bridge network. Real rclpy verification used a
separate mock publisher, see ../mock-daq -- confirmed working: mock-daq's
containerized publisher genuinely publishes over rclpy, and this image
genuinely builds, starts, imports Ros2Source, and writes evil.db to the
mounted /data volume (not the ephemeral container filesystem -- this
surfaced and fixed a real bug where run_live_compiler.py's --db default
never read EVIL_DB_PATH, so the volume mount was silently a no-op).
Publisher and compiler now have a real automated cross-container test --
../mock-daq/tests/e2e/smoke_cross_container_round_trip.sh builds both
images, seeds a real db, runs both containers on the same topic, and checks
for real ingested rows. It currently SKIPs (not fails) here: mock-daq's
own e2e test found this dev environment's DDS discovery doesn't work at
all, even for a publisher and subscriber in the same container/network
namespace (confirmed via the official ros2 topic list CLI hanging
indefinitely there too). That's a WSL2 networking limitation, not something
specific to crossing a container boundary -- expect this to behave
differently on a real Linux host (the actual NUC).
docker compose (the v2 plugin) is installed and confirmed working in this
environment now -- mcp-server, upload, and live-compiler have each been
built and started for real via docker compose up, not just syntax-checked.
schema/ SQL DDL, applied in filename order, idempotent
src/evil/
db.py connect() / connect_readonly() / apply_schema()
models.py RawSample -- the transport-independent tick shape
ingest.py writes a RawSample into raw tables + main_snapshot
geo.py haversine distance helper
registered_classifiers.py ALL_CLASSIFIERS -- the one list ingestion and the live compiler both use
ingestion/
base.py IngestionSource protocol (the ROS2/Zenoh seam)
normalize.py shared flatten + alias-match: every source routes through this
replay_source.py list-backed source, used by tests and offline replay
ros2_source.py today's live transport: rclpy subscriber on /spi_data
csv_file_source.py bulk historical ingestion from a CSV export
rosbag_file_source.py bulk historical ingestion from a rosbag2 .db3 file
classifiers/
base.py Classifier protocol + ClassifierSpec
registry.py topological sort by depends_on
runner.py incremental tick(): per-classifier cursor + margin
metrics.py shared energy/avg-speed helpers used by laps.py and straights.py
turns.py example classifier: GPS geofence turn segmentation
laps.py depends_on=["turns"]: energy, avg speed, turn count per lap
straights.py depends_on=["turns"]: the complement of turns, no open-state needed
tools/
get_turn.py "how was I in turn X"
compare_turn_instances.py "what could I have done better" (this turn, across attempts)
compare_laps.py same, across two laps (energy, avg speed, turn count, duration)
turn_name_matching.py digit-normalized fallback so "3" also resolves to "Turn 3"
list_runs.py, list_turns.py, list_laps.py, list_straights.py
browsing-shaped (paginated, no lookup needed) -- for evil-ui
nas_index.py register_nas_file() (write, CLI-only) + find_nas_files() (read, MCP tool)
read_only_sql.py constrained fallback: SELECT/WITH-only, row cap, step budget
registry.py binds tools to a connection, Ollama-shaped schemas (used by tests)
mcp_server.py exposes the tools above over MCP / Streamable HTTP
upload_server.py write-only HTTP /upload -- separate from mcp_server.py on purpose
scripts/
seed_reference_data.py load track_geometry / start_finish_line rows
ingest_recording.py bulk-load a recorded CSV or rosbag .db3 as historical data
register_nas_file.py register one NAS recording against a run
run_live_compiler.py the production entry point: live Ros2Source + ALL_CLASSIFIERS
tests/ pytest, one file per module above
tests/e2e/ bash scripts exercising the real running server, see below
Called a cursor here, not the industry-standard "watermark" (Spark/Flink's term for the same concept) -- this project also has an LLM harness, and "watermark" collides with LLM output-watermarking (SynthID and similar). Same mechanism, different name.
Each classifier declares its own lookback_margin_s (how long to wait before
it's willing to call a row "final") and its own depends_on (raw data, or
another classifier's name, forming a DAG). laps/straights depend on
turns without the runner changing. A new classifier is one file plus one
line in registered_classifiers.py -- see classifiers/turns.py for the
pattern to copy. Details and the alternatives considered (CDC, broker
offsets, dirty-flag columns) are in inference-agent/4.md section 4.
IngestionSource is the seam between however data arrives and how it's
stored. The actual wire format is JSON everywhere already -- /spi_data is
a std_msgs/String carrying a JSON snapshot, a CSV row is a flat JSON-shaped
object with string values, and a rosbag's recorded messages are the same
JSON strings CDR-encoded. ingestion/normalize.py handles all of it once:
flatten nested dicts to dotted keys, normalize casing/punctuation, alias-match
against canonical fields. Ros2Source, CsvFileSource, and
RosbagFileSource all route through it -- extending to a new field or a
future Zenoh payload shape means editing one alias table, not writing a new
per-format parser. If autonomy's Zenoh move stays inside ROS2 (rmw_zenoh),
Ros2Source likely doesn't change at all; if it goes native, a new
ZenohSource is one file that hands its JSON payload to the same
normalizer. See inference-agent/4.md section 6.
The tool functions live here (not in tern-llm or evil-ui) because
they're schema-coupled to these tables, and MCP colocated with EVIL on one
device makes this code EVIL's public API. Each tool call in mcp_server.py
opens its own short-lived connect_readonly() connection and runs the
blocking query via a thread offload -- never a connection shared across
calls, which wouldn't be safe for concurrent use. There is deliberately no
single-in-flight lock here (unlike tern-llm's /ask endpoint): that lock
protects one scarce local LLM inference slot, which doesn't exist on this
side, since external MCP clients run their own model inference elsewhere.
SQLite's WAL mode (already set in db.connect()) lets any number of
readers proceed alongside the single ingestion writer without blocking
each other.
Writes never go through MCP -- register_nas_file and bulk ingestion are
CLI-only, same reasoning as why read_only_sql is generic-but-read-only:
a bad write is a categorically worse failure than a bad read.
python3 -m venv .venv
.venv/bin/pip install -r requirements-dev.txt
.venv/bin/python -m pytest # unit, in-process (no network) -- includes a
# Hypothesis fuzz test for read_only_sql
./tests/e2e/smoke_mcp_server.sh # real server, real MCP client, real HTTP
./tests/e2e/smoke_mcp_server_tailscale.sh # same, over this machine's real tailnet IP
./tests/e2e/smoke_mcp_server_concurrency.sh # 15 real concurrent MCP clients against one server
./tests/e2e/smoke_ingest_recording.sh # real CLI subprocess, CSV
./tests/e2e/smoke_ingest_recording_rosbag.sh # same, rosbag .db3tests/test_read_only_sql_fuzz.py throws ~200 random/adversarial query
strings per run at read_only_sql (a mix of pure noise and real SQL
fragments -- DROP TABLE, ATTACH DATABASE, injection-shaped strings,
etc.), asserting the one property that actually matters: no input, however
malformed, ever changes a row. It still passes with the _DISALLOWED
keyword regex removed entirely (checked by hand) -- the connection's own
mode=ro backstops it at the SQLite driver level, genuine defense in depth,
not a gap in the test.
smoke_mcp_server_concurrency.sh fires 15 real, concurrent MCP client
sessions (mixed read_only_sql/list_runs) at one real running server,
proving the "each call opens its own connection, WAL mode lets readers
proceed alongside each other" design actually holds under concurrent load,
not just in isolation.
The tailscale variant is a genuine reachability proof where tailscale is
installed and connected (this dev machine is on CEV's real tailnet --
cev-nuc is visible in tailscale status), not a simulation. It skips
cleanly where tailscale isn't available. It proves the 0.0.0.0 bind serves
an external interface, not just loopback -- it does not prove cev-nuc
specifically can serve this, since that needs the server actually deployed
there. Not done here on purpose: deploying to or otherwise touching
cev-nuc is a real, shared team device, not something to do without it
being asked for.
- More classifiers beyond
turns/laps/straights-- "events" turned out to be a general term for whatever gets formulated next (voltage-sag / current-spike / power-mismatch anomaly detection ported fromRace-GPT/main.py'sprecheck()is the strongest concrete candidate, since that logic is already proven, just never persisted as EVIL rows) -- deliberately left out of this pass. - Running
LiveCompiler(scripts/run_live_compiler.py) against a realRos2Sourcein production. The service loop itself is built and tested (event-driven wake, bounded fallback interval, incremental classification proven via a delayedReplaySourceintests/test_run_live_compiler.py). A real mock publisher exists now (../mock-daq, matchinguc26_sensor_reader's exact telemetry shape) to exercise the liverclpyhalf specifically, but that round trip hasn't been run anywhere yet -- it needs ROS2 Humble on a matching Ubuntu 22.04 environment, seemock-daq/README.mdfor why (Humble targets Jammy, not this dev machine's Noble). - A watcher that calls
register_nas_fileautomatically oncetailscale-ros-telemetry's rosbag service finishes a recording -- the write path exists and is tested, but nothing triggers it yet besides a human (orRaceEngineerDashboard) running the CLI by hand. - Deploying this server to
cev-nucand confirming a client elsewhere on the tailnet can reach it -- verified here over this dev machine's own tailnet interface only.