SpecExec is a small Python runtime that hides read-only tool latency behind streamed LLM generation. It detects complete tool statements while an LLM is still producing a program, precomputes eligible calls, then replays the completed program authoritatively from a fresh namespace.
Exact speculative results are reused during replay. Cache misses execute normally. Speculative state is disposable and never determines final correctness.
- Incremental parsing of streamed restricted Python.
- Generic async Tool registry.
- Exact cache keys using tool name, canonical arguments, and snapshot version.
- Dependency-aware speculative execution.
- Sequential and concurrent speculation modes.
- In-flight result reuse without duplicate calls.
- Read-only safety gate for early execution.
- Final authoritative replay for all variables and writes.
- Global and per-tool concurrency limits, timeouts, budgets, and metrics.
- Optional OpenAI-compatible streaming adapter.
- Offline local RAG fixture and benchmark suite.
- Python 3.11+
- No required runtime dependencies for offline use.
- Optional openai package for live streaming.
git clone https://github.com/shabeeth2/specexec.git
cd specexec
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venv\Scripts\Activate.ps1
python -m pip install -e .
# Optional live adapter
python -m pip install -e ".[openai]"python -m specexec.cli --mode speculative_concurrent
python -m unittest discover -v
python -m specexec.benchmark --quickThe offline demo streams:
a = search("checkout latency")
b = search("payment failures")
c = fetch(a[0])
answer = cThe RAG corpus is an example fixture in examples/local_rag.py. The reusable runtime is in specexec/.
export OPENAI_API_KEY="your-api-key"
export OPENAI_MODEL="your-model"
python -m specexec.cli --query "What causes checkout latency?"Use --model and --base-url to override environment configuration. The model should return the restricted tool program and assign the final response to answer.
from specexec import Tool, run
async def search(query: str) -> list[str]:
...
tools = {
"search": Tool(
"search", search,
read_only=True,
snapshot_consistent=True,
version="search-index-v1",
timeout=10,
max_concurrency=2,
cost=0.01,
)
}
result = await run(
text_stream,
tools,
mode="speculative_concurrent",
max_concurrency=4,
speculative_budget=1.0,
)
print(result.answer)
print(result.metrics)Tool metadata includes name, handler, read_only, snapshot_consistent, version, timeout, max_concurrency, and cost. Only registered read-only tools may execute speculatively.
Supported: simple assignments, registered calls, optional await, JSON-like literals, lists, tuples, dictionaries, variable references, indexing, slicing, and one non-nested tool call per statement.
Rejected: imports, arbitrary Python, exec, eval, attribute access, control flow, function definitions, nested calls, mutation syntax, speculative writes, and semantic cache matching.
| Mode | During generation | Final execution |
|---|---|---|
| baseline | Buffer complete program | Sequential execution without speculation |
| speculative_seq | One speculative call at a time | Sequential replay with exact cache lookup |
| speculative_concurrent | Ready calls run concurrently | Sequential replay with exact cache lookup |
Legacy streaming_seq and streaming_dag names remain accepted as aliases.
- Parse complete statements from the stream.
- Precompute eligible read-only calls in an ephemeral cache.
- Parse and replay the completed program from an empty namespace.
- Perform an exact cache lookup for every final tool call.
- Reuse a valid hit; execute a miss normally.
- Treat speculative failure, timeout, cancellation, or mismatch as a miss.
- Run state-changing calls only during final replay.
- Discard the cache after the run.
SpecExec is not an operating-system sandbox. Register only handlers safe to expose to generated programs.
python -m specexec.benchmark --quick
python -m specexec.benchmark --repetitions 20 --sizes 5 10 20 --latencies 0.1 0.3 0.5 1 2 --output benchmark.jsonlThe benchmark compares independent, partially dependent, and sequential workloads with identical scripts and replay schedules. It reports p50/p95 latency, saved wall time, cache hits and misses, overlap, unused work, wasted duration, configured cost, and final-result agreement.
Measured local fixture result with five statements and 100 ms tool latency:
| Workload | Baseline | Speculative sequential | Speculative concurrent |
|---|---|---|---|
| Independent | 764 ms | 572 ms | 368 ms |
| Partially dependent | 748 ms | 483 ms | 312 ms |
| Sequential | 763 ms | 578 ms | 591 ms |
These are deterministic local measurements, not provider performance guarantees.
python -m unittest discover -v
python -m py_compile specexec/*.py examples/*.py tests/*.pyRepository layout:
specexec/ reusable runtime, CLI, adapter, benchmark
examples/local_rag.py deterministic example tool registry
tests/ runtime correctness tests
Preserve the invariant that speculation may improve latency but must never be required for correctness. Add focused tests, run the test suite and benchmark quick mode, and describe changes to safety, cache identity, timing, or final replay behavior.
Control-flow dependencies, shadow state, MCP and agent-framework adapters, resource telemetry, cancellation of provably unused work, and broader benchmark datasets.
Next-tool prediction, argument guessing, semantic cache matching, RLM recursion, and speculative writes are intentionally out of scope.
Licensed under the MIT License. See LICENSE for details.
The README follows common GitHub open-source documentation guidance by covering project purpose, installation, usage, API, development, contribution, and licensing.