Planning index only. Architecture and acceptance tests stay in spec.md. C4 views: c4.md. Do not put hostnames, LAN addresses, or SKUs here.
| Phase | Status | What it is |
|---|---|---|
| 1 — inference host | Downstairs NVIDIA path is the live target; IDE box is MCP-only | Ollama on downstairs WSL GPUs; starter tags; no paper inference on the IDE machine |
| 2 — MCP bridge | Done | local-coding-slm stdio tools on the IDE workstation; Cursor / Copilot / Claude adapters |
| 3 — measure | Protocol + corpus + apply gate + CI; live rates need downstairs powered | Layered scoring; next live rows go through downstairs |
| T12 Part A downstairs | Next hardware work (early Nov 2026) | Second PC / WSL NVIDIA via SSH — was blocked on host power; examples/downstairs-wsl-gpu.md |
| T12 Part B Copilot A8 / Claude A9 | Config ready; operator clicks pending | Same-machine only; do not block on downstairs; docs/a8-a9-operator-checklist.md |
Order (do not invert): finish Phase 3 live measurement on downstairs (T12 Part A) before treating Halo as the next lab.
- First work — downstairs: power on, WSL Ollama on the usable NVIDIA GPUs,
ssh -Lfrom the IDE workstation, A4-class check, scrubbed live harness /run_eval.py --liverows. - Then — keep measuring Part A on downstairs through November.
- Later — Halo / 395 (~Black Friday): bring-up and A13 only after downstairs is measuring.
Public-safe downstairs notes: examples/downstairs-wsl-gpu.md.
Halo / Ryzen AI Max+ 395: planned acquire around Black Friday week Nov 2026. Claim Phase 4 only after A13. Until then Phase 3 live rates stay on downstairs. The IDE workstation runs MCP only — do not treat local Ollama there as a paper or Phase 3 path. Paper track: embabel-slm paper calendar.
Paper track: Nov 2026–Mar 2027 Zenodo preprint (DOI by 31 Mar 2027). Live
pass@1 / pass@end freeze by early January (downstairs default; Halo only if
A13 green); all paper numbers freeze by late February.
Use an AMD Ryzen AI Max+ 395 / Halo-class box as another private Ollama host. Target arrival: around Black Friday week November 2026. The MCP server stays on the workstation. Premium agents still plan and review.
Until the box exists and A13 passes, do not claim this phase — keep Phase 3 live rates on the downstairs host (T12 Part A). Do not skip downstairs to chase Halo.
This is not a second product. It is the same bridge with a different
OLLAMA_BASE_URL (or the same loopback URL behind ssh -L).
Public-safe notes: examples/halo-ryzen-ai.md.
- Stdio MCP on the workstation; no listening port
- Tools:
local_status,local_code,local_refactor,local_generate_tests,local_explain,local_review - Env contract:
OLLAMA_BASE_URL,OLLAMA_FAST_MODEL,OLLAMA_STRONG_MODEL,OLLAMA_NUM_CTX - Starter tags:
qwen3.5:9b/devstral-small-2 - No public tunnels; cloud agents out of scope
- Official Ollama library tags only; output treated as untrusted
- Deployment checker (
A12)
- Validate current Ollama on the Halo with its supported AMD backend; record whether ROCm or Vulkan is used. Confirm accelerated placement after a short chat before selecting larger models.
- Pull the starter pair. Leave larger coding tags for a later benchmark.
- Reach the API from the workstation. Prefer
ssh -N -T -o ExitOnForwardFailure=yes -L 127.0.0.1:11436:127.0.0.1:11434 user@<halo-host>and keep Ollama on Halo localhost. Use any unused workstation port in place of11436. A private-interface bind plus a workstation-only firewall is optional. - Point gitignored
.envathttp://127.0.0.1:11436. Reload desktop MCP. Re-runscripts/check_deployment_safety.py. - Repeat acceptance A1–A7 and A11–A12. Run A4 (the workstation succeeds through SSH; an unauthorized LAN client fails). A8/A9 only if those clients are in use.
- After the starter pair feels usable, record a Halo benchmark (model + quant, effective context, tok/s, time to first token, peak unified memory, whether output is reviewable). Do not treat a large advertised context window as a reason to send the whole repo.
- Do not rename tools to
write_code/ollama_status.local_*is the contract already in Cursor, Copilot, and Claude configs. - Do not add
SLM_*environment aliases unless a later consumer cannot useOLLAMA_*. - Do not expose Halo in Cursor's model picker.
- Do not implement an automatic task classifier (still Phase 3).
- Do not start Halo install or wrapper changes until this phase is claimed.
These are not required to start Phase 4. Schedule them if the Halo path needs them:
- Wrapper preflight: reject a non-private
OLLAMA_BASE_URLand fail clearly when/api/tagsis down (no public fallback). - Reject oversized tool payloads (
files+taskover a configured character cap). - Keep dual fast/strong timeouts; a single
SLM_TIMEOUT_SECONDSis unnecessary unless operators ask for one knob.
- Automatic routing / classifiers (only after Phase 3 numbers exist)
- Larger-than-starter models on Halo unified memory
- Application-level shared secret, only if Ollama can do it without breaking local IDE use
- OpenRouter or Cursor OpenAI-base-URL override
- Public Ollama, ngrok, Cloudflare Tunnel
- Cloud-agent access to the private GPU in this home-lab profile
- A second MCP server just for Halo