The original problem no longer describes the system
This was filed when ingestion was one CLI command and app/Actions/Scraper/ was an empty directory waiting for an HTTP action. Neither is true now:
app/Actions/Scraper/ does not exist.
- Ingestion is
trails:ingest (app/Commands.ts:35, with an ingest:trails alias) backed by app/TrailIngestWorker.ts.
- It runs continuously in production under a systemd unit of its own — the
ingest site in config/cloud.ts — because building the US catalog is a multi-day job (~1,400 Overpass tiles at two requests a minute, plus 466 Forest Service and Park Service shards) that has to survive deploys and re-sync afterwards.
- That unit already answers
/ on loopback with live shard counts and per-source trail totals, which is how you check on it.
So the gap this issue names — "no HTTP/admin trigger for scheduled or remote scraping" — is partly filled by something better than a trigger: a worker that does not need triggering, and a status endpoint for watching it.
What is actually still missing
Only a write surface: nothing outside the box can ask for a region to be re-ingested. Whether that is worth building is a real question rather than an assumption, which is why this issue is now a decision rather than a task.
- If yes: a guarded
POST taking a region and a limit, admin-only, that enqueues shards for the existing worker rather than scraping inline — an HTTP request must not hold a multi-day job open.
- If no: close this, and note in the worker that re-ingestion is deliberately operator-driven.
Acceptance, if it is built
The original problem no longer describes the system
This was filed when ingestion was one CLI command and
app/Actions/Scraper/was an empty directory waiting for an HTTP action. Neither is true now:app/Actions/Scraper/does not exist.trails:ingest(app/Commands.ts:35, with aningest:trailsalias) backed byapp/TrailIngestWorker.ts.ingestsite inconfig/cloud.ts— because building the US catalog is a multi-day job (~1,400 Overpass tiles at two requests a minute, plus 466 Forest Service and Park Service shards) that has to survive deploys and re-sync afterwards./on loopback with live shard counts and per-source trail totals, which is how you check on it.So the gap this issue names — "no HTTP/admin trigger for scheduled or remote scraping" — is partly filled by something better than a trigger: a worker that does not need triggering, and a status endpoint for watching it.
What is actually still missing
Only a write surface: nothing outside the box can ask for a region to be re-ingested. Whether that is worth building is a real question rather than an assumption, which is why this issue is now a decision rather than a task.
POSTtaking a region and a limit, admin-only, that enqueues shards for the existing worker rather than scraping inline — an HTTP request must not hold a multi-day job open.Acceptance, if it is built
TrailIngestWorkerand returns immediately.