Agents for collecting documents from multiple sources — local filesystems, WebDAV servers, web pages, and source code management platforms (GitHub, Gitea) — and writing them to a local download directory for downstream processing. Each document is written with a .meta.json sidecar capturing its MIME type and other metadata, and synchronization state is tracked locally so subsequent runs only fetch what changed.
MIME types are detected from file content (via puremagic) rather than trusting the filename extension, and the stored file is given the extension implied by its detected type. This means files with no extension (or the wrong one) are classified correctly and filtered consistently across every source. See File Typing and Filtering.
-
Filesystem Agent (
fs): Ingest documents from local directories- Recursive directory scanning
- Content-based MIME type detection (extension-less files supported)
- Configuration validation
- Status checking to avoid re-ingesting unchanged files
-
WebDAV Agent (
webdav): Ingest documents from WebDAV servers- Support for any WebDAV-compliant server (Nextcloud, ownCloud, SharePoint, etc.)
- Recursive directory scanning
- MIME type from the server
Content-Typeheader, falling back to content sniffing - Authentication support (username/password)
- Status checking to avoid re-ingesting unchanged files
- URL export for reviewing discovered files before ingestion
- URL-based ingestion from a curated file list with per-URL error tracking
- Skip hash check option for faster ingestion when re-downloading is acceptable
-
Web Agent (
web): Ingest web pages via HTTP- Fetch and ingest HTML content from URLs
- URL list support (inline, file, or single URL)
-
SCM Agent (
scm): Ingest files and issues from Git repositories- Support for GitHub and Gitea platforms
- Content-based file type filtering (extension-less files supported)
- Issue ingestion with comments (rendered as Markdown)
- Status checking to avoid re-ingesting unchanged files
-
Manifest Runner (
manifest): Declarative multi-source ingestion- YAML-based manifest files defining ingestion components
- Supports all agent types (fs, scm, webdav, web) in a single manifest
- Shared configuration (metadata, extensions, haiku-rag load config)
- Stale document removal (
delete_stale) across all components in a manifest - Cron-based scheduling via the REST API server
- Per-component credential and extension overrides
- Directory-level execution for running multiple manifests at once
-
haiku-rag Loading: Index downloaded documents into LanceDB
- Runs
haiku-ingester run-batchafter each manifest run - One per-source
.lancedbdatabase, configurable command and config file - Globally serialized — only one load runs at a time
manifest migrate/manifest vacuummaintain those databases- Metadata providers add each document's sidecar and a PDF's page count and
document information to its haiku-rag metadata;
manifest backfill-metadatafills them in for documents already indexed
- Runs
-
REST API Server: Run agents as a web service
- FastAPI-based HTTP endpoints for all operations
- Multiple authentication methods (API key, OAuth2 proxy)
- Interactive API documentation with Swagger UI
- Health check endpoint for monitoring
- Container-ready with Docker support
Requirements:
- Python 3.13 or higher
Documents are written to the local filesystem (DOWNLOAD_DIR). Indexing
into a vector store is handled by an optional haiku-rag load step (see
haiku-rag Loading), which runs haiku-ingester
against the downloaded files.
uv add soliplex.agentspip install soliplex.agentsgit clone <repository-url>
cd ingester-agents
uv syncThe agents use environment variables for configuration. Create a .env file or export these variables:
Agents write fetched documents to the download store -- the local filesystem by default, or S3-compatible object storage (see Object Storage).
# Directory where downloaded documents are written. Each run stores files
# under <DOWNLOAD_DIR>/<source>/, preserving the source directory structure.
# Every document is accompanied by a <filename>.meta.json sidecar containing
# its MIME type and any other available metadata.
DOWNLOAD_DIR=downloads
# Directory for local synchronization state (content hashes + SCM commit
# markers), one SQLite file per source. Stays on local disk even when documents
# are written to object storage -- SQLite cannot live on S3.
STATE_DIR=sync_stateThe agents use unified authentication settings that work across all SCM providers (GitHub, Gitea, etc.):
# SCM authentication token (GitHub personal access token or Gitea API token)
scm_auth_token=your_scm_token_here
# SCM base URL (required for Gitea, optional for GitHub)
# For Gitea: Full API URL including /api/v1
# For GitHub: Defaults to https://api.github.com if not specified
scm_base_url=https://your-gitea-instance.com/api/v1Examples:
For GitHub:
export scm_auth_token=ghp_YourGitHubToken
# scm_base_url not needed for public GitHubFor Gitea:
export scm_auth_token=your_gitea_token
export scm_base_url=https://gitea.example.com/api/v1# WebDAV server URL
WEBDAV_URL=https://webdav.example.com
# WebDAV authentication
WEBDAV_USERNAME=your-username
WEBDAV_PASSWORD=your-password
# Disable TLS certificate verification (default: true)
SSL_VERIFY=trueAll WebDAV credentials can also be provided via command-line options (--webdav-url, --webdav-username, --webdav-password), which override the environment variables.
# File extensions to include (default: md,pdf,doc,docx)
EXTENSIONS=md,pdf,doc,docx
# Logging level (default: INFO)
LOG_LEVEL=INFO
# API Server Configuration
SERVER_HOST=127.0.0.1
SERVER_PORT=8001
# Authentication (for API server)
API_KEY=your-api-key
API_KEY_ENABLED=false
AUTH_TRUST_PROXY_HEADERS=false
# Manifests the server runs: on a cron (SCHEDULER_ENABLED=true) and on demand
# via POST /api/v1/manifest/run (always available)
MANIFEST_DIR=/path/to/manifests
# SCHEDULER_RECONCILE_CRON="*/1 * * * *" # how often the dir is rescanned
# Pre-process spool: where each document is staged while its pre-process
# steps run (default: system temp dir; needs room for the largest document)
# PRE_PROCESS_SPOOL_DIR=/var/lib/ingester/spool
# haiku-rag loading (runs `haiku-ingester run-batch` after each manifest run)
HAIKU_LOAD_ENABLED=false
LANCEDB_DIR=/var/lib/lancedb # read by the haiku-rag config, which
# places <source>.lancedb under it
HAIKU_PATH=/etc/haiku # base dir for haiku-rag config files
# HAIKU_DEFAULT_CONFIG=haiku.rag.default.yaml # config filename under HAIKU_PATH
# HAIKU_LOAD_COMMAND=haiku-ingester --config={haiku_cfg} run-batch
# HAIKU_LOAD_TIMEOUT=1800
# HAIKU_LOAD_CWD=/var/lib/ingester # subprocess working dir (default: inherit)
# haiku-rag maintenance (`si-agent manifest migrate` / `vacuum` / `backfill-metadata`)
# HAIKU_MAINTENANCE_COMMAND=haiku-rag --config={haiku_cfg} {verb} # not backfill-metadata
# HAIKU_MAINTENANCE_TIMEOUT=3600
# haiku subprocess output (load and maintenance) is logged in parts, inside
# the run's span: one record per chunk, split at a line break where possible
# HAIKU_OUTPUT_CHUNK_BYTES=65536 # size of one logged part
# HAIKU_OUTPUT_FLUSH_SECONDS=30 # log pending output at least this often
# HAIKU_OUTPUT_MAX_BYTES=0 # cap on logged output per stream; 0 = none
# Tracing (see Tracing with Logfire below)
# LOGFIRE_TOKEN=... # or /run/secrets/logfire_token
# LOGFIRE_SERVICE_NAME=ingester-agents
# HAIKU_TRACE_WRAPPER=false # join haiku-ingester's spans to the agent's trace
# S3-compatible storage. S3_ENDPOINT_URL is shared between urls_file reads
# and the download store; the rest are the download store's credentials.
S3_ENDPOINT_URL=https://minio.example.com:9000
# S3_ACCESS_KEY_ID / S3_SECRET_ACCESS_KEY / S3_REGION -- omit to use the AWS
# default credential chain (environment, instance role, profile).
# S3_ALLOW_HTTP=true # required for an http:// endpoint
# Setting a bucket moves the download store into object storage. Accepts a
# bare name or an s3://bucket/prefix URI. See Object Storage below.
# DOWNLOAD_S3_BUCKET=my-documentsSee haiku-rag Loading for what these settings do. The haiku-rag config file needs a further set of its own — model and service endpoints, and the bucket variables in S3 mode — which fail the load when unset; see Variables the haiku-rag config needs.
With a Logfire token (LOGFIRE_TOKEN, or the logfire_token secret at
/run/secrets/logfire_token), the server sends its logs and traces to
Logfire. Without one, nothing is sent, and the spans below cost nothing.
- Each HTTP request is a span, except
/health, which the container healthcheck polls. - Each manifest run is one trace: a
manifest runspan with acomponentspan per component, pluslist scm uris/delete stalewhen they run. Every log line an agent writes nests under the component that wrote it, so every error from one run can be found under that run. - The haiku load that a run queues stays in the run's trace: a
haiku loadspan with the exact command, exit status and time spent in the queue, and apost-processspan per step. The subprocess output is logged in parts (haiku load <source> stderr output part <n>) inside it.
The CLI is opt-in: pass --otel before the command.
si-agent --otel manifest vacuum /manifests/test.yamlThe whole command becomes one trace under a cli span, whose message is
the command line (si-agent manifest vacuum /manifests/test.yaml).
Password-like option values are recorded as [redacted]. Without --otel,
the CLI sends nothing even with a token. serve ignores --otel: the
server always configures itself when a token is available.
haiku-ingester's own spans. The agent always passes its trace context
to haiku subprocesses as TRACEPARENT / TRACESTATE. A haiku-rag that
reads it joins the agent's trace by itself. For one that doesn't, set
HAIKU_TRACE_WRAPPER=true. The agent then starts a console-script command
through python -m soliplex.agents.traced_run, which attaches the context
before the CLI runs. It only works when haiku-rag is installed in the same
environment as the agent, and a command that isn't a console script there
runs unchanged.
The download store defaults to the local filesystem. Setting a bucket moves it
to S3-compatible object storage, where DOWNLOAD_DIR becomes the key prefix
rather than a directory:
DOWNLOAD_S3_BUCKET=my-documents
DOWNLOAD_DIR=ingester/downloads # now a key prefix
S3_ENDPOINT_URL=http://minio:9000 # omit for AWS
S3_ALLOW_HTTP=true # required for an http:// endpointDocuments then land at s3://my-documents/ingester/downloads/<source>/<path>,
with the same .meta.json sidecar beside each one. No extra install step:
obstore is a plain dependency, so object storage is available in every
install and only the configuration decides whether it is used.
Credentials come from S3_ACCESS_KEY_ID / S3_SECRET_ACCESS_KEY /
S3_REGION, or are omitted entirely to use the AWS default chain (environment,
instance role, profile).
DOWNLOAD_S3_BUCKET is named deliberately rather than S3_BUCKET: the mode is
inferred from its presence, so a generic variable that some other service in
the same environment happens to export must not be able to silently redirect
every download.
The bucket accepts a full s3:// URI as well as a bare name, matching the
S3_BUCKET value the haiku-rag configs interpolate, so one deployment
variable can feed both the reader and the writer:
DOWNLOAD_S3_BUCKET |
DOWNLOAD_DIR |
Documents land at |
|---|---|---|
my-documents |
ingester/downloads |
s3://my-documents/ingester/downloads/<source>/ |
s3://my-documents |
ingester/downloads |
s3://my-documents/ingester/downloads/<source>/ |
s3://my-documents/ingester |
downloads |
s3://my-documents/ingester/downloads/<source>/ |
A prefix on the bucket nests DOWNLOAD_DIR beneath it, so all three
spellings above address the same objects -- and are treated as the same
target, sharing one state file rather than re-fetching everything because a
prefix moved between two variables. A non-s3 scheme is rejected when
configuration is read rather than at the first write.
A blank DOWNLOAD_S3_BUCKET means the same thing as an absent one: local
disk. Clearing the value is therefore how object storage is disabled from a
compose .env, where a key is always present once the compose file
references it:
DOWNLOAD_S3_BUCKET= # object storage off, downloads go to diskWhitespace counts as blank, so a stray trailing space disables the store rather than failing the boot. Leading and trailing space is stripped from a real value for the same reason.
STATE_DIR stays local. The per-source SQLite files hold the content
hashes that drive incremental ingestion, and SQLite cannot live on object
storage. Moving documents to S3 does not by itself make the agent stateless: a
persistent volume is still required for state.
Where documents go is chosen per installation, not per manifest: every
source in an installation uses the same store. To move an installation to
object storage, set DOWNLOAD_S3_BUCKET and point the haiku-rag load at a
config whose source stanza reads from S3, either for the whole installation
with HAIKU_DEFAULT_CONFIG=haiku.rag.s3.yaml or for one manifest with its
haiku_config. DOWNLOAD_URI is injected for every load and holds the
resolved base URI in both modes, so a source stanza with uri: ${DOWNLOAD_URI} works either way.
A source's SQLite state file is qualified by its target, so after a switch every source opens fresh state and re-fetches its documents from upstream into the new store. On the indexing side:
- The document URIs change (
file://…becomess3://…), and the indexer keys its own state by URI, so every document is re-converted and re-embedded. This is the real cost of a switch. - If the haiku-rag source stanza omits
id:, its identity is derived from the target, so switching replaces one source with a different one and the old documents are never cleaned up. Setid:explicitly -- then a switch is self-cleaning -- or drop and rebuild each source's.lancedb.
Manifests used to accept a per-manifest download_store block. It was
applied by temporarily rewriting the shared settings, which the haiku load
and its post-process callbacks -- running later -- never saw, so they read
the installation default instead. It has been removed: a manifest that still
sets it is rejected with a message saying so.
For large repositories or rate-limited APIs, you can use the git command-line for file synchronization instead of API calls. This clones the repository locally and reads files from the filesystem.
# Enable git CLI mode
scm_use_git_cli=true
# Optional: Custom directory for cloned repos (default: system temp directory)
scm_git_repo_base_dir=/var/lib/soliplex/repos
# Optional: Timeout for git operations in seconds (default: 300)
scm_git_cli_timeout=600How it works:
- First sync: Clones the repository to a local temp directory (shallow clone, single branch)
- Subsequent syncs: Pulls latest changes using
git pull --ff-only - Pull failure: If pull fails, deletes the local clone and re-clones
- After sync: Runs
git clean -fdto remove untracked files
Notes:
- Issues are still fetched via API (git doesn't provide issue data)
- Requires git to be installed in the runtime environment
- The Docker image includes git by default
- All credentials are masked in log output for security
Security: Git CLI mode uses strict input sanitization to prevent command injection. Only alphanumeric characters, dashes, underscores, dots, and forward slashes are allowed in repository names and paths.
All ingestion runs through manifests: YAML files that name a source
and the components (filesystem, SCM, WebDAV, web) feeding it. Run them from
the CLI with si-agent manifest run, or let the server run them on a cron
schedule or on demand (POST /api/v1/manifest/run).
The CLI tool si-agent has five command groups:
manifest: Run manifests, and maintain the haiku-rag databases they loadfs: Inspect a local directory before ingesting itscm: Inspect a Git repository and manage its incremental sync statewebdav: Inspect a WebDAV directory, and export its file list for a manifestserve: REST API server: manifest scheduler, on-demand runs, and inspection routes
Write a manifest with an fs component:
# docs.yml
id: local-docs
name: local docs
source: my-source-name
components:
- name: docs
type: fs
path: /path/to/documentsThen run it:
si-agent manifest run docs.yml --no-loadThat's it! The runner automatically:
- Scans the directory
- Builds the inventory
- Validates files
- Writes new and changed documents to the download store, and removes ones no longer in the directory
The fs commands take the document directory and build the inventory by
scanning it. To review what a run would do first:
1. Preview the inventory
si-agent fs build-config /path/to/documentsPrints the inventory as JSON — paths, hashes, sizes, and detected MIME types — without writing anything. Redirect it to a file if you want to keep a copy.
2. Validate
Check which files are supported:
si-agent fs validate-config /path/to/documents3. Check status
See which files need to be ingested:
si-agent fs check-status /path/to/documents my-source-nameAdd --detail to see the full list of files:
si-agent fs check-status /path/to/documents my-source-name --detailThe status check compares file hashes against the local sync state:
- new: File doesn't exist in local state
- mismatch: File exists but content has changed
- match: File is unchanged (will be skipped during the run)
To narrow which files are considered, set extensions in the manifest (or
EXTENSIONS) rather than editing an inventory by hand.
Ingest both files and issues from a repository with an scm manifest
component. Issues are rendered as Markdown documents with their comments.
id: my-repo
name: my repo
source: my-repo
components:
- name: repo
type: scm
platform: github # or gitea (needs base_url or SCM_BASE_URL)
owner: myorg
repo: my-repo
incremental: true # commit-based sync after the first full run
# branch: main
# content_filter: all # all | files | issuessi-agent manifest run my-repo.yml --no-loadFiles and issues are written under <DOWNLOAD_DIR>/<source>/. Issues are
saved as Markdown (.md) documents. See example-manifests/scm.yml.
With incremental: true, the first run performs a full sync and records the
latest commit; later runs process only files changed since then, which cuts
API calls and bandwidth substantially (see
Incremental Sync).
# List issues
si-agent scm list-issues github myorg/my-repo
si-agent scm list-issues gitea admin/my-repo
# List repository files
si-agent scm get-repo github myorg/my-repo
si-agent scm get-repo gitea admin/my-repoView and manage the incremental sync state for a repository:
# View current sync state
si-agent scm get-sync-state gitea admin/my-repo
# Reset sync state (forces a full sync on the next manifest run)
si-agent scm reset-sync gitea admin/my-repoThe WebDAV agent ingests documents from WebDAV servers (like Nextcloud, ownCloud, SharePoint, etc.).
Use a webdav manifest component. It takes exactly one of path (scan a
directory recursively), urls (an inline list of files) or urls_file (a
URL list file, read from a local path, S3, or the WebDAV server itself):
id: webdav-docs
name: WebDAV Documents
source: my-source-name
components:
# Scan an entire WebDAV directory recursively
- name: shared-drive
type: webdav
url: https://webdav.example.com
path: /documents
# Ingest only a curated list of files
- name: curated-files
type: webdav
url: https://webdav.example.com
urls_file: urls.txtexport WEBDAV_USERNAME=your-username
export WEBDAV_PASSWORD=your-password
si-agent manifest run webdav.yml --no-loadEach URL in a URL list is processed independently: if a file fails to
download, the error is recorded and processing continues with the rest. See
example-manifests/webdav.yml for every source form.
1. Export URLs
Scan a WebDAV directory and export discovered file URLs to a file, for review
or to use as a component's urls_file. This uses only directory listing
(PROPFIND) and does not download file content:
si-agent webdav export-urls /documents urls.txt \
--webdav-url https://webdav.example.com \
--webdav-username user \
--webdav-password passThe output file contains one absolute WebDAV path per line:
/documents/report.md
/documents/sub/readme.pdf
/documents/notes.docx
Only files matching the configured EXTENSIONS filter are included. Edit the
list down to the files you want, then point a manifest component's
urls_file at it.
2. Validate Configuration
Check if files are supported (downloads files to compute hashes):
si-agent webdav validate-config /documents \
--webdav-url https://webdav.example.com \
--webdav-username user \
--webdav-password pass3. Check Status
See which files need to be ingested:
si-agent webdav check-status /documents my-source-name \
--webdav-url https://webdav.example.com \
--webdav-username user \
--webdav-password passAdd --detail flag to see the full list of files.
The manifest runner executes declarative YAML manifests that define multi-source ingestion jobs. A single manifest can combine filesystem, WebDAV, SCM, and web components under a shared source and configuration.
Run a single manifest file:
si-agent manifest run /path/to/manifest.ymlRun all manifests in a directory:
si-agent manifest run /path/to/manifests/Output results as JSON:
si-agent manifest run /path/to/manifest.yml --jsonMaintain the databases those manifests load into (see Database Maintenance):
si-agent manifest migrate # every manifest in $MANIFEST_DIR
si-agent manifest vacuum --dry-run # print the commands without running them
si-agent manifest backfill-metadata --missing page_count --check # what a back-fill would changeA manifest file defines the ingestion source, optional shared configuration, and one or more components:
id: my-ingestion
name: My Document Ingestion
source: my-source-name
schedule:
cron: "0 0 * * *"
config:
metadata:
project: my-project
extensions:
- md
- pdf
delete_stale: true
components:
- name: local-docs
type: fs
path: /path/to/documents
- name: web-pages
type: web
urls:
- https://example.com/page1
- https://example.com/page2
- name: repo-docs
type: scm
platform: github
owner: myorg
repo: my-repo
incremental: true
- name: shared-drive
type: webdav
url: https://webdav.example.com
path: /documentsUnknown keys are rejected at every level (top level, config, schedule,
post-process steps and components), so a typo such as extentions is a
validation error rather than a setting silently left at its default. The
server logs a rejected manifest once, at ERROR, and doesn't run it; si-agent manifest run exits 1. metadata and a post-process step's kwargs are
free-form and accept any keys.
Top-level fields:
- id (required): Unique identifier for the manifest. Must be unique across all manifests when running from a directory.
- name (required): Human-readable name for display and logging.
- source (required): Source name; also the per-source folder name under
DOWNLOAD_DIR(sanitized for filesystem safety). - schedule: Optional cron schedule for automated execution via the REST API server.
- cron: Cron expression (e.g.,
"0 0 * * *"for daily at midnight).
- cron: Cron expression (e.g.,
- config: Optional shared configuration applied to all components.
- metadata: Key-value pairs attached to all ingested documents.
- extensions: File extensions to include (overrides the global
EXTENSIONSsetting). - delete_stale: Remove locally-stored documents that no longer appear in any component (default: true). See Stale Document Removal below.
- haiku_config: Override the haiku-rag config file used when loading this manifest's source. Absolute paths are used as-is; relative values resolve under
HAIKU_PATH. Defaults to${HAIKU_PATH}/haiku.rag.default.yaml. See haiku-rag Loading. - pre_run: Ordered steps run once before any component; one may skip the run. See Pre-run steps below.
- pre_process: Ordered steps run on each new or changed document before it is stored; one may skip or modify it. None run unless listed. See Pre-process steps below.
- post_process: Ordered callbacks run after the haiku-rag load completes. See Post-process callbacks below.
- components (required): List of ingestion components (see below).
Filesystem (fs):
- name (required): Component name (must be unique within the manifest).
- path (required): Path to a local directory.
- extensions: Override extensions for this component.
- metadata: Additional metadata merged with config-level metadata.
Web (web):
- name (required): Component name.
- url: Single URL to fetch.
- urls: List of URLs to fetch.
- urls_file: Path to a file containing URLs (one per line). Supports local paths,
s3://bucket/keyURLs, andhttp(s)://WebDAV URLs. - Exactly one of
url,urls, orurls_filemust be specified. - extensions: Override extensions for this component.
- metadata: Additional metadata merged with config-level metadata.
SCM (scm):
- name (required): Component name.
- platform (required):
githuborgitea. - owner (required): Repository owner or organization.
- repo (required): Repository name.
- incremental: Use commit-based incremental sync (default: false).
- branch: Branch to sync (default:
main). - content_filter: What to ingest:
all,files, orissues(default:all). - base_url: Override SCM base URL (uses
scm_base_urlenv var if not set). - auth_token: Override auth token name (resolved via Docker secrets or env vars).
- extensions: Override extensions for this component.
- metadata: Additional metadata merged with config-level metadata.
WebDAV (webdav):
- name (required): Component name.
- url (required): WebDAV server URL.
- path: WebDAV directory path to scan recursively.
- urls: List of specific WebDAV file paths to ingest.
- urls_file: Path to a file containing WebDAV URLs (one per line). Supports local paths,
s3://bucket/keyURLs, andhttp(s)://WebDAV URLs (fetched using the same WebDAV credentials). - Exactly one of
path,urls, orurls_filemust be specified. - username: Override WebDAV username (resolved via Docker secrets or env vars).
- password: Override WebDAV password (resolved via Docker secrets or env vars).
- extensions: Override extensions for this component.
- metadata: Additional metadata merged with config-level metadata.
Settings are resolved in the following order (highest priority first):
- Component-level settings (e.g.,
extensionson a component) - Manifest config-level settings (e.g.,
config.extensions) - Global environment settings (e.g.,
EXTENSIONSenv var)
For metadata, config-level and component-level values are merged, with component values taking precedence for duplicate keys.
When delete_stale: true is set in a manifest's config block, the runner removes locally-stored documents that no longer appear in any of the manifest's components. This keeps the download directory in sync with the actual source data.
How it works:
- All components execute sequentially, collecting every discovered URI and its hash.
- After all components complete, the consolidated URI set is reconciled against the actual download folder for the source (all components in a manifest share one source, hence one download folder and state DB). Reconciliation is two-pass:
- State pass: any document tracked in local state whose URI is not in the consolidated set has its file,
.meta.jsonsidecar, and state entry deleted. - Disk sweep: the source download folder is then walked, and any file (plus its sidecar) that doesn't back a surviving URI is deleted — catching orphans that were never tracked in state (e.g. files left behind by an earlier run), not just files with a state row.
- State pass: any document tracked in local state whose URI is not in the consolidated set has its file,
WebDAV 404 handling:
- A WebDAV file that returns 404 (Not Found) during download is treated as a definitive removal, not a transient error. When
delete_staleis on, its local copy is deleted (via the reconcile above) even if it still appears in a stale listing. A 404 does not block the reconcile. - This is distinct from transient errors (timeouts, 5xx) — see Safety below.
Safety:
- If any component raises, hits an unknown type, or reports a transient per-file error (timeout / 5xx),
delete_staleis skipped entirely for that manifest run. This prevents accidental deletions when the URI set may be incomplete. (A 404 is a removal signal, not a transient error, so it does not trigger this skip.) - Components that succeed still have their documents ingested normally — only the stale deletion step is skipped.
Example:
id: synced-docs
name: Synced Documentation
source: docs-source
config:
delete_stale: true
components:
- name: local-docs
type: fs
path: /data/docs
- name: shared-drive
type: webdav
url: https://webdav.example.com
path: /shared/docsIf a file is removed from /data/docs or from the WebDAV server (dropped from the listing, or returning 404 on fetch), the next manifest run detects that its URI is no longer present and deletes it — and its sidecar — from the download directory.
Note: SCM components using incremental: true only return files changed since the last sync, not the full file listing. When delete_stale is enabled with incremental SCM components, the stale detection may not have complete URI coverage for those components. Consider using full inventory mode (incremental: false) when delete_stale is needed with SCM sources.
When the REST API server is started with SCHEDULER_ENABLED=true and MANIFEST_DIR is set, manifests with a schedule block are automatically registered as cron jobs:
export SCHEDULER_ENABLED=true
export MANIFEST_DIR=/path/to/manifests
si-agent serveThe server loads all manifests from the directory at startup, validates that all manifest IDs are unique, and then:
- Manifests with a
scheduleare registered and run when their cron expression is due. - Manifests without a
scheduleare run once when first seen (a fire-and-forget task), then never again unless triggered via the API (POST /api/v1/manifest/run).
Hot-reloading schedules:
The manifest directory is rescanned on a fixed interval (every minute by
default; configurable via SCHEDULER_RECONCILE_CRON), so changes take
effect without restarting the server:
- Added files are picked up on the next scan — new schedules register, and new unscheduled manifests run once.
- Removed files are unregistered and stop firing.
- Edited
scheduleblocks are re-read; the manifest is rescheduled to its new cron expression (the next fire is computed from the change time, not backfilled). Adding ascheduleto a previously unscheduled manifest starts scheduling it; removing theschedulestops it.
A manifest that is invalid or introduces a duplicate id mid-edit is skipped for that scan (logged) and retried on the next one, so a bad save never takes down the scheduler.
Execution behavior:
- At most one manifest runs at a time. Due manifests are handed to a
single-worker FIFO queue (
server/manifest_queue.py) that drains them one at a time, so different manifests never run concurrently. This bounds resource use, since a single manifest can already fan out across its components. Serialization is structural — the worker does one thing at a time — rather than enforced by a lock each caller has to remember to take. The scheduler andPOST /api/v1/manifest/runboth feed this same queue, so an on-demand run never overlaps a scheduled one. - Manifests due at the same time all run, in order. A cron firing while another manifest is in progress is queued behind it, not skipped, so no scheduled occurrence is silently lost.
- A repeat occurrence of a still-pending manifest coalesces. If a manifest comes due again while its previous run is still queued or running, the new occurrence folds into the pending one (logged) rather than queueing a redundant second run — the pending run reloads the manifest from disk and so covers the newer occurrence anyway. This bounds the queue at the number of manifests on disk, so a cron faster than its own runs cannot grow a backlog.
- Shutdown cancels the in-flight manifest. The worker is cancelled and awaited on shutdown rather than left orphaned; queued manifests are dropped and picked up again by the reconciler on the next start.
- Single process only. Because the cron state and run queue are held in
memory, scheduling relies on the server running as a single worker (see
Starting the Server). If you run multiple server
instances, enable
SCHEDULER_ENABLEDon only one of them, and send on-demand runs to that same instance: each instance has its own queue, so two instances can run the same manifest at once.
POST /api/v1/manifest/run queues a manifest from MANIFEST_DIR, by id, and
returns 202 Accepted:
curl -X POST http://localhost:8001/api/v1/manifest/run -F manifest_id=test-scm
# {"status": "queued", "manifest_id": "test-scm"}The run happens on the queue described above, so it waits behind any
manifest already running. "status": "already_queued" means the manifest
was already queued or running and this request folded into it. The run queue
starts with the server whether or not SCHEDULER_ENABLED is set; that setting
controls only cron scheduling and the startup run of unscheduled manifests.
GET /api/v1/manifest/queue lists the ids still queued or running.
Only manifests in MANIFEST_DIR can be run this way, never an arbitrary path,
so the API can ingest only what the operator has put there.
When HAIKU_LOAD_ENABLED=true, a haiku-rag load is queued after each
manifest run (scheduled, startup, or CLI). The load indexes the documents
that the manifest just wrote to ${DOWNLOAD_DIR}/<source>/ into a
per-source LanceDB database. The default command is:
haiku-ingester --config=${HAIKU_CFG} run-batchNote what is not on that command line: a database. The haiku-rag config places it. See Where the database comes from below — a config that does not place one per source is the one misconfiguration here that fails quietly.
-
One load at a time. Inside the server, loads are drained from a single global FIFO queue by one worker, so only one
haiku-ingesterprocess runs at any moment (a capacity constraint). The CLI achieves the same by running loads sequentially after each manifest. -
Command is fully configurable via
HAIKU_LOAD_COMMAND. Supported placeholders:{haiku_cfg},{db},{source},{lancedb_dir},{haiku_path}. The template is tokenized before substitution, so values containing spaces cannot inject extra arguments. -
Config file resolves from the manifest's
config.haiku_config(absolute path used as-is; relative resolved underHAIKU_PATH), falling back to${HAIKU_PATH}/${HAIKU_DEFAULT_CONFIG}(haiku.rag.default.yaml). -
Environment. The subprocess inherits the server's environment plus three injected variables:
SOURCE— the sanitized download-folder name, so a haiku-rag config usingroot: ${DOWNLOAD_DIR}/${SOURCE}resolves to the ingested documents, and one usingdatabases: {db: ${LANCEDB_DIR}/${SOURCE}.lancedb}gets a database per source.DOWNLOAD_DIR—settings.download_dir, so the path above resolves even when it was left at its default.DOWNLOAD_URI— the resolved base URI of the download store, set in both filesystem and S3 mode so one config form works either way.
Any other
${VAR}interpolated by the haiku-rag config (LANCEDB_DIRincluded, along with e.g.OLLAMA_BASE_URL,DOCLING1_BASE_URL,DOCLING2_BASE_URL,EMBEDDINGS_BASE_URL) must be present in the server's environment.
export HAIKU_LOAD_ENABLED=true
export LANCEDB_DIR=/var/lib/lancedb
export HAIKU_PATH=/etc/haiku
export SCHEDULER_ENABLED=true
export MANIFEST_DIR=/path/to/manifests
si-agent serveThe CLI honors the same HAIKU_LOAD_ENABLED default; override per
invocation with si-agent manifest run <path> --load / --no-load.
The agent does not choose the database. It passes --config and nothing else,
and haiku-rag resolves the location from lancedb.databases in that file. So
the config must place exactly one database, and must place a different one
per source — which is what ${SOURCE} is for:
lancedb:
databases:
db: ${LANCEDB_DIR:-/lancedb}/${SOURCE}.lancedbThe entry name (db above) is arbitrary and never leaves the configuration.
${SOURCE} in the location is what separates one source from another. Four
cases, and only one of them is loud:
lancedb.databases |
Result |
|---|---|
One entry interpolating ${SOURCE} |
Correct: one database per source |
One entry, no ${SOURCE} |
Every source loads into one shared database |
| Absent | Everything lands in haiku.rag.lancedb under storage.data_dir, resolved relative to HAIKU_LOAD_CWD |
| Two or more entries | haiku-ingester refuses to start: it writes one database |
The middle two run to completion and silently merge every source into one store, so check the resolved location in the load's log line the first time a config is deployed:
Starting haiku load for source 'synced-docs' -> /var/lib/lancedb/synced-docs.lancedb
LANCEDB_DIR is read twice, and the two readers must agree: the haiku-rag
config interpolates it to place the database, and the agent resolves
${LANCEDB_DIR}/<slug>.lancedb for that log line, the run report, and the
maintenance dedupe key. It is required even though the config is what actually
places the database. The two spellings differ for a source containing
whitespace — the agent slugifies (composite source →
composite-source.lancedb) while ${SOURCE} sanitizes (composite source.lancedb) — so prefer source ids without spaces.
Working examples for both storage modes are in
example-haiku-configs/.
haiku expands ${VAR} eagerly, when the config file is read — so a variable
that is unset or empty fails the load before any document is touched:
MissingEnvVarError: Config references unset or empty environment variable
${QA_MODEL}. Set it, or use ${QA_MODEL:-default} to provide a fallback.
There is no partial start and nothing to inspect afterwards, so the whole set
has to be present up front. These are what
example-haiku-configs/ reference; a config of your
own can of course need fewer or more.
| Variable | Supplied by | Required by | Purpose |
|---|---|---|---|
SOURCE |
injected per load | both examples | Sanitized source name. Separates each source's database, queue file and documents — do not set it yourself |
DOWNLOAD_DIR |
injected per load, from settings.download_dir |
both examples | Where the manifest run wrote the documents |
DOWNLOAD_URI |
injected per load | neither example | Base URI of the download store; available for a config that wants one form across both storage modes |
STATE_DIR |
environment | both examples | Holds the per-source ingester queue file |
QA_MODEL |
environment | both examples | Vision model used to identify documents |
QA_BASE_URL |
environment | both examples | OpenAI-compatible endpoint serving QA_MODEL |
OLLAMA_BASE_URL |
environment | both examples | Ollama endpoint |
DOCLING1_BASE_URL |
environment | both examples | First docling-serve instance |
DOCLING2_BASE_URL |
environment | both examples | Second docling-serve instance |
EMBEDDINGS_BASE_URL |
environment | both examples | Endpoint serving the embedding model |
S3_BUCKET |
environment | haiku.rag.s3.yaml |
Bucket URI holding the databases |
S3_REGION |
environment | haiku.rag.s3.yaml |
Region for that bucket |
DOWNLOAD_S3_BUCKET |
environment | haiku.rag.s3.yaml |
Bucket URI holding the documents; the same value the download store wrote with |
LANCEDB_DIR |
environment | optional in both (:-/lancedb) |
Base dir, or S3 key prefix, for the database |
INGESTER_AUTH_TOKEN |
environment | only if ingester.api.auth_token is uncommented |
Bearer token for the ingester's HTTP control plane |
Two notes on that table:
LANCEDB_DIRis only optional to the config. The examples default it to/lancedb, but ingester-agents itself refuses to run a load without it — it resolves${LANCEDB_DIR}/<slug>.lancedbfor the log line, the run report and the maintenance dedupe key. Set it, and keep the two in step.- Injected variables are not yours to set.
SOURCE,DOWNLOAD_DIRandDOWNLOAD_URIare overwritten per load from the run'sLoadContext; a value exported in the server's environment is replaced, not merged. Post-process callbacks, which load the same config in-process, see all three mirrored into the process environment for the duration of the callbacks, then restored.
${VAR} inside a YAML comment is never expanded — the file is parsed
first, and expansion runs over the parsed data — so a commented-out setting
costs nothing.
Three verbs operate on the per-source LanceDB databases rather than on the downloaded documents:
si-agent manifest migrate [PATH] [--json] [--timeout N] [--dry-run]
si-agent manifest vacuum [PATH] [--json] [--timeout N] [--dry-run]
si-agent manifest backfill-metadata [PATH] [--missing KEY]... [--content-type TYPE]...
[--filter SQL] [--db-name NAME] [--batch-size N]
[--check] [--no-attachments] [--json] [--timeout N] [--dry-run]-
migrateruns pending haiku-rag schema migrations;vacuumoptimizes and compacts the tables to reclaim disk space;backfill-metadatare-runs each source's haiku-ragmetadata_providerover documents already indexed (see Back-filling metadata). -
PATHis optional and defaults toall:PATHScope all(default)Every *.yml/*.yamlin$MANIFEST_DIRa manifest YAML file That manifest only a directory Every manifest in that directory allis a reserved word, so a file or directory literally namedallcannot be addressed by name. -
Command for
migrate/vacuumis configurable viaHAIKU_MAINTENANCE_COMMAND(defaulthaiku-rag --config={haiku_cfg} {verb}).backfill-metadataalways runspython -m soliplex.agents.haiku_backfillwith the agent's own interpreter. Placeholders:{verb},{haiku_cfg},{db},{source},{lancedb_dir},{haiku_path}. As with the load command, the template is tokenized before substitution, so values containing spaces cannot inject extra arguments. -
Config file, database, and environment resolve exactly as they do for a load: the config from
config.haiku_config(or${HAIKU_PATH}/${HAIKU_DEFAULT_CONFIG}) places the database (see Where the database comes from), and the parent environment plus injectedSOURCE/DOWNLOAD_DIR/DOWNLOAD_URIlet its${VAR}references resolve. Output is streamed to the log line by line. -
One at a time. Operations run strictly sequentially, the same capacity constraint that applies to loads.
-
Deduplicated. Manifests that share a source resolve to the same database; it is processed once and the rest are reported as skipped. The dedupe key is the agent's own
${LANCEDB_DIR}/<slug>.lancedb, not the location the config resolves to, so it matches reality only while the config places the database under${LANCEDB_DIR}by source. -
Timeout defaults to
HAIKU_MAINTENANCE_TIMEOUT(3600s — higher than the load timeout because a compaction can outlast a batch load) and can be overridden per invocation with--timeout. -
Exit code is 1 if any operation failed or timed out, so the verbs can be used directly in cron or a deploy script. A skipped duplicate is not a failure.
--dry-run resolves every target and prints the command lines that would
run, one per line, in execution order — nothing is spawned. Skips and
resolution failures appear as #-prefixed comments so the block stays
paste-safe:
$ si-agent manifest vacuum --dry-run
haiku-rag --config=/etc/haiku/haiku.rag.default.yaml vacuum
haiku-rag --config=/etc/haiku/haiku.rag.web.yaml vacuum
# composite source: skipped (duplicate db)Under --dry-run the exit code reflects resolution only, so a dry run
doubles as a config check. Add --json for the full per-target detail
(argv, resolved paths, return codes).
Note: nothing coordinates a CLI maintenance run with a load already running inside the server — the FIFO queue in
server/haiku_queue.pyonly serializes loads within that process. Run maintenance during a quiet window, or with the scheduler stopped.
A manifest can run three kinds of optional, ordered steps, named for when they fire:
| Hook | Fires | Once per | Can stop |
|---|---|---|---|
pre_run |
before any component runs | manifest run | the whole run (SKIP) |
pre_process |
inside each document write, before it is stored | new or changed document | that document (SKIP) |
post_process |
after the haiku-ingester load | load | nothing |
pre_run steps ─► components ─┬─► write ─► pre_process steps ─► store
│ (per new / changed document)
└─► stale reconcile ─► haiku load ─► post_process
Every step names a method -- a dotted import path, pkg.mod:func or
pkg.mod.func, importable in the agent's environment -- plus optional
kwargs. All pre_run and pre_process methods are imported before the run
starts, so a typo fails the manifest before any step (a "starting"
notification, say) has run.
config.pre_run runs once, before any component, for notifications and
pre-checks. Each step is called as method(context, **kwargs), where
context has manifest (a copy -- a step cannot change what runs), load
(the source's resolved download target, store and sidecars) and started_at.
config:
pre_run:
- method: soliplex.agents.manifest.pre_run_steps:notify_webhook
kwargs:
url_secret: INGEST_WEBHOOK_URL # docker secret / env var
secret_headers: { Authorization: INGEST_WEBHOOK_TOKEN }
on_error: continue # a failed notification must not block ingestion
timeout: 15
- method: soliplex.agents.manifest.pre_run_steps:check_free_space
kwargs: { min_free_mb: 2048 }A step returns "continue" (or None) to carry on, or "skip" to call the
run off -- optionally with a message, as ("skip", "maintenance window"). Use
PreRunStatus from soliplex.agents.manifest.pre_run for the values.
- A skipped run does nothing: no component, no stale reconcile, no
pre-processing, and no haiku load -- so no post-process either. The next
scheduled run happens as normal.
si-agent manifest runprints the manifest asSKIPPED by <method>: <message>and still exits0. on_errordecides what a step that raises (or outlives itstimeout) means:fail(the default) fails the manifest like a crashing component;skipturns it into a skip;continuelogs it and moves on.timeout(seconds, default 300,nullfor none) keeps a slow webhook from holding up the single-worker manifest queue. An async step is cancelled; a sync step runs in a thread and cannot be interrupted, so the wait is abandoned while the thread finishes.- Outcomes are kept in the source's state DB for the last 100 runs:
si-agent manifest pre-run-report <path|all> [--status skip] [--since <iso>]answers "why didn't this source update last night?".
Built-in steps (soliplex.agents.manifest.pre_run_steps):
notify_webhookPOSTs{"event": "manifest.started", "manifest_id", "manifest_name", "source", "started_at", "download_uri"}as JSON. Giveurl, orurl_secretnaming a docker secret / env var that holds it (incoming-webhook URLs are credentials);headersare sent as written,secret_headersvalues are resolved like an SCMauth_token. A non-2xx response raises, so pair it withon_error: continue.check_free_spaceskips the run when the download directory (local store) or the pre-process spool directory has less thanmin_free_mbfree (include_spool: falsechecks only the former). An S3 store has nothing to check and is reported as such.
Writing your own is a few lines -- for example, a maintenance-window flag:
from pathlib import Path
from soliplex.agents.manifest.pre_run import PreRunStatus
def unless_paused(context, *, flag="/etc/ingester/paused"):
if Path(flag).exists():
return PreRunStatus.SKIP, f"paused by {flag}"
return PreRunStatus.CONTINUEconfig.pre_process runs on each new or changed document -- unchanged
documents are never fetched, so they are never pre-processed -- after it is
downloaded and before it is stored. Every agent writes through the same
call, so the steps apply to fs, scm, webdav and web components alike.
config:
pre_process:
- method: soliplex.agents.manifest.pre_processors:check_pdf_password
mime_types: [application/pdf]
kwargs: { skip_invalid: true, skip_owner_restricted: false }Each document is spooled to a private temp directory first, and each step is
called as method(document, **kwargs). document carries source, uri,
key (its path in the store), mime_type, path (the spooled file -- treat
it as read-only), workdir (scratch space for output), sha256 (of path),
and read_bytes(). Steps therefore see a local file whichever store the
document is headed for, so path-based tools (qpdf, ocrmypdf, pdfium by path)
work unchanged with S3. A step that also accepts context receives the
source's resolved store.
A step answers with one of:
| Return | Effect |
|---|---|
"continue" / None |
Nothing to do. |
"modified", with new content |
Later steps see the new content, and it is what gets stored (and described by the .meta.json). Logged at INFO. |
"skip" |
The document is not stored. Any version already stored at its path is deleted (the next load drops it from the index). Later steps do not run. Logged at INFO. |
Any of these may carry a message for the log and the audit:
return PreProcessStatus.SKIP, "password protected". MODIFIED needs the new
content, so it is returned as a PreProcessResult:
from soliplex.agents.manifest.pre_process import PreProcessResult
from soliplex.agents.manifest.pre_process import PreProcessStatus
def redact(document):
text = document.read_bytes().decode("utf-8")
cleaned = text.replace("CONFIDENTIAL", "")
if cleaned == text:
return PreProcessStatus.CONTINUE
return PreProcessResult(PreProcessStatus.MODIFIED, "removed markings", data=cleaned.encode("utf-8"))Give exactly one of data= (bytes) or path= (a file the step wrote,
normally under document.workdir; relative paths resolve against it). A
MODIFIED whose content is byte-for-byte unchanged counts as CONTINUE.
metadata={...} is merged into the document's .meta.json under
metadata.pre_process.<method>.
mime_typeslimits a step to documents of those detected types -- the type the document is stored under, not its URI's extension (see File Typing and Filtering). Omit it to run on everything.- Nothing runs unless listed. There are no default steps: a manifest
without
pre_processstores every document as fetched. To check PDFs or fix AsciiDoc, list the built-in steps below. on_errordecides what a step that raises (or returns something invalid) means:continue(the default) logs it and keeps the document as it stood before that step;skipskips the document;failfails that document's write, which the agent records like any download error -- no state row, retried next run, and the stale reconcile is skipped for the run.- A skipped document still gets its state row, so it is not fetched again
until it changes upstream (the reconcile tolerates its absence). After
changing a step,
si-agent manifest reprocess <path|all> [--status skip|modified|continue|error|all] [--method <dotted>] [--dry-run]forgets the matching documents so the next run fetches and checks them again (the default is--status skip). For an incremental SCM source this also costs one full listing. - Steps run one at a time per run, even while webdav downloads concurrently -- pdfium is not thread-safe, and it bounds the spool to about one document. Sync steps run in a worker thread so the event loop stays responsive.
- The spool is
PRE_PROCESS_SPOOL_DIR(default: the system temp dir) and needs room for the largest document. In containers/tmpis often tmpfs (RAM); point it at a volume for large corpora. Agents still hold each downloaded document in memory, so spooling does not lower peak memory.
Built-in steps (soliplex.agents.manifest.pre_processors):
check_pdf_passwordskips PDFs pdfium cannot open without a password (password protected). Other open failures (truncated, not a PDF) are skipped asunreadable PDF: ...unlessskip_invalid: false. A PDF encrypted with an owner password only opens fine -- printing or copying may be restricted -- so it is kept, with a message, unlessskip_owner_restricted: true.fix_asciidocrewrites AsciiDoc that docling's parser cannot handle: block attribute lines before a table, cell-format specifiers before pipes,include::/image::directives, and blank lines inside tables.
Every pre-processed write is recorded in the source's state DB:
pre_process_documents-- the latest outcome per document: status, deciding method and message,input_sha256(the downloaded bytes),output_sha256(what was stored; empty when skipped),previous_input_sha256andhash_changed_at;pre_process-- the latest outcome per (document, step), with each step's input and output hash.
The hashes are SHA-256 of the bytes pre-processing saw and stored -- not the
upstream hash the agents use for change detection (SHA3-256 for SCM, and
sometimes absent for webdav). A re-fetch with identical content keeps
hash_changed_at; a real change records the old hash and logs
content of <uri> changed (<old> -> <new>), and a changed document that is
skipped again logs new version of <uri> still skipped. Audit rows are
removed with the document's state row.
# Everything skipped, and why
si-agent manifest pre-process-report all --status skip
# Documents whose content changed since a date
si-agent manifest pre-process-report my-manifest.yml --changed-since 2026-10-01
# Every "password protected" document
si-agent manifest pre-process-report all --message "password protected" --jsonA manifest's config.post_process is an ordered list of callbacks invoked
after the haiku-rag load for that source finishes (and after its summary
has been streamed) — whether the load succeeded, failed, or timed out. Each
entry names a method (a dotted import path) and optional kwargs; the
callback is invoked as method(source, **kwargs) — source is the manifest's
source and kwargs are the configured extra args.
config:
haiku_config: haiku.rag.custom.yaml
post_process:
- method: soliplex.agents.manifest.post_processors:vacuum
kwargs: { timeout: 1800 }
- method: your_project.postprocess:identify_and_apply
kwargs: { only_missing: true, overrides: /etc/agent/overrides.json }- Dotted path:
pkg.mod:funcorpkg.mod.func. The module must be importable in the agent's environment. - Ordering: steps run sequentially in the order listed.
- Config auto-inject: when a step omits
configand the callable accepts one (an explicitconfigparameter or**kwargs), the manifest's resolved haiku config path is passed so the callback opens the store with the same config the load used. While the callbacks run,SOURCE,DOWNLOAD_DIRandDOWNLOAD_URIare set in the environment (as they are for the load subprocess) and restored afterwards, so a config interpolating them loads in-process too. Other${VAR}references must be present in the inherited environment — see Variables the haiku-rag config needs. - Context auto-inject: likewise for a
contextparameter, which receives the run'sLoadContext— the resolved download target, document store and sidecar facade for this source. A callback that needs to read what the manifest just downloaded takescontextinstead of rediscovering the storage layout from the environment. - Load outcome (
ingester): callbacks fire regardless of the load result. The load's outcome is auto-injected as aningesterkwarg for callables that accept one: aHaikuRunwithreturncode(0on success, non-zero on failure,Noneon timeout),timed_out, and the last 1 MiB ofstdout/stderr(the full output is in the log, in parts). The exit code alone is still injected asingester_exit_code, so existing callbacks keep working. - Run outcome (
run_result): likewise, the result of the manifest run that queued the load -- itssummarycounts, per-component results,pre_runandpre_processoutcomes. Under the server, loads are queued, so a later run of the same manifest may already have started by the time a load's callbacks fire. - Terminate on error: a step that raises is logged and the exception
propagates — the remaining steps do not run. The per-step outcomes are
returned under the load result's
post_processkey only when every step succeeds. In a batch/directory run the failure is isolated to that manifest (recorded ashaiku_load_erroron its result); the other manifests still run. - Requires a load: post-process only runs when a load runs — it is skipped
with
--no-load.
Built-in callbacks (soliplex.agents.manifest.post_processors):
vacuumruns LanceDB maintenance (optimize + clean up table history) on the per-source database. It shells out tohaiku-rag vacuumas a subprocess (like the load) — keeping LanceDB's async runtime out of the agent's event loop and making the pass killable via itstimeoutkwarg (default 1800s), so a stuck compaction can't hang the run. Retention comes from the haiku config'sstorage.vacuum_retention_seconds.backfill_metadataruns the same subprocess assi-agent manifest backfill-metadatafor the manifest's source, right after its load, so nothing else is writing the database. Takesmissingandcontent_types(lists),doc_filter,full,database,batch_size,attachments(defaulttrue) andtimeout(default 1800s). It must be scoped --missingordoc_filter-- or givenfull: true, since it runs after every load. Documents it could not fill are logged and do not stop the chain; the subprocess crashing or timing out raises. See Back-filling metadata.notify_webhookPOSTs{"event": "load.finished", "source", "status", "returncode", "timed_out", "summary", "manifest_id"}, wherestatusisok,failed,timed_outorno_load, plusstderr_tail(the laststderr_lines, default 20) when the load failed or timed out. It takes the sameurl/url_secret/headers/secret_headersas the pre-run version. A failed delivery raises, which stops the chain, so list it last.
Note: All commands support WebDAV credentials via environment variables (WEBDAV_URL, WEBDAV_USERNAME, WEBDAV_PASSWORD) or command-line options (--webdav-url, --webdav-username, --webdav-password).
Git Bash on Windows: If using Git Bash on Windows, use double slashes for WebDAV paths to prevent path conversion (e.g., //documents instead of /documents).
- Discovery: Files are discovered from the source (filesystem, WebDAV, SCM, or web)
- Hashing: Each file's hash is calculated
- Filesystem/WebDAV/Web sources: SHA256 hash
- SCM sources: SHA3-256 hash for files, SHA256 for issues
- Status Check: The system checks which files are new or changed against the local sync state, so only new or changed files are processed
- Write: Each file is written to
<DOWNLOAD_DIR>/<source>/<source-relative-path>, with a<filename>.meta.jsonsidecar (see Metadata Sidecars). The stored filename is given the extension implied by its detected MIME type (added when missing, replaced when it mismatches, left alone when already correct) — see File Typing and Filtering- Pre-process (manifest runs): before the write, the manifest's
pre_processsteps check or rewrite the document; a skipped document is not written, and any earlier stored version is removed (see Pre-process steps)
- Pre-process (manifest runs): before the write, the manifest's
- State Update: Content hashes (and, for SCM, the latest commit SHA) are recorded in local state
- Stale Removal (optional): When
delete_staleis enabled, the download folder is reconciled against the source — documents no longer present (dropped from the listing, or 404 on fetch) are deleted, along with untracked orphan files (see Stale Document Removal) - haiku-rag Load (optional): When
HAIKU_LOAD_ENABLEDis set, the downloaded documents are indexed into a per-source LanceDB database viahaiku-ingester(see haiku-rag Loading)
Every downloaded document is accompanied by a <filename>.meta.json sidecar
written next to it. The sidecar records:
| Field | Description |
|---|---|
mime_type |
Detected MIME type (see File Typing and Filtering) |
source |
Source identifier (the per-source folder name) |
source_uri |
Source URI the document was discovered at |
ingestion_type |
Method used to fetch the document: fs, webdav, scm, or web |
sha256 |
SHA256 of the written bytes |
size |
Size of the written bytes |
metadata |
Any additional source-specific metadata; pre-process steps add theirs under metadata.pre_process.<method> |
source_url |
Full URL the document was fetched from (see below) |
downloaded_time |
When the document was last written, ISO 8601 with a UTC offset (see below) |
Example sidecar for a WebDAV download:
{
"mime_type": "text/markdown",
"source": "webdav:docs",
"source_uri": "handbook/readme.md",
"ingestion_type": "webdav",
"sha256": "…",
"size": 1234,
"metadata": {},
"source_url": "https://dav.example.com/docs/handbook/readme.md",
"downloaded_time": "2026-09-11T14:22:05.123456+00:00"
}Every agent records one, but what it points at differs by source:
| Agent | source_url |
|---|---|
webdav |
The server URL joined with the document's path |
web |
The requested page URL. Redirects are followed when fetching, but the requested URL is what is recorded -- it is the stable identifier, and the one sync state is keyed on |
fs |
A file:// URL for the resolved source path. Only meaningful on the host that ran the ingest, which is the only address a local document ever had |
scm |
The provider's browsable html_url for the file or issue, falling back to the contents API url when the provider did not return one |
The field is omitted entirely when an agent has no URL to record -- for example an SCM provider whose response carried neither key. Consumers should treat it as optional.
The soliplex-sidecar-metadata haiku-rag metadata provider copies the
sidecar into each document's haiku-rag metadata at load time; see Document
Metadata in haiku-rag.
Records when the document's bytes were last written, not when they were last checked for changes. A document that passes its source's freshness check (an unchanged content hash, a matching WebDAV ETag) is never rewritten, so its sidecar keeps the timestamp of the fetch that did produce it.
This also means sidecars written before the field existed will never gain one, since nothing rewrites an unchanged document. Treat it as optional too.
An scm manifest component with incremental: true uses commit-based tracking for efficient synchronization:
- Sync State Check: Retrieves last processed commit SHA from local state
- Commit Enumeration: Fetches only commits since the last sync
- Change Detection: Extracts changed and removed file paths from commits
- Selective Fetch: Downloads only files that were modified
- Write: Writes changed files to
DOWNLOAD_DIRand deletes removed ones - State Update: Stores the latest commit SHA locally for subsequent syncs
This approach reduces API calls and bandwidth by 80-95% compared to full repository scans. On first run (or after si-agent scm reset-sync), a full sync is performed to establish the baseline.
MIME types are determined from file content, not the filename. Detection resolves in this order:
- Explicit
Content-Typeheader (WebDAV only — the GET response header, or the PROPFINDgetcontenttypeproperty), unless it is generic (application/octet-stream). - Content sniffing via puremagic, which recognises binary formats (PDF, PNG, Office documents, …) by their magic bytes.
- Filename extension via the standard library, plus overrides for Office and text formats.
- Plain-text default (filesystem and git only): an extension-less file
whose bytes look like UTF-8 text is treated as
text/plain. WebDAV does not apply this default — it relies on the server-provided type. - Otherwise
application/octet-stream.
Once typed, the document is written with the extension implied by its MIME
type (e.g. an extension-less PDF is stored as <name>.pdf; an extension-less
text file on the fs/git agents is stored as <name>.txt).
Detect-then-filter. Files are filtered by the EXTENSIONS configuration
against their detected type, not their original filename. The default
extensions are md, pdf, doc, docx. Extension-less files are no longer
skipped up front — they are read/downloaded, classified by content, and only
then filtered. A file survives when the extension implied by its detected MIME
type is in EXTENSIONS.
To add more types (for example, to keep extension-less text files, whose
detected type is text/plain → txt):
export EXTENSIONS=md,pdf,doc,docx,txt,rstNote: puremagic identifies binary formats by signature but cannot recognise plain text or Markdown (which have no magic bytes); those still resolve via their extension or, on the fs/git agents, the
text/plaindefault above.
The validate-config / check-status commands additionally reject files
whose recorded content type is an archive or opaque binary:
- ZIP archives
- RAR archives
- 7z archives
- Generic binary files without proper MIME types
For SCM sources, issues (including their comments) are rendered as Markdown documents and ingested alongside repository files. This enables full-text search and analysis of issue discussions.
As an example, the soliplex documentation) can be loaded using both the filesystem and via git.
Ingest a checkout's docs directory:
git clone https://github.com/soliplex/soliplex.git# soliplex-docs.yml
id: soliplex-docs
name: soliplex docs
source: soliplex-docs
components:
- name: docs
type: fs
path: <path-to-checkout>/soliplex/docs# Set up environment
export DOWNLOAD_DIR=./downloads
uv run si-agent manifest run soliplex-docs.yml --no-load
# Files land under ./downloads/soliplex-docs/, each with a .meta.json sidecar
ls ./downloads/soliplex-docsReviewing first:
# Preview the inventory as JSON (writes nothing)
uv run si-agent fs build-config <path-to-checkout>/soliplex/docs
# Check which files are supported
uv run si-agent fs validate-config <path-to-checkout>/soliplex/docs
# If there are errors, fix them now
uv run si-agent manifest run soliplex-docs.yml --no-load# soliplex-repo.yml
id: soliplex-repo
name: soliplex repo
source: soliplex-repo
components:
- name: soliplex
type: scm
platform: github
owner: mycompany
repo: soliplex# Set up environment
export DOWNLOAD_DIR=./downloads
export scm_auth_token=ghp_your_token_here
# Write repository contents
si-agent manifest run soliplex-repo.yml --no-load
# Files land under ./downloads/soliplex-repo/
ls ./downloads/soliplex-repo# webdav-docs.yml
id: webdav-docs
name: webdav docs
source: webdav-docs
components:
- name: project-docs
type: webdav
url: https://nextcloud.example.com/remote.php/dav/files/username
path: /Documents/project-docs# Set up environment
export DOWNLOAD_DIR=./downloads
export WEBDAV_USERNAME=your-username
export WEBDAV_PASSWORD=your-password
si-agent manifest run webdav-docs.yml --no-load
# Files land under ./downloads/webdav-docs/
ls ./downloads/webdav-docs# The haiku-rag config places the database; copy the example and keep its
# `lancedb.databases` entry interpolating ${SOURCE}.
mkdir -p ./haiku-config
cp example-haiku-configs/haiku.rag.default.yaml ./haiku-config/
# Set up environment. LANCEDB_DIR is what that config interpolates.
export DOWNLOAD_DIR=./downloads
export LANCEDB_DIR=./lancedb
export HAIKU_PATH=./haiku-config
export HAIKU_LOAD_ENABLED=true
# Run a manifest and load the result into ./lancedb/<source>.lancedb
si-agent manifest run /path/to/manifest.yml --loadThe example config also interpolates ${STATE_DIR} and the model/service URLs
(QA_MODEL, QA_BASE_URL, OLLAMA_BASE_URL, DOCLING1_BASE_URL,
DOCLING2_BASE_URL, EMBEDDINGS_BASE_URL); every one must be exported too, or
the load fails on the missing variable.
haiku-ingester reads only the document bytes. This package registers haiku-rag
metadata providers
that add more to each document's haiku-rag metadata. A haiku source names one
metadata_provider, so pick the one covering what you want:
| Provider | Adds |
|---|---|
soliplex-sidecar-metadata |
The document's .meta.json sidecar, flattened |
soliplex-pdf-metadata |
A PDF's page count and document information |
soliplex-metadata |
Both; the sidecar is applied last, so manifest metadata wins a clash |
haiku-rag does not call providers for the PDF attachments it extracts into documents of their own; back-fill them.
Both example configs in example-haiku-configs/ name soliplex-metadata:
sources:
- type: fs
id: ${SOURCE}
root: ${DOWNLOAD_DIR}/${SOURCE}
metadata_provider: soliplex-metadatasoliplex-sidecar-metadata reads the sidecar the manifest run wrote beside
the document and returns it flattened, as
Metadata Sidecars describes: mime_type, source,
source_uri, ingestion_type, sha256, size, source_url and
downloaded_time when set, and the manifest's metadata entries at the top
level (nested values JSON-encoded). It finds the download store the same way
the agent does: the haiku source's id is the sanitized manifest source
(${SOURCE}), and DOWNLOAD_DIR / DOWNLOAD_S3_* come from the load's
environment, so it works on the local and S3 stores alike.
A missing, unreadable or malformed sidecar gives the document no sidecar keys
and a log line; it never fails the document. A sidecar that changes while its
document does not -- new manifest metadata -- is not picked up by a load,
because haiku-rag skips an unchanged document before calling any provider
(and does not ingest *.meta.json itself). Back-fill those with a
--filter, or a pass without --missing.
| Key | Value |
|---|---|
page_count |
Number of pages (an integer) |
pdf_version |
PDF version from the header, e.g. "1.7" |
pdf_title, pdf_author, pdf_subject, pdf_keywords, pdf_creator, pdf_producer |
The PDF's document information entries |
pdf_creation_date, pdf_mod_date |
The PDF's dates, as ISO 8601 (kept as written if they do not parse) |
Only page_count is always present; an information entry the PDF leaves
empty is left out. A document is a PDF when its content type is
application/pdf or its first 1024 bytes hold a %PDF- header; anything else
gets no keys. A PDF pdfium cannot open (password protected, truncated) gets no
keys either, and a warning is logged -- the provider never fails the document.
The provider runs inside haiku-ingester, so this package must be installed
in the environment that runs it (the haiku_load_command default runs it from
the agent's own). haiku-rag calls a provider only when it fetches a new or
changed document: after enabling it, existing documents gain the keys the next
time they change, unless you back-fill them -- --missing source_uri for the
sidecar keys (every sidecar has one), --missing page_count for the PDF keys.
backfill-metadata adds a provider's keys to documents indexed before the
provider was configured, without re-ingesting them -- nothing is converted,
chunked or embedded:
# What would change, without writing anything
si-agent manifest backfill-metadata --missing page_count --check
# Every manifest in $MANIFEST_DIR, PDFs still lacking a page count
si-agent manifest backfill-metadata --missing page_count --content-type application/pdf
# One manifest; print the command instead of running it
si-agent manifest backfill-metadata /manifests/handbook.yml --missing page_count --dry-runFor each source in the haiku config that names a metadata_provider, it
lists the documents that source ingested, fetches each selected one again
through that source (so the provider sees what the ingester would hand it),
calls the provider, and merges the keys it returns into the document's
metadata. The provider is the one code path: the ingester calls it for new
documents, the back-fill for old ones.
Which documents run:
-
Scoped in LanceDB.
--missing KEYbecomes aWHEREclause on the stored metadata,(metadata IS NULL OR metadata NOT LIKE '%"KEY"%' ...), so a document that already has every key is never listed, let alone fetched.--filteradds a clause of your own (AND-ed with it), e.g.--filter "uri LIKE '%/reports/%'". A key may only contain letters, digits,_,.and-. -
Then by type. Of those, a document runs when its stored
content_typeis one of the--content-typevalues, if any are given. -
Unscoped, everything. With neither
--missingnor--filterevery document of the source is fetched again. Documents whose metadata the provider would not change are never written either way. -
Database. The haiku config must place the database (
lancedb.databases), as the load's does;--db-namepicks one when it places several.--batch-size(default 500) sizes the listing's pages, all of which are read before the first write. -
Stale documents are left alone. A document whose bytes changed since it was indexed (its stored
md5differs) is counted asstale: the next load re-ingests it, which runs the provider anyway. -
Providers can opt out. Its documents are counted as skipped and it is never called: run over stored documents it would record the time of the back-fill. Third-party providers of that kind should do the same.
-
Never alongside a load. It writes the same database the load does, so do not run the CLI verb while a load for that source is in progress. As a post-process step it runs after the load by construction, and must be scoped (
missingordoc_filter) or givenfull: true, so that a scheduled load with nothing new does not fetch every document again:config: post_process: - method: soliplex.agents.manifest.post_processors:backfill_metadata kwargs: missing: [page_count] content_types: [application/pdf]
haiku-rag stores each file embedded in a PDF as a document of its own:
<parent uri>#attachment=<percent-encoded name> (nested ones chaining
fragments), linked by parent_uri and owned by no source. It does not
call a metadata provider for them, and no source can fetch that URI, so the
back-fill derives each one from its parent instead:
- walk
parent_uriup to the top-level document, which a source owns; - fetch it once through that source, however many attachments it has;
- extract the attachment from it, following nested ones down, exactly as haiku-rag does (same URI, content type and MD5);
- call that source's provider with the attachment's bytes and
extra_metadata["parent_uri"]set, and merge as for any document.
So soliplex-pdf-metadata gives an attached PDF its own page count, and
soliplex-sidecar-metadata gives every attachment its parent's sidecar:
an attachment has no download of its own. An attachment never gains a
source_id, and no provider can change its parent_uri. An attachment whose
top-level document changed since it was indexed, or whose stored md5 no
longer matches what the parent embeds, is stale. One whose parent is no
longer indexed, or no longer embeds it, is orphaned: haiku-rag removes such
a child only when it re-ingests the parent, and not at all when the new
parent cannot be opened, so these are worth a look. --no-attachments
(attachments: false) leaves attachments alone.
The outcome line reports what it did, e.g.
(scanned 6: 3 updated (2 attachments), 1 unchanged, 0 stale, 1 orphaned, 1 skipped, 0 errors) -- under --check, would update; --json has the full
counts. skipped covers documents no source with a provider owns
(haiku-rag add-src documents and their attachments), providers that opt out,
and ones --content-type or the exact --missing check left out. A document that fails (gone from the store, provider raised) is listed
and the run continues; the subprocess then exits 3 and the verb exits 1. A PDF
the provider cannot read gets no keys and counts as unchanged, so
--missing page_count selects it again on every run.
The agents can be run as a REST API server using FastAPI. The server runs manifests (on a cron schedule, and on demand via POST /api/v1/manifest/run) and exposes read-only inspection routes for each source type, with support for authentication and interactive documentation. Ingestion over HTTP goes only through manifests.
# Basic
si-agent serve
# Custom host and port
si-agent serve --host 0.0.0.0 --port 8080
# Development mode with auto-reload
si-agent serve --reloadThe server always runs as a single worker process. The manifest
scheduler keeps its cron state and run queue in memory, so running
multiple workers would make each worker register every cron and run every
manifest independently, with no cross-process coordination. Multi-worker
mode is therefore intentionally not exposed, and any WEB_CONCURRENCY
environment variable is ignored. Scale out with multiple single-worker
instances behind a load balancer instead (note that scheduling should only
be enabled on one instance — see Scheduling).
The server supports multiple authentication methods:
si-agent serve
# All requests allowedexport API_KEY=your-api-key
export API_KEY_ENABLED=true
si-agent serveClients must include the API key in the Authorization header:
curl -H "Authorization: Bearer your-api-key" http://localhost:8001/api/v1/manifest/queueexport AUTH_TRUST_PROXY_HEADERS=true
si-agent serveThe server will trust authentication headers from a reverse proxy (e.g., OAuth2 Proxy):
X-Auth-Request-UserX-Forwarded-UserX-Forwarded-Email
The only way to ingest over HTTP. See On-demand Runs.
| Method | Endpoint | Description |
|---|---|---|
POST |
/api/v1/manifest/run |
Queue a manifest from MANIFEST_DIR by id; returns 202 |
GET |
/api/v1/manifest/queue |
Ids of manifests queued or running |
POST |
/api/v1/manifest/validate |
Validate manifests without executing |
POST /run responds 202 with "status": "queued" or "already_queued";
404 if no valid manifest in MANIFEST_DIR has that id; 409 if more than
one file declares it; 503 if MANIFEST_DIR is unset or not a directory.
Examples:
# Queue a manifest run
curl -X POST http://localhost:8001/api/v1/manifest/run \
-F "manifest_id=test-scm"
# What is still queued or running
curl http://localhost:8001/api/v1/manifest/queue
# Validate manifest files
curl -X POST http://localhost:8001/api/v1/manifest/validate \
-F "path=/path/to/manifests"Read-only inspection of a server-side directory.
| Method | Endpoint | Description |
|---|---|---|
POST |
/api/v1/fs/build-config |
Build inventory from directory |
POST |
/api/v1/fs/validate-config |
Validate the inventory built from a directory |
POST |
/api/v1/fs/check-status |
Check which files need ingestion |
Examples:
# Build configuration from directory
curl -X POST http://localhost:8001/api/v1/fs/build-config \
-F "path=/path/to/docs"
# Validate using a directory
curl -X POST http://localhost:8001/api/v1/fs/validate-config \
-F "config_file=/path/to/docs"| Method | Endpoint | Description |
|---|---|---|
GET |
/api/v1/scm/{scm}/issues |
List repository issues |
GET |
/api/v1/scm/{scm}/repo |
List repository files |
{scm} is github or gitea; both take repo_name and owner query
parameters.
Examples:
# List GitHub issues
curl "http://localhost:8001/api/v1/scm/github/issues?repo_name=my-repo&owner=myuser"
# List repository files
curl "http://localhost:8001/api/v1/scm/github/repo?repo_name=my-repo&owner=myuser"| Method | Endpoint | Description |
|---|---|---|
POST |
/api/v1/webdav/validate-config |
Validate inventory from WebDAV path |
POST |
/api/v1/webdav/check-status |
Check which files need ingestion |
Example:
# Validate using WebDAV path
curl -X POST http://localhost:8001/api/v1/webdav/validate-config \
-F "config_path=/documents" \
-F "webdav_url=https://webdav.example.com"| Method | Endpoint | Description |
|---|---|---|
GET |
/health |
Server health check |
Example:
curl http://localhost:8001/health
# Returns: {"status": "healthy"}Interactive API documentation is available at:
- Swagger UI:
http://localhost:8001/docs - ReDoc:
http://localhost:8001/redoc - OpenAPI JSON:
http://localhost:8001/openapi.json
The server is designed to run in containers. The Dockerfile is a
multi-stage build exposing two selectable targets:
| Target | Purpose | Dependencies | Default command |
|---|---|---|---|
production |
Minimal runtime image (default target) | Runtime only (uv sync --no-dev) |
si-agent serve --host=0.0.0.0 |
development |
Local dev with live reload | Runtime and dev deps (uv sync) |
si-agent serve --host=0.0.0.0 --reload |
Both stages run as a non-root appuser (uid/gid 1000 by default,
overridable via the APP_UID/APP_GID build args), include git for SCM
CLI mode, expose port 8001, and define a /health healthcheck.
production is the last stage, so it is built when no --target is given:
# Build the production image (default target)
docker build -t ingester-agents:latest .
# Run with environment variables
docker run -d \
-p 8001:8001 \
-e DOWNLOAD_DIR=/data/downloads \
-e API_KEY_ENABLED=true \
-e API_KEY=your-secret-key \
-v "$(pwd)/downloads:/data/downloads" \
ingester-agents:latest
# Check health
curl http://localhost:8001/healthThe development target includes the full toolchain and starts uvicorn
with --reload. Bind-mount the source so code changes reload live:
# Build the development image
docker build --target development -t ingester-agents:dev .
# Run with the source bind-mounted for live reload
docker run --rm -it \
-p 8001:8001 \
-v "$(pwd):/app" \
ingester-agents:devTo match file ownership on bind mounts to your host user, pass build args:
docker build --target development \
--build-arg APP_UID="$(id -u)" \
--build-arg APP_GID="$(id -g)" \
-t ingester-agents:dev .The Docker image includes:
- Non-root user for security
- Health checks for orchestration
- Proper signal handling
- Production-ready uvicorn configuration
Ensure your tokens have the required permissions:
- GitHub:
reposcope for private repositories, public access for public repos - Gitea: Access token with read permissions
For SCM and WebDAV sources, verify the source server is reachable and any
required credentials (scm_auth_token, WEBDAV_URL/WEBDAV_USERNAME/
WEBDAV_PASSWORD) are set. Downloaded files are written under DOWNLOAD_DIR.
For SCM agents, ensure the repository name and owner are correct. Use the exact repository name, not the URL.
# Clone repository
git clone <repository-url>
cd ingester-agents
# Install dependencies with dev tools
uv sync
# Run tests
uv run pytest
# Run linter
uv run ruff checkThe project uses pytest with 100% code coverage requirements:
# Run unit tests with coverage
uv run pytest
# Run specific tests
uv run pytest tests/unit/test_client.py
# Generate coverage report
uv run pytest --cov-report=htmlThe project uses Ruff for linting and code formatting:
# Check code
uv run ruff check
# Auto-fix issues
uv run ruff check --fix
# Format code
uv run ruff formatsoliplex.agents/
├── src/soliplex/agents/
│ ├── cli.py # Main CLI entry point (includes 'serve' command)
│ ├── local_store.py # Writes downloaded documents + .meta.json sidecars
│ ├── local_state.py # Per-source SQLite sync state (hashes + commit SHA)
│ ├── config.py # Configuration, settings, and manifest models
│ ├── haiku_metadata.py # haiku-rag metadata providers (sidecar, PDF, combined)
│ ├── haiku_backfill.py # Re-runs metadata providers over indexed documents
│ ├── server/ # FastAPI server
│ │ ├── __init__.py # FastAPI app initialization, scheduler
│ │ ├── auth.py # Authentication (API key & OAuth2 proxy)
│ │ └── routes/
│ │ ├── __init__.py
│ │ ├── fs.py # Filesystem API endpoints
│ │ ├── scm.py # SCM API endpoints
│ │ ├── webdav.py # WebDAV API endpoints
│ │ ├── web.py # Web API endpoints
│ │ └── manifest.py # Manifest API endpoints
│ ├── common/ # Shared utilities
│ │ ├── urls_file.py # URL list reader (local, S3, WebDAV)
│ │ ├── s3.py # S3 object reader
│ │ ├── mime.py # Content-based MIME detection + extension logic
│ │ └── config.py # Inventory read/validate helpers
│ ├── fs/ # Filesystem agent
│ │ ├── app.py # Core filesystem logic
│ │ └── cli.py # Filesystem CLI commands
│ ├── web/ # Web agent
│ │ └── app.py # Core web fetching logic
│ ├── webdav/ # WebDAV agent
│ │ ├── app.py # Core WebDAV logic
│ │ └── cli.py # WebDAV CLI commands
│ ├── manifest/ # Manifest runner
│ │ ├── runner.py # YAML loading, validation, dispatch
│ │ ├── haiku_loader.py # haiku-ingester batch load subprocess
│ │ ├── haiku_maint.py # haiku-rag migrate/vacuum/backfill-metadata subprocesses
│ │ ├── post_processors.py # Built-in post-process callbacks
│ │ └── cli.py # Manifest CLI commands
│ └── scm/ # SCM agent
│ ├── app.py # Core SCM logic
│ ├── cli.py # SCM CLI commands
│ ├── base.py # Base SCM provider interface
│ ├── github/ # GitHub implementation
│ ├── gitea/ # Gitea implementation
│ └── lib/
│ ├── templates/ # Issue rendering templates
│ └── utils.py # Utility functions
├── example-manifests/ # Example manifests (fs, scm, web, webdav, composite, delete-stale)
├── example-haiku-configs/ # Example haiku-rag configs for the load step (local, S3)
├── tests/ # Test suite
│ └── unit/
│ ├── test_server_*.py # Server API tests
│ └── ...
├── Dockerfile # Production container
├── .dockerignore # Build context exclusions
└── DOCKERFILE_CHANGES.md # Docker implementation documentation
CLI Layer:
cli.py- Main entry point withfs,web,scm,webdav,manifest, andservecommands- Agent-specific CLI commands in
fs/cli.py,webdav/cli.py,scm/cli.py, andmanifest/cli.py
Server Layer:
server/- FastAPI applicationserver/auth.py- Flexible authentication (none, API key, OAuth2 proxy)server/routes/- REST API endpoints mirroring CLI functionality
Agent Layer:
fs/app.py- Filesystem operations (shared by CLI and API)web/app.py- Web page fetching and ingestion (shared by CLI and API)webdav/app.py- WebDAV operations (shared by CLI and API)scm/app.py- SCM operations (shared by CLI and API)manifest/runner.py- Manifest loading, validation, and dispatch to agentslocal_store.py- Writes fetched documents and metadata sidecars toDOWNLOAD_DIRlocal_state.py- Local synchronization state (content hashes + SCM commit markers)
haiku-rag Layer:
manifest/haiku_loader.py- Thehaiku-ingester run-batchload after each manifest runmanifest/haiku_maint.py-migrate/vacuum/backfill-metadatasubprocesseshaiku_metadata.py- Metadata providers run insidehaiku-ingester(registered entry points)haiku_backfill.py- Back-fill subprocess re-running those providers over indexed documents
Configuration:
config.py- Pydantic settings and manifest component models- Environment variables or
.envfile for configuration - YAML manifest files for declarative multi-source ingestion
See LICENSE file for details.
For issues and questions, please open an issue on the repository.