Skip to content

Run the geospatial benchmarks on every ADBC database - #260

Draft
alxmrs wants to merge 10 commits into
mainfrom
adbc-geo-benchmarks
Draft

alxmrs wants to merge 10 commits into
mainfrom
adbc-geo-benchmarks

Conversation

@alxmrs

@alxmrs alxmrs commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Summary

Runs the geospatial benchmark cases (benchmarks/geospatial) on every ADBC database against real data (ARCO-ERA5 and WeatherBench 2): daily, by hand, and on PRs that touch this layer. Each case checks its SQL answer against an xarray reference, so a green cell means that database computed the right numbers through xarray-sql's register → SQL → to_dataset path.

Results

Latest run (36318725813): 11 databases, 53 cells, all matching the reference. Times are the SQL step, including ingest.

Database climatology zonal mean anomaly forecast skill zonal vector
DuckDB ✅ 2s ✅ 3s ✅ 2s ✅ 3s ✅ 4s
DataFusion ✅ 2s ✅ 1s ✅ 2s ✅ 2s ✅ 1s
GizmoSQL (Flight SQL) ✅ 2s ✅ 3s ✅ 2s ✅ 3s ✅ 4s
chDB ✅ 2s ✅ 3s ✅ 2s ✅ 3s ✅ 5s
ClickHouse 26.8 ✅ 1s ✅ 2s ✅ 1s ✅ 2s ✅ 80s
PostgreSQL 18 ✅ 3s ✅ 16s ✅ 4s ✅ 7s ✅ 16s
SQLite ✅ 7s ✅ 83s ✅ 23s ✅ 16s ✅ 80s
SQL Server 2022 ✅ 7s ✅ 80s ✅ 34s ✅ 13s ✅ 40s
MySQL 8.4 ✅ 22s ✅ 188s ✅ 24s ✅ 41s ✅ 188s
MariaDB 11.4 ✅ 12s ✅ 66s ✅ 32s ✅ 89s ✅ 64s
Trino 483 ✅ 203s not run ✅ 206s ✅ 1578s not run

Trino runs the regional cases only. Its driver ingests about 10k rows/s, so the two global-day cases (25M rows each) can't finish in a job. Case 01 (a 100M+ pixel Sentinel-2 scene) is too large to copy into a database per run, and 07–09 need DataFusion UDFs or Earth Engine; all of those stay n/a, as for the other non-DataFusion engines.

How

  • GEOBENCH_ENGINE=adbc-<backend> runs a case on any backend in the ADBC tests' table (tests/_adbc.py), connected the same way.
  • Windowed ingest. Registration copies data, and the cases open the whole archive, so each query ingests only the variables it names over the window its parameters bound.
  • Dialect rewrites live only in the harness (_engines._ADBC), since xarray-sql translates data, not queries:
    • date_part('hour', …), timestamp literals, DOUBLE, MySQL backticks;
    • time + <duration>, keyed to how the adapter stores durations: integer nanoseconds (SQLite datetime(), ClickHouse addNanoseconds, SQL Server DATEADD) or text (MySQL/MariaDB DATE_ADD, Trino parse_duration);
    • MariaDB gets hash joins (join_cache_level = 8).
  • Result normalization: interval, text and integer durations, and zone-aware times, are converted before comparing with the reference.
  • Workflow geospatial-adbc.yml: one parallel job per database, each starting only its own server. Trino runs from a step with a 10 GB heap and a 3 GB memory catalog. Results go in the job summary and an artifact.

Found along the way

Fixed in #255:

  • PostgreSQL had no statistics on freshly ingested tables, so the planner guessed and joins ran for over 10 minutes; the adapter now ANALYZEs (5 s and 14 s).
  • SQLite missed rows at inclusive time bounds.
  • The spilled read failed on mixed whole and fractional-second SQLite times.

Fixed in #261 (against main): the chunked round-trip lost windows over nanosecond times.

Documented in #255:

  • MariaDB joins by nested loop unless hash joins are enabled.
  • Trino's ingest rate and memory-catalog cap.

Tried and dropped: raising Trino's statement size (adbc.statement.ingest.max_query_size_bytes and query.max-length at 16 MB) made its ingest slower (277 s vs 203 s on case 02) and ran it out of heap. Turning rows into SQL text and parsing it is the bottleneck, not per-query overhead.

Fixed here: engine_suite.py --local ignored --cell-timeout (every case ran under a 600 s default), and the summary table now lists whichever engines ran.

🤖 Generated with Claude Code

https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t

@alxmrs
alxmrs force-pushed the adbc-geo-benchmarks branch 3 times, most recently from 71141d4 to 8316392 Compare September 28, 2026 15:42
Base automatically changed from adbc-adapter to main September 28, 2026 17:14
alxmrs and others added 10 commits September 28, 2026 10:15
GEOBENCH_ENGINE=adbc-<backend> runs a case on any database in the ADBC
tests' backend table (tests/_adbc.py). Registration copies data, and
the cases open the whole ARCO-ERA5 archive, so each query ingests only
the variables it names over the window its parameters bound. The case
SQL is DataFusion's; the few differences from other databases' SQL
(date_part, timestamp literals, DOUBLE, MySQL's backticks) are
rewritten explicitly in the harness, since xarray-sql translates data,
not queries. Interval and text durations and zone-aware times are
normalized before comparing with the reference.

A new workflow runs cases 02-06 on all eleven databases, one parallel
job each, daily, by hand, and on PRs that touch the engine layer. The
suite's summary table now lists whichever engines ran.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
From the first run across eleven databases:

- Case 05 joins on time + prediction_timedelta, which only works where
  durations are native intervals. The harness now rewrites time +
  duration per database, matching how the adapter stores durations
  there: integer nanoseconds (SQLite: datetime(); ClickHouse:
  addNanoseconds; SQL Server: DATEADD) or text (MySQL, MariaDB:
  DATE_ADD ... MICROSECOND). Duration columns are known from the
  registered data, and an integer duration read back under an alias is
  converted to a timedelta.
- The case aliased a column as lead, a reserved word in MySQL; it is
  quoted now.
- Trino's memory connector defaults to 128 MB, too little for a global
  ERA5 day. Trino now starts from a step that mounts a catalog with a
  4 GB limit.
- Experiment: ANALYZE each table after ingest on PostgreSQL, to test
  whether its timeouts come from planning without statistics.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
- MariaDB joins by nested loop unless hash joins are enabled, so the
  forecast-skill join timed out where MySQL finishes in 37 s. The
  harness enables them (join_cache_level = 8) for MariaDB.
- Trino's image sizes its heap to 80% of the runner's memory, which
  with the Python process holding a day of global ERA5 exhausted it
  and the server died. It now gets -Xmx6G and a 3 GB memory-connector
  limit. Its driver ingests about 10k rows/s, so it runs the regional
  cases (02, 04, 05); the global-day ones cannot finish in a job.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
Trino exited with java.lang.OutOfMemoryError (Java heap space) at 6 GB
while ingesting the regional ERA5 window; the runner itself still had
13 GB free.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
The suite's cells called run_case_cell without rep_timeout, so every
case ran under its 600 s default whatever --cell-timeout said. Cells
now get the requested limit, locally and on Coiled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
Trino's driver stores durations as text such as '43200000000000ns', so
case 05 failed on timestamp(9) + varchar. parse_duration reads exactly
that text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
Matches the adbc databases workflow: floating tags pinned to the last
green run's digests, drivers to its versions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
@alxmrs
alxmrs force-pushed the adbc-geo-benchmarks branch from 8316392 to 6308db4 Compare September 28, 2026 17:16

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant