Conversation
alxmrs
force-pushed
the
adbc-geo-benchmarks
branch
3 times, most recently
from
September 28, 2026 15:42
71141d4 to
8316392
Compare
GEOBENCH_ENGINE=adbc-<backend> runs a case on any database in the ADBC tests' backend table (tests/_adbc.py). Registration copies data, and the cases open the whole ARCO-ERA5 archive, so each query ingests only the variables it names over the window its parameters bound. The case SQL is DataFusion's; the few differences from other databases' SQL (date_part, timestamp literals, DOUBLE, MySQL's backticks) are rewritten explicitly in the harness, since xarray-sql translates data, not queries. Interval and text durations and zone-aware times are normalized before comparing with the reference. A new workflow runs cases 02-06 on all eleven databases, one parallel job each, daily, by hand, and on PRs that touch the engine layer. The suite's summary table now lists whichever engines ran. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
From the first run across eleven databases: - Case 05 joins on time + prediction_timedelta, which only works where durations are native intervals. The harness now rewrites time + duration per database, matching how the adapter stores durations there: integer nanoseconds (SQLite: datetime(); ClickHouse: addNanoseconds; SQL Server: DATEADD) or text (MySQL, MariaDB: DATE_ADD ... MICROSECOND). Duration columns are known from the registered data, and an integer duration read back under an alias is converted to a timedelta. - The case aliased a column as lead, a reserved word in MySQL; it is quoted now. - Trino's memory connector defaults to 128 MB, too little for a global ERA5 day. Trino now starts from a step that mounts a catalog with a 4 GB limit. - Experiment: ANALYZE each table after ingest on PostgreSQL, to test whether its timeouts come from planning without statistics. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
- MariaDB joins by nested loop unless hash joins are enabled, so the forecast-skill join timed out where MySQL finishes in 37 s. The harness enables them (join_cache_level = 8) for MariaDB. - Trino's image sizes its heap to 80% of the runner's memory, which with the Python process holding a day of global ERA5 exhausted it and the server died. It now gets -Xmx6G and a 3 GB memory-connector limit. Its driver ingests about 10k rows/s, so it runs the regional cases (02, 04, 05); the global-day ones cannot finish in a job. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
Trino exited with java.lang.OutOfMemoryError (Java heap space) at 6 GB while ingesting the regional ERA5 window; the runner itself still had 13 GB free. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
The suite's cells called run_case_cell without rep_timeout, so every case ran under its 600 s default whatever --cell-timeout said. Cells now get the requested limit, locally and on Coiled. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
Trino's driver stores durations as text such as '43200000000000ns', so case 05 failed on timestamp(9) + varchar. parse_duration reads exactly that text. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
Matches the adbc databases workflow: floating tags pinned to the last green run's digests, drivers to its versions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t
alxmrs
force-pushed
the
adbc-geo-benchmarks
branch
from
September 28, 2026 17:16
8316392 to
6308db4
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Runs the geospatial benchmark cases (
benchmarks/geospatial) on every ADBC database against real data (ARCO-ERA5 and WeatherBench 2): daily, by hand, and on PRs that touch this layer. Each case checks its SQL answer against an xarray reference, so a green cell means that database computed the right numbers through xarray-sql's register → SQL →to_datasetpath.Results
Latest run (36318725813): 11 databases, 53 cells, all matching the reference. Times are the SQL step, including ingest.
Trino runs the regional cases only. Its driver ingests about 10k rows/s, so the two global-day cases (25M rows each) can't finish in a job. Case 01 (a 100M+ pixel Sentinel-2 scene) is too large to copy into a database per run, and 07–09 need DataFusion UDFs or Earth Engine; all of those stay n/a, as for the other non-DataFusion engines.
How
GEOBENCH_ENGINE=adbc-<backend>runs a case on any backend in the ADBC tests' table (tests/_adbc.py), connected the same way._engines._ADBC), since xarray-sql translates data, not queries:date_part('hour', …), timestamp literals,DOUBLE, MySQL backticks;time + <duration>, keyed to how the adapter stores durations: integer nanoseconds (SQLitedatetime(), ClickHouseaddNanoseconds, SQL ServerDATEADD) or text (MySQL/MariaDBDATE_ADD, Trinoparse_duration);join_cache_level = 8).geospatial-adbc.yml: one parallel job per database, each starting only its own server. Trino runs from a step with a 10 GB heap and a 3 GB memory catalog. Results go in the job summary and an artifact.Found along the way
Fixed in #255:
ANALYZEs (5 s and 14 s).Fixed in #261 (against
main): the chunked round-trip lost windows over nanosecond times.Documented in #255:
Tried and dropped: raising Trino's statement size (
adbc.statement.ingest.max_query_size_bytesandquery.max-lengthat 16 MB) made its ingest slower (277 s vs 203 s on case 02) and ran it out of heap. Turning rows into SQL text and parsing it is the bottleneck, not per-query overhead.Fixed here:
engine_suite.py --localignored--cell-timeout(every case ran under a 600 s default), and the summary table now lists whichever engines ran.🤖 Generated with Claude Code
https://claude.ai/code/session_01AMTHTEAoyKzUJLvFmg5G6t