Skip to content

Save prediction artifacts in TabArena's result format - #41

Closed
adrian-prior wants to merge 2 commits into
mainfrom
adrian/prediction-artifacts
Closed

adrian-prior wants to merge 2 commits into
mainfrom
adrian/prediction-artifacts

Conversation

@adrian-prior

@adrian-prior adrian-prior commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator

Generated by Codex

Save per-config prediction artifacts using TabArena's raw result schema so they can support offline analysis and later storage integration. Each completed results.pkl contains validation/test predictions, labels, config metadata, metric errors and timings; temporal row identities and fit provenance live under relarena. Public loaders validate row context and array integrity, and the writer rejects unsupported task types.

Format and scope
  • Uses data/<model_config>/<dataset__task>/<repeat>_<fold>/results.pkl, with the seed identifying a repeat of temporal fold 0.
  • Labels are included in each config artifact. Loading requires trusted pickle files.
  • Validation-only outputs use validation.partial until test predictions are available, so pickle discovery does not treat them as completed results.
  • Preserves temporal holdout semantics and existing run fingerprints without hashing input tables again.
  • Checks expected row order and array shape; models must return predictions in input-table order. Plain prediction arrays cannot expose a model-side permutation.
  • Local artifact I/O only. Upload/download integration and resumable prediction caching are outside this change.
  • Runner/CLI integration and end-to-end tests are in Export prediction artifacts in TabArena's result format #42.
Testing
  • Latest lower-layer check: 38 runner, CLI, tuner and dataset tests passed. All configured pre-commit hooks passed, including pinned Ruff 0.15.13.
  • Latest complete-stack suite: 443 passed, 11 skipped. The stack covers binary/regression round trips, alignment and corruption checks, fit counts, separate extra-refit timings, and rejection of multiclass, multilabel and link-prediction task types.
  • Ran three cheap LightGBM configs (1, 2 and 3 estimators) on rel-f1/driver-dnf, seed 7, with all-config refits. Verified 566 validation rows, 702 test rows and one additional nonwinning refit.
  • Verified those three artifacts and 20 synthetic binary/regression artifacts with TabArena's actual BaselineResult.from_pickle / ConfigResult reader at revision de1be372095e23e72e2f42992eceea036b97dea6, without conversion. Recomputed metric errors and checked config metadata and simulation inputs.
  • Added workflows/verify_tabarena_predictions.py to repeat the reader check in an environment with TabArena installed.
  • No successful cloud upload/download round trip is claimed.

@adrian-prior
adrian-prior added this pull request to stack #43 September 11, 2026 15:37
@adrian-prior adrian-prior changed the title Add portable prediction and label artifacts Save prediction artifacts in TabArena's result format Sep 16, 2026
@adrian-prior
adrian-prior removed this pull request from stack #43 September 16, 2026 12:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant