Skip to content

Add derived dataset provenance to the ro-crate - #3

Merged
jswelling merged 59 commits into
mainfrom
welling/crate_builder_3
Sep 29, 2026
Merged

jswelling merged 59 commits into
mainfrom
welling/crate_builder_3

Conversation

@jswelling

Copy link
Copy Markdown
Collaborator

This version of build_crate_from_dataset.py produces a fairly complete rocrate and croissant pair for derived datasets. For primary datasets, the croissant.json is fairly complete but the provenance in the rocrate is very incomplete. Also, for primary datasets with thousands of files, both json files will be unwieldy.

Call the routine as:

env AUTH_TOK=<valid token> python build_crate_from_dataset.py -o <output-dir> HMxxx.xxxx.xxx

This will produce a file HMxxx.xxxx.xxx_crate.zip in the given output directory. Supported options include:

--debug/-d for debugging information
--include-all-files to include all files for a derived dataset, not just data product/quality control files

To validate the result, cd to that directory and do:

rocrate-validator validate HBMxxx.xxxx.xxx_crate.zip
unzip HBMxxx.xxx.xxx_crate.zip
mlcroissant validate --jsonld croissant.json

The first command validates the ro-crate json; the second validates the croissant.

Some limitations of this version:

  • For primary datasets, the ro-crate does not appropriately represent the provenance for primary datasets, but the croissant information does.
  • The program references the UUID endpoint for file information.
  • Only a subset of possible EDAM codes are supported. An unknown EDAM code will produce a warning.
  • Inferring MIME types from EDAM codes is not always possible. The code should be expanded to include the file extension.
  • Version numbers for some things (python, CWL, Airflow) are hard-coded.

@jswelling
jswelling marked this pull request as draft September 3, 2026 16:12
@jswelling
jswelling marked this pull request as ready for review September 3, 2026 16:25
@jswelling

jswelling commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator Author

Still pending:

  • Do we use the parent DOI for the Croissant entry for derived datasets that have no DOI? If so, what do we use for snare-seq?
    Deferred for the future:
  • snare-seq fails.
  • Build proper provenance for primary datasets, perhaps using protocols.io data.

@fjorka fjorka left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Five comments, all one issue: entity @ids are minted from the vocabulary namespace and
resolve to 404. This came from my code originally — I've applied the same fix on my side
and regenerated my examples.

Comment thread src/crate_builder_script/api_calls.py
Comment thread src/crate_builder_script/croissant_wrapper.py Outdated
Comment thread src/crate_builder_script/croissant_wrapper.py Outdated
Comment thread src/crate_builder_script/croissant_wrapper.py Outdated
Comment thread src/crate_builder_script/croissant_wrapper.py Outdated
@jswelling
jswelling merged commit 93f4b2a into main Sep 29, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants