Skip to content

Fix Metrics.drop corpus aggregation to use mean instead of max - #1422

Open
salignatmoandal wants to merge 1 commit into
huggingface:mainfrom
salignatmoandal:fix/drop-corpus-aggregation-mean
Open

salignatmoandal wants to merge 1 commit into
huggingface:mainfrom
salignatmoandal:fix/drop-corpus-aggregation-mean

Conversation

@salignatmoandal

@salignatmoandal salignatmoandal commented Oct 9, 2026 •

Copy link
Copy Markdown

Summary

  • Change Metrics.drop corpus aggregation from max to np.mean for both em and f1.
  • Keep per-sample max over alternate gold answers inside DropMetrics; only the across-document aggregation was wrong.
  • Add a regression test so one correct document among four reports 0.25, not 1.0.

Closes #1395.

Context

DropMetrics.compute already takes the max over gold spans for a single document. MetricsLogger.aggregate then applied corpus_level_fn={"em": max, "f1": max}, so any single exact hit made the reported task score 1.0. That diverges from lm-eval-harness DROP (aggregation: mean) and from Metrics.exact_match in the same file.

The built-in lighteval|drop task uses Metrics.exact_match, so its numbers are unaffected. Custom tasks that select Metrics.drop were reporting inflated scores.

Test plan

  • pytest tests/unit/metrics/test_drop_aggregation.py — 2 passed
  • ruff check / ruff format --check on changed files
  • CI unit tests on the PR

Per-sample max over alternate golds already lives in DropMetrics; using
max across documents made one correct sample report em/f1 of 1.0.
Closes huggingface#1395.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Metrics.drop aggregates em/f1 across documents with max, so one correct sample reports 1.0

1 participant