Skip to content

build(deps): update trl requirement from <=0.21.0 to <=1.14.0 - #763

Closed
dependabot[bot] wants to merge 1 commit into
mainfrom
dependabot/pip/trl-lte-1.14.0
Closed

dependabot[bot] wants to merge 1 commit into
mainfrom
dependabot/pip/trl-lte-1.14.0

Conversation

@dependabot

@dependabot dependabot Bot commented on behalf of github Sep 30, 2026

Copy link
Copy Markdown
Contributor

Updates the requirements on trl to permit the latest version.

Release notes

Sourced from trl's releases.

v1.14.0

Features

trl.losses is gone: DPO, KTO and GRPO stream their own log-probs

[!WARNING] from trl.losses import FusedLinearDPOLoss (or FusedLinearKTOLoss, FusedLinearGRPOLoss, FusedLinearJSDLoss) no longer works. The module introduced in v1.13 has been removed.

v1.13 vendored Liger's chunked_loss into trl.losses as a holding action, not a destination (#7063). The problem it was holding: those classes reimplement each trainer's loss math, so TRL carried two copies of every formula, and copies drift. That drift is where the bugs were. GRPO's grad_norm differed from the default path, KTO ignored the reference model under PEFT, KTO class weights and DPO label_smoothing were silently dropped. Worst of all, under plain DDP the fused path ran the loss on the unwrapped model, so DistributedDataParallel's reducer was never armed and gradients were never all-reduced: every rank silently kept its own.

The fix is to stop reimplementing. All DPO, KTO and GRPO need from the fused path is per-token log-probs of the selected tokens, and TRL already had _ChunkedLogProbFunction for exactly that: it streams the vocabulary with an online logsumexp and recomputes in the backward pass. Each trainer now runs backbone -> _ChunkedLogProbFunction -> its own existing loss code, unchanged. One implementation of each loss, and the memory win is kept, because full logits are still never materialized.

Thirteen restrictions existed only because the loss had been rewritten in a form that could not express those options. Most are gone: DPO with use_liger_kernel=True now accepts mixed loss types, f-divergences and precomputed reference log-probs, and GRPO gains entropy and off-policy masking. Still refused: use_weighting, compute_metrics, a PEFT adapter on lm_head, prompt-learning PEFT, and the MoE auxiliary loss.

Single H100, Qwen3-0.6B, bs=4, seq 512:

config median step peak
v1.13 fused, inline ref 0.2243 s 7.05 GB
v1.14 chunked, inline ref 0.2457 s (+9.5%) 7.05 GB
v1.14 chunked, precompute_ref_log_probs=True 0.1971 s (-12.1%) 4.83 GB (-31%)

The third row is the point: that configuration did not exist before, because the fused path rejected precomputed reference log-probs outright.

use_liger_kernel=True still works and still enables Liger's model kernels through transformers. In DPO, GRPO and KTO it now selects TRL's chunked log-probability path instead of Liger's fused loss.

A fused Triton logprob + entropy kernel now ships in TRL

The loss head is the only hot path TRL owns: transformers already kernelizes the model internals, and nothing on the Hub covers what happens after the decoder. On one H100 (bf16, V=151936, H=4096, 8192 tokens), selective_log_softmax plus entropy_from_logits cost 12.8 ms and 2.32 GiB over [8, 1024, 151936]. A fused Triton kernel does the same in 0.89 ms, roughly 14x, and lands closer to the fp32 reference than the bf16 path it replaces. GRPO runs this two to four times per step.

It now lives in-tree at trl.kernels and is on by default on CUDA, ROCm and XPU, with the PyTorch implementation as fallback. It was briefly loaded from the Hub; that was reversed because this kernel is pure Triton and compiles at runtime, so the Hub bought nothing while costing a version gate unrelated to whether the kernel works, a download that fails offline, trust_remote_code=True for anyone with kernels installed, a silent except Exception: pass wrapped around all of it, and 13 of 13 tests skipped in CI.

Other

... (truncated)

Commits
  • c6a6a16 Release: v1.14 (#7392)
  • cc5bb3f Agree on logged keys across ranks before flushing GRPO and RLOO metrics (#7382)
  • 5469923 Ship the logprob and entropy kernel in trl instead of the Hub (#7363)
  • 824726e Add support for vLLM 0.30.0 (#7365)
  • 78770da Document what use_liger_kernel actually does in each trainer (#7361)
  • 47cf654 Drop vLLM 0.19.1 support (#7366)
  • 2c34300 Set the pad token for vision datasets too (#7316)
  • eb5863f Test that the reward model pad token is set on the text config (#7350)
  • be9e837 Hotfix CI: Expect the Llava assistant masks tests to pass now that transforme...
  • c8ef243 Read the MoE auxiliary loss coefficient from the model config (#7248)
  • Additional commits viewable in compare view

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.


Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

  • @dependabot rebase will rebase this PR
  • @dependabot recreate will recreate this PR, overwriting any edits that have been made to it
  • @dependabot show <dependency name> ignore conditions will show all of the ignore conditions of the specified dependency
  • @dependabot ignore this major version will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this minor version will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)

Updates the requirements on [trl](https://github.com/huggingface/trl) to permit the latest version.
- [Release notes](https://github.com/huggingface/trl/releases)
- [Changelog](https://github.com/huggingface/trl/blob/main/RELEASE.md)
- [Commits](huggingface/trl@v0.2.0...v1.14.0)

---
updated-dependencies:
- dependency-name: trl
  dependency-version: 1.14.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
@dependabot @github

dependabot Bot commented on behalf of github Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

Superseded by #766.

@dependabot dependabot Bot closed this Oct 2, 2026
@dependabot
dependabot Bot deleted the dependabot/pip/trl-lte-1.14.0 branch October 2, 2026 21:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants