Integrate the coarse distribution exactly. Sample only the fine-tail residual.
Penghui Yang1,2 · Long Xing3 · Xuanlang Dai2,4 · Ziyu Liu1,2 · Kai Chen2 · Yuhang Zang2
1 Shanghai Jiao Tong University 2 Shanghai Artificial Intelligence Laboratory
3 University of Science and Technology of China 4 Fudan University
Quick Start · Results · Method · Implementation · Citation
ResOPD estimates the full-vocabulary reverse-KL gradient from the teacher's Top-k probabilities and the sampled-token score. It computes the observable coarse gradient exactly and samples only the unresolved within-tail residual, using the same sparse teacher payload and no additional teacher forward passes.
This repository provides the core implementation and its integration with verl. The paper contains the derivations and experiments. UPSTREAM.md records the pinned framework base and release scope.
| Evaluation | Reported result | Setting |
|---|---|---|
| Exact gradient audits | 52.5–75.0% lower covariance trace than sampled-token OPD | Five frozen-prefix panels; teacher Top-4 |
| Mathematical reasoning | 50.12% average score vs. 45.72% for HETS | Qwen3.5-27B → Qwen3.5-4B |
| Science and coding transfer | 44.15% average score vs. 43.63% for HETS | Qwen3-30B-A3B-Instruct → Qwen3-4B |
These are results from the paper's experimental protocols. The launcher below uses the released DeepMath training and held-out sets, with a shorter response budget for the integration example.
Use Python 3.12 and a GPU environment compatible with the pinned verl base.
git clone https://github.com/InternLM/ResOPD.git
cd ResOPD
uv sync --extra fsdp --extra vllm
source .venv/bin/activateThe separate recipe submodule is optional for this example. See
upstream installation guidance for alternative
environments.
Download the paper's processed DeepMath subset from ygyjrc/ResOPD-deepmath. The files are already in verl format and can be used directly:
python - <<'PYDATA'
from huggingface_hub import hf_hub_download
for filename in ("train.parquet", "heldout.parquet"):
hf_hub_download(
repo_id="ygyjrc/ResOPD-deepmath",
repo_type="dataset",
filename=filename,
local_dir="./data/resopd-deepmath",
)
PYDATA| File | Samples | Use |
|---|---|---|
train.parquet |
1,024 | Distillation training |
heldout.parquet |
252 | Held-out evaluation (VAL_FILE) |
The launcher defaults to these files under data/resopd-deepmath. To use another
local directory, set DATA_DIR, or override TRAIN_FILE and VAL_FILE individually.
The held-out split comes from the same screened DeepMath subset; the paper also
reports separate mathematical reasoning benchmarks. Evaluation uses verl's
boxed-answer math scorer; task rewards remain disabled for distillation.
# Print the resolved command.
DRY_RUN=1 bash examples/resopd/run_resopd.sh
# Start synchronous FSDP2/vLLM distillation.
bash examples/resopd/run_resopd.sh| Default | Value |
|---|---|
| Student / teacher | Qwen3.5-4B / Qwen3.5-27B |
| GPU allocation | 4 student GPUs + 4 separate teacher GPUs, on one node |
| Teacher support | Top-k = 4 |
| Response budget | 32,768 tokens |
| Student sampling | Temperature = 1, top-p = 1, top-k = −1 |
| Update schedule | One PPO epoch; one mini-batch per fresh rollout batch |
Adjust model IDs, GPU allocation, tensor parallelism, and sequence lengths for the available hardware. The student and teacher must share the token-ID vocabulary. Extra Hydra overrides can be passed directly to the launcher.
Add ResOPD to an existing verl OPD configuration
distillation:
enabled: true
distillation_loss:
loss_mode: resopd
topk: 4
use_task_rewards: false
use_policy_gradient: false
loss_max_clamp: null
log_prob_min_clamp: null
actor_rollout_ref:
model:
use_fused_kernels: false
actor:
use_torch_compile: false
ppo_epochs: 1
rollout:
temperature: 1.0
top_p: 1.0
top_k: -1Keep use_chunked_topk=false and set ppo_mini_batch_size equal to the
training batch size. Set loss_mode: resopd to select the estimator.
ResOPD separates an exact coarse update from a sampled fine-tail correction:
- Integrate the head. Evaluate the teacher Top-k contributions analytically.
- Account for the tail bucket. Use the residual student and teacher masses to compute the observable aggregate tail gradient.
- Sample the residual. On tail hits, add only the unresolved fine-tail correction.
At a fixed prefix, let
The code uses an algebraically equivalent centered support-event control. Score coefficients and the event-centering mass are detached: the forward value is used for logging, while backward supplies the estimator above.
Sampling matters. Unbiasedness assumes draws from the current full student distribution. The launcher uses full-support sampling and a single update per rollout batch. Truncated sampling, stale rollouts, repeated optimization passes, and sampling-altering logits processors require separate treatment, as discussed in the paper. No uniform variance ordering is claimed for every distribution.
| Component | Source |
|---|---|
| PyTorch gradient estimator | resopd.py |
| Packed tensors and sequence-parallel alignment | fsdp/losses.py |
| Loss registration and response-masked metrics | distillation/losses.py |
| Teacher Top-k and observed-token score extraction | prompt_logprobs.py |
| Held-out math scoring | deepmath_reward.py |
| Training entry point | run_resopd.sh |
| Estimator, payload, and integration checks | tests/resopd |
Supported path: vLLM teacher scoring and eager FSDP/FSDP2/VeOmni student logits; the launcher uses FSDP2. Megatron students, SGLang teachers, fused student loss kernels, and chunked Top-k are not enabled for ResOPD in this release. The student still performs full-vocabulary normalization; sparse feedback reduces teacher communication.
For the underlying framework documentation, see README.verl.md.
Standalone estimator and teacher-payload checks require only PyTorch and pytest:
python -m pytest tests/resopd/test_estimator.py tests/resopd/test_teacher_payload.py -qWith verl's runtime dependencies installed, run the CPU integration checks:
python -m pytest tests/resopd/test_verl_integration.py -qThe checks cover exact finite-action gradients, the reverse-KL expectation, variance identities, numerical precision, teacher-score alignment, and the verl data path. Release validation records 22 standalone checks and 13 integration checks, plus comparison with the original research kernel. The GitHub workflow runs the standalone checks. A full GPU training run of this extracted port is outside the recorded release validation.
If you use ResOPD, please cite the paper:
@misc{yang2026resopd,
title = {ResOPD: Tail Residualization for Sparse On-Policy Distillation},
author = {Penghui Yang and Long Xing and Xuanlang Dai and Ziyu Liu and Kai Chen and Yuhang Zang},
year = {2026},
eprint = {2610.04882},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2610.04882}
}GitHub's Cite this repository entry also points to the paper through CITATION.cff.
Released under Apache-2.0. ResOPD builds on verl and retains its original copyright and license notices. The overview and framework figures are from the ResOPD paper.

