Skip to content

Forward-merge release/26.10 into main - #2661

Merged
GPUtester merged 1 commit into
mainfrom
release/26.10
Sep 18, 2026
Merged

GPUtester merged 1 commit into
mainfrom
release/26.10

Conversation

@rapids-bot

@rapids-bot rapids-bot Bot commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Forward-merge triggered by push to release/26.10 that creates a PR to keep main up-to-date. If this PR is unable to be immediately merged due to conflicts, it will remain open for the team to manually merge. See forward-merger docs for more info.

## Summary

This PR improves data-transfer and compute overlap for host-resident out-of-core KMeans using a cyclic two-buffer pipeline.

- Batch 1 follows batch 0 on the copy stream while batch 0 compute starts as soon as its transfer completes.
- Subsequent transfers overlap computation of the current batch.
- Buffer recycling continues across iteration boundaries, allowing batch 0 of the next pass to be prefetched during the previous pass.
- Device-resident inputs remain zero-copy.

## Implementation

- Adds a private KMeans batch loader with explicit staged, acquired, and reusable buffer states.
- Uses CUDA events so compute waits for H2D completion and a buffer cannot be overwritten until all of its consumers have been submitted.
- Applies the same dependency-driven scheduling to every batch, without special APIs or states for the first two batches.
- Uses persistent device scratch to avoid per-batch deallocation and its potential device-wide synchronization.
- Computes final inertia through the regular batched assignment and reduction pipeline, accumulates it on device, and copies only the final result to host.
- Leaves the shared ANN batch iterator unchanged.

## Benchmark under similar configuration

10 GiB pinned-host FP32 dataset (10,485,760 × 256), 2,560 clusters, three iterations, ten 1 GiB out-of-core batches, and 131,072-sample assignment tiles.

| Metric | `main` | PR |
|---|---:|---:|
| Median runtime | 1.7795 s | **0.7962 s** |
| Speedup | 1.00× | **2.23×** |
| Effective bulk throughput | 22.48 GiB/s | **50.24 GiB/s** |
| Profiled GPU span | 1773.12 ms | **785.65 ms** |
| H2D time | 752.29 ms | 752.38 ms |
| Kernel time | 1003.08 ms | 760.97 ms |
| H2D/kernel overlap | 0.00 ms | **727.91 ms** |
| H2D overlapped by kernels | 0.0% | **96.75%** |
| Kernels overlapped by H2D | 0.0% | **95.66%** |


## Profile

Main branch :
<img width="1304" height="98" alt="profile_main" src="https://github.com/user-attachments/assets/fe95891e-0e5e-4feb-8081-b71c7a6524a2" />

This PR :
<img width="1340" height="137" alt="profile_pr" src="https://github.com/user-attachments/assets/262583c6-40f0-4bc5-9e81-1e4f82437650" />

This PR (multi-GPU) :
<img width="1642" height="64" alt="multi_gpu_profile" src="https://github.com/user-attachments/assets/ecfd6f45-2627-4131-84fe-cedfa2dbc24c" />

Authors:
  - Victor Lafargue (https://github.com/viclafargue)

Approvers:
  - https://github.com/irina-resh-nvda
  - Tarang Jain (https://github.com/tarang-jain)

URL: #2538
@rapids-bot
rapids-bot Bot requested a review from a team as a code owner September 18, 2026 23:14
@GPUtester
GPUtester merged commit 67885e6 into main Sep 18, 2026
1 check passed
@rapids-bot

rapids-bot Bot commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

SUCCESS - forward-merge complete.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants