Forward-merge release/26.10 into main - #2661
Merged
Merged
Conversation
## Summary This PR improves data-transfer and compute overlap for host-resident out-of-core KMeans using a cyclic two-buffer pipeline. - Batch 1 follows batch 0 on the copy stream while batch 0 compute starts as soon as its transfer completes. - Subsequent transfers overlap computation of the current batch. - Buffer recycling continues across iteration boundaries, allowing batch 0 of the next pass to be prefetched during the previous pass. - Device-resident inputs remain zero-copy. ## Implementation - Adds a private KMeans batch loader with explicit staged, acquired, and reusable buffer states. - Uses CUDA events so compute waits for H2D completion and a buffer cannot be overwritten until all of its consumers have been submitted. - Applies the same dependency-driven scheduling to every batch, without special APIs or states for the first two batches. - Uses persistent device scratch to avoid per-batch deallocation and its potential device-wide synchronization. - Computes final inertia through the regular batched assignment and reduction pipeline, accumulates it on device, and copies only the final result to host. - Leaves the shared ANN batch iterator unchanged. ## Benchmark under similar configuration 10 GiB pinned-host FP32 dataset (10,485,760 × 256), 2,560 clusters, three iterations, ten 1 GiB out-of-core batches, and 131,072-sample assignment tiles. | Metric | `main` | PR | |---|---:|---:| | Median runtime | 1.7795 s | **0.7962 s** | | Speedup | 1.00× | **2.23×** | | Effective bulk throughput | 22.48 GiB/s | **50.24 GiB/s** | | Profiled GPU span | 1773.12 ms | **785.65 ms** | | H2D time | 752.29 ms | 752.38 ms | | Kernel time | 1003.08 ms | 760.97 ms | | H2D/kernel overlap | 0.00 ms | **727.91 ms** | | H2D overlapped by kernels | 0.0% | **96.75%** | | Kernels overlapped by H2D | 0.0% | **95.66%** | ## Profile Main branch : <img width="1304" height="98" alt="profile_main" src="https://github.com/user-attachments/assets/fe95891e-0e5e-4feb-8081-b71c7a6524a2" /> This PR : <img width="1340" height="137" alt="profile_pr" src="https://github.com/user-attachments/assets/262583c6-40f0-4bc5-9e81-1e4f82437650" /> This PR (multi-GPU) : <img width="1642" height="64" alt="multi_gpu_profile" src="https://github.com/user-attachments/assets/ecfd6f45-2627-4131-84fe-cedfa2dbc24c" /> Authors: - Victor Lafargue (https://github.com/viclafargue) Approvers: - https://github.com/irina-resh-nvda - Tarang Jain (https://github.com/tarang-jain) URL: #2538
Contributor
Author
|
SUCCESS - forward-merge complete. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Forward-merge triggered by push to release/26.10 that creates a PR to keep main up-to-date. If this PR is unable to be immediately merged due to conflicts, it will remain open for the team to manually merge. See forward-merger docs for more info.