From 206d5dcf92e00df471d560dbae233d43d6d0bcef Mon Sep 17 00:00:00 2001 From: Juho Vainio Date: Wed, 30 Sep 2026 13:06:16 +0300 Subject: [PATCH 1/2] docs(ci): document the three-stage validation ladder Add the missing framing to docs/ci-hardware-testing.md: the three self-hosted GPU gates (per-PR smoke, nightly coverage, release-candidate regression) and how they differ; the validation axes (OS, hardware, channel, SDK-version); why the serve engine isn't an independently selectable matrix axis (quotes effective_serve_engine()); and a short addendum on e2e-prewarm's --version/--build-date pin flags. No lane, artifact, or job-order changes, so every contract-tested table/list/ sentence in the doc is left byte-identical. ROCMAI-431 Signed-off-by: Juho Vainio --- docs/ci-hardware-testing.md | 78 ++++++++++++++++++++++++++++++++++++- 1 file changed, 76 insertions(+), 2 deletions(-) diff --git a/docs/ci-hardware-testing.md b/docs/ci-hardware-testing.md index b59ba8f45..908abe13d 100644 --- a/docs/ci-hardware-testing.md +++ b/docs/ci-hardware-testing.md @@ -19,7 +19,39 @@ newer run's merge-required (GitHub-hosted) checks would sit pending forever (observed on PR #138). Giving the self-hosted lanes their own workflow — and thus their own concurrency group — means an offline runner can only ever stall that workflow's own supersession, never `ci.yml`'s required checks. See -`EAI-7548`. +`EAI-7548`. `xtask/src/workflow_contract.rs`'s +`self_hosted_workflow_owns_the_gpu_lanes` test pins `e2e-gpu`, +`e2e-gpu-strix-ubuntu`, and `e2e-gpu-strix-windows` by name in +`e2e-selfhosted.yml`, so this split cannot be silently undone. + +## The three-stage validation ladder + +Three self-hosted GPU gates run at different points in a change's life. None +are required-status checks (see "Blocking vs. non-blocking" below); what +differs is when each runs and what hardware it covers. + +| Rung | Workflow | Trigger | Lanes | Blocks a merge/tag? | +|---|---|---|---|---| +| Per-PR smoke gate | `e2e-selfhosted.yml` | `push`/`pull_request`/`merge_group` (see "Triggers") | The 4 jobs in the Platforms table below | No — absent from the required-status-check list | +| Nightly coverage gate | `nightly.yml` | `schedule` (06:00 UTC daily) + `workflow_dispatch` | The same 4 platforms plus R9700 (`e2e-gpu-nightly-rad3`) and MI350P (`e2e-gpu-nightly-mi350p`), each run against both the `release` and `nightly` package channel (`strategy.matrix.channel: [release, nightly]`, ROCMAI-429) | No — not part of any PR or push-to-main event | +| Release-candidate regression gate | `e2e-selfhosted.yml` | `push` to a `release/**` branch (ROCMAI-120/EAI-8761), ahead of cutting the `v*` tag `release.yml` publishes from | The same 4 per-PR lanes | No — same non-required status as the per-PR gate; a hardware signal for whoever cuts the tag, not an automated block | + +R9700 and MI350P were dropped from the per-PR gate (ROCMAI-125): still +covered, but off every PR's critical path, so they moved to the nightly-only +rung instead of being removed outright. + +The release-candidate rung's `release/**` trigger ships in PR #415; pinning +the SDK version that rung's pre-warmed runtime resolves to (`cargo xtask +e2e-prewarm --version ` / `--build-date `, so a release-branch run +can hold at `n-1`/`n-2` instead of always tracking the latest channel index) +ships in PR #464, stacked on #415. Neither is merged as of this writing, so +`e2e-selfhosted.yml` on `main` today still triggers on `push: branches: +[main]` only. Wiring an actual `sdk_version: [current, n-1, n-2]` matrix axis +into the release-branch trigger is not part of either PR: nothing in this +repo maps "n-1"/"n-2" to a concrete SDK version (`therock.rs`'s +index-version parsers are private), so the release gate's SDK version is +whatever `--version`/`--build-date` its caller passes, not an automatic +3-way matrix. ## Platforms @@ -151,7 +183,10 @@ platforms (it also runs `e2e-gpu-rad3` and `e2e-gpu-mi350p`, demoted from per-PR to nightly-only per ROCMAI-125) with the `@nightly` scenarios included. Each joins its platforms' reports — including partial or failed runs — by scenario id into one HTML report -and GitHub step summary. +and GitHub step summary. Since ROCMAI-429 the report keys each column by +`(platform_slug, channel)` rather than platform alone, so `nightly.yml`'s +release/nightly channel matrix renders two columns per platform instead of +one channel's result silently overwriting the other. The lane artifacts are named canonically (`e2e-report`, `e2e-gpu-report`, `e2e-gpu-rad3-report`, `e2e-gpu-mi350p-report`, `e2e-gpu-strix-ubuntu-report`, @@ -168,6 +203,39 @@ guessed platform on Linux, which would report a Windows lane as Linux; `xtask`'s `every_uploaded_e2e_artifact_has_a_name_the_report_can_label` guards against it. +## Engine is not an independently selectable axis + +None of the tables above vary the serve engine as a matrix dimension, because +the engine is not selectable independent of hardware and OS. `rocm serve` +picks it via `effective_serve_engine()` in +`tests/e2e-cucumber/src/capability.rs` (mirroring the product's own +`preferred_serve_engine_for_host_gpu_summary`): + +```rust +pub fn effective_serve_engine(gfx_target: Option<&str>, os_family: &str) -> String { + if os_family.eq_ignore_ascii_case("windows") { + return "lemonade".to_owned(); + } + if family_prefers_vllm(gfx_target) { + "vllm".to_owned() + } else { + "lemonade".to_owned() + } +} +``` + +Any Windows host resolves to `lemonade` (the vLLM adapter does not run +there); otherwise `vllm` is only preferred for the `*-dcgpu` families and +`gfx906`/`gfx908`/`gfx90a`. Consequences for the lanes above: + +- Strix Halo is `gfx1151` on every OS, so all three per-PR Strix lanes + (Ubuntu, Windows, WSL2) and their nightly counterparts are lemonade-only — + no matrix value turns a Strix lane into a vLLM lane. +- `e2e-gpu` (MI300X) is the only per-PR lane where vLLM is the effective + engine. +- Adding a vLLM lane means adding hardware from a vLLM-eligible family, not + adding an `engine:` value to a workflow matrix. + ## Triggers The GPU jobs (in `e2e-selfhosted.yml`) run automatically on `push`, @@ -276,6 +344,12 @@ Pre-warm then: - prunes with `rocm storage remove-old-installs` after any install, update, or repair, so the multi-version cache stays bounded. +`e2e-prewarm` also accepts mutually exclusive `--version`/`--build-date` +flags (ROCMAI-430) that pin the SDK build the pre-warm resolves to instead of +always tracking whatever the channel index currently serves; the unpinned +invocation above is unaffected, and today no caller in this repo passes +either flag yet (see "The three-stage validation ladder" above). + The runtime is always installed **in place**: `install sdk` bakes absolute paths into the runtime manifest, so a tree that is moved after installation leaves every serve pointing at a path that no longer exists. From 3b59f966832b008db6aa221aed97b174f0c2e799 Mon Sep 17 00:00:00 2001 From: Juho Vainio Date: Wed, 30 Sep 2026 19:04:18 +0300 Subject: [PATCH 2/2] docs(ci): fix e2e-prewarm tense, reword two overstated doc claims MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Review found the --version/--build-date paragraph claiming present-tense availability for a flag pair that isn't in this tree yet (ships in #464, stacked on #415, neither merged) — contradicted the ladder section's own "ships in #464" framing two paragraphs up. - e2e-prewarm --version/--build-date: reworded to future tense, matching the ladder section. - Per-PR row: note e2e-gpu-strix-ubuntu already runs both channel legs per PR, not just 4 flat jobs. - "cannot be silently undone": reworded to match what self_hosted_workflow_owns_the_gpu_lanes actually checks (job keys present as text, not a structural pin). Left the two open-decisions-for-the-author items untouched per review (non-defects). Signed-off-by: Juho Vainio --- docs/ci-hardware-testing.md | 18 ++++++++++-------- 1 file changed, 10 insertions(+), 8 deletions(-) diff --git a/docs/ci-hardware-testing.md b/docs/ci-hardware-testing.md index 908abe13d..37b8c9226 100644 --- a/docs/ci-hardware-testing.md +++ b/docs/ci-hardware-testing.md @@ -20,9 +20,10 @@ newer run's merge-required (GitHub-hosted) checks would sit pending forever thus their own concurrency group — means an offline runner can only ever stall that workflow's own supersession, never `ci.yml`'s required checks. See `EAI-7548`. `xtask/src/workflow_contract.rs`'s -`self_hosted_workflow_owns_the_gpu_lanes` test pins `e2e-gpu`, -`e2e-gpu-strix-ubuntu`, and `e2e-gpu-strix-windows` by name in -`e2e-selfhosted.yml`, so this split cannot be silently undone. +`self_hosted_workflow_owns_the_gpu_lanes` test requires `e2e-gpu`, +`e2e-gpu-strix-ubuntu`, and `e2e-gpu-strix-windows` to appear as job keys in +`e2e-selfhosted.yml`, so a PR that moves one of these lanes back into +`ci.yml` fails CI rather than merging unnoticed. ## The three-stage validation ladder @@ -32,7 +33,7 @@ differs is when each runs and what hardware it covers. | Rung | Workflow | Trigger | Lanes | Blocks a merge/tag? | |---|---|---|---|---| -| Per-PR smoke gate | `e2e-selfhosted.yml` | `push`/`pull_request`/`merge_group` (see "Triggers") | The 4 jobs in the Platforms table below | No — absent from the required-status-check list | +| Per-PR smoke gate | `e2e-selfhosted.yml` | `push`/`pull_request`/`merge_group` (see "Triggers") | The 4 jobs in the Platforms table below (`e2e-gpu-strix-ubuntu` runs both its `[release, nightly]` channel legs per PR, ROCMAI-125) | No — absent from the required-status-check list | | Nightly coverage gate | `nightly.yml` | `schedule` (06:00 UTC daily) + `workflow_dispatch` | The same 4 platforms plus R9700 (`e2e-gpu-nightly-rad3`) and MI350P (`e2e-gpu-nightly-mi350p`), each run against both the `release` and `nightly` package channel (`strategy.matrix.channel: [release, nightly]`, ROCMAI-429) | No — not part of any PR or push-to-main event | | Release-candidate regression gate | `e2e-selfhosted.yml` | `push` to a `release/**` branch (ROCMAI-120/EAI-8761), ahead of cutting the `v*` tag `release.yml` publishes from | The same 4 per-PR lanes | No — same non-required status as the per-PR gate; a hardware signal for whoever cuts the tag, not an automated block | @@ -344,11 +345,12 @@ Pre-warm then: - prunes with `rocm storage remove-old-installs` after any install, update, or repair, so the multi-version cache stays bounded. -`e2e-prewarm` also accepts mutually exclusive `--version`/`--build-date` +`e2e-prewarm` will also accept mutually exclusive `--version`/`--build-date` flags (ROCMAI-430) that pin the SDK build the pre-warm resolves to instead of -always tracking whatever the channel index currently serves; the unpinned -invocation above is unaffected, and today no caller in this repo passes -either flag yet (see "The three-stage validation ladder" above). +always tracking whatever the channel index currently serves. That flag pair +ships in PR #464, stacked on #415, neither merged as of this writing (see +"The three-stage validation ladder" above) — the unpinned invocation above is +what every lane in this tree runs today. The runtime is always installed **in place**: `install sdk` bakes absolute paths into the runtime manifest, so a tree that is moved after installation leaves every