Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
80 changes: 78 additions & 2 deletions docs/ci-hardware-testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,40 @@ newer run's merge-required (GitHub-hosted) checks would sit pending forever
(observed on PR #138). Giving the self-hosted lanes their own workflow — and
thus their own concurrency group — means an offline runner can only ever stall
that workflow's own supersession, never `ci.yml`'s required checks. See
`EAI-7548`.
`EAI-7548`. `xtask/src/workflow_contract.rs`'s

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"cannot be silently undone" overstates the test (non-blocking). self_hosted_workflow_owns_the_gpu_lanes (xtask/src/workflow_contract.rs) only asserts sh.contains(job) for the three job keys, a substring check that any mention in a comment satisfies. It also does not cover e2e-wsl, which a separate test handles. Suggest "requires the three job keys to appear in the file".

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6534052 — replaced with your suggested wording: the test requires the three job keys to appear in e2e-selfhosted.yml, so moving one back to ci.yml fails CI rather than merging unnoticed.

`self_hosted_workflow_owns_the_gpu_lanes` test requires `e2e-gpu`,
`e2e-gpu-strix-ubuntu`, and `e2e-gpu-strix-windows` to appear as job keys in
`e2e-selfhosted.yml`, so a PR that moves one of these lanes back into
`ci.yml` fails CI rather than merging unnoticed.

## The three-stage validation ladder

Three self-hosted GPU gates run at different points in a change's life. None
are required-status checks (see "Blocking vs. non-blocking" below); what
differs is when each runs and what hardware it covers.

| Rung | Workflow | Trigger | Lanes | Blocks a merge/tag? |
|---|---|---|---|---|
| Per-PR smoke gate | `e2e-selfhosted.yml` | `push`/`pull_request`/`merge_group` (see "Triggers") | The 4 jobs in the Platforms table below (`e2e-gpu-strix-ubuntu` runs both its `[release, nightly]` channel legs per PR, ROCMAI-125) | No — absent from the required-status-check list |
| Nightly coverage gate | `nightly.yml` | `schedule` (06:00 UTC daily) + `workflow_dispatch` | The same 4 platforms plus R9700 (`e2e-gpu-nightly-rad3`) and MI350P (`e2e-gpu-nightly-mi350p`), each run against both the `release` and `nightly` package channel (`strategy.matrix.channel: [release, nightly]`, ROCMAI-429) | No — not part of any PR or push-to-main event |
| Release-candidate regression gate | `e2e-selfhosted.yml` | `push` to a `release/**` branch (ROCMAI-120/EAI-8761), ahead of cutting the `v*` tag `release.yml` publishes from | The same 4 per-PR lanes | No — same non-required status as the per-PR gate; a hardware signal for whoever cuts the tag, not an automated block |

R9700 and MI350P were dropped from the per-PR gate (ROCMAI-125): still
covered, but off every PR's critical path, so they moved to the nightly-only
rung instead of being removed outright.

The release-candidate rung's `release/**` trigger ships in PR #415; pinning
the SDK version that rung's pre-warmed runtime resolves to (`cargo xtask
e2e-prewarm --version <ver>` / `--build-date <date>`, so a release-branch run
can hold at `n-1`/`n-2` instead of always tracking the latest channel index)
ships in PR #464, stacked on #415. Neither is merged as of this writing, so
`e2e-selfhosted.yml` on `main` today still triggers on `push: branches:
[main]` only. Wiring an actual `sdk_version: [current, n-1, n-2]` matrix axis
into the release-branch trigger is not part of either PR: nothing in this
repo maps "n-1"/"n-2" to a concrete SDK version (`therock.rs`'s
index-version parsers are private), so the release gate's SDK version is
whatever `--version`/`--build-date` its caller passes, not an automatic
3-way matrix.

## Platforms

Expand Down Expand Up @@ -151,7 +184,10 @@ platforms (it also runs `e2e-gpu-rad3` and `e2e-gpu-mi350p`, demoted from
per-PR to nightly-only per ROCMAI-125) with the `@nightly` scenarios
included. Each joins its platforms'
reports — including partial or failed runs — by scenario id into one HTML report
and GitHub step summary.
and GitHub step summary. Since ROCMAI-429 the report keys each column by
`(platform_slug, channel)` rather than platform alone, so `nightly.yml`'s
release/nightly channel matrix renders two columns per platform instead of
one channel's result silently overwriting the other.

The lane artifacts are named canonically (`e2e-report`, `e2e-gpu-report`,
`e2e-gpu-rad3-report`, `e2e-gpu-mi350p-report`, `e2e-gpu-strix-ubuntu-report`,
Expand All @@ -168,6 +204,39 @@ guessed platform on Linux, which would report a Windows lane as Linux;
`xtask`'s `every_uploaded_e2e_artifact_has_a_name_the_report_can_label` guards
against it.

## Engine is not an independently selectable axis

None of the tables above vary the serve engine as a matrix dimension, because
the engine is not selectable independent of hardware and OS. `rocm serve`
picks it via `effective_serve_engine()` in
`tests/e2e-cucumber/src/capability.rs` (mirroring the product's own
`preferred_serve_engine_for_host_gpu_summary`):

```rust
pub fn effective_serve_engine(gfx_target: Option<&str>, os_family: &str) -> String {
if os_family.eq_ignore_ascii_case("windows") {
return "lemonade".to_owned();
}
if family_prefers_vllm(gfx_target) {
"vllm".to_owned()
} else {
"lemonade".to_owned()
}
}
```

Any Windows host resolves to `lemonade` (the vLLM adapter does not run
there); otherwise `vllm` is only preferred for the `*-dcgpu` families and
`gfx906`/`gfx908`/`gfx90a`. Consequences for the lanes above:

- Strix Halo is `gfx1151` on every OS, so all three per-PR Strix lanes
(Ubuntu, Windows, WSL2) and their nightly counterparts are lemonade-only —
no matrix value turns a Strix lane into a vLLM lane.
- `e2e-gpu` (MI300X) is the only per-PR lane where vLLM is the effective
engine.
- Adding a vLLM lane means adding hardware from a vLLM-eligible family, not
adding an `engine:` value to a workflow matrix.

## Triggers

The GPU jobs (in `e2e-selfhosted.yml`) run automatically on `push`,
Expand Down Expand Up @@ -276,6 +345,13 @@ Pre-warm then:
- prunes with `rocm storage remove-old-installs` after any install, update, or
repair, so the multi-version cache stays bounded.

`e2e-prewarm` will also accept mutually exclusive `--version`/`--build-date`
flags (ROCMAI-430) that pin the SDK build the pre-warm resolves to instead of
always tracking whatever the channel index currently serves. That flag pair
ships in PR #464, stacked on #415, neither merged as of this writing (see
"The three-stage validation ladder" above) — the unpinned invocation above is
what every lane in this tree runs today.

The runtime is always installed **in place**: `install sdk` bakes absolute paths
into the runtime manifest, so a tree that is moved after installation leaves every
serve pointing at a path that no longer exists.
Expand Down
Loading