Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 11 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -284,17 +284,20 @@ fix` takes the id, not the position.

`fix` applies a known fix by the `id:` that `diagnose` reported — not the
ranking position noted above, which isn't a stable name. Run it with no id
to list the whole catalog. Each fix is marked AUTO (this command carries out
the change) or PRINT-ONLY (it prints the steps for you to run yourself —
usually because the right command depends on a choice only you can make,
sometimes because it also needs sudo or a reboot).
to list the whole catalog. Each fix carries a marker saying what happens on
this machine: AUTO (this command carries out the change), NEEDS-ARG (it will,
once given the argument it names), PRINT-ONLY (it prints the steps for you to
run yourself — usually because the right command depends on a choice only you
can make, sometimes because it also needs sudo or a reboot), or DIAGNOSE-ONLY
(no reliable fix exists, so nothing will be changed -- no catalog entry
carries this marker today; it is reserved for a future detect-but-cannot-repair
failure).

- `--dry-run` shows any fix's plan without changing anything.
- `--yes` skips the interactive confirmation once you've reviewed it.
- `--device-index` pins the discrete GPU index for `fix-9-igpu-dgpu`;
without it, that fix only prints the `rocminfo` (Linux) or `hipInfo.exe`
(Windows) query needed to find the index and makes no change, despite
being marked AUTO.
- `--device-index` pins the discrete GPU index for `fix-9-igpu-dgpu`, marked
NEEDS-ARG; without it, that fix only prints the `rocminfo` (Linux) or
`hipInfo.exe` (Windows) query needed to find the index and makes no change.

### ROCm installation

Expand Down
12 changes: 8 additions & 4 deletions apps/rocm/src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -160,10 +160,14 @@ enum Command {
/// `#1`/`#2` ranking position, which belongs to one report and is not a name.
/// With no id it lists the whole catalog.
///
/// Fixes are marked AUTO or PRINT-ONLY: AUTO means this command carries the
/// change out, PRINT-ONLY means it prints the steps for you to run yourself
/// (typically because they need sudo or a reboot). Use `--dry-run` to see any
/// fix's plan without changing anything.
/// Every fix carries a marker saying what happens on the machine in front of
/// you: AUTO means this command carries the change out; NEEDS-ARG means it
/// will, once told what to act on; PRINT-ONLY means it prints the steps for
/// you to run yourself (typically because they need sudo or a reboot); and
/// DIAGNOSE-ONLY means no reliable fix exists and nothing will be changed
/// (no catalog entry carries this marker today; it is reserved for a
/// future detect-but-cannot-repair failure).
/// Use `--dry-run` to see any fix's plan without changing anything.
Fix {
/// Fix id, e.g. fix-4-render-group. Omit to list available fixes.
fix_id: Option<String>,
Expand Down
224 changes: 136 additions & 88 deletions crates/rocm-core/src/diagnose.rs

Large diffs are not rendered by default.

900 changes: 681 additions & 219 deletions crates/rocm-core/src/fix.rs

Large diffs are not rendered by default.

5 changes: 3 additions & 2 deletions docs/wsl.md
Original file line number Diff line number Diff line change
Expand Up @@ -209,8 +209,9 @@ larger and adding the matching `/etc/fstab` line inside the distro.

Every WSL remedy is print-only. `rocm fix <id>` shows the commands and does not
run them: they either install packages with `sudo`, edit loader configuration, or
belong to the Windows host, and none of that meets the bar the four
auto-applicable fixes clear.
belong to the Windows host, and none of that meets the bar an auto-applied fix
has to clear. On WSL the CLI carries out exactly one catalog entry itself,
`fix-6-path`, which is not one of the WSL entries.

Two deliberate silences, so a report can be trusted:

Expand Down
86 changes: 48 additions & 38 deletions skills/rocm-doctor/reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,15 +49,25 @@ Apply a fix by id (run with no id to list). Exit codes:
| 4 | attempted but the command failed |
| 5 | user declined at the prompt |

Only four fixes are auto-applicable — `fix-2-unset-override`,
`fix-4-render-group`, `fix-6-path`, `fix-9-igpu-dgpu` — and the rest print their
plan for the user to run. Pass the **full** id (`rocm fix fix-2-unset-override`,
not `rocm fix fix-2`; a short id returns exit 2, unknown fix-id).

**Auto-fix** in the catalog below means only what the CLI reports in
`rocm fix`'s listing: the CLI has a runner for it and will carry it out itself,
rather than printing a plan for the user. It is not a promise that the runner
mutates anything.
Three ids are ever auto-applicable — never four, and `fix-9-igpu-dgpu` is not
one of them (see below) — and which of the three depends on the **host**,
because `rocm fix`'s listing marks each entry for the machine you run it on,
never for every machine the id applies to:

- **linux** — auto-applicable: `fix-4-render-group`, `fix-6-path`.
- **windows** — auto-applicable: `fix-2-unset-override`, `fix-6-path`.
- **wsl** — auto-applicable: `fix-6-path`.

The rest print their plan for the user to run. Pass the **full** id
(`rocm fix fix-2-unset-override`, not `rocm fix fix-2`; a short id returns
exit 2, unknown fix-id).

**Marker** in the catalog below is what `rocm fix`'s listing reports **on
bare-metal Linux** — AUTO means the CLI has a runner for the entry and will
carry it out itself there, rather than printing a plan for the user. It is not
a promise that the runner mutates anything, and — per the per-host list above — it is
not the same claim on Windows or WSL2 for the one entry whose behaviour
depends on more than the host alone.

The ones that do mutate print the exact command, honor `--dry-run`, refuse on a
non-interactive shell without `--yes`, and confirm first. **Two exceptions:**
Expand All @@ -67,10 +77,10 @@ non-interactive shell without `--yes`, and confirm first. **Two exceptions:**
not edit your dotfiles. So on Linux it never prompts, `--dry-run` has nothing
to preview, and `--yes` is never read. Do not tell a Linux user a dry run
previewed a change that was never going to happen.
- **`fix-9-igpu-dgpu` mutates only when you pass `--device-index N`.** Without
it — on Linux and Windows alike — the runner just prints the
`rocminfo`/`hipInfo` query that identifies which index is the discrete GPU
and returns 0: no prompt, nothing for `--dry-run` to preview, nothing
- **`fix-9-igpu-dgpu` needs `--device-index N` wherever it applies (Linux and
Windows) and is never marked AUTO.** Without the flag the runner just prints
the `rocminfo`/`hipInfo` query that identifies which index is the discrete
GPU and returns 0: no prompt, nothing for `--dry-run` to preview, nothing
pinned. Run `rocm fix fix-9-igpu-dgpu --device-index N` (not the bare id)
once you know N.

Expand All @@ -82,32 +92,32 @@ by naming `wsl`, so the bare-metal Linux entries (the `amdgpu` module,
`/dev/kfd`, the render group) stop applying there automatically rather than
reporting confident nonsense.

| id | OS | Failure mode | Typical signal | Auto-fix |
| id | OS | Failure mode | Typical signal | Marker (bare-metal Linux) |
| --- | --- | --- | --- | --- |
| `fix-1-arch` | linux/windows/wsl | GPU gfx target not in the framework's build arch list | `hipErrorNoBinaryForGpu`, `HSA_STATUS_ERROR_INVALID_ISA`, "invalid device function" | no |
| `fix-2-unset-override` | linux/windows/wsl | `HSA_OVERRIDE_GFX_VERSION` set on a GPU that now has a native wheel | page faults / `OUT_OF_REGISTERS`, override set in env | yes |
| `fix-3-rocm-kernel` | linux | ROCm + distro/kernel form an unsupported triple | ROCm installed but `amdgpu` not loaded; DKMS build failure | no |
| `fix-4-render-group` | linux | User not in `render`/`video` group (or `/dev/kfd` owned by the other group) | cannot open `/dev/kfd`, permission denied | yes |
| `fix-5-amdgpu-load` | linux | `amdgpu` module not loaded (or blacklisted) | "ROCk module is NOT loaded", blacklist entry, Secure Boot | no |
| `fix-6-path` | linux/windows/wsl | ROCm/HIP binaries not on PATH after install | `rocminfo: command not found`, `hipInfo` missing from PATH | yes |
| `fix-7-stale-repos` | linux | Stale/conflicting APT/DNF repos from prior installer runs | apt 404 `repo.radeon.com`, unmet deps, ≥2 ROCm repo files | no |
| `fix-8-wheel-rocm` | linux/windows/wsl | Framework wheel built for a different ROCm major than the system | `libamdhip64.so.X` / `amdhip64_X.dll` load failure | no |
| `fix-9-igpu-dgpu` | linux/windows | iGPU enumerated alongside dGPU, destabilising the runtime | APU + discrete AMD present, `HIP_VISIBLE_DEVICES` unset, crash/segfault | yes |
| `fix-10-container` | linux | Container can't see `/dev/kfd` or `/dev/dri/renderD*` | running in docker/podman, kfd/render devices missing | no |
| `fix-11-iommu` | linux | Multi-GPU hang with IOMMU enabled | ≥2 AMD GPUs, `iommu=` not `pt`, hang/deadlock/timeout | no |
| `fix-12-installer` | linux | `amdgpu-install` left a broken DKMS / repo state | dpkg half-configured, DKMS failed, `--accept-eula` | no |
| `fix-13-hip-sdk-missing` | windows | HIP SDK not installed | no HIP SDK under Program Files, `hipInfo` not recognized | no |
| `fix-14-adrenalin-too-old` | windows | Adrenalin / kernel-mode driver too old for the HIP SDK | `hipInfo` can't enumerate, "driver too old", HSA "no agents found" | no |
| `fix-15-msvc-redist` | windows | MSVC runtime missing (HIP DLLs can't load) | `vcruntime140.dll` / `vcruntime140_1.dll` missing | no |
| `fix-17-torch-dlpack` | linux | `torch-c-dlpack-ext` loads its CUDA prebuilt on a ROCm torch, aborting vLLM's engine start at import time | vLLM engine start fails on import; error names `torch_c_dlpack_ext` or tvm_ffi's `_optional_torch_c_dlpack` | no |
| `fix-19-shm-too-small` | linux/wsl | `/dev/shm` too small for a serving workload, which needs gigabytes where a container and WSL2 both default to 64 MB | reported under 1 GiB; a data-loader worker killed by a bus error, or a failed write to a temporary file, with nothing naming shared memory | no |
| `fix-wsl-1-gpu-not-exposed` | wsl | `/dev/dxg` absent, so the distro cannot reach the GPU at all | no `/dev/dxg`; in a container, the device was never passed through | no |
| `fix-wsl-2-dxcore-missing` | wsl | `/usr/lib/wsl/lib` DXCore shims missing, so the runtime cannot reach the host driver | `/usr/lib/wsl/lib/libdxcore.so` missing (or the directory absent entirely) | no |
| `fix-wsl-3-rocdxg-missing` | wsl | ROCDXG, the ROCm-to-DXCore shim the WSL path runs on, is not installed | `librocdxg.so` not found under any ROCm install | no |
| `fix-wsl-4-rocdxg-not-linked` | wsl | `librocdxg` installed but not in the linker cache, so it is unloadable | `librocdxg` present on disk yet absent from `ldconfig -p` | no |
| `fix-wsl-5-distro-too-old` | wsl | Distro release below the floor the WSL path requires | distro release under the supported floor (e.g. Ubuntu 22.04, whose glibc 2.35 is below the 2.38 / `GLIBCXX_3.4.32` the engines need) | no |
| `fix-wsl-6-host-driver-too-old` | wsl | Windows host driver too old or absent, with the distro side already complete | the Windows host reports no AMD display adapter | no |
| `fix-wsl-7-wsl1` | wsl | Distro running under WSL 1, which exposes no GPU device at all | the running kernel is a WSL 1 kernel | no |
| `fix-1-arch` | linux/windows/wsl | GPU gfx target not in the framework's build arch list | `hipErrorNoBinaryForGpu`, `HSA_STATUS_ERROR_INVALID_ISA`, "invalid device function" | print-only |
| `fix-2-unset-override` | linux/windows/wsl | `HSA_OVERRIDE_GFX_VERSION` set on a GPU that now has a native wheel | page faults / `OUT_OF_REGISTERS`, override set in env | print-only (auto on windows) |
| `fix-3-rocm-kernel` | linux | ROCm + distro/kernel form an unsupported triple | ROCm installed but `amdgpu` not loaded; DKMS build failure | print-only |
| `fix-4-render-group` | linux | User not in `render`/`video` group (or `/dev/kfd` owned by the other group) | cannot open `/dev/kfd`, permission denied | auto |
| `fix-5-amdgpu-load` | linux | `amdgpu` module not loaded (or blacklisted) | "ROCk module is NOT loaded", blacklist entry, Secure Boot | print-only |
| `fix-6-path` | linux/windows/wsl | ROCm/HIP binaries not on PATH after install | `rocminfo: command not found`, `hipInfo` missing from PATH | auto |
| `fix-7-stale-repos` | linux | Stale/conflicting APT/DNF repos from prior installer runs | apt 404 `repo.radeon.com`, unmet deps, ≥2 ROCm repo files | print-only |
| `fix-8-wheel-rocm` | linux/windows/wsl | Framework wheel built for a different ROCm major than the system | `libamdhip64.so.X` / `amdhip64_X.dll` load failure | print-only |
| `fix-9-igpu-dgpu` | linux/windows | iGPU enumerated alongside dGPU, destabilising the runtime | APU + discrete AMD present, `HIP_VISIBLE_DEVICES` unset, crash/segfault | needs-arg |
| `fix-10-container` | linux | Container can't see `/dev/kfd` or `/dev/dri/renderD*` | running in docker/podman, kfd/render devices missing | print-only |
| `fix-11-iommu` | linux | Multi-GPU hang with IOMMU enabled | ≥2 AMD GPUs, `iommu=` not `pt`, hang/deadlock/timeout | print-only |
| `fix-12-installer` | linux | `amdgpu-install` left a broken DKMS / repo state | dpkg half-configured, DKMS failed, `--accept-eula` | print-only |
| `fix-13-hip-sdk-missing` | windows | HIP SDK not installed | no HIP SDK under Program Files, `hipInfo` not recognized | print-only |
| `fix-14-adrenalin-too-old` | windows | Adrenalin / kernel-mode driver too old for the HIP SDK | `hipInfo` can't enumerate, "driver too old", HSA "no agents found" | print-only |
| `fix-15-msvc-redist` | windows | MSVC runtime missing (HIP DLLs can't load) | `vcruntime140.dll` / `vcruntime140_1.dll` missing | print-only |
| `fix-17-torch-dlpack` | linux | `torch-c-dlpack-ext` loads its CUDA prebuilt on a ROCm torch, aborting vLLM's engine start at import time | vLLM engine start fails on import; error names `torch_c_dlpack_ext` or tvm_ffi's `_optional_torch_c_dlpack` | print-only |
| `fix-19-shm-too-small` | linux/wsl | `/dev/shm` too small for a serving workload, which needs gigabytes where a container and WSL2 both default to 64 MB | reported under 1 GiB; a data-loader worker killed by a bus error, or a failed write to a temporary file, with nothing naming shared memory | print-only |
| `fix-wsl-1-gpu-not-exposed` | wsl | `/dev/dxg` absent, so the distro cannot reach the GPU at all | no `/dev/dxg`; in a container, the device was never passed through | print-only |
| `fix-wsl-2-dxcore-missing` | wsl | `/usr/lib/wsl/lib` DXCore shims missing, so the runtime cannot reach the host driver | `/usr/lib/wsl/lib/libdxcore.so` missing (or the directory absent entirely) | print-only |
| `fix-wsl-3-rocdxg-missing` | wsl | ROCDXG, the ROCm-to-DXCore shim the WSL path runs on, is not installed | `librocdxg.so` not found under any ROCm install | print-only |
| `fix-wsl-4-rocdxg-not-linked` | wsl | `librocdxg` installed but not in the linker cache, so it is unloadable | `librocdxg` present on disk yet absent from `ldconfig -p` | print-only |
| `fix-wsl-5-distro-too-old` | wsl | Distro release below the floor the WSL path requires | distro release under the supported floor (e.g. Ubuntu 22.04, whose glibc 2.35 is below the 2.38 / `GLIBCXX_3.4.32` the engines need) | print-only |
| `fix-wsl-6-host-driver-too-old` | wsl | Windows host driver too old or absent, with the distro side already complete | the Windows host reports no AMD display adapter | print-only |
| `fix-wsl-7-wsl1` | wsl | Distro running under WSL 1, which exposes no GPU device at all | the running kernel is a WSL 1 kernel | print-only |

Two things the numbering does not tell you. `fix-16` is a reserved handle, not a
missing row — ids are stable handles rather than positions. And the `fix-wsl-N`
Expand Down
54 changes: 45 additions & 9 deletions tests/e2e-cucumber/features/diagnose.feature
Original file line number Diff line number Diff line change
Expand Up @@ -276,16 +276,52 @@ Feature: Diagnosing failures and listing fixes
# diagnose-05 only proves the manual/zero-optional-flags wording, because
# PREVIEW_FIX_ID (fix-1-arch) needs none of sudo/reboot/re-login. The
# sudo+re-login combination only exists on a fix gated to bare-metal Linux
# (fix-4-render-group), so it needs its own scenario -- but, like
# diagnose-14, it is deliberately not OS-gated: `print_recipe` runs before
# the fix's own platform gate (see `apply` in fix.rs), so the Flags: text
# under test renders identically regardless of which lane runs it. The step
# asserts only that printed text, never the exit code -- `fix-4-render-group`
# gates its own dry-run on host state ($USER, `usermod`/`sudo` on PATH), so
# unlike PREVIEW_FIX_ID its exit code is not guaranteed to be 0 everywhere.
@id:diagnose-fix-preview-states-required-flags
Scenario: diagnose-20 - Previewing a fix that needs sudo and a re-login says so, and that it's auto-applicable
# (fix-4-render-group), so it needs its own scenario.
#
# Unlike diagnose-14, this one IS OS-gated. `print_recipe` still runs before
# the fix's own platform gate (see `apply` in fix.rs), but the Flags: line it
# prints comes from `class_here()`, which looks up the catalog entry for the
# *running* host's OS. "AUTO" only renders where `fix-4-render-group` is
# actually `Auto` -- bare-metal Linux. Everywhere else (`applies_on` has no
# other member) `class_here()` falls back to PRINT-ONLY, the generic "this
# fix does not apply here" answer diagnose-11 already covers -- not a second,
# platform-specific behaviour worth asserting under this scenario's name.
#
# The step still asserts only the printed Flags: text, never the exit code --
# `fix-4-render-group` gates its own dry-run on host state ($USER,
# `usermod`/`sudo` on PATH), so unlike PREVIEW_FIX_ID its exit code is not
# guaranteed to be 0 even on Linux.
@id:diagnose-fix-preview-states-required-flags @requires-os:linux @requires-bare-metal

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two things here.

Adding @requires-bare-metal narrows an existing scenario's coverage, and it is not mentioned in the PR description. The scenario title still ends "and that it's auto-applicable", which now reads oddly against the change.

The description's e2e claim also does not hold up. It says "the 3 failures are the @requires-bare-metal scenarios that cannot pass on a WSL2 dev host, unchanged from before this PR". But @requires-bare-metal on WSL resolves to Expectation::Skip, never a failure (tests/e2e-cucumber/src/expectation.rs:553-555), and bare_metal_skip_is_not_recorded_as_a_known_bug pins that it is not an xfail row. So those scenarios cannot be in a failure count, and whatever the 3 results actually were is still unexplained. "Unchanged from before this PR" is also false, since this PR adds the tag here and introduces diagnose-22 carrying it.

AGENTS.md:97-100 asks for gated scenarios to be named by @id: plus the lane that will exercise them. The description names neither @id:diagnose-fix-applicability-is-per-machine nor @id:diagnose-fix-needing-an-argument-says-so, and names no lane. The lanes do exist (docs/ci-hardware-testing.md:35-40), so this is a description gap, not dead coverage.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both fixed.

The scope gap: retitled diagnose-20 to "...and that it's auto-applicable here" and will call out the @requires-bare-metal addition (on diagnose-20) plus its introduction on diagnose-22 explicitly in the PR description, with the @id:s and the lane per AGENTS.md -- it was a real gap, not mentioned before.

The e2e claim: re-ran properly this time. cargo xtask e2e -- -n diagnose does bypass tag-based skip resolution exactly as you'd expect from the commit message's own warning about -n -- I reproduced that directly: with -n diagnose, diagnose-20 and diagnose-22 run for real on this WSL2 host and fail (exit 3, "Running OS is: wsl"), not skip. Using -i diagnose.feature instead (which doesn't bypass filter_run's skip resolution) gives the real picture: 22 scenarios total, 16 run, 6 skip, 0 fail -- and platform.json confirms all 6 skips are @requires-bare-metal with reason "requires a bare-metal host; this one is WSL2", diagnose-20 and diagnose-22 among them. So your read is right on both counts: the 3-failures claim doesn't hold up (they're skips, and there were never only 3), and "unchanged from before this PR" is false since this PR is what added the tag to diagnose-20 and introduced diagnose-22 carrying it. Updating the PR description with this corrected evidence and the @id:/lane callout.

Scenario: diagnose-20 - Previewing a fix that needs sudo and a re-login says so, and that it's auto-applicable here
Given a user who has chosen a fix that needs sudo and a re-login
When the user previews that fix without applying it
Then the preview states that the fix requires sudo and a re-login
And the preview states that the CLI can run it automatically

# One entry behaves differently depending on the machine: it persists the
# change on Windows, and on Linux it only reports where the value is set,
# because the code that would write it takes no options and never does.
# The listing said "the CLI will run this" on both, so a user on Linux — and
# an agent reading the same listing — was told a change was coming that never
# came. Host-independent on purpose: the assertion is that the listing agrees
# with the machine in front of it, whichever machine that is.
@id:diagnose-fix-applicability-is-per-machine
Scenario: diagnose-21 - A fix that only explains itself here is not advertised as one the CLI will run
Given a fix the CLI carries out on one kind of machine and only explains on another
When the user asks the CLI which fixes it offers
Then that fix is shown as what it does on this machine

# The other half of the same defect. This entry does have a fix and the CLI
# will carry it out, but not until it is told which device to pin; asked
# plainly it prints the query that identifies one and stops. It was marked as
# a fix the CLI applies, so the report of a change that never happened looked
# like success.
# @requires-bare-metal because the entry under test is scoped to bare-metal
# Linux and Windows. On WSL it is refused at the platform gate instead, which
# is a different contract with its own scenario — and the right one, since the
# catalog does not claim this remedy applies there.
@id:diagnose-fix-needing-an-argument-says-so @requires-bare-metal
Scenario: diagnose-22 - A fix that needs more information says what it needs and changes nothing
Given a user who has chosen a fix that cannot run until it is told what to act on
When the user asks the CLI to apply it without saying what to act on
Then the CLI names what it still needs and reports no change
Loading
Loading