diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index caa0e4e88..72326fd13 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -169,13 +169,14 @@ jobs: - 'ruff.toml' - 'PSScriptAnalyzerSettings.psd1' - '.github/workflows/**' - # Sphinx docs site: the doc sources themselves, plus README.md and - # CONTRIBUTING.md, which several pages single-source via MyST - # `{include}` directives, plus the build config/toolchain pins. + # Sphinx docs site: the doc sources themselves, plus README.md, + # CONTRIBUTING.md and docs/vllm.md, which pages single-source via + # MyST `{include}` directives, plus the build config/toolchain pins. docs: - 'docs/rocm-docs/**' - 'README.md' - 'CONTRIBUTING.md' + - 'docs/vllm.md' - '.readthedocs.yaml' - '.github/workflows/**' diff --git a/README.md b/README.md index 8249dc1b5..f7864b4fa 100644 --- a/README.md +++ b/README.md @@ -216,7 +216,7 @@ rocm serve qwen ``` `qwen` is a built-in alias for a small assistant model that serves out of the -box. You can also serve any compatible Hugging Face model directly — see +box. You can also serve any compatible Hugging Face model directly. See [Model serving](#model-serving) for the GGUF-vs-safetensors rule, since which form works depends on the engine your GPU selects. @@ -311,54 +311,84 @@ rocm update [--apply] [--runtime KEY] [--activate] [--dry-run] ``` `install sdk` downloads TheRock ROCm wheels into a Python environment managed -by rocm-cli; pass `--devel` to also install the compiler and headers needed to -build GPU code, roughly doubling the download. `--devel` is not an addition to -an existing runtime: a runtime is identified by the packages it was installed -from, so running `rocm install sdk` and later `rocm install sdk --devel` at the -same version leaves you with **two** side-by-side runtimes — the second is a -fresh full install, not a toolchain bolted onto the first — and the second one -becomes active. `rocm runtimes list` marks each one `toolchain=included` or -`toolchain=excluded`, `rocm examine` reports the active runtime's as -`active_runtime_toolchain`, and `rocm storage remove-old-installs` counts the -two kinds separately so neither evicts the other. To reclaim the space, uninstall -the one you do not want with `rocm runtimes uninstall `. -An install with no active default runtime never prompts, but once a -managed runtime is the active default every `install sdk` asks first, because -the new install takes over as the active default. That gate is not scoped to the -family or channel you are installing: a `--family` or `--channel` you have never -installed before takes over the active default just as a same-family upgrade -does, so it asks too. To approve that non-interactively — in scripts or CI, where -the prompt would otherwise refuse — pass `--approve-replacing-active-default`, -which is also what the refusal itself recommends and what ROCm CLI's own -non-interactive surfaces (chat, MCP, the dashboard) pass. `--yes` grants the same -approval *and* approves installing required system packages (such as OpenMPI for -vLLM), which means `sudo`; reach for it only where something can answer a sudo -password prompt — which an unattended job cannot, unless it has passwordless sudo -configured. In the default managed install root, the root and its manifest are -keyed by version, so an upgrade or downgrade keeps the previous install on disk -and only a same-version reinstall reuses the same root. `--prefix` opts out of -that: the folder you name is used verbatim for every version, so successive -installs into one prefix replace each other in place — and if the venv already -there no longer runs its own Python, it is removed outright and rebuilt. The -consent gate does not cover that: it asks about changing the active default -runtime, not about what a named prefix loses. `install driver` installs the AMD -kernel driver on Linux (DKMS or native package). `update` checks for a newer -ROCm package; pass `--apply` to install it, or `--dry-run` to preview what -`--apply` would do without changing anything (`--dry-run` does not require -`--apply`). `--runtime` and `--activate` require `--apply` or `--dry-run` — pass -one of those instead of naming a runtime or requesting activation on its own. -`--json` prints the check result as a single line of JSON instead of text; -`--timeout-secs` bounds its network calls (`--timeout-secs` requires `--json`; -both `--json` and `--timeout-secs` conflict with `--apply`, and `--json` also -conflicts with `--dry-run`). `update --apply` never prompts and needs no -approval flag: selecting a runtime to update is itself the approval, and it -leaves the active default alone unless you add `--activate`. `update` does -accept `--yes`, for consistency with other mutating commands, but it grants -nothing there — the approval line the update path prints never credits it. - -ROCm 10 and newer ship from a different source layout. It is opt-in, and asking -for it takes two things together: pin the version with `--version`, and name the -exact GPU arch — the raw `gfx` code, not a family label: +by rocm-cli. + +#### Compiler toolchain (--devel) + +Pass `--devel` to also install the compiler and headers needed to build GPU +code. This roughly doubles the download. + +`--devel` isn't an addition to an existing runtime. A runtime is identified by +the packages it was installed from, so running `rocm install sdk` and later +`rocm install sdk --devel` at the same version leaves you with **two** +side-by-side runtimes. The second is a fresh full install, not a toolchain +bolted onto the first, and it becomes active. + +To tell the two apart: + +- `rocm runtimes list` marks each runtime `toolchain=included` or + `toolchain=excluded`. +- `rocm examine` reports the active runtime's toolchain as + `active_runtime_toolchain`. +- `rocm storage remove-old-installs` counts the two kinds separately, so + neither evicts the other. + +To reclaim the space, uninstall the one you don't want with +`rocm runtimes uninstall `. + +#### Approval prompt + +If no managed runtime is the active default, `install sdk` doesn't prompt. +Otherwise it asks first, because the new install becomes the active default. +The prompt applies to any install, including a `--family` or `--channel` you +haven't installed before, just as it does for a same-family upgrade. + +To approve without a prompt, for example in scripts or CI, where the prompt +would otherwise refuse: + +- `--approve-replacing-active-default` approves the change of active default. + The refusal message recommends it, and ROCm CLI's own non-interactive + surfaces (chat, MCP, and the dashboard) pass it. +- `--yes` gives the same approval and also approves installing required system + packages, such as OpenMPI for vLLM. That requires `sudo`, so use it only where + something can answer a sudo password prompt. An unattended job can't, unless + it has passwordless sudo configured. + +#### Install location + +In the default managed install root, the root and its manifest are keyed by +version. An upgrade or downgrade keeps the previous install on disk. Only a +same-version reinstall reuses the same root. + +`--prefix` changes this. The folder you name is used as-is for every version, so +successive installs into one prefix replace each other in place. If the venv +already there no longer runs its own Python, it is removed outright and rebuilt. +The approval prompt doesn't cover this, because it asks only about changing the +active default runtime, not about what a named prefix loses. + +#### Driver installation + +`install driver` installs the AMD kernel driver on Linux, using DKMS or a native +package. + +#### Updates + +`update` checks for a newer ROCm package. + +| Flag | Description | +| --- | --- | +| `--apply` | Installs the update. Never prompts and needs no approval flag, because selecting a runtime to update is the approval. Leaves the active default alone unless you add `--activate`. | +| `--dry-run` | Previews what `--apply` would do without changing anything. Doesn't require `--apply`. | +| `--runtime`, `--activate` | Require `--apply` or `--dry-run`. | +| `--json` | Prints the check result as a single line of JSON instead of text. Conflicts with `--apply` and `--dry-run`. | +| `--timeout-secs` | Bounds the network calls of the check. Requires `--json`. Conflicts with `--apply`. | +| `--yes` | Accepted for consistency with other mutating commands, but grants nothing on `update`. The approval line the update path prints never credits it. | + +#### ROCm 10 and newer + +ROCm 10 and newer ship from a different source layout. You opt in by passing two +things together: pin the version with `--version`, and name the exact GPU arch, +using the raw `gfx` code rather than a family label: ``` rocm install sdk --version 10.0.0 --family gfx1200 --dry-run @@ -375,9 +405,9 @@ selected framework package carries the same ROCm build identifier before it creates or changes a managed runtime. Nothing about this happens on its own. Without a `--version` of 10 or newer, -`install sdk` resolves the same release and nightly sources it always has, and -it never quietly retries against the ROCm 10 sources when a lookup comes up -empty — it tells you what it could not find instead. +`install sdk` resolves the same release and nightly sources as before. It doesn't +quietly retry against the ROCm 10 sources when a lookup finds nothing; it tells +you what it couldn't find instead. ### Runtime management diff --git a/docs/rocm-docs/commands.md b/docs/rocm-docs/commands.md index 02e368053..e9d5f8e92 100644 --- a/docs/rocm-docs/commands.md +++ b/docs/rocm-docs/commands.md @@ -6,6 +6,10 @@ SPDX-License-Identifier: MIT # Command reference +This page describes each `rocm` command, its options, and what it does. For a +short list of the commands and what each is for, see +[Getting started](getting-started.md). + ```{include} ../../README.md :start-after: "## Commands" :end-before: "and a chat tab backed by any configured provider." diff --git a/docs/rocm-docs/engines/vllm.md b/docs/rocm-docs/engines/vllm.md new file mode 100644 index 000000000..732a57e14 --- /dev/null +++ b/docs/rocm-docs/engines/vllm.md @@ -0,0 +1,8 @@ + + +```{include} ../../vllm.md +``` diff --git a/docs/rocm-docs/getting-started.md b/docs/rocm-docs/getting-started.md index 1f0de0be8..6a25b6efd 100644 --- a/docs/rocm-docs/getting-started.md +++ b/docs/rocm-docs/getting-started.md @@ -15,10 +15,24 @@ SPDX-License-Identifier: MIT ```{include} ../../README.md :start-after: "## Configure ROCm and serve a model" +:end-before: "Running the command when a" +``` + + +Running the command when a managed runtime is already the active default asks +first, because the new install takes over as the active default; see +[ROCm installation](commands.md#rocm-installation) for that gate and the flags +that approve it without a prompt. + +```{include} ../../README.md +:start-after: "for that gate and the flags that approve it without a prompt." :end-before: "You can also serve any compatible Hugging Face model directly" ``` -You can also serve any compatible Hugging Face model directly — see +You can also serve any compatible Hugging Face model directly. See [Model serving](commands.md#model-serving) for the GGUF-vs-safetensors rule, since which form works depends on the engine your GPU selects. diff --git a/docs/rocm-docs/index.rst b/docs/rocm-docs/index.rst index f5143d006..f6ec6fbd2 100644 --- a/docs/rocm-docs/index.rst +++ b/docs/rocm-docs/index.rst @@ -17,7 +17,7 @@ adapters for Lemonade and vLLM. .. important:: **Tech Preview:** This software is provided as-is, without warranty or - guarantee of stability. APIs, commands, and behavior may change without + guarantee of stability. APIs, commands, and behavior might change without notice. Intended for experimentation and early feedback only. The ROCm CLI public repository is located at @@ -26,10 +26,6 @@ The ROCm CLI public repository is located at .. grid:: 2 :gutter: 3 - .. grid-item-card:: Demos - - * :doc:`See ROCm CLI in action ` - .. grid-item-card:: Install * :doc:`Installing ROCm CLI ` @@ -37,7 +33,9 @@ The ROCm CLI public repository is located at .. grid-item-card:: Getting started * :doc:`Getting started with ROCm CLI ` + * :doc:`See ROCm CLI in action ` - .. grid-item-card:: Commands + .. grid-item-card:: Use ROCm CLI * :doc:`Command reference ` + * :doc:`vLLM adapter ` diff --git a/docs/rocm-docs/install/installation.md b/docs/rocm-docs/install/installation.md index 5f1cb3dc8..53cabf3e1 100644 --- a/docs/rocm-docs/install/installation.md +++ b/docs/rocm-docs/install/installation.md @@ -15,7 +15,13 @@ ROCm CLI ships as a single prebuilt binary. Platform support: Live dashboard telemetry requires Linux or WSL2 (see [Interactive interfaces](../getting-started.md#interactive-interfaces)). vLLM -serving is Linux or WSL2 only (see `docs/vllm.md`). +serving is Linux or WSL2 only (see +[vLLM adapter](../engines/vllm.md)). + +```{include} ../../../README.md +:start-after: "only (see [docs/vllm.md](docs/vllm.md))." +:end-before: "> [!IMPORTANT]" +``` ```{include} ../../../README.md :start-after: "## Installation" diff --git a/docs/rocm-docs/sphinx/_toc.yml.in b/docs/rocm-docs/sphinx/_toc.yml.in index 7534a1eaa..3494fbd4b 100644 --- a/docs/rocm-docs/sphinx/_toc.yml.in +++ b/docs/rocm-docs/sphinx/_toc.yml.in @@ -1,18 +1,19 @@ root: index subtrees: - - caption: Demos - entries: - - file: demos - title: See ROCm CLI in action - caption: Install entries: - file: install/installation - caption: Getting started entries: - file: getting-started - - caption: Commands + - file: demos + title: See ROCm CLI in action + - caption: Use ROCm CLI entries: - file: commands + title: Command reference + - file: engines/vllm + title: vLLM adapter - caption: About entries: - file: about/contributing diff --git a/docs/vllm.md b/docs/vllm.md index 75c32059b..21acc6f17 100644 --- a/docs/vllm.md +++ b/docs/vllm.md @@ -4,19 +4,22 @@ Copyright © Advanced Micro Devices, Inc., or its affiliates. SPDX-License-Identifier: MIT --> -# vLLM Adapter +# vLLM adapter `rocm-engine-vllm` is a first-party adapter around an existing vLLM installation. It is intended for Linux and WSL ROCm GPU serving. The adapter does not install vLLM automatically and does not run CPU mode. Install or build vLLM in a ROCm-capable Python environment first, then make the -`vllm` command visible to rocm-cli. +`vllm` command visible to ROCm CLI. -For rocm-cli managed TheRock runtimes, prefer building vLLM from source against +Native Windows vLLM serving is skipped in this adapter. Use WSL or Linux for vLLM +ROCm serving, or choose a different engine explicitly. No CPU fallback is used. + +For ROCm CLI-managed TheRock runtimes, prefer building vLLM from source against the existing TheRock PyTorch stack. A prebuilt vLLM ROCm wheel can replace the TheRock torch packages or target a different ROCm soname set; that is not a -valid no-fallback setup for rocm-cli GPU serving. +valid no-fallback setup for ROCm CLI GPU serving. Building from source compiles HIP sources, so it needs the ROCm compiler toolchain. That is opt-in: install the runtime with `rocm install sdk --devel`, @@ -26,9 +29,9 @@ can still *serve* an already-built vLLM. ## Torch alignment on engine install Installing an engine into a managed TheRock runtime can change the torch in that -runtime. Two installers write torch into the same environment — the SDK install +runtime. Two installers write torch into the same environment: the SDK install writes TheRock's build, and the engine install then writes the build from its own -index — so `rocm engines install` settles which one stays and prints the result +index. So `rocm engines install` settles which one stays and prints the result as a `torch_alignment:` line. A torch that already executes a GPU kernel against the installed SDK is kept @@ -46,19 +49,19 @@ skip the replacement: ROCM_CLI_DISABLE_TORCH_ALIGNMENT=1 rocm engines install vllm --yes ``` -Any value works, including an empty one — the variable being set is the signal. +Any value works, including an empty one: the variable being set is the signal. The install then reports `torch_alignment: disabled`, naming both the build it would have installed and the one it kept. The device check still runs, so an opt-out that leaves the runtime unable to serve says so rather than failing later during serving. -Use it when you are deliberately running a torch the alignment would replace — a -locally built wheel, a version under test, a stack pinned for a reproduction. It +Use it when you are deliberately running a torch the alignment would replace (a +locally built wheel, a version under test, or a stack pinned for a reproduction). It is an escape hatch, not a supported configuration: the resulting combination is not validated against the supported matrix, and a runtime that cannot execute a kernel will fail at serving time. -### ROCm 10.x wheel discovery +## ROCm 10.x wheel discovery For most ROCm SDK versions, `rocm engines install vllm` pins a fixed vLLM wheel and index URL. Any ROCm SDK 10.x version is different: AMD publishes vLLM, @@ -81,11 +84,13 @@ likely install an ABI-incompatible build, so the install fails closed instead with a message naming the detected version and pointing at `ROCM_CLI_VLLM_ROCM_INDEX_URL` as the way to install anyway. +## Discovery paths and checks + Supported discovery paths: -- `ROCM_CLI_VLLM_COMMAND=/path/to/vllm` -- `ROCM_CLI_VLLM_PYTHON=/path/to/python` where a sibling `vllm` command exists -- the active rocm-cli managed TheRock runtime, if vLLM has been installed into +- `ROCM_CLI_VLLM_COMMAND=path_to_vllm`, where `path_to_vllm` is the absolute path to the `vllm` executable +- `ROCM_CLI_VLLM_PYTHON=path_to_python`, where `path_to_python` is the absolute path to a Python interpreter that has a sibling `vllm` command +- the active ROCm CLI-managed TheRock runtime, if vLLM has been installed into that Python environment - `vllm` on `PATH` @@ -98,7 +103,9 @@ rocm-engine-vllm resolve-model Qwen/Qwen3.5-4B --device-policy gpu_required python scripts/vllm_therock_gpu_test.py --self-test ``` -GPU acceptance check: +## GPU acceptance check + +Run the acceptance script against a built adapter: ```bash python3 scripts/vllm_therock_gpu_test.py \ @@ -106,14 +113,16 @@ python3 scripts/vllm_therock_gpu_test.py \ --model facebook/opt-125m ``` -The acceptance script is Linux/WSL only. It requires vLLM to be discoverable -through a rocm-cli managed TheRock runtime manifest, launches with +The acceptance script is Linux or WSL only. It requires vLLM to be discoverable +through a ROCm CLI-managed TheRock runtime manifest, launches with `gpu_required`, checks `/health` and `/v1/completions`, and verifies loaded ROCm libraries come from the managed TheRock SDK wheel directories. It rejects external vLLM command overrides and does not allow CPU fallback. It defaults to the active exact runtime key; if `--runtime-id` is passed, use an exact runtime key or an unambiguous runtime id. +### Source build notes + On WSL, the tested source build needed vLLM ROCm platform detection to use TheRock PyTorch device data when `amdsmi` is unavailable, and needed vLLM's ROCm GPTQ half-atomic compatibility path enabled for TheRock 7.13 headers. @@ -133,13 +142,15 @@ kernel. With the patch, the live acceptance harness passed on `facebook/opt-125m` and verified HIP/BLAS libraries loaded from the managed TheRock SDK wheel directories. -Serving through rocm-cli: +## Serve a model + +Serve a model through ROCm CLI: ```bash rocm serve Qwen/Qwen3.5-4B --engine vllm --device gpu_required --managed ``` -### GPU selection +## GPU selection Use `--gpu` to choose the AMD GPU vLLM runs on: @@ -151,17 +162,17 @@ rocm serve Qwen/Qwen3.5-4B --engine vllm --managed rocm serve Qwen/Qwen3.5-4B --engine vllm --gpu 1 --managed ``` -rocm-cli pins the device via `HIP_VISIBLE_DEVICES`. Serving one model across +ROCm CLI pins the device via `HIP_VISIBLE_DEVICES`. Serving one model across multiple GPUs is not supported. -### GPU memory +## GPU memory -vLLM claims a fixed fraction of each GPU's **total** VRAM — not of the free -VRAM, and not scaled to the model — for weights plus KV cache. On a large card +vLLM claims a fixed fraction of each GPU's **total** VRAM (not of the free +VRAM, and not scaled to the model) for weights plus KV cache. On a large card a small model therefore still reserves a large slice. -rocm-cli sets no `--gpu-memory-utilization` of its own, so vLLM's own default -applies unless a value comes from somewhere else — either a model's catalog +ROCm CLI sets no `--gpu-memory-utilization` of its own, so vLLM's own default +applies unless a value comes from somewhere else: either a model's catalog recipe or, taking precedence over it, the flag below: ```bash @@ -170,7 +181,7 @@ rocm serve --engine vllm --gpu-memory-utilization 0.3 --managed The value is a fraction in `(0, 1]` of total device VRAM. Lower it to leave room for a display, another workload, or a second server; raise it to give a large -model more KV cache. Applies to vLLM only — it is ignored, with a note in the +model more KV cache. Applies to vLLM only; it is ignored, with a note in the serve output, for other engines. An out-of-range or unparsable value fails the command rather than falling back silently. @@ -178,16 +189,16 @@ Earlier releases pinned this to `0.80` to leave display/WSL headroom. That pin i gone, so an unchanged command now reserves vLLM's own (higher) default. Pass `--gpu-memory-utilization 0.8` to restore the previous reservation. -### Tool calling +## Tool calling The TUI chat tab attaches tool definitions to every chat request. vLLM rejects those with HTTP 400 unless it is launched with `--enable-auto-tool-choice` **and** a matching `--tool-call-parser`. vLLM does not auto-detect the parser and it is -model-specific, so rocm-cli never guesses one: +model-specific, so ROCm CLI never guesses one: - **Built-in catalog models** carry the correct parser in their recipe metadata, - so tool calling works out of the box (e.g. Qwen family → `hermes`, - Llama 3 → `llama3_json`). + so tool calling works out of the box (for example, Qwen family → `hermes`, + Llama 3 → `llama3_json`). - **Other models** (arbitrary Hugging Face repos, or a catalog model forced onto vLLM without authored metadata) need an explicit parser: @@ -199,10 +210,7 @@ model-specific, so rocm-cli never guesses one: default, and applies to vLLM only. Common values: `hermes`, `llama3_json`, `mistral`. Without it, plain chat still works but tool calls return HTTP 400. -Native Windows vLLM serving is skipped in this adapter. Use WSL/Linux for vLLM -ROCm serving, or choose a different engine explicitly. No CPU fallback is used. - -References: +## Related resources -- vLLM ROCm installation: https://docs.vllm.ai/en/stable/getting_started/installation/gpu/ -- AMD ROCm vLLM guidance: https://rocmdocs.amd.com/en/latest/how-to/rocm-for-ai/inference/deploy-your-model.html +- [vLLM ROCm installation](https://docs.vllm.ai/en/stable/getting_started/installation/gpu/) +- [AMD ROCm AI ecosystem: vLLM](https://rocm.docs.amd.com/projects/ai-ecosystem/en/latest/inference/vllm.html)