How to verify whisper-api is actually running inference on the AMD GPU (ROCm) — and monitor it live.
uv run python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"Expected output:
True AMD Radeon RX 7900 GRE
torch.cuda.* is the correct API even on ROCm — PyTorch maps HIP to the
"cuda" device namespace.
Run these in a second terminal while requests are in flight.
watch -n 0.5 rocm-smi # util%, VRAM, temp, power
watch -n 0.5 amd-smi # newer AMD tool
rocm-smi --showuse --showmeminfo vram # one-shot snapshotrocm-smi --showpids
# or
amd-smi processLook for the python manage.py runserver process holding VRAM.
cat /sys/class/drm/card1/device/gpu_busy_percent # instant GPU util %
cat /sys/class/drm/card1/device/mem_info_vram_used # VRAM used, bytes(If you have multiple GPUs, the card number may differ — check
ls /sys/class/drm/.)
import torch
torch.cuda.is_available() # True
torch.cuda.get_device_name(0) # 'AMD Radeon RX 7900 GRE'
torch.cuda.memory_allocated() / 2**30 # GB currently used by tensorsFire a transcription in the background and sample utilization while it runs:
curl -s http://localhost:8000/v1/audio/transcriptions \
-F file=@/tmp/jfk.flac -F model=base -o /dev/null &
for i in 1 2 3 4 5 6; do
busy=$(cat /sys/class/drm/card1/device/gpu_busy_percent)
vram=$(cat /sys/class/drm/card1/device/mem_info_vram_used)
echo "sample $i: GPU busy=${busy}% VRAM used=$((vram / 1024 / 1024)) MiB"
sleep 0.3
done
wait| Signal | On GPU | CPU fallback |
|---|---|---|
| VRAM after model load | elevated, constant (~1–8 GB by model) | ~0 |
gpu_busy_percent during request |
spikes (short clips) / sustained (long audio) | 0% |
| 10 s audio transcription time | ~1–2 s (turbo) | 10–30 s, fans/CPU high |
rocm-smi --showpids |
python process listed | not listed |
Short clips with small models (
base) only produce brief utilization spikes — that's normal, not a sign of CPU fallback. Use longer audio orturbo/largefor sustained load.
If you see 0 VRAM and 0% GPU during requests:
- Confirm
torch.cuda.is_available()isTruein the server process (same venv, same env vars). - Check the server logs at startup for CUDA/HIP errors.
- Make sure no stray
CUDA_VISIBLE_DEVICES=""orHIP_VISIBLE_DEVICES=""is hiding the GPU. - Verify your user is in the
renderandvideogroups (groups | grep -E "render|video"), then re-login if you just added them.