Skip to content

Latest commit

 

History

History
102 lines (74 loc) · 3.03 KB

File metadata and controls

102 lines (74 loc) · 3.03 KB

GPU Monitoring & Verification

How to verify whisper-api is actually running inference on the AMD GPU (ROCm) — and monitor it live.

1. Quick check: does PyTorch see the GPU?

uv run python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"

Expected output:

True AMD Radeon RX 7900 GRE

torch.cuda.* is the correct API even on ROCm — PyTorch maps HIP to the "cuda" device namespace.

2. Live monitoring during transcription

Run these in a second terminal while requests are in flight.

ROCm tools (recommended)

watch -n 0.5 rocm-smi                     # util%, VRAM, temp, power
watch -n 0.5 amd-smi                      # newer AMD tool
rocm-smi --showuse --showmeminfo vram     # one-shot snapshot

Per-process view (which PID owns the VRAM)

rocm-smi --showpids
# or
amd-smi process

Look for the python manage.py runserver process holding VRAM.

Raw kernel counters (no ROCm tools needed, script-friendly)

cat /sys/class/drm/card1/device/gpu_busy_percent     # instant GPU util %
cat /sys/class/drm/card1/device/mem_info_vram_used   # VRAM used, bytes

(If you have multiple GPUs, the card number may differ — check ls /sys/class/drm/.)

From inside Python (what the server process sees)

import torch
torch.cuda.is_available()              # True
torch.cuda.get_device_name(0)          # 'AMD Radeon RX 7900 GRE'
torch.cuda.memory_allocated() / 2**30  # GB currently used by tensors

3. End-to-end verification script

Fire a transcription in the background and sample utilization while it runs:

curl -s http://localhost:8000/v1/audio/transcriptions \
     -F file=@/tmp/jfk.flac -F model=base -o /dev/null &

for i in 1 2 3 4 5 6; do
  busy=$(cat /sys/class/drm/card1/device/gpu_busy_percent)
  vram=$(cat /sys/class/drm/card1/device/mem_info_vram_used)
  echo "sample $i: GPU busy=${busy}%  VRAM used=$((vram / 1024 / 1024)) MiB"
  sleep 0.3
done
wait

4. Expected signatures of genuine GPU usage

Signal On GPU CPU fallback
VRAM after model load elevated, constant (~1–8 GB by model) ~0
gpu_busy_percent during request spikes (short clips) / sustained (long audio) 0%
10 s audio transcription time ~1–2 s (turbo) 10–30 s, fans/CPU high
rocm-smi --showpids python process listed not listed

Short clips with small models (base) only produce brief utilization spikes — that's normal, not a sign of CPU fallback. Use longer audio or turbo/large for sustained load.

5. Troubleshooting: it's running on CPU

If you see 0 VRAM and 0% GPU during requests:

  1. Confirm torch.cuda.is_available() is True in the server process (same venv, same env vars).
  2. Check the server logs at startup for CUDA/HIP errors.
  3. Make sure no stray CUDA_VISIBLE_DEVICES="" or HIP_VISIBLE_DEVICES="" is hiding the GPU.
  4. Verify your user is in the render and video groups (groups | grep -E "render|video"), then re-login if you just added them.