Skip to content

Add Snapdragon X support: Hexagon NPU and Adreno OpenCL with CPU fallback - #30

Merged
khanhnd61-vr merged 3 commits into
mainfrom
backend/hexagon-windows
Sep 27, 2026
Merged

khanhnd61-vr merged 3 commits into
mainfrom
backend/hexagon-windows

Conversation

@khanhnd61-vr

Copy link
Copy Markdown
Collaborator

What

Adds native support for Qualcomm Snapdragon X laptops on Windows on Arm: the Oryon CPU, the Adreno GPU (GGML_OPENCL=ON) and the Hexagon NPU (GGML_HEXAGON=ON).

  • CPU fallback. src/backend_fallback.cpp wraps the Hexagon or OpenCL backend. Ops the accelerator rejects run on the CPU, and the archs, gallocr and every other backend are unchanged. It is compiled only for GGML_HEXAGON and GGML_OPENCL builds.
  • Hexagon rung in src/backend.h.
    • It opens the NPU through the ggml backend registry.
    • VLA_DEVICE=cpu gives a CPU reference from the same binary.
    • It defaults F16 weights and flash attention on.
    • It routes around five ggml-hexagon kernels that return wrong results for VLA shapes: GELU, gapped broadcasts, IM2COL, F16 copy read-back, and op fusion. Each is documented with before/after max|Δ|.
  • Adreno. The xmem GEMM is off by default (precision loss, and a crash on GR00T N1.7).
  • Windows build. The code builds with Visual Studio's clang, and scripts/build_windows_snapdragon.ps1 builds each flavour, including HTP skel signing. cmake/vcpkg-triplets builds protobuf with clang-cl for the servers.
  • --weight-dtype f16.
  • VLA_BUILD_SERVER option.
  • SmolVLA loads quantize_gguf.py output. Before this, the loader refused packed weights on every platform.
  • Docs.
    • New docs/backend/hexagon-windows.md: setup, measurements, and the upstream issues.
    • docs/backend/hexagon.md rewritten around the measured results.
    • README: Hexagon column filled in, Qualcomm mentioned in the introduction.
    • CHANGELOG entries.

Why

The Hexagon column of the support matrix was empty. On a Snapdragon X (X1-26-100, Hexagon v73), SmolVLA runs in 1.23 s per action chunk on the NPU, against 2.57 s on the CPU and 3.05 s on the GPU. It is within 1.5e-3 of the CPU reference.

Verified

  • Builds clean under -Wall -Wextra (first-party code). This was checked on Windows arm64 with clang 22 for the CPU, OpenCL and Hexagon builds; Linux was not built locally.
  • ctest passes (7/7) on all three Windows builds.
  • Numeric output unchanged: SmolVLA's CPU BF16 output (vla_predict_check) is bit-identical before and after. The Hexagon/OpenCL graph changes are behind the build flags and reduce to the old calls elsewhere.

Archs and backends tested, on Windows on Arm (Snapdragon X, 16 GB):

  • Fidelity: max|Δ| of the action chunk against a CPU F32 reference, bar 2.9e-3.
  • CPU, NPU and GPU: SmolVLA, π0, π0.5, GR00T N1.5/1.6/1.7, Evo-1, VLA-Adapter, VLA-JEPA, Octo-Small and TurboVLA.
    • Nine pass on both accelerators.
    • π0 (5.6e-3) and VLA-JEPA (4.8e-3) exceed the bar on the NPU; both are marked ~. Both models lose precision on the CPU with F16 weights too.
  • Not run:
    • BitVLA: the published GGUF is CUDA-only int2.
    • OpenVLA-OFT: more than 8 GB at F16.
  • vla-server: smoke-tested on the NPU over ZeroMQ.
  • Not done: LIBERO, and a Linux build (worth watching in CI).

Built against llama.cpp build 11201 (2145525a4) via FETCHCONTENT_SOURCE_DIR_LLAMA. The b10729 pin was not tested with Hexagon.

@khanhnd61-vr
khanhnd61-vr merged commit 7abe1b4 into main Sep 27, 2026
4 checks passed
@khanhnd61-vr
khanhnd61-vr deleted the backend/hexagon-windows branch September 27, 2026 12:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant