Add Snapdragon X support: Hexagon NPU and Adreno OpenCL with CPU fallback - #30
Merged
Merged
Conversation
…nCL backends with CPU fallback
…ed README, script autodetection, style fixes
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds native support for Qualcomm Snapdragon X laptops on Windows on Arm: the Oryon CPU, the Adreno GPU (
GGML_OPENCL=ON) and the Hexagon NPU (GGML_HEXAGON=ON).src/backend_fallback.cppwraps the Hexagon or OpenCL backend. Ops the accelerator rejects run on the CPU, and the archs,gallocrand every other backend are unchanged. It is compiled only forGGML_HEXAGONandGGML_OPENCLbuilds.src/backend.h.VLA_DEVICE=cpugives a CPU reference from the same binary.scripts/build_windows_snapdragon.ps1builds each flavour, including HTP skel signing.cmake/vcpkg-tripletsbuilds protobuf with clang-cl for the servers.--weight-dtype f16.VLA_BUILD_SERVERoption.quantize_gguf.pyoutput. Before this, the loader refused packed weights on every platform.docs/backend/hexagon-windows.md: setup, measurements, and the upstream issues.docs/backend/hexagon.mdrewritten around the measured results.Why
The Hexagon column of the support matrix was empty. On a Snapdragon X (X1-26-100, Hexagon v73), SmolVLA runs in 1.23 s per action chunk on the NPU, against 2.57 s on the CPU and 3.05 s on the GPU. It is within 1.5e-3 of the CPU reference.
Verified
-Wall -Wextra(first-party code). This was checked on Windows arm64 with clang 22 for the CPU, OpenCL and Hexagon builds; Linux was not built locally.ctestpasses (7/7) on all three Windows builds.vla_predict_check) is bit-identical before and after. The Hexagon/OpenCL graph changes are behind the build flags and reduce to the old calls elsewhere.Archs and backends tested, on Windows on Arm (Snapdragon X, 16 GB):
~. Both models lose precision on the CPU with F16 weights too.vla-server: smoke-tested on the NPU over ZeroMQ.Built against llama.cpp build 11201 (
2145525a4) viaFETCHCONTENT_SOURCE_DIR_LLAMA. Theb10729pin was not tested with Hexagon.