Upgrade llama.cpp from b10173 to b10197 - #372
Merged
Merged
Conversation
GLM_DSA (GLM-5.2) NextN/MTP speculative-decoding support (upstream #25980), entirely internal to the model-architecture layer. No patches or project source needed changes. Full local build + ctest (485/485) verified.
Pure CUDA MMQ kernel-config retuning (RDNA3.5/RDNA3/RDNA4, upstream #26199). No patches or project source needed changes. Full local build + ctest (485/485) verified.
14 commits: model-specific suppress-token support (llama_vocab_get_suppress_tokens, additive), unused has_logit_bias() removal (no project usage), cosmetic slot-similarity variable renames + trace logging in server-context.cpp/ server-task.cpp (no overlap with patches 0002/0003), plus CUDA/Metal/RPC/SYCL backend internals and WebUI fixes. No patches or project source needed changes. Full local build + ctest (485/485) verified.
8 commits: ggml bumped to 0.18.0, CUDA Q2_0 support, a test-only common_get_model_or_exit() helper (unused by project), and an internal synchronize()-before-clear race fix in llama-context.cpp. No patches or project source needed changes. Full local build + ctest (485/485) verified.
bernardladenthin
had a problem deploying
to
maven-central
July 31, 2026 23:14 — with
GitHub Actions
Failure
bernardladenthin
had a problem deploying
to
startgate
July 31, 2026 23:14 — with
GitHub Actions
Error
bernardladenthin
had a problem deploying
to
maven-central
July 31, 2026 23:14 — with
GitHub Actions
Failure
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
GIT_TAGin CMake, Java version constant, README badge, and CLAUDE.md documentationDetails
This is a routine llama.cpp version bump covering 24 commits across 4 intermediate steps:
llama-context.cpp(no header signatures changed)All 7 existing patches apply cleanly across the entire range. No project source code changes required. Full local verification confirms all 485 unit tests pass.
Test plan
cmake --buildsucceeds (libjllama.so+jllama_testcompile and link,-O3)ctestpasses: 485/485 tests passingtts.cppchanges in the range)Related issues / PRs
Upstream llama.cpp releases: b10173 → b10197
Checklist
CONTRIBUTING.mdandCODE_OF_CONDUCT.mdhttps://claude.ai/code/session_018G7MFojeDt6Nt7TRKEAJrz