Skip to content

Commit b8a1d36

Browse files
committed
Upgrade llama.cpp from b11236 to b11237
One upstream commit (#29603): the OpenVINO backend marks unaligned batch-stride views unsupported. Backend-internal; no API change, no patch touches the file, no workflow change upstream. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
1 parent 674351b commit b8a1d36

5 files changed

Lines changed: 12 additions & 10 deletions

File tree

‎CLAUDE.md‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
66

77
Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI.
88

9-
Current llama.cpp pinned version: **b11236**
9+
Current llama.cpp pinned version: **b11237**
1010

1111
## Upgrading CUDA Version
1212

@@ -538,7 +538,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi
538538
ships no UI):
539539
```bash
540540
# needs node/npm + network for the asset build; the embed step is plain cmake -P
541-
git clone --depth 1 --branch b11236 https://github.com/ggml-org/llama.cpp /tmp/lc
541+
git clone --depth 1 --branch b11237 https://github.com/ggml-org/llama.cpp /tmp/lc
542542
( cd /tmp/lc/tools/ui && npm ci && npm run build )
543543
mkdir -p webui-generated /tmp/ui-gen
544544
cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \
@@ -578,7 +578,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend:
578578
- `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored
579579
as the repo secret **`DEPOT_TOKEN`**.
580580

581-
Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11236`), the
581+
Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11237`), the
582582
~280 upstream object files are byte-identical every run, so a warm cache recompiles only the
583583
*changed* files. Depot's cache is **shared across all branches** (unlike GitHub's
584584
per-branch `actions/cache`), so every branch builds incrementally; a `b<nnnn>` version bump
@@ -1795,7 +1795,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson"
17951795

17961796
#### Upstream source location (in CMake build tree)
17971797

1798-
llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11236`.
1798+
llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11237`.
17991799

18001800
**GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely
18011801
by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the

‎README.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -11,7 +11,7 @@
1111
**Build:**
1212
![Java 8+](https://img.shields.io/badge/Java-8%2B-informational)
1313
![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey)
14-
[![llama.cpp b11236](https://img.shields.io/badge/llama.cpp-%23b11236-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11236)
14+
[![llama.cpp b11237](https://img.shields.io/badge/llama.cpp-%23b11237-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11237)
1515
[![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/)
1616
![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162)
1717
[![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev)

‎docs/history/llama-cpp-breaking-changes.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -762,3 +762,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r
762762
| b11214–b11222 | patches + upstream verification | **`0001` and `0006` refreshed, six apply unchanged.** #29537 moved `SRV_INF("initializing ...")` from above to below the `common_params_parse` call in `llama_server()`, which is the context of both patches' parse-call hunk; the hunks now anchor on the preceding `server_stream_session_manager_start()` lines instead. Content unchanged, all eight apply in order on pristine b11222. **Drop-check `0001`: still required** — `common_params_parse_main` has 0 occurrences in `b11222:common/arg.h`, and `common/arg.cpp` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override. |
763763
| b11222–b11236 | Fourteen commits, 51 files, ~2.2k lines. **#29385** migrates speculative decoding, `mtmd` and the server to `batch_ext`: `common/common.h` gains the `common_batch` wrapper and `common_batch_get_one`/`common_batch_from_llama_batch`, drops `string_from(ctx, llama_batch)` and `common_batch_ext_get_one`; `common/speculative.h` adds a `common_batch` overload of `common_speculative_process`; `tools/mtmd/mtmd-helper.h` changes `mtmd_helper_post_decode_callback` to take a `mtmd_helper_embd_batch`. **None of these symbols is used by project code** (`jllama.cpp`, `tts_engine.cpp`, the test sources; `gen_audio` is unchanged). `include/llama.h` only adds `llama_get_causal_attn`. **#28876** lets RANK pooling split batches for causal-LLM rerankers (Qwen3 / Qwen3-VL) inside `server-context.cpp`, no interface change. No `tools/server/*.h`, `server-schema.cpp` or `server-task.cpp` change, so the request/response contract checks have nothing to compare. **#28362** enables a Windows ARM64 build with MSVC `cl.exe` in `ggml-cpu/CMakeLists.txt`; the arm64 job keeps `clang-cl`. No upstream `.github` change, so the ROCm and Windows-CUDA component pins stay. Version-only from this project's side. |
764764
| b11222–b11236 | patches + upstream verification | **`0001` refreshed, eight apply unchanged.** #29426 rewrote `tests/test-recurrent-state-rollback.cpp` so its `main()` strips `--models DIR` into its own `filtered_argv` before `common_params_parse(fargc, filtered_argv.data(), …)` — the same shape `test-save-load-state.cpp` took at b10679 — so by `0001`'s own rule that call site keeps `common_params_parse` and the hunk was **dropped** (37 → 36 file diffs). Verified by applying all nine patches in order to a clean b11236 worktree. **Drop checks:** `0002` still needed (`server-context.cpp` still assigns `params_base.load_progress_callback` unconditionally); `common/arg.cpp`, `src/llama-model.cpp`, `common/log.{h,cpp}` and `ggml/src/ggml-rpc/` are untouched in the range, so `0001`, `0012`, `0014` and `0015` still carry live fixes. |
765+
| b11236–b11237 | One commit, 1 file, 22 lines. **#29603** (`ggml/src/ggml-openvino/ggml-openvino.cpp`): the OpenVINO backend reports views with an unaligned batch stride as unsupported, so the scheduler falls back instead of computing them wrong. Backend-internal, no API change, no workflow change upstream. Version-only from this project's side. |
766+
| b11236–b11237 | patches + upstream verification | **All nine patches apply unchanged.** No patch-target file is in the range. |

‎llama/CMakeLists.txt‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -188,7 +188,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE)
188188
FetchContent_Declare(
189189
llama.cpp
190190
GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git
191-
GIT_TAG b11236
191+
GIT_TAG b11237
192192
PATCH_COMMAND ${CMAKE_COMMAND}
193193
-DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches
194194
-DLLAMA_SRC=<SOURCE_DIR>

‎llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -10,28 +10,28 @@
1010
* library was compiled against, exposed as a compile-time constant so callers can render a badge or
1111
* emit a startup log line without loading the native library.
1212
*
13-
* <p>{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11236"}) that mirrors the
13+
* <p>{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11237"}) that mirrors the
1414
* {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is
1515
* absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a
1616
* lightweight version badge in Android or other UIs.</p>
1717
*
1818
* <p>For the <em>authoritative</em> value that is baked into the native binary — the build number
19-
* plus the resolved upstream commit, e.g. {@code "b11236-<commit>"} — call
19+
* plus the resolved upstream commit, e.g. {@code "b11237-<commit>"} — call
2020
* {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own
2121
* {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires
2222
* the native library to be loaded).</p>
2323
*/
2424
public final class LlamaCppVersion {
2525

2626
/**
27-
* The pinned llama.cpp release tag this library was built against, e.g. {@code "b11236"}.
27+
* The pinned llama.cpp release tag this library was built against, e.g. {@code "b11237"}.
2828
*
2929
* <p>Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the
3030
* "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the
3131
* compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the
3232
* value actually linked into the native binary.</p>
3333
*/
34-
public static final String LLAMA_CPP_VERSION = "b11236";
34+
public static final String LLAMA_CPP_VERSION = "b11237";
3535

3636
// Constants holder — not instantiable.
3737
private LlamaCppVersion() {}

0 commit comments

Comments
 (0)