From af77248a5e234205c34aca9c94034f466a1fdc67 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 14:10:00 +0000 Subject: [PATCH 01/12] Update palantir-java-format to 2.100.0 and logback to 1.6.5 No formatting change under the new formatter (spotless:apply over llama/ and llama-atmosphere-agent/ left every source untouched). logback is test scope here. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- llama-atmosphere-agent/pom.xml | 2 +- llama/pom.xml | 4 ++-- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/llama-atmosphere-agent/pom.xml b/llama-atmosphere-agent/pom.xml index 20144382..8c1c575a 100644 --- a/llama-atmosphere-agent/pom.xml +++ b/llama-atmosphere-agent/pom.xml @@ -58,7 +58,7 @@ SPDX-License-Identifier: MIT 3.6.4 3.8.0 3.10.3 - 2.99.0 + 2.100.0 net.ladenthin.llama.atmosphere.LocalAgent diff --git a/llama/pom.xml b/llama/pom.xml index ffbb06ff..715a11c7 100644 --- a/llama/pom.xml +++ b/llama/pom.xml @@ -65,7 +65,7 @@ SPDX-License-Identifier: MIT 2.22.3 3.8.7 2.0.20 - 1.6.4 + 1.6.5 1.28 6.1.3 3.0 @@ -100,7 +100,7 @@ SPDX-License-Identifier: MIT 7.7.4 1.14.0 3.10.3 - 2.99.0 + 2.100.0 UTF-8 2026-09-01T07:57:45Z From 5e66f409aed6c1a0fb85745fa5103bbf4fa247b7 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 14:10:17 +0000 Subject: [PATCH 02/12] Upgrade llama.cpp from b11259 to b11275 Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index bc253083..534fc939 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11259** +Current llama.cpp pinned version: **b11275** ## Natives jars: one directory per backend (`.github/natives.csv`) @@ -690,7 +690,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11259 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11275 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -730,7 +730,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11259`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11275`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1842,7 +1842,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11259`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11275`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index badebea0..b26d699c 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11259](https://img.shields.io/badge/llama.cpp-%23b11259-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11259) +[![llama.cpp b11275](https://img.shields.io/badge/llama.cpp-%23b11275-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11275) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 4341748b..642c8e23 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -770,3 +770,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11247–b11256 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11256 with `git apply` and by a fresh configure. Two patch-target files are in the range — `tools/cli/cli.cpp` (`0001`; #29632 added two lines after the argument parse, away from the flipped call) and `tools/server/server-context.cpp` (`0002`, `0003`; only the error-message typo). Drop-checks against pristine b11256: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard (`0012` needed); no `common_log_set_callback` (`0014` needed); no `stop_server` in `ggml-rpc.cpp` (`0015` needed). | | b11256–b11259 | Three commits, 8 files, +81/−91 lines. **#29649** (`common/common.{h,cpp}`, `common/arg.cpp`, `common/preset.{h,cpp}`, `tools/rpc/rpc-server.cpp`): `fs_get_config_directory()` now returns `std::filesystem::path` instead of a `std::string` with a trailing separator, and `common_preset_context::load_from_ini()` takes a `std::filesystem::path`; both home-directory lookups share a new file-local `get_home_directory()` with the `getpwuid` fallback now also on macOS (so `` is included on every non-Windows target, Android included — bionic has it). Neither symbol is referenced by project code (`grep` over `llama/src/`), so the signature change is invisible here. **#29638** (`common/sampling.cpp`): `common_sampler_sample_and_accept_n()` stops accepting draft tokens after an EOG token (a trailing EOG from the target is still accepted) — a speculative-decoding fix inside upstream-compiled code, no project call site. **#29565** (`tools/server/server-http.cpp`): when the built-in UI is disabled or replaced (`!params.ui` or a `--path`), `init_listener()` registers `GET /sw.js` returning a self-unregistering service worker; it is in the shared listener setup, so `NativeServer`'s classic and attach modes both get it. `server-schema.cpp` / `server-task.cpp` untouched, no `release.yml` change (CUDA/ROCm/OpenVINO pins stay). Version-only from this project's side. | | b11256–b11259 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11259 with `git apply` and by a fresh configure. Two patch-target files are in the range — `common/arg.cpp` (`0001`, `0015`; #29649 touched only `common_params_apply_system_config`, away from both hunks). Drop-checks against pristine b11259: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard (`0012` needed); no `common_log_set_callback` (`0014` needed); no `stop_server` in `ggml-rpc.cpp` (`0015` needed). | +| b11259–b11275 | Sixteen commits, 47 files, ~690/202 lines, all inside upstream-compiled code. Backend work: CUDA bitonic argsort for rows wider than one block (#28957), HIP packed byte subtraction (#29478), Vulkan GDN tuning, two-at-a-time F32 A loads and MoE-aware `mat_mul_id` tiles (#29476, #29254, #29182), an Adreno `q5_K` `get_tensor` fix (#29555), Hexagon concat/GELU-ERF, AVX512-FP16 dot products accumulated in f32 (#29545). Hardening: `gguf` rejects a tensor size that wraps after padding (#26979), `get_rows_back` checks row bounds (#29575), ggml requires graph inputs to be `GGML_OP_NONE` (#29647), a C++ ODR fix (#29504). Models: rerankers honour `classifier_pooling` (#29627, `src/llama-model.cpp`), PLaMo-2/3 keep `` NORMAL (#29580). The only reviewed header touched is `ggml/include/ggml-zdnn.h` (a backend this project does not build). **No project source change.** | +| b11259–b11275 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11275 with `git apply`. Patch-target file in the range: `src/llama-model.cpp` (`0012`; #29627 changes one line, away from the split hunks). Drop-checks against pristine b11275: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index b24e3ee9..6ff2b325 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -182,7 +182,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11259 + GIT_TAG b11275 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 21c0d722..c334d148 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -9,13 +9,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11259"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11275"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11259-"} — call + * plus the resolved upstream commit, e.g. {@code "b11275-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -23,14 +23,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11259"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11275"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11259"; + public static final String LLAMA_CPP_VERSION = "b11275"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 3ffd648ce5e196064c3e1d45db12debca870199e Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 14:10:18 +0000 Subject: [PATCH 03/12] Upgrade llama.cpp from b11275 to b11278 Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 534fc939..add933c4 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11275** +Current llama.cpp pinned version: **b11278** ## Natives jars: one directory per backend (`.github/natives.csv`) @@ -690,7 +690,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11275 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11278 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -730,7 +730,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11275`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11278`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1842,7 +1842,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11275`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11278`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index b26d699c..a054a19a 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11275](https://img.shields.io/badge/llama.cpp-%23b11275-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11275) +[![llama.cpp b11278](https://img.shields.io/badge/llama.cpp-%23b11278-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11278) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 642c8e23..d94971af 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -772,3 +772,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11256–b11259 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11259 with `git apply` and by a fresh configure. Two patch-target files are in the range — `common/arg.cpp` (`0001`, `0015`; #29649 touched only `common_params_apply_system_config`, away from both hunks). Drop-checks against pristine b11259: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard (`0012` needed); no `common_log_set_callback` (`0014` needed); no `stop_server` in `ggml-rpc.cpp` (`0015` needed). | | b11259–b11275 | Sixteen commits, 47 files, ~690/202 lines, all inside upstream-compiled code. Backend work: CUDA bitonic argsort for rows wider than one block (#28957), HIP packed byte subtraction (#29478), Vulkan GDN tuning, two-at-a-time F32 A loads and MoE-aware `mat_mul_id` tiles (#29476, #29254, #29182), an Adreno `q5_K` `get_tensor` fix (#29555), Hexagon concat/GELU-ERF, AVX512-FP16 dot products accumulated in f32 (#29545). Hardening: `gguf` rejects a tensor size that wraps after padding (#26979), `get_rows_back` checks row bounds (#29575), ggml requires graph inputs to be `GGML_OP_NONE` (#29647), a C++ ODR fix (#29504). Models: rerankers honour `classifier_pooling` (#29627, `src/llama-model.cpp`), PLaMo-2/3 keep `` NORMAL (#29580). The only reviewed header touched is `ggml/include/ggml-zdnn.h` (a backend this project does not build). **No project source change.** | | b11259–b11275 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11275 with `git apply`. Patch-target file in the range: `src/llama-model.cpp` (`0012`; #29627 changes one line, away from the split hunks). Drop-checks against pristine b11275: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | +| b11275–b11278 | Three commits, 6 files: vendored BoringSSL updated to 0.20260929.0 (#29669 -- not compiled here, `CPPHTTPLIB_OPENSSL_SUPPORT` stays undefined), zDNN buffer reset and leak fixes (#29637), Hexagon f16 activation ops (#29209). **No project source change.** | +| b11275–b11278 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11278 with `git apply`. No patch-target file in the range. Drop-checks against pristine b11278: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 6ff2b325..f481f6fe 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -182,7 +182,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11275 + GIT_TAG b11278 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index c334d148..8263e6b8 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -9,13 +9,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11275"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11278"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11275-"} — call + * plus the resolved upstream commit, e.g. {@code "b11278-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -23,14 +23,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11275"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11278"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11275"; + public static final String LLAMA_CPP_VERSION = "b11278"; // Constants holder — not instantiable. private LlamaCppVersion() {} From d4c0c912b8707cf67548ce418a033d319acfa154 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 14:10:18 +0000 Subject: [PATCH 04/12] Upgrade llama.cpp from b11278 to b11279 Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index add933c4..bf89a67f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11278** +Current llama.cpp pinned version: **b11279** ## Natives jars: one directory per backend (`.github/natives.csv`) @@ -690,7 +690,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11278 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11279 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -730,7 +730,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11278`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11279`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1842,7 +1842,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11278`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11279`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index a054a19a..2fa6adad 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11278](https://img.shields.io/badge/llama.cpp-%23b11278-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11278) +[![llama.cpp b11279](https://img.shields.io/badge/llama.cpp-%23b11279-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11279) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index d94971af..b4c72693 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -774,3 +774,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11259–b11275 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11275 with `git apply`. Patch-target file in the range: `src/llama-model.cpp` (`0012`; #29627 changes one line, away from the split hunks). Drop-checks against pristine b11275: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | | b11275–b11278 | Three commits, 6 files: vendored BoringSSL updated to 0.20260929.0 (#29669 -- not compiled here, `CPPHTTPLIB_OPENSSL_SUPPORT` stays undefined), zDNN buffer reset and leak fixes (#29637), Hexagon f16 activation ops (#29209). **No project source change.** | | b11275–b11278 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11278 with `git apply`. No patch-target file in the range. Drop-checks against pristine b11278: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | +| b11278–b11279 | One commit, 32 files, ~2,377/39 lines: GLM-5.3-Flash (`glm5-next`) architecture (#27773) -- a new `src/models/glm5-next.cpp`, the arch/hparam tables in `src/llama-arch.*` / `src/llama-model.*`, the conversion script and a `llama-memory-hybrid-idx` extension. Upstream's `src/CMakeLists.txt` globs `models/*.cpp`, so the FetchContent build picks the new file up without a change here. **No project source change.** | +| b11278–b11279 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11279 with `git apply`. Patch-target files in the range: `src/llama-model.{cpp,h}` (`0012`; the new arch adds a `case` and an hparam block, away from the split helpers). Drop-checks against pristine b11279: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index f481f6fe..de72bcea 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -182,7 +182,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11278 + GIT_TAG b11279 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 8263e6b8..24b4eb00 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -9,13 +9,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11278"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11279"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11278-"} — call + * plus the resolved upstream commit, e.g. {@code "b11279-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -23,14 +23,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11278"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11279"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11278"; + public static final String LLAMA_CPP_VERSION = "b11279"; // Constants holder — not instantiable. private LlamaCppVersion() {} From b0d685700debcd381bead4faaf6a35723e7d80fd Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 14:10:18 +0000 Subject: [PATCH 05/12] Upgrade llama.cpp from b11279 to b11293 Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index bf89a67f..c49e987c 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11279** +Current llama.cpp pinned version: **b11293** ## Natives jars: one directory per backend (`.github/natives.csv`) @@ -690,7 +690,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11279 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11293 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -730,7 +730,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11279`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11293`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1842,7 +1842,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11279`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11293`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 2fa6adad..8a11000d 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11279](https://img.shields.io/badge/llama.cpp-%23b11279-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11279) +[![llama.cpp b11293](https://img.shields.io/badge/llama.cpp-%23b11293-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11293) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index b4c72693..585390c4 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -776,3 +776,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11275–b11278 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11278 with `git apply`. No patch-target file in the range. Drop-checks against pristine b11278: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | | b11278–b11279 | One commit, 32 files, ~2,377/39 lines: GLM-5.3-Flash (`glm5-next`) architecture (#27773) -- a new `src/models/glm5-next.cpp`, the arch/hparam tables in `src/llama-arch.*` / `src/llama-model.*`, the conversion script and a `llama-memory-hybrid-idx` extension. Upstream's `src/CMakeLists.txt` globs `models/*.cpp`, so the FetchContent build picks the new file up without a change here. **No project source change.** | | b11278–b11279 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11279 with `git apply`. Patch-target files in the range: `src/llama-model.{cpp,h}` (`0012`; the new arch adds a `case` and an hparam block, away from the split helpers). Drop-checks against pristine b11279: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | +| b11279–b11293 | Fourteen commits, 70 files, ~4,150/519 lines -- most of it WebUI (`tools/ui`: Hugging Face Hub data layer, model download pipeline, memory-fit estimation, #27947/#27959/#27957 and three follow-ups), which CI rebuilds from the pinned tag. Native: BF16 unary, GLU, binary and scale ops on CPU and CUDA (#29675), the CPU `mul_mat` accepts BF16 in `src1` (#28937), OpenVINO `GET_ROWS` on a weight view (#28381), MUSA defines `__CUDA_ARCH__` for device passes (#29508), SYCL allreduce via pinned host buffers (#29604), plus CI-only model checks. **No project source change.** | +| b11279–b11293 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11293 with `git apply`. No patch-target file in the range. Drop-checks against pristine b11293: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index de72bcea..899348f4 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -182,7 +182,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11279 + GIT_TAG b11293 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 24b4eb00..9f0ceea9 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -9,13 +9,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11279"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11293"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11279-"} — call + * plus the resolved upstream commit, e.g. {@code "b11293-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -23,14 +23,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11279"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11293"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11279"; + public static final String LLAMA_CPP_VERSION = "b11293"; // Constants holder — not instantiable. private LlamaCppVersion() {} From d82528e7cfa00092a34e63cc39a96ed294f0ef39 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 14:10:18 +0000 Subject: [PATCH 06/12] Upgrade llama.cpp from b11293 to b11302 Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index c49e987c..792d24c4 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11293** +Current llama.cpp pinned version: **b11302** ## Natives jars: one directory per backend (`.github/natives.csv`) @@ -690,7 +690,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11293 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11302 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -730,7 +730,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11293`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11302`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1842,7 +1842,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11293`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11302`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 8a11000d..0056531c 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11293](https://img.shields.io/badge/llama.cpp-%23b11293-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11293) +[![llama.cpp b11302](https://img.shields.io/badge/llama.cpp-%23b11302-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11302) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 585390c4..37798c43 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -778,3 +778,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11278–b11279 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11279 with `git apply`. Patch-target files in the range: `src/llama-model.{cpp,h}` (`0012`; the new arch adds a `case` and an hparam block, away from the split helpers). Drop-checks against pristine b11279: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | | b11279–b11293 | Fourteen commits, 70 files, ~4,150/519 lines -- most of it WebUI (`tools/ui`: Hugging Face Hub data layer, model download pipeline, memory-fit estimation, #27947/#27959/#27957 and three follow-ups), which CI rebuilds from the pinned tag. Native: BF16 unary, GLU, binary and scale ops on CPU and CUDA (#29675), the CPU `mul_mat` accepts BF16 in `src1` (#28937), OpenVINO `GET_ROWS` on a weight view (#28381), MUSA defines `__CUDA_ARCH__` for device passes (#29508), SYCL allreduce via pinned host buffers (#29604), plus CI-only model checks. **No project source change.** | | b11279–b11293 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11293 with `git apply`. No patch-target file in the range. Drop-checks against pristine b11293: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | +| b11293–b11302 | Nine commits, 25 files, ~456/126 lines. `llama_prefetch_rows` (#29599) is internal (`src/llama-mmap.*`, `src/llama-model.*`, gemma4/qwen4exp graphs) -- `include/llama.h` does not change. `gguf` integer-overflow fix (#29384), glm5-next scatter-row fix (#29745), Jinja coerced array attributes (#29574, `common/jinja/value.h`), MiMo dflash conversion (#29650), and `llama-cli` exits on stdin EOF (#29722). **No project source change.** | +| b11293–b11302 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11302 with `git apply`. Patch-target files in the range: `src/llama-model.{cpp,h}` (`0012`) and `tools/mtmd/mtmd-cli.cpp` (`0001`'s `common_params_parse_main` flip) -- both away from the patched hunks. Drop-checks against pristine b11302: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 899348f4..5c6e51f4 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -182,7 +182,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11293 + GIT_TAG b11302 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 9f0ceea9..29a79694 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -9,13 +9,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11293"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11302"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11293-"} — call + * plus the resolved upstream commit, e.g. {@code "b11302-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -23,14 +23,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11293"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11302"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11293"; + public static final String LLAMA_CPP_VERSION = "b11302"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 42a22dc31d12ad5e21e1485bd9356eaee5163d89 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 14:10:18 +0000 Subject: [PATCH 07/12] Upgrade llama.cpp from b11302 to b11303 Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 792d24c4..e72dd82c 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11302** +Current llama.cpp pinned version: **b11303** ## Natives jars: one directory per backend (`.github/natives.csv`) @@ -690,7 +690,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11302 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11303 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -730,7 +730,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11302`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11303`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1842,7 +1842,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11302`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11303`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 0056531c..1a0423b8 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11302](https://img.shields.io/badge/llama.cpp-%23b11302-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11302) +[![llama.cpp b11303](https://img.shields.io/badge/llama.cpp-%23b11303-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11303) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 37798c43..b3096a0d 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -780,3 +780,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11279–b11293 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11293 with `git apply`. No patch-target file in the range. Drop-checks against pristine b11293: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | | b11293–b11302 | Nine commits, 25 files, ~456/126 lines. `llama_prefetch_rows` (#29599) is internal (`src/llama-mmap.*`, `src/llama-model.*`, gemma4/qwen4exp graphs) -- `include/llama.h` does not change. `gguf` integer-overflow fix (#29384), glm5-next scatter-row fix (#29745), Jinja coerced array attributes (#29574, `common/jinja/value.h`), MiMo dflash conversion (#29650), and `llama-cli` exits on stdin EOF (#29722). **No project source change.** | | b11293–b11302 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11302 with `git apply`. Patch-target files in the range: `src/llama-model.{cpp,h}` (`0012`) and `tools/mtmd/mtmd-cli.cpp` (`0001`'s `common_params_parse_main` flip) -- both away from the patched hunks. Drop-checks against pristine b11302: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | +| b11302–b11303 | One commit, 34 files, ~513/631 lines -- **an API removal in `common/common.h` / `common/speculative.h`** (#29601, the last examples moved to `llama_batch_ext`): `common_batch_clear()`, `common_batch_add()`, `common_batch_from_llama_batch()` and the `llama_batch` overload of `common_speculative_process()` are gone; `common_batch` gains `get_sub_batch()`, `add_seq()`, a multi-sequence `add()` and a pointer overload of `common_batch_get_one()`, and `add()` no longer documents an abort. None of the removed symbols is referenced by project code (`jllama.cpp`, `tts_engine.cpp`, `train_engine.cpp` and the helpers decode through the server context or `common_opt_*`), so **no project source change**. | +| b11302–b11303 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11303 with `git apply`. The commit rewrites 20 of the `main()` files `0001` flips to `common_params_parse_main` (`examples/*`, `tools/*`, two `tests/*`); every hunk of `0001` still applies, and a sweep of the patched b11303 tree finds no unflipped `common_params_parse(argc, argv, ...)` call left outside `tools/server/server.cpp`, whose two are the embedded path `0006`/`0007` keep on purpose. Drop-checks against pristine b11303: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 5c6e51f4..d4d99bea 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -182,7 +182,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11302 + GIT_TAG b11303 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 29a79694..b61aa838 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -9,13 +9,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11302"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11303"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11302-"} — call + * plus the resolved upstream commit, e.g. {@code "b11303-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -23,14 +23,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11302"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11303"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11302"; + public static final String LLAMA_CPP_VERSION = "b11303"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 00eb723ef69f3b7f342812c9baabd3cea1e2d72c Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 14:10:18 +0000 Subject: [PATCH 08/12] Upgrade llama.cpp from b11303 to b11313 Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index e72dd82c..b7100c27 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11303** +Current llama.cpp pinned version: **b11313** ## Natives jars: one directory per backend (`.github/natives.csv`) @@ -690,7 +690,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11303 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11313 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -730,7 +730,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11303`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11313`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1842,7 +1842,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11303`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11313`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 1a0423b8..7d87b68e 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11303](https://img.shields.io/badge/llama.cpp-%23b11303-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11303) +[![llama.cpp b11313](https://img.shields.io/badge/llama.cpp-%23b11313-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11313) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index b3096a0d..1e0f41d3 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -782,3 +782,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11293–b11302 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11302 with `git apply`. Patch-target files in the range: `src/llama-model.{cpp,h}` (`0012`) and `tools/mtmd/mtmd-cli.cpp` (`0001`'s `common_params_parse_main` flip) -- both away from the patched hunks. Drop-checks against pristine b11302: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | | b11302–b11303 | One commit, 34 files, ~513/631 lines -- **an API removal in `common/common.h` / `common/speculative.h`** (#29601, the last examples moved to `llama_batch_ext`): `common_batch_clear()`, `common_batch_add()`, `common_batch_from_llama_batch()` and the `llama_batch` overload of `common_speculative_process()` are gone; `common_batch` gains `get_sub_batch()`, `add_seq()`, a multi-sequence `add()` and a pointer overload of `common_batch_get_one()`, and `add()` no longer documents an abort. None of the removed symbols is referenced by project code (`jllama.cpp`, `tts_engine.cpp`, `train_engine.cpp` and the helpers decode through the server context or `common_opt_*`), so **no project source change**. | | b11302–b11303 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11303 with `git apply`. The commit rewrites 20 of the `main()` files `0001` flips to `common_params_parse_main` (`examples/*`, `tools/*`, two `tests/*`); every hunk of `0001` still applies, and a sweep of the patched b11303 tree finds no unflipped `common_params_parse(argc, argv, ...)` call left outside `tools/server/server.cpp`, whose two are the embedded path `0006`/`0007` keep on purpose. Drop-checks against pristine b11303: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | +| b11303–b11313 | Ten commits, 21 files, ~525/336 lines. **#28520 changes training** (`src/llama-context.cpp`, `src/llama-graph.cpp`): `opt_init` now sets a training mode only when `n_ubatch == n_ctx`, in which the attention reads K and V of the current ubatch directly so the K/V projections receive gradients; otherwise it logs `n_ubatch (..) != n_ctx (..), the K and V projections will not receive gradients` and keeps the old graph (which never gave them gradients -- the warning makes an existing limit visible). It also always re-reserves the scheduler for the training graph and quadruples `graph_max_nodes` in training mode. `LlamaTrainer` behaves exactly like upstream's `finetune` here (both pass `n_ubatch` through untouched), so **no project source change**; a caller who wants full gradients sets `nUbatch` equal to `nCtx`. Also: the download path fetches the mmproj as well (#28977, `common/arg.cpp`), speculative decoding keeps the original batch order for layer inputs (#29019), qwen4exp `-sm tensor` re-enabled (#28569), CUDA `iq4_nl` short-row guard (#29683), OpenCL/WebGPU/Hexagon fixes. | +| b11303–b11313 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11313 with `git apply`. Patch-target file in the range: `common/arg.cpp` (`0001`, `0015`; #28977 adds three lines in `common_models_handler_init`, away from both). Drop-checks against pristine b11313: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index d4d99bea..9f9b4a8c 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -182,7 +182,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11303 + GIT_TAG b11313 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index b61aa838..a96c34cf 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -9,13 +9,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11303"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11313"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11303-"} — call + * plus the resolved upstream commit, e.g. {@code "b11313-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -23,14 +23,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11303"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11313"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11303"; + public static final String LLAMA_CPP_VERSION = "b11313"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 1da09f6859f035873e8d72d133a047a6d111c71f Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 14:10:19 +0000 Subject: [PATCH 09/12] Upgrade llama.cpp from b11313 to b11320 Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index b7100c27..7927b043 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11313** +Current llama.cpp pinned version: **b11320** ## Natives jars: one directory per backend (`.github/natives.csv`) @@ -690,7 +690,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11313 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11320 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -730,7 +730,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11313`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11320`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1842,7 +1842,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11313`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11320`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 7d87b68e..e4d3d683 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11313](https://img.shields.io/badge/llama.cpp-%23b11313-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11313) +[![llama.cpp b11320](https://img.shields.io/badge/llama.cpp-%23b11320-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11320) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 1e0f41d3..1454fc83 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -784,3 +784,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11302–b11303 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11303 with `git apply`. The commit rewrites 20 of the `main()` files `0001` flips to `common_params_parse_main` (`examples/*`, `tools/*`, two `tests/*`); every hunk of `0001` still applies, and a sweep of the patched b11303 tree finds no unflipped `common_params_parse(argc, argv, ...)` call left outside `tools/server/server.cpp`, whose two are the embedded path `0006`/`0007` keep on purpose. Drop-checks against pristine b11303: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | | b11303–b11313 | Ten commits, 21 files, ~525/336 lines. **#28520 changes training** (`src/llama-context.cpp`, `src/llama-graph.cpp`): `opt_init` now sets a training mode only when `n_ubatch == n_ctx`, in which the attention reads K and V of the current ubatch directly so the K/V projections receive gradients; otherwise it logs `n_ubatch (..) != n_ctx (..), the K and V projections will not receive gradients` and keeps the old graph (which never gave them gradients -- the warning makes an existing limit visible). It also always re-reserves the scheduler for the training graph and quadruples `graph_max_nodes` in training mode. `LlamaTrainer` behaves exactly like upstream's `finetune` here (both pass `n_ubatch` through untouched), so **no project source change**; a caller who wants full gradients sets `nUbatch` equal to `nCtx`. Also: the download path fetches the mmproj as well (#28977, `common/arg.cpp`), speculative decoding keeps the original batch order for layer inputs (#29019), qwen4exp `-sm tensor` re-enabled (#28569), CUDA `iq4_nl` short-row guard (#29683), OpenCL/WebGPU/Hexagon fixes. | | b11303–b11313 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11313 with `git apply`. Patch-target file in the range: `common/arg.cpp` (`0001`, `0015`; #28977 adds three lines in `common_models_handler_init`, away from both). Drop-checks against pristine b11313: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | +| b11313–b11320 | Seven commits, 18 files. LLM-jp-4.1 Harmony dialect handler (#29681, `common/parsers/*` + a chat template), PLaMo-2/3 honour the BOS/EOS settings (#29734), Metal bf16 math for mxfp4 mul-mat (#29770), `llama-bench` fixes. The +16k/-9k line count is `docs/ops/CPU.csv` (#29666, a regenerated ops matrix, ~4 MB of diff) -- which is why `llama-next-version.sh`'s byte threshold reads this range as one 4.8 MB step; excluding `docs/ops` it is 53 KiB. **No project source change.** | +| b11313–b11320 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11320 with `git apply`. No patch-target file in the range. Drop-checks against pristine b11320: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002` needed); no `split_sum == 0` guard in `load_tensors` (`0012` needed); no callback hook in `common/log.h` (`0014` needed); `add_server` still `GGML_ABORT`s on a failed connect and there is no stop function (`0015` needed); no upstream counterpart to `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or the embedded-shutdown hook (`0006`-`0008` needed). `tools/server/` is unchanged in the range, so the request-schema, bound and response-key sets are identical. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 9f9b4a8c..6a8f3a75 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -182,7 +182,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11313 + GIT_TAG b11320 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index a96c34cf..4928ee2e 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -9,13 +9,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11313"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11320"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11313-"} — call + * plus the resolved upstream commit, e.g. {@code "b11320-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -23,14 +23,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11313"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11320"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11313"; + public static final String LLAMA_CPP_VERSION = "b11320"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 1bfc3253c9b1f52b8e5be8e269039a450fba5f41 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 14:49:23 +0000 Subject: [PATCH 10/12] Pin the shared plugin versions once, in the parent's pluginManagement llama-kotlin and llama-langchain4j left maven-resources-plugin unpinned, and llama-kotlin maven-compiler-plugin, llama-langchain4j maven-jar-plugin too, so both modules -- each published to Central -- built with the defaults of whatever Maven ran them: resources 3.3.1 locally and 3.4.0 in CI, compiler 3.13.0 locally and 3.15.0 in CI, jar 3.4.1, against 3.5.0/3.16.0/3.5.1 in llama. Every plugin version more than one module uses (compiler, jar, resources, surefire, source, javadoc, gpg, central-publishing) now sits once in the root pom's pluginManagement. It replaces the copies in llama's pluginManagement, the two side modules' version properties and the literals in the parent's release profile, which named each of these versions two or three times. Plugins only llama uses stay pinned there. Checked by diffing help:effective-pom before and after, with no profile, with release and with release,natives: llama and llama-platform are unchanged, and the side modules differ only in the plugin versions above. Reactor install, llama-langchain4j verify and llama-kotlin test are green, and the release profile still loads central-publishing as a build extension. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- CLAUDE.md | 5 +++- llama-kotlin/pom.xml | 6 ---- llama-langchain4j/pom.xml | 8 ------ llama/pom.xml | 36 ------------------------ pom.xml | 59 ++++++++++++++++++++++++++++++++++++--- 5 files changed, 59 insertions(+), 55 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 7927b043..dc6c3201 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2169,7 +2169,10 @@ The repo root is a thin **aggregator/parent POM** (`net.ladenthin:llama-parent`, All modules inherit the single `` from the parent, so they **ship in lockstep by construction** (no CI guard needed). The parent also holds the shared `release` profile (GPG + -Central Publishing), so one reactor `mvn -P release,natives deploy` signs and publishes all five +Central Publishing) and a `` with every plugin version more than one module uses +(compiler, jar, resources, surefire, source, javadoc, gpg, central-publishing -- a module names such a +plugin without a version; without it a module that pins nothing builds with the default of whatever +Maven runs it, which differed between CI and a local build), so one reactor `mvn -P release,natives deploy` signs and publishes all five Maven artifacts (`llama-parent` pom, `llama` with its natives jars, `llama-langchain4j`, `llama-kotlin`, `llama-platform` pom) at the same version. diff --git a/llama-kotlin/pom.xml b/llama-kotlin/pom.xml index 96fa1b35..6b9082b5 100644 --- a/llama-kotlin/pom.xml +++ b/llama-kotlin/pom.xml @@ -64,9 +64,6 @@ SPDX-License-Identifier: MIT 1.11.0 6.1.3 3.0 - 3.6.0 - 3.4.0 - 3.5.1 @@ -148,7 +145,6 @@ SPDX-License-Identifier: MIT org.apache.maven.plugins maven-surefire-plugin - ${surefire.version} - 3.16.0 - 3.6.0 - 3.4.0 - 3.12.0 @@ -104,19 +100,16 @@ SPDX-License-Identifier: MIT org.apache.maven.plugins maven-compiler-plugin - ${compiler.plugin.version} org.apache.maven.plugins maven-surefire-plugin - ${surefire.version} org.apache.maven.plugins maven-source-plugin - ${source.plugin.version} attach-sources @@ -132,7 +125,6 @@ SPDX-License-Identifier: MIT org.apache.maven.plugins maven-javadoc-plugin - ${javadoc.plugin.version} 17 true diff --git a/llama/pom.xml b/llama/pom.xml index 715a11c7..f6b77ee9 100644 --- a/llama/pom.xml +++ b/llama/pom.xml @@ -363,40 +363,9 @@ SPDX-License-Identifier: MIT maven-assembly-plugin 3.8.0 - - org.apache.maven.plugins - maven-compiler-plugin - 3.16.0 - - - org.apache.maven.plugins - maven-gpg-plugin - 3.2.8 - - - org.apache.maven.plugins - maven-jar-plugin - 3.5.1 - - - org.apache.maven.plugins - maven-javadoc-plugin - 3.12.0 - - - org.apache.maven.plugins - maven-resources-plugin - 3.5.0 - - - org.apache.maven.plugins - maven-source-plugin - 3.4.0 - org.apache.maven.plugins maven-surefire-plugin - 3.6.0 + + + + org.sonatype.central + central-publishing-maven-plugin + 0.11.0 + + + org.apache.maven.plugins + maven-compiler-plugin + 3.16.0 + + + org.apache.maven.plugins + maven-gpg-plugin + 3.2.8 + + + org.apache.maven.plugins + maven-jar-plugin + 3.5.1 + + + org.apache.maven.plugins + maven-javadoc-plugin + 3.12.0 + + + org.apache.maven.plugins + maven-resources-plugin + 3.5.0 + + + org.apache.maven.plugins + maven-source-plugin + 3.4.0 + + + org.apache.maven.plugins + maven-surefire-plugin + 3.6.0 + + + + + release @@ -77,7 +130,6 @@ SPDX-License-Identifier: MIT org.apache.maven.plugins maven-gpg-plugin - 3.2.8 sign-artifacts @@ -97,7 +149,6 @@ SPDX-License-Identifier: MIT org.sonatype.central central-publishing-maven-plugin - 0.11.0 true central From f730fdad9229a82be4f0025085ad545b73e91015 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 15:25:53 +0000 Subject: [PATCH 11/12] Agent: pin maven-resources-plugin and maven-jar-plugin The standalone agent pom left both unpinned, so it built with the defaults of whatever Maven ran it (resources 3.3.1 and jar 3.4.1 here). Now 3.5.0 and 3.5.1, the reactor's versions, as properties next to its other plugin pins. The effective pom, with and without -P assembly, changes only in these two versions; -P assembly verify is green (277 tests, 63 model-gated skips). Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- llama-atmosphere-agent/pom.xml | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/llama-atmosphere-agent/pom.xml b/llama-atmosphere-agent/pom.xml index 8c1c575a..a327182e 100644 --- a/llama-atmosphere-agent/pom.xml +++ b/llama-atmosphere-agent/pom.xml @@ -55,6 +55,8 @@ SPDX-License-Identifier: MIT 3.0 3.16.0 3.6.0 + 3.5.0 + 3.5.1 3.6.4 3.8.0 3.10.3 @@ -182,6 +184,18 @@ SPDX-License-Identifier: MIT maven-compiler-plugin ${compiler.plugin.version} + + + org.apache.maven.plugins + maven-resources-plugin + ${resources.plugin.version} + + + org.apache.maven.plugins + maven-jar-plugin + ${jar.plugin.version} + org.apache.maven.plugins maven-surefire-plugin From 59e1a98fed308a5d9aa34858173215f79349f0cf Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 15:50:58 +0000 Subject: [PATCH 12/12] Publish the agent to Maven Central at the core's version net.ladenthin:llama-atmosphere-agent goes to Maven Central as a thin jar (with Main-Class), sources and javadoc, signed, from a step of its own in publish-snapshot and publish-release right after the reactor deploy, which has just installed the core and its natives jars it resolves against. So `jbang net.ladenthin:llama-atmosphere-agent:` starts it with no checkout and no download by hand. The GitHub-release jar without the core is unchanged. - The natives are a plain runtime dependency on llama-platform, no longer a profile: the published pom must carry them whatever a consumer's tool does with profiles. CI, which installs only the classes and now the llama-platform pom, passes -Dllama.natives=none as before; that activates a profile whose dependencyManagement excludes everything llama-platform names. - The agent's version is the reactor's (5.2.0-SNAPSHOT), llama.version defaults to ${project.version}, and check-natives.py fails when the agent pom and the reactor disagree, since versions:set does not reach it. - developers, scm and distributionManagement for Central; a release profile mirroring the parent's. Javadoc found a {@link} to a test class in ApprovalMode, now {@code}. Checked: -P release verify builds the jar, sources and javadoc jars; in an empty local repository holding only the core classes and the llama-platform pom the agent builds and tests green with -Dllama.natives=none (277 tests, 63 model-gated skips); a consumer project and JBang 0.132.1 resolving only the agent coordinates get the core, the seven CPU/Metal natives jars and one SLF4J provider, and the agent starts; the release-asset jar still carries no core and no natives. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2 --- .github/buildcheck/natives.py | 16 +- .github/buildcheck/tests/test_natives.py | 9 + .github/workflows/publish.yml | 36 +++- CHANGELOG.md | 8 + CLAUDE.md | 20 +- README.md | 14 +- docs/RELEASE.md | 8 +- llama-atmosphere-agent/CLAUDE.md | 45 +++-- llama-atmosphere-agent/README.md | 24 ++- llama-atmosphere-agent/pom.xml | 181 +++++++++++++++--- .../llama/atmosphere/ApprovalMode.java | 2 +- 11 files changed, 297 insertions(+), 66 deletions(-) diff --git a/.github/buildcheck/natives.py b/.github/buildcheck/natives.py index 7904c5a2..d127c79a 100644 --- a/.github/buildcheck/natives.py +++ b/.github/buildcheck/natives.py @@ -255,6 +255,19 @@ def check_agent_class_path(natives, agent_pom_text): return compare("llama-atmosphere-agent/pom.xml Class-Path", fatjar_targets(natives), named) +def check_agent_version(root_pom_text, agent_pom_text): + """The agent is published to Maven Central at the core's version and depends on the core of its + own version, so a release that bumps the reactor without it would publish an agent naming a + core that does not exist (or an old one). It is not a reactor module, so `versions:set` misses it.""" + def version(text): + element = ET.fromstring(text).find("m:version", NS) + return None if element is None else (element.text or "").strip() + core, agent = version(root_pom_text), version(agent_pom_text) + if core == agent: + return [] + return [f"llama-atmosphere-agent/pom.xml is version {agent}, the reactor {core}: set both to the same version"] + + def read(root, path): with open(os.path.join(root, path), encoding="utf-8") as f: return f.read() @@ -271,4 +284,5 @@ def check(root): + check_cmake(natives, read(root, "llama/CMakeLists.txt")) + check_dependency_allowlist(natives, nativedeps.ALLOWED) + check_readme(natives, read(root, "README.md")) - + check_agent_class_path(natives, read(root, "llama-atmosphere-agent/pom.xml"))) + + check_agent_class_path(natives, read(root, "llama-atmosphere-agent/pom.xml")) + + check_agent_version(read(root, "pom.xml"), read(root, "llama-atmosphere-agent/pom.xml"))) diff --git a/.github/buildcheck/tests/test_natives.py b/.github/buildcheck/tests/test_natives.py index ff6b4284..1ba0b4a8 100644 --- a/.github/buildcheck/tests/test_natives.py +++ b/.github/buildcheck/tests/test_natives.py @@ -200,6 +200,15 @@ def test_agent_class_path(self): ["llama-atmosphere-agent/pom.xml Class-Path: missing linux-x86-64"]) + def test_agent_version(self): + def pom(version): + return ('4.0.0' + f'9{version}') + self.assertEqual(natives.check_agent_version(pom("5.2.0-SNAPSHOT"), pom("5.2.0-SNAPSHOT")), []) + self.assertEqual(natives.check_agent_version(pom("5.2.0"), pom("5.2.0-SNAPSHOT")), + ["llama-atmosphere-agent/pom.xml is version 5.2.0-SNAPSHOT, the reactor 5.2.0: " + "set both to the same version"]) + class RepositoryTest(unittest.TestCase): """The checks over this repository: what code-style runs, so a red here is a red there.""" diff --git a/.github/workflows/publish.yml b/.github/workflows/publish.yml index ee021653..6d2505c2 100644 --- a/.github/workflows/publish.yml +++ b/.github/workflows/publish.yml @@ -416,7 +416,7 @@ jobs: # job — never in the dockcross cross-compilers (which have no node) or per-platform. # --------------------------------------------------------------------------- # --------------------------------------------------------------------------- - # llama-atmosphere-agent: the standalone (non-reactor, never on Maven Central) local + # llama-atmosphere-agent: the standalone (non-reactor) local # coding-agent project that wires Atmosphere's built-in OpenAI-compatible agent runtime to # this project's OpenAiCompatServer. Three jobs, all publish gates: # - model-free: unit tests + the wire-contract tests, which drive the REAL @@ -430,7 +430,10 @@ jobs: # - smoke-agent-linux (further down, after package-fatjars): the release asset itself, # started next to the real core fat jar. # The project is built with -Dllama.version= against the core that was just - # installed to the local repo, so it always tests the code of this checkout. + # installed to the local repo, so it always tests the code of this checkout. Only the classes and + # the llama-platform pom are installed; -Dllama.natives=none keeps Maven from resolving the natives + # jars that pom names (they exist only after the package job). The agent itself is published to + # Maven Central by the two publish jobs, after the reactor deploy. # --------------------------------------------------------------------------- test-java-llama-atmosphere-agent: @@ -443,8 +446,10 @@ jobs: with: java-version-file: .java-version distribution: temurin - - name: Install parent + core net.ladenthin:llama into the local repo (Java only) + - name: Install parent + core net.ladenthin:llama + the llama-platform pom into the local repo (Java only) uses: ./.github/actions/build-core + with: + modules: llama,llama-platform - name: Resolve the reactor version run: echo "VERSION=$(mvn -q -DforceStdout help:evaluate -Dexpression=project.version | tail -n1)" >> "$GITHUB_ENV" - name: Spotless check @@ -453,7 +458,8 @@ jobs: run: mvn -B --no-transfer-progress -f llama-atmosphere-agent/pom.xml "-Dllama.version=${VERSION}" -Dllama.natives=none verify # The GitHub Release asset llama-atmosphere-agent--jar-with-dependencies.jar: # the agent plus Atmosphere/JLine, WITHOUT the core (src/assembly/agent-jar.xml), so it is a - # few MB and the natives are not in the release twice. Never deployed to Maven Central. + # few MB and the natives are not in the release twice. Never deployed to Maven Central (the + # thin jar is; see the publish jobs). # smoke-agent-linux launches it next to the real core fat jar; the attach jobs sign it. - name: Build the agent release jar (without the core) run: > @@ -490,8 +496,10 @@ jobs: with: distribution: 'temurin' java-version-file: .java-version - - name: Install parent + core net.ladenthin:llama (classes; the test JVM loads the downloaded native library via lib.path) + - name: Install parent + core net.ladenthin:llama + the llama-platform pom (classes; the test JVM loads the downloaded native library via lib.path) uses: ./.github/actions/build-core + with: + modules: llama,llama-platform - name: Resolve the reactor version run: echo "VERSION=$(mvn -q -DforceStdout help:evaluate -Dexpression=project.version | tail -n1)" >> "$GITHUB_ENV" - name: Run the Atmosphere tool-loop integration test (cached Qwen2.5-1.5B tool model, CPU) @@ -2802,6 +2810,15 @@ jobs: MAVEN_USERNAME: ${{ secrets.CENTRAL_USERNAME }} MAVEN_PASSWORD: ${{ secrets.CENTRAL_TOKEN }} MAVEN_GPG_PASSPHRASE: ${{ secrets.GPG_PASSPHRASE }} + # The agent's thin jar + pom (not a reactor module, so a deploy of its own). It resolves the + # core and the natives jars llama-platform names from the local repository, where the reactor + # deploy above has just installed them; check-natives.py holds its version to the reactor's. + - name: Publish snapshot (llama-atmosphere-agent) + run: mvn --batch-mode --no-transfer-progress -f llama-atmosphere-agent/pom.xml -P release -Dmaven.test.skip=true deploy + env: + MAVEN_USERNAME: ${{ secrets.CENTRAL_USERNAME }} + MAVEN_PASSWORD: ${{ secrets.CENTRAL_TOKEN }} + MAVEN_GPG_PASSPHRASE: ${{ secrets.GPG_PASSPHRASE }} # Android AARs (llama-android, llama-android-opencl): assembled and published by # the plain-Gradle build in llama-android/ — Maven cannot deploy # aar. Natives are already on disk from the artifact @@ -2987,6 +3004,15 @@ jobs: MAVEN_USERNAME: ${{ secrets.CENTRAL_USERNAME }} MAVEN_PASSWORD: ${{ secrets.CENTRAL_TOKEN }} MAVEN_GPG_PASSPHRASE: ${{ secrets.GPG_PASSPHRASE }} + # The agent's thin jar + pom (not a reactor module, so a deploy of its own). It resolves the + # core and the natives jars llama-platform names from the local repository, where the reactor + # deploy above has just installed them; check-natives.py holds its version to the reactor's. + - name: Publish release (llama-atmosphere-agent) + run: mvn --batch-mode --no-transfer-progress -f llama-atmosphere-agent/pom.xml -P release -Dmaven.test.skip=true deploy + env: + MAVEN_USERNAME: ${{ secrets.CENTRAL_USERNAME }} + MAVEN_PASSWORD: ${{ secrets.CENTRAL_TOKEN }} + MAVEN_GPG_PASSPHRASE: ${{ secrets.GPG_PASSPHRASE }} # Android AARs (llama-android, llama-android-opencl): Maven cannot deploy # aar, so the plain-Gradle build signs + publishes them # into a local Maven-layout staging repo, which is zipped into a Central diff --git a/CHANGELOG.md b/CHANGELOG.md index bbbfb5a3..24987312 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,14 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by ## [Unreleased] +### Added +- **`net.ladenthin:llama-atmosphere-agent` on Maven Central**, at the core's version: the agent's thin jar + (with `Main-Class`), sources and javadoc, published right after the reactor. Its pom names + `llama-platform` as a runtime dependency, so `jbang net.ladenthin:llama-atmosphere-agent:` + starts it with the CPU natives of every desktop platform, no checkout and no download by hand. The + release-asset jar without the core stays on the GitHub release. The agent's version now moves with the + core's; `check-natives.py` fails when they differ. + ### Changed - **`ProcessRunner` rewritten on `ProcessBuilder`** (the helper `OSInfo` runs `uname` with): the timeout is now real -- a command that does not end in time is killed and reported as an `IOException`, where diff --git a/CLAUDE.md b/CLAUDE.md index dc6c3201..f0c4fde1 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2226,8 +2226,9 @@ missed again.) release version now appears in only ~4 spots here, not ~20 — the runtime details live once in the classifier table.) - **`llama-langchain4j/README.md`** — its own `` snippet. -- **`llama-atmosphere-agent/pom.xml`** — the `llama.version` property (the **release** version; - standalone project outside the reactor, so `versions:set` skips it), plus the fat-jar filename +- **`llama-atmosphere-agent/pom.xml`** — its own ``, which must equal the reactor's + (standalone project outside the reactor, so `versions:set` skips it; `check-natives.py` fails until + they agree), plus the JBang coordinates and the fat-jar filename `llama--jar-with-dependencies.jar` in the root README's "Local coding agent" section and the project's own README. - **`llama-android/README.md`** and **`llama-kotlin/README.md`** — their Gradle dependency @@ -2367,19 +2368,28 @@ releases as a signed Central Portal bundle upload (staging repo → zip → Publ A copy-and-run terminal agent (console, `--web`, `--acp`) pairing Atmosphere's OpenAI-compatible agent runtime with this project's `OpenAiCompatServer`. A **standalone Maven project** (Java 21, -Atmosphere's floor), **not** a reactor module and **never** on Maven Central. Everything about the +Atmosphere's floor), **not** a reactor module: an application, built without a parent so the folder +can be copied out and run. **Published to Maven Central** at the core's version as a thin jar whose +pom names `llama-platform` as a runtime dependency, so `jbang net.ladenthin:llama-atmosphere-agent:` +starts it with the CPU natives of every desktop platform. Everything about the agent itself -- the REPL, approval gate, file tools, JLine console and its eight carried JLine fixes, the three front ends -- is in **[`llama-atmosphere-agent/CLAUDE.md`](llama-atmosphere-agent/CLAUDE.md)**, which Claude Code loads when working in that directory. What matters from the rest of the repository: +- **Maven Central:** a step of its own in `publish-snapshot` / `publish-release`, right after the + reactor deploy (`-f llama-atmosphere-agent/pom.xml -P release deploy`), resolving the core and its + natives jars from the local repository that deploy just filled. Its CI jobs install only the classes + and the `llama-platform` pom and pass `-Dllama.natives=none` (a profile that excludes everything that + pom names), since no natives jar exists before `package`. - **Release asset:** `llama-atmosphere-agent--jar-with-dependencies.jar`, built **without** the core; its manifest `Class-Path` names the four `all--` fat jars and the default fat jar. Renaming a core fat jar means updating that list -- `check-natives.py` holds it to the fat-jar targets of `natives.csv`, and `smoke-agent-linux` launches the pair. - **CI:** the model-free job, the model-backed integration job and `smoke-agent-linux` all gate both publish jobs. -- **Version bump:** the pom's `llama.version` is the **release** version and `versions:set` does not - touch this standalone pom -- bump it by hand with the fat-jar filename in the two READMEs. +- **Version bump:** the pom's `` is the reactor's, and `versions:set` does not touch this + standalone pom -- bump it by hand (`check-natives.py` fails until it agrees) with the JBang + coordinates and fat-jar filename in the two READMEs; `llama.version` follows (`${project.version}`). - **The one core change it needed:** `OpenAiBackend`, `ChunkSink` and `OpenAiCompatServer(OpenAiBackend, OpenAiServerConfig)` are public, so the agent's tests can drive the real server without a model. Keep them public. diff --git a/README.md b/README.md index e4d3d683..a513b8a7 100644 --- a/README.md +++ b/README.md @@ -1114,9 +1114,17 @@ OpenCode reduced to the essentials, fully offline; it edits files and, with `--a (`docker`, `git`, build tools) — built from [Atmosphere](https://github.com/Atmosphere/atmosphere)'s built-in OpenAI-compatible agent runtime (streaming, tool loop, workspace file tools) driven **headless** against this project's OpenAI-compatible server. It is a standalone Maven project (not a -reactor module, never on Maven Central). Either download it from a release (below, JDK 21+ only), or -clone the repository and run it from that folder, which needs only JDK 21+ and Maven — the core jar -from Maven Central ships the natives: +reactor module) published to Maven Central at the core's version, so the quickest way needs only +JDK 21+ and [JBang](https://www.jbang.dev) — it resolves the agent and the core with the natives of +every desktop platform: + +```bash +jbang net.ladenthin:llama-atmosphere-agent:5.2.0 \ + --model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --workspace /path/to/project +``` + +Or download it from a release (below, JDK 21+ only), or clone the repository and run it from that +folder, which needs only JDK 21+ and Maven — the core jar from Maven Central ships the natives: ```bash # get the folder and a tool-capable model (Qwen3-4B-Instruct-2507, 2.3 GB) diff --git a/docs/RELEASE.md b/docs/RELEASE.md index 1796c7f5..c2ca71b0 100644 --- a/docs/RELEASE.md +++ b/docs/RELEASE.md @@ -21,6 +21,10 @@ mvn -q versions:set -DnewVersion={VERSION} -DgenerateBackupPoms=false from the repo root — it updates the root `` plus every child's `` at once. See the "Version bump" note in [CLAUDE.md](../CLAUDE.md) for the rationale. +**`llama-atmosphere-agent/pom.xml` is not in the reactor**, so `versions:set` skips it: set its +`` to `{VERSION}` in the same commit (`check-natives.py`, in the `code-style` job, fails while +the two differ). It is published at that version right after the reactor deploy. + ## Extra README dependency snippet Besides the root `README.md`, the `llama-langchain4j/README.md` `## Dependency` section, the @@ -33,4 +37,6 @@ One reactor `mvn -P release deploy` signs and publishes the parent pom, `llama`, `llama-langchain4j`, and `llama-kotlin` together at the same version. The **Android AARs** (`llama-android`, `llama-android-opencl`) are published by the `publish-release` job's separate Gradle step (signed Central Portal bundle upload) — no manual action, but they appear as their own -deployment named `llama-android-{VERSION}` in the Central Portal UI. +deployment named `llama-android-{VERSION}` in the Central Portal UI. The agent +(`net.ladenthin:llama-atmosphere-agent`) is likewise its own deployment, from the step after the +reactor deploy. diff --git a/llama-atmosphere-agent/CLAUDE.md b/llama-atmosphere-agent/CLAUDE.md index fe4447c8..d27479a3 100644 --- a/llama-atmosphere-agent/CLAUDE.md +++ b/llama-atmosphere-agent/CLAUDE.md @@ -8,20 +8,32 @@ this project means for the rest of the repository (release asset, CI gates, vers A **copy-and-run general-purpose terminal agent** (Claude Code / OpenCode reduced to the essentials, offline) that pairs [Atmosphere](https://github.com/Atmosphere/atmosphere)'s built-in OpenAI-compatible agent runtime with this project's `OpenAiCompatServer`. Like `android-llmservice/` it is a -**standalone Maven project, NOT a reactor module and NEVER on Maven Central** — it is an application, and it -needs Java 21 (Atmosphere's floor) while the core stays Java 8. CI builds it against the core it just -installed (`-Dllama.version=`); a user copies the folder and runs -`mvn compile exec:java -Dexec.args="…"` with no `-D` at all — the pom's `llama.version` names the -**released** core the READMEs describe (currently `5.2.0`, written as if released so the docs are -right the moment the release lands). The natives come from the -`natives` profile (`llama-platform`, active unless `-Dllama.natives=none`) plus the `gpu-natives` -profile (`-Dllama.classifier=`); CI passes `-Dllama.natives=none`, since it installs only -the classes and the model-backed job uses `-Dnet.ladenthin.llama.lib.path`. +**standalone Maven project, NOT a reactor module** — it is an application, it needs Java 21 +(Atmosphere's floor) while the core stays Java 8, and it is built without a parent so the folder can be +copied out and run on its own. **Its version is the core's** (`5.2.0-SNAPSHOT` on `main`), and +`check-natives.py` fails when the two differ: `versions:set` does not reach this pom, so a release bump +that forgot it would otherwise publish an agent naming the wrong core. `llama.version` defaults to +`${project.version}`; CI still passes `-Dllama.version=` (the same value). + +**It is published to Maven Central** as `net.ladenthin:llama-atmosphere-agent`: the thin jar (with +`Main-Class`, so `jbang ` starts it), its sources and javadoc jars and the pom, signed, by +a step of its own in `publish-snapshot` / `publish-release` right after the reactor deploy +(`-P release deploy`; the `release` profile here mirrors the parent's). It resolves the core and the +natives jars from the local repository, where that reactor deploy has just installed them. + +**The natives are a plain `runtime` dependency on `llama-platform`**, not a profile: they are part of +the published pom, and a consumer (JBang, Maven, Gradle) must get them however it treats profiles — +without them the agent fails at the first load. CI installs only the core classes and the +`llama-platform` pom (`build-core` with `modules: llama,llama-platform`), because no natives jar exists +before the `package` job, and passes `-Dllama.natives=none`: that activates the `no-natives` profile, +whose `dependencyManagement` excludes everything `llama-platform` names, so only its pom is resolved. +The model-backed job loads a downloaded library through `-Dnet.ladenthin.llama.lib.path`. The +`gpu-natives` profile (`-Dllama.classifier=`) adds a GPU backend next to the CPU natives. **Release asset: the agent jar WITHOUT the core.** `mvn -P assembly package` (the pom's `assembly` profile, descriptor `src/assembly/agent-jar.xml`) builds `llama-atmosphere-agent--jar-with-dependencies.jar` — named after the **core** version -it was built against, not the agent's own `1.0.0-SNAPSHOT`, because it only runs next to that core. +it was built against (the agent's own version is the same), because it only runs next to that core. It excludes `net.ladenthin:llama` **with its whole runtime graph** (`useTransitiveFiltering`: Jackson 2, slf4j-api, and Jackson 3's `jackson-annotations`, which resolves through the core's trail) plus `jspecify` and `slf4j-simple`, all of which every core fat jar already bundles — so the asset is ~14 MB @@ -886,10 +898,9 @@ picks its command line with `ShellTool.isWindows()` — the same detection `Shel test counts the platform's line separator. It used plain POSIX commands before and failed 4 of 5 on Windows; never skip it per OS, give a new test both command forms instead. -**Version bump note.** The pom's `llama.version` property is the **release** version, not the -reactor's `-SNAPSHOT` (CI always overrides it, so a not-yet-published default never breaks CI). -`versions:set` does not touch this standalone pom, so at release time bump it by hand together with -the fat-jar filename `llama--jar-with-dependencies.jar` in the two READMEs (`README.md` -"Local coding agent" + the project's own README) — the same class as the `llama-langchain4j/README.md` -snippet. Before the release, run it against a pre-release core with `-Dllama.version=-SNAPSHOT` -(the pom keeps the Sonatype snapshot repository for exactly that). +**Version bump note.** The pom's own `` must equal the reactor's (`check-natives.py` enforces +it); `versions:set` does not touch this standalone pom, so a release bumps it by hand, together with the +version in the two READMEs (`README.md` "Local coding agent" + the project's own README: the JBang +coordinates and the fat-jar filenames) — the same class as the `llama-langchain4j/README.md` snippet. +`llama.version` follows by itself (`${project.version}`); `-Dllama.version=-SNAPSHOT` still runs it +against another core (the pom keeps the Sonatype snapshot repository for exactly that). diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 4c41487c..1072ed30 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -32,10 +32,12 @@ offline: JetBrains IDEs and Zed (and VS Code through an ACP extension) use it as their chat agent, with their own permission dialog for writes and commands. -This folder is a **standalone Maven project**, deliberately *not* a reactor module and *never* on -Maven Central: CI builds and tests it against the core of the same checkout; you copy the folder and -run it — or download the ready-built jar from a release (see below). Its `pom.xml` pins `llama.version` to the release these instructions describe (**5.2.0**); -pass `-Dllama.version=…` to run against another core, e.g. a `-SNAPSHOT` before a release. +This folder is a **standalone Maven project**, deliberately *not* a reactor module: CI builds and tests +it against the core of the same checkout. Its version moves in lockstep with the core's (**5.2.0**), +and it ships three ways: on **Maven Central** as `net.ladenthin:llama-atmosphere-agent` (run it with +JBang, no download and no checkout, see below), as a ready-built jar on every release, and as this +folder to copy and run. It runs on the core of its own version; pass `-Dllama.version=…` to try +another, e.g. a `-SNAPSHOT` before a release. > [!WARNING] > `--allow-shell` lets the model run **any** command with **your** user's rights. By default it asks @@ -49,6 +51,20 @@ You need **JDK 21+** and **Maven** (`mvn -v` must report Java 21 or newer). No C CMake and no separate llama.cpp install: the core jar from Maven Central ships the native libraries for Windows, Linux and macOS. +**Quickest: straight from Maven Central with [JBang](https://www.jbang.dev).** JBang resolves the agent, +Atmosphere and the core with the CPU natives of every desktop platform (Metal on macOS), then starts it — +no checkout, no build, no jar to download by hand: + +```bash +jbang net.ladenthin:llama-atmosphere-agent:5.2.0 \ + --model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --workspace /path/to/project +``` + +The same coordinates work in any Maven project (`mvn exec:java` with `mainClass` +`net.ladenthin.llama.atmosphere.LocalAgent`). For a GPU backend, add the matching natives jar next to it, +e.g. `net.ladenthin:llama:5.2.0` with the classifier `cuda13-linux-x86-64` — the loader prefers it and +falls back to the CPU when its runtime is missing. + **Or skip Maven entirely.** Every [GitHub release](https://github.com/bernardladenthin/java-llama.cpp/releases) carries `llama-atmosphere-agent--jar-with-dependencies.jar` (+ `.sha256`, GPG `.asc`). It holds **no core** — that is why it is a few MB — so download it together with a core fat jar of the same diff --git a/llama-atmosphere-agent/pom.xml b/llama-atmosphere-agent/pom.xml index a327182e..e98a5a8c 100644 --- a/llama-atmosphere-agent/pom.xml +++ b/llama-atmosphere-agent/pom.xml @@ -10,17 +10,18 @@ SPDX-License-Identifier: MIT 4.0.0 net.ladenthin llama-atmosphere-agent - 1.0.0-SNAPSHOT + 5.2.0-SNAPSHOT jar ${project.groupId}:${project.artifactId} @@ -36,15 +37,34 @@ SPDX-License-Identifier: MIT + + + Bernard Ladenthin + https://github.com/bernardladenthin + + + + + scm:git:https://github.com/bernardladenthin/java-llama.cpp.git + scm:git:https://github.com/bernardladenthin/java-llama.cpp.git + https://github.com/bernardladenthin/java-llama.cpp/tree/main + + + + + central + https://central.sonatype.com/repository/maven-snapshots/ + + + UTF-8 21 - - 5.2.0 + + ${project.version} 4.0.71 4.4.6 12.1.13 @@ -57,6 +77,10 @@ SPDX-License-Identifier: MIT 3.6.0 3.5.0 3.5.1 + 3.4.0 + 3.12.0 + 3.2.8 + 0.11.0 3.6.4 3.8.0 3.10.3 @@ -64,7 +88,7 @@ SPDX-License-Identifier: MIT net.ladenthin.llama.atmosphere.LocalAgent - + sonatype-snapshots @@ -80,12 +104,24 @@ SPDX-License-Identifier: MIT + classes. --> net.ladenthin llama ${llama.version} + + + net.ladenthin + llama-platform + ${llama.version} + pom + runtime + @@ -195,6 +231,15 @@ SPDX-License-Identifier: MIT org.apache.maven.plugins maven-jar-plugin ${jar.plugin.version} + + + + + ${agent.main} + + + org.apache.maven.plugins @@ -234,27 +279,36 @@ SPDX-License-Identifier: MIT - natives + no-natives llama.natives - !none + none - - - net.ladenthin - llama-platform - ${llama.version} - pom - - + + + + net.ladenthin + llama-platform + ${llama.version} + pom + + + * + * + + + + + + + release + + + + org.apache.maven.plugins + maven-source-plugin + ${source.plugin.version} + + + attach-sources + jar-no-fork + + + + + org.apache.maven.plugins + maven-javadoc-plugin + ${javadoc.plugin.version} + + all,-missing + + + + attach-javadocs + jar + + + + + org.apache.maven.plugins + maven-gpg-plugin + ${gpg.plugin.version} + + + sign-artifacts + verify + sign + + ${gpg.keyname} + + --pinentry-mode + loopback + + + + + + + org.sonatype.central + central-publishing-maven-plugin + ${central.publishing.version} + true + + central + true + validated + 21600 + + + + + diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java index 2785d279..c55a62ef 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java @@ -58,7 +58,7 @@ public String symbol() { * gefehlt": the transport symbols are drawn two columns wide by Windows Terminal while the column * arithmetic counts them as one, so the glyph covers the single space that follows it and the name * ends up touching it. The space is genuinely in the string — the screen-backed tests in - * {@link ScreenUseCasesTest} assert the badge survives to the rendered row, and they pass — so the + * {@code ScreenUseCasesTest} assert the badge survives to the rendered row, and they pass — so the * disagreement is between the terminal's font and the width table, not in this code. A second space * is the fix that works on both: where the glyph really is two columns wide the result looks like one * space, and where it is one there is a wider gap, which is harmless. The same disagreement in a less