Skip to content

Latest commit

 

History

History
1521 lines (1240 loc) · 91.3 KB

File metadata and controls

1521 lines (1240 loc) · 91.3 KB

java-llama.cpp

Note

No IDE or local setup required. This repository is optimized for fully AI-assisted development using Claude Code. No local toolchain, no IDE, nothing to install — everything works completely through Claude.

AI:
Claude

Build:
Java 8+
Platform
llama.cpp b11320
JPMS
JUnit
JSpecify
NullAway
Checker Framework
Error Prone
Maven Enforcer
Lombok
jqwik
ArchUnit
SpotBugs
jcstress
Lincheck
vmlens
JMH
Publish
CodeQL

Build cache:
Build cache by Depot

Coverage:
Coverage Status
codecov
JaCoCo
PIT Mutation

Quality:
Quality Gate
Code Smells
Security Rating

Security:

Known Vulnerabilities
FOSSA Status
Dependencies
OSV-Scanner

Package:
Maven Central
Snapshot
Release Date
Last Commit

License:
License

Community:
OpenSSF Best Practices
Contribute with Gitpod
OpenSSF Scorecard
Dependabot
Conventional Commits
Keep a Changelog
SemVer
REUSE
Maintained?
Issues
Pull Requests
GitHub Stars
Treeware
Stand With Ukraine

Java Bindings for llama.cpp

Forked from kherud/java-llama.cpp: many thanks to @kherud for the great work!

Inference of Meta's LLaMA model (and others) in pure C/C++.

You are welcome to contribute

  1. Features
  2. Quick Start
    2.1 No Setup required
    2.2 Setup required
  3. Documentation
    3.1 Example
    3.2 Inference
    3.3 Chat Completion
    3.4 Infilling
    3.5 Embeddings & Reranking
    3.6 Raw JSON Endpoints
    3.7 Local agent: terminal, browser, IDE via ACP
  4. Android
  5. Feature Ideas

Features

  • Text completion (blocking and streaming) with full control over sampling parameters.
  • OpenAI-compatible chat completion with automatic chat-template application, including streaming and tool/function calling support via the upstream server.
  • Embeddings (single and native-batched via embed(Collection<String>)) and reranking for retrieval pipelines.
  • Runtime LoRA adapter control — list the loaded adapters and change their scales at runtime without reloading the model (getLoraAdapters() / setLoraAdapters(Map)), the typed counterpart of the upstream GET/POST /lora-adapters endpoints.
  • Text-to-speech (TextToSpeech) over llama.cpp's Qwen3-TTS pipeline (mtmd_helper::gen_audio), returning WAV audio.
  • In-JVM GGUF quantization (LlamaQuantizer) over llama.cpp's llama_model_quantize — convert a GGUF to another quantization scheme without shelling out to llama-quantize.
  • Infilling (fill-in-the-middle) for code models.
  • Tokenize / detokenize and JSON-schema → grammar conversion.
  • Raw JSON endpoint handlers mirroring the upstream llama.cpp HTTP server (/completions, /v1/completions, /embeddings, /infill, /tokenize, /detokenize).
  • Two runnable HTTP server modes, one fat-jar entry. The fat jar's Main-Class is ServerLauncher, which dispatches on the --jllama-openai-compat flag. Without it, java -jar …-jar-with-dependencies.jar -m model.gguf --port 8080 runs the full upstream llama.cpp server (embedded WebUI, every llama-server flag forwarded) hosted inside libjllama over JNI — no separate llama-server.exe. With it, java -jar … --jllama-openai-compat --model model.gguf --port 8080 runs the Java-transport, zero-extra-dependency OpenAI-compatible server (OpenAiCompatServer, streaming SSE) instead. Both are also runnable directly by class name via java -cp … net.ladenthin.llama.server.{NativeServer,OpenAiCompatServer}.
  • Model metadata access (getModelMeta()) and server management (metrics, slot save/restore, runtime thread reconfiguration).
  • Conversation checkpoints — Session.checkpoint(...) / rewind(...) / fork(...) branch and roll back a chat (KV-cache slot save/restore + transcript snapshot) without re-prefilling.
  • GGUF metadata inspection without loading the model (GgufInspector — pure Java, reads header + key/value table only, big-endian aware).
  • Distributed inference over RPC — offload a model's layers to llama.cpp RPC servers on other machines (ModelParameters.setRpcServers(...) / --rpc host:port), and serve this machine's devices to them with RpcServer (the in-JVM rpc-server). See Distributed inference over RPC.
  • Local agent (llama-atmosphere-agent, release asset, JDK 21+) — a fully offline agent on top of this library that reads and edits files and, if allowed, runs commands. One agent session, three ways to use it: a terminal (full console or line-oriented --plain, fine over SSH/PuTTY), a browser (--web, token-protected, loopback by default — reach it from elsewhere through an SSH tunnel), and IDEs over the Agent Client Protocol (--acp: JetBrains IDEs and Zed natively, VS Code through an ACP extension). See Local agent.
  • Multi-model router mode (--models-dir + per-request model selection, managed via the typed RouterClient) and attach mode (NativeServer(LlamaModel, ...) serves an already-loaded model over the full upstream HTTP frontend — one copy of the weights).
  • Pre-built native binaries for Linux (x86-64, aarch64, s390x), macOS (arm64, Metal included), Windows (x86-64, x86, arm64) and Android (arm64, x86-64), plus GPU backends (CUDA, Vulkan, OpenCL, ROCm/HIP, SYCL, OpenVINO) — one natives jar each, all loadable side by side with automatic CPU fallback; see Choosing the natives jars. Android additionally ships as the llama-android AAR with the optional llama-kotlin coroutines façade.

Quick Start

Access this library via Maven (released versions on Maven Central). llama-platform brings the classes plus the CPU natives of every desktop platform:

<dependency>
    <groupId>net.ladenthin</groupId>
    <artifactId>llama-platform</artifactId>
    <version>5.2.0</version>
    <type>pom</type>
</dependency>

(<type>pom</type> is required in Maven: llama-platform is a dependency list, not a jar. In Gradle it is just implementation("net.ladenthin:llama-platform:5.2.0").)

Note

This layout starts with 5.2.0. Up to 5.1.0, net.ladenthin:llama was one jar that carried the CPU natives of every platform, and each GPU classifier was a complete replacement for it.

There are multiple examples.

Snapshot builds

Every push to main publishes a snapshot to the Sonatype Central snapshot repository.

To use the latest snapshot, add the repository and dependency to your pom.xml:

<repositories>
  <repository>
    <id>sonatype-snapshots</id>
    <url>https://central.sonatype.com/repository/maven-snapshots/</url>
    <snapshots><enabled>true</enabled></snapshots>
    <releases><enabled>false</enabled></releases>
  </repository>
</repositories>

<dependency>
    <groupId>net.ladenthin</groupId>
    <artifactId>llama-platform</artifactId>
    <version>5.2.0-SNAPSHOT</version>
    <type>pom</type>
</dependency>

No credentials are required — the repository is publicly readable.

No Setup required

We support CPU inference for the following platforms out of the box:

  • Linux x86-64, aarch64, s390x
  • macOS aarch64 (Apple silicon, with Metal)
  • Windows x86-64, x86, aarch64
  • Android aarch64, x86-64 (see Importing in Android)

If any of these match your platform, you can include the Maven dependency and get started.

Choosing the natives jars

net.ladenthin:llama is the Java classes only. The native libraries ship as separate jars of the same artifact, one per backend and platform, selected by a Maven <classifier> of the form <backend>-<os>-<arch>. Each holds exactly one directory, net/ladenthin/llama/<OS>/<ARCH>/<backend>/, so any combination can share one classpath. At startup the loader tries the backends it finds in a fixed order — CUDA → ROCm → SYCL (fp16, fp32) → Vulkan → OpenCL → OpenVINO → Metal → MSVC → CPU — and uses the first whose library loads. A GPU jar next to the CPU jar therefore just works on a machine without that GPU or its runtime: the GPU library fails to load and the CPU one is used.

llama-platform is the classes jar plus the CPU jars of every desktop platform. For GPU acceleration, add the jar for your GPU next to it:

<dependency>
    <groupId>net.ladenthin</groupId>
    <artifactId>llama-platform</artifactId>
    <version>5.2.0</version>
    <type>pom</type>
</dependency>
<!-- Add any natives jar from the table below. Example shown: CUDA 13 on Linux x86-64. -->
<dependency>
    <groupId>net.ladenthin</groupId>
    <artifactId>llama</artifactId>
    <version>5.2.0</version>
    <classifier>cuda13-linux-x86-64</classifier>
</dependency>

To ship less, depend on net.ladenthin:llama (the classes) plus only the natives jars of the platforms you target, e.g. cpu-linux-x86-64.

Classifier Backend Target platform Runtime requirement
cpu-linux-x86-64 CPU Linux x86-64 A JDK 8+ JVM; glibc ≥ 2.17 (manylinux2014).
cpu-linux-aarch64 CPU Linux aarch64 glibc ≥ 2.39 (e.g. Ubuntu 24.04+, Debian 13+) — built natively on ubuntu-24.04-arm, matching upstream llama.cpp's own ARM binaries; older-glibc ARM hosts (Ubuntu 22.04, Debian 12, RHEL 8/9, Amazon Linux 2023) are not supported.
cpu-linux-s390x CPU Linux s390x (IBM Z, big-endian) A JDK 8+ JVM.
cpu-windows-x86-64 / cpu-windows-x86 CPU Windows x86-64 / x86 A JDK 8+ JVM. Built with Ninja Multi-Config + MSVC (static /MT CRT).
cpu-windows-aarch64 CPU Windows on ARM (Snapdragon X / Surface) A JDK 8+ JVM. Built natively on windows-11-arm with clang-cl.
metal-macos-aarch64 Metal + CPU macOS aarch64 (Apple silicon) A JDK 8+ JVM.
cpu-android-aarch64 / cpu-android-x86-64 CPU Android For Android use the llama-android AAR; these jars are the same libraries for other Android JVM setups.
msvc-windows-x86-64 / msvc-windows-x86 CPU (Visual Studio generator) Windows x86-64 / x86 Same CPU backend and MSVC toolchain as cpu-windows-*, built with the Visual Studio generator instead of Ninja — an alternate-toolchain option; tried before the cpu jar when both are present.
cuda13-linux-x86-64 CUDA 13 Linux x86-64 with NVIDIA GPU NVIDIA driver + CUDA 13 runtime libraries (libcudart.so.13, libcublas.so.13).
cuda13-windows-x86-64 CUDA 13 Windows x86-64 with NVIDIA GPU NVIDIA driver + CUDA 13 Toolkit (cudart64_13.dll, cublas64_13.dll, cublasLt64_13.dll on PATH).
vulkan-linux-x86-64 Vulkan Linux x86-64 with a Vulkan 1.2+ GPU (NVIDIA / AMD / Intel) A Vulkan runtime (libvulkan.so.1), which current GPU drivers install. The most portable Linux GPU option. glibc ≈ 2.39 (built on ubuntu-latest).
vulkan-linux-aarch64 Vulkan Linux aarch64 with a Vulkan 1.2+ GPU A Vulkan runtime (libvulkan.so.1). glibc ≥ 2.39.
vulkan-windows-x86-64 Vulkan Windows x86-64 with a Vulkan 1.2+ GPU A Vulkan runtime (vulkan-1.dll), which current GPU drivers install. The most portable Windows GPU option.
opencl-windows-x86-64 OpenCL Windows x86-64 with an OpenCL 2.0+ GPU A vendor OpenCL ICD (OpenCL.dll). The GGML OpenCL backend is Adreno-tuned; on desktop GPUs CUDA or Vulkan are better supported.
opencl-windows-aarch64 OpenCL (Adreno) Windows on ARM (Snapdragon X) The Adreno driver's OpenCL ICD (OpenCL.dll).
opencl-android-aarch64 OpenCL (Adreno) Android aarch64 with Adreno GPU A device OpenCL ICD (libOpenCL.so); see also the llama-android-opencl AAR.
rocm-linux-x86-64 ROCm / HIP Linux x86-64 with AMD GPU An AMD ROCm 10 runtime (libamdhip64.so, librocblas.so, libhipblas.so) — built against ROCm 10.0 (TheRock), like upstream llama.cpp; every GPU TheRock builds for Linux, Instinct included (gfx900/gfx906/gfx90c/gfx1153 best effort — built, but not release-ready in ROCm 10).
rocm-windows-x86-64 ROCm / HIP Windows x86-64 with AMD GPU The AMD ROCm 10 runtime DLLs (amdhip64.dll, rocblas.dll, hipblas.dll) on PATH; every Radeon target TheRock builds for Windows, gfx900 through RDNA4 (gfx900/gfx906/gfx90c/gfx1153 best effort).
sycl-fp16-linux-x86-64 SYCL (Intel oneAPI, fp16) Linux x86-64 with Intel GPU (Arc / iGPU) An Intel oneAPI / Level-Zero runtime. fp16 accumulation (faster, slightly lower precision).
sycl-fp32-linux-x86-64 SYCL (Intel oneAPI, fp32) Linux x86-64 with Intel GPU (Arc / iGPU) An Intel oneAPI / Level-Zero runtime. fp32 accumulation (higher precision).
sycl-windows-x86-64 SYCL (Intel oneAPI) Windows x86-64 with Intel GPU (Arc / iGPU) The Intel oneAPI / Level-Zero runtime DLLs on PATH.
openvino-linux-x86-64 OpenVINO Linux x86-64 (Intel GPU / NPU / CPU) An Intel OpenVINO runtime.
openvino-windows-x86-64 OpenVINO Windows x86-64 (Intel GPU / NPU / CPU) The Intel OpenVINO runtime DLLs on PATH.

Note

No vendor runtime is bundled; it comes from the GPU driver or toolkit on the host. The GPU jars are validated build-only in CI (GitHub runners have no GPU), so end-to-end GPU inference is verified locally / on self-hosted hardware. A GPU library that loads but finds no usable device (e.g. CUDA installed, no NVIDIA GPU) runs on the CPU inside that library and keeps the loader from trying the next backend; force one with -Dnet.ladenthin.llama.backend=<backend> (e.g. vulkan or cpu), which fails loud instead of falling back.

Note

On the module path each natives jar is an automatic module (net.ladenthin.llama.natives.<classifier>, with _ for -) that nothing requires, so resolve them with --add-modules ALL-MODULE-PATH (or keep the natives jars on the classpath).

Note

Android armeabi-v7a (32-bit ARM) is not published. Only 64-bit Android binaries are shipped: aarch64 (devices) and x86_64 (emulators, Chromebooks, x86-64 Android hardware) as the cpu-android-* natives jars and in the llama-android AAR, plus aarch64 as opencl-android-aarch64. 32-bit Android devices are unsupported by the released artifacts.

The minimum required Android version is API 28 (Android 9.0 Pie). Devices running Android 8.1 (API 27) or earlier are not supported.

Standalone server fat jars (GitHub Releases)

For running the embedded server without any Maven setup, every tagged GitHub Release (and the rolling snapshot pre-release from main) attaches self-contained all-backends fat jars — one download per OS/arch, runnable directly:

java -jar llama-<version>-all-linux-x86-64-jar-with-dependencies.jar -m model.gguf --port 8080
Release asset Bundled GPU backends CPU fallback
llama-<version>-all-linux-x86-64-jar-with-dependencies.jar CUDA 13, ROCm, SYCL fp16/fp32, Vulkan, OpenVINO yes
llama-<version>-all-linux-aarch64-jar-with-dependencies.jar Vulkan yes
llama-<version>-all-windows-x86-64-jar-with-dependencies.jar CUDA 13, ROCm, SYCL, Vulkan, OpenCL, OpenVINO yes
llama-<version>-all-windows-aarch64-jar-with-dependencies.jar OpenCL (Adreno / Snapdragon X) yes
llama-<version>-jar-with-dependencies.jar none (CPU only, incl. macOS Metal) —

Each all-backends jar is the classes, all Java runtime dependencies, and every natives jar of its named OS/arch (except msvc) merged into one file — pick the jar that matches your platform. (A 32-bit JVM on 64-bit Windows, for example, needs the default llama-<version>-jar-with-dependencies.jar, which carries the CPU natives of every platform.) The loader picks the backend exactly as described above, so the jar starts on every host of its OS/arch, with or without a GPU. A .sha256 checksum file and a detached GPG .asc signature accompany every jar. These fat jars are GitHub download assets only — they are not published to Maven Central.

Setup required

If none of the above listed platforms matches yours, currently you have to compile the library yourself (also if you want GPU acceleration).

This consists of two steps: 1) Compiling the libraries and 2) putting them in the right location.

Library Compilation

First, have a look at llama.cpp to know which build arguments to use (e.g. for CUDA support). Any build option of llama.cpp works equivalently for this project. You then have to run the following commands in the llama/ module directory (the native core lives there; the repository root is just the Maven reactor aggregator):

cd llama       # the native core module
mvn compile    # don't forget this line
cmake -B build # add any other arguments for your backend, e.g. -DGGML_CUDA=ON
cmake --build build --config Release

Tip

Use -DLLAMA_CURL=ON to download models via Java code using ModelParameters#setModelUrl(String).

The library is put in a directory matching your platform and backend, which appears in the cmake output. For example:

-- Backend 'cpu' - installing files to /java-llama.cpp/llama/src/main/natives/net/ladenthin/llama/Linux/x86_64/cpu

mvn test puts that directory on the test classpath; mvn -P natives package turns it into a natives jar (in CI, together with every other platform's).

Library Location

This project has to load a single shared library jllama.

Note, that the file name varies between operating systems, e.g., jllama.dll on Windows, jllama.so on Linux, and jllama.dylib on macOS.

The application will search in the following order in the following locations:

  • In net.ladenthin.llama.lib.path: Use this option if you want a custom location for your shared libraries, i.e., set VM option -Dnet.ladenthin.llama.lib.path=/path/to/directory.
  • In java.library.path: These are predefined locations for each OS, e.g., /usr/java/packages/lib:/usr/lib64:/lib64:/lib:/usr/lib on Linux. You can find out the locations using System.out.println(System.getProperty("java.library.path")). Use this option if you want to install the shared libraries as system libraries.
  • From the natives jars on the classpath: every backend directory found for your platform, in the order described in Choosing the natives jars.

System Properties Reference

Every net.ladenthin.llama.* system property recognised by the library, deep-scanned from the source. Runtime properties are resolved through LlamaSystemProperties; test-only properties are declared in the test sources (TestConstants) and consumed by individual test classes.

Property Default Scope Consumer Description
net.ladenthin.llama.lib.path unset (falls back to java.library.path) runtime LlamaLoader Directory containing the native jllama shared library. Checked first, before java.library.path. Set with -Dnet.ladenthin.llama.lib.path=/path/to/dir.
net.ladenthin.llama.tmpdir unset (falls back to java.io.tmpdir) runtime LlamaLoader Custom temporary directory used when extracting the native library from the JAR.
net.ladenthin.llama.osinfo.architecture unset (uses os.arch) runtime OSInfo Override for the architecture string used to locate the bundled library inside the JAR. Useful when os.arch reports an unexpected value (e.g. inside dockcross / chrooted environments).
net.ladenthin.llama.backend unset (auto: the first backend on the classpath whose library loads) runtime LlamaLoader Names one backend directory (e.g. cuda13, vulkan, cpu) to load exclusively — failure is then fatal instead of trying the next backend. See Choosing the natives jars.
net.ladenthin.llama.test.ngl 43 for the general suite; 0 for ToolCallingIntegrationTest test Model-backed integration tests Number of GPU layers used during testing. Pin to 0 on CPU-only hosts: mvn test -Dnet.ladenthin.llama.test.ngl=0. The tool test also selects device none at zero layers so Metal/CUDA is not initialized.
net.ladenthin.llama.tool.model models/Qwen2.5-1.5B-Instruct-Q4_K_M.gguf (test self-skips if missing) test ToolCallingIntegrationTest Path to a tool-capable GGUF used to verify required blocking and streaming tool calls. The default matches the Qwen2.5 model in upstream llama.cpp's tool-call test matrix.
net.ladenthin.llama.nomic.path models/nomic-embed-text-v1.5.f16.gguf (test self-skips if missing) test LlamaEmbeddingsTest#testNomicEmbedLoads Path to a Nomic embedding model (nomic-embed-text-v1.5.f16.gguf or a compatible BERT-family encoder). Regression test for upstream issue #98 (BERT-encoder result_output assertion).
net.ladenthin.llama.vision.model models/SmolVLM-500M-Instruct-Q8_0.gguf (test self-skips if missing) test MultimodalIntegrationTest Path to a vision-capable model GGUF. Any vision-capable GGUF works; CI default is SmolVLM-500M-Instruct-Q8_0.gguf.
net.ladenthin.llama.vision.mmproj models/mmproj-SmolVLM-500M-Instruct-Q8_0.gguf (test self-skips if missing) test MultimodalIntegrationTest Matching mmproj GGUF for the vision model.
net.ladenthin.llama.vision.image llama/src/test/resources/images/test-image.jpg (a CC-BY-4.0 / MIT-granted photo committed to the repo) test MultimodalIntegrationTest Visual prompt image. Any png/jpeg/webp/gif works; the extension drives MIME detection.
net.ladenthin.llama.audio.model unset (test self-skips) test AudioInputIntegrationTest (llama.cpp discussion #13759) Path to an audio-input model GGUF (e.g. Ultravox, Qwen2.5-Omni).
net.ladenthin.llama.audio.mmproj unset (test self-skips) test AudioInputIntegrationTest Matching audio mmproj (encoder) GGUF.
net.ladenthin.llama.audio.input src/test/resources/audios/sample.wav (committed) test AudioInputIntegrationTest .wav/.mp3 audio prompt clip; the extension drives format detection.
net.ladenthin.llama.tts.model models/Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf (test self-skips if missing) test TtsIntegrationTest Path to the Qwen3-TTS backbone (text) GGUF. Any Qwen3-TTS-family model works; CI default is Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf.
net.ladenthin.llama.tts.mmproj models/mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf (test self-skips if missing) test TtsIntegrationTest Path to the matching Qwen3-TTS mmproj GGUF (speaker encoder + code predictor + code2wav decoder); CI default is mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf.
net.ladenthin.llama.train.model models/stories260K.gguf (test self-skips if missing) test LlamaTrainerIntegrationTest Path to the model the fine-tuning smoke trains. Must be F32: llama_set_param skips every other tensor type, so a quantized model would train nothing.

The test-model defaults are exactly CI's model set (.github/models.csv, URLs included), so a model downloaded into models/ is found without any property; the properties only point a test at another file. MultimodalIntegrationTest self-skips when any of the three vision.* paths is missing, so a partial setup (just the vision model + the committed image, no mmproj) lets the test class load without erroring. AudioInputIntegrationTest self-skips the same way over the three audio.* properties. TtsIntegrationTest likewise self-skips unless both tts.model and tts.mmproj paths exist, and LlamaTrainerIntegrationTest unless train.model (default models/stories260K.gguf, an F32 model) does.

Documentation

Example

This is a short example on how to use this library:

public class Example {

    public static void main(String... args) throws IOException {
        ModelParameters modelParams = new ModelParameters()
                .setModel("models/mistral-7b-instruct-v0.2.Q2_K.gguf")
                .setGpuLayers(43);

        String system = "This is a conversation between User and Llama, a friendly chatbot.\n" +
                "Llama is helpful, kind, honest, good at writing, and never fails to answer any " +
                "requests immediately and with precision.\n";
        BufferedReader reader = new BufferedReader(new InputStreamReader(System.in, StandardCharsets.UTF_8));
        try (LlamaModel model = new LlamaModel(modelParams)) {
            System.out.print(system);
            String prompt = system;
            while (true) {
                prompt += "\nUser: ";
                System.out.print("\nUser: ");
                String input = reader.readLine();
                prompt += input;
                System.out.print("Llama: ");
                prompt += "\nLlama: ";
                InferenceParameters inferParams = new InferenceParameters(prompt)
                        .withTemperature(0.7f)
                        .withMiroStat(MiroStat.V2)
                        .withStopStrings("User:");
                for (LlamaOutput output : model.generate(inferParams)) {
                    System.out.print(output);
                    prompt += output;
                }
            }
        }
    }
}

Also have a look at the other examples.

Inference

There are multiple inference tasks. In general, LlamaModel is stateless, i.e., you have to append the output of the model to your prompt in order to extend the context. If there is repeated content, however, the library will internally cache this, to improve performance.

ModelParameters modelParams = new ModelParameters().setModel("/path/to/model.gguf");
InferenceParameters inferParams = new InferenceParameters("Tell me a joke.");
try (LlamaModel model = new LlamaModel(modelParams)) {
    // Stream a response and access more information about each output.
    for (LlamaOutput output : model.generate(inferParams)) {
        System.out.print(output);
    }
    // Calculate a whole response before returning it.
    String response = model.complete(inferParams);
    // Returns the hidden representation of the context + prompt.
    float[] embedding = model.embed("Embed this");
}

Note

Since llama.cpp allocates memory that can't be garbage collected by the JVM, LlamaModel is implemented as an AutoClosable. If you use the objects with try-with blocks like the examples, the memory will be automatically freed when the model is no longer needed. This isn't strictly required, but avoids memory leaks if you use different models throughout the lifecycle of your application.

Chat Completion

For chat models, build a list of role/content pairs and let the library apply the model's chat template. chatComplete() returns the full response, generateChat() streams tokens, and chatCompleteText() returns just the text content of the assistant message.

List<Pair<String, String>> messages = new ArrayList<>();
messages.add(new Pair<>("user", "Write a haiku about Java."));

InferenceParameters inferParams =
        new InferenceParameters("").withMessages("You are a helpful assistant.", messages);

try (LlamaModel model = new LlamaModel(modelParams)) {
    // Streaming
    for (LlamaOutput output : model.generateChat(inferParams)) {
        System.out.print(output);
    }
    // Or blocking, returns the OpenAI-compatible JSON envelope
    String json = model.chatComplete(inferParams);
    // Or just the assistant text
    String text = model.chatCompleteText(inferParams);
}

Reasoning/thinking models can receive custom Jinja template variables via ModelParameters#setChatTemplateKwargs(Map).

Vision / Multimodal Chat

Load a vision-capable GGUF with its matching projector, then place text and image parts in the same user message. Images may come from a file, raw bytes, a data URI, or an HTTP(S) URL:

ModelParameters modelParams = new ModelParameters()
        .setModel("models/SmolVLM-500M-Instruct-Q8_0.gguf")
        .setMmproj("models/mmproj-SmolVLM-500M-Instruct-Q8_0.gguf");

ChatMessage message = ChatMessage.userMultimodal(
        ContentPart.text("Describe this image in one short sentence."),
        ContentPart.imageFile(Paths.get("photo.jpg")));

try (LlamaModel model = new LlamaModel(modelParams)) {
    String answer = model.chatCompleteText(InferenceParameters.empty()
            .withMessages(Collections.singletonList(message))
            .withNPredict(64));
    System.out.println(answer);
}

The same multipart messages[].content shape works through ChatRequest and the embedded OpenAI-compatible /v1/chat/completions server. For a strictly CPU-only run, use setDevices("none").setMmprojOffload(false) in addition to setGpuLayers(0); projector offload has its own upstream default.

On a multi-GPU host the projector can be placed independently of the weights with setMmprojDevice("CUDA1") (llama.cpp --mmproj-device, added upstream in b10541). Exactly one device may be named; the literal "none" keeps the projector on the CPU. OpenAiCompatServer's CLI accepts the same flag as -mmdev/--mmproj-device, and NativeServer forwards it verbatim like every other llama-server flag.

setMmprojDevice(...) and setMmprojOffload(...) write the same upstream field (common_params::mmproj_use_gpu), so where they disagree the outcome would depend on argv order — and the rendered argv comes from a HashMap, whose order is unspecified. The builder therefore resolves the two genuinely ambiguous combinations by dropping the earlier call, and leaves the rest alone:

Combination Resolves to Builder behaviour
named device + setMmprojOffload(true) (use_gpu=true, device) in either order both kept — no clash
named device + setMmprojOffload(false) order-dependent last call wins
"none" + setMmprojOffload(true) order-dependent last call wins
"none" + setMmprojOffload(false) (use_gpu=false) in either order both kept — no clash

So a multi-GPU projector pin survives an explicit setMmprojOffload(true); only a call that would actually contradict the other is dropped. If you need a device after disabling offload, call setMmprojDevice last.

Video input — decode settings only, so far. mtmd has carried a video path since llama.cpp b9562 (#24269); b10647 (#24318) added the --video-* CLI flags and the mtmd_helper_init_opt plumbing that surfaces them. It is compiled into the shipped desktop library (MTMD_VIDEO is on by default, gated on LLAMA_SUBPROCESS, which upstream force-disables on Android and iOS). Its decode settings are exposed as setVideoFps(float), setVideoTimestampInterval(long) and setVideoFfmpegDir(String). The last one matters most in a JVM: upstream shells out to ffmpeg/ffprobe and resolves them from PATH, which an application server, an Android app or a JAR-only container frequently does not have them on — naming the directory is then the only way for video to work at all.

What is not here yet is the content part: upstream's wire type for a video is {"type":"input_video","input_video":{"data":"<base64>"}} (raw base64, not a data: URI, unlike image_url), gated server-side on mtmd_helper_support_video. ContentPart has no videoFile(...) factory emitting that shape, so these knobs currently configure a path this API cannot yet feed directly. Tracked in TODO.md.

Audio input works identically — load an audio-capable model (Ultravox, Qwen2.5-Omni, …) with its audio --mmproj and add a ContentPart.audioFile(...) (or inputAudio(bytes, "wav"|"mp3")) part. It serializes to the OpenAI input_audio content part and routes through the same mtmd pipeline:

ModelParameters modelParams = new ModelParameters()
        .setModel("models/ultravox-v0_5-llama-3_2-1b.gguf")
        .setMmproj("models/mmproj-ultravox-v0_5-llama-3_2-1b-f16.gguf");

ChatMessage message = ChatMessage.userMultimodal(
        ContentPart.text("Transcribe the audio."),
        ContentPart.audioFile(Paths.get("speech.wav")));

try (LlamaModel model = new LlamaModel(modelParams)) {
    System.out.println(model.supportsAudio()); // true
    String answer = model.chatCompleteText(InferenceParameters.empty()
            .withMessages(Collections.singletonList(message))
            .withNPredict(64));
    System.out.println(answer);
}

LlamaModel.supportsVision() / supportsAudio() report which modalities the loaded projector enables.

Tool Calling

Use a tool-aware instruct model and enable Jinja when loading it. A typed request can either return the model's tool calls through chat, or execute registered handlers until the model produces a normal assistant response through chatWithTools:

ToolDefinition weather = new ToolDefinition(
        "get_weather",
        "Get the current weather for a city",
        "{\"type\":\"object\",\"properties\":{\"city\":{\"type\":\"string\"}},"
                + "\"required\":[\"city\"]}");

ChatRequest request = ChatRequest.empty()
        .appendMessage("user", "What is the weather in Paris?")
        .appendTool(weather)
        .withToolChoice("auto")
        .withParallelToolCalls(Boolean.FALSE);

Map<String, ToolHandler> handlers = Collections.singletonMap(
        "get_weather", argumentsJson -> "{\"temperature_c\":21,\"condition\":\"sunny\"}");

try (LlamaModel model = new LlamaModel(new ModelParameters()
        .setModel("models/Qwen2.5-1.5B-Instruct-Q4_K_M.gguf")
        .enableJinja())) {
    ChatResponse response = model.chatWithTools(request, handlers);
    System.out.println(response.getFirstContent());
}

tool_choice is the OpenAI-compatible string form (auto, none, or required). Set parallel_tool_calls to false when handlers should be issued one at a time. Handler failures and unknown tool names are returned to the model as valid {"error":"..."} tool-result JSON.

Infilling

You can simply set InferenceParameters#withInputPrefix(String) and InferenceParameters#withInputSuffix(String).

Embeddings & Reranking

Load the model with enableEmbedding() (or enableReranking()) and call embed(String) to get a sentence embedding, or rerank(query, documents...) to get relevance scores.

ModelParameters modelParams = new ModelParameters()
        .setModel("/path/to/embedding-model.gguf")
        .enableEmbedding();
try (LlamaModel model = new LlamaModel(modelParams)) {
    float[] embedding = model.embed("Embed this sentence");
    // Batch form: one native dispatch for many inputs, results in request order.
    List<float[]> embeddings = model.embed(Arrays.asList("First sentence", "Second sentence"));
}

Runtime LoRA adapter control

Adapters loaded at model-load time (addLoraAdapter(...) / addLoraScaledAdapter(...), optionally setLoraInitWithoutApply() to start disabled) can be listed and re-scaled at runtime without reloading the model — the typed counterpart of the upstream GET/POST /lora-adapters endpoints:

ModelParameters modelParams = new ModelParameters()
        .setModel("models/base.gguf")
        .addLoraScaledAdapter("models/adapter.gguf", 1.0f);
try (LlamaModel model = new LlamaModel(modelParams)) {
    List<LoraAdapter> adapters = model.getLoraAdapters();      // [{id=0, path=..., scale=1.0}]
    model.setLoraAdapter(0, 0.5f);                             // re-scale at runtime
    model.setLoraAdapters(Collections.emptyMap());             // disable all adapters
}

Per the upstream contract, a scale update lists the adapters to keep active — any adapter missing from the map is set to scale 0 (disabled). The native side clears affected KV caches when the effective adapter set changes.

Text-to-Speech

TextToSpeech synthesizes audio from text over llama.cpp's upstream Qwen3-TTS pipeline (mtmd_helper::gen_audio). It is a separate AutoCloseable native type (not a LlamaModel) because TTS loads its own model pair: a backbone (text) GGUF and an mmproj GGUF bundling the speaker encoder, code predictor, and code2wav decoder. synthesize(String) returns a 24 kHz mono 16-bit WAV byte stream.

try (TextToSpeech tts = new TextToSpeech(
        "models/qwen3-tts-backbone.gguf", "models/qwen3-tts-mmproj.gguf")) {
    byte[] wav = tts.synthesize("Hello from llama dot c p p.");
    Files.write(Paths.get("out.wav"), wav);
}

Add (modelPath, mmprojPath, gpuLayers, threads) to offload to the GPU, or synthesize(text, maxFrames, topK, seed) for explicit sampling, or the full synthesize(text, speakerReferenceAudioPath, language, maxFrames, topK, seed) overload for voice cloning from a reference clip. As with LlamaModel, native memory is not GC-managed — use try-with-resources or call close().

This replaces the project's earlier two-model OuteTTS + WavTokenizer pipeline, which upstream #26254 deleted entirely in favor of Qwen3-TTS (llama.cpp b10270); there is no backward-compatible path for the old model pair.

GGUF Quantization

LlamaQuantizer converts a GGUF to another quantization scheme in-process (llama.cpp's llama_model_quantize — the llama-quantize tool without the separate binary):

LlamaQuantizer.quantize("model-f16.gguf", "model-q4_k_m.gguf", QuantizationType.Q4_K_M);
// Re-quantizing an already-quantized GGUF degrades quality and must be opted into:
LlamaQuantizer.quantize("model-q8_0.gguf", "model-q4_0.gguf", QuantizationType.Q4_0,
        /* threads */ 0, /* allowRequantize */ true);

Raw JSON Endpoints

For direct access to the upstream llama.cpp server API, the following methods take a JSON request and return a JSON response, matching the HTTP server's contract:

handleCompletions, handleCompletionsOai, handleChatCompletions, handleInfill, handleEmbeddings, handleTokenize, handleDetokenize.

Server state is exposed via getMetrics(), eraseSlot(int), saveSlot(int, String), restoreSlot(int, String), and getModelMeta().

Conversation checkpoints: rewind + fork (Session)

A Session can be snapshotted and branched — the KV-cache slot state and the transcript move together, so native state and history can never drift apart:

try (Session session = new Session(model, 0, "You are terse.")) {
    session.send("My name is Alice.");
    SessionCheckpoint cp = session.checkpoint("checkpoints/turn1.bin");

    session.send("Tell me a joke.");
    session.rewind(cp);                     // undo everything after the checkpoint
    session.send("Tell me a story instead."); // retry from the branch point

    // Branch into a second slot (model loaded with setParallel(2)+):
    try (Session forked = session.fork(1, "checkpoints/branch.bin")) {
        forked.send("Answer as a pirate.");   // both sessions continue independently
    }
}

Checkpoint files are caller-managed (KV dumps grow with context usage) and both operations are rejected while a stream is in progress. For plain transformer models a rewind is also achievable cheaply by resending a truncated history with cache_prompt (prefix reuse); checkpoints make the branch point exact and are the only reliable rollback for recurrent/hybrid models (e.g. Granite-4), whose state cannot be recomputed from a prefix.

GGUF metadata inspection (no model load)

GgufInspector reads a GGUF's header and key/value table without loading the model — pure Java, no native library, cost independent of file size (parsing stops before the tensor data). Useful for model pickers and download validators:

GgufMetadata meta = GgufInspector.read(Paths.get("models/Qwen3-0.6B-Q4_K_M.gguf"));
meta.getArchitecture();   // Optional[qwen3]
meta.getModelName();      // Optional[Qwen3 0.6B]
meta.getParameterCount(); // OptionalLong[751632384]
meta.getContextLength();  // OptionalLong[40960]  (<arch>.context_length)
meta.getFileType();       // OptionalLong[15]     (llama_ftype, cf. QuantizationType)
meta.getChatTemplate();   // Optional[{{- ... }}]
meta.getEntries();        // full decoded key/value table

Supports GGUF v2/v3, little- and big-endian (auto-detected), and fails loud on v1/corrupt files. For metadata of an already-loaded model use getModelMeta() instead.

Prompt and KV Cache Reuse

Prompt-prefix reuse is enabled by default in llama.cpp and can be controlled per request with InferenceParameters.withCachePrompt(boolean). withCacheReuse(int) enables non-prefix chunk reuse, while withSlotId(int) pins a request to a specific server slot. Session applies its slot id to every request, so generation and save/restore operate on the same KV state.

Typed results expose logical prompt, generated, cached prompt, and evaluated prompt counts through Usage. Per-request timing also remains available through Timings.getCacheN(). LlamaModel.getMetricsTyped().getSlotMetrics() reports each slot's logical, processed, cached, decoded, and remaining token counts, and the same ServerMetrics view carries the server-wide lifetime counters — including cached prompt tokens (getCumulativeCachedPromptTokens()) and the speculative-decoding tallies (getDraftTokensTotal(), getDraftAcceptedTotal(), getDraftVerifyStepsTotal(), getDraftAcceptedPerPosition(), plus the derived getDraftAcceptanceRate()), which upstream otherwise exposes only as Prometheus text.

The embedded HTTP server exposes the same native JSON at authenticated GET /metrics, with the slot array alone at GET /slots. OpenAI responses preserve usage.prompt_tokens_details.cached_tokens; Responses API output uses usage.input_tokens_details.cached_tokens; Anthropic output uses cache_read_input_tokens.

OpenAI-compatible HTTP server

net.ladenthin.llama.server.OpenAiCompatServer turns a loaded model into a local OpenAI-compatible HTTP endpoint using only the JDK's built-in com.sun.net.httpserver — no extra dependency and no separate server process. It is embeddable, and runnable via java -cp <jar> net.ladenthin.llama.server.OpenAiCompatServer … (the fat jar's default Main-Class is instead NativeServer — see "Native server with the built-in WebUI" below). It serves:

Method & path Backed by
POST /v1/chat/completions LlamaModel.streamChatCompletion (streaming SSE) / chatComplete (blocking)
POST /v1/completions LlamaModel.handleCompletionsOai
POST /v1/embeddings (requires --embedding) LlamaModel.handleEmbeddings
POST /v1/rerank (requires --reranking) LlamaModel.handleRerank (reshaped to results/data)
POST /infill LlamaModel.handleInfill (fill-in-the-middle autocomplete)
GET /v1/models the configured model id
GET /metrics native server and per-slot token/cache counters (JSON)
GET /slots native per-slot token/cache counters (JSON array)
GET /health static {"status":"ok"} (unauthenticated)

Chat completions support streaming via Server-Sent Events and non-streaming, forwarding messages/tools verbatim. The streaming path carries delta.tool_calls and (with stream_options.include_usage) a trailing usage chunk, so agent/tool-calling clients work — this is the recommended surface for VS Code Copilot agent mode, Cline, Roo Code and Continue. response_format (json_object / json_schema) is forwarded for structured outputs. Completions, embeddings, rerank and infill are non-streaming.

Every route is also reachable without the /v1 prefix, the server answers CORS preflight (OPTIONS) and stamps Access-Control-Allow-Origin (so browser/webview clients work), and POST /infill is the llama.cpp-native FIM endpoint for local ghost-text autocomplete plugins (llama.vscode, Twinny, Tabby, Continue's llama.cpp provider). Note: GitHub Copilot's inline completions cannot be served by any local endpoint — only its chat/agent surfaces — so use one of those autocomplete plugins for ghost text.

Alternative protocol surfaces. For clients that don't speak OpenAI Chat Completions, the same model is exposed through additional protocols (pure translation over the OpenAI core — no extra inference path), all supporting tools and streaming:

Surface Routes For
Ollama-native GET /api/version, GET /api/tags, POST /api/show, POST /api/chat (NDJSON streaming), POST /api/generate (prompt completion / FIM) Copilot's built-in Ollama provider; Ollama-hardcoded tools
Anthropic Messages POST /v1/messages (SSE event stream) Claude-shaped clients (Claude Code); Copilot messages apiType
OpenAI Responses POST /v1/responses (SSE event stream) Copilot responses apiType; Responses-API clients

/api/show advertises the model's capabilities (tools, insert, and vision when --mmproj is set) and context length, which Copilot's Ollama provider reads to enable agent mode. The llama.cpp-native GET /props reports default_generation_settings.n_ctx and a modalities block, which autocomplete clients such as llama.vscode read to size their context window.

Embed it in your app:

ModelParameters modelParams = new ModelParameters().setModel("models/model.gguf").setParallel(2);
OpenAiServerConfig config = OpenAiServerConfig.builder().port(8080).modelId("local-model").build();
try (LlamaModel model = new LlamaModel(modelParams);
     OpenAiCompatServer server = new OpenAiCompatServer(model, config).start()) {
    Thread.currentThread().join(); // serve until interrupted
}

…or run it standalone. The fat jar's Main-Class is the ServerLauncher dispatcher, so add --jllama-openai-compat to select this Java server (the launcher strips that flag and forwards the rest); or name the class explicitly via -cp:

# fat jar (bundles the native lib + Java deps) — select the Java server with --jllama-openai-compat
java -jar target/llama-<version>-jar-with-dependencies.jar --jllama-openai-compat \
    --model models/Qwen3-0.6B-Q4_K_M.gguf --host 0.0.0.0 --port 8080 --n-gpu-layers 99

# or name the class explicitly (fat jar or plain library jar)
java -cp target/llama-<version>.jar net.ladenthin.llama.server.OpenAiCompatServer \
  --model models/model.gguf --port 8080 --model-id local-model

Run with --help for the full option list (-m/--model, --host, -p/--port, -c/--ctx-size, -b/--batch-size, -ub/--ubatch-size, -ngl/--n-gpu-layers, -t/--threads, -tb/--threads-batch, -ctk/--cache-type-k, -ctv/--cache-type-v, --jinja, --chat-template-kwargs, --parallel, --model-id, --api-key, --mmproj, -mmdev/--mmproj-device, --embedding, --reranking). The tuning flags mirror llama.cpp's server, so an invocation like --jinja --chat-template-kwargs '{"reasoning_effort":"low"}' -ctk q8_0 -ctv q8_0 -b 4096 -ub 2048 works directly.

Verify with curl (streaming chat):

curl -N http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"local-model","stream":true,"messages":[{"role":"user","content":"hi"}]}'

VS Code Copilot setup: Command Palette → Chat: Manage Language Models → Add Models → Custom Endpoint; enter a group name, a display name and any non-empty API key, and pick API type Chat Completions. VS Code then opens chatLanguageModels.json — set the model url to your endpoint (the host/port go here, not in the form):

[
  {
    "name": "Local llama.cpp",
    "vendor": "customendpoint",
    "apiKey": "local-dummy-key",
    "apiType": "chat-completions",
    "models": [
      {
        "id": "local-model",
        "name": "Local model",
        "url": "http://127.0.0.1:8080/v1/chat/completions",
        "toolCalling": true,
        "vision": false,
        "maxInputTokens": 6144,
        "maxOutputTokens": 2048
      }
    ]
  }
]

Notes: BYOK powers the chat/agent experience only (inline completions and embeddings still require a GitHub account). On CPU, prefer a smaller model and a modest context window — the server emits SSE heartbeats so a long prompt prefill does not trip the client's stream-inactivity timeout. Agent-mode tool calling depends on the model's own tool-calling quality. Pass --api-key (or OpenAiServerConfig.apiKey(...)) to require an Authorization: Bearer token; the server binds to 127.0.0.1 by default.

Native server with the built-in WebUI (NativeServer)

OpenAiCompatServer above is a JSON API server (its / is a 404 — no web page). If you want the full upstream llama.cpp server, including its bundled Svelte WebUI, use net.ladenthin.llama.server.NativeServer. It runs the real llama_server inside libjllama over JNI — no separate llama-server.exe — and forwards the raw llama-server arguments verbatim, so every flag works exactly as it does for the standalone binary. The fat jar runs it by default (when --jllama-openai-compat is absent), forwarding its args to the native server (pass --help for the full llama-server option list):

java -jar target/llama-<version>-jar-with-dependencies.jar \
    -m models/model.gguf --host 127.0.0.1 --port 8080 -c 65536 --jinja
# then open http://127.0.0.1:8080/ for the WebUI

Or embed it:

try (NativeServer server = new NativeServer(
        "-m", "gpt-oss-20b-UD-Q4_K_XL.gguf",
        "--host", "127.0.0.1", "--port", "8080",
        "-c", "65536", "-b", "4096", "-ub", "2048",
        "--jinja", "-ngl", "0", "-t", "8", "-tb", "16",
        "-ctk", "q8_0", "-ctv", "q8_0",
        "--chat-template-kwargs", "{\"reasoning_effort\":\"low\"}",
        "--parallel", "1").start()) {
    // Open http://127.0.0.1:8080/ in a browser for the WebUI; the OpenAI API is at /v1/... too.
    Thread.currentThread().join();
}

Differences from OpenAiCompatServer: with the classic constructor it loads its own model from the arguments (an independent lifecycle, like llama-server.exe), it is single-instance per process, it serves the WebUI (in released jars — local cmake builds ship the empty-asset stub, so no UI there), and it is not available on Android (the upstream server needs posix_spawn). Readiness: poll GET /health. No SSL (plain HTTP — bind localhost or front with a TLS proxy).

Attach mode — serve an already-loaded LlamaModel

NativeServer can also attach the full upstream HTTP frontend (routes, WebUI, resumable streaming) to a LlamaModel you already loaded — one copy of the weights, shared between direct JNI calls and HTTP:

try (LlamaModel model = new LlamaModel(new ModelParameters().setModel("models/model.gguf"));
     NativeServer server = new NativeServer(model, "--host", "127.0.0.1", "--port", "8080").start()) {
    // HTTP (incl. WebUI in released jars) and direct Java calls share the same loaded model.
    String direct = model.complete(new InferenceParameters("2+2=").withNPredict(4));
    Thread.currentThread().join();
}

In attach mode the arguments carry only the HTTP-side flags (--host, --port, --api-key, …; no -m), the server reports healthy immediately (the model is already loaded), and the caller keeps ownership of the model — close the server before the model, never the other way around.

Router mode — multi-model management

Started without a model argument, the upstream server runs in router mode: it lists models from --models-dir, loads/unloads them on demand (GET /models, POST /models/load, POST /models/unload, per-request "model" selection) and serves each model from a worker subprocess. Upstream spawns workers by re-executing its own binary — inside a JVM that binary is java, so before starting an embedded router you must point the worker spawn at this library's bootstrap:

String javaBin = System.getProperty("java.home") + File.separator + "bin" + File.separator + "java";
NativeServer.setWorkerCommand(javaBin, "-cp", System.getProperty("java.class.path"),
        "net.ladenthin.llama.server.NativeServer");
try (NativeServer router = new NativeServer(
        "--host", "127.0.0.1", "--port", "8080", "--models-dir", "models").start()) {
    Thread.currentThread().join(); // each loaded model runs as a fresh worker JVM
}

Worker-command tokens may not contain whitespace (the value is whitespace-split natively).

Typed model management (RouterClient). Instead of hand-rolling HTTP+JSON against the management endpoints, use server.RouterClient — a plain-HTTP typed client (works against the embedded router above or any external llama-server router):

RouterClient client = new RouterClient(8080);
List<RouterModel> models = client.listModels();          // GET /models, typed status per entry
client.loadModel("Qwen3-0.6B-Q4_K_M");                   // POST /models/load (non-blocking)
client.awaitModelLoaded("Qwen3-0.6B-Q4_K_M", 240_000L);  // poll until LOADED; fails fast if the
                                                         // worker died (exit code in the message)
client.unloadModel("Qwen3-0.6B-Q4_K_M");                 // POST /models/unload

RouterModel carries the identifier, the lifecycle status (UNLOADED/LOADING/LOADED/SLEEPING/DOWNLOADING/DOWNLOADED), and the router's failed-worker marker. Chat requests then select a model per request via the standard "model" field on POST /v1/chat/completions.

Against a router started with --api-key, pass the key — it is sent as Authorization: Bearer <key> on every call. All of them need it: /models/load and /models/unload were always gated, and since llama.cpp b10519 the listing endpoints are too.

RouterClient client = new RouterClient(8080, System.getenv("LLAMA_API_KEY"));
// or, for a remote router: new RouterClient("router.internal", 8080, key)

Note

awaitModelLoaded waits by polling GET /models, so it cannot observe a model the router deliberately hides from that listing — a cache model deduplicated by a preset with dedup-cache-models still loads and still serves by name, but never appears. For those, skip the await and issue the request directly; with autoload the router waits for the worker itself.

Distributed inference over RPC

llama.cpp's RPC backend spreads one model over the devices of several machines: every machine that contributes runs an RPC server, and the machine that loads the model names them with --rpc. Layers are then distributed over local and remote devices exactly as over several local GPUs (setGpuLayers, setTensorSplit). Both halves are in every natives jar — CPU and GPU — with no additional runtime dependency (plain TCP over the system socket library the library already links).

Serve this machine's devices (every GPU this library found, else the CPU):

try (RpcServer server = RpcServer.startLocal(RpcEndpoint.DEFAULT_PORT)) {   // 127.0.0.1:50052
    server.awaitTermination();
}

or from the command line, with the fat jar:

java -cp llama-<version>-jar-with-dependencies.jar net.ladenthin.llama.RpcServer --port 50052

Use the servers from a model:

ModelParameters params = new ModelParameters()
        .setModel("models/big-model.gguf")
        .setGpuLayers(99)
        .setRpcServers(RpcEndpoint.parse("10.0.0.2:50052"), RpcEndpoint.parse("10.0.0.3:50052"));

The same works for both HTTP servers: --rpc 10.0.0.2:50052,10.0.0.3:50052 is forwarded to the native server as-is, and OpenAiCompatServer accepts it too. Any upstream rpc-server works as a server, and this library's RpcServer works for any llama.cpp client.

Warning

The RPC protocol has no authentication and no encryption: whoever reaches the port can use the devices and read or write the tensors on them. RpcServer.startLocal therefore binds to loopback only; RpcServer.startOnNetwork(address, …) (or --host on the command line) is the explicit opt-in for another interface and logs a warning. Across machines, use a trusted network or a tunnel (SSH, WireGuard).

What to know:

  • An unreachable server fails the load with a LlamaException naming it, instead of reaching llama.cpp. A server that disappears after the model loaded still terminates the process — llama.cpp has no error path for a device lost mid-inference.
  • One RpcServer per process. A second start while one runs throws IllegalStateException.
  • Endpoints are IPv4 addresses or host names (host:port); llama.cpp's RPC transport has no IPv6. RpcServer binds to an IPv4 literal (127.0.0.1, 0.0.0.0, an interface address).
  • Registered servers stay registered. llama.cpp keeps RPC devices in a process-wide registry with no way to remove them; this library therefore gives every later load that does not ask for a server an explicit device list without it, so a model loaded without --rpc never offloads to a server an earlier model used — the same holds for the multimodal projector, TextToSpeech and LlamaTrainer. An explicit setDevices(...) / --device is never overridden.
  • Android needs the android.permission.INTERNET permission for RPC, even over loopback — the llama-android AAR does not request it, so an app that wants RPC must declare it itself.
  • RpcServer.startLocal(port, threads, cacheDir) enables upstream's tensor cache: a client that loads the same model again sends the large tensors only once.
  • Choose the served devices when the default is wrong. RpcServer.startLocal(port, threads, cacheDir, Arrays.asList("CPU")) (or --device CPU on the command line; names as llama.cpp prints them, e.g. CUDA0, Vulkan1, MTL0) replaces the default of every accelerator. It matters because llama.cpp's RPC client treats every operation as supported by the remote device: a served GPU that cannot run one terminates the server process on the first graph that needs it. The paravirtual GPU of a macOS virtual machine is such a device — serve CPU there.

LangChain4j integration

A separate artifact, net.ladenthin:llama-langchain4j, adapts a LlamaModel to LangChain4j's ChatModel, StreamingChatModel, EmbeddingModel and ScoringModel interfaces in-process over JNI — no HTTP hop, no separate server. It is a separate artifactId (not a classifier of the core) because LangChain4j 1.x requires Java 17 while the core net.ladenthin:llama stays Java 8; keeping it separate avoids forcing that floor on every core consumer. It ships and versions in lockstep with the core.

<dependency>
    <groupId>net.ladenthin</groupId>
    <artifactId>llama-langchain4j</artifactId>
    <version>5.1.0</version>
</dependency>

From 5.2.0 on, add the natives next to it — llama-platform or the natives jars you need (see Choosing the natives jars); the core it depends on is classes only.

Each adapter borrows a LlamaModel you already loaded — it never loads or closes the native model, so you manage its lifecycle (try-with-resources), and one LlamaModel can back several adapters at once:

try (LlamaModel llama = new LlamaModel(new ModelParameters().setModel("models/qwen3-0.6b.gguf"))) {
    ChatModel chat = new JllamaChatModel(llama);
    String reply = chat.chat("Write a haiku about lazy senior devs.");
    System.out.println(reply);
}
Adapter LangChain4j interface java-llama.cpp call
JllamaChatModel ChatModel LlamaModel.chat(...)
JllamaStreamingChatModel StreamingChatModel LlamaModel.generateChat(...) (token streaming)
JllamaEmbeddingModel EmbeddingModel LlamaModel.embed(...) (model loaded with enableEmbedding())
JllamaScoringModel ScoringModel (re-ranking) LlamaModel.handleRerank(...) (model loaded with enableReranking())

See llama-langchain4j/README.md for streaming/embedding/re-ranking examples and the current mapping limitations (tool calling, JSON mode, and multimodal input are not yet forwarded).

Local agent: terminal, browser, IDE via ACP

Tip

Supported front ends — all on the same agent session (history, slash commands, approval mode):

Front end Start with Where it runs
Terminal (default) / --plain this console; --plain for piped, logged or line-only sessions
Browser --web http://127.0.0.1:8787/?token=… — over SSH: ssh -L 8787:127.0.0.1:8787 user@server
IDE via ACP --acp JetBrains IDEs and Zed natively, VS Code through an ACP extension

llama-atmosphere-agent/ is a copy-and-run general-purpose agent on the JVM — Claude Code / OpenCode reduced to the essentials, fully offline; it edits files and, with --allow-shell, runs any command on your machine (docker, git, build tools) — built from Atmosphere's built-in OpenAI-compatible agent runtime (streaming, tool loop, workspace file tools) driven headless against this project's OpenAI-compatible server. It is a standalone Maven project (not a reactor module) published to Maven Central at the core's version, so the quickest way needs only JDK 21+ and JBang — it resolves the agent and the core with the natives of every desktop platform:

jbang net.ladenthin:llama-atmosphere-agent:5.2.0 \
    --model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --workspace /path/to/project

Or download it from a release (below, JDK 21+ only), or clone the repository and run it from that folder, which needs only JDK 21+ and Maven — the core jar from Maven Central ships the natives:

# get the folder and a tool-capable model (Qwen3-4B-Instruct-2507, 2.3 GB)
git clone --depth 1 https://github.com/bernardladenthin/java-llama.cpp.git
cd java-llama.cpp/llama-atmosphere-agent
curl -L --create-dirs -o models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf \
  https://huggingface.co/unsloth/Qwen3-4B-Instruct-2507-GGUF/resolve/main/Qwen3-4B-Instruct-2507-Q4_K_M.gguf

# the agent with the model loaded in-process — a you> prompt appears (/clear, /exit)
mvn -q compile exec:java \
    -Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ctx-size 16384 --workspace /path/to/project --allow-shell"

Download instead of cloning. Every GitHub release carries llama-atmosphere-agent-<version>-jar-with-dependencies.jar (with .sha256 and a GPG .asc, like the other fat jars). It is a few MB because it holds no core: it runs next to one of the core fat jars of the same release, so the natives are downloaded once. Put both in one directory and java -jar finds the core through the agent's manifest (the all-backends jars are tried before the CPU-only default jar):

# e.g. Linux x86-64: the agent + the all-backends core fat jar of the same version
java -jar llama-atmosphere-agent-5.2.0-jar-with-dependencies.jar \
    --model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ctx-size 16384 --workspace /path/to/project --allow-shell

# or with the classpath spelled out (any directory layout; `;` instead of `:` on Windows)
java -cp llama-atmosphere-agent-5.2.0-jar-with-dependencies.jar:llama-5.2.0-all-linux-x86-64-jar-with-dependencies.jar \
    net.ladenthin.llama.atmosphere.LocalAgent --model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --workspace /path/to/project

Started without a core jar next to it, the agent stops with NoClassDefFoundError: net/ladenthin/llama/LlamaModel. CI launches exactly this pair on every run (smoke-agent-linux) before anything is published.

Warning

--allow-shell lets the model run any command with your user's rights. By default every write and every command is confirmed on the console ([y]es / [n]o / [a]uto); --auto turns that off. Use a workspace you are willing to hand to the model.

In the REPL, /help lists the commands (/status, /tools, /mode manual|auto, /compact, /clear, /exit); anything else goes to the model. A status line shows the approval mode and the context used ([manual · ctx ~3.1k/16k · 9 tools · local-model]), and the answer is rendered with headings, bullets and code spans.

Also in a browser or an editor. The same agent session has two more front ends. --web serves it to a browser on 127.0.0.1:8787 (Atmosphere's own AI console on an embedded Jetty, a random access token in the printed address; from another machine open an SSH tunnel, ssh -L 8787:127.0.0.1:8787 user@server, rather than binding to the network). --acp speaks the Agent Client Protocol on stdin/stdout, so JetBrains IDEs and Zed — and VS Code through an ACP extension — run it as their chat agent, with the editor's own permission dialog for writes and commands:

java -jar llama-atmosphere-agent-5.2.0-jar-with-dependencies.jar --model model.gguf --allow-shell --web
# JetBrains, ~/.jetbrains/acp.json: {"agent_servers": {"Local llama": {"command": "java",
#   "args": ["-jar", "/path/llama-atmosphere-agent-5.2.0-jar-with-dependencies.jar", "--acp", "--model", "/path/model.gguf"]}}}

On Windows PowerShell quote the whole argument ("-Dexec.args=--model models\… --allow-shell"); for the GPU add e.g. -Dllama.classifier=vulkan-windows-x86-64 and --ngl 99. The agent's README walks through all of it step by step. Other ways to run it, e.g. against a java-llama.cpp server that is already running (--jinja is required for tool calling; the fat jars are GitHub release assets, llama-<version>-all-<os>-<arch>-jar-with-dependencies.jar picks a GPU backend itself):

# 1. the server, e.g. from the release fat jar
java -jar llama-5.2.0-jar-with-dependencies.jar -m models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --jinja --port 8080

# 2. the agent, from the llama-atmosphere-agent/ folder — a you> prompt appears (/clear, /exit)
mvn -q compile exec:java \
    -Dexec.args="--base-url http://127.0.0.1:8080/v1 --workspace /path/to/project --allow-shell"

# a single turn instead of the prompt loop
mvn -q compile exec:java \
    -Dexec.args="--base-url http://127.0.0.1:8080/v1 --workspace /path/to/project --prompt 'Read the README and summarize it'"

# or without a separate server: load the GGUF in-process
mvn -q compile exec:java \
    -Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --workspace /path/to/project"

# everything at once: shell access plus your own system prompt (replaces the built-in one)
mvn -q compile exec:java \
    -Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --ctx-size 16384 --workspace /path/to/project --allow-shell --system 'You are a local assistant on this machine with full shell access. run_command executes any command line, including docker, git and build tools. When asked about the system, run a command instead of explaining it. Answer in the language of the user.'"

The full streaming tool-calling loop (tools → delta.tool_calls → Java tool → role:"tool" result → next turn, over several rounds) is verified on every PR against the real OpenAiCompatServer with no model, and in CI against the Qwen2.5-1.5B tool model — both gate every publish, as does the release-jar smoke above. See llama-atmosphere-agent/README.md for the options and the verified compatibility matrix.

Model/Inference Configuration

There are two sets of parameters you can configure, ModelParameters and InferenceParameters. Both provide builder classes to ease configuration. ModelParameters are once needed for loading a model, InferenceParameters are needed for every inference task. All non-specified options have sensible defaults.

ModelParameters modelParams = new ModelParameters()
        .setModel("/path/to/model.gguf")
        .addLoraAdapter("/path/to/lora/adapter");
String grammar = """
		root  ::= (expr "=" term "\\n")+
		expr  ::= term ([-+*/] term)*
		term  ::= [0-9]""";
InferenceParameters inferParams = new InferenceParameters("")
        .withGrammar(grammar)
        .withTemperature(0.8f);
try (LlamaModel model = new LlamaModel(modelParams)) {
    model.generate(inferParams);
}

Reactive integration (Reactor, RxJava, Kotlin Flow, Akka)

LlamaIterable (returned by model.generate(...) and model.generateChat(...)) implements Iterable<LlamaOutput> & AutoCloseable, so every mainstream reactive library wraps it in a few lines without java-llama.cpp pulling in a runtime reactive dependency.

Always wrap with the library's resource-management primitive — Flux.using, Flowable.using, Kotlin use {}, etc. — so that subscription cancellation flows into LlamaIterable.close() and from there into llama.cpp's native cancelCompletion. A plain Flux.fromIterable(iterable) or for (x in iter) loop will NOT close the iterable on cancel; the native task slot stays occupied until the model is closed.

Project Reactor (Spring WebFlux)

Flux<LlamaOutput> tokens = Flux.using(
        () -> model.generate(params),
        Flux::fromIterable,
        LlamaIterable::close)
    .subscribeOn(Schedulers.boundedElastic());

RxJava 3 (also for RxAndroid)

Flowable<LlamaOutput> tokens = Flowable.using(
        () -> model.generate(params),
        Flowable::fromIterable,
        LlamaIterable::close)
    .subscribeOn(Schedulers.io());

Kotlin Flow (Android / coroutines)

Ready-made: the optional net.ladenthin:llama-kotlin artifact ships generateFlow/generateChatFlow extensions (close-on-cancellation included) plus suspend wrappers whose coroutine cancellation is wired to the binding's cooperative CancellationToken:

model.generateChatFlow(params).flowOn(Dispatchers.IO).collect { print(it.text) }

Hand-rolled equivalent (no extra dependency):

fun llama(model: LlamaModel, params: InferenceParameters) = flow {
    model.generate(params).use { iterable ->
        for (output in iterable) emit(output)
    }
}.flowOn(Dispatchers.IO)

The companion Android sample LLaMAndroid demonstrates the flow { for (output in model.generate(params)) emit(output) } shape against the upstream binding. Wrap the for loop in .use { } if your collector may cancel mid-stream — otherwise the native task slot will not be released until the model is closed.

Akka Streams

val tokens: Source[LlamaOutput, NotUsed] = Source
    .fromIterator(() => model.generate(params).iterator())
    .async("blocking-io-dispatcher")

Why no built-in Publisher? Earlier snapshots of this fork shipped a hand-rolled LlamaModel.streamPublisher(...) returning a Reactive Streams Publisher<LlamaOutput>. Since every reactive library bridges blocking iterables in a few lines via its own resource-management primitive, the binding now stays free of any reactive runtime dependency — pick whichever library your app already uses. The pattern is verified end-to-end by ReactorIntegrationTest in the test sources.

Logging

Per default, llama.cpp writes its log as text to stderr (0.00.035.060 I slot … once a model is loaded): the server's own srv … / slot … lines and, from verbosity 4 on, the llama/ggml lines. All of it can be intercepted via the static method LlamaModel.setLogger(LogFormat, BiConsumer<LogLevel, String>): with a callback set, every line goes to the callback instead of the console (a setLogFile file keeps receiving them). The callback survives model loads, so set it before new LlamaModel(…) to capture the loading lines too. LogFormat.TEXT hands over the bare message, LogFormat.JSON one JSON object per line. Passing null as the callback restores the console output (always llama.cpp's own text format; the format argument only matters with a callback). Logging can be disabled by passing an empty callback. Messages arrive asynchronously from llama.cpp's log worker thread; replacing or removing the logger flushes what is queued to the previous callback first. The verbosity threshold (ModelParameters.setLogVerbosity(int), llama.cpp's -lv: 1 errors, 2 warnings, 3 info, 4 trace, 5 debug) applies before the callback: 2 keeps warnings and errors and silences the per-request INFO lines, which is what a console application sharing the terminal with its own output wants.

// Re-direct log messages however you like (e.g. to a logging library)
LlamaModel.setLogger(LogFormat.TEXT, (level, message) -> System.out.println(level.name() + ": " + message));
// Back to llama.cpp's own console output (stderr)
LlamaModel.setLogger(LogFormat.TEXT, null);
// Disable logging by passing a no-op
LlamaModel.setLogger(LogFormat.TEXT, (level, message) -> {});

The LogLevel enum values passed to the callback correspond to the native llama.cpp log levels:

Value Meaning
DEBUG Verbose diagnostic output
INFO Informational messages about model loading and inference
WARN Non-fatal warnings
ERROR Errors that may affect inference results

Importing in Android

Important

Minimum Android version: API 28 (Android 9.0 Pie). Devices running Android 8.1 (API 27) or earlier are not supported.

Option 1 (recommended): the llama-android AAR from Maven Central

One dependency line in Android Studio — no submodule, no NDK build, no manual ProGuard rules:

dependencies {
    implementation("net.ladenthin:llama-android:5.1.0")
    // or, for Qualcomm Adreno GPUs (device must provide an OpenCL ICD):
    // implementation("net.ladenthin:llama-android-opencl:5.1.0")

    // optional Kotlin coroutines facade (Flow streaming + suspend wrappers):
    implementation("net.ladenthin:llama-kotlin:5.1.0")
}

The AAR carries the full net.ladenthin:llama Java API, the CI-built native libraries for arm64-v8a (devices) and x86_64 (Android Studio emulator, Chromebooks — app bundles split per ABI so phones download only arm64), both 16 KB page-size compliant, consumer R8/ProGuard rules (applied automatically), and a manifest minSdkVersion 28 that AGP enforces against your app. CI boots an x86_64 emulator and runs real on-device inference against every AAR build. Do not also depend on the desktop net.ladenthin:llama JAR in the same app — the AAR already contains those classes, and the JAR would drag ~70 MB of desktop natives into your APK. See llama-android/README.md and llama-kotlin/README.md for details.

Runnable example app — "LLM Service". A minimal, KISS, fully-offline on-device chat app — pick a GGUF from the file system, then chat with it, tokens streaming into a Jetpack Compose UI, with a 13-language flag picker and private local save/load — lives in android-llmservice/ (net.ladenthin.android.llmservice). It builds with plain Gradle/AGP (no Android Studio required), produces a Play-shaped signed .aab, and is validated in CI by a real on-device emulator UI test. See its README for the build, signing/Play, and testing walkthrough.

Option 2 (advanced): build from source inside your app

Use this only if you need to patch the native layer or build for an ABI this project does not ship.

  1. Add java-llama.cpp as a submodule in your an droid app project directory
git submodule add https://github.com/bernardladenthin/java-llama.cpp 
  1. Declare the library as a source in your build.gradle
android {
    val jllamaLib = file("java-llama.cpp")

    // Execute "mvn compile" in the llama/ core module if its target/ doesn't exist
    // (the repository root is the Maven reactor aggregator; the native core lives in llama/).
    if (!file("$jllamaLib/llama/target").exists()) {
        exec {
            commandLine = listOf("mvn", "compile")
            workingDir = file("java-llama.cpp/llama/")
        }
    }

    ...
    defaultConfig {
	...
        externalNativeBuild {
            cmake {
		// Add an flags if needed
                cppFlags += ""
                arguments += ""
            }
        }
    }

    // Declare c++ sources
    externalNativeBuild {
        cmake {
            path = file("$jllamaLib/CMakeLists.txt")
            version = "3.22.1"
        }
    }

    // Declare java sources
    sourceSets {
        named("main") {
            // Add source directory for java-llama.cpp
            java.srcDir("$jllamaLib/src/main/java")
        }
    }
}
  1. Exclude net.ladenthin.llama in proguard-rules.pro
keep class net.ladenthin.llama.** { *; }

TODO

Open work items live in TODO.md.

  • Expand PIT mutation-testing scope. PIT is wired in pom.xml and runs on every CI build (in the test-java-linux-x86_64 job) with <mutationThreshold>100</mutationThreshold>. <targetClasses> currently covers net.ladenthin.llama.value.*, exception.*, args.* and four json parsers (295 mutations, 100% killed, hermetic — no model or fixture needed); widen it incrementally as additional classes reach mutation-test parity. Final target: <param>net.ladenthin.llama.*</param> matching the streambuffer pattern.

Feature Ideas

Forward-looking ideas being tracked for this fork:

  • Adopt feature ideas from the Kotlin Llama Stack client. Candidates (multimodal image input, typed chat messages, async API, batch inference, typed usage/timings) are inventoried with effort estimates in docs/feature-investigation-llama-stack-client-kotlin.md, derived from ogx-ai/llama-stack-client-kotlin.
  • Ship a directly Android-capable artifact — DONE. net.ladenthin:llama-android / llama-android-opencl (AAR, arm64-v8a, minSdk 28, consumer ProGuard rules, 16 KB page-size compliant) plus the optional net.ladenthin:llama-kotlin coroutines façade ship from this repo — see Importing in Android. Typed image input for VLMs is covered by ContentPart.imageBytes(...) / imageFile(...) (see the multimodal section), so downstream Android projects can drop their dependency on ogx-ai/llama-stack-client-kotlin entirely. A dedicated KISS example app — "LLM Service" (SAF model picker + Compose streaming chat, 13-language flag picker, private local save/load, plain Gradle/AGP, signed .aab, on-device emulator UI test) — ships in android-llmservice/.
  • Resolve all upstream kherud/java-llama.cpp open issues. All 37 open issues at fork time are catalogued with per-issue verdicts in docs/history/49be664_open_issues.md; fixes land in this fork as they are completed. Vision inputs (issues #103 and #34) are now wired end to end through blocking, typed, streaming, and OpenAI-compatible request surfaces.

Troubleshooting

Windows: EXCEPTION_ACCESS_VIOLATION with msvcp140.dll

If you encounter a native crash like:

EXCEPTION_ACCESS_VIOLATION (0xc0000005) at pc=0x00007ffa8f4b2f58
C [msvcp140.dll+0x12f58]

This is a known issue where the C++ runtime library (msvcp140.dll) bundled with some JDK versions is outdated.

Solution: Remove the outdated msvcp140.dll from your JDK:

# Locate and remove msvcp140.dll from JDK directory
# Example for JDK 21:
del "C:\Program Files\Java\jdk-21\bin\msvcp140.dll"
del "C:\Program Files\Java\jdk-21\bin\vcruntime140.dll"
del "C:\Program Files\Java\jdk-21\bin\vcruntime140_1.dll"

# Or on Linux with OpenJDK:
rm /usr/lib/jvm/java-21/bin/msvcp140.dll

The system's updated C++ runtime will be used instead, resolving the crash.

Contributors: do not upgrade jqwik past 1.9.3

⚠️ DO NOT UPGRADE jqwik past 1.9.3. jqwik 1.10.0 added an anti-AI prompt-injection string to test stdout; the 1.10.1 user guide states the library "is not meant to be used by any 'AI' coding agents at all." 1.9.3 is the last pre-disclosure release and is the pinned version. See CLAUDE.md section "jqwik prompt-injection in test output" for the full context. Dependabot is configured to ignore all net.jqwik updates (every version, including patches) — see the ignore rule in .github/dependabot.yml.

Similar Projects / Usage

Bindings / wrappers

  • kherud/java-llama.cpp — the upstream Java binding this project was forked from (see the note at the top of this README); development continues independently here, with the fork-time upstream issues catalogued in docs/history/49be664_open_issues.md.
  • llamacpp4j — alternative Java/JNI binding to llama.cpp (SWIG-generated facade); pre-GGUF, dormant since 2023 but historically the other Java JNI option.
  • llama-cpp-python — the Python llama.cpp binding; the de-facto feature benchmark among llama.cpp bindings (server mode, multimodal, speculative decoding).
  • LLamaSharp — C#/.NET llama.cpp binding with per-backend runtime packages (CPU/CUDA/Vulkan/Metal), the .NET analogue of this project's classifier matrix.
  • node-llama-cpp — Node.js/TypeScript llama.cpp binding (prebuilt binaries, JSON-schema-constrained output, function calling).
  • LLaMAndroid — Android app demonstrating usage of llama.cpp bindings.
  • llama-stack-client-kotlin — Kotlin client for the Llama Stack API with an ExecuTorch-backed local-inference path (the llama-android AAR + llama-kotlin façade cover the same on-device ground natively).
  • llama.cpp-android-tutorial — Step-by-step tutorial for running llama.cpp on Android.

Other local inference stacks (no llama.cpp JVM binding)

  • Ollama — llama.cpp-based local model runner with its own HTTP API and model registry. This project's OpenAI-compatible server implements the Ollama-native API surface (/api/version, /api/tags, /api/show, /api/chat, /api/generate), so Ollama-speaking clients (e.g. VS Code Copilot's Ollama provider) work against an in-process jllama model.
  • ExecuTorch — PyTorch's on-device inference runtime (.pte models, XNNPACK/NPU delegates); the engine behind llama-stack-client-kotlin's local mode and the main non-llama.cpp alternative for Android on-device inference (GGUF is not supported there — different model format ecosystem).

Pure-Java single-model inference (no JNI / no llama.cpp) — Alfonso² Peterssen's *.java family of standalone, dependency-free Java inference runtimes, one per model architecture. Useful when JNI is unavailable (e.g. some sandboxes / GraalVM native-image scenarios) or when you want a single jar with no native side at all. Different design point from this project, which prioritises GGUF compatibility and llama.cpp performance via JNI.

Pure-Java inference engines (no JNI / no llama.cpp)

  • Jlama — a full pure-Java LLM inference engine for the JVM (multiple model architectures, quantization, and distributed inference) built on the Java Vector API. A no-native alternative to the JNI approach here; different design point (pure JVM portability vs. GGUF compatibility and llama.cpp performance via JNI).

Frameworks / orchestration

  • LangChain4j — LLM-application framework for Java (chat, embeddings, RAG, tool calling, agents) over a unified provider API. This project ships a first-class in-process integration — see the llama-langchain4j module — so a llama.cpp model plugs straight into LangChain4j's ChatModel / StreamingChatModel / EmbeddingModel / ScoringModel without an HTTP hop.