Note
No IDE or local setup required. This repository is optimized for fully AI-assisted development using Claude Code. No local toolchain, no IDE, nothing to install — everything works completely through Claude.
Security:
Java Bindings for llama.cpp
Forked from kherud/java-llama.cpp: many thanks to @kherud for the great work!
Inference of Meta's LLaMA model (and others) in pure C/C++.
You are welcome to contribute
- Features
- Quick Start
2.1 No Setup required
2.2 Setup required - Documentation
3.1 Example
3.2 Inference
3.3 Chat Completion
3.4 Infilling
3.5 Embeddings & Reranking
3.6 Raw JSON Endpoints
3.7 Local agent: terminal, browser, IDE via ACP - Android
- Feature Ideas
- Text completion (blocking and streaming) with full control over sampling parameters.
- OpenAI-compatible chat completion with automatic chat-template application, including streaming and tool/function calling support via the upstream server.
- Embeddings (single and native-batched via
embed(Collection<String>)) and reranking for retrieval pipelines. - Runtime LoRA adapter control — list the loaded adapters and change their scales at runtime without reloading the model (
getLoraAdapters()/setLoraAdapters(Map)), the typed counterpart of the upstreamGET/POST /lora-adaptersendpoints. - Text-to-speech (
TextToSpeech) over llama.cpp's Qwen3-TTS pipeline (mtmd_helper::gen_audio), returning WAV audio. - In-JVM GGUF quantization (
LlamaQuantizer) over llama.cpp'sllama_model_quantize— convert a GGUF to another quantization scheme without shelling out tollama-quantize. - Infilling (fill-in-the-middle) for code models.
- Tokenize / detokenize and JSON-schema → grammar conversion.
- Raw JSON endpoint handlers mirroring the upstream llama.cpp HTTP server (
/completions,/v1/completions,/embeddings,/infill,/tokenize,/detokenize). - Two runnable HTTP server modes, one fat-jar entry. The fat jar's
Main-ClassisServerLauncher, which dispatches on the--jllama-openai-compatflag. Without it,java -jar …-jar-with-dependencies.jar -m model.gguf --port 8080runs the full upstream llama.cpp server (embedded WebUI, every llama-server flag forwarded) hosted insidelibjllamaover JNI — no separatellama-server.exe. With it,java -jar … --jllama-openai-compat --model model.gguf --port 8080runs the Java-transport, zero-extra-dependency OpenAI-compatible server (OpenAiCompatServer, streaming SSE) instead. Both are also runnable directly by class name viajava -cp … net.ladenthin.llama.server.{NativeServer,OpenAiCompatServer}. - Model metadata access (
getModelMeta()) and server management (metrics, slot save/restore, runtime thread reconfiguration). - Conversation checkpoints —
Session.checkpoint(...)/rewind(...)/fork(...)branch and roll back a chat (KV-cache slot save/restore + transcript snapshot) without re-prefilling. - GGUF metadata inspection without loading the model (
GgufInspector— pure Java, reads header + key/value table only, big-endian aware). - Distributed inference over RPC — offload a model's layers to llama.cpp RPC servers on other machines (
ModelParameters.setRpcServers(...)/--rpc host:port), and serve this machine's devices to them withRpcServer(the in-JVMrpc-server). See Distributed inference over RPC. - Local agent (
llama-atmosphere-agent, release asset, JDK 21+) — a fully offline agent on top of this library that reads and edits files and, if allowed, runs commands. One agent session, three ways to use it: a terminal (full console or line-oriented--plain, fine over SSH/PuTTY), a browser (--web, token-protected, loopback by default — reach it from elsewhere through an SSH tunnel), and IDEs over the Agent Client Protocol (--acp: JetBrains IDEs and Zed natively, VS Code through an ACP extension). See Local agent. - Multi-model router mode (
--models-dir+ per-request model selection, managed via the typedRouterClient) and attach mode (NativeServer(LlamaModel, ...)serves an already-loaded model over the full upstream HTTP frontend — one copy of the weights). - Pre-built native binaries for Linux (x86-64, aarch64, s390x), macOS (arm64, Metal included), Windows (x86-64, x86, arm64) and Android (arm64, x86-64), plus GPU backends (CUDA, Vulkan, OpenCL, ROCm/HIP, SYCL, OpenVINO) — one natives jar each, all loadable side by side with automatic CPU fallback; see Choosing the natives jars. Android additionally ships as the
llama-androidAAR with the optionalllama-kotlincoroutines façade.
Access this library via Maven (released versions on Maven Central). llama-platform brings
the classes plus the CPU natives of every desktop platform:
<dependency>
<groupId>net.ladenthin</groupId>
<artifactId>llama-platform</artifactId>
<version>5.2.0</version>
<type>pom</type>
</dependency>(<type>pom</type> is required in Maven: llama-platform is a dependency list, not a jar.
In Gradle it is just implementation("net.ladenthin:llama-platform:5.2.0").)
Note
This layout starts with 5.2.0. Up to 5.1.0, net.ladenthin:llama was one jar that
carried the CPU natives of every platform, and each GPU classifier was a complete
replacement for it.
There are multiple examples.
Every push to main publishes a snapshot to the Sonatype Central snapshot repository.
To use the latest snapshot, add the repository and dependency to your pom.xml:
<repositories>
<repository>
<id>sonatype-snapshots</id>
<url>https://central.sonatype.com/repository/maven-snapshots/</url>
<snapshots><enabled>true</enabled></snapshots>
<releases><enabled>false</enabled></releases>
</repository>
</repositories>
<dependency>
<groupId>net.ladenthin</groupId>
<artifactId>llama-platform</artifactId>
<version>5.2.0-SNAPSHOT</version>
<type>pom</type>
</dependency>No credentials are required — the repository is publicly readable.
We support CPU inference for the following platforms out of the box:
- Linux x86-64, aarch64, s390x
- macOS aarch64 (Apple silicon, with Metal)
- Windows x86-64, x86, aarch64
- Android aarch64, x86-64 (see Importing in Android)
If any of these match your platform, you can include the Maven dependency and get started.
net.ladenthin:llama is the Java classes only. The native libraries ship as separate jars of
the same artifact, one per backend and platform, selected by a Maven <classifier> of the form
<backend>-<os>-<arch>. Each holds exactly one directory,
net/ladenthin/llama/<OS>/<ARCH>/<backend>/, so any combination can share one classpath.
At startup the loader tries the backends it finds in a fixed order — CUDA → ROCm → SYCL
(fp16, fp32) → Vulkan → OpenCL → OpenVINO → Metal → MSVC → CPU — and uses the first whose
library loads. A GPU jar next to the CPU jar therefore just works on a machine without that GPU
or its runtime: the GPU library fails to load and the CPU one is used.
llama-platform is the classes jar plus the CPU jars of every desktop platform. For GPU
acceleration, add the jar for your GPU next to it:
<dependency>
<groupId>net.ladenthin</groupId>
<artifactId>llama-platform</artifactId>
<version>5.2.0</version>
<type>pom</type>
</dependency>
<!-- Add any natives jar from the table below. Example shown: CUDA 13 on Linux x86-64. -->
<dependency>
<groupId>net.ladenthin</groupId>
<artifactId>llama</artifactId>
<version>5.2.0</version>
<classifier>cuda13-linux-x86-64</classifier>
</dependency>To ship less, depend on net.ladenthin:llama (the classes) plus only the natives jars of the
platforms you target, e.g. cpu-linux-x86-64.
| Classifier | Backend | Target platform | Runtime requirement |
|---|---|---|---|
cpu-linux-x86-64 |
CPU | Linux x86-64 | A JDK 8+ JVM; glibc ≥ 2.17 (manylinux2014). |
cpu-linux-aarch64 |
CPU | Linux aarch64 | glibc ≥ 2.39 (e.g. Ubuntu 24.04+, Debian 13+) — built natively on ubuntu-24.04-arm, matching upstream llama.cpp's own ARM binaries; older-glibc ARM hosts (Ubuntu 22.04, Debian 12, RHEL 8/9, Amazon Linux 2023) are not supported. |
cpu-linux-s390x |
CPU | Linux s390x (IBM Z, big-endian) | A JDK 8+ JVM. |
cpu-windows-x86-64 / cpu-windows-x86 |
CPU | Windows x86-64 / x86 | A JDK 8+ JVM. Built with Ninja Multi-Config + MSVC (static /MT CRT). |
cpu-windows-aarch64 |
CPU | Windows on ARM (Snapdragon X / Surface) | A JDK 8+ JVM. Built natively on windows-11-arm with clang-cl. |
metal-macos-aarch64 |
Metal + CPU | macOS aarch64 (Apple silicon) | A JDK 8+ JVM. |
cpu-android-aarch64 / cpu-android-x86-64 |
CPU | Android | For Android use the llama-android AAR; these jars are the same libraries for other Android JVM setups. |
msvc-windows-x86-64 / msvc-windows-x86 |
CPU (Visual Studio generator) | Windows x86-64 / x86 | Same CPU backend and MSVC toolchain as cpu-windows-*, built with the Visual Studio generator instead of Ninja — an alternate-toolchain option; tried before the cpu jar when both are present. |
cuda13-linux-x86-64 |
CUDA 13 | Linux x86-64 with NVIDIA GPU | NVIDIA driver + CUDA 13 runtime libraries (libcudart.so.13, libcublas.so.13). |
cuda13-windows-x86-64 |
CUDA 13 | Windows x86-64 with NVIDIA GPU | NVIDIA driver + CUDA 13 Toolkit (cudart64_13.dll, cublas64_13.dll, cublasLt64_13.dll on PATH). |
vulkan-linux-x86-64 |
Vulkan | Linux x86-64 with a Vulkan 1.2+ GPU (NVIDIA / AMD / Intel) | A Vulkan runtime (libvulkan.so.1), which current GPU drivers install. The most portable Linux GPU option. glibc ≈ 2.39 (built on ubuntu-latest). |
vulkan-linux-aarch64 |
Vulkan | Linux aarch64 with a Vulkan 1.2+ GPU | A Vulkan runtime (libvulkan.so.1). glibc ≥ 2.39. |
vulkan-windows-x86-64 |
Vulkan | Windows x86-64 with a Vulkan 1.2+ GPU | A Vulkan runtime (vulkan-1.dll), which current GPU drivers install. The most portable Windows GPU option. |
opencl-windows-x86-64 |
OpenCL | Windows x86-64 with an OpenCL 2.0+ GPU | A vendor OpenCL ICD (OpenCL.dll). The GGML OpenCL backend is Adreno-tuned; on desktop GPUs CUDA or Vulkan are better supported. |
opencl-windows-aarch64 |
OpenCL (Adreno) | Windows on ARM (Snapdragon X) | The Adreno driver's OpenCL ICD (OpenCL.dll). |
opencl-android-aarch64 |
OpenCL (Adreno) | Android aarch64 with Adreno GPU | A device OpenCL ICD (libOpenCL.so); see also the llama-android-opencl AAR. |
rocm-linux-x86-64 |
ROCm / HIP | Linux x86-64 with AMD GPU | An AMD ROCm 10 runtime (libamdhip64.so, librocblas.so, libhipblas.so) — built against ROCm 10.0 (TheRock), like upstream llama.cpp; every GPU TheRock builds for Linux, Instinct included (gfx900/gfx906/gfx90c/gfx1153 best effort — built, but not release-ready in ROCm 10). |
rocm-windows-x86-64 |
ROCm / HIP | Windows x86-64 with AMD GPU | The AMD ROCm 10 runtime DLLs (amdhip64.dll, rocblas.dll, hipblas.dll) on PATH; every Radeon target TheRock builds for Windows, gfx900 through RDNA4 (gfx900/gfx906/gfx90c/gfx1153 best effort). |
sycl-fp16-linux-x86-64 |
SYCL (Intel oneAPI, fp16) | Linux x86-64 with Intel GPU (Arc / iGPU) | An Intel oneAPI / Level-Zero runtime. fp16 accumulation (faster, slightly lower precision). |
sycl-fp32-linux-x86-64 |
SYCL (Intel oneAPI, fp32) | Linux x86-64 with Intel GPU (Arc / iGPU) | An Intel oneAPI / Level-Zero runtime. fp32 accumulation (higher precision). |
sycl-windows-x86-64 |
SYCL (Intel oneAPI) | Windows x86-64 with Intel GPU (Arc / iGPU) | The Intel oneAPI / Level-Zero runtime DLLs on PATH. |
openvino-linux-x86-64 |
OpenVINO | Linux x86-64 (Intel GPU / NPU / CPU) | An Intel OpenVINO runtime. |
openvino-windows-x86-64 |
OpenVINO | Windows x86-64 (Intel GPU / NPU / CPU) | The Intel OpenVINO runtime DLLs on PATH. |
Note
No vendor runtime is bundled; it comes from the GPU driver or toolkit on the host. The GPU
jars are validated build-only in CI (GitHub runners have no GPU), so end-to-end GPU
inference is verified locally / on self-hosted hardware. A GPU library that loads but finds
no usable device (e.g. CUDA installed, no NVIDIA GPU) runs on the CPU inside that library and
keeps the loader from trying the next backend; force one with
-Dnet.ladenthin.llama.backend=<backend> (e.g. vulkan or cpu), which fails loud instead
of falling back.
Note
On the module path each natives jar is an automatic module
(net.ladenthin.llama.natives.<classifier>, with _ for -) that nothing requires, so
resolve them with --add-modules ALL-MODULE-PATH (or keep the natives jars on the
classpath).
Note
Android armeabi-v7a (32-bit ARM) is not published. Only 64-bit
Android binaries are shipped: aarch64 (devices) and x86_64
(emulators, Chromebooks, x86-64 Android hardware) as the cpu-android-*
natives jars and in the llama-android AAR, plus aarch64 as
opencl-android-aarch64. 32-bit Android devices are unsupported
by the released artifacts.
The minimum required Android version is API 28 (Android 9.0 Pie). Devices running Android 8.1 (API 27) or earlier are not supported.
For running the embedded server
without any Maven setup, every tagged GitHub Release
(and the rolling snapshot pre-release from main) attaches self-contained
all-backends fat jars — one download per OS/arch, runnable directly:
java -jar llama-<version>-all-linux-x86-64-jar-with-dependencies.jar -m model.gguf --port 8080| Release asset | Bundled GPU backends | CPU fallback |
|---|---|---|
llama-<version>-all-linux-x86-64-jar-with-dependencies.jar |
CUDA 13, ROCm, SYCL fp16/fp32, Vulkan, OpenVINO | yes |
llama-<version>-all-linux-aarch64-jar-with-dependencies.jar |
Vulkan | yes |
llama-<version>-all-windows-x86-64-jar-with-dependencies.jar |
CUDA 13, ROCm, SYCL, Vulkan, OpenCL, OpenVINO | yes |
llama-<version>-all-windows-aarch64-jar-with-dependencies.jar |
OpenCL (Adreno / Snapdragon X) | yes |
llama-<version>-jar-with-dependencies.jar |
none (CPU only, incl. macOS Metal) | — |
Each all-backends jar is the classes, all Java runtime dependencies, and every natives jar
of its named OS/arch (except msvc) merged into one file — pick the jar that matches
your platform. (A 32-bit JVM on 64-bit Windows, for example, needs the default
llama-<version>-jar-with-dependencies.jar, which carries the CPU natives of every platform.)
The loader picks the backend exactly as described above, so the
jar starts on every host of its OS/arch, with or without a GPU. A .sha256 checksum file and a
detached GPG .asc signature accompany every jar. These fat jars are GitHub download assets
only — they are not published to Maven Central.
If none of the above listed platforms matches yours, currently you have to compile the library yourself (also if you want GPU acceleration).
This consists of two steps: 1) Compiling the libraries and 2) putting them in the right location.
First, have a look at llama.cpp to know which build arguments to use (e.g. for CUDA support).
Any build option of llama.cpp works equivalently for this project.
You then have to run the following commands in the llama/ module directory (the native core lives
there; the repository root is just the Maven reactor aggregator):
cd llama # the native core module
mvn compile # don't forget this line
cmake -B build # add any other arguments for your backend, e.g. -DGGML_CUDA=ON
cmake --build build --config ReleaseTip
Use -DLLAMA_CURL=ON to download models via Java code using ModelParameters#setModelUrl(String).
The library is put in a directory matching your platform and backend, which appears in the cmake output. For example:
-- Backend 'cpu' - installing files to /java-llama.cpp/llama/src/main/natives/net/ladenthin/llama/Linux/x86_64/cpumvn test puts that directory on the test classpath; mvn -P natives package turns it into a
natives jar (in CI, together with every other platform's).
This project has to load a single shared library jllama.
Note, that the file name varies between operating systems, e.g., jllama.dll on Windows, jllama.so on Linux, and jllama.dylib on macOS.
The application will search in the following order in the following locations:
- In net.ladenthin.llama.lib.path: Use this option if you want a custom location for your shared libraries, i.e., set VM option
-Dnet.ladenthin.llama.lib.path=/path/to/directory. - In java.library.path: These are predefined locations for each OS, e.g.,
/usr/java/packages/lib:/usr/lib64:/lib64:/lib:/usr/libon Linux. You can find out the locations usingSystem.out.println(System.getProperty("java.library.path")). Use this option if you want to install the shared libraries as system libraries. - From the natives jars on the classpath: every backend directory found for your platform, in the order described in Choosing the natives jars.
Every net.ladenthin.llama.* system property recognised by the library, deep-scanned from the source. Runtime properties are resolved through LlamaSystemProperties; test-only properties are declared in the test sources (TestConstants) and consumed by individual test classes.
| Property | Default | Scope | Consumer | Description |
|---|---|---|---|---|
net.ladenthin.llama.lib.path |
unset (falls back to java.library.path) |
runtime | LlamaLoader |
Directory containing the native jllama shared library. Checked first, before java.library.path. Set with -Dnet.ladenthin.llama.lib.path=/path/to/dir. |
net.ladenthin.llama.tmpdir |
unset (falls back to java.io.tmpdir) |
runtime | LlamaLoader |
Custom temporary directory used when extracting the native library from the JAR. |
net.ladenthin.llama.osinfo.architecture |
unset (uses os.arch) |
runtime | OSInfo |
Override for the architecture string used to locate the bundled library inside the JAR. Useful when os.arch reports an unexpected value (e.g. inside dockcross / chrooted environments). |
net.ladenthin.llama.backend |
unset (auto: the first backend on the classpath whose library loads) | runtime | LlamaLoader |
Names one backend directory (e.g. cuda13, vulkan, cpu) to load exclusively — failure is then fatal instead of trying the next backend. See Choosing the natives jars. |
net.ladenthin.llama.test.ngl |
43 for the general suite; 0 for ToolCallingIntegrationTest |
test | Model-backed integration tests | Number of GPU layers used during testing. Pin to 0 on CPU-only hosts: mvn test -Dnet.ladenthin.llama.test.ngl=0. The tool test also selects device none at zero layers so Metal/CUDA is not initialized. |
net.ladenthin.llama.tool.model |
models/Qwen2.5-1.5B-Instruct-Q4_K_M.gguf (test self-skips if missing) |
test | ToolCallingIntegrationTest |
Path to a tool-capable GGUF used to verify required blocking and streaming tool calls. The default matches the Qwen2.5 model in upstream llama.cpp's tool-call test matrix. |
net.ladenthin.llama.nomic.path |
models/nomic-embed-text-v1.5.f16.gguf (test self-skips if missing) |
test | LlamaEmbeddingsTest#testNomicEmbedLoads |
Path to a Nomic embedding model (nomic-embed-text-v1.5.f16.gguf or a compatible BERT-family encoder). Regression test for upstream issue #98 (BERT-encoder result_output assertion). |
net.ladenthin.llama.vision.model |
models/SmolVLM-500M-Instruct-Q8_0.gguf (test self-skips if missing) |
test | MultimodalIntegrationTest |
Path to a vision-capable model GGUF. Any vision-capable GGUF works; CI default is SmolVLM-500M-Instruct-Q8_0.gguf. |
net.ladenthin.llama.vision.mmproj |
models/mmproj-SmolVLM-500M-Instruct-Q8_0.gguf (test self-skips if missing) |
test | MultimodalIntegrationTest |
Matching mmproj GGUF for the vision model. |
net.ladenthin.llama.vision.image |
llama/src/test/resources/images/test-image.jpg (a CC-BY-4.0 / MIT-granted photo committed to the repo) |
test | MultimodalIntegrationTest |
Visual prompt image. Any png/jpeg/webp/gif works; the extension drives MIME detection. |
net.ladenthin.llama.audio.model |
unset (test self-skips) | test | AudioInputIntegrationTest (llama.cpp discussion #13759) |
Path to an audio-input model GGUF (e.g. Ultravox, Qwen2.5-Omni). |
net.ladenthin.llama.audio.mmproj |
unset (test self-skips) | test | AudioInputIntegrationTest |
Matching audio mmproj (encoder) GGUF. |
net.ladenthin.llama.audio.input |
src/test/resources/audios/sample.wav (committed) |
test | AudioInputIntegrationTest |
.wav/.mp3 audio prompt clip; the extension drives format detection. |
net.ladenthin.llama.tts.model |
models/Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf (test self-skips if missing) |
test | TtsIntegrationTest |
Path to the Qwen3-TTS backbone (text) GGUF. Any Qwen3-TTS-family model works; CI default is Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf. |
net.ladenthin.llama.tts.mmproj |
models/mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf (test self-skips if missing) |
test | TtsIntegrationTest |
Path to the matching Qwen3-TTS mmproj GGUF (speaker encoder + code predictor + code2wav decoder); CI default is mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf. |
net.ladenthin.llama.train.model |
models/stories260K.gguf (test self-skips if missing) |
test | LlamaTrainerIntegrationTest |
Path to the model the fine-tuning smoke trains. Must be F32: llama_set_param skips every other tensor type, so a quantized model would train nothing. |
The test-model defaults are exactly CI's model set (.github/models.csv, URLs included), so a model downloaded into models/ is found without any property; the properties only point a test at another file. MultimodalIntegrationTest self-skips when any of the three vision.* paths is missing, so a partial setup (just the vision model + the committed image, no mmproj) lets the test class load without erroring. AudioInputIntegrationTest self-skips the same way over the three audio.* properties. TtsIntegrationTest likewise self-skips unless both tts.model and tts.mmproj paths exist, and LlamaTrainerIntegrationTest unless train.model (default models/stories260K.gguf, an F32 model) does.
This is a short example on how to use this library:
public class Example {
public static void main(String... args) throws IOException {
ModelParameters modelParams = new ModelParameters()
.setModel("models/mistral-7b-instruct-v0.2.Q2_K.gguf")
.setGpuLayers(43);
String system = "This is a conversation between User and Llama, a friendly chatbot.\n" +
"Llama is helpful, kind, honest, good at writing, and never fails to answer any " +
"requests immediately and with precision.\n";
BufferedReader reader = new BufferedReader(new InputStreamReader(System.in, StandardCharsets.UTF_8));
try (LlamaModel model = new LlamaModel(modelParams)) {
System.out.print(system);
String prompt = system;
while (true) {
prompt += "\nUser: ";
System.out.print("\nUser: ");
String input = reader.readLine();
prompt += input;
System.out.print("Llama: ");
prompt += "\nLlama: ";
InferenceParameters inferParams = new InferenceParameters(prompt)
.withTemperature(0.7f)
.withMiroStat(MiroStat.V2)
.withStopStrings("User:");
for (LlamaOutput output : model.generate(inferParams)) {
System.out.print(output);
prompt += output;
}
}
}
}
}Also have a look at the other examples.
There are multiple inference tasks. In general, LlamaModel is stateless, i.e., you have to append the output of the
model to your prompt in order to extend the context. If there is repeated content, however, the library will internally
cache this, to improve performance.
ModelParameters modelParams = new ModelParameters().setModel("/path/to/model.gguf");
InferenceParameters inferParams = new InferenceParameters("Tell me a joke.");
try (LlamaModel model = new LlamaModel(modelParams)) {
// Stream a response and access more information about each output.
for (LlamaOutput output : model.generate(inferParams)) {
System.out.print(output);
}
// Calculate a whole response before returning it.
String response = model.complete(inferParams);
// Returns the hidden representation of the context + prompt.
float[] embedding = model.embed("Embed this");
}Note
Since llama.cpp allocates memory that can't be garbage collected by the JVM, LlamaModel is implemented as an
AutoClosable. If you use the objects with try-with blocks like the examples, the memory will be automatically
freed when the model is no longer needed. This isn't strictly required, but avoids memory leaks if you use different
models throughout the lifecycle of your application.
For chat models, build a list of role/content pairs and let the library apply the model's chat template.
chatComplete() returns the full response, generateChat() streams tokens, and chatCompleteText() returns
just the text content of the assistant message.
List<Pair<String, String>> messages = new ArrayList<>();
messages.add(new Pair<>("user", "Write a haiku about Java."));
InferenceParameters inferParams =
new InferenceParameters("").withMessages("You are a helpful assistant.", messages);
try (LlamaModel model = new LlamaModel(modelParams)) {
// Streaming
for (LlamaOutput output : model.generateChat(inferParams)) {
System.out.print(output);
}
// Or blocking, returns the OpenAI-compatible JSON envelope
String json = model.chatComplete(inferParams);
// Or just the assistant text
String text = model.chatCompleteText(inferParams);
}Reasoning/thinking models can receive custom Jinja template variables via
ModelParameters#setChatTemplateKwargs(Map).
Load a vision-capable GGUF with its matching projector, then place text and image parts in the same user message. Images may come from a file, raw bytes, a data URI, or an HTTP(S) URL:
ModelParameters modelParams = new ModelParameters()
.setModel("models/SmolVLM-500M-Instruct-Q8_0.gguf")
.setMmproj("models/mmproj-SmolVLM-500M-Instruct-Q8_0.gguf");
ChatMessage message = ChatMessage.userMultimodal(
ContentPart.text("Describe this image in one short sentence."),
ContentPart.imageFile(Paths.get("photo.jpg")));
try (LlamaModel model = new LlamaModel(modelParams)) {
String answer = model.chatCompleteText(InferenceParameters.empty()
.withMessages(Collections.singletonList(message))
.withNPredict(64));
System.out.println(answer);
}The same multipart messages[].content shape works through ChatRequest and the embedded
OpenAI-compatible /v1/chat/completions server. For a strictly CPU-only run, use
setDevices("none").setMmprojOffload(false) in addition to setGpuLayers(0); projector offload
has its own upstream default.
On a multi-GPU host the projector can be placed independently of the weights with
setMmprojDevice("CUDA1") (llama.cpp --mmproj-device, added upstream in b10541). Exactly one
device may be named; the literal "none" keeps the projector on the CPU. OpenAiCompatServer's CLI
accepts the same flag as -mmdev/--mmproj-device, and NativeServer forwards it verbatim like
every other llama-server flag.
setMmprojDevice(...) and setMmprojOffload(...) write the same upstream field
(common_params::mmproj_use_gpu), so where they disagree the outcome would depend on argv order —
and the rendered argv comes from a HashMap, whose order is unspecified. The builder therefore
resolves the two genuinely ambiguous combinations by dropping the earlier call, and leaves the rest
alone:
| Combination | Resolves to | Builder behaviour |
|---|---|---|
named device + setMmprojOffload(true) |
(use_gpu=true, device) in either order |
both kept — no clash |
named device + setMmprojOffload(false) |
order-dependent | last call wins |
"none" + setMmprojOffload(true) |
order-dependent | last call wins |
"none" + setMmprojOffload(false) |
(use_gpu=false) in either order |
both kept — no clash |
So a multi-GPU projector pin survives an explicit setMmprojOffload(true); only a call that would
actually contradict the other is dropped. If you need a device after disabling offload, call
setMmprojDevice last.
Video input — decode settings only, so far. mtmd has carried a video path since llama.cpp
b9562 (#24269); b10647 (#24318) added the --video-* CLI flags and the
mtmd_helper_init_opt plumbing that surfaces them. It is compiled into the shipped desktop library
(MTMD_VIDEO is on by default, gated on LLAMA_SUBPROCESS, which upstream force-disables on
Android and iOS). Its decode settings are exposed
as setVideoFps(float), setVideoTimestampInterval(long) and setVideoFfmpegDir(String). The last
one matters most in a JVM: upstream shells out to ffmpeg/ffprobe and resolves them from PATH,
which an application server, an Android app or a JAR-only container frequently does not have them on —
naming the directory is then the only way for video to work at all.
What is not here yet is the content part: upstream's wire type for a video is
{"type":"input_video","input_video":{"data":"<base64>"}} (raw base64, not a data: URI, unlike
image_url), gated server-side on mtmd_helper_support_video. ContentPart has no videoFile(...)
factory emitting that shape, so these knobs currently configure a path this API cannot yet feed
directly. Tracked in TODO.md.
Audio input works identically — load an audio-capable model (Ultravox, Qwen2.5-Omni, …) with its
audio --mmproj and add a ContentPart.audioFile(...) (or inputAudio(bytes, "wav"|"mp3")) part. It
serializes to the OpenAI input_audio content part and routes through the same mtmd pipeline:
ModelParameters modelParams = new ModelParameters()
.setModel("models/ultravox-v0_5-llama-3_2-1b.gguf")
.setMmproj("models/mmproj-ultravox-v0_5-llama-3_2-1b-f16.gguf");
ChatMessage message = ChatMessage.userMultimodal(
ContentPart.text("Transcribe the audio."),
ContentPart.audioFile(Paths.get("speech.wav")));
try (LlamaModel model = new LlamaModel(modelParams)) {
System.out.println(model.supportsAudio()); // true
String answer = model.chatCompleteText(InferenceParameters.empty()
.withMessages(Collections.singletonList(message))
.withNPredict(64));
System.out.println(answer);
}LlamaModel.supportsVision() / supportsAudio() report which modalities the loaded projector enables.
Use a tool-aware instruct model and enable Jinja when loading it. A typed request can either return
the model's tool calls through chat, or execute registered handlers until the model produces a
normal assistant response through chatWithTools:
ToolDefinition weather = new ToolDefinition(
"get_weather",
"Get the current weather for a city",
"{\"type\":\"object\",\"properties\":{\"city\":{\"type\":\"string\"}},"
+ "\"required\":[\"city\"]}");
ChatRequest request = ChatRequest.empty()
.appendMessage("user", "What is the weather in Paris?")
.appendTool(weather)
.withToolChoice("auto")
.withParallelToolCalls(Boolean.FALSE);
Map<String, ToolHandler> handlers = Collections.singletonMap(
"get_weather", argumentsJson -> "{\"temperature_c\":21,\"condition\":\"sunny\"}");
try (LlamaModel model = new LlamaModel(new ModelParameters()
.setModel("models/Qwen2.5-1.5B-Instruct-Q4_K_M.gguf")
.enableJinja())) {
ChatResponse response = model.chatWithTools(request, handlers);
System.out.println(response.getFirstContent());
}tool_choice is the OpenAI-compatible string form (auto, none, or required). Set
parallel_tool_calls to false when handlers should be issued one at a time. Handler failures and
unknown tool names are returned to the model as valid {"error":"..."} tool-result JSON.
You can simply set InferenceParameters#withInputPrefix(String) and InferenceParameters#withInputSuffix(String).
Load the model with enableEmbedding() (or enableReranking()) and call embed(String) to get a sentence
embedding, or rerank(query, documents...) to get relevance scores.
ModelParameters modelParams = new ModelParameters()
.setModel("/path/to/embedding-model.gguf")
.enableEmbedding();
try (LlamaModel model = new LlamaModel(modelParams)) {
float[] embedding = model.embed("Embed this sentence");
// Batch form: one native dispatch for many inputs, results in request order.
List<float[]> embeddings = model.embed(Arrays.asList("First sentence", "Second sentence"));
}Adapters loaded at model-load time (addLoraAdapter(...) / addLoraScaledAdapter(...), optionally
setLoraInitWithoutApply() to start disabled) can be listed and re-scaled at runtime without
reloading the model — the typed counterpart of the upstream GET/POST /lora-adapters endpoints:
ModelParameters modelParams = new ModelParameters()
.setModel("models/base.gguf")
.addLoraScaledAdapter("models/adapter.gguf", 1.0f);
try (LlamaModel model = new LlamaModel(modelParams)) {
List<LoraAdapter> adapters = model.getLoraAdapters(); // [{id=0, path=..., scale=1.0}]
model.setLoraAdapter(0, 0.5f); // re-scale at runtime
model.setLoraAdapters(Collections.emptyMap()); // disable all adapters
}Per the upstream contract, a scale update lists the adapters to keep active — any adapter missing
from the map is set to scale 0 (disabled). The native side clears affected KV caches when the
effective adapter set changes.
TextToSpeech synthesizes audio from text over llama.cpp's upstream Qwen3-TTS pipeline
(mtmd_helper::gen_audio). It is a separate AutoCloseable native type (not a LlamaModel)
because TTS loads its own model pair: a backbone (text) GGUF and an mmproj GGUF bundling the
speaker encoder, code predictor, and code2wav decoder. synthesize(String) returns a
24 kHz mono 16-bit WAV byte stream.
try (TextToSpeech tts = new TextToSpeech(
"models/qwen3-tts-backbone.gguf", "models/qwen3-tts-mmproj.gguf")) {
byte[] wav = tts.synthesize("Hello from llama dot c p p.");
Files.write(Paths.get("out.wav"), wav);
}Add (modelPath, mmprojPath, gpuLayers, threads) to offload to the GPU, or
synthesize(text, maxFrames, topK, seed) for explicit sampling, or the full
synthesize(text, speakerReferenceAudioPath, language, maxFrames, topK, seed) overload for
voice cloning from a reference clip. As with LlamaModel, native memory is not GC-managed — use
try-with-resources or call close().
This replaces the project's earlier two-model OuteTTS + WavTokenizer pipeline, which upstream
#26254 deleted entirely in favor of Qwen3-TTS
(llama.cpp b10270); there is no backward-compatible path for the old model pair.
LlamaQuantizer converts a GGUF to another quantization scheme in-process (llama.cpp's
llama_model_quantize — the llama-quantize tool without the separate binary):
LlamaQuantizer.quantize("model-f16.gguf", "model-q4_k_m.gguf", QuantizationType.Q4_K_M);
// Re-quantizing an already-quantized GGUF degrades quality and must be opted into:
LlamaQuantizer.quantize("model-q8_0.gguf", "model-q4_0.gguf", QuantizationType.Q4_0,
/* threads */ 0, /* allowRequantize */ true);For direct access to the upstream llama.cpp server API, the following methods take a JSON request and return a JSON response, matching the HTTP server's contract:
handleCompletions, handleCompletionsOai, handleChatCompletions, handleInfill,
handleEmbeddings, handleTokenize, handleDetokenize.
Server state is exposed via getMetrics(), eraseSlot(int), saveSlot(int, String),
restoreSlot(int, String), and getModelMeta().
A Session can be snapshotted and branched — the KV-cache slot state and the transcript move
together, so native state and history can never drift apart:
try (Session session = new Session(model, 0, "You are terse.")) {
session.send("My name is Alice.");
SessionCheckpoint cp = session.checkpoint("checkpoints/turn1.bin");
session.send("Tell me a joke.");
session.rewind(cp); // undo everything after the checkpoint
session.send("Tell me a story instead."); // retry from the branch point
// Branch into a second slot (model loaded with setParallel(2)+):
try (Session forked = session.fork(1, "checkpoints/branch.bin")) {
forked.send("Answer as a pirate."); // both sessions continue independently
}
}Checkpoint files are caller-managed (KV dumps grow with context usage) and both operations are
rejected while a stream is in progress. For plain transformer models a rewind is also achievable
cheaply by resending a truncated history with cache_prompt (prefix reuse); checkpoints make the
branch point exact and are the only reliable rollback for recurrent/hybrid models (e.g.
Granite-4), whose state cannot be recomputed from a prefix.
GgufInspector reads a GGUF's header and key/value table without loading the model — pure
Java, no native library, cost independent of file size (parsing stops before the tensor data).
Useful for model pickers and download validators:
GgufMetadata meta = GgufInspector.read(Paths.get("models/Qwen3-0.6B-Q4_K_M.gguf"));
meta.getArchitecture(); // Optional[qwen3]
meta.getModelName(); // Optional[Qwen3 0.6B]
meta.getParameterCount(); // OptionalLong[751632384]
meta.getContextLength(); // OptionalLong[40960] (<arch>.context_length)
meta.getFileType(); // OptionalLong[15] (llama_ftype, cf. QuantizationType)
meta.getChatTemplate(); // Optional[{{- ... }}]
meta.getEntries(); // full decoded key/value tableSupports GGUF v2/v3, little- and big-endian (auto-detected), and fails loud on v1/corrupt files.
For metadata of an already-loaded model use getModelMeta() instead.
Prompt-prefix reuse is enabled by default in llama.cpp and can be controlled per request with
InferenceParameters.withCachePrompt(boolean). withCacheReuse(int) enables non-prefix chunk reuse,
while withSlotId(int) pins a request to a specific server slot. Session applies its slot id to every
request, so generation and save/restore operate on the same KV state.
Typed results expose logical prompt, generated, cached prompt, and evaluated prompt counts through
Usage. Per-request timing also remains available through Timings.getCacheN().
LlamaModel.getMetricsTyped().getSlotMetrics() reports each slot's logical, processed, cached,
decoded, and remaining token counts, and the same ServerMetrics view carries the server-wide
lifetime counters — including cached prompt tokens (getCumulativeCachedPromptTokens()) and the
speculative-decoding tallies (getDraftTokensTotal(), getDraftAcceptedTotal(),
getDraftVerifyStepsTotal(), getDraftAcceptedPerPosition(), plus the derived
getDraftAcceptanceRate()), which upstream otherwise exposes only as Prometheus text.
The embedded HTTP server exposes the same native JSON at authenticated GET /metrics, with the slot
array alone at GET /slots. OpenAI responses preserve
usage.prompt_tokens_details.cached_tokens; Responses API output uses
usage.input_tokens_details.cached_tokens; Anthropic output uses cache_read_input_tokens.
net.ladenthin.llama.server.OpenAiCompatServer turns a loaded model into a local
OpenAI-compatible HTTP endpoint using only the JDK's built-in com.sun.net.httpserver — no extra
dependency and no separate server process. It is embeddable, and runnable via
java -cp <jar> net.ladenthin.llama.server.OpenAiCompatServer … (the fat jar's default
Main-Class is instead NativeServer — see "Native server with the built-in WebUI" below). It
serves:
| Method & path | Backed by |
|---|---|
POST /v1/chat/completions |
LlamaModel.streamChatCompletion (streaming SSE) / chatComplete (blocking) |
POST /v1/completions |
LlamaModel.handleCompletionsOai |
POST /v1/embeddings (requires --embedding) |
LlamaModel.handleEmbeddings |
POST /v1/rerank (requires --reranking) |
LlamaModel.handleRerank (reshaped to results/data) |
POST /infill |
LlamaModel.handleInfill (fill-in-the-middle autocomplete) |
GET /v1/models |
the configured model id |
GET /metrics |
native server and per-slot token/cache counters (JSON) |
GET /slots |
native per-slot token/cache counters (JSON array) |
GET /health |
static {"status":"ok"} (unauthenticated) |
Chat completions support streaming via Server-Sent Events and non-streaming, forwarding
messages/tools verbatim. The streaming path carries delta.tool_calls and (with
stream_options.include_usage) a trailing usage chunk, so agent/tool-calling clients work —
this is the recommended surface for VS Code Copilot agent mode, Cline, Roo Code and Continue.
response_format (json_object / json_schema) is forwarded for structured outputs. Completions,
embeddings, rerank and infill are non-streaming.
Every route is also reachable without the /v1 prefix, the server answers CORS preflight
(OPTIONS) and stamps Access-Control-Allow-Origin (so browser/webview clients work), and
POST /infill is the llama.cpp-native FIM endpoint for local ghost-text autocomplete plugins
(llama.vscode, Twinny, Tabby, Continue's llama.cpp provider). Note: GitHub Copilot's inline
completions cannot be served by any local endpoint — only its chat/agent surfaces — so use one of
those autocomplete plugins for ghost text.
Alternative protocol surfaces. For clients that don't speak OpenAI Chat Completions, the same model is exposed through additional protocols (pure translation over the OpenAI core — no extra inference path), all supporting tools and streaming:
| Surface | Routes | For |
|---|---|---|
| Ollama-native | GET /api/version, GET /api/tags, POST /api/show, POST /api/chat (NDJSON streaming), POST /api/generate (prompt completion / FIM) |
Copilot's built-in Ollama provider; Ollama-hardcoded tools |
| Anthropic Messages | POST /v1/messages (SSE event stream) |
Claude-shaped clients (Claude Code); Copilot messages apiType |
| OpenAI Responses | POST /v1/responses (SSE event stream) |
Copilot responses apiType; Responses-API clients |
/api/show advertises the model's capabilities (tools, insert, and vision when --mmproj is set)
and context length, which Copilot's Ollama provider reads to enable agent mode. The llama.cpp-native
GET /props reports default_generation_settings.n_ctx and a modalities block, which autocomplete
clients such as llama.vscode read to size their context window.
Embed it in your app:
ModelParameters modelParams = new ModelParameters().setModel("models/model.gguf").setParallel(2);
OpenAiServerConfig config = OpenAiServerConfig.builder().port(8080).modelId("local-model").build();
try (LlamaModel model = new LlamaModel(modelParams);
OpenAiCompatServer server = new OpenAiCompatServer(model, config).start()) {
Thread.currentThread().join(); // serve until interrupted
}…or run it standalone. The fat jar's Main-Class is the ServerLauncher dispatcher, so add
--jllama-openai-compat to select this Java server (the launcher strips that flag and forwards the rest);
or name the class explicitly via -cp:
# fat jar (bundles the native lib + Java deps) — select the Java server with --jllama-openai-compat
java -jar target/llama-<version>-jar-with-dependencies.jar --jllama-openai-compat \
--model models/Qwen3-0.6B-Q4_K_M.gguf --host 0.0.0.0 --port 8080 --n-gpu-layers 99
# or name the class explicitly (fat jar or plain library jar)
java -cp target/llama-<version>.jar net.ladenthin.llama.server.OpenAiCompatServer \
--model models/model.gguf --port 8080 --model-id local-modelRun with --help for the full option list (-m/--model, --host, -p/--port, -c/--ctx-size,
-b/--batch-size, -ub/--ubatch-size, -ngl/--n-gpu-layers, -t/--threads, -tb/--threads-batch,
-ctk/--cache-type-k, -ctv/--cache-type-v, --jinja, --chat-template-kwargs, --parallel,
--model-id, --api-key, --mmproj, -mmdev/--mmproj-device, --embedding, --reranking). The tuning flags mirror
llama.cpp's server, so an invocation like
--jinja --chat-template-kwargs '{"reasoning_effort":"low"}' -ctk q8_0 -ctv q8_0 -b 4096 -ub 2048
works directly.
Verify with curl (streaming chat):
curl -N http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"local-model","stream":true,"messages":[{"role":"user","content":"hi"}]}'VS Code Copilot setup: Command Palette → Chat: Manage Language Models → Add Models →
Custom Endpoint; enter a group name, a display name and any non-empty API key, and pick API type
Chat Completions. VS Code then opens chatLanguageModels.json — set the model url to your
endpoint (the host/port go here, not in the form):
[
{
"name": "Local llama.cpp",
"vendor": "customendpoint",
"apiKey": "local-dummy-key",
"apiType": "chat-completions",
"models": [
{
"id": "local-model",
"name": "Local model",
"url": "http://127.0.0.1:8080/v1/chat/completions",
"toolCalling": true,
"vision": false,
"maxInputTokens": 6144,
"maxOutputTokens": 2048
}
]
}
]Notes: BYOK powers the chat/agent experience only (inline completions and embeddings still require a
GitHub account). On CPU, prefer a smaller model and a modest context window — the server emits SSE
heartbeats so a long prompt prefill does not trip the client's stream-inactivity timeout. Agent-mode
tool calling depends on the model's own tool-calling quality. Pass --api-key (or
OpenAiServerConfig.apiKey(...)) to require an Authorization: Bearer token; the server binds to
127.0.0.1 by default.
OpenAiCompatServer above is a JSON API server (its / is a 404 — no web page). If you want
the full upstream llama.cpp server, including its bundled Svelte WebUI, use
net.ladenthin.llama.server.NativeServer. It runs the real llama_server inside libjllama over
JNI — no separate llama-server.exe — and forwards the raw llama-server arguments verbatim, so
every flag works exactly as it does for the standalone binary. The fat jar runs it by default
(when --jllama-openai-compat is absent), forwarding its args to the native server (pass --help for the
full llama-server option list):
java -jar target/llama-<version>-jar-with-dependencies.jar \
-m models/model.gguf --host 127.0.0.1 --port 8080 -c 65536 --jinja
# then open http://127.0.0.1:8080/ for the WebUIOr embed it:
try (NativeServer server = new NativeServer(
"-m", "gpt-oss-20b-UD-Q4_K_XL.gguf",
"--host", "127.0.0.1", "--port", "8080",
"-c", "65536", "-b", "4096", "-ub", "2048",
"--jinja", "-ngl", "0", "-t", "8", "-tb", "16",
"-ctk", "q8_0", "-ctv", "q8_0",
"--chat-template-kwargs", "{\"reasoning_effort\":\"low\"}",
"--parallel", "1").start()) {
// Open http://127.0.0.1:8080/ in a browser for the WebUI; the OpenAI API is at /v1/... too.
Thread.currentThread().join();
}Differences from OpenAiCompatServer: with the classic constructor it loads its own model from
the arguments (an independent lifecycle, like llama-server.exe), it is single-instance per
process, it serves the WebUI (in released jars — local cmake builds ship the empty-asset
stub, so no UI there), and it is not available on Android (the upstream server needs
posix_spawn). Readiness: poll GET /health. No SSL (plain HTTP — bind localhost or front with a
TLS proxy).
NativeServer can also attach the full upstream HTTP frontend (routes, WebUI, resumable
streaming) to a LlamaModel you already loaded — one copy of the weights, shared between direct
JNI calls and HTTP:
try (LlamaModel model = new LlamaModel(new ModelParameters().setModel("models/model.gguf"));
NativeServer server = new NativeServer(model, "--host", "127.0.0.1", "--port", "8080").start()) {
// HTTP (incl. WebUI in released jars) and direct Java calls share the same loaded model.
String direct = model.complete(new InferenceParameters("2+2=").withNPredict(4));
Thread.currentThread().join();
}In attach mode the arguments carry only the HTTP-side flags (--host, --port, --api-key, …;
no -m), the server reports healthy immediately (the model is already loaded), and the caller
keeps ownership of the model — close the server before the model, never the other way around.
Started without a model argument, the upstream server runs in router mode: it lists models
from --models-dir, loads/unloads them on demand (GET /models, POST /models/load,
POST /models/unload, per-request "model" selection) and serves each model from a worker
subprocess. Upstream spawns workers by re-executing its own binary — inside a JVM that binary is
java, so before starting an embedded router you must point the worker spawn at this library's
bootstrap:
String javaBin = System.getProperty("java.home") + File.separator + "bin" + File.separator + "java";
NativeServer.setWorkerCommand(javaBin, "-cp", System.getProperty("java.class.path"),
"net.ladenthin.llama.server.NativeServer");
try (NativeServer router = new NativeServer(
"--host", "127.0.0.1", "--port", "8080", "--models-dir", "models").start()) {
Thread.currentThread().join(); // each loaded model runs as a fresh worker JVM
}Worker-command tokens may not contain whitespace (the value is whitespace-split natively).
Typed model management (RouterClient). Instead of hand-rolling HTTP+JSON against the
management endpoints, use server.RouterClient — a plain-HTTP typed client (works against the
embedded router above or any external llama-server router):
RouterClient client = new RouterClient(8080);
List<RouterModel> models = client.listModels(); // GET /models, typed status per entry
client.loadModel("Qwen3-0.6B-Q4_K_M"); // POST /models/load (non-blocking)
client.awaitModelLoaded("Qwen3-0.6B-Q4_K_M", 240_000L); // poll until LOADED; fails fast if the
// worker died (exit code in the message)
client.unloadModel("Qwen3-0.6B-Q4_K_M"); // POST /models/unloadRouterModel carries the identifier, the lifecycle status
(UNLOADED/LOADING/LOADED/SLEEPING/DOWNLOADING/DOWNLOADED), and the router's
failed-worker marker. Chat requests then select a model per request via the standard
"model" field on POST /v1/chat/completions.
Against a router started with --api-key, pass the key — it is sent as
Authorization: Bearer <key> on every call. All of them need it: /models/load and
/models/unload were always gated, and since llama.cpp b10519 the listing endpoints are too.
RouterClient client = new RouterClient(8080, System.getenv("LLAMA_API_KEY"));
// or, for a remote router: new RouterClient("router.internal", 8080, key)Note
awaitModelLoaded waits by polling GET /models, so it cannot observe a model the router
deliberately hides from that listing — a cache model deduplicated by a preset with
dedup-cache-models still loads and still serves by name, but never appears. For those, skip the
await and issue the request directly; with autoload the router waits for the worker itself.
llama.cpp's RPC backend spreads one model over the devices of several machines: every machine
that contributes runs an RPC server, and the machine that loads the model names them with
--rpc. Layers are then distributed over local and remote devices exactly as over several local
GPUs (setGpuLayers, setTensorSplit). Both halves are in every natives jar — CPU and GPU —
with no additional runtime dependency (plain TCP over the system socket
library the library already links).
Serve this machine's devices (every GPU this library found, else the CPU):
try (RpcServer server = RpcServer.startLocal(RpcEndpoint.DEFAULT_PORT)) { // 127.0.0.1:50052
server.awaitTermination();
}or from the command line, with the fat jar:
java -cp llama-<version>-jar-with-dependencies.jar net.ladenthin.llama.RpcServer --port 50052Use the servers from a model:
ModelParameters params = new ModelParameters()
.setModel("models/big-model.gguf")
.setGpuLayers(99)
.setRpcServers(RpcEndpoint.parse("10.0.0.2:50052"), RpcEndpoint.parse("10.0.0.3:50052"));The same works for both HTTP servers: --rpc 10.0.0.2:50052,10.0.0.3:50052 is forwarded to the
native server as-is, and OpenAiCompatServer accepts it too. Any upstream rpc-server works as a
server, and this library's RpcServer works for any llama.cpp client.
Warning
The RPC protocol has no authentication and no encryption: whoever reaches the port can use the
devices and read or write the tensors on them. RpcServer.startLocal therefore binds to loopback
only; RpcServer.startOnNetwork(address, …) (or --host on the command line) is the explicit
opt-in for another interface and logs a warning. Across machines, use a trusted network or a
tunnel (SSH, WireGuard).
What to know:
- An unreachable server fails the load with a
LlamaExceptionnaming it, instead of reaching llama.cpp. A server that disappears after the model loaded still terminates the process — llama.cpp has no error path for a device lost mid-inference. - One
RpcServerper process. A secondstartwhile one runs throwsIllegalStateException. - Endpoints are IPv4 addresses or host names (
host:port); llama.cpp's RPC transport has no IPv6.RpcServerbinds to an IPv4 literal (127.0.0.1,0.0.0.0, an interface address). - Registered servers stay registered. llama.cpp keeps RPC devices in a process-wide registry
with no way to remove them; this library therefore gives every later load that does not ask for a
server an explicit device list without it, so a model loaded without
--rpcnever offloads to a server an earlier model used — the same holds for the multimodal projector,TextToSpeechandLlamaTrainer. An explicitsetDevices(...)/--deviceis never overridden. - Android needs the
android.permission.INTERNETpermission for RPC, even over loopback — thellama-androidAAR does not request it, so an app that wants RPC must declare it itself. RpcServer.startLocal(port, threads, cacheDir)enables upstream's tensor cache: a client that loads the same model again sends the large tensors only once.- Choose the served devices when the default is wrong.
RpcServer.startLocal(port, threads, cacheDir, Arrays.asList("CPU"))(or--device CPUon the command line; names as llama.cpp prints them, e.g.CUDA0,Vulkan1,MTL0) replaces the default of every accelerator. It matters because llama.cpp's RPC client treats every operation as supported by the remote device: a served GPU that cannot run one terminates the server process on the first graph that needs it. The paravirtual GPU of a macOS virtual machine is such a device — serveCPUthere.
A separate artifact, net.ladenthin:llama-langchain4j, adapts a LlamaModel to
LangChain4j's ChatModel, StreamingChatModel,
EmbeddingModel and ScoringModel interfaces in-process over JNI — no HTTP hop, no separate
server. It is a separate artifactId (not a classifier of the core) because LangChain4j 1.x
requires Java 17 while the core net.ladenthin:llama stays Java 8; keeping it separate avoids
forcing that floor on every core consumer. It ships and versions in lockstep with the core.
<dependency>
<groupId>net.ladenthin</groupId>
<artifactId>llama-langchain4j</artifactId>
<version>5.1.0</version>
</dependency>From 5.2.0 on, add the natives next to it — llama-platform or the natives jars you need (see
Choosing the natives jars); the core it depends on is classes only.
Each adapter borrows a LlamaModel you already loaded — it never loads or closes the native
model, so you manage its lifecycle (try-with-resources), and one LlamaModel can back several
adapters at once:
try (LlamaModel llama = new LlamaModel(new ModelParameters().setModel("models/qwen3-0.6b.gguf"))) {
ChatModel chat = new JllamaChatModel(llama);
String reply = chat.chat("Write a haiku about lazy senior devs.");
System.out.println(reply);
}| Adapter | LangChain4j interface | java-llama.cpp call |
|---|---|---|
JllamaChatModel |
ChatModel |
LlamaModel.chat(...) |
JllamaStreamingChatModel |
StreamingChatModel |
LlamaModel.generateChat(...) (token streaming) |
JllamaEmbeddingModel |
EmbeddingModel |
LlamaModel.embed(...) (model loaded with enableEmbedding()) |
JllamaScoringModel |
ScoringModel (re-ranking) |
LlamaModel.handleRerank(...) (model loaded with enableReranking()) |
See llama-langchain4j/README.md for streaming/embedding/re-ranking
examples and the current mapping limitations (tool calling, JSON mode, and multimodal input are
not yet forwarded).
Tip
Supported front ends — all on the same agent session (history, slash commands, approval mode):
| Front end | Start with | Where it runs |
|---|---|---|
| Terminal | (default) / --plain |
this console; --plain for piped, logged or line-only sessions |
| Browser | --web |
http://127.0.0.1:8787/?token=… — over SSH: ssh -L 8787:127.0.0.1:8787 user@server |
| IDE via ACP | --acp |
JetBrains IDEs and Zed natively, VS Code through an ACP extension |
llama-atmosphere-agent/ is a copy-and-run general-purpose agent on the JVM — Claude Code /
OpenCode reduced to the essentials, fully offline; it edits files and, with --allow-shell, runs any command on your machine
(docker, git, build tools) — built from Atmosphere's
built-in OpenAI-compatible agent runtime (streaming, tool loop, workspace file tools) driven
headless against this project's OpenAI-compatible server. It is a standalone Maven project (not a
reactor module) published to Maven Central at the core's version, so the quickest way needs only
JDK 21+ and JBang — it resolves the agent and the core with the natives of
every desktop platform:
jbang net.ladenthin:llama-atmosphere-agent:5.2.0 \
--model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --workspace /path/to/projectOr download it from a release (below, JDK 21+ only), or clone the repository and run it from that folder, which needs only JDK 21+ and Maven — the core jar from Maven Central ships the natives:
# get the folder and a tool-capable model (Qwen3-4B-Instruct-2507, 2.3 GB)
git clone --depth 1 https://github.com/bernardladenthin/java-llama.cpp.git
cd java-llama.cpp/llama-atmosphere-agent
curl -L --create-dirs -o models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf \
https://huggingface.co/unsloth/Qwen3-4B-Instruct-2507-GGUF/resolve/main/Qwen3-4B-Instruct-2507-Q4_K_M.gguf
# the agent with the model loaded in-process — a you> prompt appears (/clear, /exit)
mvn -q compile exec:java \
-Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ctx-size 16384 --workspace /path/to/project --allow-shell"Download instead of cloning. Every GitHub release
carries llama-atmosphere-agent-<version>-jar-with-dependencies.jar (with .sha256 and a GPG .asc,
like the other fat jars). It is a few MB because it holds no core: it runs next to one of the core
fat jars of the same release, so the natives are downloaded once. Put both in one directory and
java -jar finds the core through the agent's manifest (the all-backends jars are tried before the
CPU-only default jar):
# e.g. Linux x86-64: the agent + the all-backends core fat jar of the same version
java -jar llama-atmosphere-agent-5.2.0-jar-with-dependencies.jar \
--model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ctx-size 16384 --workspace /path/to/project --allow-shell
# or with the classpath spelled out (any directory layout; `;` instead of `:` on Windows)
java -cp llama-atmosphere-agent-5.2.0-jar-with-dependencies.jar:llama-5.2.0-all-linux-x86-64-jar-with-dependencies.jar \
net.ladenthin.llama.atmosphere.LocalAgent --model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --workspace /path/to/projectStarted without a core jar next to it, the agent stops with NoClassDefFoundError: net/ladenthin/llama/LlamaModel. CI launches exactly this pair on every run (smoke-agent-linux)
before anything is published.
Warning
--allow-shell lets the model run any command with your user's rights. By default every write and
every command is confirmed on the console ([y]es / [n]o / [a]uto); --auto turns that off. Use a
workspace you are willing to hand to the model.
In the REPL, /help lists the commands (/status, /tools, /mode manual|auto, /compact,
/clear, /exit); anything else goes to the model. A status line shows the approval mode and the
context used ([manual · ctx ~3.1k/16k · 9 tools · local-model]), and the answer is rendered with
headings, bullets and code spans.
Also in a browser or an editor. The same agent session has two more front ends. --web serves it
to a browser on 127.0.0.1:8787 (Atmosphere's own AI console on an embedded Jetty, a random access
token in the printed address; from another machine open an SSH tunnel, ssh -L 8787:127.0.0.1:8787 user@server, rather than binding to the network). --acp speaks the
Agent Client Protocol on stdin/stdout, so JetBrains IDEs and Zed —
and VS Code through an ACP extension — run it as their chat agent, with the editor's own permission
dialog for writes and commands:
java -jar llama-atmosphere-agent-5.2.0-jar-with-dependencies.jar --model model.gguf --allow-shell --web
# JetBrains, ~/.jetbrains/acp.json: {"agent_servers": {"Local llama": {"command": "java",
# "args": ["-jar", "/path/llama-atmosphere-agent-5.2.0-jar-with-dependencies.jar", "--acp", "--model", "/path/model.gguf"]}}}On Windows PowerShell quote the whole argument ("-Dexec.args=--model models\… --allow-shell"); for
the GPU add e.g. -Dllama.classifier=vulkan-windows-x86-64 and --ngl 99. The agent's
README walks through all of it step by step. Other ways to run it, e.g.
against a java-llama.cpp server that is already running (--jinja is required for tool calling;
the fat jars are GitHub release assets,
llama-<version>-all-<os>-<arch>-jar-with-dependencies.jar picks a GPU backend itself):
# 1. the server, e.g. from the release fat jar
java -jar llama-5.2.0-jar-with-dependencies.jar -m models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --jinja --port 8080
# 2. the agent, from the llama-atmosphere-agent/ folder — a you> prompt appears (/clear, /exit)
mvn -q compile exec:java \
-Dexec.args="--base-url http://127.0.0.1:8080/v1 --workspace /path/to/project --allow-shell"
# a single turn instead of the prompt loop
mvn -q compile exec:java \
-Dexec.args="--base-url http://127.0.0.1:8080/v1 --workspace /path/to/project --prompt 'Read the README and summarize it'"
# or without a separate server: load the GGUF in-process
mvn -q compile exec:java \
-Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --workspace /path/to/project"
# everything at once: shell access plus your own system prompt (replaces the built-in one)
mvn -q compile exec:java \
-Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --ctx-size 16384 --workspace /path/to/project --allow-shell --system 'You are a local assistant on this machine with full shell access. run_command executes any command line, including docker, git and build tools. When asked about the system, run a command instead of explaining it. Answer in the language of the user.'"The full streaming tool-calling loop (tools → delta.tool_calls → Java tool → role:"tool" result →
next turn, over several rounds) is verified on every PR against the real OpenAiCompatServer with
no model, and in CI against the Qwen2.5-1.5B tool model — both gate every publish, as does the
release-jar smoke above. See
llama-atmosphere-agent/README.md for the options and the verified
compatibility matrix.
There are two sets of parameters you can configure, ModelParameters and InferenceParameters. Both provide builder
classes to ease configuration. ModelParameters are once needed for loading a model, InferenceParameters are needed
for every inference task. All non-specified options have sensible defaults.
ModelParameters modelParams = new ModelParameters()
.setModel("/path/to/model.gguf")
.addLoraAdapter("/path/to/lora/adapter");
String grammar = """
root ::= (expr "=" term "\\n")+
expr ::= term ([-+*/] term)*
term ::= [0-9]""";
InferenceParameters inferParams = new InferenceParameters("")
.withGrammar(grammar)
.withTemperature(0.8f);
try (LlamaModel model = new LlamaModel(modelParams)) {
model.generate(inferParams);
}LlamaIterable (returned by model.generate(...) and model.generateChat(...))
implements Iterable<LlamaOutput> & AutoCloseable, so every mainstream reactive
library wraps it in a few lines without java-llama.cpp pulling in a runtime
reactive dependency.
Always wrap with the library's resource-management primitive — Flux.using,
Flowable.using, Kotlin use {}, etc. — so that subscription cancellation
flows into LlamaIterable.close() and from there into llama.cpp's native
cancelCompletion. A plain Flux.fromIterable(iterable) or for (x in iter)
loop will NOT close the iterable on cancel; the native task slot stays
occupied until the model is closed.
Flux<LlamaOutput> tokens = Flux.using(
() -> model.generate(params),
Flux::fromIterable,
LlamaIterable::close)
.subscribeOn(Schedulers.boundedElastic());Flowable<LlamaOutput> tokens = Flowable.using(
() -> model.generate(params),
Flowable::fromIterable,
LlamaIterable::close)
.subscribeOn(Schedulers.io());Ready-made: the optional net.ladenthin:llama-kotlin artifact ships
generateFlow/generateChatFlow extensions (close-on-cancellation included) plus suspend
wrappers whose coroutine cancellation is wired to the binding's cooperative CancellationToken:
model.generateChatFlow(params).flowOn(Dispatchers.IO).collect { print(it.text) }Hand-rolled equivalent (no extra dependency):
fun llama(model: LlamaModel, params: InferenceParameters) = flow {
model.generate(params).use { iterable ->
for (output in iterable) emit(output)
}
}.flowOn(Dispatchers.IO)The companion Android sample LLaMAndroid
demonstrates the flow { for (output in model.generate(params)) emit(output) }
shape against the upstream binding. Wrap the for loop in
.use { } if your collector may cancel mid-stream — otherwise the native task
slot will not be released until the model is closed.
val tokens: Source[LlamaOutput, NotUsed] = Source
.fromIterator(() => model.generate(params).iterator())
.async("blocking-io-dispatcher")Why no built-in Publisher? Earlier snapshots of this fork shipped a
hand-rolled LlamaModel.streamPublisher(...) returning a Reactive Streams
Publisher<LlamaOutput>. Since every reactive library bridges blocking
iterables in a few lines via its own resource-management primitive, the binding
now stays free of any reactive runtime dependency — pick whichever library your
app already uses. The pattern is verified end-to-end by
ReactorIntegrationTest in the test sources.
Per default, llama.cpp writes its log as text to stderr (0.00.035.060 I slot … once a model is
loaded): the server's own srv … / slot … lines and, from verbosity 4 on, the llama/ggml lines.
All of it can be intercepted via the static method
LlamaModel.setLogger(LogFormat, BiConsumer<LogLevel, String>): with a callback set, every line goes
to the callback instead of the console (a setLogFile file keeps receiving them). The callback
survives model loads, so set it before new LlamaModel(…) to capture the loading lines too.
LogFormat.TEXT hands over the bare message, LogFormat.JSON one JSON object per line. Passing
null as the callback restores the console output (always llama.cpp's own text format; the format
argument only matters with a callback). Logging can be disabled by passing an empty callback.
Messages arrive asynchronously from llama.cpp's log worker thread; replacing or removing the logger
flushes what is queued to the previous callback first. The verbosity threshold
(ModelParameters.setLogVerbosity(int), llama.cpp's -lv: 1 errors, 2 warnings, 3 info, 4 trace,
5 debug) applies before the callback: 2 keeps warnings and errors and silences the per-request
INFO lines, which is what a console application sharing the terminal with its own output wants.
// Re-direct log messages however you like (e.g. to a logging library)
LlamaModel.setLogger(LogFormat.TEXT, (level, message) -> System.out.println(level.name() + ": " + message));
// Back to llama.cpp's own console output (stderr)
LlamaModel.setLogger(LogFormat.TEXT, null);
// Disable logging by passing a no-op
LlamaModel.setLogger(LogFormat.TEXT, (level, message) -> {});The LogLevel enum values passed to the callback correspond to the native llama.cpp log levels:
| Value | Meaning |
|---|---|
DEBUG |
Verbose diagnostic output |
INFO |
Informational messages about model loading and inference |
WARN |
Non-fatal warnings |
ERROR |
Errors that may affect inference results |
Important
Minimum Android version: API 28 (Android 9.0 Pie). Devices running Android 8.1 (API 27) or earlier are not supported.
One dependency line in Android Studio — no submodule, no NDK build, no manual ProGuard rules:
dependencies {
implementation("net.ladenthin:llama-android:5.1.0")
// or, for Qualcomm Adreno GPUs (device must provide an OpenCL ICD):
// implementation("net.ladenthin:llama-android-opencl:5.1.0")
// optional Kotlin coroutines facade (Flow streaming + suspend wrappers):
implementation("net.ladenthin:llama-kotlin:5.1.0")
}The AAR carries the full net.ladenthin:llama Java API, the CI-built native libraries for
arm64-v8a (devices) and x86_64 (Android Studio emulator, Chromebooks — app bundles
split per ABI so phones download only arm64), both 16 KB page-size compliant, consumer
R8/ProGuard rules (applied automatically), and a manifest minSdkVersion 28 that AGP
enforces against your app. CI boots an x86_64 emulator and runs real on-device inference
against every AAR build.
Do not also depend on the desktop net.ladenthin:llama JAR in the same app — the AAR
already contains those classes, and the JAR would drag ~70 MB of desktop natives into your
APK. See llama-android/README.md and
llama-kotlin/README.md for details.
Runnable example app — "LLM Service". A minimal, KISS, fully-offline on-device chat app
— pick a GGUF from the file system, then chat with it, tokens streaming into a Jetpack
Compose UI, with a 13-language flag picker and private local save/load — lives in
android-llmservice/ (net.ladenthin.android.llmservice).
It builds with plain Gradle/AGP (no Android Studio required), produces a Play-shaped signed
.aab, and is validated in CI by a real on-device emulator UI test. See its README for the
build, signing/Play, and testing walkthrough.
Use this only if you need to patch the native layer or build for an ABI this project does not ship.
- Add java-llama.cpp as a submodule in your an droid
appproject directory
git submodule add https://github.com/bernardladenthin/java-llama.cpp - Declare the library as a source in your build.gradle
android {
val jllamaLib = file("java-llama.cpp")
// Execute "mvn compile" in the llama/ core module if its target/ doesn't exist
// (the repository root is the Maven reactor aggregator; the native core lives in llama/).
if (!file("$jllamaLib/llama/target").exists()) {
exec {
commandLine = listOf("mvn", "compile")
workingDir = file("java-llama.cpp/llama/")
}
}
...
defaultConfig {
...
externalNativeBuild {
cmake {
// Add an flags if needed
cppFlags += ""
arguments += ""
}
}
}
// Declare c++ sources
externalNativeBuild {
cmake {
path = file("$jllamaLib/CMakeLists.txt")
version = "3.22.1"
}
}
// Declare java sources
sourceSets {
named("main") {
// Add source directory for java-llama.cpp
java.srcDir("$jllamaLib/src/main/java")
}
}
}- Exclude
net.ladenthin.llamain proguard-rules.pro
keep class net.ladenthin.llama.** { *; }
Open work items live in TODO.md.
- Expand PIT mutation-testing scope. PIT is wired in
pom.xmland runs on every CI build (in thetest-java-linux-x86_64job) with<mutationThreshold>100</mutationThreshold>.<targetClasses>currently coversnet.ladenthin.llama.value.*,exception.*,args.*and fourjsonparsers (295 mutations, 100% killed, hermetic — no model or fixture needed); widen it incrementally as additional classes reach mutation-test parity. Final target:<param>net.ladenthin.llama.*</param>matching the streambuffer pattern.
Forward-looking ideas being tracked for this fork:
- Adopt feature ideas from the Kotlin Llama Stack client. Candidates (multimodal image input, typed chat messages, async API, batch inference, typed usage/timings) are inventoried with effort estimates in
docs/feature-investigation-llama-stack-client-kotlin.md, derived fromogx-ai/llama-stack-client-kotlin. - Ship a directly Android-capable artifact — DONE.
net.ladenthin:llama-android/llama-android-opencl(AAR, arm64-v8a, minSdk 28, consumer ProGuard rules, 16 KB page-size compliant) plus the optionalnet.ladenthin:llama-kotlincoroutines façade ship from this repo — see Importing in Android. Typed image input for VLMs is covered byContentPart.imageBytes(...)/imageFile(...)(see the multimodal section), so downstream Android projects can drop their dependency onogx-ai/llama-stack-client-kotlinentirely. A dedicated KISS example app — "LLM Service" (SAF model picker + Compose streaming chat, 13-language flag picker, private local save/load, plain Gradle/AGP, signed.aab, on-device emulator UI test) — ships inandroid-llmservice/. - Resolve all upstream
kherud/java-llama.cppopen issues. All 37 open issues at fork time are catalogued with per-issue verdicts indocs/history/49be664_open_issues.md; fixes land in this fork as they are completed. Vision inputs (issues #103 and #34) are now wired end to end through blocking, typed, streaming, and OpenAI-compatible request surfaces.
If you encounter a native crash like:
EXCEPTION_ACCESS_VIOLATION (0xc0000005) at pc=0x00007ffa8f4b2f58
C [msvcp140.dll+0x12f58]
This is a known issue where the C++ runtime library (msvcp140.dll) bundled with some JDK versions is outdated.
Solution: Remove the outdated msvcp140.dll from your JDK:
# Locate and remove msvcp140.dll from JDK directory
# Example for JDK 21:
del "C:\Program Files\Java\jdk-21\bin\msvcp140.dll"
del "C:\Program Files\Java\jdk-21\bin\vcruntime140.dll"
del "C:\Program Files\Java\jdk-21\bin\vcruntime140_1.dll"
# Or on Linux with OpenJDK:
rm /usr/lib/jvm/java-21/bin/msvcp140.dllThe system's updated C++ runtime will be used instead, resolving the crash.
⚠️ DO NOT UPGRADE jqwik past 1.9.3. jqwik 1.10.0 added an anti-AI prompt-injection string to test stdout; the 1.10.1 user guide states the library "is not meant to be used by any 'AI' coding agents at all." 1.9.3 is the last pre-disclosure release and is the pinned version. SeeCLAUDE.mdsection "jqwik prompt-injection in test output" for the full context. Dependabot is configured to ignore allnet.jqwikupdates (every version, including patches) — see theignorerule in.github/dependabot.yml.
Bindings / wrappers
- kherud/java-llama.cpp — the upstream Java binding this project was forked from (see the note at the top of this README); development continues independently here, with the fork-time upstream issues catalogued in
docs/history/49be664_open_issues.md. - llamacpp4j — alternative Java/JNI binding to llama.cpp (SWIG-generated facade); pre-GGUF, dormant since 2023 but historically the other Java JNI option.
- llama-cpp-python — the Python llama.cpp binding; the de-facto feature benchmark among llama.cpp bindings (server mode, multimodal, speculative decoding).
- LLamaSharp — C#/.NET llama.cpp binding with per-backend runtime packages (CPU/CUDA/Vulkan/Metal), the .NET analogue of this project's classifier matrix.
- node-llama-cpp — Node.js/TypeScript llama.cpp binding (prebuilt binaries, JSON-schema-constrained output, function calling).
- LLaMAndroid — Android app demonstrating usage of llama.cpp bindings.
- llama-stack-client-kotlin — Kotlin client for the Llama Stack API with an ExecuTorch-backed local-inference path (the
llama-androidAAR +llama-kotlinfaçade cover the same on-device ground natively). - llama.cpp-android-tutorial — Step-by-step tutorial for running llama.cpp on Android.
Other local inference stacks (no llama.cpp JVM binding)
- Ollama — llama.cpp-based local model runner with its own HTTP API and model registry. This project's OpenAI-compatible server implements the Ollama-native API surface (
/api/version,/api/tags,/api/show,/api/chat,/api/generate), so Ollama-speaking clients (e.g. VS Code Copilot's Ollama provider) work against an in-process jllama model. - ExecuTorch — PyTorch's on-device inference runtime (
.ptemodels, XNNPACK/NPU delegates); the engine behindllama-stack-client-kotlin's local mode and the main non-llama.cpp alternative for Android on-device inference (GGUF is not supported there — different model format ecosystem).
Pure-Java single-model inference (no JNI / no llama.cpp) — Alfonso² Peterssen's *.java family of standalone, dependency-free Java inference runtimes, one per model architecture. Useful when JNI is unavailable (e.g. some sandboxes / GraalVM native-image scenarios) or when you want a single jar with no native side at all. Different design point from this project, which prioritises GGUF compatibility and llama.cpp performance via JNI.
- llama3.java — Llama 3 / 3.1 / 3.2 inference.
- gemma4.java — Gemma 4 (and earlier Gemma 2/3) inference.
- gptoss.java — GPT-OSS architecture inference.
- qwen35.java — Qwen 3.5 inference.
- nemotron3.java — NVIDIA Nemotron-3 inference.
Pure-Java inference engines (no JNI / no llama.cpp)
- Jlama — a full pure-Java LLM inference engine for the JVM (multiple model architectures, quantization, and distributed inference) built on the Java Vector API. A no-native alternative to the JNI approach here; different design point (pure JVM portability vs. GGUF compatibility and llama.cpp performance via JNI).
Frameworks / orchestration
- LangChain4j — LLM-application framework for Java (chat, embeddings, RAG, tool calling, agents) over a unified provider API. This project ships a first-class in-process integration — see the
llama-langchain4jmodule — so a llama.cpp model plugs straight into LangChain4j'sChatModel/StreamingChatModel/EmbeddingModel/ScoringModelwithout an HTTP hop.