Refactor natives jar distribution and build infrastructure - #472
Merged
Merged
Conversation
net.ladenthin:llama is now the Java classes only. Every native build ships as its own jar of the same artifact, classifier <backend>-<os>-<arch>, holding exactly one directory net/ladenthin/llama/<OS>/<ARCH>/<backend>/. The directories never overlap, so any combination of natives jars shares one classpath: LlamaLoader tries the backends it finds in a fixed order (GPU backends, metal, msvc, cpu) through its ClassLoader and falls through to the CPU library when a GPU runtime is missing. A GPU jar next to the CPU jar now has a CPU fallback; before, each GPU classifier was a complete replacement jar with its own compile pass and none. - llama-platform: new pom-packaging module (nothing to upload besides the pom) naming the classes jar plus the CPU/Metal jars of every desktop. - .github/natives.csv is the single list of natives jars. The merge and fat-jar scripts read it; check-natives.py (code-style job) fails when the generated pom executions, llama-platform, the workflow artifact names or BACKEND_PRIORITY disagree with it. The 16 classifier profiles (17 compile passes) became one generated `natives` profile: pom 2278 -> ~1570 lines. - CMake writes to src/main/natives/<OS>/<ARCH>/<backend>/; build jobs upload natives-<classifier>, and the package/publish jobs fetch them with one glob instead of ~20 download steps each. The merge checks every listed artifact arrived holding only its own directory. - The all-backends fat jars are plain merges carrying only their own OS/arch (all-windows-x86-64 was ~71.5 MB of other platforms' CPU natives); the jllama-backends.txt manifest is gone. Sibling files a backend needs first are listed in a per-directory jllama-extras.txt. - Module path: each natives jar declares a unique Automatic-Module-Name. Without it every natives jar derives the module name `llama` and the JVM silently keeps only the first (measured); package-fatjars.sh checks the manifest of every built jar. module-info now requires Jackson and SLF4J, without which the classes jar never worked on the module path. - smoke-natives-jars.sh loads the classes jar with all 26 natives jars at once, on the classpath and on the module path, in the package job. - Three test classes checked for the library at a hard-coded path and would have skipped silently on the new layout; one shared NativeLibraryPresence helper replaces them. - The Atmosphere agent depends on llama-platform (opt out with -Dllama.natives=none, as CI does); the langchain4j and agent integration jobs point the loader at the downloaded library via lib.path. Verified locally with a real Linux x86-64 build plus placeholder libraries for the other 25 directories: 1823 tests green, all 26 natives jars and 4 all-backends jars built and checked, the fat jar and the natives jars load on classpath and module path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
…workflow
Shared with java-llama.cpp, BitcoinAddressFinder, srcmorph and streambuffer (listed in
.github/shared-files.sha256, checked by the new shared-files job, which fails on a copy changed in
one repository alone and warns on a sibling whose copy differs):
- .github/buildcheck/{workflow,releasegate,sharedfiles}.py + unit tests (stdlib-only Python),
check-release-gate.py and check-shared-files.py;
- print-crash-logs.sh (replaces the crash-log steps pasted into every test job) and
verify-signing-key.sh (the body of the verify-signing-key job);
- the files that were already kept identical by hand, now in the manifest.
The release gate: every job of publish.yml gates both publish jobs unless
.github/release-gate-exemptions.txt names it with a reason. It found vmlens gating nothing;
vmlens and shared-files now gate both publish jobs.
This repository additionally:
- moves check-natives / verify-native-deps / verify-hip-offload-compressed into .github/buildcheck
with unit tests (synthesized ELF/PE/Mach-O, workflow fixtures). check-natives now also checks
that package waits for every natives build, that the all-backends fat jars derived from
natives.csv are each uploaded for and launched by a smoke job and named in the agent jar's
Class-Path and the README, CMake's backend names, the dependency allowlists, the README rows,
and that every *_MODEL_NAME of the workflow is in models.csv. package-fatjars.sh asks for the
targets instead of hard-coding four;
- holds the Android libraries to nativedeps (bionic-only + libOpenCL.so for the OpenCL flavour,
16 KB LOAD alignment read from the ELF headers) instead of an inline allowlist in the AAR job;
- makes the Java tests default to CI's model set (TestConstants DEFAULT_*, new
net.ladenthin.llama.train.model default stories260K.gguf), asserted equal to models.csv in both
directions, and drops the 48 -D model lines and 7 env vars from the workflow;
- adds composite actions restore-models (16 call sites) and install-sccache-windows (10),
print-host-info.sh (18 CPU-info steps), drops validate-models.bat (Windows runs the bash script,
now xxd-free), hoists USE_CACHE / SCCACHE_WEBDAV_ENDPOINT, GRADLE_VERSION and JAVA_VERSION,
and sets scorecard.yml to permissions: read-all like the siblings.
publish.yml: 4293 -> ~3530 lines, job and check names unchanged. Verified: 60+ build-check unit
tests, all check scripts, actionlint, reuse lint, package-fatjars.sh on locally built jars,
verify-signing-key.sh with a throwaway key, TestConstantsTest (mutation-checked).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
…workflow The four fat-jar smoke jobs are one smoke-fatjar matrix (one row per all-<os>-<arch> jar, fail-fast off); their single-jar artifacts are named llama-fatjar-smoke-<target>. check-natives.py holds the uploads (name and the jar inside) and the matrix rows to the targets natives.csv derives. The three macOS and two Windows Java test jobs call .github/workflows/java-tests.yml with four inputs and keep their ids, names and needs. The JDK version moved to .java-version, read by every setup-java step, since a called workflow sees none of the caller's env. Check names of these jobs change; no repository has a required status check, so no setting refers to them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
…atches to history The root CLAUDE.md is loaded into every session and had grown to 312 KB. The Atmosphere agent section (83 KB) moves to llama-atmosphere-agent/CLAUDE.md and the Android app section to android-llmservice/CLAUDE.md; Claude Code loads those only when working in that directory. The root keeps a short summary of what each means for the rest of the repository (release asset, CI gates, version bump, spec of record). The notes on the six dropped llama.cpp patches move to docs/history/dropped-llama-patches.md; the root keeps the standing drop-check and, in the s390x section, the one rule that outlived 0013 (no -DGGML_VXE=ON). Root CLAUDE.md: 312 KB -> 213 KB. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
551f527 stripped the trailing backslash from six run-script lines whose next line starts with || or &&, which makes each script a bash syntax error: build-webui (every native build needs it), the two AAR/consumer checks of package-android-aar and github-snapshot. Found by running bash -n over every run block; actionlint does not parse the scripts without shellcheck. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
…epositories An entry .github/workflows/publish.yml#<job> in .github/shared-files.sha256 hashes one job of the workflow (its header and body, not the comment lines before the next job). startgate, shared-files, verify-signing-key, check-snapshot and check-tag are identical in all four publish.yml files, and verify-signing-key-gradle, github-snapshot and github-release in the three Maven-only ones; they were identical by convention only and are now checked like files, without moving them into a reusable workflow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
…pository
check-versions.py (shared build-check library, run in the shared-files job)
compares every groupId:artifactId the POMs use -- plugins, dependencies,
annotation-processor paths and the Spotless formatter version, ${...}
resolved -- with the default branches of the three sibling repositories
and warns per difference. Dependabot bumps each repository on its own;
this is where the drift now shows instead of in a hand-kept table.
Warnings only: a bump lands in four pull requests. The repositories' own
net.ladenthin artifacts are left out. All 33 coordinates the four
repositories share agree today.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
The shared-files job now parses every run: script of .github/workflows and the composite actions that runs in bash (shell decided as the runner does: step shell, job and workflow defaults, else PowerShell on a Windows runner and bash elsewhere). A broken script -- such as a lost line continuation that leaves a line starting with || -- fails within minutes instead of in the job that runs it. New shared buildcheck module runscripts.py with tests and the check-run-scripts.py CLI, listed in the shared-files manifest of all four repositories. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Every .github file whose only copyright holder is Bernard Ladenthin gets the license header MIT OR Apache-2.0, the same in all four sibling repositories, so a shared file needs no per-repository header. Files with another copyright holder are left unchanged. CODE_OF_CONDUCT.md (and, where the copies are now identical, claude.yml, claude-code-review.yml, scorecard.yml, reuse.yml, osv-scanner.yml and dependabot.yml) joined the shared-files manifest. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
It regenerated wrappers that no longer exist (manylinux2014-x86, android-x86) from unpinned images, while the wrappers in use are pinned to a dockcross tag. Nothing referenced it; each wrapper prints the command that regenerates it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
…workflows The redesigned .github files no longer carry the upstream copyright line that was stamped onto them; they are MIT OR Apache-2.0 like every other own CI file, as are CODEOWNERS and the issue/PR templates (REUSE.toml). claude.yml, claude-code-review.yml, scorecard.yml, reuse.yml, osv-scanner.yml and dependabot.yml are now byte-identical with the sibling repositories and listed in the shared-files manifest. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Linux aarch64 builds natively on ubuntu-24.04-arm with GCC 14, so the linux-arm64-lts cross-compile wrapper has no caller; 32-bit Android was never built in CI and is not published. README, REUSE.toml, CLAUDE.md and the stdc++fs comment in CMakeLists.txt no longer name them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
llama.cpp needs GCC >= 12 since b9789, so no compiler old enough to need the explicit stdc++fs link can build the project; the last CI toolchain that old (dockcross/linux-arm64-lts) is gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Every setup-java reads the JDK from .java-version (java-version-file) instead of a JAVA_VERSION env or a literal 21, as java-llama.cpp already does. .java-version (+ its license file, now MIT OR Apache-2.0) and codeql.yml are byte-identical in all four sibling repositories and listed in the shared-files manifest. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
They are the generated output of dockcross's wrapper template (imagefiles/dockcross.sh); copyright and MIT license as in dockcross's LICENSE. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
A shared-files entry ending in ?repo is hashed with the repository name
replaced by {repo}, so SUPPORT.md, ISSUE_TEMPLATE/config.yml, CITATION.cff,
sonarqube.yml and the code-style job can be checked although they name
their repository. Newly shared: .editorconfig, .gitattributes (*.gguf
binary everywhere), FUNDING.yml, CODEOWNERS, the license texts,
.mvn/jvm.config and .mvn/settings.xml where identical, and the job
verify-signing-key-gradle, now on Gradle 9.8.0 in all four repositories.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
The line naming the upstream author was stamped onto every file when REUSE was introduced. Each file was compared with the upstream source tree at its last upstream commit (49be664) by shared token sequences, which survive reformatting and moves between files. 80 files whose upstream share is nil or only generic boilerplate no longer carry it; the 20 files that still hold upstream code or text keep it, as do README.md, models/README.md, llama/CMakeLists.txt and LICENSE. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
MainExample: blocking completion with token counts and speed, then streaming. ChatExample: a multi-turn console chat on Session, streamed. GrammarExample: a GBNF grammar, and a JSON schema bound to a Java object with completeAsJson. InfillExample: fill-in-the-middle. Each defaults to a model of .github/models.csv and accepts another GGUF as argument. Each example's body is a run(...) method taking the model parameters and the console, so ExamplesTest can run all four on every Java test job (self-skipping without the model); the old ChatExample was @disabled and nothing ever ran the others. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
.clang-format was a 230-line --dump-config of the LLVM style. It is now BasedOnStyle: LLVM plus the four options that actually differ (column limit 120, indent 4, attributes on the same line, no include sorting). The pinned clang-format 23.1.1 resolves it to the same configuration (Objective-C block indentation aside) and leaves all 27 C++ files unchanged; with IndentWidth 2 instead it reports 20 errors in one file. .clang-tidy was an older copy of llama.cpp's; it is now llama.cpp's current file, attributed to the ggml authors. Nothing runs it in CI. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
LlamaModel: class and method documentation describe the current API instead of the original four-item list, the Javadoc examples use ChatMessage, rerank sorts with a reversed comparator, decode explains why native code hands over bytes. CliParameters: argv and toString come from one arguments() method, and the options live in a LinkedHashMap, so argv follows the order the options were set instead of hash order (equals is unaffected; llama.cpp's parser does not depend on order). ModelFlag: the eight inherited descriptions now say what each flag does. What remains in common with the upstream sources is the public API and the JNI declarations, so the upstream copyright line is dropped from these three files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
…del setup ProcessRunner came from xerial/sqlite-jdbc, with a timeout overload added later whose timeout did nothing: it ignored the result of waitFor and then blocked on the output stream. It now starts the command through ProcessBuilder, kills a command that does not end in time and reports it as an IOException, and the plain call waits at most 10 s. A process library (Commons Exec, zt-exec) would be a runtime dependency for two uname calls, so it stays JDK only. ProcessRunnerTest covers output, an unknown command and the timeout. models/README.md still said the models are downloaded automatically. It now points to .github/models.csv, which CI's download-models job and the test defaults share, and shows how to fetch a model locally. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
bernardladenthin
had a problem deploying
to
startgate
October 1, 2026 10:13 — with
GitHub Actions
Error
bernardladenthin
had a problem deploying
to
maven-central
October 1, 2026 10:13 — with
GitHub Actions
Failure
bernardladenthin
had a problem deploying
to
maven-central
October 1, 2026 10:13 — with
GitHub Actions
Failure
|
This was referenced Oct 1, 2026
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Summary
llama-natives-linux-x86_64-cuda13) instead of embedding them in the main jar. Each natives jar contains only one directory structure (net/ladenthin/llama/<OS>/<ARCH>/<backend>/), allowing flexible composition on the classpath.LlamaLoader: Modified to discover native libraries from separate natives jars viaClassLoaderinstead of loading from the main jar's resources, eliminating the need for module-infoopensdirectives..github/natives.csvas the single source of truth for all natives jars; added Python build-check modules to validate consistency across pom.xml, fat-jar assembly, and CI workflows.nativesprofile; updated fat-jar packaging, native dependency verification, and model validation scripts; added new GitHub Actions for core builds and model restoration.natives.py,nativedeps.py,runscripts.py,versions.py,sharedfiles.py,hipoffload.py,models.py,releasegate.py) with unit tests to catch configuration drift early in theshared-filesCI job.CLAUDE.mdfiles forllama-atmosphere-agent/andandroid-llmservice/subdirectories; updated README and CHANGELOG; addeddocs/history/dropped-llama-patches.md.Test plan
shared-filesjob validates all build metadata)Related issues / PRs
Consolidates native jar distribution strategy and centralizes build metadata validation to prevent configuration drift.
Checklist
CONTRIBUTING.mdandCODE_OF_CONDUCT.mdhttps://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2