Skip to content

Commit 5aa95e5

Browse files
committed
ROCm: compress the embedded GPU code (--offload-compress)
The Windows ROCm jllama.dll was ~1 GB: every HIP translation unit embeds one code object per GPU target, i.e. the whole kernel set once per architecture, and clang stores those bundles uncompressed by default. The jar hid it (234 MB zipped), but LlamaLoader extracts the library to the temp dir on every start, and in the all-backends fat jar ROCm is tried right after CUDA - on nearly every machine without an NVIDIA card. ggml-hip is now compiled with --offload-compress, so each bundle is stored zstd-compressed (CCOB) and inflated by the HIP runtime at module load. The option is scoped to the ggml-hip target and to its source language (HIP on Linux, CXX on Windows, where upstream compiles HIP as C++). Both ROCm jobs run .github/verify-hip-offload-compressed.py after the build: it prints the library size and bundle counts (also into the job summary) and fails on any uncompressed bundle, or on none compressed, so a toolchain or upstream change cannot quietly bring the 1 GB library back. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FtfgoazykGcQSCR3TYBmTQ
1 parent 1b9f516 commit 5aa95e5

4 files changed

Lines changed: 123 additions & 0 deletions

File tree

Lines changed: 86 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,86 @@
1+
# SPDX-FileCopyrightText: 2026 Bernard Ladenthin <bernard.ladenthin@gmail.com>
2+
#
3+
# SPDX-License-Identifier: MIT
4+
5+
"""Fail when a ROCm/HIP build ships its GPU code uncompressed.
6+
7+
llama/CMakeLists.txt compiles ggml-hip with clang's --offload-compress, so every embedded
8+
device-code bundle is a compressed "CCOB" bundle. An uncompressed bundle starts with the magic
9+
"__CLANG_OFFLOAD_BUNDLE__"; one of those in the shipped library means the flag was lost and the
10+
library is back to carrying every GPU target's code uncompressed (~1 GB on Windows).
11+
12+
Usage: verify-hip-offload-compressed.py <dir-or-file>...
13+
14+
Scans every jllama.dll / libjllama.so found, prints its size and the bundle counts, and exits
15+
non-zero if a library has an uncompressed bundle, has no compressed bundle at all, or if no
16+
library was found. Only the standard library, so it runs on any runner.
17+
"""
18+
19+
import os
20+
import sys
21+
22+
UNCOMPRESSED = b"__CLANG_OFFLOAD_BUNDLE__"
23+
COMPRESSED = b"CCOB"
24+
NAMES = ("jllama.dll", "libjllama.so")
25+
CHUNK = 64 * 1024 * 1024
26+
27+
28+
def count(path):
29+
"""Counts both magics in one streaming pass (the file can be ~1 GB)."""
30+
overlap = len(UNCOMPRESSED) - 1
31+
uncompressed = compressed = 0
32+
tail = b""
33+
with open(path, "rb") as f:
34+
while True:
35+
block = f.read(CHUNK)
36+
if not block:
37+
break
38+
data = tail + block
39+
# a match lying entirely inside `tail` was counted in the previous round
40+
uncompressed += data.count(UNCOMPRESSED) - tail.count(UNCOMPRESSED)
41+
compressed += data.count(COMPRESSED) - tail.count(COMPRESSED)
42+
tail = data[-overlap:]
43+
return uncompressed, compressed
44+
45+
46+
def libraries(paths):
47+
for p in paths:
48+
if os.path.isfile(p):
49+
yield p
50+
for root, _, files in os.walk(p):
51+
for name in files:
52+
if name in NAMES:
53+
yield os.path.join(root, name)
54+
55+
56+
def main():
57+
if len(sys.argv) < 2:
58+
print("usage: verify-hip-offload-compressed.py <dir-or-file>...", file=sys.stderr)
59+
return 2
60+
found = failed = 0
61+
summary = []
62+
for lib in libraries(sys.argv[1:]):
63+
found += 1
64+
size = os.path.getsize(lib)
65+
uncompressed, compressed = count(lib)
66+
line = f"{lib}: {size / 1024 / 1024:.0f} MiB, {compressed} compressed / {uncompressed} uncompressed offload bundle(s)"
67+
print(line)
68+
summary.append(line)
69+
if uncompressed:
70+
print(f"::error::{lib} embeds {uncompressed} uncompressed GPU code bundle(s); is --offload-compress still set on ggml-hip?")
71+
failed += 1
72+
elif not compressed:
73+
print(f"::error::{lib} embeds no compressed GPU code bundle; was it built with GGML_HIP=ON?")
74+
failed += 1
75+
if not found:
76+
print(f"::error::no {' / '.join(NAMES)} found under {sys.argv[1:]}")
77+
return 2
78+
step_summary = os.environ.get("GITHUB_STEP_SUMMARY")
79+
if step_summary:
80+
with open(step_summary, "a", encoding="utf-8") as f:
81+
f.write("### ROCm/HIP device code\n\n" + "\n".join(f"- `{s}`" for s in summary) + "\n")
82+
return 1 if failed else 0
83+
84+
85+
if __name__ == "__main__":
86+
sys.exit(main())

‎.github/workflows/publish.yml‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2164,6 +2164,10 @@ jobs:
21642164
run: |
21652165
mvn --no-transfer-progress -f llama/pom.xml compile
21662166
.github/build.sh "-DGGML_HIP=ON -DCMAKE_HIP_COMPILER=$(hipconfig -l)/clang -DGPU_TARGETS=gfx900;gfx906;gfx908;gfx90a;gfx90c;gfx942;gfx950;gfx1010;gfx1011;gfx1012;gfx1030;gfx1031;gfx1032;gfx1033;gfx1034;gfx1035;gfx1036;gfx1100;gfx1101;gfx1102;gfx1103;gfx1150;gfx1151;gfx1152;gfx1153;gfx1200;gfx1201 -DGGML_NATIVE=OFF -DOS_NAME=Linux -DOS_ARCH=x86_64"
2167+
- name: Verify the GPU code is compressed
2168+
# ggml-hip is compiled with --offload-compress (llama/CMakeLists.txt); without it the
2169+
# library carries every target's kernels uncompressed. Fails on an uncompressed bundle.
2170+
run: python3 .github/verify-hip-offload-compressed.py llama/src/main/resources_linux_rocm
21672171
- name: Upload artifacts
21682172
uses: actions/upload-artifact@v7
21692173
with:
@@ -2245,6 +2249,10 @@ jobs:
22452249
# support for. Same rule for the extras upstream omits as on Linux.
22462250
run: |
22472251
.github\build.bat -G "Ninja Multi-Config" -DGGML_HIP=ON -DGPU_TARGETS=gfx900;gfx906;gfx90c;gfx1010;gfx1011;gfx1012;gfx1030;gfx1031;gfx1032;gfx1033;gfx1034;gfx1035;gfx1036;gfx1100;gfx1101;gfx1102;gfx1103;gfx1150;gfx1151;gfx1152;gfx1153;gfx1200;gfx1201 -DCMAKE_PREFIX_PATH="%HIP_PATH%" -DHIP_PATH="%HIP_PATH%" -DCMAKE_C_COMPILER="%HIP_PATH%\lib\llvm\bin\clang.exe" -DCMAKE_CXX_COMPILER="%HIP_PATH%\lib\llvm\bin\clang++.exe" -DCMAKE_HIP_COMPILER="%HIP_PATH%\lib\llvm\bin\clang.exe" -DCMAKE_C_FLAGS="-Wno-error=incompatible-pointer-types" -DOS_NAME=Windows -DOS_ARCH=x86_64
2252+
- name: Verify the GPU code is compressed
2253+
# Same check as the Linux ROCm job: an uncompressed bundle means ~1 GB jllama.dll again.
2254+
shell: pwsh
2255+
run: python .github/verify-hip-offload-compressed.py llama/src/main/resources_windows_rocm
22482256
- name: Upload artifacts
22492257
uses: actions/upload-artifact@v7
22502258
with:

‎CLAUDE.md‎

Lines changed: 16 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -355,6 +355,22 @@ without problems and without local patches; the moment one needs a patch or hold
355355
ROCm/llama.cpp, drop it. The two lists differ **only** by the Instinct parts
356356
(gfx908/gfx90a/gfx942/gfx950), which ROCm supports on Linux alone.
357357

358+
**The ROCm GPU code is compressed (`--offload-compress`), and CI enforces it.** Each HIP
359+
translation unit embeds one code object per GPU target, i.e. the whole kernel set (flash attention,
360+
mmq per quant type, …) once per architecture — stored **uncompressed** by default, which made the
361+
Windows `jllama.dll` ~1 GB for its 23 targets (234 MB zipped, so the jar never showed it). That size
362+
also lands on disk: `LlamaLoader` extracts the library to the temp dir on every start, and in the
363+
all-backends fat jar ROCm is tried right after CUDA, i.e. on nearly every machine without an NVIDIA
364+
card. `llama/CMakeLists.txt` therefore adds `--offload-compress` to the `ggml-hip` target only
365+
(`$<COMPILE_LANGUAGE:HIP,CXX>`: its sources are HIP on Linux and CXX on Windows, where upstream
366+
compiles HIP as C++), so clang stores every bundle zstd-compressed (a `CCOB` bundle) and the HIP
367+
runtime inflates it when the module loads. Upstream llama.cpp does **not** do this; its
368+
`ggml-hip.dll` carries the same uncompressed code (for 20 targets). `.github/verify-hip-offload-compressed.py` runs after the build in
369+
both ROCm jobs, prints the library size and bundle counts (also into the job summary), and fails on
370+
any uncompressed bundle (`__CLANG_OFFLOAD_BUNDLE__`) or on none compressed — a toolchain or upstream
371+
change that drops the flag reds the job instead of quietly shipping the 1 GB library again. The
372+
jar barely shrinks (zip already compressed the code); what shrinks is the extracted library.
373+
358374
Two routing notes mirror existing precedent: **Linux SYCL** ships two precision variants at the *same*
359375
arch, so `CMakeLists.txt` routes them to two *distinct* trees by `GGML_SYCL_F16` (fp16 vs fp32).
360376
**Windows OpenCL** now holds both `x86_64` (desktop ICD) and `aarch64` (Snapdragon/Adreno) in the one

‎llama/CMakeLists.txt‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -196,6 +196,19 @@ FetchContent_Declare(
196196
)
197197
FetchContent_MakeAvailable(llama.cpp)
198198

199+
# ROCm/HIP: compress the embedded GPU code. Without it every HIP translation unit carries one
200+
# UNcompressed code object per GPU target, so the shipped library holds the complete kernel set
201+
# (flash attention, mmq per quant type, ...) once per architecture: ~1 GB for the 23 Windows
202+
# targets, which LlamaLoader then extracts to the temp dir on every start. --offload-compress
203+
# makes clang store each bundle zstd-compressed (a "CCOB" bundle); the HIP runtime inflates it
204+
# when the module is loaded. Scoped to the ggml-hip target (the only one with device code):
205+
# its sources are LANGUAGE HIP on Linux and CXX on Windows (upstream compiles HIP as C++ there).
206+
# CI (.github/verify-hip-offload-compressed.py) fails the ROCm jobs if an uncompressed bundle
207+
# ever reappears.
208+
if(TARGET ggml-hip)
209+
target_compile_options(ggml-hip PRIVATE $<$<COMPILE_LANGUAGE:HIP,CXX>:--offload-compress>)
210+
endif()
211+
199212
# b8831 added ggml_graph_next_uid() which calls _InterlockedIncrement64 via
200213
# <intrin.h> on x86. The intrinsic only exists on x64; provide the
201214
# implementation in a compat TU so the linker resolves __InterlockedIncrement64.

0 commit comments

Comments
 (0)