Skip to content

Commit efc8c33

Browse files
Merge pull request #450 from bernardladenthin/agent-general-purpose-prompt
llama-atmosphere-agent: general-purpose default prompt; cross-platform ShellToolTest
2 parents 49ca3bf + 7dc26c0 commit efc8c33

16 files changed

Lines changed: 257 additions & 37 deletions

File tree

‎CLAUDE.md‎

Lines changed: 22 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -2140,7 +2140,7 @@ releases as a signed Central Portal bundle upload (staging repo → zip → Publ
21402140

21412141
## Local coding agent with Atmosphere (`llama-atmosphere-agent/`, standalone)
21422142

2143-
A **copy-and-run terminal coding agent** (Claude Code / OpenCode reduced to the essentials, offline)
2143+
A **copy-and-run general-purpose terminal agent** (Claude Code / OpenCode reduced to the essentials, offline)
21442144
that pairs [Atmosphere](https://github.com/Atmosphere/atmosphere)'s built-in OpenAI-compatible
21452145
agent runtime with this project's `OpenAiCompatServer`. Like `android-llmservice/` it is a
21462146
**standalone Maven project, NOT a reactor module and NOT published** — it is an application, and it
@@ -2192,7 +2192,7 @@ the moment anything runs on the module path.
21922192
lines: `AiConfig.configure` → `BuiltInAgentRuntime` → `AgentExecutionContext` + `ToolLoopPolicies`),
21932193
`ConsoleSession` (streams to stdout, prints `⚙ tool {args}` / `↳ result`, supplies the
21942194
`WorkspaceAgentFileSystem` via `injectables()`), `ShellTool` (opt-in `run_command`, `sh -c` /
2195-
`cmd /c` in the workspace, timeout kills the process tree, output tail-truncated), `LocalAgent`
2195+
`cmd /c` starting in the workspace, timeout kills the process tree, output tail-truncated), `LocalAgent`
21962196
(`--base-url` = external server, `--model` = in-process `LlamaModel` + loopback `OpenAiCompatServer`
21972197
with `enableJinja()` and `setLogVerbosity(2)` by default — llama.cpp logs to **stderr**, the console the
21982198
streamed answer shares, so the per-request `slot …` INFO lines would interleave with it; `--log-verbosity <n>`
@@ -2205,6 +2205,26 @@ a `jvm.config` takes no comments, so REUSE can only read its metadata from that
22052205
`REUSE Compliance Check` job fails on `main` — which is how it was found, the PR run having been cancelled.
22062206
Spotless (palantir) is configured in its own pom; the model-free CI job runs `spotless:check`.
22072207

2208+
**The default system prompt is general-purpose on purpose — do not narrow it back.** Every model-facing
2209+
text is a resource, not a Java literal: `src/main/resources/net/ladenthin/llama/atmosphere/` holds
2210+
`system-prompt.txt`, `system-prompt-shell.txt`, `system-prompt-no-shell.txt` and `run-command-tool.txt`
2211+
(each with a `.license` sidecar for REUSE), loaded by `LocalAgent.prompt(name)` with `{placeholder}`
2212+
substitution; `AgentOptionsTest.promptResourcesLoadAndEveryPlaceholderIsFilled` fails on a missing file or
2213+
an unfilled placeholder. `LocalAgent.systemPrompt`
2214+
and the `ShellTool` description describe `run_command` as running *any* command line through the named
2215+
shell (`ShellTool.shellName()`), not limited to the workspace, and tell the model to run a command rather
2216+
than explain one. The earlier wording ("careful *coding agent*", `run_command` "to build, test or inspect
2217+
the project" / "build, test, grep or list files") made Qwen3-4B refuse "list the docker images" — "my
2218+
tools are only for files" — with the tool registered and `docker` on `PATH`; a fresh single-turn run
2219+
refused too, so it was the prompt, not the chat history. Without `--allow-shell` the prompt says commands
2220+
are unavailable and names the flag, so the model does not invent its own limitation. Pinned by
2221+
`AgentOptionsTest.defaultSystemPromptIsGeneralPurposeAndAllowsAnyCommandWithTheShell` and
2222+
`shellToolDescriptionDoesNotNarrowItToTheProject`. `ShellToolTest` runs on every platform: each test
2223+
picks its command line with `ShellTool.isWindows()` — the same detection `ShellTool.run` uses to choose
2224+
`cmd.exe /c` over `sh -c` — so `ls`/`dir /b`, `sleep 30`/`ping -n 30 127.0.0.1 >nul`, and the truncation
2225+
test counts the platform's line separator. It used plain POSIX commands before and failed 4 of 5 on
2226+
Windows; never skip it per OS, give a new test both command forms instead.
2227+
22082228
**Version bump note.** The pom's `llama.version` property is the **release** version, not the
22092229
reactor's `-SNAPSHOT` (CI always overrides it, so a not-yet-published default never breaks CI).
22102230
`versions:set` does not touch this standalone pom, so at release time bump it by hand together with

‎README.md‎

Lines changed: 7 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1021,8 +1021,9 @@ not yet forwarded).
10211021

10221022
### Local coding agent with Atmosphere (`llama-atmosphere-agent/`)
10231023

1024-
A copy-and-run **terminal coding agent on the JVM** — Claude Code / OpenCode reduced to the
1025-
essentials, fully offline — built from [Atmosphere](https://github.com/Atmosphere/atmosphere)'s
1024+
A copy-and-run **general-purpose terminal agent on the JVM** — Claude Code / OpenCode reduced to the
1025+
essentials, fully offline; it edits files and, with `--allow-shell`, runs any command on your machine
1026+
(`docker`, `git`, build tools) — built from [Atmosphere](https://github.com/Atmosphere/atmosphere)'s
10261027
built-in OpenAI-compatible agent runtime (streaming, tool loop, workspace file tools) driven
10271028
**headless** against this project's OpenAI-compatible server. It is a standalone Maven project (not a
10281029
reactor module, not published); you copy the folder and run it. With java-llama.cpp already running
@@ -1043,6 +1044,10 @@ mvn -q compile exec:java \
10431044
# or without a separate server: load the GGUF in-process
10441045
mvn -q compile exec:java \
10451046
-Dexec.args="--model /models/Qwen2.5-7B-Instruct-Q4_K_M.gguf --ngl 99 --workspace /path/to/project"
1047+
1048+
# everything at once: shell access plus your own system prompt (replaces the built-in one)
1049+
mvn -q compile exec:java \
1050+
-Dexec.args="--model /models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --ctx-size 16384 --workspace /path/to/project --allow-shell --system 'You are a local assistant on this machine with full shell access. run_command executes any command line, including docker, git and build tools. When asked about the system, run a command instead of explaining it. Answer in the language of the user.'"
10461051
```
10471052

10481053
The full streaming tool-calling loop (tools → `delta.tool_calls` → Java tool → `role:"tool"` result →

‎llama-atmosphere-agent/README.md‎

Lines changed: 44 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -4,10 +4,12 @@ SPDX-FileCopyrightText: 2026 Bernard Ladenthin <bernard.ladenthin@gmail.com>
44
SPDX-License-Identifier: MIT
55
-->
66

7-
# llama-atmosphere-agent — a local JVM coding agent on java-llama.cpp
7+
# llama-atmosphere-agent — a local, general-purpose JVM agent on java-llama.cpp
88

9-
A minimal, copy-and-run **terminal coding agent** (think Claude Code / OpenCode, reduced to the
10-
essentials) that runs entirely on the JVM and entirely offline:
9+
A minimal, copy-and-run **general-purpose terminal agent** (think Claude Code / OpenCode, reduced to
10+
the essentials): it reads and edits files, and with `--allow-shell` it runs any command on your
11+
machine — `docker`, `git`, build tools, system information. It runs entirely on the JVM and entirely
12+
offline:
1113

1214
- **Model:** any GGUF served by java-llama.cpp's OpenAI-compatible HTTP surface — either a server
1315
you start yourself, or the GGUF loaded **in this process**.
@@ -16,7 +18,8 @@ essentials) that runs entirely on the JVM and entirely offline:
1618
workspace-confined file tools (`ls`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`,
1719
`delete`, `rename`). Driven **headless** — no Spring Boot, no servlet container, no `@Agent`
1820
scanning — through `BuiltInAgentRuntime`.
19-
- **Shell:** an opt-in `run_command` tool (`--allow-shell`) so the model can build and test.
21+
- **Shell:** an opt-in `run_command` tool (`--allow-shell`) that runs any command line through the
22+
system shell (`cmd.exe` on Windows, `sh` elsewhere).
2023

2124
This folder is a **standalone Maven project**, deliberately *not* a reactor module and *not*
2225
published: CI builds and tests it against the core of the same checkout; you copy the folder and
@@ -66,6 +69,18 @@ mvn -q compile exec:java \
6669
-Dexec.args="--model /models/Qwen2.5-7B-Instruct-Q4_K_M.gguf --ngl 99 --workspace /path/to/project"
6770
```
6871

72+
**Everything at once — shell access and your own system prompt:**
73+
74+
```bash
75+
mvn -q compile exec:java \
76+
-Dexec.args="--model /models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --ctx-size 16384 --workspace /path/to/project --allow-shell --system 'You are a local assistant on this machine with full shell access. run_command executes any command line, including docker, git and build tools. When asked about the system, run a command instead of explaining it. Read a file before you edit it. Answer in the language of the user.'"
77+
```
78+
79+
Then ask, for example, *"which docker images are available?"* or *"build the project and fix the
80+
first compiler error"*. On Windows PowerShell, quote the whole argument instead:
81+
`"-Dexec.args=--model C:\models\… --allow-shell --system '…'"`. Inside `--system '…'` avoid the
82+
apostrophe (write *the user* rather than *user's*): the value is already single-quoted.
83+
6984
GPU natives: pick the core classifier, e.g. `-Dllama.classifier=cuda13-linux-x86-64` or
7085
`vulkan-windows-x86-64` (the vendor runtime must be installed — see the root README's classifier
7186
table). Without it the default CPU jar (incl. macOS Metal) is used. In mode A the classifier is
@@ -79,8 +94,8 @@ irrelevant: inference stays in the running server, the agent's JVM loads no mode
7994
| `--model <file.gguf>` | load this GGUF in-process instead | — |
8095
| `--ngl <n>` / `--ctx-size <n>` | GPU layers / context size for `--model` | `0` / `8192` |
8196
| `--log-verbosity <n>` / `--verbose` | llama.cpp log threshold for `--model` (1 errors, 2 warnings, 3 info, 4 trace, 5 debug) / log everything | `2` / off |
82-
| `--workspace <dir>` | directory the file tools (and `run_command`) are confined to | cwd |
83-
| `--allow-shell` | register `run_command` | off |
97+
| `--workspace <dir>` | directory the file tools are confined to, and where `run_command` starts | cwd |
98+
| `--allow-shell` | register `run_command`: any command line, starting in the workspace | off |
8499
| `--system <text>` | replace the default system prompt | built-in |
85100
| `--prompt <text>`, `-p` | one turn, then exit | interactive |
86101
| `--temperature <t>` / `--max-tokens <n>` | sampling / per-call budget | `0.2` / `2048` |
@@ -103,9 +118,31 @@ code page it saw at startup, so umlauts and emoji in the answer would turn into
103118
project's `.mvn/jvm.config` pins `-Dstdout.encoding=UTF-8 -Dstderr.encoding=UTF-8` for the `mvn`
104119
JVM so both sides agree.
105120

121+
### The system prompt
122+
123+
Without `--system` the agent uses a built-in **general-purpose** prompt: it names the file tools and
124+
the workspace they work on, and — only with `--allow-shell` — states that `run_command` runs *any*
125+
command line on this machine (the shell is named, so the model writes the right syntax) and that the
126+
model should run a command rather than explain one. Without `--allow-shell` it tells the model it
127+
cannot run commands and to suggest the flag, so the model does not invent a limitation of its own.
128+
129+
The wording is plain text, not Java: [`src/main/resources/net/ladenthin/llama/atmosphere/`](src/main/resources/net/ladenthin/llama/atmosphere/)
130+
holds `system-prompt.txt` (placeholders `{workspace}` and `{shell_section}`), `system-prompt-shell.txt` /
131+
`system-prompt-no-shell.txt` (the `{shell_section}` with and without `--allow-shell`; `{shell}` is the
132+
shell's name) and `run-command-tool.txt` (the `run_command` description the model reads). Edit them
133+
there to change the default for everyone; `--system` overrides it per run.
134+
135+
This wording matters more than it looks: an earlier default called the agent a *coding agent* and
136+
described `run_command` as a way to *"build, test or inspect the project"*, and Qwen3-4B then refused
137+
*"list the docker images"* ("my tools are only for files") although the tool was registered and the
138+
command worked. `--system <text>` replaces the default **completely** — include whatever the model
139+
still needs to know (the workspace, the shell, your language) in your own text.
140+
106141
Pick a **tool-capable instruct model** (Qwen2.5/Qwen3-Instruct, Llama-3.x-Instruct, Mistral,
107142
Hermes, …). Quality of the loop is the model's: a 1.5B model calls one tool and reads its result, a
108-
7B–32B model does multi-step edit/build/test work.
143+
7B–32B model does multi-step edit/build/test work. Qwen3-4B-Instruct-2507 is a good fast default (fits
144+
an 8 GB GPU with a 16k context); Qwen2.5-Coder-7B, in contrast, wrote the call as a JSON code block
145+
into its answer instead of calling the tool.
109146

110147
## What is verified, and where
111148

‎llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -167,7 +167,7 @@ public static String usage() {
167167
"",
168168
"Agent:",
169169
" --workspace <dir> directory the file tools are confined to (default: cwd)",
170-
" --allow-shell add the run_command tool (runs shell commands in the workspace)",
170+
" --allow-shell add the run_command tool (runs any command line, starting in the workspace)",
171171
" --system <text> replace the default system prompt",
172172
" --prompt <text>, -p run one turn and exit (default: interactive; /exit to quit)",
173173
" --temperature <t> sampling temperature (default " + DEFAULT_TEMPERATURE + ")",

‎llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java‎

Lines changed: 48 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -5,8 +5,11 @@
55
package net.ladenthin.llama.atmosphere;
66

77
import java.io.BufferedReader;
8+
import java.io.IOException;
9+
import java.io.InputStream;
810
import java.io.InputStreamReader;
911
import java.io.PrintStream;
12+
import java.io.UncheckedIOException;
1013
import java.nio.charset.StandardCharsets;
1114
import java.time.Duration;
1215
import java.util.ArrayList;
@@ -23,7 +26,7 @@
2326
import org.jspecify.annotations.Nullable;
2427

2528
/**
26-
* A local, terminal coding agent in the spirit of Claude Code / OpenCode, built from two parts that
29+
* A local, general-purpose terminal agent in the spirit of Claude Code / OpenCode, built from two parts that
2730
* already exist: <b>Atmosphere</b>'s built-in OpenAI-compatible agent runtime (streaming, tool loop,
2831
* workspace file tools) and <b>java-llama.cpp</b>'s OpenAI-compatible server.
2932
*
@@ -45,6 +48,15 @@ public final class LocalAgent {
4548
/** Wall-clock bound on one user turn, including every tool round. */
4649
private static final Duration TURN_TIMEOUT = Duration.ofMinutes(30);
4750

51+
/** The default system prompt; placeholders {@code {workspace}} and {@code {shell_section}}. */
52+
static final String SYSTEM_PROMPT = "system-prompt.txt";
53+
54+
/** The {@code {shell_section}} with {@code --allow-shell}; placeholder {@code {shell}}. */
55+
static final String SHELL_PROMPT = "system-prompt-shell.txt";
56+
57+
/** The {@code {shell_section}} without {@code --allow-shell}. */
58+
static final String NO_SHELL_PROMPT = "system-prompt-no-shell.txt";
59+
4860
private static final Duration SHELL_TIMEOUT = Duration.ofSeconds(120);
4961
private static final int SHELL_MAX_OUTPUT_CHARS = 20_000;
5062

@@ -212,20 +224,47 @@ static ModelParameters modelParameters(AgentOptions options) {
212224
/**
213225
* The default system prompt, or the {@code --system} override.
214226
*
227+
* <p>The default describes a general-purpose agent on this machine, not a coding agent confined to a
228+
* project: a small model reads a narrow role or tool description as a prohibition and then refuses
229+
* requests such as "list the docker images" even though {@code run_command} could do it. With
230+
* {@code --allow-shell} the prompt therefore states that any command line is allowed and that the
231+
* model should run a command rather than explain one; without it, the prompt says so honestly
232+
* instead of letting the model invent a limitation. The text itself is in the resources
233+
* {@value #SYSTEM_PROMPT}, {@value #SHELL_PROMPT} and {@value #NO_SHELL_PROMPT} (see {@link #prompt}).
234+
*
215235
* @param options the options
216236
* @return the system prompt
217237
*/
218238
static String systemPrompt(AgentOptions options) {
219239
if (options.getSystemPrompt() != null) {
220240
return options.getSystemPrompt();
221241
}
222-
String shell = options.isAllowShell()
223-
? " Use run_command to build, test or inspect the project with shell commands."
224-
: "";
225-
return "You are a careful coding agent working in the directory " + options.getWorkspace() + "."
226-
+ " Use the tools to inspect and change files: ls, read_file, write_file, edit_file, glob,"
227-
+ " grep, delete, rename. Paths are relative to that directory." + shell
228-
+ " Work step by step: read a file before you edit it, verify the result after a change,"
229-
+ " and finish with a short summary of what you did.";
242+
String shellSection = options.isAllowShell()
243+
? prompt(SHELL_PROMPT).replace("{shell}", ShellTool.shellName())
244+
: prompt(NO_SHELL_PROMPT);
245+
return prompt(SYSTEM_PROMPT)
246+
.replace("{workspace}", options.getWorkspace().toString())
247+
.replace("{shell_section}", shellSection);
248+
}
249+
250+
/**
251+
* A prompt text from the resources next to this class, trimmed.
252+
*
253+
* <p>The wording lives in {@code src/main/resources/net/ladenthin/llama/atmosphere/*.txt} so it can
254+
* be read and edited as text; {@code {placeholders}} are filled in by {@link #systemPrompt}.
255+
*
256+
* @param name the file name, e.g. {@value #SYSTEM_PROMPT}
257+
* @return the file content without leading or trailing whitespace
258+
* @throws IllegalStateException when the resource is missing from the jar
259+
*/
260+
static String prompt(String name) {
261+
try (InputStream in = LocalAgent.class.getResourceAsStream(name)) {
262+
if (in == null) {
263+
throw new IllegalStateException("Prompt resource missing: " + name);
264+
}
265+
return new String(in.readAllBytes(), StandardCharsets.UTF_8).strip();
266+
} catch (IOException e) {
267+
throw new UncheckedIOException("Cannot read prompt resource " + name, e);
268+
}
230269
}
231270
}

0 commit comments

Comments
 (0)