| description | Test GenSwarms swarms — validate configs and run examples with mix genswarms.test and the mock backend. |
|---|
GenSwarms ships two layers of tests: fast ExUnit unit tests for individual modules
and an end-to-end harness (mix genswarms.test) that validates and runs the example
swarm configurations. A mock backend lets you exercise topologies and routing
without any LLM API calls. This page covers both, plus formatting and the
development server.
Run the ExUnit suite with mix test:
mix test # run all tests
mix test --cover # run with coverage report
mix test test/genswarms/routing/router_test.exs # a single file
mix test test/genswarms/routing/router_test.exs:42 # a single test by line
mix test --only tag_name # only tests with a given tagTests run with the synchronous EventStore.Sqlite backend (set in config/test.exs
via config :genswarms, :event_store, Genswarms.Observability.EventStore.Sqlite) so
that persist→query is deterministic, rather than the buffered default. The same
config disables the Phoenix endpoint (server: false) and lowers the log level to
:warning. It also sets config :genswarms, :load_dotenv, false, preventing
automatic .env imports both at application startup and when loading a swarm
configuration. Embedded callers can use the same setting; explicit
Genswarms.CLI.EnvManager.load/1 calls remain explicit imports.
Key test files include:
| File | Covers |
|---|---|
test/genswarms/agents/agent_protocol_test.exs |
@agent: message parsing |
test/genswarms/routing/router_test.exs |
message routing (incl. system objects) |
test/genswarms/config/loader_test.exs |
config loading (.exs/.json/.yaml/.yml) |
test/genswarms/config/swarm_config_test.exs |
config validation |
test/genswarms/agents/inbox_test.exs |
the message queue |
The bwrap lifecycle tests use a local echo executable through the real systemd/bwrap/FIFO transport, not a provider-backed agent. They require a reply round trip before claiming health and skip when the sandbox bases are absent. They do not need API keys; subzeroclaw/provider integration is a separate layer.
This manual command spends on model inference and is never run by mix test:
MIX_ENV=test GENSWARMS_ALLOWED_ENDPOINTS=router.ygr.ai \
mix run scripts/unhardcoded-smoke.exs /path/to/private.env /path/to/subzeroclawIt requires configured bwrap sandbox bases and socat/jq on PATH. On Nix,
run Mix inside nix shell nixpkgs#socat nixpkgs#jq --command ....
Only UNHARDCODED_API_KEY is read from the explicit env file; other variables
are not imported. The key is not a command argument or printed output.
The script uses profile:agent, a network-isolated bwrap sandbox, a fresh
workspace and two tool-using turns. The host verifies the file changes from
17 to 23 independently of the model's reply. It limits each turn to three
client requests, requests 512 output tokens per call, and waits at most 90
seconds per turn. Router-internal fallback and billing remain provider-controlled:
these are not a guaranteed dollar cap. Cleanup stops the sandbox and removes
the temporary workspace on normal completion/error, not on host loss/SIGKILL.
Passing establishes this small live runtime contract, not product quality.
mix format # format all source per .formatter.exsRun mix format before committing; CI and reviewers expect formatted code.
mix genswarms.test discovers, validates, and runs every example. It:
- Discovers all swarm configs (
.exs) and sim files (.sim) recursively underexamples/. - Validates each one.
- Runs each (starts the swarm, waits the full timeout, then stops it).
- Captures a per-example log.
- Reports pass/fail/skip for each, with a combined summary.
mix genswarms.test # validate + run all examples
mix genswarms.test --validate-only # only validate configs, don't run
mix genswarms.test --example tic-tac-toe # test a specific example
mix genswarms.test --mock script.json # run real agents with the mock script (no LLM)
mix genswarms.test --timeout 60000 # custom timeout per swarm (ms)
mix genswarms.test --steps 3 # steps for .sim examples
mix genswarms.test --logs-dir /tmp/logs # custom logs directory
mix genswarms.test --quiet # suppress per-example info linesAll flags use --flag value (single dash, hyphenated) form and are parsed in
strict mode — unknown flags are dropped silently.
| Flag | Type | Default | Purpose |
|---|---|---|---|
--validate-only |
boolean | off | Validate configs only; skip running. No logs directory or summary.log is written. |
--example <name> |
string | all | Keep only files whose path contains the literal substring /<name>/ |
--timeout <ms> |
integer | 60000 |
Per-swarm/per-sim run timeout in milliseconds |
--steps <n> |
integer | 3 |
Number of steps for .sim examples (ignored by .exs configs) |
--mock <path> |
string | none | Expand <path> and export it as SUBZEROCLAW_MOCK_SCRIPT for LLM-free runs |
--logs-dir <path> |
string | .test-logs |
Directory for captured run logs (expanded with Path.expand/1) |
--quiet |
boolean | off | Suppress per-example info output (failures are still printed) |
The --example filter matches a path segment, so it works against the
directory name. For instance, --example tic-tac-toe selects
examples/tic-tac-toe/tic_tac_toe_swarm.exs because the path contains
/tic-tac-toe/. The bundled example directories are: bridge, bwrap-skills,
dynamic-swarm, jev-cron, massive-swarm, party, and tic-tac-toe.
Unless --validate-only is set, each example writes a <name>.log to the logs
directory (default .test-logs/), plus a combined summary.log. The log
filename is derived from the swarm name (the config :name, not the file
path), sanitized by replacing every character outside [a-zA-Z0-9_-] with _
(so hyphens and underscores are preserved). For the bundled examples, whose
config names are tic-tac-toe and party-test, the resulting files are:
.test-logs/
├── tic-tac-toe.log
├── party-test.log
└── summary.log
A per-example .log records the relative path, swarm name, agent/object counts,
topology, timeout, and the final status (or TIMEOUT / ERROR). The
summary.log contains the N passed, N failed, N skipped header followed by one
✓/✗/⊘ line per example.
Exit codes:
0— all examples passed (or were skipped).1— at least one example failed, or no.exs/.simfiles were found underexamples/.
With --validate-only, no logs directory is created and no summary.log is
written — results are printed to the console only.
A run is a skip (⊘) when an .exs file evaluates to something that is not
a map (the (not a swarm config) case), or when a .sim file is found but
SubzeroSim is not available in the project. Note that an .exs file evaluating
to a map without a :name key is not a skip — it counts as a pass
(reported as (valid map config)); only a swarm config (a map with a :name
key) is actually started and run.
There are two distinct ways to avoid real LLM calls; they serve different goals.
backend: :mock is a stub that spawns no external process and produces no agent
output. It is for testing swarm orchestration — topology, routing, and
dynamic add/remove/scale — deterministically and instantly:
%{
name: "test-swarm",
agents: [
%{name: :researcher, backend: :mock},
%{name: :coder, backend: :mock}
],
topology: [{:researcher, :coder}]
}MockBackend.send_input/2 and deploy_skills/2 are no-ops, and handle_output/2
always returns empty output, so the backend does not exercise agent reasoning —
only the machinery around agents. An optional %{script: [...]} (via
{:mock, %{script: [...]}}) is stored on the backend struct for introspection
but is never used to generate responses. See backends.md.
To run real agents (local/docker/apple_container/bwrap) end to end without calling an LLM, give subzeroclaw a mock script:
mix genswarms.test --mock path/to/script.jsonThe task expands the path with Path.expand/1 and exports it as
SUBZEROCLAW_MOCK_SCRIPT, which is passed through to the agents (including bwrap
sandboxes). The subzeroclaw runtime — not GenSwarms — reads the script and
returns canned responses instead of calling the API. The script format is
defined by subzeroclaw.
You can also set SUBZEROCLAW_MOCK_SCRIPT directly in the environment to get the
same behavior outside the test harness.
Start the Phoenix API server (REST + WebSocket, no HTML) for local development:
mix phx.serverThe server defaults to port 4000 (override with the PORT environment variable).
A typical loop is to start the server, then drive it from the CLI or HTTP client:
genswarms start examples/tic-tac-toe/tic_tac_toe_swarm.exs # start a swarm as a daemon
genswarms events --follow # watch the event stream live- backends.md — backend types including the mock backend
- cli.md —
genswarmscommand reference - configuration.md — swarm config DSL and validation
The opt-in tests pin OpenCode 1.18.28, Codex 0.153.4 and Claude Code 2.1.260. They run installed executables, not simulated terminals. Each completes two turns with a disconnect/reattach between them. Loopback SSE providers return deterministic tool calls; the real client's tool executor reads the task and writes its completion receipt. A shared harness checks completion, explicit ACK, readiness and reattachment for all three clients.
GENSWARMS_REAL_TUI=1 mix test test/genswarms/backends/real_tui_*_test.exs
Normal runs skip these tests. Opted-in runs fail if tmux/a client is missing or its version differs: review the fixture and adapter before updating the pin. The fixtures use private homes and allowlisted environments, disable optional client integrations/telemetry, and send model traffic to loopback with no live model spend. External proxy traffic points at a closed local port; this is not OS-level network isolation. Tool permissions are explicit for controlled fixture commands. These checks prove transport compatibility, not model quality or that host execution is a sandbox.
Codex uses an unauthenticated custom Responses provider in its private config, following OpenAI configuration reference. Claude uses a local Messages gateway and a dummy fixture-only token, following Claude environment variables. Neither fixture reads or modifies the user's login/configuration.
The pinned Codex CLI accepts approval policies on-request and never;
untrusted is rejected before launch. Claude's screen-reader mode renders $,
which the adapter recognizes only together with its screen-reader and client
banner, not as a generic host-shell prompt.