Auto-tuning uses GuideLLM against the OpenAI-compatible endpoint exposed by each engine container. It is installed with the project dependencies.
First you need to setup your environment with uv.
uv venv --python 3.12
source .venv/bin/activateInstall dependencies with -e for dev mode:
uv pip install -e .This module provides a way to automatically detect the best LLM serving configuration that maximises throughput while being complient with a set of defined goodput criteria.
For a complete configuration, start with
examples/guidellm-auto-tune.yaml. The
scenario.data list maps directly to GuideLLM --data descriptors, so it can
use synthetic text, local/Hugging Face datasets, trace replay, and multimodal
data. The tuner runs GuideLLM's throughput profile for each engine sweep, then
uses a constant-rate profile at progressively lower rates until every configured
SLO is met.
For data descriptor fields, preprocessors, and dataset formats, refer to the GuideLLM datasets guide. Auto-tune owns the load profile, constraints, output location, and SLO retry policy.
The supported scenario.load.kind values are throughput, concurrent,
constant, poisson, and replay. Throughput mode uses the SLO-driven
constant-rate fallback; all other modes evaluate the declared workload directly.
Auto-tune can launch vLLM or SGLang containers. Engine base_args,
value_args_pool, and action_args_pool are passed through, so use the native
server arguments for the selected engine. Set engine.kind: sglang to use
SGLang's python3 -m sglang.launch_server entrypoint and /health_generate
readiness check; both can be overridden with engine.entrypoint and
engine.health_path.
Use YAML null in a value_args_pool list to omit that argument for a sweep
variant. JSON-valued engine arguments should be supplied as JSON strings.
The special tp-dp-combinations sweep maps to --tensor-parallel-size and
--data-parallel-size for vLLM, and to --tp-size and --dp-size for SGLang.
See examples/guidellm-sglang-auto-tune.yaml
for a runnable single-GPU configuration. SGLang's official
Docker installation guide and
server arguments
describe the engine-specific options available for additional sweeps.
Choose the traffic model separately from the condition that ends a benchmark. For example, a fixed request count is an end condition; it can be used with a concurrent-user, fixed-rate, or trace-replay workload.
Use this to find the deployment configuration with the greatest completed
throughput, while limiting the number of in-flight requests. It is a stress / capacity
test, not a model of a fixed user population: GuideLLM continually issues work as fast
as possible until it reaches max_concurrency.
scenario:
load:
kind: throughput
max_concurrency: 512This is the default auto-tuning strategy and enables the SLO-driven rate fallback.
Use this when the scenario defines a known number of active clients. Each stream sends
its next request after the prior one completes, so streams represents continuously
active clients rather than requests per second.
scenario:
load:
kind: concurrent
streams: 512Without a delay this models automated clients that immediately send another request.
For interactive chat or agent sessions, use multi-turn data and synthetic delay
(think time) between turns:
scenario:
data:
- kind: synthetic_text
prompt_tokens: 1024
output_tokens: 256
turns: 3
delay: 5
delay_min: 2
delay_max: 12
load:
kind: concurrent
streams: 512Set streams to a list to measure a concurrency curve with the same engine
configuration:
scenario:
load:
kind: concurrent
streams: [64, 128, 256, 512]
constraints:
- kind: max_requests
count: 2048
slos:
min_success_rate: 0.99
max_ttft_p99_ms: 1500GuideLLM runs one benchmark for every stream count and repeats the complete curve for every engine parameter combination. Auto-tune checks the SLOs at each point and represents that engine configuration with its highest-throughput SLO-compliant point. If no point passes, the engine configuration is rejected. The raw GuideLLM JSON report retains all measured points for plotting throughput and latency against concurrency.
Size a request-count constraint with the stream counts in mind. At a stream count
of S and a request limit of R:
R < Scannot fill every configured stream.R = Sproduces one concurrent cohort. Concurrency falls as that cohort drains because completed requests cannot be replaced.R > Sallows completed requests to be replaced and holds the workload nearSactive streams until the request budget is nearly exhausted. Each stream processes approximatelyR / Srequests, although faster streams may process more than slower ones.
Use at least the largest stream count as the request limit. For a sustained-load
measurement, a useful starting point is four to ten times the largest stream
count. For example, streams: [256, 512, 1024] with max_requests: 4096
measures approximately 16, 8, and 4 request cycles per stream, respectively.
When the data descriptor uses turns: T, each turn counts as a request; budget at
least largest_stream_count * T requests to allow one complete conversation per
stream at the largest load point.
A concurrency curve is useful when you need to locate the saturation point, observe where queueing starts to increase TTFT, compare scaling across engine configurations, or choose a deployment size before the production concurrency is known. Use a single stream count when production has one explicit concurrency contract or when the additional benchmark time is not justified. Concurrent mode evaluates the declared load directly and does not use throughput mode's automatic constant-rate fallback.
For traffic where clients arrive independently, use constant or poisson.
constant sends an even request rate; poisson introduces natural variation around
an average rate and is usually the closer production model. A useful first estimate is
arrival_rate ≈ active_users / (mean_response_time + mean_think_time).
scenario:
load:
kind: poisson # or constant
rate: 64 # requests per second
max_concurrency: 512 # optional safety capWhen production traces are available, GuideLLM's replay profile can reproduce their
timestamped arrivals. This is the most faithful workload model.
scenario:
data:
- kind: trace_synthetic
path: /absolute/path/traffic.jsonl
load:
kind: replay
time_scale: 1.0scenario.constraints maps directly to repeated GuideLLM --constraint descriptors.
Use one primary completion boundary. max_duration gives a fixed wall-clock test but
cancels in-flight work at the deadline; max_requests gives a fixed request-count test
and lets created requests reach a terminal state.
scenario:
# Capacity discovery for 60 seconds.
constraints:
- kind: max_duration
seconds: 60scenario:
# Fixed-user or fixed-rate test: 10,000 created requests reach a terminal outcome.
constraints:
- kind: max_requests
count: 10000When throughput mode needs its SLO rate fallback, rate_constraints controls the
fallback attempts. If constraints is explicitly set and rate_constraints is absent,
the same constraints are used for both. If neither constraint list nor legacy duration
setting is supplied, auto-tune defaults to max_requests: 1000 for both runs. Set
max_requests and rate_max_requests to change those defaults. Existing
throughput_duration_seconds and rate_duration_seconds configurations remain
supported as explicit legacy duration constraints.
Do not pass constraint through scenario.guidellm_options.arguments: auto-tune rejects
it to prevent conflicting stop conditions.
GuideLLM also provides max_errors, max_error_rate, max_global_error_rate, and
over_saturation constraints. They are useful safety / early-exit conditions alongside a
primary duration or request-count boundary. For example:
scenario:
constraints:
- kind: max_requests
count: 10000
- kind: max_error_rate
rate: 0.02
- kind: over_saturation
mode: enforce
min_seconds: 30When using duration-based tests, keep errored and incomplete request rates separate:
an error is a terminal backend/request failure, while an incomplete request can be a
benchmark cut-off cancellation.
SLO names use min_ or max_ followed by a normalized metric, such as
min_success_rate, max_ttft_p99_ms, max_e2e_p99_ms, or
min_output_tokens_per_second. Results include raw GuideLLM JSON reports plus
auto_tune_results.json with the selected deployment configuration.
Auto-tuning can optionally send its normalized benchmark metrics to Trackio. Each engine-parameter combination is recorded as a separate Trackio run. The primary benchmark and any fixed-rate SLO fallback attempts are logged as successive steps in that run. Trackio options are supplied through the CLI, not the benchmark YAML.
For local storage, provide only a project name:
uv run auto-tune --config examples/guidellm-auto-tune.yaml \
--trackio-project llm-tuning-toolkitAfter the auto-tune run, open the local dashboard with:
trackio show --project llm-tuning-toolkitTo log to a Hugging Face Space, add its ID. --hf-token is forwarded to
Trackio for Space authentication; it defaults to the HF_TOKEN environment
variable:
uv run auto-tune --config examples/guidellm-auto-tune.yaml \
--trackio-project llm-tuning-toolkit \
--trackio-space-id username/trackio \
--hf-token "$HF_TOKEN"To log to a self-hosted Trackio server, use its write-access URL. The write
token can be included in that URL or, preferably, supplied through the
TRACKIO_WRITE_TOKEN environment variable:
uv run auto-tune --config examples/guidellm-auto-tune.yaml \
--trackio-project llm-tuning-toolkit \
--trackio-server-url http://trackio.example:7860Start a local network-accessible server with trackio show --host 0.0.0.0.
--trackio-space-id and --trackio-server-url are mutually exclusive. Omit
--trackio-project to disable metric tracking. Use --trackio-group to
override the scenario name used to group runs.
For tool calling, provide a GuideLLM dataset with tool-call messages and add the
tool_calling_message_extractor data preprocessor through
scenario.guidellm_options.arguments. Two further settings are mandatory on
GuideLLM 0.7.3, and omitting either one silently benchmarks plain chat at 100%
"success" instead of tool calls:
load_kwargs: {split: train}on the data descriptor — required on guidellm 0.7.3: file loaders return aDatasetDictand the column mapper crashes without an explicit split.- An explicit
data-column-mapperwithcolumn_mappings: {text_column: messages, tools_column: tools}—messagesis not a default prompt column name, andcolumn_mappingsreplaces the defaults wholesale, sotoolsmust be listed too. Withouttools_columnno tool definitions reach the server.
Requests must also use backend.request_format: /v1/chat/completions, because
OpenAI tool definitions are carried in chat-completion requests. Enable the
matching vLLM tool parser in the engine's base_args.
examples/guidellm-tool-calls.yaml is a
complete, runnable configuration. For vision/audio/video, pass GuideLLM
multimodal data descriptors in the same scenario.data list.
The Typer-based CLI validates configuration paths and exposes its full option reference through built-in help:
uv run auto-tune --help
uv run auto-tune --config examples/guidellm-auto-tune.yamlThe main options are --result-dir, --dataset-id, --cache-dir,
--hf-token, --no-ui, --log-file, and the --trackio-* options described above.
--hf-token also reads HF_TOKEN when it is not supplied explicitly.
Results are stored under out/ unless --result-dir specifies another path.
By default, auto-tune uses a colored static terminal view showing sweep progress,
the active parameter combination, the best valid throughput, and GuideLLM's live
benchmark progress. GuideLLM runs in a pseudo-terminal so its interactive progress can
be embedded without including setup messages or final report tables. Use --no-ui
to run without the terminal UI: no Rich live display is installed, GuideLLM's own
console output is suppressed, and the tuner's log records are written to stdout as
plain text.
Use --log-file to save timestamped tuner logs while keeping the terminal UI:
uv run auto-tune --config examples/guidellm-auto-tune.yaml --log-file logs/auto-tune.logThe file is written as the run progresses, so it can be followed with
tail -f logs/auto-tune.log. Parent directories are created automatically, and
repeated runs append to the file. This option also works with --no-ui and includes
full tracebacks for exceptions logged by the tuner. It saves tuner logs, including
Trackio warnings; GuideLLM's progress display and Docker engine output are not
included. No log file is created unless --log-file is supplied.