Skip to content

About

Automatic tuning of LLM serving configurations.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

57 Commits

Folders and files

Repository files navigation

LLM Tuning Toolkit 🔧

Prerequisites

Auto-tuning uses GuideLLM against the OpenAI-compatible endpoint exposed by each engine container. It is installed with the project dependencies.

Installation

First you need to setup your environment with uv.

uv venv --python 3.12
source .venv/bin/activate

Install dependencies with -e for dev mode:

uv pip install -e .

Usage

This module provides a way to automatically detect the best LLM serving configuration that maximises throughput while being complient with a set of defined goodput criteria.

For a complete configuration, start with examples/guidellm-auto-tune.yaml. The scenario.data list maps directly to GuideLLM --data descriptors, so it can use synthetic text, local/Hugging Face datasets, trace replay, and multimodal data. The tuner runs GuideLLM's throughput profile for each engine sweep, then uses a constant-rate profile at progressively lower rates until every configured SLO is met.

For data descriptor fields, preprocessors, and dataset formats, refer to the GuideLLM datasets guide. Auto-tune owns the load profile, constraints, output location, and SLO retry policy.

The supported scenario.load.kind values are throughput, concurrent, constant, poisson, and replay. Throughput mode uses the SLO-driven constant-rate fallback; all other modes evaluate the declared workload directly.

Engine backends

Auto-tune can launch vLLM or SGLang containers. Engine base_args, value_args_pool, and action_args_pool are passed through, so use the native server arguments for the selected engine. Set engine.kind: sglang to use SGLang's python3 -m sglang.launch_server entrypoint and /health_generate readiness check; both can be overridden with engine.entrypoint and engine.health_path.

Use YAML null in a value_args_pool list to omit that argument for a sweep variant. JSON-valued engine arguments should be supplied as JSON strings.

The special tp-dp-combinations sweep maps to --tensor-parallel-size and --data-parallel-size for vLLM, and to --tp-size and --dp-size for SGLang. See examples/guidellm-sglang-auto-tune.yaml for a runnable single-GPU configuration. SGLang's official Docker installation guide and server arguments describe the engine-specific options available for additional sweeps.

How to simulate your scenario

Choose the traffic model separately from the condition that ends a benchmark. For example, a fixed request count is an end condition; it can be used with a concurrent-user, fixed-rate, or trace-replay workload.

Capacity discovery: maximum demand up to a concurrency limit

Use this to find the deployment configuration with the greatest completed throughput, while limiting the number of in-flight requests. It is a stress / capacity test, not a model of a fixed user population: GuideLLM continually issues work as fast as possible until it reaches max_concurrency.

scenario:
  load:
    kind: throughput
    max_concurrency: 512

This is the default auto-tuning strategy and enables the SLO-driven rate fallback.

Fixed active users: closed-loop traffic

Use this when the scenario defines a known number of active clients. Each stream sends its next request after the prior one completes, so streams represents continuously active clients rather than requests per second.

scenario:
  load:
    kind: concurrent
    streams: 512

Without a delay this models automated clients that immediately send another request. For interactive chat or agent sessions, use multi-turn data and synthetic delay (think time) between turns:

scenario:
  data:
    - kind: synthetic_text
      prompt_tokens: 1024
      output_tokens: 256
      turns: 3
      delay: 5
      delay_min: 2
      delay_max: 12
  load:
    kind: concurrent
    streams: 512

Measure several concurrent-user levels

Set streams to a list to measure a concurrency curve with the same engine configuration:

scenario:
  load:
    kind: concurrent
    streams: [64, 128, 256, 512]
  constraints:
    - kind: max_requests
      count: 2048
  slos:
    min_success_rate: 0.99
    max_ttft_p99_ms: 1500

GuideLLM runs one benchmark for every stream count and repeats the complete curve for every engine parameter combination. Auto-tune checks the SLOs at each point and represents that engine configuration with its highest-throughput SLO-compliant point. If no point passes, the engine configuration is rejected. The raw GuideLLM JSON report retains all measured points for plotting throughput and latency against concurrency.

Size a request-count constraint with the stream counts in mind. At a stream count of S and a request limit of R:

  • R < S cannot fill every configured stream.
  • R = S produces one concurrent cohort. Concurrency falls as that cohort drains because completed requests cannot be replaced.
  • R > S allows completed requests to be replaced and holds the workload near S active streams until the request budget is nearly exhausted. Each stream processes approximately R / S requests, although faster streams may process more than slower ones.

Use at least the largest stream count as the request limit. For a sustained-load measurement, a useful starting point is four to ten times the largest stream count. For example, streams: [256, 512, 1024] with max_requests: 4096 measures approximately 16, 8, and 4 request cycles per stream, respectively. When the data descriptor uses turns: T, each turn counts as a request; budget at least largest_stream_count * T requests to allow one complete conversation per stream at the largest load point.

A concurrency curve is useful when you need to locate the saturation point, observe where queueing starts to increase TTFT, compare scaling across engine configurations, or choose a deployment size before the production concurrency is known. Use a single stream count when production has one explicit concurrency contract or when the additional benchmark time is not justified. Concurrent mode evaluates the declared load directly and does not use throughput mode's automatic constant-rate fallback.

Open-arrival traffic: independent requests

For traffic where clients arrive independently, use constant or poisson. constant sends an even request rate; poisson introduces natural variation around an average rate and is usually the closer production model. A useful first estimate is arrival_rate ≈ active_users / (mean_response_time + mean_think_time).

scenario:
  load:
    kind: poisson             # or constant
    rate: 64                  # requests per second
    max_concurrency: 512      # optional safety cap

Replay real traffic

When production traces are available, GuideLLM's replay profile can reproduce their timestamped arrivals. This is the most faithful workload model.

scenario:
  data:
    - kind: trace_synthetic
      path: /absolute/path/traffic.jsonl
  load:
    kind: replay
    time_scale: 1.0

Ending a benchmark cleanly

scenario.constraints maps directly to repeated GuideLLM --constraint descriptors. Use one primary completion boundary. max_duration gives a fixed wall-clock test but cancels in-flight work at the deadline; max_requests gives a fixed request-count test and lets created requests reach a terminal state.

scenario:
  # Capacity discovery for 60 seconds.
  constraints:
    - kind: max_duration
      seconds: 60
scenario:
  # Fixed-user or fixed-rate test: 10,000 created requests reach a terminal outcome.
  constraints:
    - kind: max_requests
      count: 10000

When throughput mode needs its SLO rate fallback, rate_constraints controls the fallback attempts. If constraints is explicitly set and rate_constraints is absent, the same constraints are used for both. If neither constraint list nor legacy duration setting is supplied, auto-tune defaults to max_requests: 1000 for both runs. Set max_requests and rate_max_requests to change those defaults. Existing throughput_duration_seconds and rate_duration_seconds configurations remain supported as explicit legacy duration constraints.

Do not pass constraint through scenario.guidellm_options.arguments: auto-tune rejects it to prevent conflicting stop conditions.

GuideLLM also provides max_errors, max_error_rate, max_global_error_rate, and over_saturation constraints. They are useful safety / early-exit conditions alongside a primary duration or request-count boundary. For example:

scenario:
  constraints:
    - kind: max_requests
      count: 10000
    - kind: max_error_rate
      rate: 0.02
    - kind: over_saturation
      mode: enforce
      min_seconds: 30

When using duration-based tests, keep errored and incomplete request rates separate: an error is a terminal backend/request failure, while an incomplete request can be a benchmark cut-off cancellation.

SLO names use min_ or max_ followed by a normalized metric, such as min_success_rate, max_ttft_p99_ms, max_e2e_p99_ms, or min_output_tokens_per_second. Results include raw GuideLLM JSON reports plus auto_tune_results.json with the selected deployment configuration.

Track benchmark metrics with Trackio

Auto-tuning can optionally send its normalized benchmark metrics to Trackio. Each engine-parameter combination is recorded as a separate Trackio run. The primary benchmark and any fixed-rate SLO fallback attempts are logged as successive steps in that run. Trackio options are supplied through the CLI, not the benchmark YAML.

For local storage, provide only a project name:

uv run auto-tune --config examples/guidellm-auto-tune.yaml \
  --trackio-project llm-tuning-toolkit

After the auto-tune run, open the local dashboard with:

trackio show --project llm-tuning-toolkit

To log to a Hugging Face Space, add its ID. --hf-token is forwarded to Trackio for Space authentication; it defaults to the HF_TOKEN environment variable:

uv run auto-tune --config examples/guidellm-auto-tune.yaml \
  --trackio-project llm-tuning-toolkit \
  --trackio-space-id username/trackio \
  --hf-token "$HF_TOKEN"

To log to a self-hosted Trackio server, use its write-access URL. The write token can be included in that URL or, preferably, supplied through the TRACKIO_WRITE_TOKEN environment variable:

uv run auto-tune --config examples/guidellm-auto-tune.yaml \
  --trackio-project llm-tuning-toolkit \
  --trackio-server-url http://trackio.example:7860

Start a local network-accessible server with trackio show --host 0.0.0.0. --trackio-space-id and --trackio-server-url are mutually exclusive. Omit --trackio-project to disable metric tracking. Use --trackio-group to override the scenario name used to group runs.

For tool calling, provide a GuideLLM dataset with tool-call messages and add the tool_calling_message_extractor data preprocessor through scenario.guidellm_options.arguments. Two further settings are mandatory on GuideLLM 0.7.3, and omitting either one silently benchmarks plain chat at 100% "success" instead of tool calls:

  • load_kwargs: {split: train} on the data descriptor — required on guidellm 0.7.3: file loaders return a DatasetDict and the column mapper crashes without an explicit split.
  • An explicit data-column-mapper with column_mappings: {text_column: messages, tools_column: tools} — messages is not a default prompt column name, and column_mappings replaces the defaults wholesale, so tools must be listed too. Without tools_column no tool definitions reach the server.

Requests must also use backend.request_format: /v1/chat/completions, because OpenAI tool definitions are carried in chat-completion requests. Enable the matching vLLM tool parser in the engine's base_args. examples/guidellm-tool-calls.yaml is a complete, runnable configuration. For vision/audio/video, pass GuideLLM multimodal data descriptors in the same scenario.data list.

The Typer-based CLI validates configuration paths and exposes its full option reference through built-in help:

uv run auto-tune --help
uv run auto-tune --config examples/guidellm-auto-tune.yaml

The main options are --result-dir, --dataset-id, --cache-dir, --hf-token, --no-ui, --log-file, and the --trackio-* options described above. --hf-token also reads HF_TOKEN when it is not supplied explicitly. Results are stored under out/ unless --result-dir specifies another path.

By default, auto-tune uses a colored static terminal view showing sweep progress, the active parameter combination, the best valid throughput, and GuideLLM's live benchmark progress. GuideLLM runs in a pseudo-terminal so its interactive progress can be embedded without including setup messages or final report tables. Use --no-ui to run without the terminal UI: no Rich live display is installed, GuideLLM's own console output is suppressed, and the tuner's log records are written to stdout as plain text.

Use --log-file to save timestamped tuner logs while keeping the terminal UI:

uv run auto-tune --config examples/guidellm-auto-tune.yaml --log-file logs/auto-tune.log

The file is written as the run progresses, so it can be followed with tail -f logs/auto-tune.log. Parent directories are created automatically, and repeated runs append to the file. This option also works with --no-ui and includes full tracebacks for exceptions logged by the tuner. It saves tuner logs, including Trackio warnings; GuideLLM's progress display and Docker engine output are not included. No log file is created unless --log-file is supplied.

About

Automatic tuning of LLM serving configurations.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages