Run A2UI agents on free LLM providers, with automatic fallback across ~90 live models.
A2UI lets agents declare UI as JSON instead of describing it in prose. This repo makes A2UI runnable on free models from NVIDIA NIM, Groq, Cerebras, Mistral, Cloudflare Workers AI, and more, with a weekly cron probe that keeps the live-models list fresh as providers rotate or sunset models.
git clone https://github.com/Kiragu-Maina/a2ui-free.git
cd a2ui-free/A2UI/samples/agent/adk/restaurant_finder
uv sync
# Recommended A2UI config — verified passing end-to-end:
export LLM_MODEL=hybrid/synthesize-best
export TOOL_MODEL=nvidia/meta/llama-3.3-70b-instruct
export COUNCIL_JUDGE=nvidia/meta/llama-3.3-70b-instruct
export COUNCIL_PROVIDERS=nvidia,groq,cerebras,mistral
export NVIDIA_NIM_API_KEY=... # https://build.nvidia.com
export GROQ_API_KEY=... # https://console.groq.com
export CEREBRAS_API_KEY=... # https://cloud.cerebras.ai
export MISTRAL_API_KEY=... # https://console.mistral.ai
uv run __main__.pyThe agent now answers user queries with valid, schema-validated A2UI JSON, generated by a council of 3 fast free models and merged by NVIDIA Llama 3.3 70B as the judge.
A2UI agents need an LLM to generate the UI tree. Paid models work but cost real money. Free providers each have:
- different rate limits, quotas, and outages
- models that get sunset or renamed without warning (NVIDIA NIM rotated 2 of shellwire's 6 models in May)
- some that emit visible chain-of-thought tokens that break strict JSON parsers
- different OpenAI-compat shapes, auth headers, and base URLs
This repo papers over all of that with one env var (LLM_MODEL=...) and a fallback pool that almost always returns a result.
Across 11 supported providers (one free API key each), a recent probe found:
| Provider | Live free chat models |
|---|---|
| Mistral | 43 |
| NVIDIA NIM | 38 |
| Cloudflare Workers AI | up to 37 (daily Neuron quota) |
| Groq | 5 |
| Cerebras | 4 |
| SambaNova | 3 |
| Cohere, AI21, Google, DeepSeek, Fireworks | varies by tier |
The set updates weekly: discover_free_models.py probes each provider's /v1/models catalog with a tiny ping and writes a manifest of live / slow / dead. Apps consume the manifest to build their council from whatever's currently working.
| Spec | What it does |
|---|---|
gemini/gemini-2.5-flash |
Native Google Gemini |
nvidia/meta/llama-3.3-70b-instruct |
NVIDIA NIM via LiteLLM |
cloudflare/@cf/meta/llama-3.3-70b-instruct-fp8-fast |
Cloudflare Workers AI |
openai/gpt-4o-mini, anthropic/claude-haiku-4-5, ollama_chat/llama3.1, ... |
Any LiteLLM-supported provider |
council/<strategy> |
Fan out to COUNCIL_MODELS and reduce |
auto/<strategy> |
Build a council from discovered_models.json (round-robin per provider, fastest first) |
hybrid/<strategy> |
Tool turns -> TOOL_MODEL (single), post-tool turns -> council |
single-query— just the first childranked-fallback— try in order, return first successmajority-vote— N in parallel, return the most-common normalized textsynthesize-best— N in parallel, then a judge merges into one clean responseresilient— ranked-fallback over the entire live pool with per-process quarantine on rate-limit / 5xx (skips offenders for 60–300s)
Full env reference: A2UI/samples/agent/adk/.env.example.
Set up the probe to keep the pool fresh:
# Local probe
python A2UI/samples/agent/adk/_shared/discover_free_models.py \
--out discovered_models.json --timeout 25 --parallel 6
# Cron (Sundays 03:17 UTC), see council-test/cron_discover.sh
17 3 * * 0 /path/to/cron_discover.sh >> /var/log/discover.log 2>&1The script sources your .env for provider keys, archives the previous manifest, and emits a new one. Any app pointing at the same discovered_models.json picks up changes on next boot.
A2UI/ Customized fork of google/A2UI@main
samples/agent/adk/
_shared/ The factory + council code
model_factory.py
council_llm.py
discover_free_models.py
.env.example Full reference for every LLM_MODEL mode
restaurant_finder/, orchestrator/, mcp_app_proxy/, rizzcharts/python/,
custom-components-example/, gemini_enterprise/*, personalized_learning/
(all sample agents wired to build_model())
council-test/ Local test workspace
test_hybrid_judge.py The recommended config, against a real A2UI agent
test_resilient_fallback.py Forced-failure test of the quarantine map
test_*.py More end-to-end tests
cron_discover.sh Cron wrapper for the weekly probe
slides-download/ Lightning-talk slide deck (PNGs)
What's verified working end-to-end (against restaurant_finder with real free API calls):
single,council/*,auto/*,hybrid/*,resilientmodes- Cross-provider council (NVIDIA + Groq + Cerebras + Mistral, in one request)
- Schema validation passing on first attempt with
hybrid/synthesize-best+ NVIDIA Llama 3.3 70B judge - Quarantine + fallback under simulated rate-limit / 5xx errors
- Weekly cron probe writing manifest, consumed by both this repo and a NestJS sibling project
MIT for code in this repo (see LICENSE). The A2UI/ subdirectory is Apache 2.0 (see A2UI/LICENSE).