All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
Developer-experience release. No 0.9.0 API was removed, renamed, or changed — every existing program runs unchanged. See Migrating to 0.10.
@tooldecorator - any Python function becomes an agent tool. The signature becomes the Pydantic input model, the Google-style docstring becomes the description and per-parameter descriptions, and the return value is coerced into aToolResult. Sync functions run off the event loop; aToolUseContextparameter is injected rather than exposed in the schema;Annotatedconstraints survive into the JSON Schema. The decorated function stays directly callable.as_tool_definition()normalizes aToolDefinition, a decorated function, a plain function, or a built-in tool name.- Keyword agent construction -
Agent(name=..., instructions=..., tools=[...])wires its own registry and executor, so one working agent is one object instead of three. Provider and model are auto-detected from whichever credentials are present (ANTHROPIC_API_KEY,OPENAI_API_KEY,GOOGLE_API_KEY,AZURE_OPENAI_API_KEY,AWS_DEFAULT_REGION,OLLAMA_BASE_URL), overridable withANYCODE_DEFAULT_PROVIDER/ANYCODE_DEFAULT_MODEL. CrewAI-stylerole/goal/backstorycompose into a system prompt.tools=accepts functions,ToolDefinitions, and built-in names. - Blocking entry points -
Agent.run_sync,prompt_sync,stream_sync(incremental, via a worker loop),call_tool/call_tool_sync, plusCrew.run_syncandCompiledWorkflow.run_sync/stream_sync. Crew- a one-import multi-agent entry point over the existing wavefront scheduler. Tasks accept aTaskSpec, a dict, or a bare string title;process="sequential"chains tasks that declare no dependency of their own; with no tasks,run(goal)falls through to coordinator decomposition.CrewResultexposessuccess,output,outputs,usage, andcostwhile keeping the fullTeamRunResultattached.Workflow- a state-graph runtime for branching, looping, retry, and fan-out. Frozen Pydantic state (or a dict), nodes returning patches,Commandfor dynamic routing, reducers (add,merge,keep_first,keep_last), conditional edges with or without a path map, compile-time validation that aggregates every structural problem, amax_stepscap surfaced as the newmax_stepsstop reason, event streaming, andto_mermaid()/to_dict(). Nodes may be async or sync functions, anAgent, aCrew, or another compiled workflow.- Long-horizon capabilities on
Agent-planning=Trueregisterswrite_todosand exposesagent.todos;subagents=[SubAgentSpec(...)]registers adelegatetool whose sub-agents run on a fresh conversation with depth capped at one and usage merged into the parent result;workspace=creates a directory and confines the file tools to it. Each is inert unless switched on. There is no separate deep-agent class. TaskSpecergonomics -agent=,expected_output=, andcontext=, withdescriptiondefaulting to the title.expected_outputreaches the agent prompt as an explicit line, andTaskcarries it through the queue and checkpoints.- Machine-readable API map -
anycode.describe()and theanycode apicommand (--core,--compact,--json, or a single symbol) render the public surface as data so an AI coding agent can learn the API without reading the source. AnyCode.register_agent(agent)- adopt a pre-built agent into an engine so it keeps the tool registry it was constructed with.- Documentation - new guides for function tools, crews, workflows, and long-horizon agents; a recipes page of runnable snippets; an orientation page for AI coding agents; a core-surface table in the public API reference; and a migration page.
- Examples -
45_function_tools.py,46_crew_quickstart.py,47_workflow_graph.py, and48_long_horizon_agent.py, all verified against a live provider.
- Lazy public API -
import anycodedropped from ~1913 ms and 1312 modules to ~33 ms and 73 modules. Symbols resolve on first attribute access and cache into the module namespace; aTYPE_CHECKINGblock keeps every name statically resolvable. Optional dependencies (chromadb, redis, OpenTelemetry, the MCP SDK, provider clients) are imported only when their code path runs. Unknown attributes raiseAttributeErrorwith a did-you-mean suggestion. First use ofAgentfell from 1758 ms to ~414 ms, and ofCrewfrom 1855 ms to ~476 ms. Agent(tools=[])means no tools. An empty list was previously indistinguishable fromNone;tools=Noneremains the default and still means every built-in tool.- Missing optional memory backends report the extra to install instead of silently vanishing from
anycode.memory.__all__.
- Latent circular import -
anycode.types→identity→contracts→helpers→anycode.typesmeantimport anycode.typesfailed on its own; the eager package__init__had been masking it. Each ofanycode.types,.helpers,.contracts,.identity,.core, and.memoryis now importable standalone. Agent.call_toolcould not invoke a side-effecting tool - it now supplies a fresh idempotency key when the arguments do not carry one.- Clock-resolution flake in
FilesystemRunStore.mark_interrupted_runs- the staleness cutoff is inclusive, sostale_after_seconds=0means "every running run is stale" even when the platform clock has not ticked since the heartbeat was written.
0.9.0 - 2026-07-24
- Expanded sandbox provider catalog - E2B, Modal, Runloop, Vercel Sandbox, and LangSmith backends now implement the
SandboxProviderprotocol alongside Daytona, each behind its own install extra (sandbox-e2b,sandbox-modal,sandbox-runloop,sandbox-vercel,sandbox-langsmith) with lazy SDK imports, honest capability reports, evidence digests, and fail-closed handling of unsupported network modes, snapshot restores, and secret schemes. A newcreate_sandbox_provider(name)factory builds any backend by name. - Robust Ollama integration - the Ollama adapter now supports thinking (
reasoning_effort/thinking_budget_tokensmap to the nativethinkparameter, withThinkingBlockcontent andthinkingstream events), structured outputs viachat(..., response_format=...)translated to Ollama'sformat, sampling options (max_tokens→num_predict, plus adefault_optionspassthrough forseed,top_p,num_ctx, and friends),keep_alive, and ollama.com cloud authentication throughOLLAMA_API_KEYorapi_key=. - Sandbox and Ollama examples -
examples/41_sandbox_catalog.py(offline provider catalog, capability reports, fail-closed guards),examples/42_vercel_sandbox.pyandexamples/43_modal_sandbox.py(full lifecycle verified against live Vercel and Modal sandboxes), andexamples/44_ollama_robustness.py(thinking, structured outputs, streaming, tool calls, and error handling verified against a live Ollama server).
- Provider-prefixed sandbox secrets -
SandboxSpec.secret_referencesnow accepts any<provider>:<name>reference instead of onlydaytona:; each backend validates its own prefix at create time and returnssandbox_secret_reference_invalidfor foreign prefixes. Existingdaytona:references keep working unchanged.
- Ollama image input and error reporting - image blocks are now sent as the native base64
imagesarray (previously an unsupportedimage_urlshape the server ignored), mid-stream NDJSON error objects surface as terminalerrorstream events instead of a silently truncated answer,done_reason: "length"maps tostop_reason="max_tokens", and a404names the missing model with the exactollama pullcommand. - Vercel sandbox on Windows - the Vercel SDK imports Unix-only pty modules (
termios/tty) at import time for an interactive-shell helper the adapter never calls; the adapter now stubs them on Windows so the sandbox API loads instead of failing withModuleNotFoundError. - Modal filesystem API - sandbox file transfer now prefers
Sandbox.filesystem.write_bytes/read_bytesover the deprecatedSandbox.open(), falling back toopen()on older Modal SDKs.
0.8.2 - 2026-07-19
- Runnable, searchable documentation - added complete copyable programs across core guides, expanded CLI and public API coverage, normalized page descriptions for search snippets, removed duplicate social metadata, and added regression checks for page structure and version-switcher deployment.
- Release metadata recovery - synchronized package and documentation release metadata for the required
v0.8.2tag after the malformed0.8.1GitHub release failed before publishing a package or versioned documentation.
0.8.0 - 2026-07-16
- Portable agent infrastructure preview - added pluggable in-memory, SQLite, and Dapr durability backends with leases, fencing, signals, migration, conformance, and failure-soak coverage; execution identity and fail-closed external policy enforcement; versioned OpenTelemetry GenAI mapping and capture profiles; companion and Daytona sandbox adapters; policy-constrained multi-provider routing; a browser/Node TypeScript service client; and container/Kubernetes hosting profiles with graceful drain and endpoint-specific A2A Agent Cards.
- Versioned semantic contract preview - added strict JSON-only models and checked-in schemas for runs, tasks, messages, artifacts, events, checkpoints, policy decisions, verification results, and capability descriptors. Deterministic state, cancellation, retry, dependency, resume, projection, leased claim/fencing, and artifact-integrity semantics include golden histories, exhaustive state checks, race coverage, and a credential-free end-to-end example.
- Maintainer and contributor governance - added authoritative maintainer, contribution, security-reporting, and release policies covering roles, branches, review evidence, compatibility, deprecation, versioning, backports, Trusted Publishing, release verification, and recovery. Versioned site guides expose the development, governance, and release workflows to contributors.
- Source-linked documentation validation - added a generated inventory for every package-root public export and
scripts/check_docs.pychecks for API coverage, registered built-in tools, numbered examples, page metadata, and curatedllms.txtlinks. Documentation CI and package publication gates now run the strict site build and consistency check. - Executable runtime contract and baseline - documented the current capability matrix, lifecycle transition table, verification attachment points, persisted local formats, supported resume scenarios, side-effect boundary, ADR template, and contract-test conventions. A deterministic example now records task admission, execution, checkpoint size, event volume, and context growth in local or CI evidence, and a real child-process exit test proves cleanup-independent durable resume.
- Responsive documentation and discovery - rebuilt the documentation home and content layouts for mobile navigation, narrow code blocks, scrollable tables, accessible focus states, and balanced desktop grids; removed promotional hero badges; corrected duplicate canonical and heading metadata; added dedicated durability-backend, execution-identity, policy-routing, sandbox-provider, service-hosting, and GenAI-telemetry guides; and synchronized the README,
llms.txt, release runbook, feature guides, and TypeScript client coverage. - Release-bound documentation publishing - pushes to
mainnow validate documentation without overwriting released pages. Final release tags publish theX.Ydocs and movelatest; pre-release tags publish a candidate version without movinglatest. Package publishing resolves locked dependencies and runs repository-wide quality, test, documentation, build, and metadata gates before Trusted Publishing. - Team verification lifecycle -
run_team()now evaluatesafter_teamexactly once against coordinator and task output, preserves lifecycle and verification evidence onTeamRunResult, and returns a recoverable failure for team-level retry decisions. Passing tool-boundary gates return fromverifyingtoexecuting, allowingbefore_toolandafter_toolsensors to coexist in one legal lifecycle.
0.7.0 - 2026-07-11
- MCP, plugin, and tool trust hardening - agent tool allowlists are now enforced again at execution, including an explicit empty list, so a provider cannot invoke an unadvertised registered tool. MCP tools require exact per-agent server opt-in, remain bound to discovery ownership across prebuilt agents and reconnects, always use side-effect idempotency, fail closed when configured auth is missing, and clean up partial initialization without masking cancellation. Plugin entry points are filtered before import and plugin contributions are preflighted before shared registry mutation. The security threat model and production-readiness checklist define the remaining host, network, identity, storage, and operational responsibilities.
- Cross-platform compatibility CI - pull requests and
mainnow run the complete non-integration suite on Linux, Windows, and macOS across Python 3.12 and 3.13. Separate jobs enforce locked quality checks, core-only and per-extra dependency isolation, Redis/ChromaDB integration coverage, and built wheel smoke tests for both core and CLI installations. Repository tests keep the optional-extra matrix synchronized with package metadata. - Enforced compatibility contracts - the complete v0.6 top-level Python API now has an additive CI baseline, and duplicate public declarations fail tests. Declarative YAML/TOML files use format v1, preserve unversioned v1 compatibility, and reject future versions with
UnsupportedConfigVersionError. The compatibility reference defines public import boundaries, semantic-version rules, checkpoint and durable-run reader ranges, plus upgrade and rollback procedures. - Operational observability - agent runs now share durable run/trace correlation across turn, LLM, tool, and terminal spans, with task-local async parenting and deterministic per-trace sampling. Completed spans automatically feed bounded latency, first-token, token, estimated-cost, retry, outcome, and error metrics plus redacted structured events. The new
jsonlexporter emits one correlated completion record per sampled span for container log collectors; OTLP spans preserve runtime timing, carry explicit AnyCode correlation attributes, and expose flush/shutdown lifecycle controls. Configurable span/event/series/histogram retention prevents long-lived telemetry growth and exposes drop counters, while exporter failures remain isolated from run behavior. - End-to-end cancellation ownership — caller cancellation now propagates through
AgentRunner, high-level run and stream APIs, orchestrator waves, provider waits, parallel tools, and shell process trees without being converted into an ordinary result. Agents settle in an explicitcancelledstate, durable runs persist auser_cancelledstop and checkpoint, semaphore accounting remains balanced, andAnyCode.close()cancels and awaits all tracked standalone, coordinator, team, reflection, and handoff operations before resource teardown. - Fail-closed side-effect idempotency — tools can opt into atomic claim-before-execute semantics with
side_effecting=True. Explicit business keys take precedence over deterministic run/turn/call fallbacks; completed calls replay, mismatched input conflicts, and in-progress or unrecorded outcomes terminate the run withside_effect_unknowninstead of being retried. Post-invocation errors default to non-retryable unless the tool explicitly proves otherwise, and uncertain claims are retained during pruning for operator reconciliation. Public in-memory and SQLite stores support process-local or restart-safe coordination, hashed storage keys, redacted result persistence, and pruning. Mutating built-ins and every discovered MCP tool are protected by default. - Provider capacity controls —
ResilientAdapternow applies a provider-scoped concurrency bulkhead (default8) and optional evenly pacedrequests_per_minutelimit to chat and streaming attempts, including retries. Capacity is shared across adapters in the same event loop and scope; conflicting limits for one scope fail clearly. Queue waits load-shed withProviderCapacityErrorafter a configurable timeout, while cancellation and early stream closure release slots immediately.ProviderResilienceConfigis configurable globally throughOrchestratorConfig/YAML or perAgentConfig, with per-agent precedence. - Protected and bounded durable storage —
AgentRunnerandRunSchedulernow accept the publicRunStoreprotocol so production backends can replace the local filesystem implementation.FilesystemRunStoreaccepts aRunPayloadProtectorfor versioned, fail-closed protected payload envelopes while retaining read compatibility with legacy plaintext stores. Run records, transcript events, and turn checkpoints now carry an explicit schema version; legacy unversioned artifacts remain readable and unsupported future versions fail clearly. Workflow checkpoints now default correctly to format v2 while retaining v1 compatibility.RunRetentionPolicyprunes only terminal runs by age and count, and can be applied throughsweep_once,RunScheduler, oranycode runs sweep. Run IDs are constrained to one path segment to prevent storage-root traversal. - Default-on credential redaction — centralized
redact_text,redact_sensitive, andsafe_exception_messagehelpers now protect built-in telemetry exports, exception surfaces, workflow and turn checkpoints, run records and transcripts, context artifacts, session-chain files, persistent memory backends, eval reports, and harness artifacts. Structured sensitive keys and common provider/cloud token formats are replaced with<redacted-secret>while token-usage metrics and other non-secret fields retain their shape. Persistence and exporter configs expose explicitredact_sensitive_data=Falseopt-outs for independently protected stores that require exact replay.
0.6.0 - 2026-07-10
- Provider resilience & prompt caching — new
anycode.providers.resiliencemodule:ResilientAdapterwraps anyLLMAdapterwith classified retry/backoff (429/5xx/timeouts/connection errors, honoringRetry-After), a wall-clock deadline per call, and a per-provider circuit breaker that fails fast withProviderUnavailableErrorwhile open.create_adapterwraps every built-in and plugin provider by default (ProviderResilienceConfig(enabled=False)opts out). Newprovider_unavailablestop reason surfaces exhausted retries as a structured, recoverableRunResultinstead of a raw error. The Anthropic adapter now requests prompt caching (cache_controlbreakpoints on the stable system+tools prefix) whenever the resolved model profile supports it. Newtokensextra declarestiktokenso token accounting can upgrade past the chars/4 heuristic. Tests:tests/test_resilience.py. - Durable run store & mid-run checkpoints — new
anycode.runstorepackage: one directory per run with an atomicmeta.json(RunRecord— status, heartbeat, wake condition), an append-onlytranscript.jsonlevent log (TranscriptEvent, torn tails tolerated), and prunedTurnCheckpoints carrying full conversation, budget, cost, loop-detector window, lifecycle events, verification results, gate decisions, and context manifests.AgentRunneracceptsdurability=DurabilityConfig(...)(opt-in; default behavior unchanged) plusresume_from=to continue a killed run from its last turn boundary with accounting intact, andBudgetTracker/LoopDetectorgained snapshot/restore.CheckpointDatabumped to format v2: serialized agent results now retain lifecycle/verification/gate/manifest state (v1 files still load). Team workflows checkpoint after every completed task, so a mid-wave crash resumes at the first incomplete task instead of re-running the wave. Tests:tests/test_runstore.py. - Automatic context reset & session chaining — handoff artifacts upgraded to a five-layer structure (typed state, narrative, decisions, next steps, warnings) and, with
ContextPolicy(auto_reset_on_handoff=True), the runner now rebuilds its conversation from the artifact mid-run athandoffpressure — same run identity, budget, and audit trail, fresh window — re-injecting task-state invariants as a maximally recent message. NewGoalContract/GoalCriterion(criteria flip only through an external verifier, never the agent's own claim) andSessionChain(anycode.core.session_chain), which drives fresh-context sessions over a persisted contract plus append-onlyprogress.md. Tests:tests/test_session_chain.py. - Tiered persistent memory — the orchestrator now honors
MemoryConfig.vector_backendthrough the newcreate_vector_storefactory instead of hardcoding the in-memory TF-IDF store, so RAG memory survives restarts withvector_backend="chromadb"(a loud warning fires when long-term memory is volatile). Newanycode.memory.knowledgemodule:KnowledgeStorepersists curated "what was learned" entries as human-editable Markdown+frontmatter files with provenance (source, author, timestamp, content hash) and append-plus-supersede curation;build_knowledge_toolsexposes opt-inknowledge_save/memory_searchagent tools;apply_retentiongives rolling logs FIFO retention. Tests:tests/test_knowledge.py. - Context lifecycle hardening — new
maskpressure stage (betweentrimandoffload) replaces aged tool results with short restorable pointers while protecting the recency window. Compaction is now archive-first: the untouched history is written to disk before any summarization and the archive path plus an artifact index are injected into the summary, so compaction is never lossy (ContextManifest.archive_path).ContextManageraccepts an optionalsummarizercallable (e.g. an LLM) with the deterministic extractive path as both default and failure fallback, re-injects preserved task-state invariants after every compaction boundary, and calibrates pressure classification against provider-actual token counts vianote_actual(EMA, clamped). Tests:tests/test_context_hardening.py. - Scheduling, heartbeats & watchdogs — runs can now pause with a persisted
WakeCondition(at_time,on_approval,on_provider_recovery,manual) and be woken by an idempotent, concurrency-safe sweep (sweep_once, per-run lock with stale takeover) from cron or the in-processRunSchedulertick loop. Watchdog semantics separate liveness from progress: stale heartbeat →interrupted(crash), fresh heartbeat without progress → astall_warningaudit event, never an automatic kill. A durable run that exhausts provider retries now pauses with a timedon_provider_recoverywake instead of failing. Newanycode.schedulepackage also shipsScheduledTaskmodes (notification/script/agent/hybrid) so recurring work spends tokens on judgment, not mechanics. Tests:tests/test_schedule.py. - Runs operator CLI — new
anycode runscommand group over the run store:list(status/turns/cost),show(record, wake condition, accounting, recent events),tail(events after a sequence number),audit(deterministic digest of a time window: event counts, tools used, stops/pauses/stalls), andsweep(one watchdog pass). All views derive from the same append-only transcript the runner writes. Tests:tests/test_runs_cli.py. - New examples:
examples/28_durable_runs.py(kill-and-resume),examples/29_session_chain.py(goal contract across fresh contexts),examples/30_scheduled_wakeups.py(pause/wake sweeps + scheduled task modes) — all runnable withFakeAdapter, no API keys. - Plugin / extension ecosystem — new
anycode.pluginspackage introducingPluginManifest, thePluginProtocol, aPluginBaseno-op default, andPluginRegistry. Plugins bundle custom tools, async provider factories, verification sensors, and turn hooks into a single object.AnyCode.register_plugin(...)installs a plugin into the engine;AnyCode.load_installed_plugins()discovers and installs every plugin published under theanycode.pluginsentry-point group.create_adapternow dispatches unknown provider names through the plugin-registered provider-factory registry.anycode inspect pluginsand the augmentedanycode inspect providerssurface discovered plugins and plugin-contributed providers. New exampleexamples/27_plugin_ecosystem.pyand teststests/test_plugins.py. - Handoff chain recursion —
HandoffExecutor.executenow follows multi-hop chains: when the target agent itself emits ahandoff_request, the executor recurses to the next agent up tomax_handoff_depth, appending each hop (and any depth-limit short-circuit) to the optionalchainargument.AnyCode._run_wave_taskpasses the chain list soTeamRunResult.handoffsrecords the full multi-hop path. YAML/TOML configs accept a top-levelmax_handoff_depthinteger. - Context engineering for huge-context models — new
anycode.contextpackage with a built-inModelContextProfileregistry (Anthropic 200k/1M, OpenAI 128k/1M, Google 1M/2M, plus unbounded fallback), profile resolution chain (override → custom → built-in → provider default → unbounded), and pluggableTokenizerProtocol with heuristic + optionaltiktokenbackends.ContextPolicygainsmode(disabled/manual/auto),reserved_response_tokens, typedsectionsbudgets, andcustom_profiles/model_profilefields so any future window size — including 5M+ tokens — is supported without code changes.ContextManagernow classifies content intoContextSectionKinds, applies per-sectionoverflowstrategies (trim/summarize/offload/drop/error), emits aContextUsageReporton every manifest, and exposesContextManager.reconcile(...)so provider-actual token counts replace heuristic estimates after the call.TokenUsageaddscache_creation_input_tokensandcache_read_input_tokens;CostTracker/calculate_costbill cache reads atcached_input_cost_per_1kwhen available. The Anthropic adapter extracts cache token classes natively; the YAML/TOML config loader recognises a top-levelcontext_engineeringblock plus per-agentcontext_policyoverrides. New helpersformat_usage_reportandrender_usage_report_tableship Markdown-ready summaries. New exampleexamples/26_context_engineering.py, configexamples/config/context_engineering.yaml, docdocs/context-engineering.md, and teststests/test_context_engineering.py,tests/test_model_profiles.py,tests/test_token_accounting.py,tests/test_huge_context.py. - Runtime telemetry & cancellation —
AgentRunner.streamnow catchesasyncio.CancelledError, emits a terminalcancelledlifecycle phase with auser_cancelledStopReason, and records finalphase/stop_reason/recoverableattributes on a dedicatedanycode.agent.{name}.terminalspan before re-raising. - Adaptive context lifecycle —
ContextPolicygainsprovider_overrides,preserved_task_state, andpreserved_verification_failures.ContextManageraccepts aprovider=kwarg, resolves the matching override viaContextPolicy.for_provider(), and emits aContextManifestthat includes the resolved provider plus preserved state/failure sections during compaction. - Declarative quality gates — new
anycode.verification.registryexposesregister_sensor_factory,build_sensor, andbuild_sensors. Built-in factories coverruff,pyright,pytest, and a pure-Pythonregexsensor.AgentConfig.verification,RunnerOptions.verification, andOrchestratorConfig.verificationplumbVerificationSensorConfigtuples. The runner instantiates aQualityGateand evaluates it atbefore_tool,after_tool, andafter_taskphases; the orchestrator builds a separate team-level gate and evaluates it atafter_team.block/escalateoutcomes translate into averification_failedStopReasononAgentRunResultandTeamRunResult, whileretryoutcomes feed sensor feedback back into the agent loop. The YAML config loader reads top-level and per-agentverification:blocks. - Team lifecycle aggregation —
TeamRunResultnow exposes aggregatedlifecycle_events,verification_results,gate_decisions, and a top-levelstop_reasonso callers can inspect every task's lifecycle trail and any team-level gate outcome. - Deterministic evaluation suite — new
anycode.providers.fake.FakeAdapter(andFakeResponse) replays a scripted reply sequence with no LLM credentials.EvalScenariogainsdeterministic,fake_responses, andfake_tool_failuresfields;run_scenarionow branches into a deterministic harness when the flag is set.EvalScenarioResultandEvalReportaggregatecost_usd,retries, andverification_failuresso CI can track the new metrics. - New examples:
examples/22_deterministic_eval.py,examples/23_context_pressure.py,examples/24_verification_gates.py,examples/25_runtime_cancellation.py. - New deterministic eval fixture
tests/fixtures/eval/runtime_reliability_deterministic.yaml. - New tests in
tests/test_harness_runtime.pycovering provider overrides, the sensor registry, the deterministic suite, runner cancellation telemetry, multi-phase quality gates (before_tool/after_tool), and orchestrator team-level (after_team) gating.
- Declarative config loading now rejects unknown root, agent, task, context-engineering, and nested model fields with
UnknownConfigFieldErrorinstead of silently ignoring them. Programmatic model construction is unchanged. - Handoff sentinel encoding — the built-in
handofftool now encodes its payload as__HANDOFF__:<json>instead of a colon-delimited string. Free-formsummary/reasontext containing:(or any other character) now round-trips losslessly throughAgentRunner._detect_handoff. New helpersencode_handoff_payload/decode_handoff_payloadare exported fromanycode.handoff.tool. AgentStateusesField(default_factory=...)formessagesandtoken_usageto guarantee per-instance defaults.AgentRunner.streamhandoff path now appends every executedToolCallRecordfrom the same batch and emits atool_resultevent for each before yielding thehandoff/doneevents. Previously, sibling tool calls executed alongside a handoff were dropped fromRunResult.tool_callsand never streamed to observers.
validate_config()no longer reports a missingsystem_promptas a validation error — it was only ever a soft recommendation, so callers likeanycode inspect configno longer fail hard on configs that omit it.- The PyYAML
ImportErrorraised byanycode.config.loadernow points at the correct extras:pip install "anycode-py[cli]"(the package on PyPI isanycode-py, notanycode).
0.5.0 - 2026-05-06
- Agent Handoff (orchestrator integration) —
Teamworkflows now route handoff requests throughHandoffExecutor, validate handoff targets against team membership, support policy-driven handoffs viaOrchestratorConfig.handoff_policy, and emitHandoffrecords onTeamRunResult.handoffs. - Intelligent Routing (orchestrator integration) —
Router.route()decisions are applied per task before execution; the resolved model/provider override is layered onto the agent config without mutating the original.TeamRunResult.route_decisionsexposes the decision trail. - CLI Toolkit — new
anycodeCLI (built ontyper+rich):anycode init <dir>scaffolds a project (team.yaml,main.py,.env.example,tools/,.gitignore).anycode run <config.yaml>loads a team config and runs it end-to-end.anycode inspect tools|providers|team <path>|config <path>introspects the runtime.anycode versionprints package + Python info.- Available via
pip install anycode-py[cli].
- Declarative YAML/TOML config (
src/anycode/config/) —load_config(path)+validate_config(path)parse.yaml/.yml/.tomlfiles into typedLoadedConfigwith${ENV_VAR}substitution. NewAnyCode.from_config(path)andengine.run_team_from_config(goal=...)classmethod/method. - Examples cookbook — five new end-to-end examples (
13_cost_tracking.py,14_self_reflection.py,15_rag_memory.py,16_dag_visualization.py,17_yaml_config.py). - Self-Reflection / Critic Loop (
src/anycode/reflection/) —LLMCritic,parse_critic_json,ReflectionLoopwithself/peer/custommodes. Configured viaOrchestratorConfig.reflection = ReflectionConfig(...). Tracksreflections_countandquality_scoreonAgentRunResult. - Cost-Aware Execution Engine (
src/anycode/cost/) —CostTracker,build_cost_report,DEFAULT_PRICING,find_pricing(with wildcard fallback),calculate_cost. Configured viaOrchestratorConfig.cost = CostConfig(budget_usd=..., on_budget_exceeded="stop"|"warn"|"continue"). Emits cost-alert events at the configured threshold and stops execution when the budget is exhausted (whenon_budget_exceeded="stop").TeamRunResult.cost_reportexposes per-agent and per-model breakdown. - DAG Visualization (
src/anycode/viz/) —render_dag(queue, format="mermaid"|"dot"|"json"|"ascii", show_status=True)andrender_timeline(team_result, width=40). Mermaid output includesclassDefstyling per task status. - RAG Memory (
src/anycode/memory/rag.py,src/anycode/memory/indexer.py) —RAGRetriever(dedup, namespace filtering, relevance/token caps) andRAGIndexer(paragraph-aware chunking, optional tool-result indexing). Configured viaOrchestratorConfig.rag = RAGConfig(...). RAG context is auto-injected into every task prompt and outputs are auto-indexed to the configuredVectorStore(defaults toInMemoryVectorStore).
AgentRunResultgainedhandoff_request,reflections_count, andquality_scorefields (all optional / defaulted).RunResultgainedhandoff_requestfield (optional).TeamRunResultgainedhandoffs,route_decisions, andcost_reportfields (all optional).OrchestratorConfiggainedcost,reflection, andragfields (all optional).- Public exports added to
anycode.__init__:CostTracker,build_cost_report,DEFAULT_PRICING,calculate_cost,find_pricing,render_dag,render_timeline,LLMCritic,ReflectionLoop,parse_critic_json,RAGRetriever,RAGIndexer, plus the new typesModelPricing,CostConfig,CostBreakdown,CostReport,CriticResult,Critic,ReflectionConfig,RAGConfig,RAGContext,RAGEntry.
- 43 new tests across
tests/test_cost.py,tests/test_viz.py,tests/test_config.py,tests/test_reflection.py,tests/test_rag.py,tests/test_cli.py. Total suite: 343 passing.
0.4.0 - 2026-06-10
- Additional LLM Providers — 4 new provider adapters implementing the
LLMAdapterProtocol.GeminiAdapter— Google Gemini viagoogle-genaiSDK with function calling and streaming support.OllamaAdapter— Local Ollama models via HTTP (httpx), zero external SDK dependencies, OpenAI-compatible tool format.BedrockAdapter— AWS Bedrock for Claude models viaboto3, Anthropic message format, streaming via response streams.AzureOpenAIAdapter— Azure OpenAI via the officialopenaiSDK with Azure-specific auth and deployment configuration._openai_compatshared helper module — extracted common OpenAI mapping logic (messages, tools, stop reasons) for reuse across OpenAI, Azure, and Ollama adapters.- Extended
create_adapter()factory to resolve all 6 providers with lazy imports.
- MCP Integration module (
src/anycode/mcp/) — Model Context Protocol support for external tool servers.MCPClient— manages connection lifecycle (stdio, SSE, streamable-http transports) via the officialmcpSDK, with tool discovery and tool execution.schema_to_pydantic_model()— dynamic Pydantic model generation from JSON Schema for MCP tool inputs.mcp_tool_to_definition()— converts MCP tools into AnyCodeToolDefinitionwith prefixed naming (mcp_{server}_{tool}).discover_and_register()— batch discovery and registration of MCP tools into theToolRegistry.validate_server_config()— transport-aware configuration validation.ToolRegistry.register_from_mcp()andToolRegistry.deregister_prefix()for MCP tool lifecycle management.
- Agent Handoff module (
src/anycode/handoff/) — context-preserving agent-to-agent task delegation.HANDOFF_TOOL_DEF— built-in sentinel tool that agents call to request a handoff (returns__HANDOFF__:to:summary:reason).HandoffExecutor— orchestrates context transfer with conversation trimming, system/user prompt generation, and configurable depth limiting.trim_context(),build_handoff_system_prompt(),build_handoff_user_message()— protocol helpers for handoff payloads.- Runner integration:
AgentRunnerdetects handoff sentinels in tool results and yieldsStreamEvent(type="handoff").
- Intelligent Routing module (
src/anycode/routing/) — zero-cost heuristic task routing.classify_task()— microsecond complexity classification (5 levels: trivial, simple, moderate, complex, expert) based on description length and dependency count.match_rule()/evaluate_rules()— declarative rule engine supporting complexity conditions, keyword-in checks, and regex patterns with priority ordering.DefaultRouter—RouterProtocol implementation combining classifier + rules engine with default model fallback.- Orchestrator integration: routing decisions applied before task wave execution.
- New Pydantic types (all
frozen=True):MCPServerConfig,MCPToolInfo,HandoffRequest,Handoff,HandoffPolicyProtocol,ComplexityLevel,RoutingRule,RoutingConfig,RouteDecision,RouterProtocol. - Extended
AgentConfig.providerliteral to include"google" | "ollama" | "bedrock" | "azure". - Extended
OrchestratorConfigwithmcp_servers,handoff_policy,max_handoff_depth,routingfields. - Extended
TeamRunResultwithhandoffsfield. - Optional dependency groups in
pyproject.toml:google(google-generativeai>=0.8),bedrock(boto3>=1.34),azure(openai>=1.50),mcp(mcp>=1.0). - Examples:
examples/09_multi_provider.py,examples/10_mcp_tools.py,examples/11_agent_handoff.py,examples/12_intelligent_routing.py. - Test suites:
tests/test_providers.py(35 tests),tests/test_mcp.py(24 tests),tests/test_handoff.py(19 tests),tests/test_routing.py(18 tests).
- Orchestrator —
AnyCodenow manages MCP client lifecycles (connect/disconnect) as an async context manager, registers the handoff tool for agents that opt in, injects per-agent MCP tools into the tool registry, and applies routing decisions before task wave execution. - AgentRunner — detects handoff sentinel results in the tool loop; on detection, yields a
StreamEvent(type="handoff")and terminates the turn. - ToolRegistry — added
register_from_mcp()for batch MCP tool registration andderegister_prefix()for cleanup on server disconnect. - providers/openai.py — refactored to import shared mapping logic from
_openai_compat.py(no behavior change).
0.3.0 - 2026-04-05
- Pluggable Memory module (
src/anycode/memory/) — layered memory system with persistent KV stores and semantic vector search.SQLiteStore— async SQLite-backedMemoryStorewith WAL mode, metadata tracking, andcreated_at/updated_attimestamps.RedisStore— Redis-backedMemoryStorefor distributed deployments (optional[redis]extra).InMemoryVectorStore— TF-IDF + cosine similarity vector search with zero external dependencies.ChromaDBVectorStore— embedding-backed vector search via ChromaDB (optional[vector]extra).CompositeMemory— unified interface querying both KV and vector stores with auto-indexing support.create_memory_store()factory for config-driven backend creation fromMemoryConfig.
- Workflow Checkpointing module (
src/anycode/checkpoint/) — crash recovery for long-running DAG-based agent workflows.CheckpointManager— automatic checkpoint creation after each execution wave, spec-change detection via SHA-256 hash, and configurable auto-pruning.FilesystemCheckpointStore— human-readable JSON checkpoint files with atomic writes (tmp → rename).SQLiteCheckpointStore— WAL-mode SQLite backend for high-concurrency checkpoint storage.serialize_checkpoint()/deserialize_checkpoint()— deterministic round-trip serialization supporting all LLM message content types (TextBlock,ToolUseBlock,ToolResultBlock,ImageBlock).
- Human-in-the-Loop module (
src/anycode/hitl/) — approval gates for enterprise-grade agent workflows.ApprovalManager— config-driven approval enforcement with tool/task filtering and audit history tracking.CallbackApprovalGate— programmatic approval via user-provided async callable.StdinApprovalGate— interactive console approval with box-formatted prompts for CLI workflows.WebhookApprovalGate— HTTP webhook + polling approval for async and remote approval flows.format_approval_request()— box-formatted console output for approval prompts.
- New Pydantic types (all
frozen=True):VectorSearchResult,VectorStoreProtocol,MemoryConfig,CheckpointConfig,CheckpointData,CheckpointStoreProtocol,ApprovalConfig,ApprovalRequest,ApprovalResponse,ApprovalGateProtocol. - Optional dependency groups in
pyproject.toml:persistence(aiosqlite>=0.20),redis(redis[hiredis]>=5.0),vector(chromadb>=0.5). - Examples:
examples/06_pluggable_memory.py(SQLite, Redis, vector search, composite memory, SharedMemory DI),examples/07_checkpointing.py(filesystem/SQLite stores, serialization, spec-change detection, crash/resume),examples/08_hitl_approval.py(callback/stdin/webhook gates, config enforcement, timeouts, audit trail). - Test suites for the modules — unit tests (
test_memory.py,test_checkpoint.py,test_hitl.py) and integration tests (test_checkpoint_stores.py,test_composite_memory.py,test_full_pipeline.py).
- Orchestrator —
AnyCodenow saves checkpoints automatically after each execution wave viaCheckpointManager, supportsresume_fromparameter (accepts"latest"or a specific checkpoint ID) for crash recovery, and enforces task-level approval gates viaApprovalManagerbefore execution. - SharedMemory — accepts any
MemoryStorebackend via constructor injection; defaults toInMemoryStorefor full backward compatibility. - TeamConfig — new optional
memory_storeparameter for pluggable team memory backends. - Types — expanded
types.pywith 11 new Pydantic models for memory, checkpoint, and approval subsystems. All models remain frozen (immutable).
0.2.0 - 2025-04-05
- Telemetry module — OpenTelemetry-integrated tracing with
Tracer,Span, andConsoleExporterfor full lifecycle visibility across agent runs. IncludesMetricsCollectorwithTimerfor latency tracking andEventEmitterfor structured telemetry events. - Guardrails module — Runtime safety layer with
BudgetTrackerfor token/cost budget enforcement,HookRunnerwithLoggingHookfor turn-level lifecycle hooks, and composable content validators (MaxLengthValidator,ContainsValidator,BlocklistValidator). - Structured output module — Schema-constrained LLM responses via Pydantic models. Includes
schema_to_tool_defandschema_to_openai_response_formatfor cross-provider schema conversion,parse_structured_outputfor validated extraction, andbuild_retry_promptfor automatic recovery on malformed responses. - Production features example (
examples/05_production_features.py) demonstrating telemetry, guardrails, and structured output working together in a real workflow. - Test suite — Initial tests for guardrails, structured output, and telemetry modules.
- Dev scripts —
scripts/setup.sh,scripts/lint.sh,scripts/test.shfor reproducible local development. TraceConfig,SpanAttributes,GuardrailConfig,BudgetStatus,ValidationResult,OutputValidator,TurnHook,StructuredOutputConfig,StructuredRunResult, andStructuredAgentResulttypes.- Optional
telemetrydependency group for OpenTelemetry packages.
- Runner — Expanded
AgentRunnerwith guardrail integration, structured output support, trace context propagation, and improved turn-level error handling. Significant internal refactor for extensibility. - Orchestrator —
AnyCodeorchestrator now supports telemetry hooks, budget-aware scheduling, and structured task results. Task execution flow refactored for better observability. - Agent —
Agentclass extended with guardrail config, trace context, and structured output options. Agent state management improved. - Scheduler — Enhanced scheduling strategies with budget-aware task prioritization.
- Providers —
AnthropicAdapterandOpenAIAdapterupdated with streaming improvements, better error propagation, and structured output pass-through. - Tools —
ToolRegistryandToolExecutorrefined for safer execution, improved validation, and better error messages.bash,file_write, andgreptools hardened. - Collaboration —
Team,MessageBus, andSharedMemoryrefined with tighter type contracts and improved concurrency safety. - Types — Expanded
types.pywith all new Pydantic models. All models remain frozen (immutable).
- Pool concurrency edge case in
AgentPoolunder high parallelism. - Task dependency validation now catches circular references earlier.
0.1.0 - 2025-03-20
- Initial release of the AnyCode Python orchestration framework.
- Core agent system with
Agent,AgentRunner,AgentPool, andScheduler. AnyCodehigh-level orchestrator withTaskSpecdeclarative API.- Provider-agnostic LLM integration via
LLMAdapterprotocol with Anthropic and OpenAI adapters. - Team collaboration primitives:
Team,MessageBus,SharedMemory,InMemoryStore. - Dependency-aware task scheduling with topological sort.
- Built-in tool system:
bash,file_read,file_edit,file_write,grep. ToolRegistrywithdefine_toolfor runtime tool registration.Semaphore-based concurrency gating.- Token usage tracking with
merge_usage. - Four examples: solo worker, crew workflow, staged pipeline, hybrid tooling.
- Pydantic-based immutable type system (
frozen=Trueon all models).