Replies: 3 comments 10 replies
|
@iluv7 Great point. When designing the new The OpenAI Response API supports mask-based tool choice — it sends the full JSON schema upfront but controls which tools are active via masks during inference, preserving the KV cache. This is actually why we evolved That said, mask-based caching is currently only supported by the Response API, which is why calling Bringing it back to your proposal — I find the idea genuinely interesting. Delegating tool discovery to the agent itself is a thoughtful application-layer solution. That said, I have a couple of concerns:
That said, I think this direction is worth pursuing. Let me sketch a more complete design, grounded in two principles: the agent should have deterministic awareness of its global tool inventory — a stable understanding of its capability boundaries — and deterministic awareness of currently active tool schemas at runtime. Assuming the total tool count is fixed and tools are pre-organized into named groups (dynamic registration is a separate discussion), the This avoids any dependency on To address both, we could replace It's worth being upfront about what this demands from the model:
很好的观察。在设计新的 OpenAI Response API 支持基于掩码的工具选择——它一次性发送完整的 JSON schema,但在推理过程中通过掩码控制哪些工具处于激活状态,从而保留 KV cache。这也是我们将 不过,掩码缓存目前仅 Response API 支持,这也是为什么对其他 provider 调用 回到你的提案——我觉得这个思路很有意思,把工具发现的能力委托给 agent 自身,是一种很好的应用层解法。但我有几点顾虑: 首先,将 JSON schema 后置到工具执行结果中,本质上和放在 system prompt 里的差异并不大,但对模型行为的实际影响目前还不确定:
这个方向值得深入探索。我想在此基础上提出一个更完整的设计思路,核心围绕两条原则:
前提:假设总工具数量固定,且已按组预先分类(动态新增工具的情况单独讨论)。
当 这样做的好处是:不依赖 但这个方案存在两个明显问题:
针对这两个问题,可以进一步改进:
需要坦诚的是,这套方案对模型的要求不低:
|
|
The two concerns you called out are the right acceptance gates: a stable top-level tool surface is useful only if discovery recall and the extra request path are measured independently. Disclosure: I maintain MCP Lens, a DeepSeek Harness implementation of the fixed That evidence is deliberately narrow: covered-call lexical retrieval, no model calls, not the official MCP-Atlas benchmark, and not an AgentScope result. It does show why cache shape and search quality should be separate benchmark axes. For AgentScope I would freeze four arms on the same tool catalog and tasks:
Then record: first-request tool-schema JSON bytes; prefix hash continuity; target Recall@1/5 before any model run; model request count; successful target execution; argument-validation failures; and post-compaction recovery. A search result must not grant authority: For global capability awareness, I would keep only bounded inventory metadata stable (groups, counts, short descriptions), not every tool name/schema. Exact schemas should be refreshed near the tail when needed; after compaction, retain discovered identities and make re-search explicit rather than replaying all schemas. If useful, I can contribute a keyless AgentScope-native fixture for this matrix. It should make no token, latency, cache-cost, or model-quality claim unless those are measured in a separate provider run. |
|
Update: I noticed that both Anthropic's and OpenAI's Response APIs support mid-conversation system messages that can add/remove tools on the fly. This means tool group management could potentially be implemented via a special control message, rather than updating the tools schema directly. However, the current tool context is stored independently from the message context ( That said, this might not necessarily be a bad thing. We've already run into a related issue on the frontend side: when the frontend subscribes to agent events and needs to be notified of changes like That approach does introduce a new problem, though: after compressing the conversation history, the tools' current schema would be lost. We'd need to either recompute the latest tools state right after compression, or compress the entire tools change history into a single consolidated change. This ties directly into how we implement context compression as a whole. |
Uh oh!
There was an error while loading. Please reload this page.
Problem
AgentScope's
ToolGroupusesreset_toolsto switch active tool groups dynamically, reducing the number of tools visible to the model in each call.reset_toolsdoes not return or append the newly activated tools' schemas itself. It only updates the activation state and returns the activated groups' usage instructions as a normal tool result.The complete schemas are added separately when AgentScope prepares the next model call:
The relevant implementation points are:
ResetTools.call()updatesactivated_groupsand returns group usage instructions, not tool schemas;Agent._prepare_model_input()regeneratestoolsfrom the active groups;Toolkit._get_available_tools()collects basic tools, tools from active groups, and MCP tools.AgentScope eventually constructs the following input:
{ "messages": messages, "tools": tools, }Although
toolsappears aftermessagesin this Python dictionary, that does not mean tool definitions are placed at the end of the model's effective input. Each provider parses and organizes these fields independently.The appended
reset_toolsresult is cache-friendly by itself. The cache boundary is introduced by the separate reconstruction of the top-leveltoolsbefore the next request. The core issue is therefore:Why this affects caching
Prompt caching relies on a continuous, identical input prefix. Appending new content to the end of the conversation is cache-friendly:
Changing tool definitions may instead introduce a difference before the conversation history:
Anthropic's documentation explicitly describes its cache hierarchy as:
Changing tool definitions invalidates the tools, system, and messages caches. Therefore, with Anthropic, the current
reset_toolsbehavior creates an explicit cache boundary.References:
Proposed approach
With on-demand tool loading, newly discovered tool information can be appended to the conversation history without changing the top-level
tools. I suggest keeping the current behavior while adding a provider-independent stable mode:native: keep usingToolGroup + reset_toolsto expose real tools dynamically;stable: expose only a fixedsearch_tools + execute_toolpair at the top level.search_toolssearches AgentScope's complete tool registry and appends the candidates as a tool result.execute_toolresolves the real tool by name and reuses the existing Toolkit execution path.Tool recall in long contexts
As a conversation grows, earlier tool search results move toward the middle of the context, where the model may use their schemas less reliably. Context compaction may also remove these results entirely. The
stablemode should therefore not depend on the model remembering discovered tools indefinitely.search_toolsshould support repeated keyword search and direct loading by tool name;execute_toolargument validation fails, it should return a clear hint telling the model to search again and retry;Repeated searches only append content to the conversation and do not modify the existing cached prefix.
The top-level
toolstherefore remain unchanged before and after discovery:This approach only changes how tools are exposed to the model. It does not change AgentScope's existing tool execution logic and does not require support for any provider-specific deferred-loading protocol. The default remains
nativefor backward compatibility.This is my current understanding of the problem and a possible optimization direction. I may have missed something or misunderstood part of the behavior, so feedback and discussion are very welcome.
All reactions