Summary
ChatService._run_impl's in-run inbox continuation does not reuse the parked guard that the dispatcher wake entry already has. When an agent is parked on an ASKING (user confirmation) or SUBMITTED (external execution) tool call AND its inbox happens to hold a queued payload, the run re-enters the loop with input_msg=None and calls Agent.reply_stream(inputs=None), which Agent._check_incoming_event rightly rejects with ValueError. The fallout is a spurious ReplyEnd(ERROR), a misleading team-leader failure reminder (the leader then self-answers a failure report), and an empty error reply in the worker transcript — while the parked state itself stays perfectly clean (confirm/resume still works).
Affected version: agentscope 2.0.8 (agentscope/app/_service/_chat.py, agentscope/app/_bus_ops.py).
Code contrast
Dispatcher wake entry HAS the guard (_chat.py ~L1123):
if self._skip_parked_wakeup(session_id, agent, input_msg):
return
_skip_parked_wakeup's docstring (~L711-723) states the intent:
If the agent is parked on an ASKING or SUBMITTED tool call ... another None run would hit Agent._check_incoming_event, which rightly rejects None when there is something to confirm, and fail noisily. The inbox content is safe to leave queued: whenever the user does confirm ..., the resuming run's next reasoning step lets InboxMiddleware drain it naturally.
The in-run continuation point has NO such check (_chat.py ~L1359-1365):
if not await has_pending_inbox_or_release(self._message_bus, session_id):
released = True
break
input_msg = None # <- re-enters Case A reply_stream(inputs=None) while parked
Two entry points, same illegal re-entry (parked + None), inconsistent handling: the wake entry silently drops it, the continuation entry blows up with ERROR.
Reproduction (team scenario, minimal)
- Leader invites a worker (AgentInvite returns immediately; leader turn ends asynchronously).
- Worker calls a permission-gated tool -> tool_call enters ASKING, REQUIRE_USER_CONFIRM emitted, worker reply_stream ends normally, agent parked.
- Before the user confirms, another member delivers a team message into the worker's inbox (
deliver_to_inbox; worker is the registered consumer so the producer only pushes, no wake).
- Worker run reaches the turn-tail
has_pending_inbox_or_release -> True -> input_msg = None -> loop re-entry.
- Case A
reply_stream(inputs=None) -> _check_incoming_event raises ValueError.
- Observed: spurious ReplyEnd(ERROR) + leader failure reminder + empty error reply;
derive_parked_status still AWAITING_PERMISSION (parked state intact).
Non-team repro: any single session parked on a confirmation card + any inbox payload (e.g. a channel message) triggers the same path.
Suggested fix
Fold the wake entry's parked check into the in-run continuation condition (matches the docstring intent):
if not await has_pending_inbox_or_release(self._message_bus, session_id) or self._is_parked(agent):
released = True
break
input_msg = None
_is_parked(agent) can reuse the existing predicate agent.state.has_awaiting_tool_calls(agent.name) (same source as _check_incoming_event's raise condition; _skip_parked_wakeup's context-tail check could converge onto it too).
Note: with released = True the finally skips abandon_inbox_consumer, so the consumer registration would dangle — the upstream fix should release the registration in the parked branch (or route through abandon_inbox_consumer, whose wake is correctly dropped by _skip_parked_wakeup).
Our interim mitigation (patch layer, for reference)
team_parked_guard.py in our deployment wraps Agent.reply_stream (marks session parked when the wrapped generator ends with awaiting tool calls; clears on any non-None inputs) and wraps has_pending_inbox_or_release (returns False for parked sessions, releasing the run and returning the consumer registration under the inbox lock). 14 unit tests green; eliminated a ~50% crash rate in our team HITL flows.
Happy to contribute a PR with the fix + regression tests if welcome.
Summary
ChatService._run_impl's in-run inbox continuation does not reuse the parked guard that the dispatcher wake entry already has. When an agent is parked on anASKING(user confirmation) orSUBMITTED(external execution) tool call AND its inbox happens to hold a queued payload, the run re-enters the loop withinput_msg=Noneand callsAgent.reply_stream(inputs=None), whichAgent._check_incoming_eventrightly rejects withValueError. The fallout is a spuriousReplyEnd(ERROR), a misleading team-leader failure reminder (the leader then self-answers a failure report), and an empty error reply in the worker transcript — while the parked state itself stays perfectly clean (confirm/resume still works).Affected version: agentscope 2.0.8 (
agentscope/app/_service/_chat.py,agentscope/app/_bus_ops.py).Code contrast
Dispatcher wake entry HAS the guard (
_chat.py~L1123):_skip_parked_wakeup's docstring (~L711-723) states the intent:The in-run continuation point has NO such check (
_chat.py~L1359-1365):Two entry points, same illegal re-entry (parked + None), inconsistent handling: the wake entry silently drops it, the continuation entry blows up with ERROR.
Reproduction (team scenario, minimal)
deliver_to_inbox; worker is the registered consumer so the producer only pushes, no wake).has_pending_inbox_or_release-> True ->input_msg = None-> loop re-entry.reply_stream(inputs=None)->_check_incoming_eventraises ValueError.derive_parked_statusstill AWAITING_PERMISSION (parked state intact).Non-team repro: any single session parked on a confirmation card + any inbox payload (e.g. a channel message) triggers the same path.
Suggested fix
Fold the wake entry's parked check into the in-run continuation condition (matches the docstring intent):
_is_parked(agent)can reuse the existing predicateagent.state.has_awaiting_tool_calls(agent.name)(same source as_check_incoming_event's raise condition;_skip_parked_wakeup's context-tail check could converge onto it too).Note: with
released = Truethefinallyskipsabandon_inbox_consumer, so the consumer registration would dangle — the upstream fix should release the registration in the parked branch (or route throughabandon_inbox_consumer, whose wake is correctly dropped by_skip_parked_wakeup).Our interim mitigation (patch layer, for reference)
team_parked_guard.pyin our deployment wrapsAgent.reply_stream(marks session parked when the wrapped generator ends with awaiting tool calls; clears on any non-None inputs) and wrapshas_pending_inbox_or_release(returns False for parked sessions, releasing the run and returning the consumer registration under the inbox lock). 14 unit tests green; eliminated a ~50% crash rate in our team HITL flows.Happy to contribute a PR with the fix + regression tests if welcome.