Skip to content

[Bug] In-run inbox continuation lacks parked guard: parked agent + queued inbox => ValueError, spurious ERROR reply and misleading leader failure reminder #2842

Description

@iamsee

Summary

ChatService._run_impl's in-run inbox continuation does not reuse the parked guard that the dispatcher wake entry already has. When an agent is parked on an ASKING (user confirmation) or SUBMITTED (external execution) tool call AND its inbox happens to hold a queued payload, the run re-enters the loop with input_msg=None and calls Agent.reply_stream(inputs=None), which Agent._check_incoming_event rightly rejects with ValueError. The fallout is a spurious ReplyEnd(ERROR), a misleading team-leader failure reminder (the leader then self-answers a failure report), and an empty error reply in the worker transcript — while the parked state itself stays perfectly clean (confirm/resume still works).

Affected version: agentscope 2.0.8 (agentscope/app/_service/_chat.py, agentscope/app/_bus_ops.py).

Code contrast

Dispatcher wake entry HAS the guard (_chat.py ~L1123):

if self._skip_parked_wakeup(session_id, agent, input_msg):
    return

_skip_parked_wakeup's docstring (~L711-723) states the intent:

If the agent is parked on an ASKING or SUBMITTED tool call ... another None run would hit Agent._check_incoming_event, which rightly rejects None when there is something to confirm, and fail noisily. The inbox content is safe to leave queued: whenever the user does confirm ..., the resuming run's next reasoning step lets InboxMiddleware drain it naturally.

The in-run continuation point has NO such check (_chat.py ~L1359-1365):

if not await has_pending_inbox_or_release(self._message_bus, session_id):
    released = True
    break
input_msg = None   # <- re-enters Case A reply_stream(inputs=None) while parked

Two entry points, same illegal re-entry (parked + None), inconsistent handling: the wake entry silently drops it, the continuation entry blows up with ERROR.

Reproduction (team scenario, minimal)

  1. Leader invites a worker (AgentInvite returns immediately; leader turn ends asynchronously).
  2. Worker calls a permission-gated tool -> tool_call enters ASKING, REQUIRE_USER_CONFIRM emitted, worker reply_stream ends normally, agent parked.
  3. Before the user confirms, another member delivers a team message into the worker's inbox (deliver_to_inbox; worker is the registered consumer so the producer only pushes, no wake).
  4. Worker run reaches the turn-tail has_pending_inbox_or_release -> True -> input_msg = None -> loop re-entry.
  5. Case A reply_stream(inputs=None) -> _check_incoming_event raises ValueError.
  6. Observed: spurious ReplyEnd(ERROR) + leader failure reminder + empty error reply; derive_parked_status still AWAITING_PERMISSION (parked state intact).

Non-team repro: any single session parked on a confirmation card + any inbox payload (e.g. a channel message) triggers the same path.

Suggested fix

Fold the wake entry's parked check into the in-run continuation condition (matches the docstring intent):

if not await has_pending_inbox_or_release(self._message_bus, session_id) or self._is_parked(agent):
    released = True
    break
input_msg = None

_is_parked(agent) can reuse the existing predicate agent.state.has_awaiting_tool_calls(agent.name) (same source as _check_incoming_event's raise condition; _skip_parked_wakeup's context-tail check could converge onto it too).

Note: with released = True the finally skips abandon_inbox_consumer, so the consumer registration would dangle — the upstream fix should release the registration in the parked branch (or route through abandon_inbox_consumer, whose wake is correctly dropped by _skip_parked_wakeup).

Our interim mitigation (patch layer, for reference)

team_parked_guard.py in our deployment wraps Agent.reply_stream (marks session parked when the wrapped generator ends with awaiting tool calls; clears on any non-None inputs) and wraps has_pending_inbox_or_release (returns False for parked sessions, releasing the run and returning the consumer registration under the inbox lock). 14 unit tests green; eliminated a ~50% crash rate in our team HITL flows.

Happy to contribute a PR with the fix + regression tests if welcome.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions