Skip to content

fix(sidecar): queue bursts instead of refusing them at the accept backlog - #40

Open
linhdmn wants to merge 3 commits into
mainfrom
fix/sidecar-accept-queue
Open

linhdmn wants to merge 3 commits into
mainfrom
fix/sidecar-accept-queue

Conversation

@linhdmn

@linhdmn linhdmn commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Two commits, one root cause

The sidecar that runs on this Mac (launchd job ai.hermes.laya-sidecar, port 8092) and the copy in this repo had drifted, and the drift was not cosmetic. Three of the differences are the bugs this sidecar was originally fixed for, so anyone who ran the repo copy — or refreshed the deployed file from the repo — got the old behaviour back.

1. Queue bursts instead of refusing them at the accept backlog

socketserver's default TCPServer.request_queue_size is 5, and that is the LISTEN backlog, not the worker pool: connections past the fifth are refused by the kernel before any worker thread starts. Measured against the deployed sidecar on this M2 Pro box, 120 concurrent requests lost 40 at ~9 rps, while 60 concurrent lost none at ~35 rps — the loss was the backlog, not the model.

class Server(ThreadingHTTPServer): request_queue_size = 128 turns that loss into the queue the server already provides, and stays bounded so a saturated backlog still sheds load rather than growing threads without limit.

scripts/test_laya_sidecar.py asserts the constant and that main() builds Server rather than a plain ThreadingHTTPServer. The second assertion is the one that matters: the class can survive untouched while nothing uses it, which no reader or constant-only check would catch. Both reverts were verified to fail the test.

2. Converge the repo copy with the deployed one

repo copy (before) deployed
router Router(preload=True) — loads all three checkpoints, 5.5 GB footprint, peak 6.1 GB lazy, LAYA_MAX_LOADED=1, 2.4 GB
predict unlocked one _PREDICT_LOCK
state body.get("state", "") only falls back to text

On a 16 GB Mac that already swaps, the repo copy costs more than twice the memory for the same answers. The missing lock is worse than a slowdown: two threads calling router.predict at once hit the same Metal command buffer and kill the process (A command encoder is already encoding to this command buffer — measured, 8 concurrent requests died after 2 replies, 20 sequential were clean). And because System One callers send text while Laya's own API calls it state, the repo copy answered a blank string for every one of them.

Also brought over: the /health payload reporting loaded + preload (so a silent fallback away from the expected head is visible from outside the process), the 404 branch for unknown GET paths, and --host / LAYA_HOST / PORT support.

The AST is now identical to the deployed copy with docstrings stripped, so the next refresh overwrites nothing. Verified against a stub Router on a throwaway port: health payload, 404 branch, text input, and a 60-request burst all 60/60. The live sidecar on 8092 was not touched.

Note on the test

Two measurement traps, both hit while writing it:

  • A fast handler hides the backlog bug entirely. With a 10 ms stub the same 120-request burst passed 120/120 on the unfixed code — the accept loop kept up, so the backlog never filled. The check needs a stand-in for the slow work (300 ms reproduced it reliably).
  • runpy.run_path cannot be patched this way. It returns a copy of the module globals, so mock.patch.object(mod, "Server", …) raises AttributeError and patching the returned dict never reaches main()'s real global — which silently produced a test asserting nothing. Load it as a real module with importlib.util.spec_from_file_location instead, under a name that is not __main__.

…klog

socketserver's default `request_queue_size` is 5, and that is the LISTEN
backlog, not the worker pool: connections past the fifth are refused by the
kernel before any worker thread starts. Measured against the deployed sidecar
on this M2 Pro box, 120 concurrent requests lost 40 at ~9 rps while 60 lost
none at ~35 rps -- the loss was the backlog, not the model.

`class Server(ThreadingHTTPServer): request_queue_size = 128` turns that loss
into the queue the server already provides. It stays bounded, so a saturated
backlog still sheds load rather than growing threads without limit.

The test asserts both the constant AND that `main()` builds `Server` rather
than a plain `ThreadingHTTPServer` -- the class can survive while nothing uses
it, which no reader or constant-only assertion would catch. Both reverts were
verified to fail the test.

Also: the module docstring now says this file is the SOURCE copy and that the
installed `~/.local/share/laya-sidecar/laya-sidecar.py` is the one launchd
runs on 8092, because the two have drifted and editing only the repo copy
changes nothing until it is copied over.
The copy in the repo had drifted from the file launchd actually runs
(`~/.local/share/laya-sidecar/laya-sidecar.py`), and the drift was not
cosmetic -- three of the differences are the bugs this sidecar was fixed for:

- `Router(preload=True)` loads ALL THREE checkpoints: measured phys_footprint
  5.5 GB, peak 6.1 GB, on a 16 GB Mac that already swaps. The deployed copy is
  lazy with LAYA_MAX_LOADED=1 and measures 2.4 GB.
- no predict lock. Two threads calling `router.predict` at once hit the same
  Metal command buffer and kill the process ("A command encoder is already
  encoding to this command buffer"; measured 8 concurrent requests died after
  2 replies, 20 sequential were clean).
- `state = body.get("state", "")` only. System One callers send `text`, so
  every such request answered a blank string.

Also brought over: the `/health` payload reporting `loaded` + `preload` (so a
silent fallback away from the expected head is visible from outside), the
404 branch for unknown GET paths, and `--host` / LAYA_HOST / PORT env
support.

The AST is now identical to the deployed copy with docstrings stripped, so
the next refresh overwrites nothing. Verified against a stub Router on a
throwaway port: health payload, 404 branch, `text` input, and a 60-request
burst all 60/60.
The docstring claimed "one resident checkpoint is ~2.4 GB" and nothing about
what the launchd plist actually decides. Measured on this M2 Pro / 16 GB box
with the deployed configuration (LAYA_EAGER=1, LAYA_MAX_LOADED=2,
LAYA_PRELOAD=english): 3.2 GB phys_footprint with english + multilingual hot,
peak 4.3 GB. A reader sizing a second instance off "2.4 GB" would be off by a
third.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant