Running several agents in parallel against one BYOM worker keeps running into limits that are sized for a shared server sandbox, not for work running on the user's own machine. This collects what one deployment hit over two days (a remote-bridge Code API behind a Cloudflare tunnel, one trusted-VM worker with six registered repositories, LibreChat agents with background subagents), with suggested defaults for each. #257 covers the related single-broad-root layout. This issue assumes one registered root per repository.
What we hit
| # |
Symptom |
Cause |
Evidence |
| 1 |
26% of workspace calls failed 503 WORKSPACE_QUEUE_TIMEOUT |
Queue wait was a fixed 30 s even when the caller's deadline was 55–125 s |
598 of ~2,300 calls in one day, up to 26 per minute. Fixed by #264 + #265 with a 90 s LibreChat budget: 12 of 141 afterwards, all in one burst |
| 2 |
The whole machine ran one call at a time across all six repositories |
CODEAPI_BRIDGE_MAX_WORKSPACE_LEASE_SLOTS and LIBRECHAT_CODE_WORKSPACE_LEASE_SLOTS both default to 1; with 1 slot the bridge takes one worker-wide lock |
Parallel agents in different repositories queued behind each other |
| 3 |
25% of workspace calls failed 429 RATE_LIMITED |
The exec limiter (MAX_REQUESTS=20 per RATE_LIMIT_WINDOW=30000, per principal) also counts every remote-bridge workspace call (read, list, search, edit, command) |
18 of 72 calls after slots were raised; peak 42 workspace calls in one minute from one user |
| 4 |
Agents working in different linked worktrees of one repository never overlap |
The isolation key is the registered root (plus an optional conversation instance). cwd and paths are never part of it, so <root>/.worktrees/a and .../b share one lane |
All subagents of one conversation also share one key, because they carry the parent's conversation identity |
| 5 |
Occasional 503 COMMAND_UNAVAILABLE ("Native executor setup failed before dispatch") |
The native pool throws when its root is busy instead of waiting |
2 in one day |
| 6 |
Contention is hard to diagnose |
The Workspace tool request completed log has workerId, operation, status, durationMs and deadlineBudgetMs, but no workspace, queue-wait time, slot or isolation key |
Attributing timeouts to a repository or burst needed guesswork |
| 7 |
Skill files fail to prime into the workspace |
File reference authorization rejected: Unauthorized file reference, reason upload_missing, for a skill file after a skill version bump; Batch upload file failed for skill files |
Seen in the BYOM API log; not investigated further |
Suggested defaults
| Setting |
Today |
Suggested |
| Exec limiter on remote-bridge workspace tools |
20 per 30 s per principal, shared with sandbox exec |
A separate limiter for /v1/workspace-tools/* on remote-bridge backends, default 120 per 30 s per principal. Keep 20 for server-side sandbox exec, where the server does the work |
CODEAPI_BRIDGE_MAX_WORKSPACE_LEASE_SLOTS |
1 |
8. It is only a ceiling, because the worker negotiates down to its own value |
LIBRECHAT_CODE_WORKSPACE_LEASE_SLOTS |
1 |
min(4, CPUs / 2), and a startup warning when more than one root is registered with 1 slot |
Queue wait without X-LibreChat-Workspace-Queue-Wait-Ms |
30 s |
Keep, as the safe default for older clients |
JOB_TIMEOUT behind a proxy |
5 min |
Document that a Cloudflare-proxied route cuts requests at ~100 s: use ≤ 90 s and set the client budget to match. We verified 90 s end to end: a 77.5 s command returned 200, and a read queued 49 s returned 200 |
| Conversation worktrees |
Opt-in |
Recommend them in the multi-agent docs, and log at startup when they are off but more than one agent can target a root |
Proposals
- Separate the remote-bridge workspace-tool limiter from sandbox exec, with the default above. Optionally weight reads/lists/searches below commands and edits. Keep the typed body (
rate_limited, retry_after_seconds), and document that a 429 guarantees the operation was not started, so clients can retry it safely. LibreChat will retry a 429 within its budget the way it already retries WORKSPACE_QUEUE_TIMEOUT.
- Raise the lease-slot defaults as above.
- Per-linked-worktree lanes. Add an explicit
worktree field (one path segment) to workspace-tool requests. The worker verifies that <root>/.worktrees/<name> is a real linked worktree of the root: its .git file points to <root>/.git/worktrees/<name>, commondir resolves to <root>/.git, realpath containment holds, and its identity is pinned. The worker then re-bases cwd/paths onto the worktree and narrows the sandbox to it plus the shared object store. Scheduling becomes hierarchical: a worktree lane conflicts only with itself and its root, while root-level requests, git worktree add/remove/prune and gc/maintenance stay exclusive. An explicit field is better than inferring from cwd, because a command can cd anywhere in the root, and the Code API has to know the key before admission. docs/remote-bridge/projects.md already lists the shared-metadata hazards this has to handle.
- Queue on a busy native pool instead of throwing
COMMAND_UNAVAILABLE. If it must fail, fail with the typed not-started 503 so clients retry.
- Log contention fields on every workspace-tool completion: workspace id, a hashed isolation key,
queueWaitMs, slot index and admission outcome. Also expose slot occupancy as a metric.
- Investigate the skill-file priming failures (
upload_missing after a skill version bump).
Acceptance
- Two agents in two repositories on one worker overlap with default settings.
- Two agents in two linked worktrees of one repository overlap when they pass
worktree, while a root-level command waits for both.
- A burst of 40 workspace calls per minute from one user does not return 429 under the new remote-bridge default.
- Contention can be attributed from the completion logs alone.
Running several agents in parallel against one BYOM worker keeps running into limits that are sized for a shared server sandbox, not for work running on the user's own machine. This collects what one deployment hit over two days (a remote-bridge Code API behind a Cloudflare tunnel, one trusted-VM worker with six registered repositories, LibreChat agents with background subagents), with suggested defaults for each. #257 covers the related single-broad-root layout. This issue assumes one registered root per repository.
What we hit
503 WORKSPACE_QUEUE_TIMEOUTCODEAPI_BRIDGE_MAX_WORKSPACE_LEASE_SLOTSandLIBRECHAT_CODE_WORKSPACE_LEASE_SLOTSboth default to 1; with 1 slot the bridge takes one worker-wide lock429 RATE_LIMITEDexeclimiter (MAX_REQUESTS=20perRATE_LIMIT_WINDOW=30000, per principal) also counts every remote-bridge workspace call (read, list, search, edit, command)cwdand paths are never part of it, so<root>/.worktrees/aand.../bshare one lane503 COMMAND_UNAVAILABLE("Native executor setup failed before dispatch")Workspace tool request completedlog hasworkerId,operation,status,durationMsanddeadlineBudgetMs, but no workspace, queue-wait time, slot or isolation keyFile reference authorization rejected: Unauthorized file reference, reasonupload_missing, for a skill file after a skill version bump;Batch upload file failedfor skill filesSuggested defaults
exec/v1/workspace-tools/*on remote-bridge backends, default 120 per 30 s per principal. Keep 20 for server-side sandbox exec, where the server does the workCODEAPI_BRIDGE_MAX_WORKSPACE_LEASE_SLOTSLIBRECHAT_CODE_WORKSPACE_LEASE_SLOTSX-LibreChat-Workspace-Queue-Wait-MsJOB_TIMEOUTbehind a proxyProposals
rate_limited,retry_after_seconds), and document that a 429 guarantees the operation was not started, so clients can retry it safely. LibreChat will retry a 429 within its budget the way it already retriesWORKSPACE_QUEUE_TIMEOUT.worktreefield (one path segment) to workspace-tool requests. The worker verifies that<root>/.worktrees/<name>is a real linked worktree of the root: its.gitfile points to<root>/.git/worktrees/<name>,commondirresolves to<root>/.git, realpath containment holds, and its identity is pinned. The worker then re-basescwd/paths onto the worktree and narrows the sandbox to it plus the shared object store. Scheduling becomes hierarchical: a worktree lane conflicts only with itself and its root, while root-level requests,git worktree add/remove/pruneand gc/maintenance stay exclusive. An explicit field is better than inferring fromcwd, because a command cancdanywhere in the root, and the Code API has to know the key before admission.docs/remote-bridge/projects.mdalready lists the shared-metadata hazards this has to handle.COMMAND_UNAVAILABLE. If it must fail, fail with the typed not-started 503 so clients retry.queueWaitMs, slot index and admission outcome. Also expose slot occupancy as a metric.upload_missingafter a skill version bump).Acceptance
worktree, while a root-level command waits for both.