Thanks to visit codestin.com
Credit goes to github.com

Skip to content

🤖 perf: bind xum server listener before startup recovery; stop per-task config.json reloads - #4058

Merged
ibetitsmike merged 28 commits into
mainfrom
mike/server-bind-before-recovery
Sep 3, 2026
Merged

🤖 perf: bind xum server listener before startup recovery; stop per-task config.json reloads#4058
ibetitsmike merged 28 commits into
mainfrom
mike/server-bind-before-recovery

Conversation

@ibetitsmike

@ibetitsmike ibetitsmike commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Summary

xum server now binds its port after the fast core init plus agent-task restart recovery, and runs the O(workspaces)/O(reported tasks) startup housekeeping in the background, so time-to-listen depends on the number of active tasks rather than on deployment size. The housekeeping loops also stop re-parsing config.json once per reported task, the terminal-attention pending scan is bounded-concurrent, slow startup passes are logged at warn, and AnalyticsService.dispose() waits for its worker so a signal during the startup DuckDB sync no longer aborts the process. Fixes #4055.

Background

On a deployment with 1,535 workspaces (343 reported tasks, 18 GB sessions dir) xum server took 13 minutes to bind its port because src/cli/server.ts awaited serviceContainer.initialize() before startServer(), and TaskService.initialize spent ~776 s in sequential recovery: a sequential scan of every session dir for pending terminal-attention records (306 s), patch-generation recovery with a synchronous 2.47 MB config parse per reported task (205 s), and reported-task cleanup with another parse per task (131 s). Clients got connection refused the whole time and Coder marked the app unhealthy. See #4055 for the measured timeline.

The part of startup that actually has to precede clients is small: reconciling execution handles, fixing stale starting tasks, draining the queue, and prompting the handful of awaiting_report/running tasks to continue. Everything that scales with deployment size (patch artifacts and cleanup for every reported task, the session-dir scan, workflow garbage sweeps, chat restart retries, orphan sweeps) is housekeeping that is idempotent and already runs against live state at runtime.

Implementation

  • TaskService.initialize() is split into recoverInterruptedTasks() (execution-handle reconciliation, stale-starting fixup, queue drain, inactive-workflow-owner prepass, awaiting_report/running restart prompts; bounded by the number of active tasks) and runStartupHousekeeping({ signal }) (patch-artifact recovery, best-of delivery, reported-task cleanup, workflow archive sweep, terminal-attention sweeps and drains). initialize() still runs both for desktop startup and tests.
  • ServiceContainer.initialize() is split the same way: initializeCore() (extensionMetadata, telemetry, policy, experiments, then taskService.recoverInterruptedTasks()) and runStartupHousekeeping() (workspaceService.initialize(), taskService.runStartupHousekeeping(), then the idle-compaction/heartbeat/agent-status starts and the completion log). server.ts awaits initializeCore(), starts the server, then runs runStartupHousekeeping() in the background with an error handler; dispose() aborts the housekeeping signal and joins the in-flight step (bounded by STARTUP_HOUSEKEEPING_JOIN_TIMEOUT_MS, 500 ms) so a shutdown mid-housekeeping stops at the next step boundary before the services it uses are torn down, and never starts periodic services against disposed dependencies; the transient chat-recovery sessions that housekeeping scheduled are disposed right after that join (their chains re-check disposed before every dispatch), with backgroundProcessManager.beginShutdown() still the first teardown step so none of those disposals can erase persisted armed-monitor records, and workspaceService.initialize({ signal }) schedules none once shutdown began. Desktop startup is unchanged (initialize()).
  • Because task recovery completes before the listener exists, no client can race its status transitions or sends, so it needs no compare-and-set or admission probes. Only the post-listen housekeeping can overlap with clients, and every step there re-checks live state before it mutates:
    • Reported-task cleanup screens candidates on the loop's config snapshot and confirms eligibility on fresh config as a beforeRemove precondition that WorkspaceService.remove() evaluates inside its task-tree lifecycle lock (the lock reactivation, re-parenting, and task_stop all mutate under). That confirmation also rejects a task whose execution mirror (taskExecutionStatus) is starting or running, which is the state an existing-workspace turn leaves before its stream registers, and the lineage walk continues from the parent the live confirmation saw (a re-parent since the screen).
    • Best-of finalization of a crash-left parent partial (tryFinalizePendingTaskToolCallInPartial, also the runtime child-report path) is now a compare-and-set: HistoryService.updatePartialIfMessageIdMatches() re-reads the partial under the per-workspace file lock and writes only while it is still the same message with the task call still pending and no stream running. A parent turn that starts meanwhile (client send, or a resumed child's task_send_message) commits that partial and writes its own under a new id, so the finalization is dropped instead of resurrecting or overwriting it. commitPartial() is now one transaction under the same pair of locks the CAS holds (the workspace mutex and the cross-process history write lock): snapshot, history append/update/delete, and partial delete cannot interleave with the CAS or with another backend's commit of the same partial, so a finalization is either included in the commit or declined, never appended pre-update and deleted. The lock-held bodies of appendToHistory/updateHistory/deleteMessage/getHistoryFromLatestBoundary were extracted for that (their public wrappers are unchanged).
    • Patch-artifact recovery passes the loop snapshot via maybeStartGeneration(..., { config }); a snapshot child with no artifact yet, or with a crash-left pending one, is re-checked on live config before generation runs, and skipped when it was removed or reactivated (active workspace-turn execution) or is streaming. Reactivated tasks get their artifact from the existing continuation refresh when that execution settles.
    • Chat restart recovery re-reads config right before scheduling (skipping workspaces archived or removed since the metadata read), and archive() disposes a workspace's still-pending transient recovery session once archivedAt is durable; a recovery scheduled on a session a client had already created is not disposed by archive, so the two internal dispatch points of startup recovery (pending compaction follow-up, startup auto-retry) re-read the durable archived state right before dispatching, mirroring the guard WorkspaceService.sendMessage already applies; the orphan scratch-workdir sweep gained the session sweep's grace window and fresh-config recheck; the DevTools-log sweep re-checks live archive state under the task-tree lock before each deletion.
  • TerminalAttentionStore.listPendingOwnerWorkspaceIds() scans session dirs through an AsyncSemaphore(16) and short-circuits per owner on the first pending record.
  • [startup] ... completed logs use log.warn when totalMs exceeds SLOW_STARTUP_WARN_THRESHOLD_MS (30 s, src/constants/startup.ts).
  • AnalyticsService.dispose() waits for the analytics worker to exit instead of tearing it down mid-native call. With the graceful SIGINT handler installed while the startup DuckDB sync may still be running, exiting the process mid-sync tore the worker down inside native DuckDB and aborted the process (Napi::Error -> SIGABRT); UAT reproduced this 7/8 times before the fix.

Validation

Remote dogfood UAT ran on a Coder Agents workspace against an earlier head of this branch (one where all recovery, including the task prompts, ran post-listen) and main with identical synthetic roots (config.json with 1,500 to 8,000 workspaces, 300 to 2,000 reported tasks, session dirs with terminal-attention records), driving the real UI with a real model:

  • Time-to-first-/health 200 at N=1,500: 2.7 to 3.9 s (branch) vs 12.8 s (main); at N=8,000 with 1,500 reported tasks: 2.2 s, while background housekeeping took ~57 s. The current head additionally awaits the task recovery phase before binding; that phase is bounded by the number of active tasks (0 to 2 in the measured deployments, sub-second) rather than by N.
  • UI, /api/docs, and a 150-request hammer stayed healthy during housekeeping; a workspace created during housekeeping persisted with a clean 2-turn chat; kill -9 mid-housekeeping restarted cleanly; a second server on the same root was refused by the lockfile.
  • cleanupReportedTasksMs stayed flat as config grew 16x (111 to 173 ms vs 701 to 8,479 ms on main); terminalAttentionDrainMs over 8,001 session dirs was 209 ms; pendingTerminalAttentionOwnerWorkspaceCount was exact (3 of 3), including with a corrupt record and a stray file in sessions/.
  • SIGINT/SIGTERM during the startup analytics sync: 0/26 aborts after the dispose fix (7/8 before), all exit 0.
  • Unit coverage for the split: initializeCore waits for task recovery and does not run housekeeping; recoverInterruptedTasks resumes running tasks without touching reported ones while runStartupHousekeeping prunes reported tasks without resuming anything; best-of finalization leaves a parent partial that a live turn replaced mid-finalization untouched (red-green); dispose aborts in-flight housekeeping before the periodic services start; the snapshot patch path skips removed and reactivated children (red-green).

Risks

  • Requests can arrive while housekeeping runs. Its steps are the same passes the runtime already executes against live state (cleanup rechecks, terminal-attention sweeps on a timer, chat retry sessions) plus the live re-checks listed above; task recovery itself finishes before the listener opens. Desktop behavior is unchanged.
  • The early lockfile check in server.ts remains a fast-fail nicety, as on main: task recovery runs before startServer() acquires the lock, so two servers racing on one root could both attempt recovery; startServer() still refuses the second one.
  • UAT (Medium, not fixed here): SIGINT/SIGTERM while the startup analytics sync is mid-checkpoint now exits cleanly but slowly (11.5 to 13.6 s at N=1,500 and N=8,000). The main thread parks while one worker thread finishes the DuckDB checkpoint/close, and the 5 s Cleanup timed out, forcing exit path does not take effect in that state. A docker stop with a 10 s grace period would SIGKILL mid-checkpoint. Bounding or interrupting the checkpoint is a follow-up.
  • Residual: when patch generation actually runs at startup (artifacts left pending by a crash), each completion still triggers a cleanup recheck that reloads config, so patchGenerationRecoveryMs still grows with config size in that case. It is off the connect path now.
  • WorkspaceService.initialize chat-restart recovery with thousands of non-task workspaces is noisy (MaxListenersExceededWarning); pre-existing and unchanged.
  • Pre-existing lock nesting, unchanged by this PR: runtime cleanup callers reach cleanupReportedLeafTask under the workspace event lock and WorkspaceService.remove() then takes the task-tree lifecycle lock (event -> tree), while task_send_message takes tree -> event. This PR adds no acquisition on that path.

Pains

Eleven Codex rounds found one race after another between background recovery and early clients while the whole of TaskService.initialize ran post-listen (stop vs restart nudge, reactivation vs patch artifact, unarchive vs DevTools sweep, workflow resume vs child interrupt, ...), each fixed with another compare-and-set or lock. Moving the small state-mutating recovery phase back in front of the listener removed that whole class along with the admission probes, idle-only sends, refunds, run-keyed workflow admissions, and settlement-lock compare-and-set they had accumulated (net about 120 production and 360 test lines), and left only idempotent housekeeping behind the listener.


Generated with xum • Model: anthropic:claude-fable-5-1 • Thinking: xhigh • Cost: $__COST__

… config reloads (#4055)

- ServiceContainer.initialize() split into initializeCore() (fast) and
  runStartupRecovery() (O(workspaces) task/workspace recovery); `xum server`
  binds its port after core init and runs recovery in the background.
- Recovery snapshots config synchronously before its first await so tasks
  created by clients after listen are never treated as stale `starting`
  tasks; running-task resume skips tasks that are already streaming.
- Startup patch-generation and reported-task cleanup loops reuse one config
  snapshot instead of re-parsing config.json per task.
- TerminalAttentionStore.listPendingOwnerWorkspaceIds scans session dirs
  with bounded concurrency and short-circuits on the first pending record.
- [startup] completion logs switch to warn above SLOW_STARTUP_WARN_THRESHOLD_MS.
@chatgpt-codex-connector

This comment has been minimized.

@ibetitsmike

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

This comment has been minimized.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5f89fb0cec

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/cli/server.ts Outdated
Comment thread src/node/services/serviceContainer.ts Outdated
Comment thread src/node/services/serviceContainer.ts Outdated
- Capture the recovery config snapshot before startServer() opens the
  listener (it accepts requests while awaiting the lockfile/mDNS), and pass
  it into runStartupRecovery().
- Orphan scratch workdir sweep skips recently touched dirs and re-checks
  fresh config before deleting, so createScratch during recovery is safe.
- Isolate workspace/task recovery failures so the periodic services
  (idle compaction, heartbeat, agent status) always start.
@ibetitsmike

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

This comment has been minimized.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 314818f428

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/taskService.ts Outdated
Comment thread src/node/services/taskService.ts
Comment thread src/cli/server.ts Outdated
…anges

- Stale-starting recovery compare-and-sets: only entries still `starting`
  are rewritten, so a task_stop that persisted `interrupted` wins.
- Running-task resume re-reads the live status before nudging, so a task
  stopped after the snapshot is not restarted.
- WorkspaceService.initialize re-reads config right before scheduling chat
  startup recovery, skipping workspaces archived or removed since the
  metadata read.
@ibetitsmike

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6aa18ab264

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/taskService.ts Outdated
Comment thread src/node/services/serviceContainer.ts Outdated
Comment thread src/node/services/serviceContainer.ts
…p and unarchive

Startup recovery now runs while clients are connected, so every one-shot read it
makes before a mutation can go stale:

- The pending-guidance replay, restart nudge, and awaiting_report completion
  prompt carry an admissionStale probe. A task_stop that lands between the live
  status read and turn admission now refuses the send instead of resurrecting
  the stopped task through markInterruptedTaskRunning.
- reconcileAgentTaskExecutionIds compare-and-sets its mirror write under the
  per-handle settlement lock, so a stop that settled the handle during the
  liveness check keeps its terminal mirror and no stale live registration is
  installed.
- cleanupArchivedDevToolsLogs re-checks the live archive state under the
  task-tree lock before each deletion. DevToolsService.hasWorkspaceData gates
  that fresh config read so it only happens for actual deletions, not once per
  archived workspace.

Copy link
Copy Markdown
Contributor Author

@codex review

1 similar comment
@ibetitsmike

Copy link
Copy Markdown
Contributor Author

@codex review

Now that the graceful SIGINT handler is installed while the startup DuckDB
sync may still be running, exiting the process mid-sync tore the worker
down inside native DuckDB and aborted the process (Napi::Error ->
SIGABRT, 7/8 reproductions in UAT). AnalyticsService.dispose() now waits
(bounded by ANALYTICS_WORKER_SHUTDOWN_TIMEOUT_MS) for the worker to exit
after posting the shutdown message.
@ibetitsmike

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

This comment has been minimized.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f095888084

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/analytics/analyticsService.ts Outdated
Comment thread src/node/services/workspaceTurnManager.ts Outdated
…nstalled execution handle

- AnalyticsService.dispose() now waits for the worker's exit event with no
  local bound. A 2 s timeout let dispose resolve while an ingest still held
  DuckDB open, so the caller's process.exit() reproduced the SIGABRT the wait
  exists to prevent. The outer quit budgets in cli/server.ts and
  desktop/main.ts already race the whole dispose and own the hard-exit
  decision. A worker that already exited short-circuits the wait.
- reconcileAgentTaskExecutionIds also compare-and-sets the workspace mirror
  pointer against its snapshot. A client that installs a new handle while
  reconciliation awaits the old handle's liveness check writes the mirror
  under a different settlement lock and leaves the old record unchanged, so
  the record-status check alone let the stale handle overwrite it.

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

This comment has been minimized.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 02cfc0a53a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/cli/server.ts Outdated
Comment thread src/node/services/taskService.ts
…and removed tasks

- interruptTaskRecoveryForInactiveWorkflowOwner leaves a child alone when a
  workflow resume is admitted for the run's workspace (checked synchronously
  inside the config edit via the workflow admission registry), so recovery
  cannot mark a freshly resumed run's child interrupted.
- maybeStartGeneration re-reads config before creating a pending artifact for
  a snapshot child that has no artifact yet, so a task removed since the
  startup snapshot does not get a new artifact.
@ibetitsmike

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

This comment has been minimized.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b460a3e682

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/taskService.ts
Comment thread src/node/services/taskService.ts Outdated
…ity on live config

- Workflow admissions are tracked per run id as well as per workspace;
  startup recovery skips a stale child only when a resume of its own run is
  in flight, so unrelated runs in the same workspace no longer shield it.
- canCleanupReportedTask treats a caller snapshot as a screen only: any
  positive verdict is re-evaluated on freshly loaded config before the
  workspace is removed.
@ibetitsmike

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

This comment has been minimized.

@chatgpt-codex-connector

This comment has been minimized.

Copy link
Copy Markdown
Contributor Author

@codex review

Retrying round 21 on 1d4d77e: the previous request returned "Something went wrong".

@chatgpt-codex-connector

This comment has been minimized.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1d4d77ecd9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/taskService.ts
Comment thread src/node/services/taskService.ts
A reported best-of child that a client reawakened keeps its previous report
artifact while its continuation runs, and the grouped builder only checked that
every artifact exists. Revalidate each non-reporting sibling's execution and
stream state before assembling the grouped output, and treat such a sibling as
recoverable for the synthetic fallback so the parent waits for the new report.

Copy link
Copy Markdown
Contributor Author

@codex review

Round 22 (433e2bd): best-of finalization now revalidates sibling execution/stream state (round-21 P1 #2, fixed with a red-green test); the lock-order finding (round-21 P1 #1) is pre-existing on main and tracked in #4072 with line-level evidence in the thread.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 433e2bda4f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/taskService.ts
Comment thread src/node/services/agentSession.ts
Comment thread src/node/services/agentSession.ts
Comment thread src/node/services/taskService.ts
…st-of assembly

Removal holds turn admission on the workspace's sessions for its whole
duration (released on failure), so startup recovery resuming from a disk read
cannot start work against a checkout being deleted. sendMessage refuses at the
pre-persist gate once the shutdown latch is set, so no follow-up row can read
as a dispatched turn on the next startup. Best-of assembly re-reads sibling
state after its artifact reads so a reactivation during those awaits cannot
ship the stale report it already collected.

Copy link
Copy Markdown
Contributor Author

@codex review

Round 23 (b795316): three of the round-22 findings are fixed (removal admission hold, pre-persist shutdown gate, post-read sibling revalidation), each with a red-green test; the CAS-supersession finding is answered in its thread as the pre-existing busy-parent delivery path.

@chatgpt-codex-connector

This comment has been minimized.

Copy link
Copy Markdown
Contributor Author

@codex review

Retrying round 23 on b795316: the previous request returned "Something went wrong".

@chatgpt-codex-connector

This comment has been minimized.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b795316c1e

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/agentSession.ts
Comment thread src/node/services/taskService.ts
@ibetitsmike
ibetitsmike added this pull request to the merge queue Sep 3, 2026
Merged via the queue into main with commit e2c3ab1 Sep 3, 2026
99 of 110 checks passed
@ibetitsmike
ibetitsmike deleted the mike/server-bind-before-recovery branch September 3, 2026 22:30
yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Sep 6, 2026
…r-step timeouts (Wave 4 PR 3) (coder#4076)

## Summary

Wave 4 PR 3 (plan §2 D6 / §3 "PR 3"):
`ServiceContainer.initializeCore()` is now a Promise facade over a
**startup effect run on the app runtime** — one `Effect.gen` over an
ordered table of the five hard startup steps, each `Effect.tryPromise`
(identity catch) bounded by
`Effect.timeoutOrElse(STARTUP_STEP_TIMEOUT_MS)` on the runtime's
`Clock`. A hung step no longer pins the splash screen / listener bind
forever: after 60 s it fails startup with a `StartupStepTimeoutError`
(`"<step> exceeded <ms> ms"`, `step`/`timeoutMs` fields) through the
roots' **unchanged** failure paths. Step errors keep their identity;
later steps do not run; the five step names/order and the `[startup]
ServiceContainer.initialize completed { stepDurationsMs }` payload are
unchanged. Every process root now runs the bounded `dispose()` before
exiting on a rejected startup (D6 abandon-and-quit safety):
`cli/server.ts` gains it, and the ACP root's existing catch **did not**
dispose after a rejected `initialize()` (`if (initialized)` guard — the
plan's claim that it did was wrong), so it does now.

## Background

- Plan: `~/.xum/plans/mux/effect-wave4-lifecycle-core-plan.md`
(collapsed below); PR 1 = coder#4070; the plan predates coder#4058, which split
`initialize()` into `initializeCore()` (hard steps, gate the server
listener) + `runStartupHousekeeping()` (best-effort, abort-signal
cancellable). This PR targets `initializeCore()` only.
- `runStartupHousekeeping()` is deliberately untouched: its steps are
already cancellable through the dispose abort signal and non-fatal by
policy, so per-step timeouts there would be a policy change the plan
forbids (recorded in the `appRuntime.ts` startup contract).

## Implementation

- `serviceContainer.ts`: `startupCoreSteps: readonly StartupStep[]`
(names asserted unique in the constructor — they key `stepDurationsMs`);
`initializeCore()` = `assert(not disposed)` +
`runtime.managed.runPromise(startupCoreEffect())`; `timedStartupStep` =
`Effect.suspend` → `tryPromise({ try: async () => run(), catch: identity
})` → `timeoutOrElse` → `Effect.ensuring` (duration recorded when the
*wait* ends — settled, failed, or abandoned — never by the abandoned
step later). `timeoutOrElse` rather than `timeout` + `catchTag` because
the error channel is `unknown`. No `forkDetach`: the zero-arity thunk
gets no AbortSignal, so the step promise keeps running; its late
settlement is a no-op on the exited fiber and `tryPromise` keeps a
rejection handler attached (pinned by test).
- `StartupStepTimeoutError extends Error` with `name` set in the
constructor, so the desktop `Startup Failed` dialog and `Failed to
initialize server:` print class + step without formatting changes.
- `src/constants/terminationTimeouts.ts`: `STARTUP_STEP_TIMEOUT_MS = 60
s` (evidence below) and `SERVICE_TEARDOWN_BUDGET_MS = 5 s` (the outer
teardown budget `cli/server.ts` already used as a literal, now shared
with the ACP root).
- `cli/server.ts`: `main().catch` runs `dispose()` bounded by the same 5
s budget as the SIGTERM cleanup when the container was constructed, then
exits 1. `acp/serverConnection.ts`: the catch disposes unconditionally
(bounded). `desktop/main.ts` unchanged — the catch at `:1249–1265` quits
and the `before-quit` listener (`:1271–1305`) races `dispose()` against
5 s.
- `appRuntime.ts`: new "Startup" contract section; "Deliberately not
done" no longer lists `initialize()`.
- Decisions pinned by tests: re-entry is **not** guarded (parity with
the promise chain); `initializeCore()` after `dispose()` fails fast with
an assertion instead of the disposed runtime's bare `"ManagedRuntime
disposed"` string defect.

## Validation

- `serviceContainer.test.ts` (+5): TestClock timeout — rejects with
`StartupStepTimeoutError` naming the step **exactly** at
`TestClock.adjust(STARTUP_STEP_TIMEOUT_MS)` (still pending at −1 ms),
later steps not called, the abandoned step's late rejection is not
unhandled and changes nothing on the container; error identity (`toBe`)
for a rejecting and for a synchronously throwing step with later steps
skipped; happy path records the five keys in order and a second call
re-runs the steps; after `dispose()` fails fast without running a step.
**Red check** against the old promise chain: the timeout test (hangs → 5
s test timeout) and the after-dispose test fail; the
identity/sync/re-run tests are parity pins (green on both, by design).
- Gates: `serviceContainer.test.ts` (23), `di/*.test.ts` +
`coreServicesRoot.test.ts` + `acp/*.test.ts` (55), `TEST_INTEGRATION=1
bun x jest
tests/ipc/{doubleRegister,windowTitle,savedQueries,acp.sessionMethods,acp.disconnectCleanup}`
(32), `bun test src/cli/` (185), `make static-check`.
- Dogfooding (first comment, with screenshots): 3 + 3 `xum server` cold
starts on `main` vs branch — same ten `stepDurationsMs` keys in the same
order, slowest core step `taskService.recoverInterruptedTasks` 49–61 ms;
a throwaway build with the constant forced to 1 ms on `xum server` and
`xum acp` shows `Failed to initialize server: StartupStepTimeoutError:
extensionMetadata.initialize exceeded 1 ms`, the full bounded
`[shutdown]` dispose sequence (28 ms / 36 ms), exit code 1.
- `cli/server.ts` has no unit seam for `main().catch` (`server.test.ts`
exercises `createOrpcServer`); the throwaway transcript is the evidence
for that path.

## Risks

- False timeout on a pathologically slow host turns a slow-but-fine
start into a crash — bounded at 60 s, ≥ 1000× the measured maximum and
above the policy service's own 10 s fetch timeout; the only potentially
unbounded step, `taskService.recoverInterruptedTasks`, scales with
active agent tasks, not deployment size.
- An abandoned step (e.g. mid-`editConfig`) keeps running until the
process exits; the bounded `dispose()` applies the same latches as a
quit (`beginShutdown`s) and config writes are lock/journal protected —
identical exposure to a SIGTERM mid-startup today.
- Disposing the runtime concurrently with an in-flight
`initializeCore()` does not interrupt the startup fiber (root fibers are
not scope children) — same as the promise chain it replaces; documented,
not changed.

OFF-RAMP did not fire: error identity, `tests/ipc` and the ACP path are
preserved.

---

<details>
<summary>📋 Implementation Plan (Wave 4 —
effect-wave4-lifecycle-core-plan.md)</summary>

# Effect migration — Wave 4: finish the concurrency/lifecycle core

Bounded wave: **4 PRs (PR 4 optional), explicit STOP criterion, explicit
OFF-RAMPs.** Plan only; nothing here is implemented.

> **Review status:** Independently reviewed (adversarial Reviewer
sub-agent, advisor unavailable): APPROVE WITH REQUIRED EDITS — both
edits applied; verified claims: Effect.promise 0-arity thunk allocates
no AbortController (internal/effect.js:741–776); closed-scope
forkIn+startImmediately runs onInterrupt (2237–2274, 391–409); forkIn
observer removes the scope finalizer on exit (2270–2271);
Effect.timeoutOrElse exists (Effect.d.ts:7833); 0 line drift at
b87f627; 11 settleWorkspaceTurn callers confirmed; 'aborted' ∈
NON_RETRYABLE_STREAM_ERRORS. Line references are to `main` @
`b87f62729`.

## 0. Thesis check (coordinator's judgment vs. evidence)

**Thesis:** Effect's payoff in this app is structured concurrency +
interruption-safe lifecycles in the orchestration core (still Promise +
AbortController).

**Verdict: holds for the stream engine; only half-holds for turn
handles.**

- Stream engine — **holds.** `ServiceContainer.dispose()` never stops or
awaits in-flight streams (`serviceContainer.ts:478–537` has no
`streamManager` step); an in-flight stream dies with the process and is
recovered on next load from `partial.json` (≤ 500 ms stale,
`PARTIAL_WRITE_THROTTLE_MS`, `streamManager.ts:776`). `AppFiberScope`
exists precisely for this and has no occupant. A supervised per-stream
fiber is the right tool.
- Turn handles — **half-holds.** The 7× "superseded by an uncorrelated
workspace stream-end" false-settle is a *correlation-predicate* bug
(`interruptWorkspaceTurnFromUncorrelatedStreamEnd`,
`workspaceTurnManager.ts:4220–4307`: any uncorrelated stream-end after
the prompt index settles the handle `interrupted`), not a
Promise-vs-fiber structure bug. Turn handles are **persisted records**
(`taskHandleStore.upsertWorkspaceTurn`) spanning **multiple streams**
(tool-call continuations are deferred via `hasSameTurnContinuation`,
`:4491`) and surviving restarts; a fiber/Deferred can only model the
in-process waiter and would not fix correlation. Open PR **coder#3949** fixes
the predicate in Promise idiom and is Codex-green. Wave 4's turn-handle
PR therefore becomes **"codify the settlement invariant + prove the
class is gone"**, not "fiberize handles" (D5 below).

Corrected baseline numbers (measured this workspace): 46/470 `src/node`
non-test files import `effect` (coordinator said 35); 113 direct
`Effect.run*` sites outside `di/` in 16 files; 226 `Effect.gen`; 9
`TaggedError` classes; effect `4.0.0-rc.112`, `@orpc/*` `1.14.11`;
**effect v4 is not GA** (rc line still current).

## 1. Verified current state (evidence the design rests on)

<details>
<summary>Stream engine (streamManager.ts)</summary>

- `startStream` (`:4723–4901`): per-workspace mutex → `new
AbortController()` + `linkAbortSignal` (`:4771–4772`) → `resourceScope =
Scope.makeUnsafe()` (`:4777`) → temp-dir `Effect.acquireRelease`
(`:4802–4824`) → `createStreamAtomically` → `streamText` (`:2244`,
`abortSignal: abortController.signal` `:2250`) → registered in
`workspaceStreams` (`:2463`) → **`streamInfo.processingPromise =
this.processStreamWithCleanup(...)` fire-and-forget (`:4876–4882`)** →
returns `Ok({ messageId, completion })`.
- `processStreamWithCleanup` (`:3331–4089`, plain async): `while(true)`
retry loop; `for await (part of fullStream)` (`:3358–3837`) with abort
check at loop head (`:3361`); post-loop `if (!signal.aborted)` gate
(`:3849`) → completion path (`deletePartial` `:3981`, `updateHistory`
`:3989`, `recordSessionUsage` `:4001`, `state = COMPLETED` `:4017`, emit
`stream-end` `:4023`, `terminalCompletion` `:4024`); error path →
`handleStreamFailure` (`:4094–4112`) → `persistStreamError` writes error
partial; `finally` (`:4052–4088`): release MCP lease,
`Effect.runFork(Scope.close(resourceScope))` (`:4064–4066`), unlink
abort, `workspaceStreams.delete`, `eventSpine.emit("stream.end")`,
`completionController.settle`.
- Cancellation: `stopStream` (`:5043–5111`) → `cancelStreamSafely`
(`:1766–1800`): `if (state === COMPLETED) { await processingPromise;
return }` → `state = STOPPING` → `flushPartialWrite` →
`abortController.abort()` → `cleanupAbortedStream` (`:1828–1951`):
`await processingPromise` → usage → `writePartial` (`:1876–1910`) →
`emitStreamAbort` → `settle({status:"aborted"})`. **No completed-guard
after the await** (verified `:1838–1951`): a cancel landing between
`:3849` and `:4017` re-writes `partial.json` after `deletePartial` and
emits `stream-abort` after `stream-end` (pre-existing window; dispose()
will widen its exposure). `cancelStreamSafely` is also not idempotent
for concurrent callers (only `COMPLETED` is checked).
- AIService on `stream-abort` (`aiService.ts:355–377`): `abandonPartial
? deletePartial : commitPartial → deletePartial` (fire-and-forget
listener).
- Crash recovery: `HistoryService.commitPartial`
(`historyService.ts:1963–2061`) — strips error metadata,
`hasCommitWorthyParts`, stale-epoch check, update-or-append by
`historySequence`, delete partial; invoked from `agentSession.init`
(`:5002`), `aiService.streamMessage` (`:886`), stream-abort (`:364`),
`duplicateWorkspace`.
- `StreamAbortReason = "user" | "startup" | "system"`
(`src/common/orpc/schemas/stream.ts:295`).
- Pinned seams: chaos test `Reflect.set(streamManager, "tokenTracker" |
"createStreamResult")` (`streamManager.chaos.test.ts:130–134, 238–242`);
`streamManager.test.ts` `Reflect.set` on `processStreamWithCleanup`
(`:2787`), `createStreamAtomically` (`:2783`), `createTempDirForStream`,
`cleanupStreamTempDir`, `Reflect.get` on `workspaceStreams`,
`schedulePartialWrite`, …; `modelOnlyNotifications.test.ts` calls
`processStreamWithCleanup` directly (`:93, :187`); `aiService.test.ts`
spies `startStream`, `generateStreamToken`, `createTempDirForStream`,
`isResponseIdLost`. Constructor: `(historyService, sessionUsageService?,
getProvidersConfig?, eventSink = noop, runner = defaultEffectRunner)`
(`:801–813`); `effectRunner` used at `:1152, :1154, :1169` only.
- Every stream event carries `workspaceId` + `messageId`;
`stream-end`/`stream-abort`/`error` carry `metadata.muxMetadata` when
the prompt had it.
</details>

<details>
<summary>DI / shutdown / startup</summary>

- `AppFiberScopeLive` is in `CoreLive`'s `runtimeSeams`
(`di/layers/core.ts:644`), so both roots have it; `StreamManagerLive`
(`core.ts:226–238`, stage S2b) already yields `EffectRunnerTag`. CLI
cleanup lists include `appFiberScope.close` (`cli/run.ts:1579`,
`cli/workflow.ts:286`).
- Bounds: `APP_FIBER_SCOPE_CLOSE_TIMEOUT_MS = 2000`,
`APP_RUNTIME_DISPOSE_TIMEOUT_MS = 2000`; outer budgets are **5000 ms**
on both desktop (`desktop/main.ts:1297` `Promise.race` vs
`setTimeout(5000)`) and `xum server` (`cli/server.ts:236–243`
force-exit). **The scope bound cannot grow** without changing outer
budgets.
- rc.112 semantics verified in
`node_modules/effect/dist/internal/effect.js:2264`: `forkIn` registers a
scope finalizer and **removes it when the fiber completes** (no leak),
and **interrupts immediately if the scope is already closed** (streams
starting mid-shutdown fail closed). `Effect.promise(evaluate: (signal)
=> PromiseLike)`, `Effect.onInterrupt`, `Effect.forkIn(_, scope, {
startImmediately? })`, `Stream.toAsyncIterableWith(context)`,
`Stream.provideContext` all exist.
- `ServiceContainer.initialize()` (`serviceContainer.ts:297–362`): six
awaited `initialize()`s wrapped in `recordStep` (durations only, no
catch, no timeout) + three sync `start()`s + two fire-and-forget sweeps.
Failure handling: desktop `Startup Failed` dialog + `app.quit()`
(`desktop/main.ts:1249–1265`); `cli/server.ts:136` uncontained; ACP
`serverConnection.ts:205–216` dispose + rethrow; `tests/ipc/setup.ts:85`
no catch. No outer timeout anywhere.
- `streamBridge.subscriptionIterable` (`orpc/streamBridge.ts:176`) →
`Stream.toAsyncIterable(...)` on the global runtime, 19 call sites in
`routerSubscriptions.ts`; heartbeat via `Effect.sleep` in `forkScoped`
(`:145–152`). `streamBridge.test.ts` has 11 real-time waits, but only
**3 are clock-bound** (`:207` 1 ms initial delay, `:241` heartbeat 10
ms, `:255` 10 ms laziness); 8 are `waitFor(listenerCount…)` readiness
polls that TestClock cannot replace.
</details>

<details>
<summary>Turn handles + open PRs</summary>

- Handle record `{ handleId "wst_…", ownerWorkspaceId, workspaceId,
turnId, messageId, status, attentionPolicy, disposableWorkspace }`;
prompt carries `muxMetadata: { type:"workspace-turn-task", taskHandleId,
ownerWorkspaceId, turnId }` (`workspaceTurnManager.ts:1420`).
`TaskService` forwards `aiService` `stream-end`/`stream-abort`/`error`
to `finalizeWorkspaceTurnFromStreamEnd` (`:4442–4544`): correlated
branch matches `record.workspaceId && record.turnId` (`:4472`);
**uncorrelated branch** (`metadata == null`, not `agentId ===
"compact"`) → `interruptWorkspaceTurnFromUncorrelatedStreamEnd` →
settles `interrupted` whenever `streamEndIndex >= promptIndex`
(`:4293–4305`). Producers of such uncorrelated ends: bash-monitor wake
continuations, child terminal-attention deliveries, heartbeat, peer
messages, parent auto-resume.
- Cascade: disposable child → `cleanupDisposableWorkspaceTurn` →
`workspaceService.remove(…, true)` kills its background processes;
persistent child → parent sees `interrupted` → `task_stop` →
`backgroundProcessManager.stopMonitor(…, "canceled")`. This is the
observed "monitors died afterwards".
- Settlement chokepoint: `settleWorkspaceTurn(params)` (`:2085`), **11
callers**, guarded by `workspaceTurnSettlementLocks.withLock(handleId)`;
waiters in `pendingWorkspaceTurnWaitersByHandleId` with `setTimeout`
timeouts (`:2425–2496`).
- **coder#3949** "preserve turns across synthetic wake ends" (coadler):
rewrites the uncorrelated branch — walks history from the turn anchor to
the stream-end and settles *only if a manual child input intervened*
(`isManualChildWorkspaceInput`); otherwise ignores the end. Touches
`:281–295, :4217–4355` + tests (+286/−31). Codex: "Didn't find any major
issues" + clean security on `f9baa2fc9`. `mergeable: MERGEABLE`, but
`Test / Unit` and `Codex Comments` red, 19 commits behind main.
- **coder#3915** "correlate workspace-turn liveness" (coadler): creation
reservations +
`getWorkspaceTurnLiveness`/`getWorkspaceTurnRuntimeActivity`
(identity-matches the active stream's `muxMetadata` against the record)
for staleness/capacity. Touches `:442–486, :1298, :2573, :3761,
:3800–4064` (+494/−60). `BLOCKED`, latest Codex review has open
findings, `Test / Unit` red, 19 behind.
- Together they are the identity-correlated model the coordinator wants:
coder#3915 = identity-correlated *liveness*, coder#3949 = identity-gated
*settlement*.
</details>

## 2. Design decisions

**D1 — Fibers WRAP the AbortController; they do not replace it.**
The AI SDK is cancelled only via `AbortSignal`; the `for await` loop,
soft-interrupt at step boundaries, retry/fallback re-creation of
`streamResult`, and ~30 abort touchpoints (coder#4032) all key off the
signal. Converting the 750-line loop to `Stream.fromAsyncIterable` +
fiber interruption would touch hundreds of `WorkspaceStreamInfo`
transitions and break the
`processStreamWithCleanup`/`createStreamResult` spy seams. Instead:
**the fiber is the ownership/supervision unit; the signal stays the
cancellation transport.** The dual-cancellation glue coder#4032 feared is
confined to **one** point — the supervisor's `onInterrupt` — which
routes through the existing user-stop path (`cancelStreamSafely`), so
shutdown ≡ "user pressed stop" semantically (partial flushed with usage,
`stream-abort` emitted, `completion` settles `aborted`, AIService
commits the partial).

**D2 — Supervisor topology: one supervisor fiber per stream in
`AppFiberScope`, wrapping the already-started `processingPromise`.**
`streamInfo.processingPromise = this.processStreamWithCleanup(...)`
stays byte-identical (sync-start preserved;
`Reflect.set(processStreamWithCleanup)` seam preserved;
`cleanupAbortedStream`'s `await processingPromise` unchanged).
Immediately after it:

```ts
// startStream, after processingPromise is assigned (unsupervised path unchanged when no scope)
this.superviseEngine(typedWorkspaceId, streamInfo);

private superviseEngine(workspaceId: WorkspaceId, streamInfo: WorkspaceStreamInfo): void {
  if (this.engineScope === undefined) return;               // direct construction / CLI tests: today's behavior
  assert(streamInfo.engineFiber === undefined, "engine already supervised");
  // Zero-arity thunk on purpose: rc.112 allocates an internal AbortController only
  // when `evaluate.length !== 0`; the stream's own controller stays the sole signal.
  const supervisor = Effect.promise(() => streamInfo.processingPromise).pipe(
    Effect.onInterrupt(() =>
      Effect.uninterruptible(   // explicit, per house doctrine (finalizers are already uninterruptible)
        Effect.promise(async () => this.cancelStreamSafely(workspaceId, streamInfo, "system"))
      )
    ),
    Effect.catchDefect((d) => Effect.sync(() => log.warn("[stream] engine supervisor defect", { workspaceId, error: d })))
  );
  streamInfo.engineFiber = this.effectRunner.runSync(
    Effect.forkIn(supervisor, this.engineScope, { startImmediately: true })
  );
}
```
- `Effect.promise` is interruptible while suspended
(`internal/effect.js:741–801`, Async op); `onInterrupt` =
`onErrorFilter(causeFilterInterruptors, …)` (`:1762`); `forkIn`
registers `fiberInterrupt(fiber)` as the scope finalizer (`:2264–2275`),
`fiberInterrupt` awaits the fiber (`:635–642`), and parallel
`scopeClose` awaits all finalizers via `fiberAwaitAll` (`:1590–1601`) →
`closeScopeBounded` at dispose step 2 gives "interrupt **and** await"
while `historyService`/`sessionUsage`/`eventSink → AIService → bridge
servers` are still alive (bridges stop in step 3, so clients receive
`stream-abort`).
- Normal completion: fiber exits → `forkIn`'s observer removes the scope
finalizer (verified) → no per-stream residue.
- Stream started after step 2: `forkIn` on a closed scope calls
`fiber.interruptUnsafe` synchronously and returns the fiber
(`:2272–2274`, `runSync` does not defect). With `startImmediately:
true`, `forkUnsafe` runs `child.evaluate` synchronously (`:2233–2247`)
up to the `Effect.promise` Async op (`:772–801`), so the fiber is
suspended (`_running=false`) when the interrupt lands and
`interruptUnsafe` (`:391–409`) unwinds the stack through the
`onInterrupt` handler → the stream is aborted as `system` (fail-closed
during shutdown). Verified in rc.112 internals (`effect.js:2233–2247,
391–409`); **pin with a test** ("stream started after scope close is
aborted") so an RC bump cannot silently change it.
- `"system"` is semantically exact: `"user"`/`"startup"` suppress
next-startup recovery (`retryEligibility.ts:114–118, 284–287`),
`"system"` marks an involuntary backend interruption (as
`taskService.ts:8100, 8223` use it). **No in-session retry loop is
possible:** the `stream-abort` handler (`agentSession.ts:6010`) routes
`{ type: "aborted" }` to `retryManager.handleStreamFailure`, and
`"aborted"` is in `NON_RETRYABLE_STREAM_ERRORS`
(`retryEligibility.ts:49–59, 106`) → `retryManager.ts:99–104` abandons
immediately, never schedules a fiber. Dogfooding still checks the
*restart* UX (the recovered partial is shown as interrupted; note
whether any next-startup recovery re-sends — same class as today's
`system` aborts from `taskService`).
- `engineScope` arrives as an **optional 6th constructor parameter**
(`engineScope?: Scope.Closeable`), wired from `AppFiberScopeTag` in
`StreamManagerLive` (`core.ts:226`). Default `undefined` keeps every
direct-construction test and `aiService.ts:174` path identical (I4).
`AppFiberScopeLive` already sits beneath S2b in `runtimeSeams`, so no
staging change (I6).
- Abort reason: reuse **`"system"`** — no wire/schema change; UI copy
for `system` already exists.
- Pending-start window (`pendingStreamStarts`, before registration) is
**not** supervised: nothing is persisted for it yet, and `stopStream`
already aborts pending controllers. Documented, not fixed.

**D3 — Fix the two adjacent cancel races in the same PR (closely-related
bugs, not deferrals).**
(a) `cleanupAbortedStream`: after `await processingPromise`, if
`streamInfo.terminalCompletion !== undefined` (completed/failed while
the cancel was in flight) → return without abort bookkeeping (prevents
`partial.json` resurrection after `deletePartial` and a `stream-abort`
after `stream-end`). (b) `cancelStreamSafely` (`:1766`): latch a
per-stream `cancelPromise` so concurrent cancellers (user stop racing
dispose) join one cleanup → exactly one `stream-abort`, one `settle`.
**Zero-suspension requirement:** the latch must be checked and assigned
**synchronously at function entry, before any `await`** (the current
first await is `flushPartialWrite` at `:1789`) — otherwise racing
callers can both enter `cleanupAbortedStream`. Shape:

```ts
if (streamInfo.cancelPromise) return streamInfo.cancelPromise;
streamInfo.cancelPromise = (async () => { /* existing body, unchanged */ })();
return streamInfo.cancelPromise;
```
Both are ≤ 10 LoC and get behavioral tests.

**D4 — Shutdown bound stays 2 s; the finalizer must be fast or
abandoned.**
Outer budgets are 5 s; 2 s + 2 s already consume 4 s. A flowing stream
aborts within one chunk; a wedged provider (no chunks, ignores abort)
hits the existing `boundedTeardown` timeout: warning, continue, process
exit — identical to today's outcome. Dogfooding measures the actual
`[shutdown] AppFiberScope closed { ms }` with a live stream.

**D5 — Turn handles: codify the settlement invariant; do not fiberize.**
Invariant: *a workspace-turn handle settles terminally only by (i) a
stream terminal event whose `muxMetadata` correlates `{taskHandleId,
ownerWorkspaceId, turnId}` to the record; (ii) an explicit interrupt
(`task_stop`/`interruptWorkspaceTurn`); (iii) manual supersession — a
manual child input after the turn anchor; (iv) stale-liveness
reconciliation.* An uncorrelated stream-end is **never** terminal by
itself. coder#3949 makes (iii) the only uncorrelated outcome; coder#3915
implements (iv) by identity. Wave 4 adds a `cause` discriminant to
`settleWorkspaceTurn` (the single chokepoint) with a runtime assertion,
plus the regression harness. Rationale for not converting waiters to
`Deferred`/fibers: no behavioral gain, 4.9k-line file, and the
coordinator's "settle only on the owning stream's termination" is
over-specified — a turn owns *several* streams.

**D6 — Startup: `initialize()` stays a Promise facade over a runtime-run
startup effect; timeout ⇒ same failure path as a thrown step.**
Each step is `Effect.tryPromise({ try: async () => step(), catch:
identity }).pipe(Effect.timeoutOrElse({ duration:
STARTUP_STEP_TIMEOUT_MS, orElse: () => Effect.fail(new
StartupStepTimeoutError(name, ms)) }))` (`timeoutOrElse` exists in
rc.112, `Effect.d.ts:7833`; chosen over `timeout` + `catchTag` because
the step's error channel is `unknown`, which `catchTag` cannot narrow).
No `forkDetach` needed: a Promise step keeps running on its own when the
waiting fiber times out (not inside an uninterruptible region, so the
timeout interrupts the wait directly). `StartupStepTimeoutError extends
Error` with `name = "StartupStepTimeoutError"` set in the constructor
and message `"<step> exceeded <ms> ms"` (so the desktop dialog's error
formatting shows both the class and the step name) → desktop shows it in
the existing `Startup Failed` dialog; CLI/ACP/tests paths unchanged.
Step **errors keep their identity** (v4 `runPromise` rejects with the
raw failure). Downgrading any step to best-effort is a **policy change,
out of scope** (audit of the six implementations:
extensionMetadata/telemetry/experiments are local fs, <50 ms; policy has
its own 10 s fetch timeout; workspaceService bounds its sync internally;
only `taskService.initialize` — config scan + `editConfig` + recovery
`sendMessage`s — is potentially unbounded). The three `start()`s stay
sync (`Effect.sync`), the two fire-and-forget sweeps stay outside the
effect. `stepDurationsMs` is preserved.
**Abandon-and-quit safety:** an abandoned `taskService.initialize` may
be mid-`editConfig` when the root exits. Parity requirement for PR 3:
after a rejected `initialize()`, every root runs the bounded `dispose()`
before exiting. Verified: desktop already does — `services` is assigned
before the await (`main.ts:653–656`), the catch calls `app.quit()`, and
the `before-quit` listener (`:1271–1305`, guard `if (isDisposing ||
!services) return`) races `services.dispose()` against 5 s; ACP does
(`serverConnection.ts:205–216`); **`cli/server.ts` does not**
(`:133–136` awaited at top level, `main().catch` at `:282` only logs) →
PR 3 adds a bounded `dispose()` there (≤ 10 LoC, same 5 s budget).

**D7 — streamBridge: thread the runtime context, not a runner.**
`subscriptionIterable` gains `context?: Context.Context<never>` →
`Stream.toAsyncIterableWith(context)`; `routerSubscriptions` passes the
handler's `"effect/context"`. Production behavior identical; heartbeat
sleeps on the runtime `Clock`; tests can run the 3 clock-bound waits on
`TestClock`. Honest scope: the 8 readiness polls stay.

## 3. PRs (ordered by value ÷ risk; each independently mergeable)

### PR 1 — StreamManager engine core becomes the first `AppFiberScope`
occupant
**Value:** high (the only remaining shutdown data-integrity gap; the
reason `AppFiberScope` exists). **Risk:** medium → low with D1/D2. **Net
product LoC ≈ +55** (`superviseEngine` ~25, ctor param/field ~5,
`engineFiber` field ~2, D3 guards ~12, `core.ts` wiring ~2, doc updates
in `appRuntime.ts`/`appFiberScope.ts` "occupant" text ~10).

Files: `src/node/services/streamManager.ts`,
`src/node/services/di/layers/core.ts`,
`src/node/services/di/appRuntime.ts` + `appFiberScope.ts` (docs), tests
below.

Pre-work (before writing product code; each yields a note in the PR
body):
1. Confirm `processStreamWithCleanup` never rejects (try/catch/finally
shape `:3331–4089`); else the supervisor must fold rejections (it
already `catchDefect`s).
2. Confirm `Effect.promise` interruption + `onInterrupt` await ordering
under `Scope.close` in a 20-line probe test (pattern of
`appFiberScope.test.ts:27–47`).
3. Enumerate abort observers that run *after* the finalizer resolves
(AIService `stream-abort` listener → `commitPartial`; agentSession
completion continuations) and confirm the durable order (`writePartial`
→ commit → `deletePartial`) makes a mid-flight `process.exit`
recoverable on next load (it is: partial survives until commit
completes).
4. Measure: `[shutdown] AppFiberScope closed { ms }` with a live stream
in the sandbox (D4).
5. Pin (same probe test): a fiber forked with `startImmediately: true`
into an already-closed scope still runs its `onInterrupt` finalizer
(reviewer-verified in rc.112 internals; the test guards RC bumps).

Acceptance (behavioral tests only):
- `streamManager.test.ts` (new cases; existing cases untouched): with
`engineScope = Scope.makeUnsafe("parallel")` and a fake
`createStreamResult` whose `fullStream` yields one `text-delta` then
blocks until its `AbortSignal` fires — `closeScopeBounded(engineScope)`
resolves; `writePartial` was called with the streamed text; exactly one
`stream-abort` (`abortReason: "system"`) and zero `stream-end`;
`completion` settles `{status:"aborted"}`; `workspaceStreams` is empty.
- Wedged provider (fullStream never yields, ignores abort):
`closeScopeBounded` resolves within the bound, never rejects, warns once
(assert the returned promise resolves and no throw; do **not** assert
log text).
- No-scope construction: identical event sequence to today (guards
existing suites; no new assertions needed beyond the unchanged suites
passing).
- D3(a): cancel issued after the loop exits but before `COMPLETED` →
history has exactly one final message, `partial.json` absent, event
order `stream-end` only.
- D3(b): `stopStream` + `closeScopeBounded` racing on one stream →
exactly one `stream-abort`, one settle.
- Fiber residue: after 50 completed streams,
`closeScopeBounded(engineScope)` emits zero `stream-abort` and completes
in the same tick class as an empty scope (assert no aborts and
`workspaceStreams.size === 0`).
- `streamManager.chaos.test.ts` — existing cases byte-identical; **one
new fuzz variant** constructs with an engine scope and closes it at a
random iteration: every stream settles **exactly once** (count terminal
events per `messageId` ≤ 1, all `completion` promises settle).
- `serviceContainer.test.ts`: "dispose() aborts and awaits an in-flight
stream before `desktopBridgeServer.stop()`" (extend the ordering harness
at `:295–323`); `coreServicesRoot.test.ts`: `xum run` cleanup list does
the same via `appFiberScope.close`.

Gate suites: `streamManager.test.ts`, `streamManager.chaos.test.ts`,
`streamManager.modelOnlyNotifications.test.ts`, `aiService.test.ts`,
`agentSession.disposeRace.test.ts`,
`agentSession.sinceReplayContract.test.ts`, `serviceContainer.test.ts`,
`coreServicesRoot.test.ts`, `di/*.test.ts`, `taskService.test.ts`,
`workspaceService.test.ts`, `turnRequestBuilder.test.ts`; `make
static-check`.

House pre-review audits: interruption posture (supervisor's only
suspension is the promise; finalizer uninterruptible end-to-end incl.
`cancelStreamSafely` → `cleanupAbortedStream`); no defect escapes
(`catchDefect` on the supervisor; `Effect.promise` thunks `async`);
spy-seam check (`processStreamWithCleanup`, `createStreamResult`,
`createStreamAtomically`, `startStream` signatures unchanged;
constructor arity unchanged, trailing optional); sync-start
(`processingPromise` assigned before fork; `runSync(forkIn)` completes
synchronously); no constructor side-effects added; zero-suspension check
on D3(b) latch (`cancelPromise` checked-and-assigned synchronously at
`cancelStreamSafely` entry, before the first `await` at `:1789`; review
the diff for any inserted `await`/lookup ahead of the assignment).

Rollback: revert the `core.ts` wiring line → `engineScope` undefined →
today's behavior; D3 guards can stay (independent bug fixes).

### PR 2 — Turn-settlement invariant + false-settle regression harness
(gated on coder#3949)
**Value:** high (7× production race). **Risk:** low. **Net product LoC ≈
+40** (`WorkspaceTurnSettlementCause` union + `cause` on
`settleWorkspaceTurn` params + assert ~10; 11 call sites × 1–3 lines).

Relationship to open PRs — explicit:
- **coder#3949 is the fix and a hard prerequisite.** PR 2 rebases on it,
changes none of its logic
(`interruptWorkspaceTurnFromUncorrelatedStreamEnd`,
`isWorkspaceTurnAnchorForRecord`, `isManualChildWorkspaceInput`), and
adds the invariant + proof on top. If coder#3949 has not merged when PRs 1/3
are done: **do not fork a competing fix**; report to the coordinator,
offer the regression test file to coder#3949's author as a review artifact,
and hold PR 2 (it is not on any other PR's critical path).
- **coder#3915 is a soft prerequisite.** Its diff (`:3800–4064` incl.
`settleStaleWorkspaceTurn`, a `settleWorkspaceTurn` caller) overlaps PR
2's one-line-per-caller change. Prefer landing after it; if PR 2 must go
first, the conflict is a one-line `cause:` addition per caller. PR 2
never edits liveness/reservation code.
- Wave 4 does not otherwise touch `workspaceTurnManager.ts`.

Design: `type WorkspaceTurnSettlementCause` enumerated from the 11
callers (audited at `main` @ `b87f62729`): `:1373` creation validation
failure; `:1476` pre-stream interrupt during launch; `:1500`/`:1518`
pre-stream send failure; `:3834`/`:3863` stale-liveness recovery /
restart timeout (`settleStaleWorkspaceTurn` — coder#3915's region); `:4301`
**uncorrelated-stream-end manual supersession** (the only uncorrelated
settle in the codebase; the path coder#3949 rewrites); `:4529` correlated
terminal; `:4571` stream-abort; `:4675` deferred stream error; `:4736`
terminal stream error. `settleWorkspaceTurn` asserts `params.cause` is a
member and, for `manual-supersession`, that the superseding input's
`messageId` is supplied — turning D5 into an exhaustive `Record<Cause,
…>` check rather than prose, so a future "settle on uncorrelated end"
cannot be added without naming (and justifying) a cause.

Acceptance:
- New `workspaceTurnManager.uncorrelatedStreamEnd.test.ts` (real
`WorkspaceTurnManager` + `TaskHandleStore` + fake `aiService` emitter,
following the existing suite's harness): (1) create turn → correlated
`stream-start` → **synthetic wake stream** on the same child ends
uncorrelated after the anchor → handle stays `running`, waiter
unresolved, no disposable cleanup, no terminal attention → correlated
`stream-end` → `completed`. (2) same with `finishReason:"tool-calls"`
continuation in between. (3) manual child input between anchor and end →
`interrupted` with `cause: manual-supersession`. (4) explicit
`interruptWorkspaceTurn` → `interrupted`, `cause: explicit-interrupt`.
Case (1) is the **scripted reproduction**: it must **fail on the
pre-coder#3949 merge-base** (run the file from a sibling worktree at `git
merge-base origin/main <coder#3949 head>`; record the failing assertion in
the PR body) and pass after.
- Existing 113 `workspaceTurnManager.test.ts` cases and
`taskService.test.ts` turn cases unchanged.

Gate suites: `workspaceTurnManager.test.ts`, `taskService.test.ts`,
`taskHandleStore.test.ts`, `tools/task*.test.ts`; `make static-check`.

Audits: spy-seam (`getWorkspaceTurn`, `listAllWorkspaceTurns`,
`enqueueTerminalAttention`, `deliverPersistentChildWorkspaceTurnResult`
untouched); settlement lock held across the assert; no new suspension
inside `withLock`.

Rollback: revert; the test file stays valid against coder#3949 alone (drop
the `cause` assertions).

### PR 3 — `ServiceContainer.initialize()` as a runtime-run startup
effect with per-step timeouts
**Value:** medium (a hung `taskService.initialize()` currently pins the
splash screen forever; deterministic TestClock tests of startup).
**Risk:** low–medium. **Net product LoC ≈ +80** (step table ~20,
timed-step helper ~15, `StartupStepTimeoutError` ~8, constant ~3, facade
~10, root dispose-on-failure parity ≤ 10, doc update ~10).

Files: `serviceContainer.ts`, `src/constants/terminationTimeouts.ts`
(keep with the termination constants so the budget doc stays in one
place), `cli/server.ts` (dispose in the startup catch if missing),
`di/appRuntime.ts` doc ("Deliberately not done" → remove the
initialize() line; add startup contract).

Design (D6): `initialize(): Promise<void>` →
`this.runtime.managed.runPromise(this.startupEffect())`. `startupEffect
= Effect.gen` over an ordered `readonly steps: ReadonlyArray<{ name,
run: () => Promise<void> }>` (assert names unique); each step:
`recordStep` timing kept, `Effect.tryPromise({ try: async () => run(),
catch: identity }).pipe(Effect.timeoutOrElse({ duration:
STARTUP_STEP_TIMEOUT_MS, orElse: () => Effect.fail(new
StartupStepTimeoutError(name, STARTUP_STEP_TIMEOUT_MS)) }))`. Then
`Effect.sync` for the three `start()`s; the sweeps remain after
`runPromise`. Constant `STARTUP_STEP_TIMEOUT_MS` — pre-work measures
`[startup] <step> { ms }` across sandbox cold starts and picks ≥ 10× the
slowest observed (propose 60 s; must be generous — a false timeout turns
a slow-but-fine start into a crash). Roots dispose after a rejected
`initialize()` (D6 abandon-and-quit safety).

Acceptance (all in `serviceContainer.test.ts`, TestClock via the
existing `AppLive` spy at `:355–395`):
- A step that never resolves → `initialize()` rejects with
`StartupStepTimeoutError` naming the step after exactly
`TestClock.adjust(STARTUP_STEP_TIMEOUT_MS)`; later steps did not run.
- A rejecting step → `initialize()` rejects with **the same error
object** (identity), later steps did not run (parity with today).
- Happy path → `stepDurationsMs` has all six keys; `start()`s called
once each; second `initialize()` call behavior unchanged from today
(verify whether re-entry is guarded today; preserve).
- `tests/ipc` harness and ACP entry still pass unchanged.

Gate suites: `serviceContainer.test.ts`, `coreServicesRoot.test.ts`,
`di/*.test.ts`, `src/node/acp/*.test.ts`, `TEST_INTEGRATION=1 bun x jest
tests/ipc` (smoke subset); `make static-check`.

Audits: I1 untouched (no layer body changes); I2 (only the composition
root touches the runtime); error identity preserved (no wrapping);
abandoned-step safety — every root runs bounded `dispose()` after a
rejected `initialize()` (D6; `cli/server.ts` gains it in this PR); no
sweep moved into the effect; the six-step order and `[startup] <step>`
names unchanged.

Rollback: revert; constant removal.

### PR 4 (optional — cut if budget is exhausted) — `streamBridge` on the
runtime context
**Value:** low (closes the last documented "global runtime" exception;
enables TestClock for the heartbeat). **Risk:** low. **Net product LoC ≈
+30** (`context?` option + `toAsyncIterableWith` ~8; 19 call sites × 1
line via one shared helper in `routerSubscriptions.ts` ~3).

Acceptance: the 3 clock-bound waits (`:207, :241, :255`) run on
`TestClock`; the heartbeat test asserts N heartbeats after
`TestClock.adjust(N × interval)` with zero real time; existing
behavioral assertions unchanged; `tests/ipc` subscription tests pass. Do
not rewrite the 8 readiness polls.

Gate suites: `streamBridge.test.ts`, `routerSubscriptions*.test.ts`,
`orpc/*.test.ts`, `TEST_INTEGRATION=1 bun x jest tests/ipc`
(subscription subset); `make static-check`.

Audits: `Stream.toAsyncIterableWith` preserves double-close safety (pin
with the existing test); `Cause.Done` typing unchanged; no
`Scope`/`MemoMap`/`Scheduler` captured (pass the oRPC `effect/context`,
which the DI layer already strips per `EffectRunnerLive`); `context`
stays optional so direct callers/tests without a runtime keep today's
global-runtime path.

Rollback: revert; the optional `context` default (`Context.empty()`) is
exactly today's `toAsyncIterable`, so a partial revert of call sites is
also safe.

### Execution order and size
Net product LoC for the wave ≈ **+205** (PR 1 ≈ +55, PR 2 ≈ +40, PR 3 ≈
+80, PR 4 ≈ +30); tests ≈ +600–800. PR 1 starts immediately. PR 2 starts
the moment coder#3949 merges (parallel with PR 1/3 — disjoint files). PR 3
after PR 1 merges (both touch `appRuntime.ts` docs; PR 3 also touches
`serviceContainer.ts`). PR 4 last, only if PRs 1–3 landed and no
OFF-RAMP fired. Each PR: Codex dual review, `Codex Comments`
minimization, merge queue; commit WIP early (`/tmp` wipes).

## 4. STOP criterion (measurable) and OFF-RAMPs

Wave 4 is **done** — and the Effect migration line **stops** without a
new RFC — when all hold:
1. **dispose() awaits in-flight streams:** `serviceContainer.test.ts`
ordering test + `coreServicesRoot.test.ts` pass on main; a sandbox
`script -f` transcript of `xum server` receiving SIGTERM mid-stream
shows `stream-abort` → `[shutdown] AppFiberScope closed { ms }` →
`[shutdown] desktopBridgeServer.stop`, and immediately after exit
`partial.json` is absent while `chat.jsonl` contains the interrupted
assistant message (baseline on `main`: `partial.json` present, message
absent until next load). `{ ms }` < 2000 in the flowing-stream case.
2. **False-settle class eliminated:** the scripted reproduction fails on
the pre-coder#3949 merge-base and passes on main after PR 2;
`settleWorkspaceTurn` rejects any settlement without an enumerated
cause; the coordinator's own Mux sessions show zero "superseded by an
uncorrelated workspace stream-end" in the two weeks after PR 2 (soft
signal, logged in the wave summary).
3. **Startup:** timeout and error-identity tests pass under TestClock;
`[startup]` per-step lines unchanged in the sandbox transcript; a
throwaway build with the constant set to 1 ms shows `Startup failed:
StartupStepTimeoutError: <step> exceeded 1 ms` and a clean exit.
4. **No new lifecycle flakes:** 0 failures attributable to the touched
suites across **N = 20** consecutive *completed* `Test / Unit` runs on
`main` after the last Wave 4 merge — query `gh run list --workflow
pr.yml --branch main --limit 60 --json
databaseId,status,conclusion,event,headSha` (note: `gh run list --json`
serializes these fields in **lowercase**, e.g.
`{"status":"completed","conclusion":"success"}`, unlike
`statusCheckRollup`), keep `status === "completed"` (pending runs have
an empty `conclusion`, not null), take the newest 20, and for any run
with `conclusion !== "success"` (case-insensitive normalization
acceptable) inspect the failing job's log for the touched suite names
(job `timeout`/`cancelled` from the 15-min budget is not a flake); plus
green merge-queue runs for each PR. Any attributable flake → fix or
revert before declaring done.

**OFF-RAMP (PR 1):** fires if pre-work 1–3 shows (a) routing shutdown
through `cancelStreamSafely` cannot preserve crash-recovery semantics
without changing `cleanupAbortedStream`'s contract beyond D3, (b) the
chaos variant exposes a double-settle not closable by D3(b), or (c) the
finalizer cannot fit the 2 s bound for flowing streams. Then: stop PR 1,
keep `AppFiberScope` unoccupied, update `appRuntime.ts` "Deliberately
not done" with the concrete blocker and the measured evidence, land D3
alone as a bug-fix PR. PRs 2–4 are independent and proceed.
**OFF-RAMP (PR 3):** if error identity or the `tests/ipc`/ACP paths
cannot be preserved, keep `initialize()` as is and record why.
**OFF-RAMP (PR 2):** coder#3949 not merged → hold (see PR 2).

## 5. Risk register

| Risk | Likelihood | Mitigation |
|---|---|---|
| Crash-recovery regression: double commit / partial resurrection when
shutdown-abort races completion | medium (window widens with dispose())
| D3(a) guard + test; `commitPartial`'s `historySequence`
update-or-append is idempotent (`historyService.ts:2036–2041`) |
| Provider abort emits an `error` chunk → error path instead of abort
path | low | identical to today's user-stop path (parity); chaos variant
covers hostile streams |
| Wedged provider pins the 2 s bound → warning every shutdown | low |
`boundedTeardown` already bounds; transcript measures; no budget change
possible (5 s outer) |
| AIService `stream-abort` listener (`commitPartial`) still in flight
when `process.exit` runs | low | durable order writePartial → commit →
deletePartial; next-load recovery; PR 1 transcript checks `partial.json`
is already gone when `cli/server.ts` logs its final cleanup line before
`process.exit(0)` (`:252–266`) |
| Chaos-test seams (`createStreamResult`, `tokenTracker`) | none if
scope-less construction stays default | new variant added, old cases
untouched |
| Collision with coder#3915/coder#3949 | medium | PR 2 gated; no edits to their
regions; one-line `cause:` conflicts only |
| RC churn (rc.113+ renames
`forkIn`/`onInterrupt`/`toAsyncIterableWith`) | low | all Effect imports
already in `streamManager.ts`/`streamBridge.ts`; pins fixed; GA upgrade
is a separate lockstep PR (§6) |
| Startup false timeout on slow hosts | medium if constant too small |
measure first; ≥ 10× slowest observed; generous default (60 s) |
| Sync-start assumptions in tests that
`Reflect.set(processStreamWithCleanup)` | low | promise assigned before
fork; forkIn `runSync` synchronous |
| `shutdown()` (desktop second `before-quit` listener) still does not
await streams | accepted | contract says `shutdown()` never touches the
runtime; desktop's dispose race is the covered path |
| Streams starting during shutdown | covered | `forkIn` on closed scope
interrupts immediately → `system` abort (`startImmediately` semantics
verified in rc.112; pinned by test) |
| `system` abort triggers an in-session RetryManager retry during
shutdown | **unreachable** (verified) | `"aborted"` ∈
`NON_RETRYABLE_STREAM_ERRORS` → `retryManager.ts:99–104` abandons; no
fiber scheduled |
| PR 3: abandoned `taskService.initialize` mid-`editConfig` when the
root exits after a timeout | low | desktop/ACP already dispose on
startup failure; PR 3 adds the missing `cli/server.ts` dispose; config
writes are lock/journal-protected |
| `startImmediately`/`onInterrupt` semantics differ in a later RC | low
| pinned by the probe test in `appFiberScope.test.ts` style; RC bumps
are a separate lockstep PR |

## 6. Standing item — effect v4 GA + `@orpc/experimental-effect`
lockstep (analysis only)

v4 is **not GA** (rc.112 is current; v3 `3.x` remains the stable line).
No PR this wave. When GA ships: one lockstep PR bumping `effect` + all
`@orpc/*` (`1.14.11` today; check the GA-compatible
`@orpc/experimental-effect`), canary gates = `di/*.test.ts`,
`streamBridge.test.ts`, `streamManager.test.ts`,
`serviceContainer.test.ts`, `TEST_INTEGRATION=1 bun x jest tests/ipc`,
`make static-check`. The `Context → ServiceMap` rename risk is
firewalled: `Context.Service` tags, `Context.omit/get`, `Layer`,
`ManagedRuntime`, `TestClock` live only under `di/` +
`orpc/effectContext.ts`; `streamManager.ts`/`streamBridge.ts` use
`Effect`/`Scope`/`Fiber`/`Exit`/`Stream`/`Queue`/`Cause` only. PR 4 adds
one `Context.Context<never>` type reference to `streamBridge.ts` — keep
it as a type-only import so a rename is a one-line fix.

## 7. Dogfooding (per PR; evidence attached to the PR with `gh …
--attach`)

Common setup: `make dev-server-sandbox
DEV_SERVER_SANDBOX_ARGS="--clean-projects"` (or `xum server` on a temp
`XUM_ROOT`) with `XUM_LOG_LEVEL=debug`, run under `script -f
~/wave4-scratch/<pr>-<scenario>.log`; scratch under
`$HOME/wave4-scratch/` (never `/tmp`). Drive the UI with `agent-browser`
(`open` → `snapshot -i` → click the explicit "Send message" ref;
re-snapshot after typing). Screenshots are primary evidence; record WebM
and finalize with `ffmpeg -c copy`.

- **PR 1:** (1) start a long stream (prompt that streams ~30 s), wait
3–5 s, `kill -TERM <server pid>`; transcript must show the order in STOP
coder#1 and `[shutdown] AppFiberScope closed { ms }`; (2) `ls
<XUM_ROOT>/sessions/<ws>/partial.json` (absent) + `tail -n 1 chat.jsonl`
(interrupted assistant message); (3) restart, open the workspace in
agent-browser, screenshot the persisted interrupted message; (4) same
scenario on `main` for the baseline diff; (5) `xum run` Ctrl-C
mid-stream transcript (CLI root parity); (6) quality gate between
phases: gate suites green before the sandbox run, sandbox evidence
before requesting review.
- **PR 2:** primary evidence is the regression file run on both the
merge-base worktree (failing output) and the branch (passing).
Secondary: sandbox parent workspace delegates via `task` kind=workspace
to a child that arms a background bash monitor firing within ~10 s and
keeps working ~60 s; screenshot the parent's task result (baseline
`main`: `interrupted … uncorrelated workspace stream-end`; after:
`completed`) and the child's log lines.
- **PR 3:** `xum server` cold start transcript with `[startup] <step> {
ms }` for the six steps (parity); throwaway worktree build with
`STARTUP_STEP_TIMEOUT_MS = 1` → transcript of `Startup failed:
StartupStepTimeoutError …` and exit code (do not ship); desktop dialog
cannot be shown headless — cite the unchanged
`desktop/main.ts:1249–1265` catch.
- **PR 4:** agent-browser session left idle 60 s with the connection
indicator visible (heartbeats keep it green) + screenshot;
`streamBridge.test.ts` on TestClock.

## 8. Non-goals (restated; out of this wave)

Typed-error propagation sweep / removing the ~113 facades; converting
services to `yield*`-based Effect services; PubSub for the internal
EventEmitter bus; Schema at persistence boundaries; Effect
observability; converting sync read paths, AI-SDK per-request callbacks,
cross-process lock interiors, or deterministic try-lock funnels;
replacing `AbortController` as the SDK cancellation transport;
converting the `fullStream` loop to `Stream`; fiberizing turn-handle
waiters; downgrading startup steps to best-effort; changing outer quit
budgets; supervising the pre-registration stream-start window;
`shutdown()` semantics; the effect GA bump (standing analysis only).


</details>

---

_Generated with `xum` • Model: `anthropic:claude-fable-5-1` • Thinking:
`xhigh` • Cost: `$13.39`_

<!-- mux-attribution: model=anthropic:claude-fable-5-1 thinking=xhigh
costs=13.39 -->
yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Sep 9, 2026
…its (coder#4154)

## Summary

`AgentSession.sendMessage` checked the shutdown latch once, before
awaiting snapshot materialization and the turn's history appends. A
shutdown that began inside those awaits left the appended rows durable
while the turn was still refused, so the next startup read the
answer-less user row as an interrupted turn and auto-retried it. The
latch is now re-checked at the rollback horizon, where the rows are
still rollback-eligible, and the refusal goes through the existing
rollback path. Edits are the one exception: once their truncation has
happened, shutdown lets the replacement row land instead of losing both
versions of the input.

Fixes coder#4073

## Background

Codex flagged this on coder#4058 (round 23, P2). The pre-persist gate added
there refuses shutdown before any row lands, but everything after it
(`materializeAgentSkillSnapshots`, `materializeMcpPromptSnapshots`, the
snapshot/user-row appends) is awaited with no further latch check.
`isCurrentTurn` stays true while the coordinator is merely closing, so
the existing rollback-horizon checkpoint did not observe the shutdown
either; `coordinator.prepare()` later rejected the turn with `closing`,
returning `Err` to the caller with the row already durable.

## Implementation

- One new check at the rollback horizon (after the caller-probe
staleness checkpoint, before `markRowsDurable()`) that refuses through
`refuseBeforeAcceptance`, the same helper the pre-persist gate uses: it
rolls back this attempt's rows and marks them durable only if the
rollback verifiably failed. Covering the horizon rather than only the
two materialize awaits closes the append-I/O windows too, and every path
(ordinary, pre-turn batch, on-send compaction row, token-budget batch)
funnels through it.
- Edits: `truncateAfterMessage` runs before any of these checks and is
irreversible. Every pre-persist shutdown refusal an edit can reach
afterwards (the pre-persist gate, the two token-budget checks before the
batch append, and the new horizon check) now goes through
`shutdownRefusesBeforePersist()`, which stays false once the truncation
succeeded. The replacement row lands, the PREPARING gate refuses the
turn with the row retained, and startup recovery resumes the edit.
Before truncation, and for the missing-target no-op truncation, shutdown
still refuses as before.
- The `scopedLifetimes` join test previously pinned the retained row
after a shutdown landing inside an ordinary send's append; its join
assertions are unchanged and it now asserts the rolled-back outcome.

## Validation

- New `startupAutoRetry` cases run shutdown inside the stubbed
`materializeAgentSkillSnapshots` await and inside the user row's own
append (row durable, then removed). Both were red on `origin/main` at
the history assertion (refusal returned, one extra durable row) and are
green with the fix; seeded earlier rows survive the rollback.
- Edit cases: the same two windows with `editMessageId` keep `[u0, a0,
replacement]` and start no provider; both fail when the truncation flag
is not set.
- All 34 `agentSession*.test.ts` suites and `make static-check` pass
locally.

## Risks

Low and confined to shutdown ordering. The checks only fire when
`beginShutdown()` has already run, and the turn was already going to be
refused at the PREPARING gate; the change is that ordinary sends roll
their rows back instead of retaining them, while post-truncation edits
proceed to persist their replacement instead of refusing (previously the
pre-persist gate could refuse a truncated edit and leave neither
version). The emergency rollover path (`rolloverAfterBudgetFailure`) is
separate and untouched.

---

_Generated with `xum` • Model: `anthropic:claude-fable-5-1` • Thinking:
`xhigh` • Cost: `$0.00`_

<!-- mux-attribution: model=anthropic:claude-fable-5-1 thinking=xhigh
costs=0.00 -->
yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Sep 9, 2026
…message (coder#4161)

## Summary

`task_send_message` to a descendant task nested the task-tree lifecycle
lock outside the workspace event lock, while reported-task cleanup nests
them the other way (event lock held, then `WorkspaceService.remove()`
takes the tree lock). When a send to a task interleaved with that task's
cleanup, each side waited on the other's lock forever. The send path now
acquires the event lock first, so every path that holds both locks uses
one order, and a deterministic regression test pins the interleaving.

Fixes coder#4072

## Background

Surfaced by Codex on coder#4058, which added the live recheck inside
`remove()` and documented the inversion without changing it; the
ordering predates that PR.

## Implementation

- `dispatchTrustedDescendantMessage` now nests
`workspaceEventLocks.withLock(taskId, () =>
withTaskTreeLifecycleLock(taskId, ...))`. The joint critical section is
unchanged (both locks are still held for the whole body); only the
acquisition order flips.
- This direction was chosen over deferring `remove()` out of the
event-locked section because the event lock is the outermost lock
everywhere else it appears: the stream-end, stream-abort and error
listeners take it first, and cleanup rechecks, hard-timeout termination
and workspace-turn finalization all nest inside it. Moving cleanup would
have meant moving every tree-lock acquisition out from under
`handleStreamEnd`.
- The `workspaceEventLocks` declaration now documents the order (event
lock before task-tree lock) and the cleanup comment no longer describes
an inversion.

## Validation

- New test `reported-task cleanup and task_send_message never deadlock
on the event and task-tree locks`: drives the real
`requestReportedTaskCleanupRecheck` on a reported workflow-owned child,
parks a `sendMessageToDescendantAgentTask` on its lock acquisition
inside the window before the (lock-mirroring) `remove()` takes the tree
lock, then lets removal proceed. Red on the previous nesting (`Expected
"settled", Received "deadlocked"`), green after the swap; whole
`taskService.test.ts` 490/490.
- Remote dogfood UAT at this SHA (Bun 1.3.5) reproduced the same
red/green and full-suite results, and a real-model smoke confirmed
`task_send_message` to a reported child reactivates it, to a live child
queues, and stop-then-send reactivates, with no hang. The exact
cleanup/send overlap window is covered by the deterministic test only;
it is not observable by hand.

## Risks

Low. The send path holds the same two locks over the same body; only
which one it waits for first changes. Other tree-lock holders (create,
stop, retitle, remove, archive, unarchive) never take the event lock, so
the new order adds no wait edge.

---

_Generated with `xum` • Model: `anthropic:claude-fable-5-1` • Thinking:
`xhigh` • Cost: `$36.01`_

<!-- mux-attribution: model=anthropic:claude-fable-5-1 thinking=xhigh
costs=36.01 -->
yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Sep 9, 2026
…coder#4162)

## Summary

Best-of grouped finalization could commit the parent's task tool output
on a sibling report artifact that the sibling was already replacing.
Publication of a best-of child's report now runs under the same
per-parent lock as the grouped assembly, a resumed sibling counts as
executing again, and the partial commit re-checks the siblings
synchronously right before the write.

## Background

Fixes coder#4074 (Codex round 23 P1 on coder#4058).
`buildBestOfCompletedTaskToolOutput` read every sibling's report
artifact, then re-sampled only current activity
(`isBestOfSiblingExecutingAgain`: active workspace-turn status or
streaming) before handing the grouped output to the partial commit. Two
windows let a stale artifact through:

- A sibling's continuation that started after its artifact was read and
finished before the re-check was idle again at the re-check, so the
assembly finalized on the old artifact already in `reports`.
- `finalizeAgentTaskReport` flipped the child to `reported` before
writing its artifact, both outside `deferredBestOfLocks`. An assembly
that ran between the flip and the write saw a reported, idle sibling
with its old artifact.

In both cases the sibling's later delivery found the grouped output
already finalized, so its replacement report was dropped from the tool
result.

## Implementation

- `finalizeAgentTaskReport` moves the status flip plus artifact writes
into `publishAgentTaskReport` and, for best-of children, runs it under
`deferredBestOfLocks(parent)`, the lock the assembly already holds. The
lock is released before `deliverReportToParent` re-acquires it, so the
critical section stays the publication only.
- `isBestOfSiblingExecutingAgain` also treats `taskStatus` `running` /
`awaiting_report` as executing again: a resumed sibling (for example
`markInterruptedTaskRunning`) has no workspace-turn mirror and is not
streaming between its stream end and its publication.
`shouldDeferBestOfFallback` now relies on the shared predicate instead
of repeating those statuses.
- `areBestOfSiblingsStillSettled` is a synchronous check (sibling
identities plus executing-again) used by the assembly's live re-check
and again inside the `updatePartialIfMessageIdMatches` transform, where
no await separates the check from the write. When it fails there,
`tryFinalizePendingTaskToolCallInPartial` returns `not_ready` so the
reporting sibling defers rather than appending a synthetic fallback; the
resumed sibling's own report finalizes the group.

## Validation

- New test `best-of finalization never ships a report a resumed sibling
is replacing`: a barrier on the parent's partial commit resumes the
already-reported sibling and drives its replacement report through the
production stream-end path while the first sibling's assembly is
mid-commit; the barrier lifts when the sibling requests the group lock.
Red on main (grouped output carried the pre-continuation report); green
with the fix.
- Inverted toggles over the whole `taskService.test.ts` file: removing
the publication lock, the status-aware predicate, or the commit-time
re-check each turns the new test red again.
- `grouped partial finalization exposes each terminal report exactly
once` seeded a `running` sibling with an on-disk artifact, which is now
by definition a sibling whose report is being replaced; its fixture
models the settled state (`interrupted` with the artifact) so its
exactly-once assertion still holds.

## Risks

Moderate, scoped to best-of report delivery. The new lock hold in
`finalizeAgentTaskReport` covers config edits, metadata emission, and
artifact writes; it acquires `deferredBestOfLocks(parent)` before
`workspaceFileLocks(parent)`, the same order the assembly uses, and is
released before delivery re-acquires it. A best-of sibling with a stale
artifact whose status is `running` or `awaiting_report` now defers
grouped finalization until it reports again, where it previously
finalized on the stale artifact.

---

_Generated with `xum` • Model: `anthropic:claude-fable-5-1` • Thinking:
`xhigh` • Cost: `$0.00`_

<!-- mux-attribution: model=anthropic:claude-fable-5-1 thinking=xhigh
costs=0.00 -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

🤖 perf: xum server takes 13 min to bind its port on large deployments; TaskService.initialize runs O(workspaces) sequential recovery before listen()

1 participant