Chatd has 4 main pieces:
- core state machine: describes how a chat's state in the database can change over time. It defines the valid states and transitions for committed chat data: status, messages, queued messages, pending actions, worker ownership, and the fields used to reject stale work. It's a specification implemented by chatstate/machine.go. Runtime components, such as the HTTP endpoints and the chat worker, use it to ensure that they modify the state only in valid ways.
- API surface: the HTTP endpoints that coderd exposes. Responsible for: creating chats, sending messages, editing messages, updating metadata, managing the queue, interrupting active work, and submitting tool results. These are used by the client, usually via the browser, to interact with chats.
- chat worker: lives inside every coderd replica. It acquires chats, calls the LLM API, executes tools, handles interrupts and tool-result waits, and commits completed outcomes through the core state machine.
- stream loop: powers
GET /api/experimental/chats/{chat}/stream, the WebSocket endpoint that the UI uses to consume a live chat. It combines two kinds of data: messages committed to the database and streaming message parts emitted by the chat worker. It receives notifications over pubsub whenever the chat state is updated, fetches messages from the database, and connects to the coderd replica that currently owns the chat to relay the streaming message parts to the client.
Chatd attributes AI Gateway requests with a synthetic API key owned by the chat owner, one key per user. There is no mapping table: the key is found in api_keys by its deterministic token name, chatd_<owner_id>_session_token, excluding login_type = 'token' rows. Token names are unvalidated user input, so the login type filter ensures chatd never picks up (or extends) a real bearer token a user created with the colliding name. Synthetic keys are minted with the owner's login type, which is never 'token'. All chatd AI Gateway attribution resolves the key from chats.owner_id; callers do not provide the key ID.
Synthetic keys expire after 30 days. When less than 24 hours remain, chatd extends the expiry of the existing row in place instead of replacing it, because an in-flight generation may have already delegated the current key ID to the gateway. The key ID is therefore stable for the lifetime of the user. Mints and extensions are serialized with a per-user advisory lock, since the partial unique index on token names only covers login_type = 'token' rows. The generated token is discarded, so the stored key cannot be used as a bearer credential, and it carries a minimal scope as defense in depth.
Messages and queued messages no longer carry api_key_id columns; attribution is resolved solely from chats.owner_id. The drop migration discards any IDs stamped by older replicas, and its rollback restores the columns as nullable without backfilling them.
Deleting a synthetic key (password reset, explicit key deletion, dbpurge of long-expired keys) does not touch chat messages, queued messages, or their version fields. Chatd mints a replacement on the next request without mutating history. User suspension and deletion still block delegated gateway authorization.
The core state machine describes how a chat's execution state in the database can change over time. A fundamental component of the state machine is the set of valid states it can be in. We will consider 2 kinds of states: execution states and ownership states. These states let us describe what the runtime components of chatd can do with a chat at a given point in time.
We say that the following data constitutes a chat's execution state:
- chat status on the
chatstable, such aswaiting,running,interrupting,requires_action, orerror; - the
archivedmarker on thechatstable; - message history in the
chat_messagestable, including therevisionfield; - queued user messages in the
chat_queued_messagestable, including thepositionandcreated_byfields; worker_idandrunner_idfields on thechatstable (ownership fields);- the
last_errorfield on thechatstable (last error message from the agent loop); - the
retry_statefield on thechatstable, a JSONB object that stores the last error message encountered by the agent loop, and information about when the next retry will be attempted; - the
snapshot_version,history_version,queue_version,generation_attempt,retry_state_versionfields on thechatstable, defined later in the document; - the
requires_action_deadline_atfield on thechatstable (pending-action deadline, defined later in the document);
There is other data that is held in the database and is associated with a chat, but it's not part of the execution state:
- title;
- labels;
- pin order;
- workspace binding;
- model configuration;
- plan mode;
- file links.
We call it metadata. The core state machine concerns itself with execution state. As a general guideline, a piece of data is execution state if the core state machine needs it to decide what the next state transition may be, or if it's directly modified by a state transition. For example, a queued message is part of the execution state because it impacts what the next action of the agent loop can be. If the agent loop finishes processing a user message and would otherwise stop, but there's a queued message, the agent loop will start processing the queued message instead. On the other hand, a chat's title does not impact the agent loop at all - it's just a label that helps the user identify the chat.
File links are metadata, but they are written inside transitions: if a transition persists message content that references uploaded files (chat create, message send, queued send, or message edit), it records the file links in the same transaction. Two invariants are enforced when links are written. A file belongs to at most one chat: attaching a file that another chat already holds is refused in the same way as attaching a file that no longer exists. There is an upper bound on the number of files a chat holds. When a message's files would push the chat over the cap, the oldest files on the chat are deleted to make room, and their links go with them. Files in the same message are never evicted by that message, so only a message that is on its own larger than the cap is rejected. Files created by tools during a run take the same path, so a tool's attachment can evict a user's upload and vice versa.
Eviction means that a persisted message may reference a file that no longer exists. That's expected: the UI shows the attachment as expired, and when the history is sent to the model, an evicted user upload is replaced with a short placeholder saying the content has expired, while evicted assistant and tool files are dropped. Editing a message that still references an evicted file is refused until the attachment is removed from the edit.
If the distinction isn't completely clear to you at this point, don't worry. It should become clearer as you learn more about the core state machine.
A chat's execution state lets the chat worker and the HTTP endpoints decide what they can do with the chat. In total, there are 13 execution states. The states are decided by what's in the database:
- By whether a chat exists;
- By all the chat statuses on the
chatstable:waiting,running,interrupting,requires_action, anderror; - By the
archivedmarker on thechatstable; - By the queued messages in the
chat_queued_messagestable.
The shorthands in the table below use the convention that the first 1 or 2 letters indicate the status, and then 1 or 0 indicate the presence or absence of queued messages.
| Shorthand | Status | Queue | Archived | Meaning |
|---|---|---|---|---|
N |
- | - | - | Chat does not exist |
W |
waiting |
empty | false |
There's no work to be done by the chat worker |
E0 |
error |
empty | false |
The worker encountered an unrecoverable error while processing the chat. There's no more work to be done by the chat worker |
E1 |
error |
non-empty | false |
The worker encountered an unrecoverable error while processing the chat, and there's currently no work to be done by the chat worker. There's a queued message that should be processed once the error is cleared |
R0 |
running |
empty | false |
Running state with no queued messages: a chat worker should be processing the chat |
R1 |
running |
non-empty | false |
Running state with queued messages: a chat worker should be processing the chat, and there's a queued message that should be processed next |
I0 |
interrupting |
empty | false |
The chat was interrupted by the user, and the chat worker should commit any partial message that had been generated before the interruption |
I1 |
interrupting |
non-empty | false |
The chat was interrupted by the user, and the chat worker should commit any partial message that had been generated before the interruption, and there's a queued message that should be processed next |
A0 |
requires_action |
empty | false |
The chat worker is waiting until the user submits tool results; this state is used only by the “dynamic tools” feature |
A1 |
requires_action |
non-empty | false |
The chat worker is waiting until the user submits tool results, and there's a queued message that should be processed next; this state is used only by the “dynamic tools” feature |
XW |
waiting |
empty | true |
The chat was archived while it was in the waiting state, it will go back to waiting once unarchived |
XE0 |
error |
empty | true |
The chat was archived while it was in the error state, it will go back to error once unarchived |
XE1 |
error |
non-empty | true |
The chat was archived while it was in the error state, it will go back to error once unarchived, and there's a queued message that should be processed once the error is cleared |
If these states seem arbitrary and abstract at this point, that's expected. Each one of these states is needed by some runtime component of chatd for some specific use case, and their purpose will emerge as we discuss the implementation of the HTTP endpoints and the chat worker.
At a high-level, these states let us reason about what should be possible to happen with a chat at a given point in time. For example, a chat in the R0 state can be picked up by a chat worker, an LLM message can be appended to its history. On the other hand, a chat in the XW state must be ignored by the chat worker, and most of the HTTP endpoints must refuse to interact with it. We'll define precisely what is possible in each state in the Transitions section.
A chat's ownership state lets the chat worker decide whether a chat can be acquired or not. It's decided by the worker_id field on the chats table. In total there are 2 ownership states.
| Shorthand | Worker ID | Meaning |
|---|---|---|
U |
null | Unowned chat |
O |
not null | Owned chat |
Now that we've defined the states, we can define the transitions between them. In practice, a transition is just a sequence of SQL queries that modify the database state in a transaction. That transaction first takes a row lock on the chat to ensure that it's serialized with respect to other transactions that modify the chat. Multiple transitions can be executed atomically in a single transaction.
Remember!
a transition is just a sequence of SQL queries that modify the database state in a transaction
We will not define the SQL queries that correspond to each transition - it'd take too much space and it's not central to the document's purpose. Instead, we focus on what each transition does to the database state, and how it affects the execution and ownership states.
Each transaction that applies one or more transitions advances the snapshot_version field on the chats table by 1 immediately after locking the chat row and before mutating any tables. This lets us version the chat's execution state. The chat worker and the stream loop rely on it to ensure they do not process outdated or out of order notifications.
Chat-message changes update history_version on the chats table and the revision fields on the chat_messages table automatically via Postgres triggers described in Message revisions and history version. history_version stores the latest snapshot_version in which chat message history changed. The chat runner and the stream loop rely on it to ensure they are fully aware of the chat's history changes. See Event processing for how the runner uses history_version differently from snapshot_version.
Queue changes update queue_version automatically via Postgres triggers described in Queue version.
I don't recommend reading the rest of section thoroughly if this is your first time reading this document. It's an information dump that only makes sense once you pair it with a specific runtime component of chatd. Give it a cursory look, and treat it as a reference that you can return to later when you're analyzing how an HTTP endpoint or a chat worker implements a specific feature.
Create(initialMessages)creates a new chat, initializessnapshot_versionto 1, inserts its initial history, and lands inrunning. The inserted initial history setshistory_versionto 1. Since the queue has not changed,queue_versionremains 0. This transition is a special case: since the chat does not exist at the time it's run, the chat row cannot be locked before the transition is applied.SetArchived(archived)sets or clears the archived marker for one chat.SendMessage(m, busy_behavior)inserts a user message directly when the chat is idle, or queues it when the chat is busy.busy_behaviormust be eitherqueueorinterrupt. Withbusy_behavior=interrupt, it also requests interruption or cancels a pending dynamic-tool action as needed.EditMessage(k, replacement)clears queued messages, cancels or obsoletes active work, marks the truncated active-history suffix as deleted, inserts the replacement turn followed by any caller-provided suffix messages, and lands inrunning.DeleteQueuedMessage(qid)removes one queued message without changing the active history.PromoteQueuedMessage(qid)makes a queued message the next message to process. It reorders the queue, interrupts active work, cancels pending dynamic-tool action, or promotes into history immediately as required by the input state.Interrupt(reason)requests cancellation of an active generation or closes pending dynamic-tool action. It preserves queued backlog.CompleteRequiresAction(results)inserts submitted tool-result messages followed by any caller-provided suffix messages, clearsrequires_action_deadline_at, and lands inrunning. It preserves queued messages.RequestCompactionrecords a manual compaction request on an idle or errored chat by settingcompaction_requested_atand landing inrunningwithout inserting any message. It clearslast_errorper the leave-error rule, advanceshistory_versionto the transaction's newsnapshot_version, and resetsgeneration_attempt, so the compaction turn gets a full retry budget and message part episode keys that cannot collide with episodes retained from the previous turn. The chat worker picks the chat up like any other running chat and consumes the request. See Manual compaction.ClearContext(messages)commits a manual context reset synchronously, without involving the chat worker. It inserts the caller-built compressed clear boundary triplet (a hidden model-only sentinel user row, plus visible syntheticchat_clearedtool-call and tool-result messages), clearslast_errorand any pendingcompaction_requested_at, leaves ownership untouched, and lands inwaiting. No worker turn or model call follows; the message insert trigger advanceshistory_versionand resetsgeneration_attempt.E1is rejected because no waiting-with-queue state exists and a synchronous clear has no turn after which the queue would drain.
Acquire(worker_id, runner_id)locks the chat row, setschats.worker_idandchats.runner_id, and inserts an initial heartbeat row for(chat_id, runner_id).Abandonclearsworker_idandrunner_idon the chat row.CommitStep(step)inserts one durable message suffix while remainingrunning. A committed step may insert ordinary assistant/tool messages, and a compaction step may insert a compressed summary boundary plus visible compaction tool-call and tool-result messages, optionally followed by uncompressed model-only user rows replaying the pending-user segment (see Manual compaction).EnterRequiresActionrecords a pending-action episode by relying on the committed assistant tool-call messages as the durable call set, setsrequires_action_deadline_at, which is a timestamp 5 minutes in the future, and lands inrequires_action.FinishInterruption(optionalPartialStep)inserts one final interrupted assistant/tool suffix if present, or finalizes interruption without a suffix if none is available, clears the interrupting state, and lands inwaitingif no queued message is promoted. If interrupt finalization also promotes the queue head, it inserts the promoted queued message into history and lands inrunning.RecordGenerationAttemptverifies the chat is stillrunning, incrementsgeneration_attempt, and returns the updated chat snapshot.RecordRetryState(payload)verifies the chat is stillrunning, stores the retry payload sent to clients asretry_state, and returns the updated chat snapshot.FinishTurncompletes the current generation turn atomically. If the queue is empty, it lands inwaiting. If the queue is non-empty, it removes the queue head, inserts it into history as a user turn, and lands inrunning.FinishError(err)parks the chat inerrorand persistslast_error = err, replacing any previously stored error. It is allowed when an unarchived chat is waiting or running.CancelRequiresAction(reason)closes pending dynamic tool calls with synthetic cancellation tool results, satisfies the pending-action projection, clearsrequires_action_deadline_at, and lands inrunning.ReconcileInvalidStatereconciles a chat in an invalid state by setting it to a valid state. Defined in the Invalid states section.
Now comes maybe the densest part of this document. It's a diagram that shows all the possible transitions between all the execution states. Again, I don't recommend reading the diagram thoroughly at first. Take a quick look to get a sense of what it's about and treat is as a reference you can return to later. I recommend reading the diagram as text and not looking at the rendered visual. The text is clearer.
A transition between input state A and output state B is allowed only if it's listed in the diagram below (A --> B: Transition Name). If a transition is not allowed, the core state machine implementation must reject it.
stateDiagram-v2
direction LR
[*] --> N
N --> R0: Create
W --> R0: SendMessage
W --> R0: EditMessage
W --> R0: RequestCompaction
W --> W: ClearContext
W --> E0: FinishError
W --> XW: SetArchived(true)
E0 --> R0: SendMessage
E0 --> R0: EditMessage
E0 --> R0: RequestCompaction
E0 --> W: ClearContext
E0 --> XE0: SetArchived(true)
E1 --> R1: SendMessage
E1 --> R0: EditMessage
E1 --> R1: RequestCompaction
E1 --> E0: DeleteQueuedMessage / removed last queued
E1 --> E1: DeleteQueuedMessage / queue still non-empty
E1 --> R0: PromoteQueuedMessage / promoted last queued
E1 --> R1: PromoteQueuedMessage / queue still non-empty
E1 --> XE1: SetArchived(true)
R0 --> R0: RecordGenerationAttempt
R0 --> R0: RecordRetryState
R0 --> R0: CommitStep
R0 --> A0: EnterRequiresAction
R0 --> I0: Interrupt
R0 --> I1: SendMessage(interrupt)
R0 --> R0: EditMessage
R0 --> W: FinishTurn / queue empty
R0 --> E0: FinishError
R0 --> R1: SendMessage(queue)
R1 --> R1: RecordGenerationAttempt
R1 --> R1: RecordRetryState
R1 --> R1: CommitStep
R1 --> A1: EnterRequiresAction
R1 --> I1: Interrupt
R1 --> I1: SendMessage(interrupt)
R1 --> R0: EditMessage
R1 --> E1: FinishError
R1 --> R1: SendMessage(queue)
R1 --> R0: DeleteQueuedMessage / removed last queued
R1 --> R1: DeleteQueuedMessage / queue still non-empty
R1 --> I1: PromoteQueuedMessage
R1 --> R0: FinishTurn / promoted last queued
R1 --> R1: FinishTurn / queue still non-empty after promoting head
I0 --> I1: SendMessage
I0 --> R0: EditMessage
I0 --> W: FinishInterruption
I1 --> I1: SendMessage
I1 --> R0: EditMessage
I1 --> I0: DeleteQueuedMessage / removed last queued
I1 --> I1: DeleteQueuedMessage / queue still non-empty
I1 --> I1: PromoteQueuedMessage
I1 --> R0: FinishInterruption / promoted last queued
I1 --> R1: FinishInterruption / queue still non-empty after promoting head
A0 --> R0: CompleteRequiresAction
A0 --> R0: Interrupt
A0 --> R0: CancelRequiresAction
A0 --> A1: SendMessage(queue)
A0 --> R1: SendMessage(interrupt)
A0 --> R0: EditMessage
A1 --> R1: CompleteRequiresAction
A1 --> R1: Interrupt
A1 --> R1: CancelRequiresAction
A1 --> A1: SendMessage(queue)
A1 --> R1: SendMessage(interrupt)
A1 --> R0: EditMessage
A1 --> A0: DeleteQueuedMessage / removed last queued
A1 --> A1: DeleteQueuedMessage / queue still non-empty
A1 --> R0: PromoteQueuedMessage / promoted last queued
A1 --> R1: PromoteQueuedMessage / queue still non-empty
XW --> W: SetArchived(false)
XE0 --> E0: SetArchived(false)
XE1 --> E1: SetArchived(false)
[Invalid] --> E0: ReconcileInvalidState / no queued messages
[Invalid] --> E1: ReconcileInvalidState / queued messages
The ownership state transition diagram is much simpler. It shows all the possible transitions between all the ownership states.
stateDiagram-v2
direction LR
[*] --> U
U --> O: Acquire(worker_id, runner_id)
O --> O: Acquire(worker_id, runner_id)
O --> U: Abandon
Notice that the Acquire and Abandon transitions only affect ownership state, and not execution state. They are fully orthogonal to the execution state transitions and have separate diagrams. This means that the chat's execution state can change independently of its ownership state, and vice versa. An archived chat may be acquired by a chat worker, and the core state machine's data model does not prevent that. The actual implementation of the chat worker will ignore chats that are in execution states that don't need processing, but it's not a concern of the core state machine.
- Any transition that's not
CompleteRequiresActionwhich supportsA0orA1as input states, and lands in output states different fromA0andA1, must insert synthetic, cancellation tool-call results for pending dynamic tool calls to avoid corrupting the message history. - Any transition that inserts a new user message into active history must answer outstanding tool calls in active history before inserting the user message. It may do this by inserting synthetic cancellation tool-call results.
- Any transition leaving
E0orE1(exceptSetArchived(true)) should clear thelast_errorfield.
Right after the refactor described in this document is complete, some chats may be in invalid states. For example, a chat may have archived set to true and status to running, which isn't allowed by the new state machine. To get the chat out of an invalid state, the ReconcileInvalidState transition is used, which does the following:
- Increment
snapshot_versionby 1. - Set
archived = false. - Set
status = 'error'. - Set
last_errorto an error message describing the chat was in an invalid state and a new message should be submitted or the message history should be edited to continue. - Set
requires_action_deadline_at = null. - If the chat has pending dynamic tool calls, insert synthetic cancellation results for them.
This will land the chat in either E0 or E1, depending on whether it has any queued messages.
Users can reconcile a chat's state by calling the POST /api/experimental/chats/{chat}/reconcile-invalid endpoint.
Each row in chat_messages has a revision column. It stores the chats.snapshot_version of the transition that last inserted or meaningfully updated that message row. revision is mutable, trigger-managed, and not unique. Multiple message rows can share the same revision when they are changed in the same transaction.
chats.history_version stores the latest snapshot_version in which chat message history changed. It starts at 0, remains unchanged for non-history transitions, and is set to the current snapshot_version whenever a message is inserted or meaningfully updated. A newly created chat starts with snapshot_version = 1; because Create inserts initial history in that snapshot, the created chat's history_version becomes 1. No-op message updates do not advance message revision, advance history_version, or reset generation_attempt. Whenever history_version changes, generation_attempt is reset to 0; generation attempts are scoped to the current history version.
Message revision triggers depend on the transition invariant that snapshot_version is allocated immediately after the chat row is locked and before any message mutation happens. Runtime code must not assign chat_messages.revision directly, and every chat_messages insert or update must go through a state machine transition: the triggers advance history_version on any write, so an out-of-band write (even of a hidden or soft-deleted row) moves history_version without a matching snapshot_version bump and breaks the fence of an in-flight generation task.
A BEFORE INSERT trigger assigns the current chat snapshot_version to the inserted message row and records the same value as the chat's latest history version:
CREATE FUNCTION set_chat_message_revision()
RETURNS trigger AS $$
DECLARE
chat_snapshot_version bigint;
BEGIN
IF TG_OP = 'INSERT' AND NEW.revision IS NOT NULL THEN
RAISE EXCEPTION 'chat_messages.revision must be assigned by trigger';
END IF;
IF TG_OP = 'UPDATE' THEN
IF OLD.chat_id IS DISTINCT FROM NEW.chat_id THEN
RAISE EXCEPTION 'chat_messages.chat_id is immutable';
END IF;
IF OLD.revision IS DISTINCT FROM NEW.revision THEN
RAISE EXCEPTION 'chat_messages.revision must be assigned by trigger';
END IF;
IF OLD IS NOT DISTINCT FROM NEW THEN
RETURN NEW;
END IF;
END IF;
UPDATE chats
SET
history_version = snapshot_version,
generation_attempt = 0
WHERE id = NEW.chat_id
RETURNING snapshot_version INTO chat_snapshot_version;
IF chat_snapshot_version IS NULL THEN
RAISE EXCEPTION 'chat % does not exist', NEW.chat_id;
END IF;
NEW.revision = chat_snapshot_version;
RETURN NEW;
END;
$$ LANGUAGE plpgsql;
CREATE TRIGGER trigger_set_chat_message_revision_on_insert
BEFORE INSERT ON chat_messages
FOR EACH ROW
EXECUTE FUNCTION set_chat_message_revision();A BEFORE UPDATE trigger uses the same function for message row updates:
CREATE TRIGGER trigger_set_chat_message_revision_on_update
BEFORE UPDATE ON chat_messages
FOR EACH ROW
EXECUTE FUNCTION set_chat_message_revision();chats.queue_version stores the latest snapshot_version in which the queue changed. It starts at 0, remains unchanged for non-queue transitions, and is set to the current snapshot_version whenever a queued message is inserted, updated, reordered, or deleted. A newly created chat with no queued messages has queue_version = 0.
An AFTER INSERT, AFTER UPDATE, and AFTER DELETE trigger records that the queue changed:
CREATE FUNCTION bump_chat_queue_version_on_queued_message_change()
RETURNS trigger AS $$
DECLARE
changed_chat_id uuid;
BEGIN
IF TG_OP = 'DELETE' THEN
changed_chat_id = OLD.chat_id;
ELSE
changed_chat_id = NEW.chat_id;
END IF;
UPDATE chats
SET queue_version = snapshot_version
WHERE id = changed_chat_id;
IF TG_OP = 'DELETE' THEN
RETURN OLD;
END IF;
RETURN NEW;
END;
$$ LANGUAGE plpgsql;
CREATE TRIGGER trigger_bump_chat_queue_version_on_queued_message_insert
AFTER INSERT ON chat_queued_messages
FOR EACH ROW
EXECUTE FUNCTION bump_chat_queue_version_on_queued_message_change();
CREATE TRIGGER trigger_bump_chat_queue_version_on_queued_message_update
AFTER UPDATE OF content, model_config_id, position, created_by
ON chat_queued_messages
FOR EACH ROW
EXECUTE FUNCTION bump_chat_queue_version_on_queued_message_change();
CREATE TRIGGER trigger_bump_chat_queue_version_on_queued_message_delete
AFTER DELETE ON chat_queued_messages
FOR EACH ROW
EXECUTE FUNCTION bump_chat_queue_version_on_queued_message_change();chats.retry_state_version stores the latest snapshot_version in which retry_state changed. It starts at 0, remains unchanged for transitions that do not affect retry state, and is set to the current snapshot_version whenever retry_state changes. A newly created chat starts with retry_state = null and retry_state_version = 0.
Retry state is scoped to the current generation attempt. Whenever generation_attempt changes, retry_state is cleared automatically. If that clear changes the value of retry_state, retry_state_version is set to the current snapshot_version.
A single BEFORE UPDATE trigger handles both clearing retry_state on generation-attempt changes and bumping retry_state_version on retry-state changes. The trigger mutates NEW directly and does not run an UPDATE chats ... statement, so it does not recursively trigger itself:
CREATE FUNCTION sync_chat_retry_state()
RETURNS trigger AS $$
BEGIN
IF OLD.retry_state_version IS DISTINCT FROM NEW.retry_state_version THEN
RAISE EXCEPTION 'chats.retry_state_version must be assigned by trigger';
END IF;
IF NEW.generation_attempt IS DISTINCT FROM OLD.generation_attempt THEN
NEW.retry_state = NULL;
END IF;
IF NEW.retry_state IS DISTINCT FROM OLD.retry_state THEN
NEW.retry_state_version = NEW.snapshot_version;
END IF;
RETURN NEW;
END;
$$ LANGUAGE plpgsql;
CREATE TRIGGER trigger_sync_chat_retry_state
BEFORE UPDATE OF retry_state, retry_state_version, generation_attempt
ON chats
FOR EACH ROW
EXECUTE FUNCTION sync_chat_retry_state();This section maps the public endpoints that mutate chat state to the transitions they use.
Chat routes are registered once by registerChatAPIRoutes and mounted under both /api/experimental and /api/v2 during a compatibility window. Paths in this document are written with one prefix or the other, but every promoted route answers on both. The routes that were not promoted answer only on /api/experimental; at the time of writing these are the providers and user-provider-configs collections under /chats, the computer-use-provider and advisor routes under /chats/config, GET /chats/{chat}/stream/desktop, and GET /chats/{chat}/debug/runs with GET /chats/{chat}/debug/runs/{debugRun}. Each mount reserves the top-level /chats/<segment> collection paths it does not serve (model-configs, providers, and user-provider-configs on /api/v2; models and model-configs on /api/experimental) so they return 404 instead of matching the {chat} wildcard and failing UUID parsing.
Clients discover models through GET /api/v2/organizations/{organization}/chats/models. The handler requires either full API token scope or chat model configuration read scope, then queries only configs in the requested organization that pass the caller's RBAC filter.
Provider configuration remains deployment-scoped and is read under Chatd's restricted system context. The response projects providers to redacted descriptors rather than exposing credentials, endpoints, or custom headers. It evaluates availability for the caller from deployment credentials and user-provided keys, and returns the readable model configs, provider availability, and unsupported provider types.
Model configuration writes are serialized per organization by inChatModelConfigWriteTx. The helper opens a ReadCommitted transaction and takes a Postgres advisory lock keyed by the organization ID before re-reading or changing model configs. ReadCommitted is required so reads after acquiring the lock observe writes committed by the previous lock holder.
The write paths maintain exactly one default whenever an organization has at least one live model config. The partial unique index permits at most one default per organization, while the locked write logic self-promotes the first config, unsets an old default before replacing it, and elects a replacement when the current default is demoted or deleted. Election prefers an enabled config whose provider is enabled, then falls back to another live config. If the only config is explicitly demoted, it is promoted again to preserve the invariant.
This endpoint uses Create(initialMessages):
N -> Create(initialMessages) -> R0
No other input states are supported.
When archiving or unarchiving a root chat, the operation applies SetArchived(archived) to the root and all descendants atomically. If any chat in the family cannot apply the requested archived-state transition, the whole operation fails without changing any chat. Unarchiving an individual child chat remains guarded: it must fail while its parent is archived
For archived updates, the supported input and output states are:
W -> SetArchived(true) -> XWE0 -> SetArchived(true) -> XE0E1 -> SetArchived(true) -> XE1XW -> SetArchived(false) -> WXE0 -> SetArchived(false) -> E0XE1 -> SetArchived(false) -> E1
If the request does not change archived, this endpoint doesn't emit any state transitions.
Other execution-state classes are not supported for archive/unarchive.
For busy_behavior=queue, SendMessage(m, queue) supports:
W -> SendMessage(m, queue) -> R0E0 -> SendMessage(m, queue) -> R0E1 -> SendMessage(m, queue) -> R1: this appendsmto the end of the queue, promotes the current queue head, and clears the error. The scenario where this happens is:- the user queued some messages
- the chat ran into an error and stopped, for example because of an unretriable problem with the LLM provider
- the user then sends a new message, but there is a non-empty queue. As defined here, the UX will be “add the new message to the end of the queue and promote the queue head.” Arguably, a better UX could be “add the new message to the chat immediately and start running it, even though there's a non-empty queue.” I think the former is better because it's more consistent with the behavior of the endpoint in other cases.
R0 -> SendMessage(m, queue) -> R1R1 -> SendMessage(m, queue) -> R1I0 -> SendMessage(m, queue) -> I1I1 -> SendMessage(m, queue) -> I1A0 -> SendMessage(m, queue) -> A1A1 -> SendMessage(m, queue) -> A1
For busy_behavior=interrupt, SendMessage(m, interrupt) supports:
W -> SendMessage(m, interrupt) -> R0E0 -> SendMessage(m, interrupt) -> R0E1 -> SendMessage(m, interrupt) -> R1R0 -> SendMessage(m, interrupt) -> I1R1 -> SendMessage(m, interrupt) -> I1I0 -> SendMessage(m, interrupt) -> I1I1 -> SendMessage(m, interrupt) -> I1A0 -> SendMessage(m, interrupt) -> R1A1 -> SendMessage(m, interrupt) -> R1
When SendMessage(m, interrupt) lands in I1, the queued message is promoted later by FinishInterruption(partial?) after the interrupted suffix is finalized.
Other input states are not supported.
This endpoint uses EditMessage(k, replacement):
W -> EditMessage(k, replacement) -> R0E0 -> EditMessage(k, replacement) -> R0E1 -> EditMessage(k, replacement) -> R0R0 -> EditMessage(k, replacement) -> R0R1 -> EditMessage(k, replacement) -> R0I0 -> EditMessage(k, replacement) -> R0I1 -> EditMessage(k, replacement) -> R0A0 -> EditMessage(k, replacement) -> R0A1 -> EditMessage(k, replacement) -> R0
EditMessage clears queued messages, cancels or obsoletes active work without preserving partial output, clears pending dynamic-tool action if present, marks the truncated active-history suffix as deleted, inserts the replacement turn, and lands in running.
Other input states are not supported.
This endpoint uses DeleteQueuedMessage(qid):
E1 -> DeleteQueuedMessage(qid) -> E0if removing the last queued messageE1 -> DeleteQueuedMessage(qid) -> E1if the queue remains non-emptyR1 -> DeleteQueuedMessage(qid) -> R0if removing the last queued messageR1 -> DeleteQueuedMessage(qid) -> R1if the queue remains non-emptyI1 -> DeleteQueuedMessage(qid) -> I0if removing the last queued messageI1 -> DeleteQueuedMessage(qid) -> I1if the queue remains non-emptyA1 -> DeleteQueuedMessage(qid) -> A0if removing the last queued messageA1 -> DeleteQueuedMessage(qid) -> A1if the queue remains non-empty
No other input states are supported.
This endpoint uses PromoteQueuedMessage(qid):
E1 -> PromoteQueuedMessage(qid) -> R0if promoting the last queued messageE1 -> PromoteQueuedMessage(qid) -> R1if the queue remains non-emptyR1 -> PromoteQueuedMessage(qid) -> I1I1 -> PromoteQueuedMessage(qid) -> I1A1 -> PromoteQueuedMessage(qid) -> R0if promoting the last queued messageA1 -> PromoteQueuedMessage(qid) -> R1if the queue remains non-empty
PromoteQueuedMessage reorders qid to the queue head internally when needed. From E1 and A1, it removes the queued message and inserts it into history immediately. From R1 and I1, it leaves the message queued at the head so FinishInterruption(partial?) can promote it after finalizing the interrupted suffix.
No other input states are supported.
This endpoint uses Interrupt(user_cancel):
R0 -> Interrupt(user_cancel) -> I0R1 -> Interrupt(user_cancel) -> I1A0 -> Interrupt(user_cancel) -> R0A1 -> Interrupt(user_cancel) -> R1
When Interrupt(user_cancel) lands in I0 or I1, the chat is later picked up by a ChatRunner to apply FinishInterruption(partial?).
No other input states are supported.
This endpoint uses CompleteRequiresAction(results):
A0 -> CompleteRequiresAction(results) -> R0A1 -> CompleteRequiresAction(results) -> R1
No other input states are supported.
This endpoint uses RequestCompaction:
W -> RequestCompaction -> R0E0 -> RequestCompaction -> R0E1 -> RequestCompaction -> R1
No other input states are supported: generating chats get a conflict error, and archived chats are rejected. Requesting compaction from an error state clears last_error, so a context-overflowed chat can recover by compacting instead of re-running the same oversized prompt. The endpoint is owner-only because the compaction runs LLM inference with the owner's delegated credentials. Inside the same transaction, after the transition succeeds, the endpoint verifies there is at least one uncompressed assistant message after the latest compaction boundary and rolls back with a "nothing to compact" conflict otherwise, so no LLM call is ever started for an empty or already-compacted chat. See Manual compaction for how the worker consumes the request.
This endpoint uses ClearContext:
W -> ClearContext -> WE0 -> ClearContext -> W
No other input states are supported: generating chats and chats with queued messages get a conflict error, and archived chats are rejected. Unlike /compact, there is no worker round-trip and no model call: the endpoint builds the boundary triplet itself and commits it synchronously inside the API transaction. The transcript is preserved; only future prompts stop seeing pre-clear history. Clearing from an error state clears last_error, so a context-overflowed chat gets an instant recovery path that discards the oversized history instead of summarizing it. The prompt-assembly query needs no changes because the clear boundary reuses the compressed model-only anchor shape produced by compaction. Boundary detection (latestContextBoundaryIndex) recognizes both chat_summarized and chat_cleared boundaries, so clear and compaction never reach across each other's boundary. If no active model-visible non-system message follows the latest boundary, the transaction rolls back with a "nothing to clear" conflict, so an empty or already-cleared chat never gains a duplicate boundary. The endpoint is owner-only for symmetry with /compact. The web UI surfaces it as the /clear slash command.
The chat worker and the stream loop need real-time notifications when the chat state changes to ensure they are responsive. To achieve this, we use pubsub.
As with the transitions section, I don't recommend reading the rest of this section thoroughly at first. Give it a cursory look, and treat it as a reference that you can return to later when you're analyzing the GET /api/experimental/chats/{chat}/stream endpoint or the chat worker.
There are 2 notification channels:
-
chat:ownershipis a global channel consumed by chat workers. Its payload is:chat_idsnapshot_versionIt notifies chat workers about chats that need processing by a chat worker, but aren't owned by a chat worker. A worker then picks the chat up.
-
chat:update:{chat_id}is a per-chat channel consumed by a chat worker that owns the chat and by active stream loops. Its payload is:snapshot_versionworker_idrunner_idhistory_versionqueue_versionretry_state_versiongeneration_attemptstatusarchived
It notifies receivers that a chat's execution state changed. Receivers use the payload as a hint to decide whether they should fetch the latest state from the database.
chat:update:{chat_id}is emitted after every successful transition bundle that advancessnapshot_version.- The
chat:update:{chat_id}payload contains the committed post-transition values for the notification fields. chat:ownershipis emitted when a transition leaves the chat in a runnable state, defined in Acquisition loop, and no worker owns it. That means eitherworker_idorrunner_idis NULL, or there is no fresh heartbeat row for the current(chat_id, runner_id).- Notifications are post-commit, best-effort, and versioned via the
snapshot_versionfield. - The current pubsub API is not assumed to provide transaction atomicity or commit-order delivery. Receivers must tolerate duplicates, drops, and reordering.
- Every receiver tracks the highest
snapshot_versionit has processed per chat. Notifications withsnapshot_versionless than or equal to that watermark are discarded.
A chat worker lives inside every coderd replica. It acquires chats, calls the LLM API, executes tools, handles interrupts and tool-result waits, and commits completed outcomes through the core state machine.
The chat worker is responsible for:
- acquiring chats when the chat is in a runnable state and no worker owns it, by listening to
chat:ownershipnotifications and doing periodical checks via database queries; - spawning a chat runner for each acquired chat: the chat runner is scoped to a single chat and is responsible for driving a chat forward by calling the LLM API and executing tools;
- upserting heartbeat rows in
chat_heartbeatsfor runners owned by the worker; - maintaining in-memory buffers of in-flight message parts for each chat;
- cleaning up runners when a chat is no longer owned by the runner, which it detects by inspecting
chat:update:{chat_id}notifications and database sync results.
A chat worker is identified by a worker ID, which is regenerated on worker startup.
A chat may use model configs only from its own organization. Chat creation, message sends, and message edits reject an explicit config that is disabled, unavailable to the caller, or belongs to another organization.
Queued-message promotion revalidates the stored model with daemon authorization. It keeps the model when the model and its provider are enabled and the model belongs to the chat's organization. Worker generation preparation revalidates with the chat owner's authorization context, so it also requires the model to remain readable by the owner. When those checks make the stored model unavailable, the path selects that organization's enabled default. If no local default is available, processing fails with ErrNoDefaultChatModelConfig; it never falls back to another organization's model.
The acquisition loop is a simple component that greedily acquires unowned or lease-expired chats from the database anytime it has a chance. It's driven by two triggers:
- a periodic timer that wakes up every second.
- a pubsub message on the
chat:ownershipchannel.
It finds suitable chats by fetching every chat that:
- is in a runnable execution state, meaning one of:
R0,R1,I0,I1,A0,A1; and - doesn't have an owner, meaning
worker_idorrunner_idis null, or there is nochat_heartbeatsrow for the current(chat_id, runner_id)newer than the lease expiry threshold of 5 minutes.
For every matching chat, it locks it, checks if the chat still meets the aforementioned conditions, and performs the Acquire(worker_id, runner_id) transition on it. The runner_id is a random UUID generated by the acquisition loop.
Both the 1-second timer interval and the 5-minute lease expiry threshold are defaults applied by chatd.New (DefaultPendingChatAcquireInterval and DefaultInFlightChatStaleAfter); coderd does not override them and no deployment flag exposes them.
When a chat is successfully acquired, the acquisition loop requests the Runner manager to spawn a chat runner for it.
The design doesn't attempt to distribute load between workers fairly. Whenever a chat needs an owner, all replicas race to acquire it. If there's a coder replica that has a lower latency to the database, it'll tend to acquire chats more frequently than other replicas.
The runner manager is responsible for the lifecycle of chat runners. For every chat runner that the acquisition loop requests to be spawned, it:
- spawns the runner as a goroutine;
- includes the chat in a periodic database sync operation;
- forwards chat state updates from the database sync to the chat runner;
The manager supports the existence of multiple runners for the same chat. This is possible when a runner abandons the chat, the acquisition loop on the same replica acquires it again, and the manager hasn't yet cleaned up the old runner.
The manager is comprised of 4 loops.
The main loop listens on 3 go channels:
- a channel for chat runner spawn requests;
- a channel for chat runner cleanup requests.
- a channel for chat runner cleanup completion notifications.
It processes one request at a time. When it receives a spawn request, it spawns a new runner as a goroutine. When it receives a cleanup request, it cancels the runner's goroutine, but it does not wait for it to finish and does not clean up the runner's resources synchronously. Instead, it spawns a goroutine that waits for the runner to finish and sends a cleanup completion notification when it does. The loop cleans up the resources of the runner that finished when it processes the cleanup completion notification.
Events sent on all the aforementioned channels have the following shape:
{
ChatID string,
RunnerID string,
}The database sync loop is responsible for fetching the current database state of all chats registered with the runner manager. On an interval, it runs a single query like this:
SELECT ... FROM chats WHERE id = ANY($1::uuid[]);where $1::uuid[] is the list of chat IDs registered with the runner manager. It then forwards the results to runners for the corresponding chats. There may be more than one runner for a given chat, each keyed by a different runner_id value, so the results must be forwarded to all of them.
Heartbeats are stored in a dedicated table:
CREATE UNLOGGED TABLE chat_heartbeats (
chat_id uuid NOT NULL REFERENCES chats(id) ON DELETE CASCADE,
runner_id uuid NOT NULL,
heartbeat_at timestamp with time zone NOT NULL,
PRIMARY KEY (chat_id, runner_id)
);
CREATE INDEX chat_heartbeats_heartbeat_at_idx
ON chat_heartbeats (heartbeat_at);For every runner registered with the runner manager, the heartbeat loop upserts the corresponding row in chat_heartbeats every 30 seconds (DefaultChatHeartbeatInterval, applied by chatd.New the same way as the acquisition defaults). Rows are keyed by (chat_id, runner_id). Against the 5-minute lease expiry threshold, a lease survives ten heartbeat intervals after the last successful heartbeat write; a worker whose heartbeat writes stop or fail for that long loses the lease, and the acquisition loop on any replica can acquire the chat under a new runner_id.
The loop uses this query:
INSERT INTO chat_heartbeats (chat_id, runner_id, heartbeat_at)
SELECT chat_id, runner_id, now()
FROM unnest($1::uuid[], $2::uuid[]) AS runners(chat_id, runner_id)
ON CONFLICT (chat_id, runner_id)
DO UPDATE SET heartbeat_at = EXCLUDED.heartbeat_at;Updating heartbeat rows does not advance snapshot_version and does not emit pubsub notifications.
The heartbeat cleanup loop runs every 30 seconds and removes heartbeat rows older than the lease expiry threshold used by the acquisition loop:
DELETE FROM chat_heartbeats
WHERE heartbeat_at < NOW() - (INTERVAL '1 second' * $1::int);Heartbeat rows are also removed automatically when their chat is deleted via the chat_heartbeats.chat_id foreign key.
The message part buffer is a global, scoped to a single replica, in-memory store of streaming message parts for each chat. Whenever an LLM API call returns a streaming message part, the runner synchronously adds it to the buffer. LLM responses are identified by episodes, which are tuples of (chat_id, history_version, generation_attempt).
The buffer maintains a mapping of episodes to arrays of message parts. Each array is capped at 1MB of content, calculated by serializing the parts to JSON and inspecting the size of the resulting strings. If the array is full, attempts to add parts to it return an error.
The buffer exposes the following API:
CreateEpisode(chat_id, history_version, generation_attempt): creates a new, empty message part array for an episode. May only be called once for a given episode, subsequent calls will return errors. It must be called before adding parts to the episode.CloseEpisode(chat_id, history_version, generation_attempt): closes an episode, preventing further parts from being added to it. May be called multiple times for a given episode, subsequent calls will be no-ops. Calling it on a non-existent episode creates the episode and closes it immediately. Concurrent parts of the system may race to create the episode and close it, so creating and closing in one operation prevents race conditions.AddPart(chat_id, history_version, generation_attempt, content): adds a message part to the buffer. Returns a predefined error if the episode is not found or the array is full.GetParts(chat_id, history_version, generation_attempt): returns the message parts for an episode. Returns a predefined error if the episode is not found.StartModelInvocation(chat_id, history_version, generation_attempt): stamps the instant the episode opens its provider stream. Returns a predefined error if the episode is not found or already closed. Episodes that never invoke a model, such as local tool execution batches, are never stamped.ModelInvokedAt(chat_id, history_version, generation_attempt): returns the instant stamped byStartModelInvocation, or the zero time when the episode is unknown or never opened a provider stream. It must be read beforeCloseEpisode, because closed episodes are garbage collected and reading afterwards races the cleanup loop. The interrupt goroutine reads it just before closing the episode and uses the span between that instant and the interrupt as the interrupted attempt's billable runtime.RecordToolStart(chat_id, history_version, generation_attempt, call_index, started_at)andRecordToolCompletion(chat_id, history_version, generation_attempt, call_index, completed_at): record when each call occurrence starts and finishes.ToolCompletions(chat_id, history_version, generation_attempt): returns when each tool call started and finished.SubscribeToEpisode(chat_id, history_version, generation_attempt): returns a go channel that will receive all message parts for the episode. It spawns a goroutine that delivers parts to the channel. It's live until the episode is closed or until a subscriber requests that the channel be closed. Once the goroutine delivers all message parts for a closed episode, it closes the channel and exits. If the episode is already closed at the time of the call, the goroutine delivers all message parts for the episode, closes the channel, and exits.SubscribeToEpisodedoes not return an error if the episode is not found: it waits for it to be created instead.
Closed episodes are garbage collected after at least 15 seconds since they were closed and when they have no active subscribers. The message part buffer maintains a garbage collection goroutine.
Subscribers must accept parts within 10 seconds of them being sent on the channel. If a subscriber does not accept a part within that timeframe, the subscription channel is closed.
A chat runner is responsible for driving a chat forward by calling the LLM API and executing tools. It is scoped to a single chat and is responsible for committing results through the core state machine. It is also responsible for handling interrupts.
The runner is implemented as a single event loop that listens on a go channel with state updates. The loop does not perform any side effects nor does it query the database by itself. Instead, it spawns goroutines to communicate with the outside world.
State updates processed by the loop come from:
chat:update:{chat_id}pubsub notifications;- database sync results forwarded by the runner manager;
- goroutines notifying the loop after applying core state machine transitions (fast path to avoid waiting for pubsub notifications); and
- right after being spawned, from a single database call the runner makes to fetch the initial state of the chat.
The runner is responsible for subscribing to the chat:update:{chat_id} pubsub channel. During bootstrap, it must first subscribe to the channel and then fetch the initial state of the chat from the database to avoid missing any updates.
Every event that the runner loop processes has the following shape:
{
WorkerID *string,
RunnerID *string,
SnapshotVersion int64,
HistoryVersion int64,
GenerationAttempt int64,
Archived bool,
Status string,
}The runner maintains the following local state:
- latest processed event's
SnapshotVersion; - latest processed event's
HistoryVersion; - latest processed event's
GenerationAttempt; - latest processed event's
Status; - its own
WorkerIDandRunnerID. - a list of goroutines it has spawned to perform side effects, each identified by a unique ID, together with cancellation handles and go channels that the goroutines use to notify they have finished.
- the ID of the currently active goroutine, if there is one.
The main idea behind the event processing logic is that a chat's status and its history version determine the work that the runner should be performing at any given time. Let's go through an example:
- A chat's status is
running, and history version is42. The runner is calling the LLM API or executing tools using the message history identified by42. - The runner sees an event with chat status still
running, but history version changed to46. The runner should still be calling the LLM API, but it should be using the message history identified by46. So if the runner sees that it has an active goroutine doing work onhistory_version=42, it should cancel that goroutine, and spawn a new one to do work onhistory_version=46. - Then if the history version is still
46, but the status changed tointerrupting, the runner should cancel the active goroutine, and spawn a new one to handle the interrupt on the core state machine level - that is, submit theFinishInterruptiontransition.
The runner processes one event at a time. We call processing an event a loop iteration. For each event, it does the following things in order:
- If the event's
SnapshotVersionis less than or equal to the runner's latest processed event'sSnapshotVersion, ignore the event and stop. - If the event's
WorkerIDis not the runner'sWorkerID, or the event'sRunnerIDis not the runner'sRunnerID, send a cleanup request to the runner manager, and stop. - If the event's
HistoryVersionis equal to the runner's latest processed event'sHistoryVersion, and if the event'sStatusis equal to the runner's latest processed event'sStatus, stop. - If the event's
HistoryVersionis greater than the runner's latest processed event'sHistoryVersion, or if the event'sStatusis different from the runner's latest processed event'sStatus, cancel the currently running goroutine if there is one, remove its ID from the active goroutine ID, and continue to the next step. Do not wait for the goroutine to finish. - Iterate through the goroutine list and remove all goroutines that have finished. This will likely not include the goroutine that may have just been cancelled: that's okay, a future loop iteration will clean it up.
- If the event's
Archivedistrue(core state machine is inXW,XE0, orXE1), spawn a goroutine to abandon the chat, mark it as active, and stop. We call this the abandon chat goroutine. - If the event's
Statusisrunning(R0orR1), spawn a goroutine to call the LLM API and execute tools, and mark it as active. We call this the generation goroutine. - If the event's
Statusisinterrupting(I0orI1), spawn a goroutine to handle the interrupt, and mark it as active. We call this the interrupt goroutine. - If the event's
Statusisrequires_action(A0orA1), spawn a goroutine to wait for the dynamic tool timeout to pass, and mark it as active. We call this the dynamic tools timeout goroutine. - Otherwise, spawn an abandon chat goroutine and mark it as active.
After stopping, the iteration updates the runner's local state to the event's values, unless it ignored the event because of a stale SnapshotVersion.
The runner spawns goroutines to perform side effects. There are 4 kinds: generation, interrupt, dynamic tools timeout, and abandon chat.
Goroutines perform core state machine transitions. The goroutines are spawned knowing their intended history version, chat status, and runner ID. Whenever they access the database, either for reading or writing, they must lock the chat and ensure that values in the database match the intended values. If they do not, they must exit to prevent performing stale work and applying stale transitions.
Locks are paramount: goroutines must not mix database reads from states with differing history versions, statuses, or runner IDs. However, locks must not be held for extended periods of time, such as during calling the LLM API. They must be obtained only for database operations.
In addition to database locks, each goroutine must also obtain a local, in-memory lock scoped to its intended history_version and status. In case 2 goroutines with the same intended values are ever active at the same time, this lock prevents them from racing with each other.
Goroutines described in this section may only exit in controlled ways. If they encounter an unexpected error, they must retry the operation they were meant to perform. Retries are not limited, but are governed by bounded exponential backoff.
Expected exit conditions include, but are not limited to:
- successful completion of the operation the goroutine was meant to perform;
- stale fence failure;
- context cancellation;
- chat deleted.
Retriable conditions include, but are not limited to:
- database connection error;
- LLM API request error, with the exception of hitting the generation attempt limit, which is considered to be a successful completion of the operation the goroutine was meant to perform.
The generation goroutine is responsible for calling the LLM API and executing tools. It is spawned when the event indicates the core state machine is in R0 or R1 (status is running).
It inspects the chat's message history, and decides what's the next step to take. The result of that step is the application of one of the following core state machine transitions:
CommitStep: applied when an LLM API call returns a response.FinishTurn: applied when the chat processing logic determines that there's no more work to do for the current message history (no pending tool calls, user message is not the last message in the history, etc.).FinishError: applied when the LLM API call fails and the retry limit is reached, determined by thegeneration_attemptvalue.EnterRequiresAction: applied when there are pending dynamic tool calls.
The generation goroutine also applies the RecordGenerationAttempt transition every time before calling the LLM API. It may apply this transition multiple times in case of retries. When an LLM API call fails with a retryable error and the goroutine will retry after a backoff, it applies RecordRetryState(payload) with the retry payload that should be sent to clients.
When receiving streaming message parts from the LLM API, the generation goroutine adds them to the Message part buffer in real time. Whenever it starts a new generation attempt, it must start a new episode in the buffer, and mark it as closed when the attempt is finished; either because the LLM API call returned a response, or the attempt was cancelled. If AddPart returns an error, the goroutine ignores it. Storing parts in the buffer is best-effort: if the buffer is full, or the episode is closed, the parts are dropped. A stale generation goroutine may keep on adding parts to the buffer until it is cancelled or exits.
Since the runner doesn't wait for goroutines to finish when it cancels them, and spawns new goroutines to perform new work immediately, the runner does not guarantee that any interrupted tool calls are fully stopped before continuing. Tool call interrupts are best-effort.
Tool calls have at least once semantics: if the goroutine executes a tool call, and the replica crashes before the result is persisted, another replica will execute the tool call again later. Future work may include adding a mechanism to ensure at most once semantics.
Parallel tool call results must be inserted in bulk after all parallel tool calls finish in a single CommitStep transition so that the generation goroutine only increments history_version once, since a change to the history_version interrupts the gorotuine. This is consistent with the existing chatd implementation.
The generation goroutine supports:
- chat compaction (automatic and manual, see Manual compaction)
- MCP tools
- subagents (
spawn_agent,wait_agent,message_agent,interrupt_agent,list_agents,list_subagent_models)close_agentis a deprecated alias that dispatches tointerrupt_agent, so historical tool calls in chat history still resolve
- file links
- workspace binding
- plan mode
- respecting model configuration
- provider-specific tools like web search and computer use
- turn limit after a user message (the LLM shouldn't be able to spin forever in loop)
- and other things
Model configs may carry a reasoning_effort config ({default, max}) inside chat_model_configs.options. Users select a per-turn effort when sending or editing a message; the value is stored on chat_messages.reasoning_effort and on chat_queued_messages.reasoning_effort for queued messages. Queued messages carry the value through promotion, and chats.last_reasoning_effort tracks the most recent message that set one, mirroring last_model_config_id.
Subagent spawning is a second source of both values. Organization admin overrides are stored in chat_organization_model_overrides, keyed by organization and subagent context. Personal overrides are stored in chat_user_model_overrides, keyed by user, organization, and context. Model references are typed UUIDs constrained to configs in the same organization.
Subagent model and effort resolution follows this precedence:
- Explicit
spawn_agentarguments. Optionalmodel_config_idandreasoning_effortargs are discoverable vialist_subagent_models. An explicit model becomes the child chat'slast_model_config_id, and an explicit effort is stored on the child's initial message. Each explicit value wins over personal overrides, organization admin overrides, and parent inheritance. The values are validated at spawn time for an enabled config and provider, matching organization, usable credentials, and effort on the global scale. Invalid values produce a tool error before child creation.computer_usespawns reject both arguments because their model routing is specialized. - Personal member override. When personal overrides are enabled, Chatd reads the row for the chat owner, organization, and subagent context. Mode
chat_defaultstops override resolution and preserves the subagent type's default or inheritance behavior. Modemodelselects the row's model and optional effort; if that selection becomes unusable, resolution falls through softly to the organization admin override. Modedeployment_defaultretains its legacy name but also defers to the organization admin override. - Organization admin override. Chatd reads the row for the chat's organization and the
generalorexplorecontext. An unset or unusable row falls through softly to the subagent type's default. - Subagent type default. Both
generalandexploresubagents inherit the parent chat's current model.
During generation preparation, the effective effort is resolved as the chat's last_reasoning_effort if set, else the config's default; clamped to the config's max on the global scale none < minimal < low < medium < high < xhigh < max; and passed through to the provider. The provider verifies whether the configured value is valid for that model at runtime. If the model config has no reasoning_effort, any user-selected value is ignored. The resolved value is injected into the provider-native options by chatprovider.ProviderOptionsForCall, which converts the model config and applies the effort in one step. For Anthropic, the fantasy provider converts effort into enabled budget thinking on models older than Claude 4.6, which reject adaptive thinking.
applyReasoningEffort clamps none and minimal to low for GPT-6 Astra and its dated snapshots (chatopenai.IsGPT6Astra, a case-insensitive prefix match on gpt-6-astra), because that model rejects none with HTTP 400 and lists no minimal effort.
OpenAI models speak either the Responses API or Chat Completions. The provider SDK picks per model from a static known-model list, so a newly released model absent from that list falls back to Chat Completions. Model configs may override the choice with openai_config.use_responses_api inside chat_model_configs.options: unset keeps the known-model list, true forces Responses, false forces Chat Completions. It sits in openai_config rather than provider_options.openai because it is applied once when the client is built, while provider_options holds per-request parameters.
chatopenai.UsesResponsesAPI owns the unset-override decision for both the client (WithResponsesAPIFunc) and TransportFor. When the override is nil it consults the SDK's known-model list, except that GPT-6 Astra defaults to Responses because its function calling is Responses-only.
The transport is resolved exactly once, when the client is built, and carried on chatprovider.Model as a chatopenai.Transport. Model wraps the fantasy client with that resolved fact; its fields are unexported and only its constructor sets the transport, deriving it from the client, so no caller can pick a transport that disagrees with the client. TransportInvalid is the zero value and panics when read rather than defaulting to a wire format. A nil client yields that invalid zero value, which the construction path reports as an error.
Request preparation reads the transport from the model instead of recomputing it. Three places depend on it, and each fails silently when it disagrees with the client:
- Provider option conversion chooses between the Responses and Chat Completions option structs. The SDK type-asserts the concrete struct, so a mismatch discards every OpenAI provider option rather than failing.
- Reasoning effort injection creates those option structs when a config has no OpenAI options of its own.
- File part conversion (
Model.AcceptsFilePartMediaType) gates attachments, because the Responses API natively accepts only images and PDFs. A mismatch here drops text attachments.
The first two happen together in chatprovider.ProviderOptionsForCall, the only entry point in chatprovider that builds provider options for a call; it delegates transport-aware OpenAI conversion to chatopenai.ProviderOptionsFromChatConfig. Config conversion and effort injection cannot pick different option types because one function owns both.
Debug recording replaces the wrapped client and preserves the resolved transport. Computer-use turns substitute a hardcoded default model that has no config of its own; it carries its own transport, so the chat model's openai_config does not follow it.
Azure is deliberately exempt: its provider always enables the Responses API for known models and exposes no equivalent per-model hook, so the transport keeps following the known-model list for Azure. Ignoring the override there is what keeps the decisions above in agreement with the Azure client. The exemption is narrower than it appears, because chatd never builds an azure-typed provider as a fantasy azure client: fantasyConfigForAIBridge folds every provider type other than anthropic, bedrock, and openai into openai-compat, which always speaks Chat Completions. Bedrock is the exception within that set: its fantasy client depends on the model ID, so anthropic.* bedrock models fold to the Anthropic Messages client while non-anthropic bedrock models fold to the OpenAI Responses client.
Both transports read the same provider_options.openai config, but not every field applies to both wire formats. The table below records, per field, which transport honors it; TestProviderOptionsTransportParity fails when a field is honored on one transport and silently ignored on the other without being recorded there as intentional.
provider_options.openai field |
Responses | Chat Completions |
|---|---|---|
include |
yes | no |
instructions |
yes | no |
logit_bias |
no | yes |
log_probs |
yes | yes |
top_log_probs |
yes | yes |
max_tool_calls |
yes | no |
parallel_tool_calls |
yes | yes |
user |
yes | yes |
reasoning_summary |
yes | no |
max_completion_tokens |
no | yes |
text_verbosity |
yes | yes |
prediction |
no | yes |
store |
yes | yes |
metadata |
yes | yes |
prompt_cache_key |
yes | yes |
safety_identifier |
yes | yes |
service_tier |
yes | yes |
structured_outputs |
no | yes |
strict_json_schema |
yes | no |
web_search_enabled |
no | no |
search_context_size |
no | no |
allowed_domains |
no | no |
Three asymmetries are deliberate near-equivalents rather than gaps. max_completion_tokens is the Chat Completions cap; on Responses the transport-neutral max_output_tokens config bounds output instead. structured_outputs and strict_json_schema are the per-API strictness switches, each honored only by its own API. On Responses, top_log_probs wins over log_probs because that API takes a single logprobs value. The trailing web search fields configure tool wiring rather than per-request provider options, so neither transport reads them during option conversion.
The model editor scopes the field to openai-typed providers with a providers struct tag, which the option schema generator emits as visible_for_providers. Gating on the raw provider type rather than the alias table keeps the control out of editors for provider types that cannot honor it.
Compaction is an auxiliary LLM call: when the conversation approaches the context limit, the generation goroutine asks a model to summarize the history, commits the summary as a compressed boundary, and continues the turn on the chat model.
By default the summary is generated with the chat model. Organization admins can select a dedicated compaction model via PUT /api/v2/organizations/{organization}/chats/model-overrides/compaction. The selection is stored as a typed chat_organization_model_overrides row and resolved using the chat's organization. Its composite foreign key binds the model config UUID to that organization, so cross-organization and malformed string references cannot be stored. The override affects only the summary call; compressed-message storage and the post-compaction assistant generation keep using the chat model.
Details that follow from the override:
- Context limits: the compaction trigger uses the stricter of the chat model's and the compaction model's context limits, because the history must also fit the summarizer's window. The post-compaction "still over limit" check uses that same stricter limit; otherwise a smaller compaction-model window could trigger repeated compactions instead of a terminal error.
- Failure semantics: an unset override uses the chat model. A stored config that later becomes deleted or disabled, whose provider becomes disabled, or whose required credentials become unavailable is logged and falls back to the chat model during generation preparation. Failure to read the override row or load provider credentials stops preparation. A failure while resolving the referenced model config or provider is logged and falls back to the chat model. A usable override that fails at use (route or client construction, provider call failure) fails the generation visibly through the normal error path; there is no silent fallback. The override model client is constructed inside the compact generation action, not at prepare time, so a broken override cannot fail turns that finish without compacting (including turns over the threshold whose last assistant step already completed).
- Prompt safety: the prompt is built and sanitized for the chat model, so when the override points at a different provider the compaction copy of the prompt is re-sanitized: provider-executed tool history is flattened into plain text parts (keeping its content while dropping the provider-specific wire shape), file parts the compaction model rejects are replaced with text placeholders, and Anthropic provider-tool sanitization is re-run for the compaction provider. The assistant generation prompt is never mutated.
- Observability: compaction metrics and chat debug runs record the provider and model that actually generated the summary. This includes the "still over limit" terminal error, which is recorded before the override client is built: prepare-time resolution keeps the override's provider/model identity so that error lands on the same metric series as the compact action's own events.
The interrupt goroutine is responsible for handling interrupts. It is spawned when the event indicates the core state machine is in I0 or I1 (status is interrupting).
The goroutine does the following in order:
- It fetches the generation attempt number from the database.
- It closes the episode corresponding to its history version and generation attempt by calling the
CloseEpisodemethod on the Message part buffer. - It reads the buffered parts for that episode by calling the
GetPartsmethod on the message part buffer. - It applies the
FinishInterruption(partial?)transition on the core state machine. If there are no buffered parts for that episode, or the episode is not found, it passesnilas thepartialargument.
The dynamic tools timeout goroutine is responsible for waiting for the dynamic tool timeout to pass, which is determined by the requires_action_deadline_at field on the chat. It is spawned when the event indicates the core state machine is in A0 or A1 (status is requires_action). The goroutine fetches the deadline value from the database. When the timeout passes, it applies the CancelRequiresAction transition on the core state machine.
The abandon chat goroutine is responsible for abandoning the chat. It is spawned whenever the event processing logic determines that the chat no longer needs to be owned by the runner. It applies the Abandon transition on the core state machine after checking that the chat is still owned by the runner.
When the manager cleans up a runner, the runner must cancel all goroutines it has spawned and unsubscribe from pubsub.
By default, chatd runs up to five top-level chats and ten subagent chats at once. Each limit applies across the entire deployment. Enterprise deployments can remove these limits when their plan permits it. Extra chats wait for capacity, but users can still interrupt active chats.
The worker periodically archives old, unused chats.
Compaction reduces the LLM prompt size by summarizing older history into a compressed boundary. It normally runs automatically: while preparing a generation, the worker compares the latest known token usage against the model's compaction threshold, and when the threshold is exceeded it makes a non-streaming LLM call to produce a summary and commits it as a compressed message triplet (a hidden model-only summary boundary, a visible chat_summarized tool call, and its tool result). Prompt queries prune history at the newest boundary.
Trailing user messages the assistant has not answered yet are not summarized: they are excluded from the summarizer's input and re-committed after the triplet as model-only user rows, so the pruned prompt keeps them verbatim instead of relying on summary fidelity.
Users can also request a compaction on demand via POST /api/experimental/chats/{chat}/compact (surfaced in the web UI as the /compact slash command). Manual compaction is a durable one-shot request executed through the normal worker loop rather than synchronously in the HTTP handler. This reuses the worker's lock fencing, retry accounting, streamed "Summarizing..." progress parts, metrics, and debug runs, and it survives replica crashes. The flow:
- The endpoint applies the
RequestCompactiontransition: allowed fromW,E0, andE1, it setschats.compaction_requested_at = now(), clearslast_error, lands inR0(orR1fromE1, preserving the queue) without inserting any message, and publishes a status-change pubsub event to wake workers. Because the transition inserts no history, it advanceshistory_versionto the transaction's newsnapshot_versionand resetsgeneration_attemptitself, granting the fresh retry budget and episode keys a history change would otherwise provide. A timestamp is used instead of a boolean for debuggability. AI Gateway attribution needs no per-request key: generation preparation resolves the owner's synthetic API key like any other turn. - The generation goroutine's decision logic checks
compaction_requested_atafter the unresolved local/dynamic tool guards but before the history-completeness check (an idle chat's history is otherwise complete, which would end the turn). If the marker is set and at least one uncompressed assistant message exists after the latest compaction boundary, it selects a forced compaction; if there is nothing to compact, the marker is ignored and the turn finishes normally, clearing it. - A forced compaction bypasses the automatic threshold gates (usage below threshold, unknown context window, and the threshold=100 disable) and stamps
source: "manual"instead ofsource: "automatic"into thechat_summarizedtool call arguments, tool result JSON, and streamed parts so clients can render manual compactions distinctly. - The compaction
CommitStepconsumes the request by clearingcompaction_requested_atin the same transaction that commits the summary triplet. The next decision pass finds the history complete and finishes the turn, so a chat with an empty queue returns towaitingwith no assistant follow-up; a chat compacted fromE1proceeds to its queued messages instead. Apost_compacthook effect is the one exception: because the decision reads user-visible history, an effect that commits a user-visible message leaves the history incomplete and the turn continues with an assistant response. A model-only effect such asmodel_contextreaches the model without resuming generation.
The compaction_requested_at marker is one-shot: transitions that keep an active turn alive (Acquire, Abandon, SetArchived, queueing a message on a busy chat) carry it forward, while every other transition that rewrites the execution state (FinishTurn, FinishError, Interrupt, EditMessage, PromoteQueuedMessage, CancelRequiresAction, ReconcileInvalidState, and so on) clears it by construction, so a stale request can never replay on a later turn.
When the agent-lifecycle-hooks experiment is enabled and a hook URL is configured, chatd sends events to an external consumer at key points in a conversation: session start, prompt submission, tool use, compaction, and turn completion. The event types are session_start, user_prompt_submit, pre_tool_use, post_tool_use, pre_compact, post_compact, and stop.
The consumer can observe activity, add model-only or user-visible context, replace supported prompt or tool input, and deny prompts or tool calls. Only user_prompt_submit and pre_tool_use accept a permission decision or input override; a response carrying one on any other event is rejected as an invalid response. Prompt submission is evaluated once when the submission is accepted, including queued messages and subagent prompts. Returned context becomes part of the conversation for its intended audience, except that context returned before a compaction guides the compaction summary instead.
Lifecycle hooks fail closed. If the consumer cannot be reached or returns an invalid response, Coder stops the triggering operation rather than continuing without the consumer's decision. Affected chats can enter an error state until the consumer recovers or hooks are disabled.
Concurrent dispatches are capped per replica, and each dispatch declares whether it admits new work into a chat or belongs to work a chat already admitted. Admission can hold only part of the cap, so a burst of new submissions cannot consume the capacity that already-admitted work depends on. The caller declares this, because the event type does not determine it: a subagent spawn submits a prompt from inside a running turn, and editing a message starts a session at admission time.
Coder stores no hook-specific dispatch or decision state. Delivery is best-effort and can duplicate, and a failed dispatch is never redelivered, so the consumer owns durable policy state, audit records, and deduplication based on stable event identifiers.
The stream loop powers the GET /api/experimental/chats/{chat}/stream endpoint. It is scoped to one chat and one client WebSocket. It's responsible for delivering a stream of chat updates to the client, including:
- messages committed to the database; and
- streaming message parts emitted by the chat worker via the relay mechanism.
The following chat stream events, delivered to the client over WebSocket, are supported:
message_part: a streaming message part emitted by the chat worker. Each carries thehistory_versionandgeneration_attemptof the episode it belongs to, so a client knows which episode a message part comes from.message: a committed chat message present in the database.status: the chat's status.error: the chat's persisted error payload.queue_update: the full current queued-message list.action_required: a dynamic tool call was issued by the chat worker, the client must execute it and submit the result.retry: emitted when the chat worker is waiting before retrying a failed generation attempt.preview_reset: a reset of the stream's preview state (message parts), emitted when the worker's LLM call fails mid-way for whatever reason.history_reset: a reset of the stream's history state (committed messages), emitted when the message history is edited and some messages are removed from the history.
When a client connects, the endpoint:
- Subscribes to
chat:update:{chat_id}pubsub channel and starts buffering notifications. - Registers with the sync poller so the chat is included in the replica's periodic database sync.
- Initializes the stream loop to the null local state.
- Starts the stream loop. Its first action is an initial database fetch whose result is applied to establish the baseline state. The loop then handles:
- triggering Sync operations from pubsub notifications and the sync poller;
- instructing the Relay mechanism which streaming parts to forward;
- emitting client events;
- updating local state;
When the endpoint exits, it deregisters from the sync poller, unsubscribes from pubsub, stops the stream loop, stops the relay forwarder, and closes request resources.
The stream loop stores:
- latest synchronized
snapshot_version; - latest synchronized
history_version; - latest synchronized
queue_version; - latest synchronized
retry_state_version; - known committed messages:
- message ID;
- latest message revision sent to the client;
- latest status sent to the client;
history_versionfor the last sent error;- latest
history_versionfor whichaction_requiredwas sent; - latest
worker_id; - latest
generation_attempt; - last accepted preview part
seq;
Initial null state:
- synchronized versions are
0; - known committed-message revision map is empty;
worker_idis null;generation_attemptis0;- status is unset;
- last sent error history version is
0; - action-required cursor is
0; - preview part sequence is
0;
The loop has two operations:
| Operation | Description |
|---|---|
Sync(hints) |
Maybe fetch database state. If newer state is observed, emit required client events, update local cursors, and configure the relay target. Triggered by pubsub notifications and the sync poller. |
Part(history_version, generation_attempt, seq, content) |
Emit one live preview part. The operation succeeds only if the part matches local watermarks (history version, generation attempt, and seq). Triggered by the relay forwarder. |
The loop processes one operation at a time. It must not process another input halfway through a Sync or Part.
Sync input fields include:
snapshot_versionoptional int64;history_versionoptional int64;queue_versionoptional int64;retry_state_versionoptional int64;statusoptional string;worker_idoptional string;generation_attemptoptional int64.
Sync uses its hints to decide whether to fetch from the database:
- If
snapshot_versionis present andsnapshot_version <= local.snapshot_version, return no-op. - Compare each present input field to local state:
history_version > local.history_version;queue_version > local.queue_version;retry_state_version > local.retry_state_version;status != local.status;worker_id != local.worker_id;generation_attempt != local.generation_attempt.
- If none of the above conditions indicate that local state may be stale, return no-op.
- Otherwise fetch from the database.
After fetching, if db.snapshot_version <= local.snapshot_version, return no-op. Otherwise apply the database result.
The endpoint's initial bootstrap fetch skips the hint check in steps 1 to 3, fetches unconditionally, and applies the database result the same way. Because it runs against the null local state, every database field is newer and the full state is emitted.
Applying the database result means, in deterministic order:
- If
db.history_version > local.history_version, run message synchronization. - If
db.queue_version > local.queue_version, run queue synchronization. - If
db.status != local.status, run status synchronization. - If
db.status = erroranddb.history_version > local.error_history_version, run error synchronization. - If
db.status = requires_actionanddb.history_version > local.action_required_history_version, run action-required synchronization. - If
db.retry_state_version > local.retry_state_version, run retry-state synchronization. - If
db.history_version != local.history_versionor (db.generation_attempt != local.generation_attemptanddb.generation_attempt != 0), set last accepted preview partseq = 0and emitpreview_reset. (generation_attempt = 0is the initial value for thegeneration_attemptfield after the message history changes, and there are never any preview parts associated with it. Episodes created by the Generation goroutine always have a generation attempt number greater than 0.) - If
db.generation_attempt > 0, configure the relay forwarder withdb.worker_id,db.history_version, anddb.generation_attempt. - Save
snapshot_version,history_version,queue_version,retry_state_version,status,worker_id, andgeneration_attemptfrom the db to local state.
If Sync fetches from the database, all database reads for that Sync must happen in the same read transaction. This includes reading the chat row, changed messages, full-history refresh messages, queued messages, retry state, error data, and pending dynamic tool-call data.
A pubsub notification is converted into:
Sync(
snapshot_version?,
history_version?,
queue_version?,
retry_state_version?,
status?,
worker_id?,
generation_attempt?,
)
Each field is optional. Missing fields are ignored. worker_id uses a tri-state representation:
- absent means the check is ignored;
- present with a null worker means compare against a null worker;
- present with a worker ID means compare against that worker ID.
Notification fields do not directly cause client events. They only help decide whether Sync should fetch from the database. Only the database result decides what to emit.
The sync poller, a helper component described in full in the Sync poller section below, periodically reads the database state for every registered chat and delivers a hint-based Sync to each subscriber:
Sync(
snapshot_version,
history_version,
queue_version,
retry_state_version,
status,
worker_id,
generation_attempt,
)
Message sync happens inside Sync.
Flow:
Syncobservesdb.history_version > local.history_version.- Fetch rows from
chat_messageswhererevision > local.history_version. - Inspect the fetched rows.
If no fetched rows are soft-deleted:
- Emit
messageevents for rows whose revision is newer than the local known-message revision. - Update the known-message revision map.
If any fetched row is soft-deleted, mirror the current stream endpoint's full-refresh behavior with the addition of emitting the history_reset event:
- Fetch all current non-deleted messages from the beginning in client-visible order.
- Emit the
history_resetevent. - Resend all messages as
messageevents. - Replace the known-message revision map with the revisions from the resent messages.
Required invariant:
- every client-visible message-history change advances the changed row revision and chat
history_version.
Queue sync happens inside Sync.
Flow:
Syncobservesdb.queue_version > local.queue_version.- Fetch the full current queue in client-visible order.
- Emit one
queue_updateevent with the full queue.
Required invariant:
- every client-visible queue insert, update, reorder, or delete advances
queue_version.
Retry-state sync happens inside Sync.
Flow:
Syncobservesdb.retry_state_version > local.retry_state_version.- If
db.retry_stateis null, emit nothing. - If
db.retry_stateis non-null, emit oneretryevent withdb.retry_stateas the payload.
Required invariant:
- every client-visible retry-state change advances
retry_state_version.
Status sync happens inside Sync.
Flow:
- Compare database status to local status.
- If they differ, emit
status.
Error sync happens inside Sync.
Flow:
- If database status is not
error, stop. - If database status is
erroranddb.history_version > local.error_history_version, emiterror. - Set
local.error_history_version = db.history_version.
Action-required sync happens inside Sync.
Flow:
- If database status is not
requires_action, stop. - If database status is
requires_actionanddb.history_version > local.action_required_history_version, emitaction_required. - Set
local.action_required_history_version = db.history_version.
The sync poller is a replica-global helper, scoped to a single replica and shared by every stream loop running on it. Every 10 seconds, it fetches the current database snapshots for all registered chats and delivers them to the stream loops.
Stream loops register and deregister through a mutex-backed map keyed by chat_id. Each map entry holds the set of subscribers for that chat, because one replica may serve multiple stream loops for the same chat. Each subscriber carries a handle the helper uses to deliver Sync operations into that stream loop's input.
Every 10 seconds, the helper:
- Acquires the mutex, snapshots the registered chat IDs and their subscribers, and releases the mutex.
- Reads the current version columns for those chats in a single query.
- Delivers a hint-based
Syncto every subscriber of each returned chat, using the row's version columns as the hints.
The query is:
SELECT id, snapshot_version, history_version, queue_version,
retry_state_version, generation_attempt, status, worker_id
FROM chats
WHERE id = ANY($1::uuid[]);We make use of a relay mechanism when there are multiple coderd replicas. If a client connects to the stream endpoint on replica A, but the chat worker that owns the chat is on replica B, the endpoint will connect to replica B and relay streaming message parts.
There exists a GET /api/v2/chats/{chat}/stream/parts endpoint that is responsible exclusively for streaming message parts. That endpoint talks to the chat worker on the same replica to obtain the message parts and relay them to the client.
The flow is:
- Client connects to the
GET /api/v2/chats/{chat}/streamendpoint. - The endpoint checks the database to see which replica owns the chat and resolves the replica's address.
- The endpoint connects to the
GET /api/v2/chats/{chat}/stream/partsendpoint on that replica. - The stream endpoint relays both the full chat state and the streaming message parts to the client.
Some edge cases:
- In the case where the chat worker is on the same replica as the stream endpoint, the endpoint obtains message parts directly from memory over go channels; it doesn't dial itself over WebSocket. From the perspective of the relay mechanism, this doesn't matter. The transport method is abstracted away.
- When the chat's ownership changes, the stream endpoint detects it and reconnects to the new replica.
Each streaming message part belongs to an episode, identified by history_version and generation_attempt. Within an episode, each part has a seq number. The first part has seq=1, and each later part increments seq by 1.
Sync configures the relay forwarder with:
worker_id;history_version;generation_attempt.
If worker_id is null, the relay forwarder stops forwarding messages and, if connected to a parts endpoint, closes the connection.
If worker_id changes, the relay forwarder connects to the new worker's parts endpoint. If only history_version or generation_attempt changes while worker_id stays the same, the relay forwarder keeps the existing WebSocket connection and sends a control message over that connection to select the new episode. It should not tear down and recreate the connection just because the requested episode changed.
The forwarder must only pass parts for the currently requested episode to the stream loop.
The parts endpoint is a WebSocket endpoint.
Connection setup:
- the URL identifies the chat ID, for example
GET /api/v2/chats/{chat}/stream/parts; - the endpoint accepts the connection regardless of whether the local replica owns the chat;
- after connecting, the client sends control messages over the WebSocket to choose which episode it wants.
Control message shape:
{
"history_version": 12,
"generation_attempt": 3
}Behavior after receiving an episode-selection message:
- stop sending parts for the previously selected episode;
- start serving only parts for the requested
history_versionandgeneration_attempt; - send buffered parts for that episode if any exist;
- stream future parts for that episode as they are produced;
- if the worker has not produced parts for that episode yet, keep the connection open and emit nothing until matching parts become available;
- missing episode state is not an error;
- enforce contiguous
seqdelivery for the selected episode.
Only worker changes require a new parts WebSocket connection.
The endpoint uses the Message part buffer to fetch parts for each episode.
The relay forwarder sends Part operations to the stream loop:
Part(history_version, generation_attempt, seq, content)
The operation succeeds only if:
history_version == local.history_version
generation_attempt == local.generation_attempt
seq == local.last_part_seq + 1
If accepted:
emit message_part
local.last_part_seq = seq
If any check fails, the operation is rejected.
A sequence gap is an invariant violation because the parts endpoint must enforce contiguous delivery for each requested episode.
The loop returns when the endpoint returns.
Expected termination causes:
- client disconnect;
- request context cancellation;
- WebSocket write failure;
- database failure after bounded retry;
- internal invariant violation.
The loop should not terminate just because:
- a stale sync hint arrives;
- a duplicate sync hint arrives;
- pubsub drops a notification;
- the relay has no parts yet for the requested episode;
- the relay connection reconnects.