Thanks to visit codestin.com
Credit goes to github.com

Skip to content

feat(agents): thread-aware reads, in-thread posts, and the threading cue (TASK-052) - #1176

Merged
lilyshen0722 merged 4 commits into
mainfrom
feat-agent-thread-reads
Aug 23, 2026
Merged

feat(agents): thread-aware reads, in-thread posts, and the threading cue (TASK-052)#1176
lilyshen0722 merged 4 commits into
mainfrom
feat-agent-thread-reads

Conversation

@lilyshen0722

Copy link
Copy Markdown
Contributor

Completes TASK-052: the read gap that held the cue (sprint-review's degrades-open finding), the write half (in-thread without pinging, resolver-validated, fail-loud on both Mongo-fallback edges), and the cue itself — three verbs + quote-thread independence + Sam's overflow rule, shipped only now that every promise in it is kernel-enforced. 10 executing tests, 698/698 across the neighboring suites, ThreadRootError propagation pinned in the failure direction. Fleet is halted on Sam's order; ships on operator self-review with the evidence above.

🤖 Generated with Claude Code

https://claude.ai/code/session_01TdEJoXUmbHmW5TFk7hfkbK

lilyshen0722 and others added 2 commits August 23, 2026 03:30
…cue (TASK-052)

The cue teaching agents the three verbs was held on one finding: peers
reading context could not see thread structure, so teaching agents to put
continuations in threads would hide the work from its audience. Both
halves now exist, so the cue ships with them:

READ — getRecentMessages' PG mapping carries thread_root_id,
reply_to_message_id and the formatted replyTo (explicit null = not in a
thread; the Mongo fallback omits the keys = server cannot say, same
convention as the frontend). ?threadRootId= on the runtime messages read
scopes to one thread (root + rooted rows); the Mongo fallback REFUSES a
thread-scoped read rather than returning the pod as if it were the thread.

WRITE — threadRootId on the agent post body continues a thread WITHOUT
addressing anyone (constraint 5), through the same threadRootResolver and
five rejections as the human composer; ThreadRootError resolves BEFORE the
PG try so it can never be swallowed into the Mongo fallback as a silently
un-threaded row, and the route maps it to a 400 with the resolver's code.

CUE — the pod-context frame teaches the three verbs, quote-thread
independence, and Sam's overflow rule (prose continuations go in a thread;
attachments are for genuine artifacts), promising only what the kernel
now enforces.

10 new executing tests across three suites; two existing call-contract
pins widened for the new trailing argument; 698/698 across the service +
mention + route suites.

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01TdEJoXUmbHmW5TFk7hfkbK
…njection

The id arrives from agent-supplied metadata; String() at the entrance
collapses operator-object payloads into literals that match nothing.
Surfaced by this PR's re-traced dataflow, same class as #1094's lesson.
Comment thread backend/services/agentEventService.ts Fixed
@lilyshen0722
lilyshen0722 merged commit d13901a into main Aug 23, 2026
11 checks passed
@lilyshen0722
lilyshen0722 deleted the feat-agent-thread-reads branch August 23, 2026 10:48
lilyshen0722 added a commit that referenced this pull request Aug 24, 2026
… (0.3.4) (#1187)

The backend shipped agent threading in #1176 and the pod-context cue has
been advertising it since, but the MCP tool never gained the parameter —
seats were being taught a capability their tool physically lacked, so
they answered with consecutive top-level messages instead (observed live
on the SEO squad's first day). commonly_post_message gains threadRootId
(continue a thread without pinging; prose overflow goes in a thread) and
commonly_get_messages gains the matching read param. Schema-presence
pinned in tests: an undeclared param is one a model can never pass.


Claude-Session: https://claude.ai/code/session_01TdEJoXUmbHmW5TFk7hfkbK

Co-authored-by: Claude Fable 5 <[email protected]>
lilyshen0722 added a commit that referenced this pull request Aug 26, 2026
* fix(cli): prose overflow continues in a thread — the wrapper was the one attaching

Sam (57691) asked to fold "prose overflow goes in a thread, not an attachment"
into the cue. The cue already says it, verbatim, since #1176. The behaviour he
watched all day continued anyway, and the reason is that the cue was aimed one
layer above the thing doing it: agents were not choosing to attach.

`deliverChatReply` decides the delivery mode, and its ladder ended at upload:

    fits in one message         -> post
    splits into <= maxChunks    -> post the chunks
    longer than that            -> upload the whole text as a file

Every `<agent>-reply-<eventId>.md` card in this pod came from that last rung,
including two of mine today. The function predates threads and had no concept
of one. So the cue promised a behaviour the wrapper actively contradicted —
which is @sprint-review's degrades-open caveat exactly: the cue should promise
what the kernel enforces, and here it could not be obeyed at all.

Adds the rung Sam specified — post the headline to the channel, continue the
rest under it with threadRootId — and keeps attach for the case it was always
right for: a single indivisible unit over attachThreshold (a long fence, an
unbreakable run). That is a document by construction; prose that outgrew a
message is not.

Two of this suite's existing tests caught a real bug in the first draft. The
attach rung also leads with chunks[0], so falling through after a successful
headline post duplicated the opening line in the room. The recovery now tracks
whether the headline landed: if it did, post the REMAINDER top-level
(thread-fallback) rather than the whole text again; if the headline itself
failed, nothing reached the room and attach is free to lead as before.

Two existing tests changed meaning rather than being bent to pass, and both
say so at their call site. Their fixtures are prose, which is precisely the
case this reclassifies — their INTENT (nothing is cut; flood beats truncation
or silence) is preserved and now carried by the thread.

Probe: reverting the rung reddens the six behavioural tests and leaves the
indivisible-oversize control green. Suite 61 passed.

Version 0.1.18 -> 0.1.19. Note this collides with #1215's bump if both land —
whichever merges second needs a re-bump, since the guard compares against main.

Co-Authored-By: Claude Opus 5 <[email protected]>

* fix(cli): resume the thread fallback from what actually posted, not from 1

`headlinePosted` was a boolean, so the recovery always resumed at
`chunks.slice(1)`. A boolean can only distinguish "nothing posted" from
"something posted"; it cannot say how much.

Fail a continuation at chunk 3 and chunks 1 and 2 are already in the thread —
the fallback then posts them again top-level, and the reader sees them twice.
That is the duplicate-opening bug this rung exists to prevent, one index
further along.

The suite did not catch it because both existing failure tests throw at the
root-id step, before any continuation has posted. At that instant "something
posted" and "one thing posted" are the same statement, which is exactly when a
boolean stands in for a count without looking wrong.

Now a counter, incremented after each successful post, with the fallback
resuming at `chunks.slice(posted)`. New test fails the 4th post and asserts
every chunk arrives exactly once; reverting to `slice(1)` reddens it alone
(1 of 62).

Found by @sprint-review gating #1217. Pushed onto this branch rather than a
second PR — the head had not moved in three hours.

Co-Authored-By: Claude Opus 5 <[email protected]>

* fix(cli): record runtime post refusals

* fix(cli): wire lint into local and CI checks

---------

Co-authored-by: Claude Opus 5 <[email protected]>
samxu01 pushed a commit that referenced this pull request Aug 29, 2026
…claim

Two corrections from sprint-review's gate, both verified here rather than
accepted:

- /pulls/:n/comments (inline review comments) is a third collection and does
  carry commit_id. The rule stands — every inline comment's
  pull_request_review_id resolves to an event /pulls/:n/reviews returns
  (#1312, #1302, #1260) — but the entry's surface count was wrong, in an
  entry about getting a surface count wrong. Also: they are not rare here;
  a repo-wide sweep finds them on #1312/#1302/#1297/#1274/#1260/#1176/#1094/#1022.
  The 0-across-five-PRs sample was all docs rows.

- The entry claimed the comments collection is "what gh pr view N prints
  without flags". False. Bare gh pr view prints neither. --comments prints
  BOTH interleaved, split only by a status: line and with no sha on either;
  --json comments returns half. On #1338: 2 vs 1.

Co-Authored-By: Claude Opus 5 <[email protected]>
lilyshen0722 added a commit that referenced this pull request Aug 30, 2026
… a commit (#1338)

* docs(ax): entry 51 — a PR's two comment surfaces, and the one without a commit_id

`gh pr view --json comments` and `/pulls/:n/reviews` are disjoint sets, not a
set and a subset: `gh pr review --comment` files a review event that never
appears in the comments collection. The comments surface is the default
projection and the obvious one to reach for, so an agent asking "has anyone
gated the tree that would press?" reads it, sees nothing, and concludes nobody
has — which is what produced a false published warning against pressing a
ready PR.

The sharper half is that an issue comment carries no `commit_id` at all, so
that surface cannot answer the question even when it does show a gate.
Measured across eight open PRs: one with a live gate a comments read omits,
one with a gate at a dead sha, and one correctly gated with zero review
events, where the only thing binding the approval to a tree is that the
reviewer typed the sha into the prose.

Co-Authored-By: Claude Opus 5 <[email protected]>

* docs(ax): entry 51 — third collection, and correct the gh-projection claim

Two corrections from sprint-review's gate, both verified here rather than
accepted:

- /pulls/:n/comments (inline review comments) is a third collection and does
  carry commit_id. The rule stands — every inline comment's
  pull_request_review_id resolves to an event /pulls/:n/reviews returns
  (#1312, #1302, #1260) — but the entry's surface count was wrong, in an
  entry about getting a surface count wrong. Also: they are not rare here;
  a repo-wide sweep finds them on #1312/#1302/#1297/#1274/#1260/#1176/#1094/#1022.
  The 0-across-five-PRs sample was all docs rows.

- The entry claimed the comments collection is "what gh pr view N prints
  without flags". False. Bare gh pr view prints neither. --comments prints
  BOTH interleaved, split only by a status: line and with no sha on either;
  --json comments returns half. On #1338: 2 vs 1.

Co-Authored-By: Claude Opus 5 <[email protected]>

* docs(ax): entry 51 — the gate check built from it is prefix-width-sensitive

The #1330 case forces a prose-sha query; that query has a free width
parameter. This repo writes 8-char shas, so a 9-char prefix returns zero
across all 12 open PRs measured — indistinguishable from an arm that never
ran. At 8 it finds a gate at head on 9 of 12. Prescribe 7 (git's minimum
abbreviation) plus a positive control for any arm that returns an
all-population zero.

Co-Authored-By: Claude Opus 5 <[email protected]>

* docs(ax): entry 51 — delete the prefix width, don't retune it

sprint-review's review of 4ce6e8a is right twice. "This repo writes 8"
is a majority habit, not a rule — #1322 and a #1325 comment write 9
(re-derived, not borrowed). And "cut to 7 so it catches any convention
shorter than 8" is self-refuting: grep 'a1607e8' does not match a1607e,
so 7 relocates the threshold and tells the next reader the check is safe.

Replace the width with a width-free comparison: extract hex tokens from
the body and test whether the head STARTS WITH the token. Verified on the
same population (a1607e8 on #1330, 35e4a1a on #1327). The residual
minimum-token-length knob fails by over-reporting, which is visible,
rather than to zero, which reads as an answer. Promote the positive
control above the width advice — it is what catches the class.

Co-Authored-By: Claude Opus 5 <[email protected]>

---------

Co-authored-by: Claude Opus 5 <[email protected]>
lilyshen0722 added a commit that referenced this pull request Sep 1, 2026
… what the fields do (#1216)

* feat(agents): the three-verb cue tells agents how to CHOOSE, not just what the fields do

Sam's ask (57672) was "teach agents when to use reply, or in thread, or
quote." #1176 shipped the mechanics — what a plain post, `replyToMessageId`
and `threadRootId` each do — and that is the other question. A description of
three fields does not answer a choice, so an agent that has read the whole
paragraph still re-derives which verb its next message wants, every time,
from field semantics.

Adds @ux-lead's decision rule (57678), close to their phrasing on purpose:

    Rule of thumb: if your message answers one person, reply; if it
    continues a topic, thread; if it starts one, post. A reply inside a
    thread is allowed and still addresses its author.

It is written as a test the agent applies to its own draft rather than as
three more facts. The trailing clause is load-bearing: without it the rule
reads as three mutually exclusive branches and an agent concludes it must
pick between quoting and threading, when the two fields are independent.

Verified rather than taken on the copy's word — ux-lead's framing says each
verb "says who is woken", and that claim is checkable. It holds:
threadWakeScopeService.narrowToThread scopes ambient thread activity to the
thread's effective followers and can only NARROW an already-computed opt-in
list, so "wakes followers only" is the real behaviour, not aspirational.

Three tests in the existing inline-cue suite, pinning the decision rule
rather than the paragraph around it — the cue ships as one opaque string, so
"the frame mentions threads" stays green on the mechanics clauses alone.
The third is a control proving the assertions can tell the two halves apart.
Probe: replacing the rule with a mechanics-only tail reddens exactly the two
behavioural tests and leaves the control green. Suite 110 passed; tsc clean
for this file.

Co-Authored-By: Claude Opus 5 <[email protected]>

* fix(agents): the overflow cue must not send substance where nobody wakes

@sprint-review (57706) found the hole in the prose-overflow sentence, and it
is the expensive kind — the cue was obeyable and wrong.

"Post your headline to the channel, continue under your own root" reads as
license to make the top-level message a pointer. It cannot be.
`effectiveFollowerIds` derives `participants` as authors only — `SELECT
DISTINCT user_id FROM messages WHERE thread_root_id = $1 OR id = $1`. At the
instant you open a thread under your own root you are its only author, so you
are its only follower, and `narrowToThread` empties the wake list for every
peer. An agent following the cue literally broadcasts a title and writes the
substance where zero agents are woken.

Two clauses close it, both naming kernel mechanisms rather than preferences:
the top-level message must stand alone (the channel post is the only delivery
the room is guaranteed), and an @mention inside the thread reaches a named
peer regardless of scope — the mention path runs before this narrowing, and
`followMentionedThreadUsers` then writes `following IS TRUE` for that target,
enrolling them for the ambient remainder.

The comment recording the pre-ship verification is corrected too. "Wakes
followers only" was true and insufficient: it confirmed the SET the wake is
narrowed to and never asked what that set contains on the path the cue tells
agents to take. Confirming a predicate is not confirming its extension.

Four guards, including a control that pins the exact unqualified sentence
that shipped before this — so a revert reddens rather than passing on the
shared "thread, not an attachment" phrase. Negative control: dropping the two
clauses reddens exactly 2 of 69, the other 67 stay green.

Co-Authored-By: Claude Opus 5 <[email protected]>

* fix(agents): name what threading does, not only what follows from it

@sprint-review (57707): "continue in-thread" reads as RELOCATING a message
when it is actually UN-ADDRESSING it. That is the intuition behind the
mistake, and the two clauses added in the previous commit do not correct it —
they state mechanisms, and a mechanism does not dislodge a wrong model.

One sentence, guarded separately so a future trim cannot read it as a
flourish on clauses that already "cover it". It is the only line in the frame
that tells an agent threading REMOVES something rather than moving it.

Co-Authored-By: Claude Opus 5 <[email protected]>

* fix(agents): the @mention escape does not survive a mute, and the cue said it did

@sprint-review (58348). The clause added two commits ago promised that
addressing a peer inside a thread "enrols them for the rest of it", flat.
It does not when they have muted the thread: `followByParticipation` writes
only `WHERE thread_user_state.following IS NULL`, and `effectiveFollowerIds`
subtracts `muted` last, so an explicit mute survives both paths.

The mention itself still wakes them — addressing outranks a mute, by design.
What fails is the subscription, which is exactly the half the cue was selling.

Their diagnosis is the reusable part and it is the same shape as the bug it
corrects: I checked that the write HAPPENS and not the condition it is
guarded on. `followByParticipation`'s own docstring names the case outright
("muting a thread and then being mentioned in it is the ordinary case, not an
edge one") and I read past it.

Co-Authored-By: Claude Opus 5 <[email protected]>

* feat(agents): a human is addressed by handle, and the frame never said so (#1244)

* feat(agents): a human is addressed by handle, and the frame never said so

Sam observed 2026-08-25 that seats write about him by name and nothing
routes. The pod-context frame taught three addressing verbs — plain post,
replyToMessageId, threadRootId — and all three move attention between
AGENTS. None reaches a person, and the paragraph never said so, so an
agent that had read it correctly could still conclude that naming a human
was a way of addressing one.

Verified rather than assumed, because the cue is only worth shipping if
the escape it teaches actually works:

- activityService.ts:517-521 builds `mentionNeedle = '@' + lowerUsername`
  and sets `isMention` from `content.includes(needle)`; :591 is the
  `mentions` filter that reads it. Substring on the literal handle.
- resolveHumanMentionUserIds (agentMentionService.ts:1033) extracts
  handles from `[a-z0-9_-]` after an `@`, anchored and case-insensitive.

So `@handle` surfaces in the human's mentions filter and a bare name
matches neither test. The failure is silent — nothing errors, the message
posts, no attention routes — which is why the cue names the outcome and
not just the prescription.

This is the human-facing twin of the gap ADR-018 D6.3 closed for bots: a
message plainly ABOUT someone still has to be addressed TO them before
anything routes. There the fix was a missing implicit-reply wake; here
only the author can supply the handle.

Deliberately teaches the escape and not a heuristic. Whether a bare name
SHOULD route is an open decision (TASK-070b) precisely because name
matching is fuzzy — every message about Sam is not for Sam — so the cue
must not imply that writing the name is enough.

Tests pin the two halves separately (prescription, and the silent-failure
outcome) plus a control built from the pre-change clauses most likely to
keep a loose assertion green: the frame already contains "human" twice and
"@" many times. Mutation-checked — softening "A bare name notifies nobody"
fails the second test and leaves the other two green.

Stacked on #1216, which edits the same frame string; based on its head
rather than main so the two clauses do not conflict.

Co-Authored-By: Claude Opus 5 <[email protected]>

* fix(cue): the handle is necessary, not sufficient — state the ceiling too

@sprint-review's review of #1244: every clause about the failure was
precise and nothing stated the ceiling of the remedy, so an agent reads
"a bare name notifies nobody" as "and the @handle notifies somebody". It
does not. Humans have no AgentEvent delivery row, so the handle buys the
`isMention` flag on the activity feed — a pull surface ADR-017 keeps off
the push channel. Re-derived the narrower half myself rather than
borrowing it: `resolveHumanMentionUserIds` is called only inside
`if (threadRootId)` (:1743), so a plain channel post gets the flag alone
and not even the thread follow.

That would have been a new false model replacing an old one, and harder
to catch — the message now looks correctly addressed while the seat sits
waiting on an answer nobody was told to give.

Two assertions, both mutation-checked; the control gains the same pair.

Co-Authored-By: Claude Opus 5 <[email protected]>

---------

Co-authored-by: Claude Opus 5 <[email protected]>

* test(mentions): pin the human-handle mechanism and put a budget on the wake frame

TASK-074. The pod-context frame makes assertions to agents about how the
kernel behaves, to a reader who cannot falsify them: a seat acts on the cue
and has no view of `enqueueMentions`. Every test on #1216/#1244 is a
string-presence assertion, so a cue can become FALSE while its text is
untouched and the suite stays green.

Two files, both mutation-checked against the pre-existing 113.

**Claim 4 — "the handle is necessary and not sufficient; nothing pushes."**
`agentMentionService.humansAreNotWoken.test.js`, 6 cases, each negative
paired with a control:

- a human @handle enqueues no AgentEvent of any type; the same sentence to
  an installed seat does; one message naming both routes only to the seat.
- the thread-follow half is guarded: a plain channel post makes no
  `followByParticipation` call and does not even run the lookup; the same
  message inside a thread does follow that human; and a follow is not a
  wake — the threaded case still enqueues nothing.

Blind-mutation baseline, run with the new file REMOVED, per @pod-architect's
method on #1249:

| mutation | pre-existing 113 | with this file |
|---|---|---|
| enqueue a chat.mention per resolved human handle (TASK-070b answered "push it") | **113 green** | 4 red |
| hoist `resolveHumanMentionUserIds` out of `if (threadRootId)` | **113 green** | 1 red |

Both are the realistic future edit, not a crude break. The first is the
literal open decision in TASK-070b; the second reads as a consistency fix.

**The frame's own size.** `agentMentionService.frameBudget.test.js` measures
the rendered `chat.mention` content for a reference wake — plain chat pod,
one seat, explicit mention, no thread, no wake-on-message — currently 2,875
chars, and asserts it two-sided against 2,600/3,000. A ceiling alone is
satisfied by deleting the frame, and the copy assertions elsewhere pin
sentences one at a time; neither notices a section going missing. Verified in
both directions: +200 chars fails the ceiling, gutting the Collaboration
block fails the floor.

Not a cap. Raising `BUDGET_MAX` is one line, and that line is the point — it
turns an invisible per-wake, fleet-wide spend into a deliberate one a
reviewer can argue with. **#1216 will fail this and should raise it in its
own diff**; that is the mechanism working, not a conflict.

**Two corrections to the task row I filed, both found by running it.**

Claims 2 and 3 were already pinned, behaviourally, on the shipped SQL —
`threadWakeScope.test.js` runs `effectiveFollowerIds` against pg-mem with the
real DDL, 24 cases. Dropping `OR id = $1` fails 15; dropping the muted
subtraction fails 6; dropping `following IS NULL` from
`followByParticipation` fails exactly the one test written for it. The row's
claim that "every test on both PRs is a string-presence assertion" was wrong
about those two, and nothing here re-covers them.

And #1244 is NOT on main — it merged into #1216's branch, which is still
open. The human-handle cue is unshipped; these tests pin the mechanism at
main, so they hold either way and become that cue's missing companion when
#1216 lands.

122/122 green across all seven agentMentionService suites on Node 22.

Co-Authored-By: Claude Opus 5 <[email protected]>

* test(mentions): carry #1265's budget and raise it for the three-verbs clause

@sprint-review sharpened the merge-order note correctly: order was necessary,
not sufficient. `BUDGET_MAX` lives only on #1265's branch, so this PR could not
raise a constant it did not have — which meant a bulk press turned `main` red in
EITHER order (this first, then #1265 lands on an over-budget frame; #1265 first,
then this one lands red).

Merging #1265's branch here removes the ordering hazard instead of documenting
it. The raise now travels with the growth that caused it, so this PR is safe to
merge in any order, and #1265 stays mergeable on its own.

The band is 3550/4100, kept as tight around the new 3,935-character reference as
2600/3000 was around 2,877. Leaving MIN at 2,600 would have let a third of the
frame disappear without failing — the exact hole the lower bound was added to
close.

What the fleet buys for the extra ~1,058 characters (+37%), per the constant's
own instruction to state the trade: the three addressing verbs, spelled out.
Agents were choosing between plain post / replyToMessageId / threadRootId with
no statement of what each one does to attention, and picking wrong in both
directions — broadcasting what should have been threaded, and threading what
needed a ping.

135 passing across `agentMentionService`.

Co-Authored-By: Claude Opus 5 <[email protected]>

* docs(mentions): cite the call site by symbol, not by line

@sprint-review caught that this comment's `:1743` had drifted to `:1773` — my
own #1265 merge moved the call and left the citation pointing 30 lines short.
Inside the paragraph arguing that claims decay, which is a fair place to be
caught.

Their call was that it is not worth a push of its own, and for a line-number
correction I agree. This is not that: a raw line number in a comment is a
citation that expires on the next edit above it, so fixing the number restores
the same defect for the next person. `resolveHumanMentionUserIds` has exactly
one call site and it is inside the `if (threadRootId)` branch of
`enqueueMentions` — both of which survive an edit that moves the line.

The reason for the change is left in the comment, so the next author sees why
the form is a symbol rather than a number and does not helpfully convert it
back.

Comment-only; the budget test measures string literals on non-comment lines, so
the frame is unchanged. 135 passing across `agentMentionService`.

Co-Authored-By: Claude Opus 5 <[email protected]>

* style(mentions): reflow the over-long comment line left by the citation fix

Comment-only. 4bb0e6d replaced the stale `:1743` citation but left one
line running well past the wrap the rest of the paragraph keeps.

Co-Authored-By: Claude Opus 5 <[email protected]>

* test(cues): give the mechanics half of the three-verb cue a live reader

Self-review gap in this PR, found by mutation on 2026-08-29.

This PR pinned the CHOOSING half of the three-verb cue against the live
frame and left the MECHANICS half unread. The existing control test is
not a reader of it: `mechanicsOnly` is a literal in the test file
asserted against itself, which is the right shape for proving the
choosing assertions discriminate and the wrong shape for noticing that
the cue changed.

Measured before: deleting any mechanics clause from the live cue left all
140 tests green. So did INVERTING one — rewriting the frame to tell every
woken agent that a threaded continuation "pings every member of the pod,
loudly", which is the exact opposite of what `effectiveFollowerIds` does.
Deleting a choosing clause reds 1, so the instrument worked and the gap
was one half of one sentence.

Adds three tests that read the live frame:

  names all three verbs
  states that replyToMessageId pings the author it addresses
  states that threadRootId does NOT ping, and never claims it does

The third excludes the contradiction as well as asserting the negative,
because a cue can carry both sentences at once.

Mutation table, each anchor asserted at exactly one occurrence before
applying, each restored after:

  drop the threadRootId defining clause      was 140 pass -> now 1 FAILED
  drop "its author is pinged"                was 140 pass -> now 2 FAILED
  drop the three-verb opener                 was 140 pass -> now 1 FAILED
  invert to "pings every member of the pod"  was 140 pass -> now 1 FAILED
  copy-edit: colon -> semicolon, reworded         143 pass (unchanged)
  copy-edit: reword the plain-post clause         143 pass (unchanged)

The two copy-edit controls are the point: these are clause-level rather
than whole-paragraph, so ordinary editing does not red the build while a
claim reversal does.

One assertion was tightened after its own mutation came back green.
`toContain('threadRootId')` passes even when the clause defining that
verb is deleted, because the name appears again later in the same frame
("continue the detail under your own root with threadRootId"). It now
asserts 'continues a thread', which reds. A bare name match on a string
that repeats is not a reader of the sentence you meant.

Suite: 140 -> 143, 8 suites, all passing.

Co-Authored-By: Claude Opus 5 <[email protected]>

---------

Co-authored-by: Claude Opus 5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants