
<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <generator uri="https://jekyllrb.com/">Jekyll</generator>
  <link href="https://declare-lab.github.io/feed.xml" rel="self" type="application/atom+xml" />
  <link href="https://declare-lab.github.io/" rel="alternate" type="text/html" />
  <title type="html">DeCLaRe Lab</title>
  <subtitle>Research notes by DeCLaRe Lab members at Nanyang Technological University.</subtitle>
  <id>https://declare-lab.github.io/feed.xml</id><updated>2026-09-09T00:00:00+00:00</updated><author><name>DeCLaRe Lab</name></author><entry>
    <title type="html">ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step</title>
    <link href="https://declare-lab.github.io/lab-notes/scrambletoolbench/" rel="alternate" type="text/html" title="ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step" />
    <published>2026-09-09T00:00:00+00:00</published>
    <updated>2026-09-09T00:00:00+00:00</updated>
    <id>https://declare-lab.github.io/lab-notes/scrambletoolbench/</id>
    <summary type="html">An agent discovers what its tools do. Their names change, and it starts searching again—even when its earlier observations point to the next call.</summary>
    <content type="html" xml:base="https://declare-lab.github.io/lab-notes/scrambletoolbench/">&lt;section class=&quot;lab-note-hero&quot;&gt;
  &lt;div&gt;
    &lt;p class=&quot;work-kicker&quot;&gt;Lab note&lt;/p&gt;
    &lt;h1&gt;ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step&lt;/h1&gt;
    &lt;p&gt;An agent discovers what its tools do. Their names change, and it starts searching again—even when its earlier observations point to the next call.&lt;/p&gt;
    

&lt;div class=&quot;note-byline&quot;&gt;
  &lt;img src=&quot;/assets/images/people/vernon.jpeg&quot; alt=&quot;Vernon Toh&quot; width=&quot;56&quot; height=&quot;56&quot; decoding=&quot;async&quot; /&gt;
  &lt;div&gt;
    &lt;p class=&quot;note-byline__name&quot; data-type-role=&quot;item-title&quot;&gt;Vernon Toh&lt;/p&gt;
    &lt;p class=&quot;note-byline__meta&quot; data-type-role=&quot;meta&quot;&gt;September 9, 2026 · 10 min read · Learning from Interaction&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

    &lt;div class=&quot;project-links&quot;&gt;
      &lt;a href=&quot;https://arxiv.org/abs/2608.02358&quot;&gt;Paper&lt;/a&gt;&lt;a href=&quot;https://github.com/declare-lab/ScrambleToolBench&quot;&gt;GitHub&lt;/a&gt;
    &lt;/div&gt;
  &lt;/div&gt;
  &lt;figure class=&quot;figure-panel note-figure note-figure--compact&quot;&gt;
    &lt;a href=&quot;/assets/images/lab-notes/scrambletoolbench/cryptex-1.png&quot; aria-label=&quot;Open the ScrambleToolBench episode figure at full size&quot;&gt;
      &lt;img src=&quot;/assets/images/lab-notes/scrambletoolbench/cryptex-1.png&quot; alt=&quot;A ScrambleToolBench episode: an agent probes scrambled commands, follows responses to a passcode, submits it and reuses tool knowledge in the next task&quot; width=&quot;2932&quot; height=&quot;828&quot; decoding=&quot;async&quot; fetchpriority=&quot;high&quot; /&gt;
    &lt;/a&gt;
    &lt;figcaption&gt;The agent first has to discover what the commands do. A useful response leads it to the next command and eventually to a solution; that experience carries into the next task. &lt;a href=&quot;/assets/images/lab-notes/scrambletoolbench/cryptex-1.png&quot;&gt;View full-size figure&lt;/a&gt;.&lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/section&gt;

&lt;aside class=&quot;lab-note-share&quot; aria-label=&quot;Share this lab note&quot;&gt;
  &lt;span&gt;Share&lt;/span&gt;
  &lt;div&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--facebook&quot; href=&quot;https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fscrambletoolbench%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on Facebook&quot; title=&quot;Share on Facebook&quot;&gt;
      &lt;i class=&quot;fa-brands fa-facebook-f&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--x&quot; href=&quot;https://x.com/intent/post?text=ScrambleToolBench%3A+Agents+Search+Exhaustively+Even+When+Their+Own+Map+Points+to+the+Next+Step%20https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fscrambletoolbench%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on X&quot; title=&quot;Share on X&quot;&gt;
      &lt;span class=&quot;lab-note-share__x-mark&quot; aria-hidden=&quot;true&quot;&gt;X&lt;/span&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--linkedin&quot; href=&quot;https://www.linkedin.com/shareArticle?mini=true&amp;amp;url=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fscrambletoolbench%2F&amp;amp;title=ScrambleToolBench%3A+Agents+Search+Exhaustively+Even+When+Their+Own+Map+Points+to+the+Next+Step&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on LinkedIn&quot; title=&quot;Share on LinkedIn&quot;&gt;
      &lt;i class=&quot;fa-brands fa-linkedin-in&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
  &lt;/div&gt;
&lt;/aside&gt;

&lt;div class=&quot;lab-note-article&quot; data-prose-align=&quot;left&quot;&gt;

  &lt;p&gt;Suppose an agent has worked out which command reads a file and which one queries a database. It has written the names down and used both successfully. Then the interface changes. The command that used to read a file now calls a different function.&lt;/p&gt;

  &lt;p&gt;The agent could try every command again. But the unexpected response contains a clue: it tells the agent which function has moved into that name. Its old map can tell it where that function used to be. Together, those two observations suggest a specific next call.&lt;/p&gt;

  &lt;p&gt;We built ScrambleToolBench to examine what agents do at moments like this. The strongest models can discover unfamiliar tools and complete tasks with them. Recovering cheaply after a change is harder. Even with the relevant entries in memory, they rarely follow the chain those entries imply.&lt;/p&gt;

  &lt;p&gt;This connects to a question from my &lt;a href=&quot;/lab-notes/mnist-pro/&quot;&gt;MNIST-PRO note&lt;/a&gt;: when an agent has already gathered useful evidence, what stops it from using that evidence? Here the evidence is a tool’s behavior rather than a fragment of an image. We can inspect the next action to see whether the agent has made the connection.&lt;/p&gt;

  &lt;h2 id=&quot;what-the-tool-names-normally-tell-you&quot;&gt;What the tool names normally tell you&lt;/h2&gt;

  &lt;p&gt;A name such as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;read_file&lt;/code&gt; gives an agent a useful starting hypothesis. It suggests both the operation and the kind of argument to supply. That is helpful in practice, but it makes it difficult to tell how much an agent has learned from the current environment and how much it brought with it.&lt;/p&gt;

  &lt;p&gt;ScrambleToolBench removes these cues from 28 tools for file operations, network diagnostics and data processing. A tool receives an arbitrary name such as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fn_3d8a&lt;/code&gt;. Its parameter names are replaced with tags such as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;arg_0&lt;/code&gt;, and the initial schema does not list them. A malformed call reveals the required keys and types. Output fields and success indicators are also obfuscated, so the agent has to interpret the responses it receives.&lt;/p&gt;

  &lt;p&gt;The environment is a Python simulator, with 20 procedural task templates. An episode contains five tasks, each with a budget of 100 inference steps. The tasks share a history: a discovery made while investigating a service can be useful later when locating a file or retrieving a value. Three control commands, including &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;submit_solution&lt;/code&gt;, remain readable so that submitting an answer does not itself require tool discovery.&lt;/p&gt;

  &lt;p&gt;We then change the conditions in three ways. &lt;strong&gt;Mapping drift&lt;/strong&gt; moves seven of the 28 tool identifiers in a cycle between tasks. &lt;strong&gt;Action failure&lt;/strong&gt; makes valid executions return a timeout with probability 0.15. &lt;strong&gt;Execution windows&lt;/strong&gt; require certain procedures to finish within ten actions of a trigger; an expired window resets the state and regenerates intermediate values. We test each change separately and all three together.&lt;/p&gt;

  &lt;p&gt;These changes require different responses. A timeout can justify retrying a correct call. A changed mapping requires revising which command to use. An expired window requires restarting the procedure rather than reusing values from the previous attempt. An agent that treats all three as the same kind of error will waste calls or carry an incorrect assumption forward.&lt;/p&gt;

  &lt;h2 id=&quot;discovery-can-succeed-while-adaptation-fails&quot;&gt;Discovery can succeed while adaptation fails&lt;/h2&gt;

  &lt;p&gt;The main evaluation uses 20 five-task episodes per model and condition. Episode completion means solving all five tasks. With 20 episodes, one completed episode changes the reported completion rate by five percentage points.&lt;/p&gt;

  &lt;p&gt;The selected results below show why the changing environment matters. Claude Sonnet 5, Gemini 3.1 Pro and Gemini 3.5 Flash all complete every episode in the static scrambled setting. Under the three combined changes, their completion rates fall to 0%, 20% and 25%.&lt;/p&gt;

  &lt;div class=&quot;lab-note-table&quot; role=&quot;region&quot; aria-label=&quot;Selected ScrambleToolBench episode completion results&quot; tabindex=&quot;0&quot;&gt;

    &lt;table&gt;
      &lt;thead&gt;
        &lt;tr&gt;
          &lt;th&gt;Model&lt;/th&gt;
          &lt;th&gt;Named tools&lt;/th&gt;
          &lt;th&gt;Scrambled&lt;/th&gt;
          &lt;th&gt;+ Drift&lt;/th&gt;
          &lt;th&gt;+ Failure&lt;/th&gt;
          &lt;th&gt;+ Window&lt;/th&gt;
          &lt;th&gt;+ All&lt;/th&gt;
        &lt;/tr&gt;
      &lt;/thead&gt;
      &lt;tbody&gt;
        &lt;tr&gt;
          &lt;td&gt;Qwen 3.6 27B&lt;/td&gt;
          &lt;td&gt;95%&lt;/td&gt;
          &lt;td&gt;55%&lt;/td&gt;
          &lt;td&gt;25%&lt;/td&gt;
          &lt;td&gt;35%&lt;/td&gt;
          &lt;td&gt;15%&lt;/td&gt;
          &lt;td&gt;0%&lt;/td&gt;
        &lt;/tr&gt;
        &lt;tr&gt;
          &lt;td&gt;Qwen 3.6 27B + Memory&lt;/td&gt;
          &lt;td&gt;—&lt;/td&gt;
          &lt;td&gt;60%&lt;/td&gt;
          &lt;td&gt;40%&lt;/td&gt;
          &lt;td&gt;35%&lt;/td&gt;
          &lt;td&gt;35%&lt;/td&gt;
          &lt;td&gt;0%&lt;/td&gt;
        &lt;/tr&gt;
        &lt;tr&gt;
          &lt;td&gt;Claude Sonnet 5&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;70%&lt;/td&gt;
          &lt;td&gt;0%&lt;/td&gt;
        &lt;/tr&gt;
        &lt;tr&gt;
          &lt;td&gt;Gemini 3.1 Pro&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;90%&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;80%&lt;/td&gt;
          &lt;td&gt;20%&lt;/td&gt;
        &lt;/tr&gt;
        &lt;tr&gt;
          &lt;td&gt;Gemini 3.1 Pro + Memory&lt;/td&gt;
          &lt;td&gt;—&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;90%&lt;/td&gt;
          &lt;td&gt;50%&lt;/td&gt;
        &lt;/tr&gt;
        &lt;tr&gt;
          &lt;td&gt;Gemini 3.5 Flash&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;90%&lt;/td&gt;
          &lt;td&gt;85%&lt;/td&gt;
          &lt;td&gt;65%&lt;/td&gt;
          &lt;td&gt;25%&lt;/td&gt;
        &lt;/tr&gt;
        &lt;tr&gt;
          &lt;td&gt;Gemini 3.5 Flash + Memory&lt;/td&gt;
          &lt;td&gt;—&lt;/td&gt;
          &lt;td&gt;100%&lt;/td&gt;
          &lt;td&gt;95%&lt;/td&gt;
          &lt;td&gt;95%&lt;/td&gt;
          &lt;td&gt;80%&lt;/td&gt;
          &lt;td&gt;30%&lt;/td&gt;
        &lt;/tr&gt;
        &lt;tr&gt;
          &lt;td&gt;Mean across all 15 models, without memory&lt;/td&gt;
          &lt;td&gt;93%&lt;/td&gt;
          &lt;td&gt;32%&lt;/td&gt;
          &lt;td&gt;23%&lt;/td&gt;
          &lt;td&gt;26%&lt;/td&gt;
          &lt;td&gt;19%&lt;/td&gt;
          &lt;td&gt;3%&lt;/td&gt;
        &lt;/tr&gt;
      &lt;/tbody&gt;
    &lt;/table&gt;

  &lt;/div&gt;

  &lt;p&gt;These are selected rows from &lt;a href=&quot;https://arxiv.org/html/2608.02358v1#S2.T2&quot;&gt;Table 2 of the paper&lt;/a&gt;. The mean includes the full set of 15 no-memory models, not just the four displayed here. A dash denotes an unreported condition, not zero performance.&lt;/p&gt;

  &lt;p&gt;The memory variant maintains two structured stores: task recipes and tool knowledge. One records reusable procedures; the other records inferred tool behavior, arguments and confidence. The agent updates these stores as it acts. For Gemini 3.1 Pro, memory raises completion in the combined condition from 20% to 50%. That is a useful gain, but it leaves another question: does the agent use its remembered map to reason about the change, or does it simply search more successfully?&lt;/p&gt;

  &lt;h2 id=&quot;a-wrong-call-can-tell-you-where-to-look-next&quot;&gt;A wrong call can tell you where to look next&lt;/h2&gt;

  &lt;p&gt;Consider a small illustrative example. Before a change, command A reads files, B queries the database and C lists interfaces. After the names rotate, A lists interfaces, C queries the database and B reads files.&lt;/p&gt;

  &lt;p&gt;The agent wants to read a file, so it starts with A. The response identifies the interface-listing function. The old map says that function used to be at C, so C is the next place to look. C now returns the database function, whose old name was B. Calling B reaches the file reader. The route is A → C → B; the agent did not need to scan the remaining commands.&lt;/p&gt;

  &lt;p&gt;That is &lt;strong&gt;cycle tracing&lt;/strong&gt;. After a mismatch, identify the function that answered, look up its previous identifier, and call that identifier next. Repeat until the required function appears. The stable, function-specific response fields let the agent match a response to an earlier observation even when the command name changes.&lt;/p&gt;

  &lt;p&gt;For the benchmark’s seven-identifier cycle, this takes six extra calls relative to an ordinary successful tool call. That guarantee assumes the stored map was current immediately before the drift and the returned functions can be identified. The same calls reveal the changed part of the map, allowing it to be repaired before the next task. If the map mixes unresolved observations from several changes, the six-call bound no longer follows.&lt;/p&gt;

  &lt;p&gt;The paper also gives a reference for a task requiring four of the 28 functions. With seven identifiers chosen uniformly for drift, the probability that at least one required function moved is one minus the probability that all four stayed unchanged. Multiplying that probability by six gives &lt;strong&gt;4.25 expected extra actions&lt;/strong&gt; at one fresh task boundary. This is a recovery reference under the stated assumptions, not a bound on the cost of discovering the whole environment or handling the combined stressors. The derivation is in &lt;a href=&quot;https://arxiv.org/html/2608.02358v1#S4.SS2&quot;&gt;Section 4.2&lt;/a&gt;.&lt;/p&gt;

  &lt;h2 id=&quot;do-agents-follow-that-clue&quot;&gt;Do agents follow that clue?&lt;/h2&gt;

  &lt;p&gt;We measure whether, after observing a changed mapping, the agent calls the next identifier in the chain within three actions. The comparison is with random selection among identifiers the agent has already mapped.&lt;/p&gt;

  &lt;div class=&quot;lab-note-table&quot; role=&quot;region&quot; aria-label=&quot;Use of the recovery chain after mapping drift&quot; tabindex=&quot;0&quot;&gt;

    &lt;table&gt;
      &lt;thead&gt;
        &lt;tr&gt;
          &lt;th&gt;Model&lt;/th&gt;
          &lt;th&gt;Reasoning&lt;/th&gt;
          &lt;th&gt;Opportunities&lt;/th&gt;
          &lt;th&gt;Follows within three actions&lt;/th&gt;
          &lt;th&gt;Random baseline&lt;/th&gt;
        &lt;/tr&gt;
      &lt;/thead&gt;
      &lt;tbody&gt;
        &lt;tr&gt;
          &lt;td&gt;Gemini 3.1 Pro&lt;/td&gt;
          &lt;td&gt;High&lt;/td&gt;
          &lt;td&gt;484&lt;/td&gt;
          &lt;td&gt;11.0%&lt;/td&gt;
          &lt;td&gt;10.8%&lt;/td&gt;
        &lt;/tr&gt;
        &lt;tr&gt;
          &lt;td&gt;Gemini 3.1 Pro + Memory&lt;/td&gt;
          &lt;td&gt;High&lt;/td&gt;
          &lt;td&gt;444&lt;/td&gt;
          &lt;td&gt;12.8%&lt;/td&gt;
          &lt;td&gt;10.9%&lt;/td&gt;
        &lt;/tr&gt;
        &lt;tr&gt;
          &lt;td&gt;Claude Sonnet 5&lt;/td&gt;
          &lt;td&gt;Low&lt;/td&gt;
          &lt;td&gt;564&lt;/td&gt;
          &lt;td&gt;14.0%&lt;/td&gt;
          &lt;td&gt;10.6%&lt;/td&gt;
        &lt;/tr&gt;
        &lt;tr&gt;
          &lt;td&gt;Claude Sonnet 5&lt;/td&gt;
          &lt;td&gt;Medium&lt;/td&gt;
          &lt;td&gt;590&lt;/td&gt;
          &lt;td&gt;11.9%&lt;/td&gt;
          &lt;td&gt;10.6%&lt;/td&gt;
        &lt;/tr&gt;
        &lt;tr&gt;
          &lt;td&gt;Claude Sonnet 5&lt;/td&gt;
          &lt;td&gt;High&lt;/td&gt;
          &lt;td&gt;604&lt;/td&gt;
          &lt;td&gt;14.1%&lt;/td&gt;
          &lt;td&gt;10.6%&lt;/td&gt;
        &lt;/tr&gt;
      &lt;/tbody&gt;
    &lt;/table&gt;

  &lt;/div&gt;

  &lt;p&gt;The opportunity counts are diagnostic events, not independent episodes; the analysis excludes early-exit episodes as described in &lt;a href=&quot;https://arxiv.org/html/2608.02358v1#S4.T5&quot;&gt;Table 5&lt;/a&gt;. Gemini’s follow rates are close to the random baseline, including with memory. Sonnet is modestly above it at low and high reasoning, but increasing reasoning does not make it follow the chain more consistently. These results do not mean that no useful deduction ever occurs. They show that this particular recovery rule is not used reliably.&lt;/p&gt;

  &lt;p&gt;One trajectory makes the alternative behavior easy to recognize. Qwen 3.6 27B had used &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fn_fc40&lt;/code&gt; for file search. After drift, it kept varying the arguments to that name over a seven-step loop, trying different search patterns instead of reconsidering which function the name now referred to. The agent was changing its query while keeping the more consequential assumption fixed.&lt;/p&gt;

  &lt;p&gt;Memory can also be revised sensibly. In another episode, Gemini 3.1 Pro explicitly cleared its old tool mappings while retaining task progress and critical details. Under an execution window, Gemini 3.5 Flash stopped making exploratory calls and focused on the required sequence. These examples show useful local adaptation. They also explain why a single success rate cannot tell us whether the agent found the cheaper recovery procedure.&lt;/p&gt;

  &lt;h2 id=&quot;more-reasoning-improves-completion-but-not-the-shortcut&quot;&gt;More reasoning improves completion, but not the shortcut&lt;/h2&gt;

  &lt;p&gt;We vary the reasoning effort for Gemini 3.1 Pro and Claude Sonnet 5 in both the static scrambled environment and the drift condition. Higher effort improves the number of tasks completed, especially for Gemini. The chain-following results above show that this improvement should not be mistaken for reliable cycle tracing.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure note-figure--small&quot;&gt;
  &lt;a href=&quot;/assets/images/lab-notes/scrambletoolbench/reasoning-tavg-1.png&quot; aria-label=&quot;Open the reasoning effort and task completion plot at full size&quot;&gt;
    &lt;img src=&quot;/assets/images/lab-notes/scrambletoolbench/reasoning-tavg-1.png&quot; alt=&quot;Mean tasks completed per episode at low, medium and high reasoning: performance approaches five tasks at high effort for Gemini 3.1 Pro and Claude Sonnet 5, in Base and Drift&quot; width=&quot;450&quot; height=&quot;321&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;/a&gt;
  &lt;figcaption&gt;Mean tasks completed per episode, out of five. Blue circles denote the static scrambled Base condition; orange squares denote Drift. &lt;a href=&quot;/assets/images/lab-notes/scrambletoolbench/reasoning-tavg-1.png&quot;&gt;View full-size figure&lt;/a&gt;.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;Completion tokens per solved task count tokens spent across all episodes, including failed ones, and divide by the total number of tasks solved. At high reasoning under drift, Sonnet uses 11,652 tokens per solved task, compared with Gemini’s 3,332—about 3.5 times as many at similar task completion.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure note-figure--small&quot;&gt;
  &lt;a href=&quot;/assets/images/lab-notes/scrambletoolbench/reasoning-token-cost-1.png&quot; aria-label=&quot;Open the reasoning effort and token cost plot at full size&quot;&gt;
    &lt;img src=&quot;/assets/images/lab-notes/scrambletoolbench/reasoning-token-cost-1.png&quot; alt=&quot;Completion tokens per solved task: Gemini 3.1 Pro uses roughly 1,600 in Base and 3,300 in Drift across reasoning levels; Claude Sonnet 5 uses more, with its largest drift cost at low effort&quot; width=&quot;458&quot; height=&quot;317&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;/a&gt;
  &lt;figcaption&gt;Completion tokens per solved task, in thousands. The cost includes unsuccessful episodes. Lower reasoning effort is not consistently cheaper. &lt;a href=&quot;/assets/images/lab-notes/scrambletoolbench/reasoning-token-cost-1.png&quot;&gt;View full-size figure&lt;/a&gt;.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;The cost does not rise monotonically with reasoning effort. Sonnet’s low-effort drift setting is the most expensive per solved task; Gemini’s cost changes little across effort levels. The distinction is between completing more tasks and adopting a more economical strategy. A larger reasoning budget can improve the former while leaving the latter largely unchanged.&lt;/p&gt;

  &lt;h2 id=&quot;the-connection-to-mnist-pro&quot;&gt;The connection to MNIST-PRO&lt;/h2&gt;

  &lt;p&gt;In &lt;a href=&quot;/lab-notes/mnist-pro/&quot;&gt;MNIST-PRO&lt;/a&gt;, an agent moves a small window over a handwritten digit. It may collect many useful fragments and still give the wrong answer. In ScrambleToolBench, it may identify a changed function and retain the relevant earlier mapping, yet choose an unrelated command next. Both failures happen after useful evidence has reached the agent.&lt;/p&gt;

  &lt;p&gt;The two benchmarks let us investigate that gap in different ways. For MNIST-PRO, we keep a completed trajectory fixed and assemble its observed glimpses into a coordinate-aligned canvas. No new region becomes visible. If the final answer improves, a better search path cannot explain the gain: organizing the existing observations has helped. For example, on Claude Opus 5’s Textual State trajectories, this changes Level 1 accuracy from 41% to 76% and Level 2 accuracy from 13% to 67%.&lt;/p&gt;

  &lt;p&gt;For ScrambleToolBench, we inspect a decision rather than supply a consolidated view. Once a changed call returns a known function, does the agent use its old map to choose the next identifier? The low follow rates show that storing the relevant entries does not reliably produce that action. We have not shown here that a different memory format would fix the problem; the MNIST-PRO canvas result is not evidence that the same intervention works for tool use.&lt;/p&gt;

  &lt;p&gt;There is also an important difference in what changes. The digit in MNIST-PRO stays still. An early guess may be wrong, or the agent may have failed to join its views together, but a previously observed stroke does not move. ScrambleToolBench deliberately changes the interface: a mapping that was correct can become stale. The first task calls for constructing and interpreting a state from partial views; the second additionally calls for revising that state when the environment changes.&lt;/p&gt;

  &lt;p&gt;This is why I would not treat either result as a general argument for more memory. A history can preserve the right observations without exposing their spatial relationship. A tool dictionary can preserve a correct old mapping without prompting the reverse lookup that would make it useful now. The question is what the stored information allows the agent to do at the next decision.&lt;/p&gt;

  &lt;h2 id=&quot;what-i-would-test-next&quot;&gt;What I would test next&lt;/h2&gt;

  &lt;p&gt;For tool use, I would separate three checks: whether the map contains the needed entries, whether the agent recognizes a mismatch, and whether it follows the implication of that mismatch. A controlled intervention could provide the current pre-drift map, then test a reverse-lookup rule only when a response identifies a displaced function. That would help distinguish incomplete discovery from a failure to use information already available. It is a proposed test, not a result reported here.&lt;/p&gt;

  &lt;p&gt;For perception, the corresponding check is to hold the path fixed before changing how its observations are represented. If the agent never saw the second digit, a better arrangement of its existing glimpses cannot supply it. If it did see the distinguishing strokes, the next question is whether it can use them together.&lt;/p&gt;

  &lt;p&gt;The practical aim is to locate the missed step. In ScrambleToolBench, an error response can be the observation that tells the agent where to call next. In MNIST-PRO, several individually ambiguous crops can become a recognizable digit when placed together. An agent needs to carry those relationships into its decisions, not just carry the observations into its context.&lt;/p&gt;

  &lt;h2 id=&quot;citation&quot;&gt;Citation&lt;/h2&gt;

  &lt;div class=&quot;language-bibtex highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nc&quot;&gt;@misc&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;toh2026scrambletoolbenchagentssearchexhaustively&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;title&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;author&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{Vernon Toh and Navonil Majumder and Zhengyuan Liu and Nancy F. Chen and Soujanya Poria}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;year&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{2026}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;eprint&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{2608.02358}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;archivePrefix&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{arXiv}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;primaryClass&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{cs.CL}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;url&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{https://arxiv.org/abs/2608.02358}&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;  &lt;/div&gt;

&lt;/div&gt;</content>
    <author><name>Vernon Toh</name></author>
    <category term="AI Agents" />
    <category term="Learning from Interaction" />
    <category term="Memory Representations" />
    <media:content medium="image" url="https://declare-lab.github.io/assets/images/lab-notes/scrambletoolbench/cryptex-1.png" xmlns:media="http://search.yahoo.com/mrss/" />
  </entry><entry>
    <title type="html">MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents</title>
    <link href="https://declare-lab.github.io/lab-notes/mnist-pro/" rel="alternate" type="text/html" title="MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents" />
    <published>2026-09-08T00:00:00+00:00</published>
    <updated>2026-09-08T00:00:00+00:00</updated>
    <id>https://declare-lab.github.io/lab-notes/mnist-pro/</id>
    <summary type="html">When an agent gets a digit wrong, did it miss the evidence or struggle to put the glimpses together? We use MNIST-PRO to tell these failures apart.</summary>
    <content type="html" xml:base="https://declare-lab.github.io/lab-notes/mnist-pro/">&lt;section class=&quot;lab-note-hero&quot;&gt;
  &lt;div&gt;
    &lt;p class=&quot;work-kicker&quot;&gt;Lab note&lt;/p&gt;
    &lt;h1&gt;MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents&lt;/h1&gt;
    &lt;p&gt;When an agent gets a digit wrong, did it miss the evidence or struggle to put the glimpses together? We use MNIST-PRO to tell these failures apart.&lt;/p&gt;
    

&lt;div class=&quot;note-byline&quot;&gt;
  &lt;img src=&quot;/assets/images/people/vernon.jpeg&quot; alt=&quot;Vernon Toh&quot; width=&quot;56&quot; height=&quot;56&quot; decoding=&quot;async&quot; /&gt;
  &lt;div&gt;
    &lt;p class=&quot;note-byline__name&quot; data-type-role=&quot;item-title&quot;&gt;Vernon Toh&lt;/p&gt;
    &lt;p class=&quot;note-byline__meta&quot; data-type-role=&quot;meta&quot;&gt;September 8, 2026 · 7 min read · Agentic Perception&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

    &lt;div class=&quot;project-links&quot;&gt;
      &lt;a href=&quot;https://arxiv.org/abs/2608.31022&quot;&gt;Paper&lt;/a&gt;&lt;a href=&quot;https://github.com/declare-lab/MNIST-PRO&quot;&gt;GitHub&lt;/a&gt;
    &lt;/div&gt;
  &lt;/div&gt;
  &lt;figure class=&quot;figure-panel note-figure note-figure--compact&quot;&gt;
    &lt;a href=&quot;/assets/images/lab-notes/mnist-pro/trajectory_episode_0.png&quot; aria-label=&quot;Open the MNIST-PRO trajectory figure at full size&quot;&gt;
      &lt;img src=&quot;/assets/images/lab-notes/mnist-pro/trajectory_episode_0.png&quot; alt=&quot;A successful MNIST-PRO episode: four glimpses along a handwritten zero, ending with the correct answer at step 9&quot; width=&quot;1536&quot; height=&quot;456&quot; decoding=&quot;async&quot; fetchpriority=&quot;high&quot; /&gt;
    &lt;/a&gt;
    &lt;figcaption&gt;The full digit on the left is hidden from the agent. The selected glimpses show a successful search that ends with the answer 0 at step 9. &lt;a href=&quot;/assets/images/lab-notes/mnist-pro/trajectory_episode_0.png&quot;&gt;View full-size figure&lt;/a&gt;.&lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/section&gt;

&lt;aside class=&quot;lab-note-share&quot; aria-label=&quot;Share this lab note&quot;&gt;
  &lt;span&gt;Share&lt;/span&gt;
  &lt;div&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--facebook&quot; href=&quot;https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fmnist-pro%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on Facebook&quot; title=&quot;Share on Facebook&quot;&gt;
      &lt;i class=&quot;fa-brands fa-facebook-f&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--x&quot; href=&quot;https://x.com/intent/post?text=MNIST-PRO%3A+MNIST+is+Back+as+a+Partially+Observable+World+for+AI+Agents%20https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fmnist-pro%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on X&quot; title=&quot;Share on X&quot;&gt;
      &lt;span class=&quot;lab-note-share__x-mark&quot; aria-hidden=&quot;true&quot;&gt;X&lt;/span&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--linkedin&quot; href=&quot;https://www.linkedin.com/shareArticle?mini=true&amp;amp;url=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fmnist-pro%2F&amp;amp;title=MNIST-PRO%3A+MNIST+is+Back+as+a+Partially+Observable+World+for+AI+Agents&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on LinkedIn&quot; title=&quot;Share on LinkedIn&quot;&gt;
      &lt;i class=&quot;fa-brands fa-linkedin-in&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
  &lt;/div&gt;
&lt;/aside&gt;

&lt;div class=&quot;lab-note-article&quot; data-prose-align=&quot;left&quot;&gt;

  &lt;p&gt;Look at the zero above. With the whole image in view, recognizing it is easy. Now cover everything except a small square near the bottom-right stroke. That fragment could belong to several digits. To answer, you have to move the window, remember the earlier fragments, and work out how they fit together.&lt;/p&gt;

  &lt;p&gt;In the illustrated episode, the agent moves up along the stroke, reaches the top of the digit, and eventually answers 0. The interesting part is not the label. It is the work between the first glimpse and that answer: deciding where to look, keeping track of position, and deciding when the observations support a conclusion.&lt;/p&gt;

  &lt;p&gt;We built MNIST-PRO to study that work. In this note, I want to focus on what a wrong answer can tell us—and why collecting more evidence is not always the same as making better use of it.&lt;/p&gt;

  &lt;h2 id=&quot;why-return-to-handwritten-digits&quot;&gt;Why return to handwritten digits?&lt;/h2&gt;

  &lt;p&gt;Imagine navigating a room with only a flashlight. You need to combine successive views into a useful picture of your surroundings. But if a robot fails in that room, the cause may be difficult to identify: it could have misread an object, forgotten a location, collided with something, or failed to execute an action.&lt;/p&gt;

  &lt;p&gt;MNIST gives us a simpler setting. There is no physics to simulate and no object to manipulate. The agent moves a window over a static image and reports a digit. We can record exactly which parts it saw before it answered.&lt;/p&gt;

  &lt;p&gt;The window is 64 × 64 pixels and moves 32 pixels at a time, up, down, left or right. Level 1 places one digit on a 224 × 224 canvas. Level 2 places two digits side by side on a 224 × 448 canvas and asks for their ordered sequence: seeing a 5 and an 8 is not enough if the answer should be &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;58&lt;/code&gt;, not &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;85&lt;/code&gt;.&lt;/p&gt;

  &lt;p&gt;This keeps the recognition problem familiar while changing how the evidence arrives. It also gives us two separate questions to ask of an unsuccessful episode: &lt;strong&gt;did the agent look in the right places, and could it use what it had already seen?&lt;/strong&gt;&lt;/p&gt;

  &lt;h2 id=&quot;what-survives-after-the-window-moves&quot;&gt;What survives after the window moves?&lt;/h2&gt;

  &lt;p&gt;Keeping every image in context sounds like an obvious solution. That is our Image Only baseline: the agent retains its previous glimpses and actions. But a sequence of crops is not a map. The model still has to work out which views overlap and where each fragment belongs.&lt;/p&gt;

  &lt;p&gt;We compare this with two ways of writing information down. In Textual State, only the current image is visible; descriptions of earlier observations carry information forward. In Metric Grid Map, the agent also records features against relative coordinates in a structured map. A description tells it what it saw. Coordinates are intended to help it remember where.&lt;/p&gt;

  &lt;p&gt;These representations impose different demands. A text description can omit a small visual detail. A map can contain an incorrect position. A long visual history can contain all the fragments without making their spatial relationship clear. None of these formats guarantees that the agent has a reliable picture of the digit.&lt;/p&gt;

  &lt;p&gt;The full-image control makes the contrast concrete. Gemini 3.1 Pro reaches 99% accuracy on the single-digit control but 38% in the Image Only condition. The digit has not become a different object; the model now has to gather and combine views of it.&lt;/p&gt;

  &lt;h2 id=&quot;did-it-actually-see-enough&quot;&gt;Did it actually see enough?&lt;/h2&gt;

  &lt;p&gt;Accuracy alone does not answer this. We also measure &lt;strong&gt;stroke coverage&lt;/strong&gt;: how much of the digit’s foreground was exposed through the glimpse window. This is different from the number of moves. An agent can spend many steps returning to the same region.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;a href=&quot;/assets/images/lab-notes/mnist-pro/evidence_acquisition_vs_native_use.svg&quot; aria-label=&quot;Open the evidence coverage and accuracy plot at full size&quot;&gt;
    &lt;img src=&quot;/assets/images/lab-notes/mnist-pro/evidence_acquisition_vs_native_use.svg&quot; alt=&quot;Stroke coverage versus task accuracy: Claude Fable 5 has 76.5% coverage and 29.5% accuracy, while Gemini 3.7 Flash has 64.0% coverage and 36.8% accuracy&quot; width=&quot;1440&quot; height=&quot;900&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;/a&gt;
  &lt;figcaption&gt;Evidence acquired versus task accuracy. Each point averages the four conditions formed by Level 1 and Level 2 × Image Only and Textual State; it does not include Metric Grid Map. &lt;a href=&quot;/assets/images/lab-notes/mnist-pro/evidence_acquisition_vs_native_use.svg&quot;&gt;View full-size figure&lt;/a&gt;.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;The plot shows why coverage is worth measuring separately. Claude Fable 5 covers more of the strokes than Gemini 3.7 Flash in this comparison, yet answers fewer episodes correctly. In the Textual State condition specifically, Fable and Opus exceed 80% stroke coverage, while their Level 1 accuracies are 38% and 41%; their Level 2 accuracies are 18% and 13%.&lt;/p&gt;

  &lt;p&gt;That does not make coverage useless. Nor does it prove that memory is the sole cause of failure. A large covered area may still leave out the part that distinguishes two digits. What it tells us is that “the agent should look more” is not a complete diagnosis.&lt;/p&gt;

  &lt;p&gt;To go further, we need an intervention that changes how the evidence is presented without giving the agent another chance to explore.&lt;/p&gt;

  &lt;h2 id=&quot;keep-the-path-fixed-put-the-glimpses-together&quot;&gt;Keep the path fixed; put the glimpses together&lt;/h2&gt;

  &lt;p&gt;We take completed Image Only and Textual State trajectories and arrange the already observed glimpses on a coordinate-aligned canvas. The model then makes a final prediction from this consolidated view. The route stays fixed. No new region of the digit is revealed.&lt;/p&gt;

  &lt;p&gt;This matters because a better answer can no longer be attributed to a better search path. We have changed the organization of the evidence available at decision time.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;a href=&quot;/assets/images/lab-notes/mnist-pro/evidence_representation_rescue.svg&quot; aria-label=&quot;Open the fixed-trajectory canvas comparison at full size&quot;&gt;
    &lt;img src=&quot;/assets/images/lab-notes/mnist-pro/evidence_representation_rescue.svg&quot; alt=&quot;With the same collected glimpses assembled into a canvas, average accuracy increases for all eight plotted model settings; Claude Fable 5 rises from 29.5% to 66.5%&quot; width=&quot;1440&quot; height=&quot;900&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;/a&gt;
  &lt;figcaption&gt;Open circles show the original predictions; filled circles show predictions with the consolidated canvas. Coverage stays fixed. Both endpoints use the same four-condition averaging as the preceding plot. These are offline final-answer interventions, not new exploration runs. &lt;a href=&quot;/assets/images/lab-notes/mnist-pro/evidence_representation_rescue.svg&quot;&gt;View full-size figure&lt;/a&gt;.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;All eight model settings plotted improve on average. Looking at the Textual State trajectories separately makes the size of the change easier to see:&lt;/p&gt;

  &lt;table&gt;
    &lt;thead&gt;
      &lt;tr&gt;
        &lt;th&gt;Model&lt;/th&gt;
        &lt;th&gt;Level 1: original → canvas&lt;/th&gt;
        &lt;th&gt;Level 2: original → canvas&lt;/th&gt;
      &lt;/tr&gt;
    &lt;/thead&gt;
    &lt;tbody&gt;
      &lt;tr&gt;
        &lt;td&gt;Claude Fable 5&lt;/td&gt;
        &lt;td&gt;38% → 81%&lt;/td&gt;
        &lt;td&gt;18% → 69%&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
        &lt;td&gt;Claude Opus 5&lt;/td&gt;
        &lt;td&gt;41% → 76%&lt;/td&gt;
        &lt;td&gt;13% → 67%&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
        &lt;td&gt;Gemini 3.6 Flash&lt;/td&gt;
        &lt;td&gt;47% → 89%&lt;/td&gt;
        &lt;td&gt;15% → 49%&lt;/td&gt;
      &lt;/tr&gt;
    &lt;/tbody&gt;
  &lt;/table&gt;

  &lt;p&gt;These results support a narrower, useful conclusion: for these trajectories, much of the error can be reduced by organizing the observations differently. They do not show that exploration no longer matters. A canvas cannot recover a stroke that the agent never saw. And the intervention supplies spatial alignment, so it does not separate every possible failure in remembering, aligning and interpreting the fragments. The per-condition results are in &lt;a href=&quot;https://arxiv.org/html/2608.31022v1#S3.T2&quot;&gt;Table 2 of the paper&lt;/a&gt;.&lt;/p&gt;

  &lt;h2 id=&quot;some-agents-stop-before-reaching-the-second-digit&quot;&gt;Some agents stop before reaching the second digit&lt;/h2&gt;

  &lt;p&gt;The other side of the diagnosis is visible in the trajectories. In an audit of 1,600 Level 2 episodes using Textual State or Metric Grid Map, 323 predictions contained only one digit. In 238 of those episodes, the agent never looked at the second digit at all. In 266, it exposed less than a quarter of that digit. These single-digit answers arrived after roughly 15 steps on average, out of an available 78.&lt;/p&gt;

  &lt;p&gt;Here, reorganizing the existing evidence cannot solve the missing-observation problem. The agent has committed before checking the whole task.&lt;/p&gt;

  &lt;p&gt;We also see redundant exploration and failures to revise early hypotheses. For Gemini 3.7 Flash, adding a Metric Grid Map reduces the revisit rate from 45.1% to 13.6%. That is a substantial change in where it spends its moves, but a more efficient route is not itself a correct answer. Likewise, an early guess must remain revisable: later observations are useful only if they can change the prediction.&lt;/p&gt;

  &lt;p&gt;These failures call for different tests. An episode that never reaches the second digit needs a different explanation from one that observes both digits and still answers incorrectly.&lt;/p&gt;

  &lt;h2 id=&quot;what-if-the-agent-can-write-its-own-tools&quot;&gt;What if the agent can write its own tools?&lt;/h2&gt;

  &lt;p&gt;The offline canvas is something we construct for the model. We also tested a harness in which agents can write and execute Python and manage local files. Could they build useful representations during the task themselves?&lt;/p&gt;

  &lt;p&gt;In the Gemini 3.7 Flash runs, the agent wrote scripts such as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;stitch.py&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;render_mnist.py&lt;/code&gt; to assemble glimpses into a coordinate-aligned matrix, inspect an ASCII rendering, and guide further exploration. Its harness accuracy was 88% on Level 1 and 63% on Level 2.&lt;/p&gt;

  &lt;p&gt;This is a different experiment from replaying a fixed trajectory: tools can change both the representation and the search that follows. The result is encouraging, but it should not be read as an isolated measurement of memory quality. It shows a model using computation to make its observations easier to work with.&lt;/p&gt;

  &lt;h2 id=&quot;when-the-observations-are-tool-responses&quot;&gt;When the observations are tool responses&lt;/h2&gt;

  &lt;p&gt;My &lt;a href=&quot;/lab-notes/scrambletoolbench/#the-connection-to-mnist-pro&quot;&gt;ScrambleToolBench note&lt;/a&gt; asks a related question about tool use. An agent learns which commands perform which operations, then some command names change. A response from a displaced function, together with the old map, can point to the next command to try. Agents often search elsewhere instead.&lt;/p&gt;

  &lt;p&gt;The distinction matters. In MNIST-PRO, the digit stays fixed, and assembling the same glimpses into a canvas tests whether their organization affects the answer. In ScrambleToolBench, the interface changes, and we inspect whether the agent uses the resulting mismatch to revise its next action. Both ask what happens after useful evidence arrives. Keeping that evidence available is only part of the work; the agent still has to use the relationship between observations.&lt;/p&gt;

  &lt;h2 id=&quot;what-i-would-check-before-adding-more-context&quot;&gt;What I would check before adding more context&lt;/h2&gt;

  &lt;p&gt;For an agent that fails on a partially observed task, I would first inspect what it had actually seen at the time it answered. Did it reach every relevant object? Did it revisit the same area? Did it stop while much of the task remained unseen?&lt;/p&gt;

  &lt;p&gt;Then I would hold that evidence fixed and change its representation. If the answer improves when the same observations are organized into a map, collecting additional observations is not the only route to improvement. If crucial evidence is absent, the search and stopping policy need attention too.&lt;/p&gt;

  &lt;p&gt;MNIST-PRO lets us make those checks in a small, controlled world. Handwritten digits do not establish how an agent will behave on a website or in a physical room. They do give us a way to tell apart errors that a single accuracy score would otherwise hide.&lt;/p&gt;

  &lt;h2 id=&quot;citation&quot;&gt;Citation&lt;/h2&gt;

  &lt;div class=&quot;language-bibtex highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nc&quot;&gt;@misc&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;toh2026mnistpromnistpartiallyobservable&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;title&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;author&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{Vernon Toh and Navonil Majumder and Zhengyuan Liu and Nancy F. Chen and Soujanya Poria}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;year&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{2026}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;eprint&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{2608.31022}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;archivePrefix&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{arXiv}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;primaryClass&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{cs.AI}&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;url&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{https://arxiv.org/abs/2608.31022}&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;  &lt;/div&gt;

&lt;/div&gt;</content>
    <author><name>Vernon Toh</name></author>
    <category term="Agentic Perception" />
    <category term="Multimodal Agents" />
    <category term="Memory Representations" />
    <category term="Learning from Interaction" />
    <media:content medium="image" url="https://declare-lab.github.io/assets/images/lab-notes/mnist-pro/trajectory_episode_0.png" xmlns:media="http://search.yahoo.com/mrss/" />
  </entry><entry>
    <title type="html">IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation</title>
    <link href="https://declare-lab.github.io/lab-notes/ideagent/" rel="alternate" type="text/html" title="IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation" />
    <published>2026-08-12T00:00:00+00:00</published>
    <updated>2026-08-12T00:00:00+00:00</updated>
    <id>https://declare-lab.github.io/lab-notes/ideagent/</id>
    <summary type="html">IDEAgent searches for a diverse set of candidate research ideas and records how each candidate changes during refinement.</summary>
    <content type="html" xml:base="https://declare-lab.github.io/lab-notes/ideagent/">&lt;section class=&quot;lab-note-hero&quot;&gt;
  &lt;div&gt;
    &lt;p class=&quot;work-kicker&quot;&gt;Lab note&lt;/p&gt;
    &lt;h1&gt;IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation&lt;/h1&gt;
    &lt;p&gt;IDEAgent asks whether a fixed generation budget can produce more ideas that meet stated quality thresholds without repeating one another.&lt;/p&gt;
    

&lt;div class=&quot;note-byline&quot;&gt;
  &lt;img src=&quot;/assets/images/people/varun-gumma.jpg&quot; alt=&quot;Varun Gumma&quot; width=&quot;56&quot; height=&quot;56&quot; decoding=&quot;async&quot; /&gt;
  &lt;div&gt;
    &lt;p class=&quot;note-byline__name&quot; data-type-role=&quot;item-title&quot;&gt;Varun Gumma&lt;/p&gt;
    &lt;p class=&quot;note-byline__meta&quot; data-type-role=&quot;meta&quot;&gt;August 12, 2026 · 7 min read · AI for Science&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

    &lt;div class=&quot;project-links&quot;&gt;
      &lt;a href=&quot;https://arxiv.org/abs/2607.22375&quot;&gt;Paper&lt;/a&gt;
      &lt;a href=&quot;https://github.com/declare-lab/IDEAgent&quot;&gt;GitHub&lt;/a&gt;
      &lt;a href=&quot;https://medium.com/@varun.gumma/ideagent-agentic-quality-diversity-search-for-research-idea-generation-bcea0b63600a&quot;&gt;Medium&lt;/a&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/section&gt;

&lt;aside class=&quot;lab-note-share&quot; aria-label=&quot;Share this lab note&quot;&gt;
  &lt;span&gt;Share&lt;/span&gt;
  &lt;div&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--facebook&quot; href=&quot;https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fideagent%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on Facebook&quot; title=&quot;Share on Facebook&quot;&gt;
      &lt;i class=&quot;fa-brands fa-facebook-f&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--x&quot; href=&quot;https://x.com/intent/post?text=IDEAgent%3A+Agentic+Quality-Diversity+Search+for+Research+Idea+Generation%20https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fideagent%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on X&quot; title=&quot;Share on X&quot;&gt;
      &lt;span class=&quot;lab-note-share__x-mark&quot; aria-hidden=&quot;true&quot;&gt;X&lt;/span&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--linkedin&quot; href=&quot;https://www.linkedin.com/shareArticle?mini=true&amp;amp;url=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fideagent%2F&amp;amp;title=IDEAgent%3A+Agentic+Quality-Diversity+Search+for+Research+Idea+Generation&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on LinkedIn&quot; title=&quot;Share on LinkedIn&quot;&gt;
      &lt;i class=&quot;fa-brands fa-linkedin-in&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
  &lt;/div&gt;
&lt;/aside&gt;

&lt;div class=&quot;lab-note-article&quot; data-prose-align=&quot;left&quot;&gt;

  &lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;

  &lt;p&gt;Most ideation systems optimize one proposal at a time. This can produce a high-scoring idea while the full set remains repetitive. IDEAgent instead treats ideation as a &lt;strong&gt;quality-diversity search&lt;/strong&gt; problem: under a fixed budget, it tries to increase the number of ideas that clear stated quality thresholds and remain distinct from those already accepted.&lt;/p&gt;

  &lt;h2 id=&quot;how-ideagent-works&quot;&gt;How IDEAgent works&lt;/h2&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/ideagent/search-overview.png&quot; alt=&quot;IDEAgent sequentially generates ideas, checks diversity, and repairs or refines eligible proposals&quot; width=&quot;3004&quot; height=&quot;1164&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Figure 1: IDEAgent generates ideas sequentially for diversification and applies repair or refinement to eligible proposals. Descendants retain the lineage of their parent.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;The system has four main pieces:&lt;/p&gt;

  &lt;ul&gt;
    &lt;li&gt;Generate ideas &lt;em&gt;sequentially&lt;/em&gt;, using the context of prior ideas to explicitly avoid repetition and encourage diversification.&lt;/li&gt;
    &lt;li&gt;Decompose each free-form idea into specific fields, providing a more interpretable representation for pairwise comparison between ideas.&lt;/li&gt;
    &lt;li&gt;Treat improvements to an idea as children of a common parent and assign them the same &lt;em&gt;lineage&lt;/em&gt; identifier, so they are not mistakenly counted as separate “new discoveries.”&lt;/li&gt;
    &lt;li&gt;Maintain &lt;em&gt;archives&lt;/em&gt; of currently accepted ideas, historical/retired ideas (superseded by newer versions), and rejected ideas and families. Rejected ideas are clustered into families to capture broader avoidance patterns and help identify similar ideas in the future.&lt;/li&gt;
  &lt;/ul&gt;

  &lt;figure class=&quot;figure-panel note-figure note-figure--small&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/ideagent/archives.png&quot; alt=&quot;Active, historical and rejected IDEAgent archives and the transitions between them&quot; width=&quot;928&quot; height=&quot;1080&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Figure 2: An incoming idea can replace its closest active duplicate if it has a higher quality score. The other version moves to the historical archive.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;Six LLM agents handle generation, summarisation and evaluation:&lt;/p&gt;

  &lt;ul&gt;
    &lt;li&gt;&lt;strong&gt;Ideator&lt;/strong&gt;: The primary generator that produces new ideas or refines existing ones.&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;Stenographe&lt;/strong&gt;: Compresses each idea into a structured representation by identifying the &lt;em&gt;problem addressed, central mechanism, novel addition over the background literature, key assumptions, and expected measurable effect&lt;/em&gt; from the prose.&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;Quality Judge&lt;/strong&gt;: Grades the idea ($I$) on &lt;em&gt;Non-obviousness&lt;/em&gt; ($N_I$), &lt;em&gt;Clarity&lt;/em&gt; ($C_I$), and &lt;em&gt;Feasibility&lt;/em&gt; ($F_I$) on a scale of 0–100.&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;Soundness Panel&lt;/strong&gt;: Assesses the logical and mathematical rigor of the idea and its assumptions using five independent judgments for broader coverage. The judgments are given on a scale of 0–9 and then aggregated into a single score on a 0–100 scale ($S_I$).&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;Diversity Judge&lt;/strong&gt;: Compares the idea against all &lt;em&gt;archives&lt;/em&gt; and identifies the nearest duplicates from each. It also assigns a single score ($D_I$) on a 0–100 scale to capture the overall uniqueness, novelty, and diversity of the current idea with respect to prior generations.&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;Critic&lt;/strong&gt;: Assimilates the judgments and scores from the three aforementioned judges to produce a single, targeted improvement directive for the Ideator, guiding it to improve the current generation.&lt;/li&gt;
  &lt;/ul&gt;

  &lt;p&gt;We define the overall quality score as: $Q = 0.7N_I + 0.2S_I + 0.1C_I$, with ties decided by $F_I$.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/ideagent/overall-framework.png&quot; alt=&quot;IDEAgent framework showing ideation, quality and diversity evaluation, archives, repair and refinement&quot; width=&quot;1600&quot; height=&quot;1127&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Figure 3: IDEAgent&apos;s agents, archives and evaluation loop.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/ideagent/idea-lifecycle.png&quot; alt=&quot;State transitions in an idea&apos;s lifecycle from generation through rejection, repair, acceptance and refinement&quot; width=&quot;1600&quot; height=&quot;844&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Figure 4: Transitions between generation, rejection, repair, acceptance and refinement.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;h3 id=&quot;search-procedure&quot;&gt;Search procedure&lt;/h3&gt;

  &lt;ul&gt;
    &lt;li&gt;An idea generated by the Ideator is first compressed by the Steno into a structured summary, which is used for future comparisons.&lt;/li&gt;
    &lt;li&gt;The full idea is then scored by all three judges. The Diversity Judge also produces a list of near-duplicates from each &lt;em&gt;archive&lt;/em&gt;.&lt;/li&gt;
    &lt;li&gt;If the idea fails to clear the required threshold for any metric (Non-obviousness, Clarity, Soundness, or Diversity) by more than a permissible margin, it is rejected and added to the rejected &lt;em&gt;archive&lt;/em&gt;. It is also assigned to one or more rejection families based on its failure pattern.&lt;/li&gt;
    &lt;li&gt;If the idea matches an idea in the historical &lt;em&gt;archive&lt;/em&gt; but has no match in the active set, it is discarded because it has replicated an older idea that has already been superseded.&lt;/li&gt;
    &lt;li&gt;If the idea falls below a threshold but remains within the permissible margin, it gets an opportunity for &lt;em&gt;repair&lt;/em&gt;. Based on the Critic’s feedback, which incorporates the evaluation results, the Ideator corrects the bottleneck or problematic aspects of the idea. The repaired idea is then evaluated again. If it still fails to cross the required thresholds, it is rejected.&lt;/li&gt;
    &lt;li&gt;If the idea, either before or after repair, clears all the thresholds, it becomes eligible for acceptance. If the active set (which has a fixed capacity) is not full, the idea is added directly. If the active set is full, the incoming idea competes against its closest match in the active set, as identified by the Diversity Judge. The idea with the higher overall (Q) score is retained, while the other is moved to the historical &lt;em&gt;archive&lt;/em&gt;. If the incumbent is replaced, the newcomer inherits its lineage because it was identified as a duplicate of the incumbent, thereby preserving the lineage.&lt;/li&gt;
    &lt;li&gt;Once an idea is accepted into the active set, it has up to two additional &lt;em&gt;refinement&lt;/em&gt; opportunities, provided that one of them was not already used for repair. Refinement is triggered when the idea falls below a higher &lt;em&gt;soft target&lt;/em&gt; above the acceptance thresholds. The Critic’s feedback is again provided to the Ideator, and the resulting variant is accepted if it preserves Non-obviousness while improving the overall quality score. If the refinement does not meet these criteria, the new child is discarded.&lt;/li&gt;
    &lt;li&gt;Once the search budget is exhausted, the ideas remaining in the active set are treated as the final variants and passed on for downstream evaluation.&lt;/li&gt;
  &lt;/ul&gt;

  &lt;h2 id=&quot;evaluation&quot;&gt;Evaluation&lt;/h2&gt;

  &lt;p&gt;We evaluate our ideas at both the individual and group levels using LLM-based judges that are distinct from those used during generation (i.e., the internal evaluation).&lt;/p&gt;

  &lt;p&gt;At the individual level, we evaluate each idea for &lt;em&gt;Non-obviousness, Soundness&lt;/em&gt;, and &lt;em&gt;Clarity&lt;/em&gt; as our primary quality metrics, each on a 0–9 scale. At the group level, we compute &lt;em&gt;Diversity&lt;/em&gt; pairwise between ideas, also on a 0–9 scale.
Finally, &lt;strong&gt;Yield&lt;/strong&gt; filters ideas through predefined quality gates and measures the largest subset that also satisfies a pairwise diversity threshold.&lt;/p&gt;

  &lt;h2 id=&quot;results&quot;&gt;Results&lt;/h2&gt;

  &lt;p&gt;We compare IDEAgent with four alternatives and ablations:&lt;/p&gt;

  &lt;ul&gt;
    &lt;li&gt;&lt;strong&gt;Stateless&lt;/strong&gt;: Naive parallel generation of N ideas, with no information shared between them.&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;Single-Shot&lt;/strong&gt;: Generates all N ideas at once in a single output. For a thinking-based Ideator, the reasoning space is shared across all N ideas, allowing them to share a common memory block for reference and diversification.&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;Sequential Memory&lt;/strong&gt;: Sequentially generates ideas while providing summaries of prior generations as references for diversification. In other words, this is IDEAgent without the quality-improvement component.&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;NOVA&lt;/strong&gt;: A NOVA-inspired version of Sequential Memory that uses iterative &lt;em&gt;seed-pool-based&lt;/em&gt; germination and replacement for diversification. In each round, N ideas are generated, after which an internal judge selects the K most promising ideas as seeds for generating the next N ideas. After three rounds, the internal judge selects the most diverse N ideas from the resulting 3N ideas.&lt;/li&gt;
  &lt;/ul&gt;

  &lt;figure class=&quot;figure-panel note-figure note-figure--small&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/ideagent/results-table.png&quot; alt=&quot;IDEAgent results table comparing soundness, clarity, diversity, non-obviousness and Yield across methods&quot; width=&quot;1448&quot; height=&quot;1152&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Figure 5: Results across 32 topics. A successful topic has non-zero Yield. S: soundness; C: clarity; D: diversity; NB: non-obviousness.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/ideagent/yield-zoomed.png&quot; alt=&quot;Detailed Yield surface comparing IDEAgent with sequential-memory generation across quality thresholds&quot; width=&quot;1246&quot; height=&quot;738&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Figure 6: Yield for IDEAgent and Sequential Memory at $S \in [6, 8]$, $NB \in [6, 7]$ and $D \geq 7$.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/ideagent/yield-full.png&quot; alt=&quot;Full Yield surface across soundness, non-obviousness and diversity thresholds&quot; width=&quot;2438&quot; height=&quot;948&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Full 10x10 Yield surface: $S \in [6, 8]$, $NB \in [6, 7]$, $D \geq 7$.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;h2 id=&quot;limits-and-next-steps&quot;&gt;Limits and next steps&lt;/h2&gt;

  &lt;p&gt;IDEAgent evaluates ideation at the level of a set, not only one proposal at a time. Yield makes the quality and diversity thresholds explicit. Both the search and the evaluation depend on model-based judges, so the results do not show that the generated proposals are scientifically correct or useful. Expert review, literature checks and experiments are still required.&lt;/p&gt;

  &lt;h2 id=&quot;citation&quot;&gt;Citation&lt;/h2&gt;

  &lt;ul&gt;
    &lt;li&gt;&lt;strong&gt;Code&lt;/strong&gt;: &lt;a href=&quot;https://github.com/declare-lab/IDEAgent&quot;&gt;https://github.com/declare-lab/IDEAgent&lt;/a&gt;&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;Paper&lt;/strong&gt;: &lt;a href=&quot;https://arxiv.org/abs/2607.22375&quot;&gt;https://arxiv.org/abs/2607.22375&lt;/a&gt;&lt;/li&gt;
  &lt;/ul&gt;

  &lt;div class=&quot;language-latex highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;@misc&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;gumma2026ideagentagenticqualitydiversitysearch,
      title=&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;,
      author=&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;Varun Gumma and Navonil Majumder and Soumitra Sinhahajari and Soujanya Poria&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;,
      year=&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;2026&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;,
      eprint=&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;2607.22375&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;,
      archivePrefix=&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;arXiv&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;,
      primaryClass=&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;cs.AI&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;,
      url=&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;https://arxiv.org/abs/2607.22375&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;,
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;  &lt;/div&gt;

&lt;/div&gt;</content>
    <author><name>Varun Gumma</name></author>
    <category term="AI for Science" />
    <category term="Agents" />
    <category term="Scientific discovery" />
    <media:content medium="image" url="https://declare-lab.github.io/assets/images/lab-notes/ideagent/search-overview.png" xmlns:media="http://search.yahoo.com/mrss/" />
  </entry><entry>
    <title type="html">Toward Efficient Data-Centric Training, Part II: Beyond What to Select - How Much Data Should We Use?</title>
    <link href="https://declare-lab.github.io/lab-notes/data-centric-training-part-ii/" rel="alternate" type="text/html" title="Toward Efficient Data-Centric Training, Part II: Beyond What to Select - How Much Data Should We Use?" />
    <published>2026-05-14T00:00:00+00:00</published>
    <updated>2026-05-14T00:00:00+00:00</updated>
    <id>https://declare-lab.github.io/lab-notes/data-centric-training-part-ii/</id>
    <summary type="html">PODS changes how much data is used at each stage of training instead of fixing one selection ratio.</summary>
    <content type="html" xml:base="https://declare-lab.github.io/lab-notes/data-centric-training-part-ii/">&lt;section class=&quot;lab-note-hero&quot;&gt;
  &lt;div&gt;
    &lt;p class=&quot;work-kicker&quot;&gt;Lab note · Part II&lt;/p&gt;
    &lt;h1&gt;Toward Efficient Data-Centric Training, Part II: Beyond What to Select - How Much Data Should We Use?&lt;/h1&gt;
    &lt;p&gt;Part I asked which examples a model should use. PODS asks whether the amount of selected data should also change during training.&lt;/p&gt;
    

&lt;div class=&quot;note-byline&quot;&gt;
  &lt;img src=&quot;/assets/images/people/suorong-yang.png&quot; alt=&quot;Suorong Yang&quot; width=&quot;56&quot; height=&quot;56&quot; decoding=&quot;async&quot; /&gt;
  &lt;div&gt;
    &lt;p class=&quot;note-byline__name&quot; data-type-role=&quot;item-title&quot;&gt;Suorong Yang&lt;/p&gt;
    &lt;p class=&quot;note-byline__meta&quot; data-type-role=&quot;meta&quot;&gt;May 14, 2026 · 5 min read · Efficiency&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

    &lt;div class=&quot;project-links&quot;&gt;
      &lt;a href=&quot;https://arxiv.org/abs/2605.14773&quot;&gt;Paper&lt;/a&gt;
      &lt;a href=&quot;/lab-notes/data-centric-training-part-i/&quot;&gt;Part I&lt;/a&gt;
    &lt;/div&gt;
  &lt;/div&gt;
  &lt;figure class=&quot;figure-panel note-figure note-figure--compact&quot;&gt;
    &lt;img src=&quot;/assets/images/lab-notes/data-centric-training-part-ii/figure-01.png&quot; alt=&quot;PODS oscillatory data-volume scheduling overview&quot; width=&quot;1124&quot; height=&quot;798&quot; decoding=&quot;async&quot; fetchpriority=&quot;high&quot; /&gt;
    &lt;figcaption&gt;PODS changes the selected data volume over the course of training.&lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/section&gt;

&lt;aside class=&quot;lab-note-share&quot; aria-label=&quot;Share this lab note&quot;&gt;
  &lt;span&gt;Share&lt;/span&gt;
  &lt;div&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--facebook&quot; href=&quot;https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fdata-centric-training-part-ii%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on Facebook&quot; title=&quot;Share on Facebook&quot;&gt;
      &lt;i class=&quot;fa-brands fa-facebook-f&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--x&quot; href=&quot;https://x.com/intent/post?text=Toward+Efficient+Data-Centric+Training%2C+Part+II%3A+Beyond+What+to+Select+-+How+Much+Data+Should+We+Use%3F%20https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fdata-centric-training-part-ii%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on X&quot; title=&quot;Share on X&quot;&gt;
      &lt;span class=&quot;lab-note-share__x-mark&quot; aria-hidden=&quot;true&quot;&gt;X&lt;/span&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--linkedin&quot; href=&quot;https://www.linkedin.com/shareArticle?mini=true&amp;amp;url=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fdata-centric-training-part-ii%2F&amp;amp;title=Toward+Efficient+Data-Centric+Training%2C+Part+II%3A+Beyond+What+to+Select+-+How+Much+Data+Should+We+Use%3F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on LinkedIn&quot; title=&quot;Share on LinkedIn&quot;&gt;
      &lt;i class=&quot;fa-brands fa-linkedin-in&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
  &lt;/div&gt;
&lt;/aside&gt;

&lt;div class=&quot;lab-note-article&quot; data-prose-align=&quot;left&quot;&gt;

  &lt;h2 id=&quot;why-a-fixed-selection-ratio-is-limiting&quot;&gt;Why a fixed selection ratio is limiting&lt;/h2&gt;

  &lt;p&gt;Most data selection methods focus on deciding &lt;strong&gt;what&lt;/strong&gt; to select. They may dynamically change the identity of selected samples, but the selected data volume is usually fixed by a target ratio such as 50% or 70% throughout training.&lt;/p&gt;

  &lt;p&gt;That assumption is convenient, but training is not static. Early, middle, and late training phases can have different optimization needs. Sometimes, seeing less data can act like regularization by forcing the model to focus on informative or harder samples. At other times, seeing more data is important for coverage and stable optimization.&lt;/p&gt;

  &lt;p&gt;If the ratio is too low, training may be efficient but unstable or biased. If the ratio is too high, training is more stable but loses much of the efficiency and regularization benefit. PODS treats the selected data volume as a temporal control signal rather than a fixed cost knob.&lt;/p&gt;

  &lt;h2 id=&quot;how-pods-works&quot;&gt;How PODS works&lt;/h2&gt;

  &lt;p&gt;PODS stands for Plug-and-play Oscillatory Data-volume Scheduling. It alternates between low- and high-ratio phases during training.&lt;/p&gt;

  &lt;p&gt;Low-ratio phases expose the model to a smaller amount of selected data. These phases strengthen the regularization effect induced by data selection and reduce computation. High-ratio phases expose the model to more data, improving coverage and helping optimization recover.&lt;/p&gt;

  &lt;p&gt;The important constraint is that PODS works under the same cumulative data budget as the fixed-ratio baseline. The total amount of data used across training does not exceed the target budget. What changes is &lt;strong&gt;when&lt;/strong&gt; the data is used.&lt;/p&gt;

  &lt;p&gt;&lt;strong&gt;Low-ratio phase -&amp;gt; High-ratio recovery phase -&amp;gt; Low-ratio phase -&amp;gt; High-ratio recovery phase&lt;/strong&gt;&lt;/p&gt;

  &lt;h2 id=&quot;why-alternate-the-ratio&quot;&gt;Why alternate the ratio?&lt;/h2&gt;

  &lt;p&gt;The intuition comes from viewing data selection through optimization. When training on selected data, the stochastic gradient differs from the full-data gradient. This difference can introduce an implicit regularization effect, and PODS shows that the strength of this effect is modulated by the instantaneous selection ratio.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/data-centric-training-part-ii/figure-02.png&quot; alt=&quot;PODS implicit regularization proposition&quot; width=&quot;1650&quot; height=&quot;432&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;The selected-data objective can be understood as introducing a ratio-dependent implicit regularization effect.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;A lower selection ratio leads to stronger selection-induced regularization, which can reduce overfitting and improve generalization. But if the ratio remains too low for too long, the model may lose data coverage and optimization can become unstable.&lt;/p&gt;

  &lt;p&gt;A higher selection ratio weakens this regularization effect, but improves coverage and makes optimization closer to full-data training. PODS is designed to exploit both sides of this trade-off rather than choosing one fixed point.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/data-centric-training-part-ii/figure-03.png&quot; alt=&quot;PODS phase-aligned oscillatory training dynamics&quot; width=&quot;1280&quot; height=&quot;590&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Low-ratio and high-ratio phases produce phase-aligned oscillatory training dynamics.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;h2 id=&quot;pods-does-not-change-the-scoring-rule&quot;&gt;PODS does not change the scoring rule&lt;/h2&gt;

  &lt;p&gt;PODS does not propose a new sample-importance metric. It is orthogonal to selectors based on loss, uncertainty, gradient norm, clustering, coverage, or learned data policies such as Data Agent.&lt;/p&gt;

  &lt;p&gt;A selector decides &lt;strong&gt;what&lt;/strong&gt; to select. PODS decides &lt;strong&gt;how much&lt;/strong&gt; to select at a given training stage. This makes it a lightweight module that can be placed on top of existing static or dynamic data selection methods.&lt;/p&gt;

  &lt;h2 id=&quot;reported-results&quot;&gt;Reported results&lt;/h2&gt;

  &lt;p&gt;PODS is evaluated across image classification, fine-grained recognition, long-tailed classification, out-of-distribution generalization, object detection, and LLM instruction tuning.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/data-centric-training-part-ii/figure-04.png&quot; alt=&quot;PODS classification results across selection methods&quot; width=&quot;696&quot; height=&quot;444&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;PODS improves the efficiency-generalization trade-off across several existing selection methods.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;On ImageNet-1k, PODS reduces training cost while maintaining or improving accuracy, showing that data-volume scheduling can scale beyond small benchmarks. The paper reports 50% ImageNet-1k training-cost reduction with improved accuracy.&lt;/p&gt;

  &lt;p&gt;PODS also generalizes to more challenging recognition settings. On fine-grained classification, it preserves or improves full-data performance with reduced training costs. On long-tailed recognition, it improves results especially for medium-shot and few-shot categories, suggesting that scheduled data exposure can help models learn more balanced representations.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/data-centric-training-part-ii/figure-06.png&quot; alt=&quot;PODS LLM instruction-tuning results&quot; width=&quot;1446&quot; height=&quot;352&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Reported instruction-tuning cost and downstream benchmark results.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;In large-scale fine-tuning experiments with Qwen and LLaMA models, PODS accelerates instruction tuning while maintaining competitive performance on benchmarks such as MMLU and BBH. This suggests that scheduling data volume is useful for both vision and language model training.&lt;/p&gt;

  &lt;h2 id=&quot;what-the-two-studies-show&quot;&gt;What the two studies show&lt;/h2&gt;

  &lt;p&gt;Data Agent and PODS address two separate choices in efficient data-centric training.&lt;/p&gt;

  &lt;p&gt;&lt;strong&gt;Part I:&lt;/strong&gt; learn what data to select.&lt;/p&gt;

  &lt;p&gt;&lt;strong&gt;Part II:&lt;/strong&gt; schedule how much data to use.&lt;/p&gt;

  &lt;p&gt;Both make data selection part of the training loop: Data Agent changes the examples, while PODS changes their volume. The papers report the effects on training cost and performance across the evaluated vision and language tasks.&lt;/p&gt;

&lt;/div&gt;</content>
    <author><name>Suorong Yang</name></author>
    <category term="Efficiency" />
    <category term="Data-centric training" />
    <media:content medium="image" url="https://declare-lab.github.io/assets/images/lab-notes/data-centric-training-part-ii/figure-01.png" xmlns:media="http://search.yahoo.com/mrss/" />
  </entry><entry>
    <title type="html">Toward Efficient Data-Centric Training, Part I: Can Models Learn What Data They Need?</title>
    <link href="https://declare-lab.github.io/lab-notes/data-centric-training-part-i/" rel="alternate" type="text/html" title="Toward Efficient Data-Centric Training, Part I: Can Models Learn What Data They Need?" />
    <published>2026-05-13T00:00:00+00:00</published>
    <updated>2026-05-13T00:00:00+00:00</updated>
    <id>https://declare-lab.github.io/lab-notes/data-centric-training-part-i/</id>
    <summary type="html">Data Agent updates its data-selection policy as the target model trains.</summary>
    <content type="html" xml:base="https://declare-lab.github.io/lab-notes/data-centric-training-part-i/">&lt;section class=&quot;lab-note-hero&quot;&gt;
  &lt;div&gt;
    &lt;p class=&quot;work-kicker&quot;&gt;Lab note · Part I&lt;/p&gt;
    &lt;h1&gt;Toward Efficient Data-Centric Training, Part I: Can Models Learn What Data They Need?&lt;/h1&gt;
    &lt;p&gt;Not every example is equally useful throughout training. Data Agent learns a selection policy alongside the model it is training.&lt;/p&gt;
    

&lt;div class=&quot;note-byline&quot;&gt;
  &lt;img src=&quot;/assets/images/people/suorong-yang.png&quot; alt=&quot;Suorong Yang&quot; width=&quot;56&quot; height=&quot;56&quot; decoding=&quot;async&quot; /&gt;
  &lt;div&gt;
    &lt;p class=&quot;note-byline__name&quot; data-type-role=&quot;item-title&quot;&gt;Suorong Yang&lt;/p&gt;
    &lt;p class=&quot;note-byline__meta&quot; data-type-role=&quot;meta&quot;&gt;May 13, 2026 · 5 min read · Efficiency&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

    &lt;div class=&quot;project-links&quot;&gt;
      &lt;a href=&quot;https://arxiv.org/abs/2603.07433&quot;&gt;Paper&lt;/a&gt;
      &lt;a href=&quot;https://github.com/Jackbrocp/Data-Agent&quot;&gt;GitHub&lt;/a&gt;
      &lt;a href=&quot;/lab-notes/data-centric-training-part-ii/&quot;&gt;Part II&lt;/a&gt;
    &lt;/div&gt;
  &lt;/div&gt;
  &lt;figure class=&quot;figure-panel note-figure note-figure--compact&quot;&gt;
    &lt;img src=&quot;/assets/images/lab-notes/data-centric-training-part-i/figure-02.png&quot; alt=&quot;Data Agent training-loop overview&quot; width=&quot;1280&quot; height=&quot;886&quot; decoding=&quot;async&quot; fetchpriority=&quot;high&quot; /&gt;
    &lt;figcaption&gt;Data Agent updates its selection policy as the target model trains.&lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/section&gt;

&lt;aside class=&quot;lab-note-share&quot; aria-label=&quot;Share this lab note&quot;&gt;
  &lt;span&gt;Share&lt;/span&gt;
  &lt;div&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--facebook&quot; href=&quot;https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fdata-centric-training-part-i%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on Facebook&quot; title=&quot;Share on Facebook&quot;&gt;
      &lt;i class=&quot;fa-brands fa-facebook-f&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--x&quot; href=&quot;https://x.com/intent/post?text=Toward+Efficient+Data-Centric+Training%2C+Part+I%3A+Can+Models+Learn+What+Data+They+Need%3F%20https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fdata-centric-training-part-i%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on X&quot; title=&quot;Share on X&quot;&gt;
      &lt;span class=&quot;lab-note-share__x-mark&quot; aria-hidden=&quot;true&quot;&gt;X&lt;/span&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--linkedin&quot; href=&quot;https://www.linkedin.com/shareArticle?mini=true&amp;amp;url=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fdata-centric-training-part-i%2F&amp;amp;title=Toward+Efficient+Data-Centric+Training%2C+Part+I%3A+Can+Models+Learn+What+Data+They+Need%3F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on LinkedIn&quot; title=&quot;Share on LinkedIn&quot;&gt;
      &lt;i class=&quot;fa-brands fa-linkedin-in&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
  &lt;/div&gt;
&lt;/aside&gt;

&lt;div class=&quot;lab-note-article&quot; data-prose-align=&quot;left&quot;&gt;

  &lt;h2 id=&quot;why-data-selection-needs-to-adapt&quot;&gt;Why data selection needs to adapt&lt;/h2&gt;

  &lt;p&gt;Data selection is often treated as a cost-reduction tool: use fewer samples, train faster, and preserve performance if possible. That view is useful, but incomplete. The usefulness of a sample is not fixed.&lt;/p&gt;

  &lt;p&gt;Early in training, a model may need examples that help build broad representations. Later, it may need samples near the decision boundary, uncertain samples, or cases that reveal what the model still misunderstands. A sample that is useful at one stage can become redundant at another.&lt;/p&gt;

  &lt;p&gt;Many existing methods rely on predefined scoring rules such as loss, gradient norm, clustering distance, or other handcrafted metrics. These rules can be effective in specific settings, but they are often static, snapshot-based, or tied to a particular task. Data Agent reframes the problem: data selection should be an adaptive part of optimization, not only a preprocessing step.&lt;/p&gt;

  &lt;h2 id=&quot;how-data-agent-works&quot;&gt;How Data Agent works&lt;/h2&gt;

  &lt;p&gt;Data Agent places a lightweight agent inside the training loop. At each stage, the agent observes the current state of the target model and outputs sample-wise selection weights. The selected data is then used to update the model. After the update, the model’s new state becomes fresh feedback for the agent.&lt;/p&gt;

  &lt;p&gt;This creates a closed loop:&lt;/p&gt;

  &lt;p&gt;&lt;strong&gt;Model state -&amp;gt; Data Agent -&amp;gt; Selected data -&amp;gt; Model update -&amp;gt; New model state&lt;/strong&gt;&lt;/p&gt;

  &lt;p&gt;Instead of using one fixed importance score, Data Agent learns a policy that changes with the target model. The model trains on selected samples; its updated state then provides feedback for the next selection step.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/data-centric-training-part-i/figure-01.png&quot; alt=&quot;Heuristic-driven data selection compared with optimization-driven selection&quot; width=&quot;1640&quot; height=&quot;1048&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Data Agent shifts selection from heuristic-driven scoring to optimization-driven policy learning.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;More concretely, Data Agent formulates dynamic data selection as a sequential decision-making problem. The agent observes feature representations from the target model as its state, outputs continuous selection weights as actions, and prioritizes samples with higher weights during training.&lt;/p&gt;

  &lt;h2 id=&quot;training-signals&quot;&gt;Training signals&lt;/h2&gt;

  &lt;p&gt;A central question is how the agent knows which samples are useful. Data Agent uses two complementary training-aware signals.&lt;/p&gt;

  &lt;p&gt;&lt;strong&gt;Difficulty&lt;/strong&gt; captures what the model has not learned well. It is based on sample loss: high-loss samples often expose under-learned regions of the data distribution and can have larger optimization impact.&lt;/p&gt;

  &lt;p&gt;&lt;strong&gt;Uncertainty&lt;/strong&gt; captures where the model is unsure. A sample may be classified correctly but still lie near a decision boundary. Predictive entropy helps identify these samples and encourages the agent to refine regions where confidence remains low.&lt;/p&gt;

  &lt;p&gt;The relative importance of difficulty and uncertainty changes over training. Data Agent therefore introduces an adaptive weighting mechanism that balances the two signals automatically, rather than relying on a fixed manually tuned reward weight.&lt;/p&gt;

  &lt;h2 id=&quot;reported-results&quot;&gt;Reported results&lt;/h2&gt;

  &lt;p&gt;The paper evaluates Data Agent on image classification, object detection, semantic segmentation, LLM instruction tuning and noisy-data settings, using several model families.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/data-centric-training-part-i/figure-03.png&quot; alt=&quot;Data Agent ImageNet and LLM instruction tuning results&quot; width=&quot;1844&quot; height=&quot;740&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Across large-scale vision and instruction-tuning settings, Data Agent improves efficiency while preserving or improving performance.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;On ImageNet-1k, the paper reports substantial training-cost reduction while maintaining or improving accuracy. The method also remains effective across ResNet, ViT, and Swin Transformer backbones. Beyond classification, Data Agent extends to MS-COCO object detection with YOLOv8 and ADE20K semantic segmentation with UperNet.&lt;/p&gt;

  &lt;p&gt;The method also applies to LLM instruction tuning. With only part of the training data, Data Agent improves performance on benchmarks such as MMLU and AlpacaEval compared with full-data training in the reported experiments.&lt;/p&gt;

  &lt;figure class=&quot;figure-panel note-figure note-figure--small&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/data-centric-training-part-i/figure-04.png&quot; alt=&quot;Data Agent detection and segmentation results&quot; width=&quot;908&quot; height=&quot;574&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;Data Agent also transfers to dense prediction tasks such as detection and segmentation.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;h2 id=&quot;what-changes&quot;&gt;What changes&lt;/h2&gt;

  &lt;p&gt;Data Agent moves selection from preprocessing into the training loop. The reported experiments test which samples to use; &lt;a href=&quot;/lab-notes/data-centric-training-part-ii/&quot;&gt;Part II&lt;/a&gt; asks whether the amount of selected data should also change over time.&lt;/p&gt;

&lt;/div&gt;</content>
    <author><name>Suorong Yang</name></author>
    <category term="Efficiency" />
    <category term="Data-centric training" />
    <media:content medium="image" url="https://declare-lab.github.io/assets/images/lab-notes/data-centric-training-part-i/figure-02.png" xmlns:media="http://search.yahoo.com/mrss/" />
  </entry><entry>
    <title type="html">δ-mem: Giving Large Language Models Lightweight, Online, and Dynamically Evolving Memory</title>
    <link href="https://declare-lab.github.io/lab-notes/delta-mem/" rel="alternate" type="text/html" title="δ-mem: Giving Large Language Models Lightweight, Online, and Dynamically Evolving Memory" />
    <published>2026-05-12T00:00:00+00:00</published>
    <updated>2026-05-12T00:00:00+00:00</updated>
    <id>https://declare-lab.github.io/lab-notes/delta-mem/</id>
    <summary type="html">δ-mem keeps a compact state during inference and uses it to modify attention in a frozen language model.</summary>
    <content type="html" xml:base="https://declare-lab.github.io/lab-notes/delta-mem/">&lt;section class=&quot;lab-note-hero&quot;&gt;
  &lt;div&gt;
    &lt;p class=&quot;work-kicker&quot;&gt;Lab note&lt;/p&gt;
    &lt;h1&gt;δ-mem: Giving Large Language Models Lightweight, Online, and Dynamically Evolving Memory&lt;/h1&gt;
    &lt;p&gt;Long-running assistants and agents need to reuse earlier information without placing their entire history in every new prompt. δ-mem tests a small state that updates during inference.&lt;/p&gt;
    

&lt;div class=&quot;note-byline&quot;&gt;
  &lt;img src=&quot;/assets/images/people/jingdi-lei.jpg&quot; alt=&quot;Jingdi Lei&quot; width=&quot;56&quot; height=&quot;56&quot; decoding=&quot;async&quot; /&gt;
  &lt;div&gt;
    &lt;p class=&quot;note-byline__name&quot; data-type-role=&quot;item-title&quot;&gt;Jingdi Lei&lt;/p&gt;
    &lt;p class=&quot;note-byline__meta&quot; data-type-role=&quot;meta&quot;&gt;May 12, 2026 · 7 min read · Efficiency&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

    &lt;div class=&quot;project-links&quot;&gt;
      &lt;a href=&quot;https://arxiv.org/abs/2605.12357&quot;&gt;Paper&lt;/a&gt;
      &lt;a href=&quot;https://github.com/declare-lab/delta-Mem&quot;&gt;GitHub&lt;/a&gt;
    &lt;/div&gt;
  &lt;/div&gt;
  &lt;figure class=&quot;figure-panel note-figure note-figure--compact&quot;&gt;
    &lt;img src=&quot;/assets/images/lab-notes/delta-mem/architecture.png&quot; alt=&quot;δ-mem architecture and write granularities&quot; width=&quot;1210&quot; height=&quot;608&quot; decoding=&quot;async&quot; fetchpriority=&quot;high&quot; /&gt;
    &lt;figcaption&gt;A compact online state supplies low-rank corrections to frozen attention layers.&lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/section&gt;

&lt;aside class=&quot;lab-note-share&quot; aria-label=&quot;Share this lab note&quot;&gt;
  &lt;span&gt;Share&lt;/span&gt;
  &lt;div&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--facebook&quot; href=&quot;https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fdelta-mem%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on Facebook&quot; title=&quot;Share on Facebook&quot;&gt;
      &lt;i class=&quot;fa-brands fa-facebook-f&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--x&quot; href=&quot;https://x.com/intent/post?text=%CE%B4-mem%3A+Giving+Large+Language+Models+Lightweight%2C+Online%2C+and+Dynamically+Evolving+Memory%20https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fdelta-mem%2F&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on X&quot; title=&quot;Share on X&quot;&gt;
      &lt;span class=&quot;lab-note-share__x-mark&quot; aria-hidden=&quot;true&quot;&gt;X&lt;/span&gt;
    &lt;/a&gt;
    &lt;a class=&quot;lab-note-share__button lab-note-share__button--linkedin&quot; href=&quot;https://www.linkedin.com/shareArticle?mini=true&amp;amp;url=https%3A%2F%2Fdeclare-lab.github.io%2Flab-notes%2Fdelta-mem%2F&amp;amp;title=%CE%B4-mem%3A+Giving+Large+Language+Models+Lightweight%2C+Online%2C+and+Dynamically+Evolving+Memory&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; aria-label=&quot;Share this lab note on LinkedIn&quot; title=&quot;Share on LinkedIn&quot;&gt;
      &lt;i class=&quot;fa-brands fa-linkedin-in&quot; aria-hidden=&quot;true&quot;&gt;&lt;/i&gt;
    &lt;/a&gt;
  &lt;/div&gt;
&lt;/aside&gt;

&lt;div class=&quot;lab-note-article&quot; data-prose-align=&quot;left&quot;&gt;

  &lt;h2 id=&quot;why-memory-is-not-just-longer-context&quot;&gt;Why memory is not just longer context&lt;/h2&gt;

  &lt;p&gt;Putting more history into a prompt lets a model see the past, but increases attention cost and does not guarantee that the relevant information will be used. The memory problem also includes deciding what to retain, updating it as new information arrives and retrieving it for the current input.&lt;/p&gt;

  &lt;p&gt;Standard Transformer attention becomes more expensive as context grows. Long contexts can also disperse attention and make relevant details harder to recover.&lt;/p&gt;

  &lt;p&gt;δ-mem therefore keeps memory outside the visible prompt and updates it while the model is running.&lt;/p&gt;

  &lt;h2 id=&quot;how--mem-works&quot;&gt;How δ-mem works&lt;/h2&gt;

  &lt;p&gt;δ-mem stores history in an Online State of Associative Memory beside a frozen full-attention backbone.&lt;/p&gt;

  &lt;p&gt;When a token or interaction segment arrives, the model projects it into a low-dimensional memory space and updates the state with a delta rule. The update is residual: it depends on the state’s prediction error rather than simply adding each new value.&lt;/p&gt;

  &lt;p&gt;The state changes online without retraining the backbone. In the main setting it is an 8 × 8 matrix.&lt;/p&gt;

  &lt;p&gt;During generation, the current input reads a signal from the previous state. That signal becomes low-rank corrections to the query and output sides of attention. Earlier information can therefore affect computation without reappearing as prompt tokens.&lt;/p&gt;

  &lt;h2 id=&quot;comparison-with-other-memory-mechanisms&quot;&gt;Comparison with other memory mechanisms&lt;/h2&gt;

  &lt;p&gt;Textual memory methods such as RAG, MemoryBank and prompt compression return memory to the context as text. δ-mem instead writes directly to a numerical state.&lt;/p&gt;

  &lt;p&gt;Unlike methods with a separate retriever, reader and fusion path, its state directly produces attention corrections.&lt;/p&gt;

  &lt;p&gt;LoRA, prefix tuning and model editing usually produce parameter changes that are fixed after training or editing. δ-mem also uses a low-rank interface, but its corrections depend on the current memory state and can change from one history to another.&lt;/p&gt;

  &lt;h2 id=&quot;three-write-schedules&quot;&gt;Three write schedules&lt;/h2&gt;

  &lt;p&gt;The paper studies three write granularities.&lt;/p&gt;

  &lt;ul&gt;
    &lt;li&gt;&lt;strong&gt;Token-State Write (TSW):&lt;/strong&gt; the state is updated at every token. This captures fine-grained information changes, but it can be sensitive to formatting symbols, repeated expressions, and local noise.&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;Sequence-State Write (SSW):&lt;/strong&gt; the hidden states of a message or segment are averaged and written once per segment. This reduces redundant writes and makes state evolution smoother.&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;Multi-State Write (MSW):&lt;/strong&gt; multiple parallel memory states are maintained, allowing different states to carry different information types and reducing overwriting within a single state.&lt;/li&gt;
  &lt;/ul&gt;

  &lt;h2 id=&quot;reported-results&quot;&gt;Reported results&lt;/h2&gt;

  &lt;figure class=&quot;figure-panel note-figure&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/delta-mem/results-table.png&quot; alt=&quot;δ-mem benchmark results table&quot; width=&quot;1256&quot; height=&quot;432&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;δ-mem improves memory-heavy and general benchmarks over the frozen Qwen3-4B-Instruct backbone and memory baselines.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;On Qwen3-4B-Instruct, TSW raises the reported overall average from 46.79% for the frozen backbone to 51.66%. Its average is 6.76 points above Context2LoRA in the same table.&lt;/p&gt;

  &lt;p&gt;The gains are especially visible on memory-intensive evaluations. On MemoryAgentBench, δ-mem increases the average score from 29.54% to a maximum of 38.85%. On LoCoMo, MSW reaches the highest score of 49.12%. On HotpotQA, TSW improves EM/F1 from 42.35% / 56.00% to 49.41% / 63.66%.&lt;/p&gt;

  &lt;h2 id=&quot;the-no-context-test&quot;&gt;The no-context test&lt;/h2&gt;

  &lt;figure class=&quot;figure-panel note-figure note-figure--small&quot;&gt;
  &lt;img src=&quot;/assets/images/lab-notes/delta-mem/no-context-recovery.png&quot; alt=&quot;δ-mem no-context recovery results&quot; width=&quot;768&quot; height=&quot;326&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; /&gt;
  &lt;figcaption&gt;In no-context recovery, δ-mem can recover part of the task-relevant signal from its compressed online state.&lt;/figcaption&gt;
&lt;/figure&gt;

  &lt;p&gt;The paper also removes the original historical context and asks the model to answer using only the compressed memory state.&lt;/p&gt;

  &lt;p&gt;In that setting, δ-mem scores above the no-context baseline on HotpotQA and LoCoMo. On HotpotQA, overall EM rises from 0.08% to 6.48% and F1 from 8.27% to 15.20%. The state recovers part, but not all, of the information lost when the context is removed.&lt;/p&gt;

  &lt;h2 id=&quot;parameter-and-memory-cost&quot;&gt;Parameter and memory cost&lt;/h2&gt;

  &lt;p&gt;In the main experiments, SSW and TSW add 4.87M trainable parameters, about 0.12% of the Qwen3-4B backbone. The multi-state MSW version uses 19.47M parameters, about 0.48%.&lt;/p&gt;

  &lt;p&gt;Reported memory use is close to the base model. Decoding is slightly slower because the state must be updated.&lt;/p&gt;

  &lt;h2 id=&quot;what-the-experiments-show&quot;&gt;What the experiments show&lt;/h2&gt;

  &lt;p&gt;The results show that a small recurrent state can retain useful information on the evaluated tasks and influence a frozen model through attention corrections. They do not establish unlimited memory or complete recovery: the no-context scores show that substantial information is still lost. Testing other backbones, longer histories and interactive tasks remains open work.&lt;/p&gt;

&lt;/div&gt;</content>
    <author><name>Jingdi Lei</name></author>
    <category term="Efficiency" />
    <category term="Memory" />
    <category term="Agents" />
    <media:content medium="image" url="https://declare-lab.github.io/assets/images/lab-notes/delta-mem/architecture.png" xmlns:media="http://search.yahoo.com/mrss/" />
  </entry>
</feed>
