
<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>DeviceMark — On-Device LLM Leaderboard changelog</title>
  <link>https://devicemark.github.io/changelog.html</link>
  <atom:link href="https://devicemark.github.io/feed.xml" rel="self" type="application/rss+xml"/>
  <description>Board updates: new models, new device measurements, methodology changes.</description>
  <language>en</language>

  <item>
    <title>New track: Anchored ops (preview) — device chores, checked deterministically</title>
    <link>https://devicemark.github.io/</link>
    <guid isPermaLink="false">devicemark-2026-07-22-anchored-ops-preview</guid>
    <pubDate>Wed, 22 Jul 2026 12:00:00 GMT</pubDate>
    <description>A new preview track measures what a pocket model can DO with device data: 48 single-turn chores in 7 skill clusters (time-zone math, bill splitting, item selection, ownership filtering, note merging, slot copying, batch multiplicity), compiled from the agent task suite and scored by deterministic checkers — exact amounts, ISO-instant equality, exact item sets; unparsable replies fail. Results (MACRO over clusters): Youtu-LLM-2B 50%, Granite-4.0-H-1B 25%, Qwen3.5-2B 24%, LFM2.5-1.2B 20%, Nanbeige4.1-3B 10%, Nemotron-3-Nano-4B 5%, Qwen3.5-0.8B 4%. The ops ranking decouples from the intelligence composite in both directions. Mac engine, greedy, cap 4096; not in the composite; iPhone runs pending.</description>
  </item>

  <item>
    <title>Gemma on its native LiteRT-LM runtime; a cap-unit bug fixed; editorial framing removed</title>
    <link>https://devicemark.github.io/</link>
    <guid isPermaLink="false">devicemark-2026-07-22-gemma-litert-capfix</guid>
    <pubDate>Wed, 22 Jul 2026 06:00:00 GMT</pubDate>
    <description>Gemma 4 E2B is now measured on its native LiteRT-LM runtime (greedy). A measurement bug was found and fixed: the LiteRT-LM harness capped output in characters (4096) while every other row capped in tokens (4096) — a ~4x tighter budget. Re-measured at a token-matched budget, Gemma's MMLU-Pro no-answer dropped 69→5 and composite rose 48.1%→52.9%; all rows now share the same 4096-token budget. Editorial "findings" removed — the board shows results and methodology only, with per-row diagnostics (answered %, median generated tokens) so a low score reads as "delivered fewer answers within budget," not "less capable." Runtime-validity: Gemma scores 85/100 on GSM8K on this runtime (matching the LiteRT int8 reference), confirming the activation path is correct.</description>
  </item>

  <item>
    <title>Two new vendors: NVIDIA Nemotron-3-Nano-4B + Nanbeige4.1-3B (9 vendors, 12 rows)</title>
    <link>https://devicemark.github.io/</link>
    <guid isPermaLink="false">devicemark-2026-07-17-nemotron-nanbeige</guid>
    <pubDate>Thu, 17 Jul 2026 06:00:00 GMT</pubDate>
    <description>NVIDIA Nemotron-3-Nano-4B (Mamba2 hybrid) and Nanbeige4.1-3B (32-layer GQA) join with full 596-item quality, float retention, and an iPhone 17 Pro token-exact gate (nat 24/24, oracle 24/24). Composite 64.9% / 63.1%, decode 14.7 / 16.9 tok/s. Both heavy reasoners (cap 4096); Nanbeige is the most extreme won't-stop-thinking model — top completed-only MMLU/MATH yet IFEval ≈1%. MiniCPM5-1B held: its Core AI export degenerates (GGUF of the same weights is clean).</description>
  </item>

  <item>
    <title>Built-in FM decode speed — honest estimate, on-device (▵ ~70–86 tok/s, iPhone 17 Pro)</title>
    <link>https://devicemark.github.io/methodology.html</link>
    <guid isPermaLink="false">devicemark-2026-07-15-fm-speed</guid>
    <pubDate>Wed, 15 Jul 2026 02:00:00 GMT</pubDate>
    <description>The FoundationModels API exposes no token counts, so the FM speed cell was n/a. Now measured with streamed chunk timing (prefill excluded, n=8 each, TTFC 0.42s): 311.4 chars/s sustained on iPhone 17 Pro (on-device) and 220.3 chars/s on M4 Max, published as ▵ ~70–86 tok/s (iPhone) via the open ports' empirical 3.61–4.43 chars/token band. Curiosity disclosed: the phone measures faster than the Mac. Raw timings in the dataset raw/.</description>
  </item>

  <item>
    <title>Cloud sea-level lines added — how far is the pocket from the data center</title>
    <link>https://devicemark.github.io/methodology.html#cloud</link>
    <guid isPermaLink="false">devicemark-2026-07-14-cloud</guid>
    <pubDate>Tue, 14 Jul 2026 10:00:00 GMT</pubDate>
    <description>Gemini Flash and Pro run the identical 596-item battery (same scorers, temp 0), both ~93% composite. Best pocket port ~72%, built-in FM 76%, cloud ceiling ~93% — the gap is the price of fitting on the phone. Drawn as reference lines; raw outputs in the dataset.</description>
  </item>

  <item>
    <title>Scoring made bit-reproducible; two upstream IFEval items disclosed</title>
    <link>https://devicemark.github.io/methodology.html#ci</link>
    <guid isPermaLink="false">devicemark-2026-07-14-determinism</guid>
    <pubDate>Tue, 14 Jul 2026 02:00:00 GMT</pubDate>
    <description>The official IFEval checkers are non-reproducible unseeded (langdetect, hash order, and upstream items 1122/1129 whose non-alphabet letters the checker replaces with a random letter). All pinned; the board now regenerates bit-identically from the published raw outputs. Values shifted at most 0.3pp, inside every CI.</description>
  </item>

  <item>
    <title>v0 launch — 8 entries, full battery, device-measured</title>
    <link>https://devicemark.github.io/</link>
    <guid isPermaLink="false">devicemark-2026-07-13-v0</guid>
    <pubDate>Sun, 13 Jul 2026 06:00:00 GMT</pubDate>
    <description>8 entries (5 open-model vendors + Apple's built-in Foundation Model), full 596-item battery, 95% CIs, 7/8 rows device-measured on iPhone 17 Pro. Findings: five-way statistical tie at the top; Google's official QAT int4 at parity on MMLU-Pro/MATH; quantization loss not monotone with size. Dataset: https://huggingface.co/datasets/devicemark/results</description>
  </item>

  <item>
    <title>Per-model badges</title>
    <link>https://devicemark.github.io/changelog.html</link>
    <guid isPermaLink="false">devicemark-2026-07-13-badges</guid>
    <pubDate>Sun, 13 Jul 2026 09:00:00 GMT</pubDate>
    <description>Embeddable per-row badges at /badge/&lt;slug&gt;.svg — intelligence composite + device tok/s, regenerated on every deploy.</description>
  </item>

</channel>
</rss>
