Thanks to visit codestin.com
Credit goes to github.com

Skip to content

Quality benchmark (the one you asked for): int4 GLM-5.2 scores 62.5% mean acc_norm — but the scoring protocol is a confound; here's the experiment that would settle it #108

Description

@dnnspaul

You flagged this as "the single most valuable thing a faster machine can contribute," so here it is: a full ./coli bench run on the int4 container.

Headline: the int4 model scores ~62.5% mean acc_norm, well below the 85–95% the README anticipates for full-precision GLM-5.2. The gap is larger than the sampling error — but I cannot attribute it to quantization yet, and I don't think anyone should until the control below is run.

Result

task                  n     acc  acc_norm
hellaswag            40   60.0%     67.5%
arc_challenge        40   62.5%     60.0%
mmlu                 40   60.0%     60.0%

MEAN acc_norm: 62.5% across 3 tasks

With Wilson 95% intervals (correct for n=40; the normal approximation is not):

task acc_norm 95% CI
hellaswag 67.5% [52.0%, 79.9%]
arc_challenge 60.0% [44.6%, 73.7%]
mmlu 60.0% [44.6%, 73.7%]
pooled (n=120) 62.5% [53.6%, 70.6%]

Random baseline is 25% (4 options), so the model is clearly working — it is not broken, just well short of published figures. The pooled upper bound (70.6%) still sits ~15 points under the low end of your 85–95% expectation, and the deficit is consistent across three independent tasks, which is harder to explain as noise.

Why I am NOT concluding "int4 costs 25 points"

Two confounds have to die first, and neither can be killed with this run alone:

1. Scoring method vs published numbers. This harness does 0-shot log-likelihood scoring (EleutherAI style). Published frontier-model scores of 85–95% are typically few-shot and/or CoT-prompted. GLM-5.2 is a reasoning model, and reasoning models are known to underperform their headline scores badly under raw log-likelihood MC scoring, because you never let them think. Some — possibly most — of this gap may be the scoring protocol, not the quantization. Notably REFERENCE in tools/eval_glm.py is still unfilled, so the harness itself has no calibrated comparison point.

2. n=40 is small. ±14pp per task. I'd want --limit 200+ before quoting any single task's number.

The decisive experiment isolates quantization from everything else: run this same harness, on the same machine, against a model small enough to evaluate at BOTH fp16 and int4. OLMoE (which colibrì already supports, tools/convert_olmoe.py) is ideal — a few GB, so both precisions fit. The delta between fp16-OLMoE and int4-OLMoE under this exact harness is the quantization cost, with the scoring protocol held constant and cancelling out. If that delta is ~2 points, then int4 is fine and the GLM gap above is a scoring artifact. If it's ~20 points, your int4 container is genuinely lossy and mixed/grouped-scale quantization becomes the priority.

I'm happy to run that A/B — say the word and I'll do it. I'd rather hand you a clean causal answer than a scary number with a confound in it.

Run details (for the record)

  • Runtime: 4h30m for 480 forwards (40 questions × 3 tasks × answer options)
  • 75.3 TB read from disk, sustained ~4.8 GB/s
  • Config: CPU reference pathCOLI_CUDA=0 MTP=0 IDOT=1, full expert routing (no --topp), RAM_GB=65, OMP tuning on
  • Deliberately not using MTP, the CUDA tier, or --topp: per Greedy decoding is not reproducible: MTP (3/5 prompts) and the CUDA expert tier (2/5) both change the output #100 all three can flip near-tie tokens, and none of them belong anywhere near an accuracy measurement
  • Machine: Ryzen 7 9800X3D · 70 GB RAM (WSL2) · Samsung 9100 PRO PCIe 5.0 (10.51 GB/s O_DIRECT) · RTX 5090 (unused here)
  • Model: mateogrgic/GLM-5.2-colibri-int4-with-int8-mtp
  • colibrì f38e74e

Two small harness papercuts, easy fixes:

  • ./coli bench needs tokenizers and datasets, which aren't mentioned in the README's quick-start pip line (torch safetensors huggingface_hub numpy). It dies with ModuleNotFoundError after downloading the datasets.
  • It prints nothing for hours — the engine gets all 480 requests in one subprocess.run and only reports at the end. On a slow disk that's a day of silence with no way to tell progress from a hang. A per-question counter to stderr would help a lot. (I resorted to reading the request file's offset out of /proc/<pid>/fdinfo to get a progress percentage.)

Relation to #100

This may be the same story from the other end. In #100 we found the int4 model sits close enough to argmax ties that any rounding change (batched MTP verification, the CUDA tier, even swapping the kernel family via IDOT=0) flips tokens, occasionally into degenerate text. A model with healthy logit margins would not be that fragile. A genuine int4 quality deficit and that tie-fragility would be two symptoms of one cause: the int4 container leaving too little margin. The OLMoE A/B above would test that directly.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    qualityQualità del modello / quantizzazione

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions