Invariant · Technical Evidence · 18 August 2026
Our 3-bit model beats their 4-bit model.
Qwen3-0.6B PTQ vs. stock GPTQModel 7.1.0 · quant_report.pdf
32.4%
our 3-bit beats their 4-bit (29.6%)
0.4%
GPTQModel 3-bit — collapses
~10M
tokens — vs. 35,000×+ for QAT routes
Invariant (ours)
GPTQModel (stock)
fp16 reference
GSM8K · think mode
higher is better
GSM8K · nothink mode
higher is better
WikiText-2 perplexity
lower is better
KL divergence to fp16
lower is better
GPTQModel collapses at 3-bit — our 3-bit build stays intact and beats their 4-bit build outright.
All figures, group size 256
| Build | Perplexity ↓ | KL ↓ | GSM8K think ↑ | GSM8K nothink ↑ |
| fp16 teacher | 20.96 | — | 59.6% | 34.8% |
| Invariant, 4-bit | 25.34 | 0.252 | 58.0% | 36.0% |
| Invariant, 3-bit | 31.37 | 0.501 | 32.4% | 26.4% |
| GPTQModel, 4-bit | 32.40 | 0.486 | 29.6% | 7.2% |
| GPTQModel, 3-bit | 146.57 | 2.008 | 0.4% | 1.2% |
Route to a low-bit model — tokens spent
| Route | Tokens | vs. Invariant |
| Invariant PTQ (this note) | ~10M | 1× |
| BitCPM4 QAT phase (ModelBest) | 350B | 35,000× |
| BitNet ternary, from scratch | 4T | 400,000× |
| BitCPM4 end to end (pretrain + QAT) | 8.35T | 835,000× |
The correction holds across the recipe sweep
Production floats above the control everywhere — every point on the grid, both modes.
uniform-grid quantizer + correction (control)
bit-plane quantizer + correction (production)
GSM8K think %
each point: one regularization × correction-damping setting
GSM8K nothink %
each point: one regularization × correction-damping setting