Invariant · Technical Evidence · 18 August 2026

Our 3-bit model beats their 4-bit model.

Qwen3-0.6B PTQ vs. stock GPTQModel 7.1.0 · quant_report.pdf

32.4%
our 3-bit beats their 4-bit (29.6%)
0.4%
GPTQModel 3-bit — collapses
~10M
tokens — vs. 35,000×+ for QAT routes
Invariant (ours) GPTQModel (stock) fp16 reference

GSM8K · think mode

higher is better
fp16 59.6
58.0
Invariant
4-bit
32.4
Invariant
3-bit
29.6
GPTQModel
4-bit
0.4
GPTQModel
3-bit

GSM8K · nothink mode

higher is better
fp16 34.8
36.0
Invariant
4-bit
26.4
Invariant
3-bit
7.2
GPTQModel
4-bit
1.2
GPTQModel
3-bit

WikiText-2 perplexity

lower is better
fp16 20.96
25.3
Invariant
4-bit
31.4
Invariant
3-bit
32.4
GPTQModel
4-bit
146.6
GPTQModel
3-bit

KL divergence to fp16

lower is better
0.252
Invariant
4-bit
0.501
Invariant
3-bit
0.486
GPTQModel
4-bit
2.008
GPTQModel
3-bit

GPTQModel collapses at 3-bit — our 3-bit build stays intact and beats their 4-bit build outright.

All figures, group size 256
BuildPerplexity ↓KL ↓GSM8K think ↑GSM8K nothink ↑
fp16 teacher20.96—59.6%34.8%
Invariant, 4-bit25.340.25258.0%36.0%
Invariant, 3-bit31.370.50132.4%26.4%
GPTQModel, 4-bit32.400.48629.6%7.2%
GPTQModel, 3-bit146.572.0080.4%1.2%
Route to a low-bit model — tokens spent
RouteTokensvs. Invariant
Invariant PTQ (this note)~10M1×
BitCPM4 QAT phase (ModelBest)350B35,000×
BitNet ternary, from scratch4T400,000×
BitCPM4 end to end (pretrain + QAT)8.35T835,000×

The correction holds across the recipe sweep

Production floats above the control everywhere — every point on the grid, both modes.

uniform-grid quantizer + correction (control) bit-plane quantizer + correction (production)

GSM8K think %

each point: one regularization × correction-damping setting

GSM8K nothink %

each point: one regularization × correction-damping setting

Qwen3-0.6B · GPTQModel 7.1.0 · group 256 · GSM8K n=250