Every year, one more bit falls. We are building for the floor.
Invariant — July 2026 · London, UK · → to navigate · every cyan link opens the derivation
The bit-width for full-precision parity falls every year; weights lead, activations trail by ~two years. Ternary weights match today, binary weights ~2027–28, full binary (W1A1) by 2028–30.
Anchored on MMLU: GPT-3's 175B scored ~44 in 2020; by 2023 a 7B (Mistral) beat it; by 2024–25, 3–4B models (Phi-3, Qwen-class) clear it by 25+ points — GPT-3-class capability compressed ~50–80× in five years. GPT-4-class (~86 MMLU) went from a rumored ~1.8T MoE in 2023 to 32–70B dense in about two years.
Measured rates: capability density ×2 ≈ every 3.3 months[1] · compute for fixed capability halves ≈ every 8 months[2]. Multiply by the main trend — bits per weight halving — and capability per byte doubles every few months: the fit boundary runs toward the chip from both directions.
The research world has noticed — quietly. Quantisation titles are 4.6× on arXiv since 2019 and 6× at the top venues, accelerating every year. ICML alone: 5 → 38 since 2023.
Title search, CS-classified arXiv & 11 top venues (NeurIPS · ICML · ICLR · CVPR · ACL …), July 2026. UAI: five papers in seven years.
Our binary chip will enable far lower latency and more throughput than can be achieved in the current paradigm.
Log scale — each gridline is 10× · binary (W1A1) design point · click to switch binary ⇄ ternary
The same job, drawn as silicon. Hover any part for its cost.
Amber = supply-chain chokepoints · dashed = parts our design doesn't have.
Regions are stylised, costs are order-of-magnitude BOM estimates — full accounting one click deep. All numbers are the binary (W1A1) design point; ternary variants live in the appendix. Our prices are input costs at scale (BOM) — pricing and margin deliberately TBD; theirs are street prices. Edge shown is the value tier (mature 28 nm — always several times cheaper than a Jetson); the sub-watt N5 flagship ($399, 500 tok/s) fights Thor on the Edge duel slide. Datacenter spec → Edge spec → Cost model → Rack math → Future ladder → vs Cerebras → vs CPUs → Op energy → Token energy →
AlphaGo out-searched, not out-smarted, Lee Sedol. o1 rediscovered this trick for language. Capability is parameters × attempts[1][2]. The bottleneck is latency. A 250-attempt search takes ~18 min today; on our architecture, ~22 s.
Planning is imagining futures faster than they happen. At 11,400 tok/s per stream, a world model rolls out thousands of candidate futures per second — model-predictive control becomes model-predictive reasoning.
A 100 Hz control loop leaves 10 ms per decision. Their best stream fits ~2 tokens in that window. Ours fits ~114 — a full thought between steps, on-device, deterministic, no network.
At ~10 mJ/token on $14 silicon, thinking becomes ambient: appliances, wearables and sensors running a continuous inner monologue — offline, private, on a nightlight's power. That's the Super-Edge chip, one slide back.
You can buy the GPUs. You can't plug them in. Grid interconnection queues run years. We draw 100× less.
The heat has nowhere to go without water. At 132 kW per rack, air physically cannot remove the heat: chiller plants, cooling towers, and ~1.9 litres evaporated per kWh. Water permits and local politics now decide where AI gets built. Our water bill rounds to zero.
The GPU is mostly a memory-delivery vehicle. HBM stacks and CoWoS are sold out on multi-year allocation. Our weights live in on-die SRAM on mature nodes — no HBM, no CoWoS, nothing on allocation. The queue doesn't exist for us.
Remove the multiply · fit the model on-die · and the cheapness that follows. The far corner (★) is best — every rival maxes one edge. Drag to rotate; click a point for detail.
Analog compute (Mythic, memristor crossbars, IBM PCM). Strengths: physics does the summation — extraordinary efficiency on paper. The wall: the weights drift, the devices vary, precision stalls at ~4–6 bits, and digitising every edge costs ~58% of the power. Great physics; unshippable determinism.
Pure digital logic. Strengths: the fastest inference physically possible — ~0.25 ns per layer, no memory fetch, no adder tree, zero standby state; unbeatable for compiled boolean rules. The wall: no continuous accumulation — it has never scaled to language.
In-memory compute (analog IMC; digital IMC like d-Matrix and IBM NorthPole). Strengths: kills the von-Neumann shuttle — compute moves to the data; the digital flavours are deterministic and actually shipping. The wall: analog flavours inherit the ADC tax and the noise; digital flavours give back most of the density win — and every flavour is still accelerating multiply-accumulate. Fixes the shuttle, keeps the multiply.
Digital binary near-memory — our lane. Bits stay bits: no ADCs, no DACs, no drift, logic-grade yield on mature CMOS with V-Cache-proven bonding. The multiply isn't accelerated — it's deleted, and the model lives next to the logic, so the memory wall dies in the same move. The only approach where both assassinations succeed — deterministically, on boring silicon.
The scoreboard that matters: the largest model each approach has ever demonstrated end to end. Three of the four lines have been flat for years — at toy scale.
Only one post-multiply approach has ever trained a language model end to end — and it's already at parity: BitNet b1.58, 2B params / 4T tokens[1] · TII ships ternary Falcon-E at 1B & 3B[2]. Analog CIM's flagship: a 45M-param inference demo[3]. Logic-gate networks: CIFAR-scale[4]. NorthPole runs a 3B model — trains nothing[5].
Double first in physics and ML PhD, Oxford — spectral machine learning under Stephen Roberts. Huawei AI Theory & Amazon ML. NeurIPS / ICML / JMLR ×2.
Represented Uzbekistan at the International Mathematical Olympiad and the informatics olympiads · NUS data science · published with Diego.
Ten years as CTO of sureCore, leading its low-power SRAM and in-memory-compute IP. Technical Design Authority at Fractile. Tape-out experience to 6nm.
We solve the mathematically intractable optimal brain damage problem.
Our novel method safeguards reasoning capacity in LLMs when going from 16-bit to 3-bit.
Quantisation as a Service — live and selling. Send the checkpoint; we put precision where the model is fragile and ship in days.
Oscar Wang and Stephen Lin are closing engagements now and building partnerships in Taiwan.
16 → 3.8 bits, sold today → 1.58 → 1 in silicon.
The same binary near-memory fabric — the multiply deleted, the model resident on-die, no HBM — scaled by die size and process node from a datacenter package down to a $14 chip.
Silicon TAM — chip content only, never end-market revenue. Datacenter figure is batched peak; the rest are single-stream. Input costs are order-of-magnitude BOM at scale. Binary (W1A1) design point. Full floorplans, part by part →
The floor is published. The bet is ours. — investment@invariant.fyi
| Operation | Published silicon | ~4–5 nm estimate | vs binary |
|---|---|---|---|
| FP32 multiply-accumulate | 4.6 pJ @45nm | ~1.5 pJ | ~500× |
| FP16 multiply-accumulate | ~1.5 pJ @45nm | ~0.5 pJ | ~170× |
| INT8 multiply-accumulate | 0.2 pJ @45nm | ~0.07 pJ | ~25× |
| Ternary op (sign-gated INT8 add) | — | ~12 fJ | ~4× |
| Binary XNOR + popcount | 21.6 fJ @22nm — silicon-proven | ~3 fJ | 1× |
| 8 KB SRAM access (word) | 10 pJ @45nm | ~2 pJ | — |
| Off-chip DRAM access (word) | 1–2 nJ | HBM ~6 pJ/bit | — |
| GPU today | Invariant-1 | Ratio | |
|---|---|---|---|
| Where weights live | off-package HBM3e | on-package SRAM | — |
| Energy per bit moved | ~6 pJ | ~0.2 pJ | 30× |
| Bits per weight | 8–16 (FP8/FP16) | 1 (1.6 ternary) | 8–16× |
| Energy per weight touched | ~48–96 pJ | ~0.2 pJ | 240–480× |
| 70B model footprint | 70–140 GB — needs HBM | 8.75 GB binary / 14 GB ternary — fits in SRAM | 10–16× |
| Today (100 MW AI site) | Invariant (1.5 MW) | |
|---|---|---|
| Rack density | 132 kW — air cooling physically impossible; 100% direct liquid, no fans in the rack | 18 kW — ordinary air-cooled racks |
| Plant | chillers, CDUs, cooling towers, 25–45°C facility water loops | dry coolers — no chillers, no towers, no facility water |
| Cooling energy share | 7–40% of facility power (PUE 1.1–1.54) | ~13% of a number 65× smaller |
| Water | WUE ~1.9 L/kWh → 4.56M L/day (range 2–5M by WUE) | ≈0 L/day |
| Heat rejected | ~100 MW thermal — a small power station | ~1.5 MW — an office building's HVAC |
| Today — 100 MW GB300 site | $ |
|---|---|
| Facility: AI fit-out $20–25M/MW (liquid cooling, high-density power) | $2.0–2.5B |
| Hardware: 606 racks × ~$3.5M | $2.1B |
| Total | ~$4.5B |
| Invariant — same 667M tok/s | $ |
|---|---|
| Facility: ~2 MW ordinary shell at $10.7M/MW | ~$21M |
| Hardware: 635 chips × $12k | $7.6M |
| Hosts, network, dry coolers | ~$5M |
| Total (>100× less; we state ">50×" on the main slide) | ~$34M |
| Invariant-1 power budget | Binary @ 1.05M tok/s | W | Ternary @ 650k tok/s | W |
|---|---|---|---|---|
| MAC fabric | 1.47×10¹⁷ op/s × 3 fJ | 441 | 9.1×10¹⁶ op/s × 12 fJ | 1,092 |
| Weight SRAM reads | 95.7 TB/s × 8 × 0.2 pJ | 153 | 94.8 TB/s × 8 × 0.2 pJ | 152 |
| KV-cache reads (MLA + 4-bit) | 1.05M × 32.8 MB | 55 | 650k × 32.8 MB | 34 |
| INT8 vector unit (~1% of ops) | softmax/norm/rotary | 118 | same | 100 |
| NoC, partial sums, control | allowance | 200 | allowance | 200 |
| Leakage (power-gated) + IO | 6 dies | 100 | 9 dies | 130 |
| Total | ≈1.07 kW → 1.02 mJ/tok | 118× | ≈1.7 kW → 2.6 mJ/tok | 46× |
| Cost breakdown (volume BOM) | Binary (6 dies) | Ternary (9 dies) |
|---|---|---|
| SRAM dies, N5 HD (~$300 ea; redundancy-repairable → high yield) | $1,800 | $2,700 |
| Compute die, N4P (~500 mm², mature yields) | ~$400 | ~$400 |
| Hybrid bonding + package + test | $1,500–2,500 | $2,000–3,000 |
| Board, power delivery, LPDDR | ~$1,000 | ~$1,100 |
| BOM → sell price (≈2× BOM) | ~$5.5k → $12k | ~$7k → $16k |
| Technology | Who ships it today | Our exposure |
|---|---|---|
| Hybrid bonding, ~9 µm pitch | AMD 3D V-Cache — in production since 2022 (TSMC SoIC) | none — buying a mature flow |
| N5 high-density SRAM (0.021 µm²/cell) | every N5 product since 2020; SRAM stopped scaling at N3 — the cheap node IS the optimal node | none |
| N4P logic | mainstream mobile/DC silicon | none |
| XNOR-popcount fabric | silicon-proven at 22nm (XNE); trivially synthesizable | low — ours is bigger, not different |
| MLA + 4-bit KV | DeepSeek — shipping in production models | low — adopt, don't invent |
| W1A1 models at parity | nobody — this is the bet, and the moat | the company — this is the bet |
| GB300 NVL72 | Invariant-1 (binary) | Ratio | Invariant-1T (ternary) | |
|---|---|---|---|---|
| Throughput (70B decode) | 1,100,000 tok/s | 1,050,000 tok/s | 0.95× | 650,000 tok/s |
| Power | 132,000 W | ~1,070 W | 123× | ~1,700 W |
| Energy per token | 120 mJ | 1.02 mJ | 118× | 2.6 mJ (46×) |
| Cooling | 100% liquid + facility water | air / closed loop | — | same |
| Memory | 20.7 TB HBM3e off-package | 12 GB SRAM on-package | — | 18 GB SRAM |
| Mass | ~1,360 kg | ~2 kg | ~700× | ~2 kg |
| Price | ~$3,500,000 | ~$12,000 | ~290× | ~$16,000 |
| 100 TB/s — three independent feasibility legs |
|---|
| 1 · Throughput mode already sustains 95.7 TB/s (batch 96 × 8.75 GB × 10,937 steps/s) — no new assumption |
| 2 · Areal: 100 TB/s ÷ 4,800 mm² = 21 GB/s/mm² — AMD ships V-Cache at ~55–70 GB/s/mm². We need less than half of shipped. |
| 3 · AMD MI300X: 17 TB/s from one 256 MB layer, in production since 2023. We tile 6 far larger dies. |
| Today | Binary | Ternary | |
|---|---|---|---|
| Racks | 606 × 132 kW, liquid | 40 × 18 kW, air | 64 × 18 kW, air |
| Facility power | 100 MW | 1.1–1.5 MW | ~2.7 MW |
| Grid interconnect | years in the queue | a commercial feeder | same |
| Device | 8B decode | Power | Price | tok/J | tok/s/$k |
|---|---|---|---|---|---|
| Jetson AGX Orin 64GB — 204.8 GB/s LPDDR5 | ≤51 ceiling · ~40 real | 15–60 W | $1,999 | ~0.9 | 20 |
| Jetson AGX Thor — 273 GB/s LPDDR5X, 2,070 TFLOPS FP4 | ≤68 ceiling · ~55 real | 40–130 W | $3,499 | ~0.6 | 16 |
| Invariant-E — 1 GB W1A1 on-die SRAM, 2 TB/s | 500 spec · 2,000 ceiling | 3 W | ~$399 | ~167 | 1,250 |
| Invariant-E ternary variant — 1.6 GB trit-packed | 250 spec · 1,250 ceiling | ~2.5 W | ~$449 | ~100 | 557 |
| Invariant-E Value — one monolithic 28 nm die, 1 GB on-die | 250 spec · ~750 ceiling | ~3 W | ~$69 | ~83 | 3,600 |
| Invariant-E energy per token (batch 1) | mJ |
|---|---|
| Weights: 8×10⁹ bits × 0.3 pJ/bit (mobile stacked SRAM) | 2.4 |
| XNOR compute: 16 Gop × 10 fJ | 0.16 |
| KV (windowed + MLA, LPDDR) | ~2 |
| SoC overhead, DRAM standby | 1–2 |
| Total → 6–9 mJ/tok → 100 tok/s @ 0.9 W · 500 tok/s @ ~3 W | 6–9 |
| Buyer | Why they move |
|---|---|
| Hyperscalers / frontier labs | power-constrained on every continent; every MW we free is a MW for training. A 72× token-cost floor is not a procurement decision — it's a board-level emergency for whoever moves second. |
| Neoclouds & API resellers | their entire P&L is $/token margin. First mover resets the market price and prints margin while GPU fleets serve at a loss. |
| Sovereign AI programs | a national inference grid for $35M and 1.5 MW — no US-scale grid, no exotic supply chain, sited domestically, running open or domestic models. |
| Enterprises (on-prem) | one 18 kW rack in an ordinary server room replaces the cloud contract; data never leaves the building. Compliance departments buy this before engineers do. |
| Result | What it showed |
|---|---|
| Snell et al. 2024 arXiv 2408.03314 | compute-optimal test-time scaling lets a smaller model beat one 14× larger, FLOPs-matched (MATH); adaptive allocation is 4× more efficient than best-of-N |
| Large Language Monkeys arXiv 2407.21787 | coverage scales log-linearly over 4 orders of magnitude of samples; DeepSeek-V2-Coder: 15.9% → 56% on SWE-bench Lite at 250 samples vs 43% frontier single-attempt SOTA |
| The reasoning era o1/o3, R1 | frontier gains now come from spending more tokens at inference — the industry already concedes that tokens are intelligence |
| Date | Method | Bits/weight at ~parity |
|---|---|---|
| 2021 | FP16 deployed baseline | 16 |
| Aug 2022 | LLM.int8() | 8 |
| Oct 2022 – Jun 2023 | GPTQ · AWQ | 4 |
| Feb 2024 | QuIP# · AQLM — first Pareto-optimal <2-bit | 2 |
| Apr 2025 | BitNet b1.58 2B4T — native ternary, 4T tokens, matches FP16 peers. Open weights. | 1.58 |
| 2025 | ParetoQ (Meta) — sub-4-bit QAT scaling laws keep improving | 1.58–2 |
| Date | Method | Bits/activation |
|---|---|---|
| 2021 | FP16 baseline | 16 |
| Nov 2022 | SmoothQuant — W8A8 by migrating outlier scale into weights | 8 |
| Apr 2024 | QuaRot — rotate the network; outliers vanish; W4A4KV4 | 4 |
| 2024 | SpinQuant — learned rotations, ~2.9 pts from FP at W4A4KV4 | 4 |
| Nov 2024 | BitNet a4.8 — 4-bit activations on 1.58-bit weights, no loss vs b1.58 | 4 |
| proj. 2026 | A2 on ternary weights (rotations + native training) | 2 |
| proj. 2027–28 | A1 small-scale → W1A1 | 1 |
| Player | Why they don't build this |
|---|---|
| GPU vendors | binary-in-SRAM deletes the HBM bill of materials they monetise. You do not disrupt your own margin structure; you get disrupted out of it. |
| Frontier model labs | no silicon teams; GPU-portable roadmaps by design. BitNet is Microsoft Research — shipped as a paper, because shipping it as a product requires a chip. |
| Chip startups | can't train models. A binary chip without binary-native weights is an empty socket. (See Groq: superb silicon, borrowed models, brutal economics.) |
| Hyperscaler TPU teams | fleet homogeneity rules; a W1A1 bet breaks their fleet years before it pays. |
| Milestones — each de-risks the next |
|---|
| 1 · 8B W1A1 at parity — our own run; the scientific result that prices everything |
| 2 · FPGA prototype — the decode law measured, not modelled |
| 3 · Invariant-1 tape-out — the spec in this deck, on mature nodes, no exotic supply |
FOUNDER & LEAD SCIENTIST
Oxford MPhys, double first; ML PhD at Oxford’s MLRG under Stephen Roberts — spectral machine learning. Huawei AI Theory & Amazon ML — deployed under real product constraints, including on-device. Two JMLR papers on the geometry of training (batch-size scaling via random matrix theory, 2022; iterate averaging, 2024) the RMT × deep-learning line with Keating & Baskerville, and the Safety–Efficacy Trade-off (ICML 2026, with Ulug’bek Abdimanobov).
CO-FOUNDER & LEAD ENGINEER
Represented Uzbekistan at the International Mathematical Olympiad; informatics olympiad. National University of Singapore, data science. Co-author, Hessian Spectral Analysis at Foundation Model Scale (2026, with Diego). Algorithms & systems — and the olympiad bench: nine engineers he trained or recruited, on call.
FOUNDING ADVISOR
Fellow of the Royal Society; Fröhlich Prize. Quantum chaos and random matrix theory × the Riemann zeta function. Wills Professor at Bristol, then Sedleian Professor of Natural Philosophy at Oxford; President of the LMS; chaired the Heilbronn Institute (the GCHQ partnership) to 2020. Random matrices & spectral problems: the foundations of our compiler.
FOUNDING ENGINEER
CTO of sureCore for a decade, leading design of the company’s low-power SRAM and in-memory-compute IP. Technical Design Authority at Fractile, the Gelsinger-backed in-memory-compute inference accelerator startup. Transistor-level and custom circuit design; tape-out experience across process nodes down to 6nm.
The mathematics that beats heuristics at 3 bits is the mathematics that gets to 1. This team wrote it.
| Stage | Product | Economics |
|---|---|---|
| 0 · Now | Precision-compiler optimisation on customers' existing GPU fleets — the compiler that beats GPTQ/AWQ-class at 3-/2-bit ships as a service today | priced as a share of measured savings — revenue and design-win relationships before any silicon exists |
| 1 · Edge first | Invariant-E + SDK — OEM design wins in phones, vehicles, robotics, defence | value tier ~$69 (28 nm monolithic) to flagship $399–449 (N5, sub-watt) — both ≈2× BOM, vs $249–3,499 incumbents |
| 2 · Datacenter | Invariant-1 systems or capacity — sell 18 kW racks, or sell tokens per megawatt | capacity priced under market ($0.60–0.90/M) and >30× above our floor — ~99% gross margin on capacity at market price |
| 3 · Sovereign | Design + compiler licensing for national inference grids | mature-node manufacturable; standard export-control compliance; licensing revenue on others' capex |
CUDA sells flexible GPU programming: every precision, graphics, scientific computing, whichever model is frontier this month. That breadth is what customers pay for.
CUDA schedules multiplies. HBM streams 16-bit weights. At one bit there is nothing to schedule and nothing to stream — the flexibility has nothing left to sell.
Kernels for it are certain. But the weights still stream from HBM: best case on a GB300 rack is 457 tok/s per stream at 0.06 J/token — 25× slower, 57× more energy than ours.
Hundreds of billions in revenue, chasing $1T on $3M racks at ~75% margin. A cheaper, simpler $5.5k part is not a market they reorganise for — the innovator's dilemma, on schedule.
Their ceiling assumes perfect 2-bit packed kernels on GB300-class HBM — the best case for them.
Cerebras made our diagnosis first: weights must live in SRAM, or you die on the haul. Their fix is to grow the chip 57× — one un-diced wafer. Ours is to shrink the model 16× — one bit. Hover the parts.
Same thesis, opposite lever — and ~1,000× apart on cost per resident GB. Their scaling axis (more wafer) has stalled with SRAM itself; ours (fewer bits) halves yearly. The honest asymmetry, stated: they run every FP16 checkpoint today; we need binary-native models — the dated bet this deck prices. Cerebras is validation, not refutation: they prove customers pay a premium for SRAM-resident latency.
bitnet.cpp proved ternary models run on silicon you already own[1] — free marketing for the binary future. Then volume arrives, and the physics bill comes due. Hover the parts.
Same power, three orders of magnitude apart on delivered tokens — because the CPU’s bandwidth ceiling is the DRAM haul, and ours doesn’t exist. The framing for the room: at 16 bits models needed GPUs; at 1 bit they merely run on CPUs — and belong on silicon where the weights never move. CPUs could run graphics in 1995, too. GPUs happened anyway.
Every number below is a physics ceiling — SRAM density × stacked area, nothing else. Constant at every stop: no CoWoS, no HBM, no N3/N2 allocation — and no kernel NVIDIA can write moves any of it, because their weights still stream from off-package memory at fixed bandwidth. Models are trainable to any of these sizes.
| Configuration | Capacity | Max binary model | BOM (est.) | $ / B params |
|---|---|---|---|---|
| 9-die N5 stack — the shipping design | 18 GB | ~144B | ~$7k | ~$50 |
| 16-die tall stack — SoIC roadmaps run 8–12+ high; same bonding flow | 32 GB | ~250B | ~$9–10k | ~$40 |
| 2 packages, 2.5D stitch — only 1–2-bit activations cross | 64 GB | ~500B | ~$18–20k | ~$38 |
| 4 packages, one board | 128 GB | ~1T | ~$36–40k | ~$38 |
| Gain-cell / oxide-semiconductor memory — 2–4× SRAM density; ~2028–30 | 60–120 GB | 0.5–1T | ~$8–12k | ~$10–20 |
| CFET-era SRAM — stacked transistors resume cell scaling; 2030+ | ~50–60 GB | ~450B | ~$12–15k | ~$30 |
| Wafer-scale at N5 density — binary-native, speculative | ~115 GB | ~900B | ~$50–100k system | ~$60–110 |
Node-by-node manufacturing picture: slide the node →
Accessed July 2026. MLPerf figures labelled verified/unverified as published. Floor-to-floor comparisons throughout; every derate applied to our side.
Correct — that's the deck's stated frame, and every headline number re-prices cleanly on ternary, which does exist (open weights, 2B, FP16 parity). Bit-width at parity has halved yearly for 4 years. Our first milestone is our own 8B W1A1 — the experiment that prices the company.
It does — that's our free marketing. A ~$25k server moves ~0.6 TB/s from DRAM: ~65 tok/s single-stream on a 70B binary, ~1.5k batched at a kilowatt. Our package moves weights zero millimetres: 11,400 and 1.05M at the same power — ~400× tokens/W, ~700× $/throughput. CPUs educate the market and serve low-volume private inference; volume forces the ASIC. Chip-to-chip →
Those laws fit post-hoc/fixed-budget quantisation; native training keeps bending them (BitNet 2B4T, ParetoQ). And even a 2× effective-param tax leaves ≥55× energy and ~25× latency.
On-package: MLA-latent + 4-bit, ~3.1 GB at batch 96 × 2k context. Long context pages to LPDDR at a disclosed ~2× derate at 32k. It's a budgeted line item with its own power row.
It's 21 GB/s/mm² areal across 4,800 mm² of bonded SRAM — less than half of what AMD ships in V-Cache today. MI300X pulls 17 TB/s from one 256 MB layer. Throughput mode already needs 96 TB/s, so latency mode adds no new assumption.
At FP8/16 — needing ~576 chips or a 23 kW wafer for 70B. They proved the latency market and the economics of not having small weights. Binary is 16× denser: one 1 kW package. They validated the demand; compression changes the answer.
The honest risk — it has its own card in the Full Stack slide. The win isn't the MAC: it's weights-in-SRAM (deletes the HBM economics they monetise), binary-native frontier weights (they don't have), and a compiler. Clean-sheet redesign + retraining = a window measured in years.
Disclosed on-slide — and it favours them. Offline is their best case; deployed serving runs ~18× worse than the floor we grant them. We compare against their press release, not their reality.
Replacing a 100 MW site = 635 chips ≈ ~60 N5 wafers + SoIC capacity — noise against GPU volumes. No CoWoS, no HBM allocation. SRAM dies are redundancy-repairable: high yield on a 2020-era node.
BitNet is Microsoft Research — as a paper. Shipping it needs clean-sheet silicon + native pretraining + a compiler, against every incumbent's GPU-portability roadmap. Vertically-integrated fabless labs are the historically successful shape for exactly this move.
Two paths. Path A: native QAT with straight-through estimators — how ternary reached parity; we extend that toolchain. Path B: skip backprop entirely — evolution strategies (EGGROLL, 2025) train with forward passes only, in pure integer arithmetic, population-averaging away the noise STE fakes, at ~91% of inference throughput. Path B's workload is millions of cheap forward passes — the exact thing our silicon does 46–118× cheaper. If B scales, the training market lands on our chip too. (Honestly: demonstrated at 1.5B-class fine-tuning, not 70B pretraining.)
Correct, and the plan is shaped around it: FPGA-first retires the physics risk before the first big cheque; physical design runs with an established design-services partner (the standard fabless path); first silicon hires are a lead SoC architect, an SRAM/memory designer and a packaging lead. The team today is the part nobody can hire: the mathematics and the compiler.
The compiler. It already beats GPTQ/AWQ-class heuristics at 3-/2-bit and ships as an optimisation service on customers' existing GPU fleets, priced as a share of measured savings — cash and design-win relationships years before tape-out. Then edge silicon, then datacenter capacity. Full sequencing in the Business Model appendix.
Scale-out, not scale-up: only 1–2-bit activations ever cross a package boundary, so a multi-chip fabric is cheap wires, not HBM — the 2-die 2.5D stitch is the proven pattern. And MoE helps us: more parameters per FLOP is exactly what cheap on-die capacity rewards; experts partition naturally one-per-die.
Distillation shrinks the parameter count; it does not change the format. A distilled 8B at FP16 still pays 16× the memory physics of the same model binarised. The two compound — distill, then drop the bits — they do not compete. Our floor applies to whatever size the distillers produce.
No — and the deck says so: parity is proven at 2B (open weights); 70B is the ~2027 point on a four-year trend. What is already fact in ternary mode is the hardware: the format's energy, cost and capacity numbers. The dated claim is model quality at scale, and we ship datacenter capacity only when open evals pass.
We are a fabless company on mature and N5-class nodes with standard export-control compliance; sovereign deals are licensed designs manufactured where lawful — a regulated go-to-market lane, not a gray market. The attraction for the buyer is that nothing in our supply chain is on the EUV frontier.
Serving is an OpenAI-compatible API; edge is an SDK; training stays PyTorch. Binary is an inference-format problem we own end-to-end — not a CUDA-replacement problem. Nobody rewrote code for FP8 either.
If a question isn't here, we want it — investment@invariant.fyi