01 / 09
Binary Intelligence Lab

Binary
Intelligence

Every year, one more bit falls. We are building for the floor.

Invariant — July 2026 · London, UK · → to navigate · every cyan link opens the derivation

The Trend nobody is talking about.

High performance ≠ High precision arithmetic.

The bit-width for full-precision parity falls every year; weights lead, activations trail by ~two years. Ternary weights match today, binary weights ~2027–28, full binary (W1A1) by 2028–30.

Bits at full-precision parity · log₂ scale · ▬ weights  ▬ activations · solid = published · dashed = projected
Appendix · The Field

Low-bit research is accelerating.

The research world has noticed — quietly. Quantisation titles are 4.6× on arXiv since 2019 and 6× at the top venues, accelerating every year. ICML alone: 5 → 38 since 2023.

Quantisation papers · ▬ arXiv  ▬ top-tier conferences

Title search, CS-classified arXiv & 11 top venues (NeurIPS · ICML · ICLR · CVPR · ACL …), July 2026. UAI: five papers in seven years.

The Numbers

Faster and more throughput.

Our binary chip will enable far lower latency and more throughput than can be achieved in the current paradigm.

1M 100k 10k 1k 100
227
50× faster11,400
15,200
69× more1.05M
Single-stream speed
one user · batch = 1 · tok/s
Peak throughput
fully batched · tok/s
NVIDIA B300 · one GPU package Invariant-1 · one package, 70B on-die

Log scale — each gridline is 10× · binary (W1A1) design point · click to switch binary ⇄ ternary

Technical Appendix · The Chip Designs

What will the new chip designs look like?

The same job, drawn as silicon. Hover any part for its cost.
Amber = supply-chain chokepoints · dashed = parts our design doesn't have.

NVIDIA B300 · one GPU package (rack share)
LIQUID COLD PLATE~$400 — mandatory at 1,400 W. Ties the GPU to a chilled-water facility.
HBM~$5,500 for the 8 stacks — the sold-out commodity, ~$15–20/GB on multi-year allocation. Exists only because the model doesn't fit on the die.
HBM~$5,500 for the 8 stacks — the sold-out commodity, ~$15–20/GB on multi-year allocation. Exists only because the model doesn't fit on the die.
HBM~$5,500 for the 8 stacks — the sold-out commodity, ~$15–20/GB on multi-year allocation. Exists only because the model doesn't fit on the die.
HBM~$5,500 for the 8 stacks — the sold-out commodity, ~$15–20/GB on multi-year allocation. Exists only because the model doesn't fit on the die.
HBM~$5,500 for the 8 stacks — the sold-out commodity on multi-year allocation.
HBM~$5,500 for the 8 stacks — the sold-out commodity on multi-year allocation.
HBM~$5,500 for the 8 stacks — the sold-out commodity on multi-year allocation.
HBM~$5,500 for the 8 stacks — the sold-out commodity on multi-year allocation.
COMPUTE DIE 1
N4 · TENSOR CORES
~$2,500 for the pair — the world's best multiply-accumulate arrays. The only part about intelligence.
COMPUTE DIE 2~$2,500 for the pair — reticle-limit N4 silicon.
CoWoS-L SILICON INTERPOSER~$1,200 — the other chokepoint: TSMC CoWoS capacity, not wafers, caps world GPU supply.
12-LAYER SUBSTRATE · VRM FOR 1.4 kW · NVLINK~$1,500 — power delivery scales with the wattage.
VS
Invariant-1 · one package · 70B loaded — holds 96B
AIR HEATSINK~$40 — ~1 kW over 4,800 mm² of package: a big fan, not a water plant.
COMPUTE DIE
N4P · XNOR FABRIC
~$400 — multipliers deleted: sign-match gates, popcount trees, integer thresholds. The binary MACs are <20 mm² of this die.
6× SRAM DIES — 12 GB SRAM, 6 stacked dies
~$1,800 dies + ~$800 hybrid bonding — the whole model lives here, bonded on the logic (AMD V-Cache flow, shipping since 2022).
HBM:
NONE
−$5,500 — deleted. The model is on-die; there is nothing to stream.
NO SILICON INTERPOSER−$1,200 — deleted. No HBM → no CoWoS → no allocation queue.
ORGANIC SUBSTRATE · BOARD · LPDDR SIDE-PORT~$1,200 — commodity packaging; LPDDR is long-context paging only, never the token path.

Regions are stylised, costs are order-of-magnitude BOM estimates — full accounting one click deep. All numbers are the binary (W1A1) design point; ternary variants live in the appendix. Our prices are input costs at scale (BOM) — pricing and margin deliberately TBD; theirs are street prices. Edge shown is the value tier (mature 28 nm — always several times cheaper than a Jetson); the sub-watt N5 flagship ($399, 500 tok/s) fights Thor on the Edge duel slide. Datacenter spec → Edge spec → Cost model → Rack math → Future ladder → vs Cerebras → vs CPUs → Op energy → Token energy →

Deep Research · Test-Time Search

Instant deep research.

AlphaGo out-searched, not out-smarted, Lee Sedol. o1 rediscovered this trick for language. Capability is parameters × attempts[1][2]. The bottleneck is latency. A 250-attempt search takes ~18 min today; on our architecture, ~22 s.

SWE-bench Lite — exact results · Brown et al., 2024 (arXiv:2407.21787)
The Payoff · AGI Bottlenecks

Solving AGI bottlenecks.

You can buy the GPUs. You can't plug them in. Grid interconnection queues run years. We draw 100× less.

Watts per rack — NVIDIA's own roadmap
Rack power, 2022→2027+15×
Grid interconnection queueyears
The Plan · Staged Execution

Roadmap, staging & what each step costs

QaaS in customer trials today; edge silicon and datacenter silicon follow in sequence.
0612182430 · Series A3642
1 · Softwaremo 0–6
W3A8 live & shipping · W2A4 in progress · 3-bit PTQ →
2 · Chip Architecturemo 6–12
matmul W2A4 complete · live transformer demo in progress
3 · Chip designmo 12–18
not started
4 · Edge Siliconmo 18–30
not started
5 · Datacenter Siliconpost-Series A
not started
Software & licensing Pre-silicon R&D Manufacturing & bring-up Post-Series A scale
The Field · Competitor Map

Two levers, one payoff — one corner.

Remove the multiply · fit the model on-die · and the cheapness that follows. The far corner (★) is best — every rival maxes one edge. Drag to rotate; click a point for detail.

◐ drag · click a point
Appendix · The Field

Competing approaches — in detail.

Analog compute (Mythic, memristor crossbars, IBM PCM). Strengths: physics does the summation — extraordinary efficiency on paper. The wall: the weights drift, the devices vary, precision stalls at ~4–6 bits, and digitising every edge costs ~58% of the power. Great physics; unshippable determinism.

Analog crossbar — where the power goes
DAC
DRIVERS
Row drivers — every input needs digital→analog conversion before the array can use it.
MEMRISTIVE CROSSBAR
CONDUCTANCE = WEIGHT
The physics is the pitch: summation for free by Kirchhoff's law. The physics is also the problem: conductance drifts with temperature and age, device variability corrupts weights, and precision tops out ~4–6 bits (Sebastian et al., Nat. Nanotech 2020).
CAL / TRIM
LOOPS
Forever-calibration — compensating drift is a permanent tax, not a one-off.
ADC READOUT ROWThe chokepoint: every column's analog sum must be digitised. ADCs ≈ 58% of power and ~31% of area in the canonical design (ISAAC, ISCA 2016).
The People · The Proof

The Team.

Diego Granziol
Diego Granziol
Founder

Double first in physics and ML PhD, Oxford — spectral machine learning under Stephen Roberts. Huawei AI Theory & Amazon ML. NeurIPS / ICML / JMLR ×2.

Khurshid Juraev
Khurshid Juraev
Co-founder

Represented Uzbekistan at the International Mathematical Olympiad and the informatics olympiads · NUS data science · published with Diego.

Tony Stansfield
Tony Stansfield
Founding Engineer

Ten years as CTO of sureCore, leading its low-power SRAM and in-memory-compute IP. Technical Design Authority at Fractile. Tape-out experience to 6nm.

Jon Keating
Jon KeatingFounding Advisor · Fellow of the Royal Society · Sedleian Professor, Oxford
Appendix · The Technique

Already a technique.

brain
FP16 Invariant SOTA
The hard problem

We solve the mathematically intractable optimal brain damage problem.

Our novel method safeguards reasoning capacity in LLMs when going from 16-bit to 3-bit.

Appendix · Traction

Team with traction.

Quantisation as a Service — live and selling. Send the checkpoint; we put precision where the model is fragile and ship in days.

3.8
average bits per weight, in production
94%
original quality retained
2.7×
memory reduction
1.8×
faster inference, same silicon

Oscar Wang and Stephen Lin are closing engagements now and building partnerships in Taiwan.

16 → 3.8 bits, sold today → 1.58 → 1 in silicon.

Appendix · The Product Line

One architecture. Four markets.

The same binary near-memory fabric — the multiply deleted, the model resident on-die, no HBM — scaled by die size and process node from a datacenter package down to a $14 chip.

The shared architecture XNOR + popcount fabric weights bonded on-die (SRAM) no HBM, no interposer mature CMOS, logic-grade yield
Datacenter
Invariant-1
one package · vs NVIDIA B300
70Bmodel on-die
NodeN4P · 5 nm
Peak throughput1.05M tok/s
Power1.1 kW
Input cost~$5.5k
Power user
Invariant-W
desktop card · vs RTX 5090
30Bmodel on-die
NodeN4P · 5 nm
Single-stream13,300 tok/s
Power≤150 W
Input cost~$1,350
Edge
Invariant-E
one chip · vs Jetson
8Bmodel on-die
Node28 nm
Single-stream250 tok/s
Power~3 W
Input cost~$34
Super-edge
Invariant-SE
one die · vs cloud API
3Bmodel on-die
Node45 nm
Single-stream~30 tok/s
Power~0.3 W
Input cost~$14
Total addressable market $500B → $1T/yr datacenter silicon by 2030 $50–100B/yr consumer & edge · 1.5B devices/yr sovereign inference grids · no-EUV chain new always-on categories — created, not taken all at a 72× cost floor

Silicon TAM — chip content only, never end-market revenue. Datacenter figure is batched peak; the rest are single-stream. Input costs are order-of-magnitude BOM at scale. Binary (W1A1) design point. Full floorplans, part by part →

Invariant · Binary Intelligence Lab

The future is
binary

The floor is published. The bet is ours. — investment@invariant.fyi

Technical Appendix · Physics

Energy per operation

OperationPublished silicon~4–5 nm estimatevs binary
FP32 multiply-accumulate4.6 pJ @45nm~1.5 pJ~500×
FP16 multiply-accumulate~1.5 pJ @45nm~0.5 pJ~170×
INT8 multiply-accumulate0.2 pJ @45nm~0.07 pJ~25×
Ternary op (sign-gated INT8 add)—~12 fJ~4×
Binary XNOR + popcount21.6 fJ @22nm — silicon-proven~3 fJ1×
8 KB SRAM access (word)10 pJ @45nm~2 pJ—
Off-chip DRAM access (word)1–2 nJHBM ~6 pJ/bit—
A binary op is one XNOR gate and a counter. No exponent logic, no mantissa alignment, no carry-propagate array.
The op gets so cheap that the budget moves to wires and memory — which the memory-wall math then guts.

Assumptions

  • 21.6 fJ/op @22nm is measured, synthesizable digital silicon (XNOR Neural Engine) — not a projection
  • 22nm → 4–5nm scaling of ~5–7× (V² + capacitance across ~4 nodes)
  • 3 fJ covers XNOR + count + operand latch; weight SRAM reads are a separate budget line — no double counting
  • ternary op = zero-gated add of an INT8 activation, ~12 fJ incl. control
  • 45nm figures: canonical Horowitz ISSCC 2014
Technical Appendix · The Three Walls

Power, water, memory — in detail

Decoding a token touches every weight in the model. Energy = (bits per weight) × (energy per bit moved):
GPU todayInvariant-1Ratio
Where weights liveoff-package HBM3eon-package SRAM—
Energy per bit moved~6 pJ~0.2 pJ30×
Bits per weight8–16 (FP8/FP16)1 (1.6 ternary)8–16×
Energy per weight touched~48–96 pJ~0.2 pJ240–480×
70B model footprint70–140 GB — needs HBM8.75 GB binary / 14 GB ternary — fits in SRAM10–16×
The footprint collapse changes the architecture, not just the coefficient: at 8.75 GB the model fits in bonded SRAM, so HBM is deleted from the bill of materials — its energy, its cost, and its supply constraint. NVIDIA's margin lives in that line item. Ours doesn't.

Assumptions

  • HBM ~6 pJ/bit (≈200 pJ per 32-bit word incl. PHY + controller)
  • stacked-SRAM read ~0.2 pJ/bit incl. array + hybrid-bond vertical hop
  • batch amortisation applies equally to both sides — ratios are batch-invariant
  • 70B dense reference model throughout
Technical Appendix · Energy

Energy per token — both sides, line by line

// Baseline — best published GPU run (Llama-2-70B, offline)
GB300 NVL72: 132,000 W ÷ 1,100,000 tok/s = 0.120 J/token (IT power)
× PUE 1.25 = 0.150 J/token at the meter
// check 1: prior gen GB200 = 120 kW ÷ 865k = 0.139 J — Blackwell Ultra improved 16% ✓
// check 2: H100 measured 0.39 J/tok — consistent generational curve ✓
// check 3: deployed frontier serving ≈0.31 Wh/query ≈ 2.2 J/tok — 18× worse than the floor we grant them
Invariant-1 power budgetBinary @ 1.05M tok/sWTernary @ 650k tok/sW
MAC fabric1.47×10¹⁷ op/s × 3 fJ4419.1×10¹⁶ op/s × 12 fJ1,092
Weight SRAM reads95.7 TB/s × 8 × 0.2 pJ15394.8 TB/s × 8 × 0.2 pJ152
KV-cache reads (MLA + 4-bit)1.05M × 32.8 MB55650k × 32.8 MB34
INT8 vector unit (~1% of ops)softmax/norm/rotary118same100
NoC, partial sums, controlallowance200allowance200
Leakage (power-gated) + IO6 dies1009 dies130
Total≈1.07 kW → 1.02 mJ/tok118×≈1.7 kW → 2.6 mJ/tok46×
Binary: 0.120 ÷ 0.00102 = 118×. Ternary — today's published science, nothing speculative — still lands 46×. Flip the toggle on any slide: the story survives either way.

Assumptions

  • 1.1M tok/s is offline + unverified MLPerf v5.1 (Azure, observed by Signal65) — the incumbent's best foot, granted
  • rack at 132 kW — the bottom of Supermicro's 132–140 kW operating range: again favours them
  • binary op 3 fJ, ternary op 12 fJ, stacked-SRAM 0.2 pJ/bit (op-energy table)
  • decode batch 96, avg context 2,048, MLA + 4-bit KV
  • softmax/norm/residual INT8 — standard in BitNet-class models
  • Invariant-1 is specified, not taped out
Technical Appendix · Silicon

Invariant-1 — diagram, nodes, cost, maturity

Package — plan view (one organic substrate, no interposer, no CoWoS)
Compute die · ~500 mm² · TSMC N4PXNOR fabric 126 P-op/s (<20 mm² of it) · NoC · INT8 vector unit · SRAM PHY
3D SRAM · 6×800 mm² · TSMC N5 HD12 GB @ 2 GB/die · hybrid-bonded stack · 100 TB/s sustained
Side ports2× LPDDR5X (long-context paging) · PCIe/CXL host
Heatsink / closed-loop cold plate — ~1.1 kW package
SRAM dies ×6 (N5) — hybrid bond, ~9 µm pitch, sub-pJ/bit vertical hops
Compute die (N4P) — placed edge-adjacent (SoIC-X style) so logic heat never crosses the full stack
Organic substrate — commodity packaging, no silicon interposer
// the decode law — one line prices every claim in this deck
tok/s = batch × bandwidth ÷ weight-bytes
// their rack: 72 × 8 TB/s ÷ 35 GB (FP4) × batch/eff ≈ 1.10M ✓
// our chip (binary): 96 × 100 TB/s ÷ 8.75 GB = 1.097M → spec 1.05M tok/s
// our chip (ternary, trit-packed 1.6 b/w): 96 × 100 ÷ 14 GB = 686k → spec 650k tok/s
Cost breakdown (volume BOM)Binary (6 dies)Ternary (9 dies)
SRAM dies, N5 HD (~$300 ea; redundancy-repairable → high yield)$1,800$2,700
Compute die, N4P (~500 mm², mature yields)~$400~$400
Hybrid bonding + package + test$1,500–2,500$2,000–3,000
Board, power delivery, LPDDR~$1,000~$1,100
BOM → sell price (≈2× BOM)~$5.5k → $12k~$7k → $16k
TechnologyWho ships it todayOur exposure
Hybrid bonding, ~9 µm pitchAMD 3D V-Cache — in production since 2022 (TSMC SoIC)none — buying a mature flow
N5 high-density SRAM (0.021 µm²/cell)every N5 product since 2020; SRAM stopped scaling at N3 — the cheap node IS the optimal nodenone
N4P logicmainstream mobile/DC siliconnone
XNOR-popcount fabricsilicon-proven at 22nm (XNE); trivially synthesizablelow — ours is bigger, not different
MLA + 4-bit KVDeepSeek — shipping in production modelslow — adopt, don't invent
W1A1 models at paritynobody — this is the bet, and the moatthe company — this is the bet
Read that table again: every ingredient except the model is bought off the shelf. No bleeding-edge node, no CoWoS queue, no HBM allocation fight. The only hard thing is the thing we're the lab for.

Memory budget

  • binary: 8.75 GB weights + 3.1 GB KV (batch 96 × 2k ctx × ~16 KB MLA-latent 4-bit) + 0.15 workspace = 12 GB / 6 dies
  • ternary: 14 GB weights (trit-packed, 1.6 b/w — BitNet.cpp's 5-trits-per-byte, shipping) + KV → 18 GB / 9 dies
  • raw GQA-70B KV is 82 KB/tok at 4-bit → MLA ≈4–5× compression is an explicit, cited assumption
  • long context pages to LPDDR at a disclosed ~2× derate at 32k — we never claim "all-SRAM" past ~4k average
  • >70B and MoE: multi-package scale-out — only 1–2-bit activations cross packages, so the fabric is cheap wires; 2-die 2.5D stitch is the proven pattern, and MoE's params-per-FLOP favours on-die capacity (experts partition one-per-die)
  • stack yield: assembled from known-good dies; SRAM redundancy repairs post-bond — the V-Cache flow, in consumer production since 2022

Thermal & risk, stated

  • ~1.1 kW package: compute die sits beside the stack, not under it — logic heat never crosses 6 SRAM layers
  • SRAM dies dissipate ~mW/mm² — stacking memory is thermally boring; stacking logic is not, and we don't
  • closed-loop cold plate is the fallback; the claim is no facility water, no chillers
  • Invariant-1 is specified, not taped out — printed on every chip slide on purpose
Technical Appendix · Head-to-Head

One chip vs one GB300 NVL72

GB300 NVL72Invariant-1 (binary)RatioInvariant-1T (ternary)
Throughput (70B decode)1,100,000 tok/s1,050,000 tok/s0.95×650,000 tok/s
Power132,000 W~1,070 W123×~1,700 W
Energy per token120 mJ1.02 mJ118×2.6 mJ (46×)
Cooling100% liquid + facility waterair / closed loop—same
Memory20.7 TB HBM3e off-package12 GB SRAM on-package—18 GB SRAM
Mass~1,360 kg~2 kg~700×~2 kg
Price~$3,500,000~$12,000~290×~$16,000
// per-GPU sanity check, same law
1.1M ÷ 72 = 15,278 tok/s per B300 at 1,833 W of rack share = 0.120 J/tok — same answer ✓
// ternary framing: one 1.7 kW ternary chip ≈ 42 B300 GPUs' worth of decode — on published science alone.

Assumptions

  • 1.1M tok/s: Llama-2-70B, offline, unverified MLPerf v5.1 (Azure ND GB300 v6, FP4, TensorRT-LLM) — their best possible number, granted in full
  • 132 kW = bottom of the published operating range (132–140)
  • decode-dominated serving; prefill-heavy mixes narrow the gap — covered by the 3× derate in the cost model
  • Invariant-1: specified, not taped out
Technical Appendix · Head-to-Head

Single-stream latency — the physics of interactivity

// batch-1 decode is weight-bandwidth-bound on every architecture, theirs included
B300, granted FP4 (its best case): 35 GB ÷ 8 TB/s = 4.375 ms → 229 tok/s
B300 at FP8: 70 GB ÷ 8 TB/s = 8.75 ms → 114 tok/s
Invariant-1 binary: 8.75 GB ÷ 100 TB/s = 87.5 µs → 11,400 tok/s (+1.1 µs compute)
Invariant-1 ternary: 14 GB ÷ 100 TB/s = 140 µs → 7,100 tok/s
// binary = 50× their best case · ternary = 31× — on published science
What 11,400 tok/s single-stream buys: a 10,000-token chain of thought in under a second. Reasoning that feels instant. Voice with zero dead air. Robot control loops inside the model. Test-time search in real time.
100 TB/s — three independent feasibility legs
1 · Throughput mode already sustains 95.7 TB/s (batch 96 × 8.75 GB × 10,937 steps/s) — no new assumption
2 · Areal: 100 TB/s ÷ 4,800 mm² = 21 GB/s/mm² — AMD ships V-Cache at ~55–70 GB/s/mm². We need less than half of shipped.
3 · AMD MI300X: 17 TB/s from one 256 MB layer, in production since 2023. We tile 6 far larger dies.

Groq / Cerebras — the obvious objection, answered

  • they proved SRAM inference wins latency — at FP8/16
  • 70B at those widths needs ~576 Groq chips or a ~23 kW wafer
  • binary is 16× denser: the same idea collapses into one 1 kW package — they validated the market; compression changes the answer

Assumptions

  • weights re-read once per token at batch 1 — true of all decode
  • 100 TB/s sustained / 120 peak across 6 bonded dies
  • B300 batch-1 figures are theoretical BW ceilings — real GPUs land below them

Sources

AMD V-Cache 2+ TB/s MI300X 17 TB/s B300 8 TB/s HBM3e Cerebras, chip-to-chip → CPUs & bitnet.cpp →
Technical Appendix · Datacenter

The 100 MW walk — iso-throughput

// today's 100 MW AI site, GB300 fleet
100 MW ÷ PUE 1.25 = 80 MW IT ÷ 132 kW/rack = 606 racks × 1.1M = 667M tok/s

// the same tokens on Invariant-1 (binary)
667M ÷ 1.05M = 635 chips × 1.15 kW (chip + fan share) = 0.73 MW
+ 28% hosts & network = 0.93 MW IT × PUE 1.15 (air) = ~1.1 MW best estimate
// conservative: chips at 850k, +40% overhead, PUE 1.2 → 1.5 MW — the headline survives the miss case

// ternary (published science): 667M ÷ 650k = 1,026 chips × 1.8 kW → ≈2.7 MW total = 37×
TodayBinaryTernary
Racks606 × 132 kW, liquid40 × 18 kW, air64 × 18 kW, air
Facility power100 MW1.1–1.5 MW~2.7 MW
Grid interconnectyears in the queuea commercial feedersame
// reconciling 118× (chip physics) with 65× (facility): hosts, network, PUE and deliberate rounding spend half the physics win. Both numbers are real; they answer different questions.

Assumptions

  • PUE 1.25 today (liquid AI build; the industry average 1.54 would flatter us — we don't use it)
  • PUE 1.15 air replacement; 18 kW racks are catalogue equipment
  • hosts/network +28% best / +40% conservative
  • iso-throughput, 70B-class decode-dominated serving
Technical Appendix · Economics

Cost per token — both floors, line by line

// GB300 NVL72 floor (100% utilisation, 4-year straight-line)
$3,500,000 ÷ 35,040 h = $99.87/h
+ power: 132 kW × PUE 1.25 × $0.08/kWh = $13.20/h
= $113.07/h ÷ 3,960M tok/h = $0.029 / M tokens

// Invariant-1 floor, binary (same accounting)
$12,000 ÷ 35,040 h = $0.342/h + power $0.101/h + host share $0.06/h
= $0.503/h ÷ 3,780M tok/h = $0.000133 / M
× 3 derate (prefill share, <100% util, batch dilution) = $0.0004 / M → 72×
// per billion tokens (the main-slide unit): $29 them · 40¢ binary · 85¢ ternary

// ternary: $16k, 1.7 kW, 650k tok/s → $0.66/h ÷ 2,340M = $0.00028 ×3 = $0.00085 / M → 34×
// what this does to the market
70B-class output tokens sell for $0.60–0.90 / M today. Against our binary floor that is a 1,500–2,200× markup — labeled context, never our headline ratio. First mover on this floor sets the price and still runs 99% gross margin.

Assumptions — every derate lands on our side

  • electricity $0.08/kWh industrial, both sides
  • 4-year amortisation, no financing, both sides
  • 3× serving derate applied only to us; their floor stays theoretical-perfect
  • their rack priced at $3.5M (est. range $3–4M)
  • BOM detail in chip spec
Technical Appendix · Edge

Invariant-E vs Jetson — batch 1, nothing hidden

Device8B decodePowerPricetok/Jtok/s/$k
Jetson AGX Orin 64GB — 204.8 GB/s LPDDR5≤51 ceiling · ~40 real15–60 W$1,999~0.920
Jetson AGX Thor — 273 GB/s LPDDR5X, 2,070 TFLOPS FP4≤68 ceiling · ~55 real40–130 W$3,499~0.616
Invariant-E — 1 GB W1A1 on-die SRAM, 2 TB/s500 spec · 2,000 ceiling3 W~$399~1671,250
Invariant-E ternary variant — 1.6 GB trit-packed250 spec · 1,250 ceiling~2.5 W~$449~100557
Invariant-E Value — one monolithic 28 nm die, 1 GB on-die250 spec · ~750 ceiling~3 W~$69~833,600
// the ceiling column is just the decode law — nobody escapes it
Jetson: LPDDR bandwidth ÷ 4 GB of FP4 weights. 2,070 TFLOPS of compute cannot help batch-1 decode.
Invariant-E: 2 TB/s on-die ÷ 1.0 GB = 2,000 tok/s ceiling — we spec 500 and spend the slack on power.
Invariant-E energy per token (batch 1)mJ
Weights: 8×10⁹ bits × 0.3 pJ/bit (mobile stacked SRAM)2.4
XNOR compute: 16 Gop × 10 fJ0.16
KV (windowed + MLA, LPDDR)~2
SoC overhead, DRAM standby1–2
Total → 6–9 mJ/tok → 100 tok/s @ 0.9 W · 500 tok/s @ ~3 W6–9
First token: 512-tok prompt ÷ 30 Top/s burst = ~270 ms, zero network. Battery: 1M tokens ≈ 2.2 Wh ≈ 15% of a phone battery for a day of heavy use. And it works in a tunnel, on a plane, in a war zone.

Assumptions

  • Jetson "real" figures: public Jetson AI Lab / developer benchmarks for 7–8B at 4-bit; ceilings are bandwidth ÷ bytes — generous to Jetson (no efficiency loss assumed)
  • Invariant-E: 1 GB SRAM die (~400 mm² N5-class, or 2×200 mm²), 0.3 pJ/bit; specified, not taped out
  • KV windowed to 1–2k recent tokens + MLA — assistant workloads
  • value tier = the 28 nm monolithic port: memory + logic on one fully-depreciated die, ~$34 BOM → ~$69 — always under a Jetson by a wide margin; the N5 flagship exists for sub-watt phone sockets
  • the same architecture ports to 22–28 nm mature nodes for cost-down and sovereign supply (see roadmap)

Why this market is unwinnable for them

  • their silicon must stream weights over LPDDR — the decode ceiling is physics, not engineering
  • a 16× smaller model is the only way through, and it requires binary-native models they don't have
  • always-on frontier AI at <1 W: phones, wearables, vehicles, robots, defence — the entire ambient-AI category is ours by default
Technical Appendix · Market

Iso-power: what 100 MW buys under binary

// method mirrors the iso-throughput walk, inverted
Today: 100 MW → 667M tok/s (606 racks × 1.1M — see power walk)
Binary: (100 ÷ 1.5) × 667M = 44.5B tok/s (57B on the 1.1 MW best estimate — stated once, used never)
Ternary: (100 ÷ 2.7) × 667M = 24.7B tok/s

44.5×10⁹ × 86,400 = 3.8 quadrillion tokens/day ÷ 8.1×10⁹ humans ≈ 5.5 tok/s per human, continuously
Same substation. Same permits. Same building. 65× the sellable intelligence. For an operator whose binding constraint is megawatts — which is all of them — this isn't an upgrade. It's a different business with the same address.

Assumptions

  • conservative 1.5 MW replacement basis (65×)
  • 70B-class decode-dominated serving mix throughout
  • world population 8.1B
Technical Appendix · Market

Who buys it — and why they can't not

BuyerWhy they move
Hyperscalers / frontier labspower-constrained on every continent; every MW we free is a MW for training. A 72× token-cost floor is not a procurement decision — it's a board-level emergency for whoever moves second.
Neoclouds & API resellerstheir entire P&L is $/token margin. First mover resets the market price and prints margin while GPU fleets serve at a loss.
Sovereign AI programsa national inference grid for $35M and 1.5 MW — no US-scale grid, no exotic supply chain, sited domestically, running open or domestic models.
Enterprises (on-prem)one 18 kW rack in an ordinary server room replaces the cloud contract; data never leaves the building. Compliance departments buy this before engineers do.
Wedge: sell tokens per megawatt to whoever is most power-starved. The hardware sale follows the token economics; the token economics follow the physics; the physics is on the previous nine slides.

Context

  • AI datacenter power demand is the industry's stated #1 constraint (IEA; utility interconnect queues run years)
  • inference, not training, is where token volume and power spend compound
  • Invariant sells chips/systems or capacity on owned fleets — both priced against a 72× floor advantage
Technical Appendix · Test-Time Scaling

The evidence — and its limits, stated

ResultWhat it showed
Snell et al. 2024
arXiv 2408.03314
compute-optimal test-time scaling lets a smaller model beat one 14× larger, FLOPs-matched (MATH); adaptive allocation is 4× more efficient than best-of-N
Large Language Monkeys
arXiv 2407.21787
coverage scales log-linearly over 4 orders of magnitude of samples; DeepSeek-V2-Coder: 15.9% → 56% on SWE-bench Lite at 250 samples vs 43% frontier single-attempt SOTA
The reasoning era
o1/o3, R1
frontier gains now come from spending more tokens at inference — the industry already concedes that tokens are intelligence
// the arithmetic that matters
250 attempts × $0.0004/M ≈ the cost of 3.4 attempts at NVIDIA's own floor — or 0.1 attempts at API prices.
Token price is the exchange rate between money and intelligence. We moved it 72×.
// the honest caveat, where it belongs
Without automatic verifiers, sample-selection (voting, reward models) plateaus beyond a few hundred samples (Monkeys, limitations). Search pays best where answers can be checked — code, math, tools, agents — which is exactly where token volume is exploding.
Problems solved vs attempts · log scale

Why this compounds with binary

  • parity-per-token is the deck's premise; parity-per-dollar doesn't even need it — a weaker model with 72× the search budget already competes
  • at 11,400 tok/s single-stream, a 10k-token reasoning chain completes in <1 s — search becomes interactive
  • our cost floor converts directly into capability headroom no GPU fleet can match at iso-cost
Technical Appendix · The Collapse

Weight quantisation — timeline & method

DateMethodBits/weight at ~parity
2021FP16 deployed baseline16
Aug 2022LLM.int8()8
Oct 2022 – Jun 2023GPTQ · AWQ4
Feb 2024QuIP# · AQLM — first Pareto-optimal <2-bit2
Apr 2025BitNet b1.58 2B4T — native ternary, 4T tokens, matches FP16 peers. Open weights.1.58
2025ParetoQ (Meta) — sub-4-bit QAT scaling laws keep improving1.58–2

Extrapolation method — two axes, printed, falsifiable

  • Axis 1 — bit-width: parity bit-width has halved every ~12 months for 4 consecutive years (log₂ falls ~1/yr)
  • Axis 2 — scale: the size at which each low-bit result reaches parity climbs ~10× params per 1–2 yrs
  • 70B is 1.5 orders above 2B → ternary parity at 70B-class ≈ 2027 (±1 yr)
  • trend extrapolation, not a theorem — a bet stated at full strength, not a guarantee
Provenance: our compiler already beats heuristic SOTA at 3- and 2-bit. We are a point on this curve, not a spectator of it.

Sources — every point on the chart

LLM.int8() GPTQ AWQ QuIP# AQLM BitNet b1.58 BitNet 2B4T weights ParetoQ
Technical Appendix · The Collapse

Activation quantisation — the harder axis

DateMethodBits/activation
2021FP16 baseline16
Nov 2022SmoothQuant — W8A8 by migrating outlier scale into weights8
Apr 2024QuaRot — rotate the network; outliers vanish; W4A4KV44
2024SpinQuant — learned rotations, ~2.9 pts from FP at W4A4KV44
Nov 2024BitNet a4.8 — 4-bit activations on 1.58-bit weights, no loss vs b1.584
proj. 2026A2 on ternary weights (rotations + native training)2
proj. 2027–28A1 small-scale → W1A11
Why activations were "impossible": a handful of outlier channels carry huge values and wreck uniform quantisation. Rotations (Hadamard, learned) spread that energy flat — the outliers were an artifact of basis, not information. That's why the curve bent in 2024 and keeps bending.

Add the scale lag (~1–2 yrs per 10× params) → W1A1 at 70B-class lands 2028–2030, median 2029. The lab's job is to land it early — and the ternary mode of this deck pays the bills while we do.

Assumptions

  • same two-axis method as the weight timeline
  • original BitNet (2023) used A8 — activations lag weights ~2 yrs consistently; the projection preserves that lag
  • softmax/norm stay INT8 even at W1A1 (~1% of ops — budgeted in the power table)
Technical Appendix · The Lab

Why the loop must close in one lab

PlayerWhy they don't build this
GPU vendorsbinary-in-SRAM deletes the HBM bill of materials they monetise. You do not disrupt your own margin structure; you get disrupted out of it.
Frontier model labsno silicon teams; GPU-portable roadmaps by design. BitNet is Microsoft Research — shipped as a paper, because shipping it as a product requires a chip.
Chip startupscan't train models. A binary chip without binary-native weights is an empty socket. (See Groq: superb silicon, borrowed models, brutal economics.)
Hyperscaler TPU teamsfleet homogeneity rules; a W1A1 bet breaks their fleet years before it pays.
The co-design is literal: the model's quantisation fixes the compiler's lattice, which fixes the fabric's dataflow, which fixes the SRAM floorplan. One loop, one set of decisions, one team. Whoever holds all three compounds; whoever holds two waits on a partner — and the partner is the moat.
Milestones — each de-risks the next
1 · 8B W1A1 at parity — our own run; the scientific result that prices everything
2 · FPGA prototype — the decode law measured, not modelled
3 · Invariant-1 tape-out — the spec in this deck, on mature nodes, no exotic supply

Provenance

  • our precision compiler already beats heuristic SOTA (GPTQ/AWQ-class) at 3- and 2-bit on coding benchmarks
  • the same mathematics — loss-landscape curvature, not per-layer heuristics — is what goes below 2 bits

The honest risk

  • NVIDIA can add binary tensor cores. What they cannot do quickly: abandon HBM economics, retrain frontier models binary-native, and clean-sheet the architecture. That cycle is measured in years — that's the window, and the moat is the loop, not the gate.
Technical Appendix · The Lab

Team

Diego Granziol

Diego Granziol

FOUNDER & LEAD SCIENTIST

Oxford MPhys, double first; ML PhD at Oxford’s MLRG under Stephen Roberts — spectral machine learning. Huawei AI Theory & Amazon ML — deployed under real product constraints, including on-device. Two JMLR papers on the geometry of training (batch-size scaling via random matrix theory, 2022; iterate averaging, 2024) the RMT × deep-learning line with Keating & Baskerville, and the Safety–Efficacy Trade-off (ICML 2026, with Ulug’bek Abdimanobov).

Khurshid Juraev

Khurshid Juraev

CO-FOUNDER & LEAD ENGINEER

Represented Uzbekistan at the International Mathematical Olympiad; informatics olympiad. National University of Singapore, data science. Co-author, Hessian Spectral Analysis at Foundation Model Scale (2026, with Diego). Algorithms & systems — and the olympiad bench: nine engineers he trained or recruited, on call.

Jon Keating

Jon Keating

FOUNDING ADVISOR

Fellow of the Royal Society; Fröhlich Prize. Quantum chaos and random matrix theory × the Riemann zeta function. Wills Professor at Bristol, then Sedleian Professor of Natural Philosophy at Oxford; President of the LMS; chaired the Heilbronn Institute (the GCHQ partnership) to 2020. Random matrices & spectral problems: the foundations of our compiler.

Tony Stansfield

Tony Stansfield

FOUNDING ENGINEER

CTO of sureCore for a decade, leading design of the company’s low-power SRAM and in-memory-compute IP. Technical Design Authority at Fractile, the Gelsinger-backed in-memory-compute inference accelerator startup. Transistor-level and custom circuit design; tape-out experience across process nodes down to 6nm.

The mathematics that beats heuristics at 3 bits is the mathematics that gets to 1. This team wrote it.

Technical Appendix · Commercial

What we sell, in what order

StageProductEconomics
0 · NowPrecision-compiler optimisation on customers' existing GPU fleets — the compiler that beats GPTQ/AWQ-class at 3-/2-bit ships as a service todaypriced as a share of measured savings — revenue and design-win relationships before any silicon exists
1 · Edge firstInvariant-E + SDK — OEM design wins in phones, vehicles, robotics, defencevalue tier ~$69 (28 nm monolithic) to flagship $399–449 (N5, sub-watt) — both ≈2× BOM, vs $249–3,499 incumbents
2 · DatacenterInvariant-1 systems or capacity — sell 18 kW racks, or sell tokens per megawattcapacity priced under market ($0.60–0.90/M) and >30× above our floor — ~99% gross margin on capacity at market price
3 · SovereignDesign + compiler licensing for national inference gridsmature-node manufacturable; standard export-control compliance; licensing revenue on others' capex
// pricing doctrine
We price just under the market, never near the floor. The floor is the moat, not the price. Every dollar between our floor and the market price is margin nobody else can follow us down to.

Where the models come from

  • customers bring their own weights; open models are converted by our compiler (ternary today)
  • our W1A1 pretraining runs are the science bet, not a frontier-lab product — we are not competing with OpenAI for models
  • we own the format layer: the compiler, the kernels, the silicon it lands on

Wedge logic

  • edge ships first: smaller die, mature nodes, cheaper tape-out, faster design wins (roadmap →)
  • datacenter follows with the proven fabric scaled up
  • every stage is gated by the previous one's proof — no leap-of-faith capex
The Question Every VC Asks

Why wouldn't NVIDIA do this?

Their moat is generality

CUDA sells flexible GPU programming: every precision, graphics, scientific computing, whichever model is frontier this month. That breadth is what customers pay for.

Generality dies at one bit

CUDA schedules multiplies. HBM streams 16-bit weights. At one bit there is nothing to schedule and nothing to stream — the flexibility has nothing left to sell.

They will support ternary

Kernels for it are certain. But the weights still stream from HBM: best case on a GB300 rack is 457 tok/s per stream at 0.06 J/token — 25× slower, 57× more energy than ours.

The segment is beneath them

Hundreds of billions in revenue, chasing $1T on $3M racks at ~75% margin. A cheaper, simpler $5.5k part is not a market they reorganise for — the innovator's dilemma, on schedule.

457 vs 11,400tok/s single-stream — their ternary ceiling vs our chip
57×energy per token — set by off-die weights, not software
~75%gross margin on $3M racks they must protect
$1Tthe market they are built to chase — ours looks like noise

Their ceiling assumes perfect 2-bit packed kernels on GB300-class HBM — the best case for them.

Appendix · Head-to-Head

Cerebras: same physics, opposite lever.

Cerebras made our diagnosis first: weights must live in SRAM, or you die on the haul. Their fix is to grow the chip 57× — one un-diced wafer. Ours is to shrink the model 16× — one bit. Hover the parts.

Cerebras WSE-3 — one 300 mm wafer, un-diced[1][2]
84 RETICLE FIELDS
STITCHED IN THE SCRIBE LINES
·
900K CORES · FP16/FP8 MAC
Litho exposes one ~858 mm² reticle at a time — Cerebras wires across the scribe lines (plus spare cores and redundant routing for defects) to fuse 84 fields into one 46,225 mm² fabric. 4T transistors — still multiply-accumulate: the op we delete.
44 GB SRAM
21 PB/s
The point of the wafer — SRAM residency. But N5 SRAM is wafer-priced: ~1,000× HBM per resident GB. And it stopped scaling: WSE-2 → WSE-3 grew logic +54%, SRAM 40 → 44 GB (+10%).
MEMORYX
SWARMX IO
70B at FP16 is 140 GB > 44 GB — weights stream from external memory boxes, or the model shards across ~4 systems. Off-wafer, the haul returns.
23 kW · LIQUID LOOP · ~$2–3M PER SYSTEM (REPORTED)Custom power delivery and a liquid cooling loop per wafer — the price of fitting 16-bit numbers by force.
Invariant-1 — one package, mature nodes
XNOR FABRIC +
INTEGER THRESHOLDS
The multiply is deleted, not accelerated — sign-match gates, popcount trees, 12-bit accumulators. Logic-grade yield on boring, available CMOS.
BONDED SRAM — 12 GB · MODEL ON-DIE
V-Cache-proven stacking: a 70B binary model is 8.75 GB — weights + KV fit in-package. No interposer exotica, no allocation queue.
NO HBM
NO CoWoS
Nothing to queue for — the supply chain NVIDIA and Cerebras fight over simply isn’t in the design.
AIR-COOLED · 1.1 kW · MATURE NODES · $5.5K BOMInput cost at scale — margins decided later. A microwave’s power for a rack’s output.

Same thesis, opposite lever — and ~1,000× apart on cost per resident GB. Their scaling axis (more wafer) has stalled with SRAM itself; ours (fewer bits) halves yearly. The honest asymmetry, stated: they run every FP16 checkpoint today; we need binary-native models — the dated bet this deck prices. Cerebras is validation, not refutation: they prove customers pay a premium for SRAM-resident latency.

Appendix · Head-to-Head

The CPU: educator, not competitor.

bitnet.cpp proved ternary models run on silicon you already own[1] — free marketing for the binary future. Then volume arrives, and the physics bill comes due. Hover the parts.

Server CPU — EPYC-class, ~$25k dual-socket[2]
64–128 GENERAL CORES
OoO · BRANCH PREDICTORS
AVX-512 · POPCNT
Flexibility, not tokens: most of the silicon exists to run anything. Binary popcount rides along at a ~0.2 P-op/s ceiling — the batched limit of ~1.5k tok/s per socket.
12× DDR5
~0.6 TB/s
The ceiling. Decode re-streams all 8.75 GB of a 70B binary model per token → ~65 tok/s single-stream. At one bit the DRAM haul is 16× lighter — but never deleted.
L3 · ~0.4 GB~4% of the model — caches cannot hold a working set that is re-read wholesale every token.
~$25k SERVER · ~1 kW · AVAILABLE TODAYThe honest win: zero NRE, ships today, bitnet.cpp runs out of the box — the correct answer for low-volume private inference.
Invariant-1 — one package, mature nodes
XNOR FABRIC +
INTEGER THRESHOLDS
Fixed-function: every transistor either holds a weight or counts sign-matches. Nothing rides along.
BONDED SRAM — 12 GB · MODEL ON-DIE
Weights move microns, not modules: the whole 70B binary model lives in-package — no DRAM anywhere in the token path.
NO DRAM
IN PATH
The haul isn’t lighter here — it’s gone.
AIR-COOLED · 1.1 kW · MATURE NODES · $5.5K BOMInput cost at scale — margins decided later.

Same power, three orders of magnitude apart on delivered tokens — because the CPU’s bandwidth ceiling is the DRAM haul, and ours doesn’t exist. The framing for the room: at 16 bits models needed GPUs; at 1 bit they merely run on CPUs — and belong on silicon where the weights never move. CPUs could run graphics in 1995, too. GPUs happened anyway.

Technical Appendix · Supply Chain

Slide the node. See what it builds.

Every number below is a physics ceiling — SRAM density × stacked area, nothing else. Constant at every stop: no CoWoS, no HBM, no N3/N2 allocation — and no kernel NVIDIA can write moves any of it, because their weights still stream from off-package memory at fixed bandwidth. Models are trainable to any of these sizes.

5nm 7nm DUV 14nm 16nm 28nm
Who can fab it
max model · trivial KV
single-stream
batch decode
tape-out · first silicon
unit BOM
fab queue

Assumptions

  • capacity = SRAM weight budget ÷ bits per weight; single-stream = bandwidth ÷ weight-bytes; batched = × concurrent streams (KV-limited on small stacks)
  • max model quoted at trivial KV (batch-1 / short context, KV <1% of weights); dense serving (96 streams × 2k ctx) draws ~25–35% of weight bytes from the same pool — resident model shrinks accordingly
  • beyond one package — taller stacks, multi-package, 1T-class: the capacity ladder →
  • the ~100 TB/s hybrid-bond port is treated as node-independent — bond pitch is packaging, not lithography; unverified below 7nm-class
  • figures marked ~ are engineering estimates; the rest anchor to the chip spec and roadmap NRE classes; ternary capacities (97B · 50B · 15B) carried in the data

The edge line rides the same ladder

  • N5 → Invariant-E: 8B on 1 GB on-die, 500 tok/s at 3 W
  • 28nm monolithic → E Value: ~1–2B per die, 250 tok/s, ~$34 BOM
  • 22–28nm → SE: sub-1B always-on at 0.3 W, ~$14 BOM
Technical Appendix · Scaling

The capacity ladder — resident models beyond 70B

ConfigurationCapacityMax binary modelBOM (est.)$ / B params
9-die N5 stack — the shipping design18 GB~144B~$7k~$50
16-die tall stack — SoIC roadmaps run 8–12+ high; same bonding flow32 GB~250B~$9–10k~$40
2 packages, 2.5D stitch — only 1–2-bit activations cross64 GB~500B~$18–20k~$38
4 packages, one board128 GB~1T~$36–40k~$38
Gain-cell / oxide-semiconductor memory — 2–4× SRAM density; ~2028–3060–120 GB0.5–1T~$8–12k~$10–20
CFET-era SRAM — stacked transistors resume cell scaling; 2030+~50–60 GB~450B~$12–15k~$30
Wafer-scale at N5 density — binary-native, speculative~115 GB~900B~$50–100k system~$60–110
// the two facts that matter
Capacity scale-out is linear in money: a resident 1T-parameter machine is ~$40k of BOM — four of today’s packages and copper, not a new technology.
$/token stays ~flat as packages multiply (bandwidth scales with capacity) — the cost of size and the cost of tokens are independent dials.

Assumptions

  • basis: 2 GB per 800 mm² N5 HD die (~$300); bonding, package and test per the cost model; binary weights at trivial KV; ternary sizes = ÷1.6
  • the first four rows need no new technology — taller stacks are the V-Cache flow continued; multi-package interconnect is cheap wires because only sign-class activations cross
  • bond-pitch scaling (9 → ~3 µm, demonstrated) multiplies bandwidth, not capacity — tok/s upside on every row, independent of this table
  • the gain-cell row is the one memory technology that improves $/parameter ~3×; pre-production — the technology worth tracking
  • every row keeps weights resident: the no-kernel-fixes-this argument holds at 1T exactly as at 70B, while a streaming architecture’s ceiling grows only with HBM stack height — the queue they are already in

Node-by-node manufacturing picture: slide the node →

Technical Appendix

What the sharp money asks

Existence

"W1A1 at parity doesn't exist."

Correct — that's the deck's stated frame, and every headline number re-prices cleanly on ternary, which does exist (open weights, 2B, FP16 parity). Bit-width at parity has halved yearly for 4 years. Our first milestone is our own 8B W1A1 — the experiment that prices the company.

CPUs

"Why not just CPUs — bitnet.cpp runs ternary on a laptop."

It does — that's our free marketing. A ~$25k server moves ~0.6 TB/s from DRAM: ~65 tok/s single-stream on a 70B binary, ~1.5k batched at a kilowatt. Our package moves weights zero millimetres: 11,400 and 1.05M at the same power — ~400× tokens/W, ~700× $/throughput. CPUs educate the market and serve low-volume private inference; volume forces the ASIC. Chip-to-chip →

Scaling laws

"Precision laws say ~8-bit is optimal."

Those laws fit post-hoc/fixed-budget quantisation; native training keeps bending them (BitNet 2B4T, ParetoQ). And even a 2× effective-param tax leaves ≥55× energy and ~25× latency.

KV cache

"Where does the KV cache live?"

On-package: MLA-latent + 4-bit, ~3.1 GB at batch 96 × 2k context. Long context pages to LPDDR at a disclosed ~2× derate at 32k. It's a budgeted line item with its own power row.

Bandwidth

"100 TB/s of SRAM is fantasy."

It's 21 GB/s/mm² areal across 4,800 mm² of bonded SRAM — less than half of what AMD ships in V-Cache today. MI300X pulls 17 TB/s from one 256 MB layer. Throughput mode already needs 96 TB/s, so latency mode adds no new assumption.

Competition

"Groq and Cerebras already do SRAM inference."

At FP8/16 — needing ~576 chips or a 23 kW wafer for 70B. They proved the latency market and the economics of not having small weights. Binary is 16× denser: one 1 kW package. They validated the demand; compression changes the answer.

NVIDIA

"NVIDIA just adds binary tensor cores."

The honest risk — it has its own card in the Full Stack slide. The win isn't the MAC: it's weights-in-SRAM (deletes the HBM economics they monetise), binary-native frontier weights (they don't have), and a compiler. Clean-sheet redesign + retraining = a window measured in years.

Benchmarks

"Your 1.1M baseline is unverified/offline."

Disclosed on-slide — and it favours them. Offline is their best case; deployed serving runs ~18× worse than the floor we grant them. We compare against their press release, not their reality.

Supply

"Fab and packaging capacity?"

Replacing a 100 MW site = 635 chips ≈ ~60 N5 wafers + SoIC capacity — noise against GPU volumes. No CoWoS, no HBM allocation. SRAM dies are redundancy-repairable: high yield on a 2020-era node.

Incumbents

"Why hasn't Google or Microsoft done this?"

BitNet is Microsoft Research — as a paper. Shipping it needs clean-sheet silicon + native pretraining + a compiler, against every incumbent's GPU-portability roadmap. Vertically-integrated fabless labs are the historically successful shape for exactly this move.

Training

"You can't even train W1A1 — backprop dies at sign()."

Two paths. Path A: native QAT with straight-through estimators — how ternary reached parity; we extend that toolchain. Path B: skip backprop entirely — evolution strategies (EGGROLL, 2025) train with forward passes only, in pure integer arithmetic, population-averaging away the noise STE fakes, at ~91% of inference throughput. Path B's workload is millions of cheap forward passes — the exact thing our silicon does 46–118× cheaper. If B scales, the training market lands on our chip too. (Honestly: demonstrated at 1.5B-class fine-tuning, not 70B pretraining.)

Team

"Three founders and no silicon veteran?"

Correct, and the plan is shaped around it: FPGA-first retires the physics risk before the first big cheque; physical design runs with an established design-services partner (the standard fabless path); first silicon hires are a lead SoC architect, an SRAM/memory designer and a packaging lead. The team today is the part nobody can hire: the mathematics and the compiler.

Revenue

"What do you sell before the chip exists?"

The compiler. It already beats GPTQ/AWQ-class heuristics at 3-/2-bit and ships as an optimisation service on customers' existing GPU fleets, priced as a share of measured savings — cash and design-win relationships years before tape-out. Then edge silicon, then datacenter capacity. Full sequencing in the Business Model appendix.

Scale

"The frontier is MoE at 1T params — isn't a 70B die irrelevant?"

Scale-out, not scale-up: only 1–2-bit activations ever cross a package boundary, so a multi-chip fabric is cheap wires, not HBM — the 2-die 2.5D stitch is the proven pattern. And MoE helps us: more parameters per FLOP is exactly what cheap on-die capacity rewards; experts partition naturally one-per-die.

Distillation

"Why not just distill to a small FP16 model?"

Distillation shrinks the parameter count; it does not change the format. A distilled 8B at FP16 still pays 16× the memory physics of the same model binarised. The two compound — distill, then drop the bits — they do not compete. Our floor applies to whatever size the distillers produce.

Parity

"Is 70B ternary at parity actually real today?"

No — and the deck says so: parity is proven at 2B (open weights); 70B is the ~2027 point on a four-year trend. What is already fact in ternary mode is the hardware: the format's energy, cost and capacity numbers. The dated claim is model quality at scale, and we ship datacenter capacity only when open evals pass.

Geopolitics

"Can you actually sell to sovereigns?"

We are a fabless company on mature and N5-class nodes with standard export-control compliance; sovereign deals are licensed designs manufactured where lawful — a regulated go-to-market lane, not a gray market. The attraction for the buyer is that nothing in our supply chain is on the EUV frontier.

Software

"Software ecosystem?"

Serving is an OpenAI-compatible API; edge is an SDK; training stays PyTorch. Binary is an inference-format problem we own end-to-end — not a CUDA-replacement problem. Nobody rewrote code for FP8 either.

If a question isn't here, we want it — investment@invariant.fyi