C O H E R A T E C H

More work from every tick of the clock. Less electricity follows.

COHERA builds the HPU, a Hybrid Processing Unit: a new kind of helper chip that works beside the processors the world already runs on. Here is what it is, who needs it, and what it did on real hardware.

23 September 2026 · draft and confidential · every figure recorded on our own hardware unless marked ESTIMATE or DESIGN INTENT · the engineering detail is folded at the foot of the page
Part 1 · What it is

Three helpers you already know. One you don’t.

Every computer has a main processor. Over fifty years, each new kind of work brought a new kind of helper chip that plugs in beside it. None replaced the others; they work side by side.

CPUCentral Processing Unit
the generalist · 1970s

Can do anything, one step at a time. It runs the computer.

GPUGraphics Processing Unit
the crowd · 1990s

Thousands of simple workers doing the same thing at once. Built for graphics, then drafted into AI.

TPU · NPUTensor / Neural Processing Unit
the AI specialist · 2010s

Built for the arithmetic of AI: Google’s TPU in its data centres, NPUs inside phones and laptops.

And now, a fourth kind
HPU
Hybrid Processing Unit · by COHERA

The efficient one. The same work and the same answers as the other three, for a fraction of the electricity. Let the numbers below speak.

Like an NPU or a GPU, the HPU works alongside the CPU. It does not put more load on your processors: it takes work off them and does it for a fraction of the electricity. DESIGN INTENT

Measured todayThe HPU's engines, recorded on our laboratory chip doing image recognition. The numbers in Part 3 come from here.
AI inference · in progressWork in progress: our card has already worked beside a GPU, choosing what a large AI model reads from its memory; every one of its choices matched the reference. Its speed and electricity on that job are the next measurement.
Analogue · in developmentComputing directly with electrical signals rather than digital numbers, for the lowest-power jobs.
Part 2 · Who needs it

Every industry that makes decisions from data. Which is every industry.

Inside all of these, a small program looks at incoming data and decides: pass or reject, keep or discard, alert or ignore, approve or block. Millions of times a second, all day. Every one of those decisions costs electricity.

These are uses, not measurements. Inside each of them the chip does the same kind of arithmetic whatever the data was; the numbers in Parts 3 and 4 are for one measured model, and a model of a different size is measured again.

Part 3 · The fair fight

One core against one core.

Every crate is a pile of finished decisions. The glass column is electricity. Same work: both towers grow together; watch whose glass fills up. Same electricity: both glasses empty together; watch whose tower grows. Pick a question and a rival, and press play.

Why this comparison. One core each. Same work. Same answers. Same electricity meter. The only difference is the design.
The question
The rival
The work
Same workloadidentical job and identical answers on both sides
drag to look around

Part 4 · Against real computer chips

Against NVIDIA’s GB10: one HPU engine measured, then counted up to match its speed.

Same work: both towers grow together; watch whose glass fills up. Same electricity: both glasses empty together; watch whose tower grows. The rival here is a finished commercial chip in our own lab. We measured one HPU engine and worked out how many it would take to keep pace: 15 against its GPU, 4 against its CPU. Then we compare the electricity.

Why this comparison. Real processors and graphics chips bundle many cores into one package, so a single engine against a whole package is not a fair comparison. We take the speed and electricity of one measured HPU engine and multiply it to the count that matches the rival. A chip with several HPU engines has not been built yet, so that side is an estimate; the rival’s side is measured. Our engine runs on a development board; the rival is finished production silicon.
The question
The rival
The work
Same workloadidentical job and identical answers on both sides
drag to look around

For the engineers · sources and conditions

Four designs ran the same digit classifier (MNIST, 784→32→10, int8, 25,408 multiply-adds per answer) on one AMD ZCU104 board (Zynq UltraScale+ ZU7EV), 1,000 images checked against the same golden (957/1000) on every batch, electricity read at the board's 12 V input (INA226) at 20 samples a second. HPU engine = MXU v2 at 150 MHz; "AI chip" = VTA-class micro-NPU gen-2 at 150 MHz; "graphics chip" = 8-lane SIMT soft GPU at 115 MHz; "standard processor" = NaxRiscv rv64gc soft core at 125 MHz. All four are built in the same programmable fabric; none is a commercial CPU, GPU or NPU part. Each is one core: the processor benchmark runs single-threaded on one soft core; the graphics-style design is one 8-lane compute unit; the AI chip and the HPU engine are one engine each. Run files (sha256, first twelve): mnist-mxu 7da436e99069, mnist-gpu 410232328880, mnist-softcpu 63e293cd3848 (14 September 2026), mnist-npu 7483ed79d97e (16 September 2026); results file 0fcab0fa59e9. Whole-board electricity per decision: HPU 17.38 µJ, AI chip 60.87, graphics chip 9,207, standard processor 12,784 µJ; mean board power 10.30 · 10.20 · 10.19 · 13.05 W. Part 3A: whole-board electricity per decision for each design (recorded); "same work" = 600,000 decisions; "same electricity" = one kilowatt-hour. Part 3B, measured on the NVIDIA GB10 in our lab on 24 September 2026 with the identical model, weights, test set and arithmetic, every arm reproducing checksum 0x78c65cd1 and 957/1000 (raws: tree5 raws/real-opponent-2026-09-24, pre-registered before the run): GB10 graphics chip, batches of 65,536, 8,570,747 decisions/s, 21.82 W above its idle on its own power sensor = 2.546 µJ per decision added; GB10 processor, all 20 cores, 2,199,242 decisions/s; its electricity is not measurable on this box and is estimated at 80 W added from ServeTheHome's DGX Spark review (40–45 W idle, 120–130 W with the processor loaded, whole system) = 36.4 µJ per decision; range 75–110 W added (34.1–50.0 µJ) including Jeff Geerling's Dell Pro Max GB10 review (about 30 W idle, processor maxing out around 140 W). Measuring points differ: the processor figures are whole-system at the wall (including its power supply's loss), the graphics chip is its own sensor, the HPU is the 12 V board input after its adapter; this favours the HPU on the processor comparison. HPU engine: 592,272 decisions/s and 0.4355 µJ per decision added above the board floor (recorded); engines needed = rival speed ÷ 592,272, rounded up (15 against its GPU, 4 against its CPU), a calculation, not a built chip. "Same work" in 3B = one hour at the rival's speed; "same electricity" = one kilowatt-hour of added electricity. One real processor core (Cortex-X925) measured 460,535 decisions/s. Industry labels are illustrations: each assumes a decision the same size as the measured one, which the record shows can change the ranking when the job grows (MNIST scale-up, 450× the arithmetic). "1 second for 600,000 decisions": HPU 1.0 s against 9 min 48 s for the processor, both inside their recorded ten minutes. AMD accelerator: Vitis AI DPU B4096, two cores, as shipped, driven by AMD's runtime, our measurement of 3 August 2026, 111× whole board against that campaign's COHERA figure. AI inference stage: on 2 September 2026 a production FPGA card on AWS made every KV-selection call for Qwen2.5-14B on a GPU host, 2,448/2,448 live layer receipts and 120,960/120,960 replayed selections equal to the certified reference, at 61–68 s per token over the internet; a correctness result, not a speed or energy result. RECORDED measured on hardware · ESTIMATE computed from recorded figures under stated assumptions · DESIGN INTENT what the product is built to do, not yet measured.