Most AI hardware projects are closed, enormous, or both. openTPU is the opposite: one small monorepo holding the whole stack, from the instruction set and SystemVerilog RTL to a bit-exact simulator, a kernel language with its compiler, and a profiler that replays runs in the browser. Its authors frame it as a question rather than a product: how far can AI agents get at hardware design, and can they build the accelerator that runs their own inference?

The design is deliberately simple. A sequencer issues one instruction per cycle to a small set of units: DMA moves data, a systolic matrix unit multiplies int8 weights streamed from DRAM, a vector unit does fp32 math, and a quantizer writes results back as int8. There is no cache and no hidden scheduling, so every byte moved is an instruction and every cycle can be explained from a trace — a debugging property conventional NPUs rarely offer.

On an Inspur YPCB-00338 card carrying a Xilinx Kintex-7 xc7k480t and two DDR3 channels, the project reports ten modern models running with their real weights, producing tokens identical to the simulator's, bit for bit. LFM2.5-230M decodes at 59 tokens/second in int8 and 85.8 in 4-bit; Qwen3-0.6B at 21.6; Qwen3.5-4B at 5.88 in 4-bit. DRAM is the bottleneck, and the card already reaches 82–94% of the DDR3-1066 peak.

Models larger than the card's 4 GB of memory run by expert offload: a mixture-of-experts model keeps its experts in per-layer slots, the card routes each token and computes every expert, and the host streams only the missing experts over PCIe. LFM2.5-8B-A1B (8.5B parameters, 1.7B active) reaches 10.6 tokens/second with a 98.5% slot hit rate; Qwen3.5-35B-A3B (34.7B parameters, 3B active) reaches 3.95 tokens/second.

The project is Apache-2.0 and readable end to end, which is its real contribution. It is a teaching artifact as much as a benchmark: the simulator serves as the specification the RTL must match, tests keep the two aligned, and the ISA documentation plus kernel examples let a developer follow one matmul from Python down to the wires. Everything but the card runs on a laptop — `otpu-chat` can talk to a model through the ISA simulator with no FPGA at all.

Why it matters: if agent-driven hardware design can produce a working, verifiable accelerator for an FPGA that shipped years ago, the same loop becomes a plausible way to explore accelerator design faster — and to run small models on hardware nobody would otherwise write a compiler for. It is also a clean answer to the question of whether 'AI-designed silicon' has anything behind it beyond press releases: here is the RTL, the simulator, and the bit-exact match.