
An open-source project called openTPU is testing whether AI coding agents can contribute to the design of the hardware used to run AI models, with a working accelerator now executing modern language models on an inexpensive Xilinx Kintex-7 FPGA board.
Developed by Felipe Sens Bonetto, the openTPU repository brings the accelerator’s hardware and software stack into a single project. It includes the SystemVerilog RTL, a custom instruction set, an instruction-level simulator, a kernel language, compiler, host software and FPGA build files.
The project describes its central question in more direct terms: how far AI agents can go in hardware design, and whether they can build the hardware that runs their own inference. The current implementation is an FPGA accelerator rather than a fabricated AI chip, but it has been tested on physical hardware with real model weights.
A complete inference stack on one FPGA
The openTPU hardware targets an Inspur YPCB-00338 accelerator card built around a Xilinx Kintex-7 XC7K480T-FFG1156-2 FPGA. The board provides two 2 GiB DDR3 memory channels for 4 GiB of total local memory and connects to the host system through PCIe.
The accelerator’s production configuration uses a 128-wide data dimension, four matrix-unit columns and eight vector-processing lanes. It runs at 133.33 MHz and has a theoretical combined DDR3 bandwidth of about 17.1 GB/s.
Rather than relying on a large software stack outside the project, openTPU exposes the path from model kernels to hardware instructions. Its compiler turns kernels written in the project’s small kernel language into the accelerator’s instruction set, while the simulator provides an executable reference for checking the hardware’s behavior.
The repository also includes host software and utilities for controlling and profiling the accelerator. The design does not use a cache or a hidden scheduler; memory movement is explicitly represented in the program.
Qwen, Gemma and LFM2.5 run on the board
The project has moved beyond simulation and demonstrated inference using actual model weights on the Kintex-7 card. Its published benchmark set covers models including Qwen3, Qwen3.5, LFM2.5 and Gemma 4, along with several larger models.
For example, the project reports a device decode rate of 85.8 tokens per second for the 4-bit version of LFM2.5-230M, compared with 59.0 tokens per second using int8 weights. The 4-bit Qwen3-0.6B configuration reaches 31.3 tokens per second, while Qwen3.5-0.8B reaches 24.5 tokens per second.
Gemma 4 E2B is reported at 12.14 tokens per second using the project’s 4-bit configuration. The benchmark measurements were carried out on the physical accelerator rather than being estimates derived only from simulation.
The reported figures distinguish accelerator execution from complete host-to-device wall-clock performance. In the project’s benchmark, decoding uses 64 greedy tokens after a 512-token prompt, with some host-side work included in the wall-time measurement.
The project also reports models that exceed the board’s 4 GiB local memory capacity. For mixture-of-experts models, it can keep selected experts on the FPGA and stream others from host storage over PCIe.
That approach has been demonstrated with LFM2.5-8B-A1B, which the project reports at 10.6 tokens per second, and Qwen3.5-35B-A3B, reported at 3.95 tokens per second. The larger model has 34.7 billion total parameters but about 3.0 billion active parameters, according to the project’s measurements.
AI-generated hardware is being tested against physical constraints
The question behind openTPU is more demanding than asking an AI model to generate SystemVerilog that compiles. Hardware changes have to survive simulation, synthesis, implementation and timing constraints before they can be useful on a real FPGA.
The project draws on the methodology used in Bonetto’s related auto-arch-tournament project. That system uses AI coding agents to propose and implement architectural changes, then evaluates each candidate with a fixed pipeline involving hardware checks, benchmarks and implementation tests. A candidate only becomes the new design when it produces a measurable improvement while remaining valid.
That earlier experiment evaluated 73 hardware hypotheses, with 10 winning changes eventually merged. The reported result improved its measured fitness from about 301 iterations per second to roughly 577 iterations per second.
For openTPU, the same general approach shifts the target from a CPU experiment to an AI inference accelerator. The important distinction is that the agents operate within a human-defined development and verification environment. The available evidence does not show an AI system independently conceiving, fabricating and validating a commercial AI chip from scratch.
The project instead demonstrates an AI-assisted hardware engineering loop in which generated changes can be checked against an instruction-level simulator and then taken through real FPGA implementation.
Performance is limited by the memory system
The Kintex-7 implementation is also constrained by the hardware it runs on. The project reports that inference decoding is largely DRAM-bound, because model weights have to be streamed repeatedly for each generated token.
During decoding, the accelerator is reported to use roughly 82% to 94% of the available DRAM bandwidth, depending on the model and configuration. The project identifies memory efficiency and timing margin as continuing engineering constraints.
The production design reportedly closes timing at 133.33 MHz with only about 0.032 ns of worst-negative-slack margin, showing how tightly the implementation is operating within the FPGA’s limits.
The project’s 4-bit configurations reduce weight traffic further. Its documentation describes a custom 4-bit representation averaging about 4.25 bits per weight, with the language-model head generally remaining at int8. The trade-off is lower memory traffic and higher decode throughput at the cost of some model-quality degradation.
For the larger models that cannot fit entirely into local memory, PCIe becomes another constraint. The project reports that Qwen3.5-35B-A3B requires substantial expert transfers between host storage and the accelerator during decoding.
openTPU’s repository states that hardware execution and the project’s ISA simulator produce matching tokens in its tests, providing a direct software reference for the physical implementation.
Discover more from Aree Blog
Subscribe now to keep reading and get access to the full archive.



