
Developers are turning retired FPGA hardware originally built for cryptocurrency mining into custom accelerators for Qwen3.5, including a Qwen3.5-9B implementation running on an SQRL Forest Kitten 33 board that sells for roughly $270 to $350 on the secondary market.
The Forest Kitten 33, or FK33, uses a Xilinx Virtex UltraScale+ VU33P FPGA with 8GB of HBM2. In the llm.vhdl project, developers built a VHDL inference engine specifically around Qwen3.5-9B rather than relying on CUDA or a conventional GPU runtime.
An early hardware run on the FK33 produced about 2 tokens per second at a 75MHz clock, according to the project’s developer, who described the board as an approximately $280 purchase. Current secondary-market listings have also placed FK33 boards in the same general price range.
Qwen3.5-9B runs on a second mining-era FPGA
A separate project has produced a faster Qwen3.5-9B result on another piece of SQRL hardware. The fable5_llm project runs the model on an SQRL BCU-1525 built around an XCVU9P FPGA and reports a best measured decode rate of 8.18 tokens per second.
That figure comes from an actual hardware measurement rather than a projected result. The project documents a progression from 6.77 tokens per second in its original schedule to 7.62, 7.93 and finally 8.18 tokens per second after execution reordering and double-buffering changes.
The BCU-1525 is a different board from the roughly $300 FK33. It was also considerably more expensive when FPGA mining hardware was in demand, with historical records showing prices around $3,350 to $3,600 during the cryptocurrency mining boom.
The fable5 implementation uses INT4 weights and streams the model through the FPGA’s memory system. Its documentation reports about 3,902 MiB for the quantized weight pack and four DDR4 channels delivering a measured aggregate bandwidth of 70.70GB/s. The project estimates a memory-bandwidth ceiling of about 17.3 tokens per second for the workload.
The hardware is built around Qwen3.5’s hybrid architecture
The FPGA implementations are designed around the structure of Qwen3.5 rather than trying to reproduce a general-purpose GPU execution environment. The 9B model contains 24 Gated DeltaNet layers and eight conventional attention layers, requiring different computation paths for the two layer types.
The VHDL implementation includes custom hardware for INT4 matrix-vector operations, Gated DeltaNet processing, attention, normalization, rotary position embeddings, KV-cache handling and a transformer sequencer. Model weights, activations and state are arranged around the FPGA’s on-board high-bandwidth memory.
The fixed-point approach also reduces the amount of hardware needed to represent the model. The fable5 project uses integer arithmetic and verifies its FPGA output against an integer reference implementation, with the repository documenting bit-for-bit hardware checks and repeated state verification.
Performance remains constrained by memory movement. The fable5 documentation identifies weight streaming and the ability to overlap data movement with computation as significant limits on throughput, with parts of the execution schedule leaving the compute hardware idle.
The 27B project has not reached a completed hardware demonstration
The more ambitious part of the VHDL project involves a two-FPGA SQRL Jungle Cat configuration based on two VU35P devices. The project describes this system as a target for Qwen3.8-27B, not Qwen3.5-27B.
The repository estimates that the INT4 27B model could fit within the available FPGA resources, but the multi-FPGA implementation still has unresolved hardware and data-transfer requirements. The planned high-speed Aurora link requires a clocking solution that is not currently populated on the carrier board, while the existing loading path is too slow for moving a model of this size efficiently.
The project’s tensor-parallel communication layer also remains incomplete. As a result, Qwen3.8-27B should be treated as a development target rather than a completed benchmark on the Jungle Cat.
The broader hardware effort traces back to SQRL’s cryptocurrency-mining products. The company’s surviving product site identifies the FK33 and BCU-1525 as FPGA mining hardware, while mining software documentation also lists the boards among supported FPGA systems.
A separate FPGA project has taken a similar approach at a much smaller model scale. The OpenTPU project reports running Qwen3.5-0.8B on an Inspur YPCB-00338 Kintex-7 board at about 14 tokens per second at 120MHz.
Discover more from Aree Blog
Subscribe now to keep reading and get access to the full archive.


