Apex Inference Chip Review: 0.56 tok/s and the Truth About FPGA LLM Inference

Apex Inference Chip Review: 0.56 tok/s and the Truth About FPGA LLM Inference

The Apex Inference Chip (SigmanticAI’s APEX, or tinyNPU) does run a real LLM on an FPGA, and it does so at 0.56 tokens per second on its fastest measured image — roughly 1.78 seconds per token, confirmed on two separate builds. That number is deliberately slow, honestly published, and the entire point of the project is the verification trail behind it. That paragraph is the honest summary, and this review exists because almost every other sentence about this repository can be made misleading with one small edit. Strip the words “measured on silicon” and APEX looks like a hardware failure. Strip “one decoder layer” and it looks like a chip. Strip “projected from an analytic model” and its 7B performance claims look like results. The repository itself refuses to let you do any of those things, which is why it is worth your time even if you never build a bitstream. ...

October 1, 2026 · 19 min · baeseokjae