RTL-validated performance.
Measured, not estimated.
Not modeled. Not estimated. Measured on RTL.
VLEN=1024b · Validated on RTL — not simulated, not estimated.
Qwen 2.5 7B — Full batch-1 FP16.
Real-time LLM inference under a 100-cycle DRAM penalty · batch 1 · full FP16, no quantization.
Decode throughput · 34.7 ms/token
Core power
Compute die area
| Platform (FP16, batch 1) | tok/s | ms/token | SoC power | Compute die |
|---|---|---|---|---|
| SimplEx Micro (core sim) | 28.8 | 34.7 | ~150 mW core | ~1.5 mm² |
| SimplEx Micro (SoC est.) | 28.8 | 34.7 | ~5 W | ~15 mm² |
| Apple M4 Max (40-core) | 35.8 | 28 | 70 W | 410 mm² |
| Apple M4 (base) | 8 | 125 | 40 W | 167 mm² |
| NVIDIA L4 | 12 | 80 | 72 W | 294 mm² |
| AMD EPYC 8534P | 8.2 | 121 | 110 W | 70mm²×4 |
| NVIDIA Jetson Orin Nano | N/A | N/A | 7 W | 200 mm² |
Simulated / RTL / synthesis. Competitor decode = batch-1 FP16, public specs. SoC estimate uses HBM4 TSV I/O in 7nm.
YOLOv10-M FP16 — Real-time detection.
640×640 with 32×32 GEMM option · and DeepSORT 500-object tracking.
| Platform tier | Latency / frame | FPS | Class |
|---|---|---|---|
| Server / edge CPU class | 211–347 ms | 2.9–4.7 | not real-time |
| SimplEx + 32×32 GEMM | ~15 ms | 50–70 | real-time edge ✓ |
| Jetson Orin Nano class | 16–22 ms | 45–62 | GPU + tensor |
| NVIDIA T4 / desktop | ~4.7 ms | ~211 | datacenter ceiling |
| Ambarella | — | — | no FP16 support |
TinyML ResNet INT8 — Sub-watt, sub-microsecond.
First inference — TinyML ResNet
Energy per task
FP16 matmul efficiency
Performance under real DRAM latency.
Faster at low latency (10 cycles) vs. best-in-class RISC-V CPU/VPU
Faster at high latency (100 cycles) — where competitors stall
SimplEx performance from 10 → 100 cycles of latency
Performance stays high when memory gets slow. Exactly what matters at the edge.
Benchmark: MatMul FP16 latency sweep · SimplEx vs. best-in-class RISC-V CPU/VPU. Measured on RTL.
vs. ARM Ethos NPU — the edge AI standard.
| Metric | ARM + Ethos NPU | SimplEx VPU |
|---|---|---|
| Cost & power | High — two chips | Fraction of cost & power |
| CPU↔Accel link | MMIO — serialized | RVV — direct, deterministic |
| Memory model | DMA + dedicated SRAM | Latency-tolerant, no preload |
| Dynamic scaling | Fixed-function, idle between | Cycle-by-cycle, no idle silicon |
How do we know these numbers are real?
RTL correlation methodology, cache studies, and simulation rigor — see the Performance Model.
View the Performance Model