RTL-validated performance.
Measured, not estimated.

Not modeled. Not estimated. Measured on RTL.

Measured on RTL Performance Model Correlated Projected at 5nm Industry reference
400 GFLOPS
Matrix multiply @ VLEN 1024
2.56 TOPS
Peak — single core @ 5nm
49.2 TFLOPS
64-core system, modeled
Zero
Speculative execution overhead

VLEN=1024b · Validated on RTL — not simulated, not estimated.

Qwen 2.5 7B — Full batch-1 FP16.

Real-time LLM inference under a 100-cycle DRAM penalty · batch 1 · full FP16, no quantization.

28.8 tok/s

Decode throughput · 34.7 ms/token

~150 mW

Core power

~1.5 mm²

Compute die area

Platform (FP16, batch 1)tok/sms/tokenSoC powerCompute die
SimplEx Micro (core sim)28.834.7~150 mW core~1.5 mm²
SimplEx Micro (SoC est.)28.834.7~5 W~15 mm²
Apple M4 Max (40-core)35.82870 W410 mm²
Apple M4 (base)812540 W167 mm²
NVIDIA L4128072 W294 mm²
AMD EPYC 8534P8.2121110 W70mm²×4
NVIDIA Jetson Orin NanoN/AN/A7 W200 mm²

Simulated / RTL / synthesis. Competitor decode = batch-1 FP16, public specs. SoC estimate uses HBM4 TSV I/O in 7nm.

YOLOv10-M FP16 — Real-time detection.

640×640 with 32×32 GEMM option · and DeepSORT 500-object tracking.

Platform tierLatency / frameFPSClass
Server / edge CPU class211–347 ms2.9–4.7not real-time
SimplEx + 32×32 GEMM~15 ms50–70real-time edge ✓
Jetson Orin Nano class16–22 ms45–62GPU + tensor
NVIDIA T4 / desktop~4.7 ms~211datacenter ceiling
Ambarellano FP16 support

TinyML ResNet INT8 — Sub-watt, sub-microsecond.

365 ns

First inference — TinyML ResNet

54 nJ

Energy per task

2.67 TFLOPS/W

FP16 matmul efficiency

Performance under real DRAM latency.

2.8×

Faster at low latency (10 cycles) vs. best-in-class RISC-V CPU/VPU

16×

Faster at high latency (100 cycles) — where competitors stall

~Flat

SimplEx performance from 10 → 100 cycles of latency

Performance stays high when memory gets slow. Exactly what matters at the edge.

Benchmark: MatMul FP16 latency sweep · SimplEx vs. best-in-class RISC-V CPU/VPU. Measured on RTL.

vs. ARM Ethos NPU — the edge AI standard.

MetricARM + Ethos NPUSimplEx VPU
Cost & powerHigh — two chipsFraction of cost & power
CPU↔Accel linkMMIO — serializedRVV — direct, deterministic
Memory modelDMA + dedicated SRAMLatency-tolerant, no preload
Dynamic scalingFixed-function, idle betweenCycle-by-cycle, no idle silicon

How do we know these numbers are real?

RTL correlation methodology, cache studies, and simulation rigor — see the Performance Model.

View the Performance Model