NP-1 “Apex” · Edge Inference Accelerator · Rev 1.2

One die. 512 INT8 TOPS.
96 MB on-chip SRAM, 256 GB/s.

NP-1 “Apex” is a 4 nm neural-processor for sustained on-device inference. It runs transformer and convolutional workloads at the edge — no host offload, no network round-trip. Specifications below are engineering-silicon figures.

NP-1 · FLOORPLAN PRE-PRODUCTION REV 1.2 · 2026-08
NP-1 Apex AI accelerator chip PCIe Gen5 x16 PHY LPDDR5X PHY LPDDR5X PHY Global SRAM · 80 MB · 16 banks Command processor · scheduler · 2-D mesh NoC PE 00 PE 01 PE 02 PE 03 PE 10 PE 11 PE 12 PE 13 PE 20 PE 21 PE 22 PE 23 PE 30 PE 31 PE 32 PE 33
INT8 throughput
512 TOPS
FP16
256 TFLOPS
FP4 / INT4
1,024 TOPS
On-chip SRAM
96 MB
Memory bandwidth
256 GB/s
Power (typical)
85 W
Process
4 nm
Package
45×45 mm FC-BGA
All figures are “up to” or “typical” and measured on engineering silicon with PeakStack 3.2. Values subject to change before production release. See datasheet for full conditions.
01 / Product line

Three parts, one toolchain.

All datasheets →
02 / Measured, not claimed

Throughput on engineering silicon.

Benchmarks & methodology →

ResNet-50 · INT8 · batch 64

NP-1 Apex
48,000 img/s
NP-2 Crest
12,000 img/s
NP-4 Ridge
1,800 img/s

PeakStack 3.2 · batch 64 · median of 1,000 runs · 25 °C ambient.

Llama-3-8B · INT4 · decode

NP-1 Apex
180 tok/s
NP-2 Crest
55 tok/s
NP-4 Ridge
8 tok/s

Weight-only INT4 · FP16 KV cache · batch 1 · single decode token.

03 / Under the lid

A memory-bound design, by intent.

Architecture →

Dataflow, from host to MAC array

Host CPUPCIe Gen5 x16
Command processorscheduler + DMA
Global SRAM80 MB · 2 TB/s agg.
16 × PEtile SRAM · MAC array

Why SRAM, not HBM

The memory wall, not compute, bounds most edge-inference latency. NP-1 keeps weights resident in on-chip SRAM so the MAC array is never idle waiting on an external bus. LPDDR5X handles the long tail: activations, KV cache, and multi-tenant context.

Global SRAM ~2 TB/s, tile SRAM ~12 TB/s — versus 256 GB/s off-chip. Full subsystem in the architecture document.

04 / Latest

Announcements.

News & changelog →
2026-08-14

PeakStack 3.3 — FP4 sparsity kernels and ONNX Runtime bridge

Toolchain release adds 2:4 structured-sparsity kernels for FP4, a scheduler pass for the 2-D mesh, and a drop-in ONNX Runtime execution provider.

2026-06-30

NP-2 “Crest” reaches general availability

Mid-range part moves from sampling to GA across commercial and industrial temperature grades.

2026-04-09

NP-1 first silicon boots; bring-up complete

Engineering samples pass functional bring-up across the full die: PCIe, memory PHY, and all 16 processing elements.

05 / In production

What design partners say.

The SRAM-first architecture is what closed the deal for us. Our latency budget stopped being a negotiation with the memory controller.

Robert Chen
Robert ChenVP Platform · Edge Robotics

We moved from host offload to fully on-device and the power envelope barely moved. That is the number our customers actually care about.

Nina Johansson
Nina JohanssonHead of AI · Industrial Automation