Architecture / NP-1

SRAM is the product. The MAC array is plumbing.

A neural accelerator is a memory system with arithmetic attached. NP-1 is organized around keeping weights and activations on-die so the array never waits on a bus.

01

Top-level block diagram

NP-1 Apex · die blocks

control plane on the left, memory at the top, the compute array center, memory PHY at the right
PCIe Gen5 x16 PHY host interface Command processor boot · dispatch Scheduler + DMA mesh placement Global SRAM · 80 MB 16 banks · ~2 TB/s Memory ctrl LPDDR5X LPDDR5X PHY 2× 32-bit 16 × processing element (PE) PE 00 PE 01 PE 02 PE 03 PE 10 PE 11 PE 12 PE 13 PE 20 PE 21 PE 22 PE 23 PE 30 PE 31 PE 32 PE 33
02

Processing element

Each PE is self-contained: it holds its tile of weights, executes against them, and reduces locally before ever touching global SRAM.

PE microarchitecture

one of sixteen · systolic MAC array at the center
Tile SRAM 1 MB · ~750 GB/s Systolic MAC array INT4 / INT8 / FP8 / FP16 / FP4 Vector engine softmax · norm Quant / dequant · 2:4 sparsity decode
03

Memory subsystem

Three tiers, three orders of magnitude of latency and bandwidth. The compiler's job is to keep traffic in the top two tiers.

TierCapacityAggregate bandwidthLatencyPurpose
Tile SRAM16 MB · 1 MB × 16 PE~12 TB/s~2 nsWeights and activations local to one PE
Global SRAM80 MB · 16 banks~2 TB/s~8 nsCross-PE reduction, KV cache, staging
LPDDR5Xup to 64 GB256 GB/s~80 nsLong context, multi-tenant working set

Why no HBM

HBM buys bandwidth at the cost of latency, thermals, and a co-packaged stack that fights the edge power budget. Edge inference is latency-bound, not bandwidth-bound at the margin. A large on-chip SRAM pool plus commodity LPDDR5X covers the workload envelope without the package cost.

Dataflow, by workload

  • Convolution — weight-stationary; weights pinned in tile SRAM, activations stream in.
  • Attention — KV cache resident in global SRAM; softmax and norm in the vector engine.
  • Fully-connected — tensor-sliced across the 16 PEs, reduced over the mesh.
  • Sparse — 2:4 metadata decoded on-chip; zero tiles skipped, not fetched.
04

Interconnect & precision

2-D mesh NoC

The sixteen PEs are wired as a two-dimensional mesh. Neighbor-to-neighbor transfers are single-hop; global reductions are folded across rows then columns. The scheduler places dependent operators on adjacent PEs to keep hops minimal.

Quantization datapath

One MAC array, five precisions: INT4, INT8, FP8, FP16, and FP4. Conversion to and from higher precision happens in the quant/dequant block, not in software. Sparsity is 2:4 structured — two of every four weights are skipped with metadata, so the speedup is deterministic rather than data-dependent.