Edge inference silicon · 6 nm-class NPU

Neural compute that runs where the data is.

NeuralPeak builds neural-processing chips for AI inference at the edge — low latency, low power, and deterministic, so robotics, cameras, and industrial devices can think on their own hardware.

neuralpeak · NP-1 · die floorplannode · 6 nm-class
SRAM · L2 vector engine control tensor cluster SRAM · L2 NoC
Peak compute
40 TOPS INT8
Memory bandwidth
68 GB/s
Typical power
8 W
Inference latency
<1 ms
Why edge

The cloud is too far away for the things that move.

Robots, cameras, and machines on factory floors can't wait for a round trip to a data center — and they can't afford the power or the connectivity dependence. NeuralPeak puts inference on the device itself.

01 / LATENCY

Deterministic, not "best effort"

Sub-millisecond inference with a fixed timing budget. A control loop that closes in one millisecond behaves very differently from one that closes in fifty.

02 / POWER

Watts, not kilowatts

Edge devices run on batteries and tight thermal envelopes. Our chips do the work of a server-class accelerator inside a single-digit power budget.

03 / SOVEREIGNTY

Data that never leaves the device

Inference on-device means video, telemetry, and sensor streams stay local — no upload, no air gap to cross, no link to drop.

Product line

One architecture, two form factors.

A single silicon architecture scaled across power and performance points — so a camera and a robot arm can run the same models, the same toolchain.

FLAGSHIP

NP-1

edge inference processor

The full-size neural engine for robotics, autonomous platforms, and multi-camera systems that need headroom for transformer-class models.

  • Up to 40 TOPS INT8 · 68 GB/s memory bandwidth
  • Typical 8 W · 6 nm-class process node
  • Transformer, CNN, and small-language-model support
  • MIPI, PCIe, and Ethernet host interfaces
Full specifications →
COMPACT

NP-C1

vision processor

A power-sipping variant for smart cameras, doorbells, and sensor nodes where the whole board has to fit inside a lens housing.

  • Up to 8 TOPS INT8 · 12 GB/s memory bandwidth
  • Typical 1.5 W · sub-4 mm package
  • Single-camera vision pipelines: detection, tracking, embedding
  • MIPI CSI-2 in, USB or SPI out
Full specifications →
Architecture

A die built for one job: running models well.

General-purpose silicon wastes energy moving data. NP-1 is a purpose-built neural engine — a tiled compute array fed by on-die SRAM, with a lightweight control plane and a deterministic network-on-chip.

host · SoC MIPI / PCIe / Eth DRAM · LPDDR 68 GB/s neural engine tiled compute array SRAM · NoC · vector
  • INT8, INT4, and FP16 — precision per layer, not per model
  • On-die SRAM keeps weights resident; no refetch per inference
  • Deterministic NoC with bounded, predictable routing latency
  • Hardware scheduler that skips zero-weight operations automatically
  • Toolchain: PyTorch and ONNX in, one quantization pass, flash
Performance

Real models, measured on real silicon.

Representative figures from our own evaluation boards, batch 1. Full methodology and model list on the performance page.

ResNet-50 · batch 1
4,800 img/s

INT8, 224×224, single chip

YOLO-class detector
120 fps

INT8, 1080p, end-to-end pipeline

Energy efficiency
5.0 TOPS/W

Typical, INT8 sustained

See full benchmark methodology →
Where it runs

Built for machines that have to get it right the first time.

01

Robotics & automation

Perception, manipulation, and navigation running on-board — so a robot keeps working when the network drops.

02

Smart cameras & vision

Detection, tracking, and recognition at the sensor, without streaming raw video to the cloud.

03

Industrial & manufacturing

Defect inspection and process monitoring with deterministic latency, inside a cabinet on the line.

04

Drones & mobile platforms

Low-power perception for platforms where every watt is range and every millisecond is control authority.

From the lab

Notes on what we're shipping.

See all news & releases →
Who it's for

"We moved detection and tracking off the host and onto the NP-1, and the whole system got quieter — lower power, lower latency, and no more waiting on a round trip to the server. It's the first edge part our team didn't have to fight."

Marta Kowalski · Director of Vision Systems, industrial automation