NeuralPeak builds neural-processing chips for AI inference at the edge — low latency, low power, and deterministic, so robotics, cameras, and industrial devices can think on their own hardware.
Robots, cameras, and machines on factory floors can't wait for a round trip to a data center — and they can't afford the power or the connectivity dependence. NeuralPeak puts inference on the device itself.
Sub-millisecond inference with a fixed timing budget. A control loop that closes in one millisecond behaves very differently from one that closes in fifty.
Edge devices run on batteries and tight thermal envelopes. Our chips do the work of a server-class accelerator inside a single-digit power budget.
Inference on-device means video, telemetry, and sensor streams stay local — no upload, no air gap to cross, no link to drop.
A single silicon architecture scaled across power and performance points — so a camera and a robot arm can run the same models, the same toolchain.
The full-size neural engine for robotics, autonomous platforms, and multi-camera systems that need headroom for transformer-class models.
A power-sipping variant for smart cameras, doorbells, and sensor nodes where the whole board has to fit inside a lens housing.
General-purpose silicon wastes energy moving data. NP-1 is a purpose-built neural engine — a tiled compute array fed by on-die SRAM, with a lightweight control plane and a deterministic network-on-chip.
Representative figures from our own evaluation boards, batch 1. Full methodology and model list on the performance page.
INT8, 224×224, single chip
INT8, 1080p, end-to-end pipeline
Typical, INT8 sustained
Perception, manipulation, and navigation running on-board — so a robot keeps working when the network drops.
Detection, tracking, and recognition at the sensor, without streaming raw video to the cloud.
Defect inspection and process monitoring with deterministic latency, inside a cabinet on the line.
Low-power perception for platforms where every watt is range and every millisecond is control authority.
Our next-generation die is powered up and passing bring-up. Early results, and what it means for transformer inference at the edge.
2026-06-24On-device LLM inference joins the toolchain — quantized, sparse, and running in a single-digit watt envelope.
2026-04-15Why we rebuilt the reference design around thermal headroom, and what that unlocks for sealed industrial housings.
2025-11-19The compact vision processor moves from engineering samples to first design-ins with camera and sensor partners.
"We moved detection and tracking off the host and onto the NP-1, and the whole system got quieter — lower power, lower latency, and no more waiting on a round trip to the server. It's the first edge part our team didn't have to fight."