Positron AI Raises $875 Million to Take On Nvidia’s Inference Chips

On this page
Positron AI just raised $875 million at a $5 billion valuation to build an inference chip that skips Nvidia’s most expensive bottleneck: high-bandwidth memory. The Series C, announced September 10, 2026, is one of the largest AI-chip funding rounds of the year, and it’s built around a simple bet — that most AI workloads today don’t need the priciest memory on the planet, they need a lot of ordinary memory wired together well.
I’ve been tracking the inference-chip race on this site since we covered Cerebras’s CS-4 launch, and Positron’s pitch stands out because it isn’t trying to out-spec Nvidia on raw FLOPs. It’s trying to make inference cheaper by using memory nobody is fighting over.
What Positron Actually Announced
The round was co-led by NEA, Atreides Management, Valor Equity Partners, Andra Capital, SemiAnalysis Capital, and Jim Clark — the Silicon Graphics founder, which is a fitting name to see on a chip funding round. The $875 million pushes Positron’s post-money valuation to $5 billion, a huge jump from the $1 billion-plus valuation it hit with its $230 million Series B back in February 2026.
The money isn’t going toward a moonshot. Positron says it will fund three concrete things: the tapeout of its next chip, called Asimov, on TSMC’s N3P process; production ramp of its Titan inference systems; and a new 2 MW-plus engineering data center and emulation platform to test chips before they ship. That’s a company spending on manufacturing capacity, not just research headcount.
Why Skipping HBM Is the Interesting Part
Every high-end AI chip on the market right now — Nvidia’s Blackwell and Vera Rubin lines, AMD’s MI300 series, Google’s TPUs — leans on high-bandwidth memory (HBM) stacked directly on the chip package. HBM is fast, but it’s also scarce, expensive, and single-sourced from basically three suppliers (SK Hynix, Samsung, Micron). That scarcity is a big reason GPU allocations have been so tight industry-wide.
Asimov takes a different path. It pairs Positron’s compute architecture with commodity LPDDR5X memory — the same kind of memory used in phones and laptops — and scales it up to between 288 GB and 2,304 GB per chip, depending on configuration. That’s a staggering amount of memory capacity per chip, achieved by using cheap, abundant parts instead of scarce, expensive ones.
The tradeoff is bandwidth-per-byte: LPDDR5X isn’t as fast as HBM on a pound-for-pound basis. But for inference — running an already-trained model rather than training a new one — memory capacity often matters more than raw bandwidth, especially for models with huge context windows or mixture-of-experts architectures that need to keep a lot of parameters resident at once. Positron is betting that “enough bandwidth, way more capacity, far lower cost” beats “maximum bandwidth, tight capacity, high cost” for a large share of real-world inference traffic.
The Titan System: Built for Trillion-Parameter Models
Asimov chips don’t work alone. Positron’s Titan system bundles four to eight Asimov chips into a single node, and the company says it’s designing Titan to handle models beyond 16 trillion parameters with context windows exceeding 10 million tokens — in one node, before you even start scaling across a rack or a data center. Whether models that large become common is still an open question, but the direction of travel — bigger context windows, more experts, more resident parameters — has been consistent for two years running.
Asimov is scheduled to tape out on TSMC’s N3P node at the end of 2026, with production targeted for the second half of 2027. That’s a long runway, and it means Positron’s $875 million has to last through a year-plus of no chip revenue before Titan systems built on real Asimov silicon start shipping.
Quick Reference
| Detail | Figure |
|---|---|
| Round size | $875 million (Series C) |
| Post-money valuation | $5 billion |
| Previous valuation (Feb 2026 Series B) | ~$1 billion+ on $230 million raised |
| Chip name | Asimov |
| Process node | TSMC N3P |
| Memory per chip | 288 GB – 2,304 GB (commodity LPDDR5X) |
| Tapeout target | End of 2026 |
| Production target | H2 2027 |
| System product | Titan (4–8 Asimov chips per node) |
Why This Matters Beyond One Funding Round
The bigger story here is that the AI infrastructure market is starting to fragment on purpose. For the last three years, almost every serious AI workload assumed you’d eventually need Nvidia HBM-based silicon, whether that was an H100, a Blackwell, or now a Vera Rubin system — the kind of order we covered when AM Intelligence placed a $8 billion order for 9,000 Vera Rubin systems. Positron, and companies like Cerebras before it, are betting there’s a large enough second market — inference at scale, not frontier training — where a fundamentally different memory strategy wins on cost per token served rather than peak performance. We saw a related bet on the compute side when we wrote about Cerebras’s CS-4 hitting 30x faster inference without a new chip; Positron is approaching the same problem from the memory side instead of the compute side.
If Positron delivers Asimov on schedule, it gives cloud providers and large AI labs a real alternative for inference-heavy workloads at a moment when HBM supply is the single tightest constraint in the entire AI buildout. If it slips — and first-generation chip startups slip constantly — the $5 billion valuation will look like it got ahead of the silicon. Six co-lead investors and a founder of Jim Clark’s stature putting real money behind the bet is a signal worth watching, not a guarantee.
Frequently Asked Questions
What does Positron AI actually make?
Positron designs AI inference chips — hardware built to run already-trained AI models efficiently, rather than to train new ones. Its current focus is the Asimov chip and the Titan system that bundles multiple Asimov chips into one node.
How is Asimov different from an Nvidia GPU?
The core difference is memory. Nvidia’s high-end chips use high-bandwidth memory (HBM) stacked on the chip package for maximum speed. Asimov instead uses commodity LPDDR5X memory — the type found in phones and laptops — to reach much higher memory capacity per chip at a lower cost, trading some raw bandwidth for far more room to hold large models.
When will Asimov chips actually be available?
Positron expects Asimov to tape out on TSMC’s N3P process at the end of 2026, with production targeted for the second half of 2027. That means real, shipping Titan systems built on Asimov silicon are still roughly a year out from this funding announcement.
Who invested in Positron’s $875 million round?
The Series C was co-led by NEA, Atreides Management, Valor Equity Partners, Andra Capital, and SemiAnalysis Capital, with participation from Silicon Graphics founder Jim Clark.
