Inference Chip Funding Wave Reaches New High: Positron AI Closes $875M Series C at $5B Valuation, Betting LPDDR5X Will Disrupt HBM

Positron AI has raised $875 million in Series C funding at a $5 billion post-money valuation, betting that LPDDR5X memory can replace HBM in AI inference. The company plans to use the funds for its Asimov custom chip tape-out, a 2 MW+ engineering data center and simulation platform, and mass production of its Titan inference servers.

On September 10, 2026, AI inference chip company Positron AI announced the close of an $875 million Series C round at a $5 billion post-money valuation. The round was split into two tranches: the first, a $375 million Series C, was co-led by NEA, Atreides Management, Valor Equity Partners, Andra Capital, and SemiAnalysis Capital, founded by Dylan Patel, at an estimated valuation of $3.5 billion; the second, a Series C-1 of up to $500 million, was led by NEA and Jim Clark, founder of Silicon Graphics and Netscape. Qatar Investment Authority (QIA), DFJ Growth, Cisco Investments, Hudson River Trading, and other institutions participated.

The funds will be used to pay for tape-out of the Asimov custom chip, build a 2 MW+ engineering data center and simulation platform, and drive mass production of Titan inference servers. Asimov is expected to complete tape-out by the end of 2026, and Titan systems will enter mass production in H2 2027.

The Real Technical Bet: Bypassing the HBM Supply Chain

Positron’s financing rationale is based on replacing the high-bandwidth memory (HBM) commonly used in AI chips with commercial LPDDR5X memory. The HBM supply bottleneck is the real choke point in today’s AI hardware supply chain. HBM production depends heavily on TSMC’s CoWoS advanced packaging capacity, which must simultaneously meet demand for NVIDIA’s H-series and B-series GPUs, leaving supply extremely tight. In September 2025, AMD was forced to discontinue some of its Versal product lines due to HBM2E supply issues.

Positron’s logic for choosing LPDDR5X is that the performance bottleneck for inference workloads is not compute power but memory capacity and bandwidth. During large language model inference, model weights must be repeatedly read from memory; memory bandwidth utilization is the key variable determining performance per watt and performance per dollar. The company’s system achieves over 90% utilization of available memory bandwidth.

Asimov’s memory specification scales from 288 GB to 2,304 GB per chip, enabled by LPDDR5X rather than HBM. This allows Titan systems to build multi-terabyte memory inference nodes and handle ultra-long contexts and very large parameter models without relying on scarce packaging resources.

Reference Points for the Inference Financing Wave

Between January and August 2026, the AI chip sector completed 12 disclosed financings totaling about $5.37 billion, of which inference-oriented companies accounted for 8 deals and about 67% of the capital. Positron alone accounts for about 16% of total sector financing. Groq raised $650 million in June 2026 to continue expanding its inference cloud service; Cerebras completed its IPO in May 2026, with a first-day fully diluted valuation of about $56 billion.

Atlas Is Already Running, Not Just a PowerPoint

Positron’s first-generation product, Atlas, has already been deployed in real production environments. More than 50 racks of Atlas systems are currently running on Oracle Cloud Infrastructure, and inference provider Parasail uses this capacity to power its own inference services; quantitative trading firm Jump Trading and gaming infrastructure company i3d.net are also production customers of Atlas. Quantitative trading firms are extremely sensitive to inference latency and energy efficiency; winning such customers indicates that Atlas’s actual performance has at least passed internal technical evaluations.

“Speed matters in this market—both how quickly you launch a new generation of chips and how quickly you get them into customers’ hands. Deploying Atlas at scale gave us a deep understanding of what inference customers really need, and those lessons have directly informed the design of Asimov and Titan.” — Mitesh Agrawal, CEO of Positron AI

What Dylan Patel’s Participation Signals

In this round’s investor lineup, Dylan Patel participated as a lead investor through SemiAnalysis Capital and will join Positron’s board. SemiAnalysis is known for in-depth teardown reports on the chip supply chain and AI compute costs.

Unanswered Questions

Asimov’s measured bandwidth data has not yet been made public. LPDDR5X is already widely used in mobile and edge devices, but at data center inference scale, whether its interconnect latency and bandwidth stability can support the company’s claimed “90% bandwidth utilization” requires independent benchmark testing. Samsung’s LPDDR5X-PIM research shows that standard LPDDR5X interface bandwidth is about 76.8 GB/s, while HBM4 can already reach more than 2 TB/s.

Titan’s mass production timing (H2 2027) means there is still a window of more than a year before formal delivery to large-scale customers. At a time when AI hardware iterates extremely quickly, successors to NVIDIA’s Blackwell architecture and AMD’s next-generation inference chips could both change the competitive landscape within this window.

Assessment

The rationale for Positron’s financing lies in betting on a real structural gap in the industry. The reality that HBM supply is constrained by TSMC’s CoWoS packaging capacity will not disappear in the short term; inference cost and throughput efficiency are currently the variables with the highest real weight in cloud providers’ and inference service providers’ procurement decisions.

The deployment of 50 racks of Atlas on Oracle Cloud, along with validation from technically demanding customers such as Jump Trading, at least proves that this architecture is not a lab concept. When Titan lands in H2 2027, whether the market is still waiting for it and whether competitors’ alternatives have become more attractive will be the real test of whether this $5 billion valuation can be realized.