The B200 is the flagship GPU of NVIDIA's Blackwell architecture. Per our coverage: 208 billion transistors, roughly 2.5x the H100; a 5x jump in FP8 training performance for large models; and about 25% better energy efficiency. NVIDIA later unveiled a B200 Blackwell Ultra edition at GTC 2026 aimed at AGI workloads, claiming up to 30x inference over the H100. Treat all such multipliers as vendor numbers under specific configurations — they vary widely across precisions and cluster scales.
On the software side, Blackwell-native low-precision formats like MXFP8 and NVFP4 are being fully exploited by training and inference stacks — see Miles' Blackwell-native 8/4-bit RL work, and Nous Research training the open NousCoder-14B coding model on 48 B200s in just four days.
The B200 has been oversubscribed since its debut, with first deliveries once delayed (early report: B200 orders full, deliveries pushed back). By April 2026 the shortage went mainstream: daily mentions of B200 stockouts on X surged from 8,000 to 58,000 (+625%), supply-chain investigations by Reuters and Bloomberg confirmed the gap, and NVIDIA executives admitted to "unprecedented demand" — see our global shortage report.
The buyer list tells the story: Oracle issued over $25B in bonds this year to build H100/B200 clusters (while laying off 21,000); Reflection signed a $1B five-year compute deal with Nebius built on H100 and B200 clusters; and six-month-old Hark listed "secured NVIDIA B200 data-center resources" as a highlight of its $700M Series A.
The B200's performance leap comes with higher power draw — and liquid cooling itself adds pump and heat-exchanger loads. The macro consequence landed in Texas: ERCOT's queue held about 1,800 projects, 90% of them data centers, with 474 GW of cumulative demand — beyond the grid's existing peak capacity. A single planned AI data center often draws 500 MW, the full output of a mid-size gas plant. Texas responded by freezing data-center grid-connection approvals.
Regulators had seen it coming: the state utility commission asked 377 data-center firms for water and power data, only 28 responded, and the audit expanded to some 300 projects (see the first grid-freeze report). For teams siting their own compute, power and cooling are now as binding a constraint as the chips themselves.
The B200, like the H100, sits at the center of US export controls on China, which the Commerce Department tightened further in 2026. The flip side is a black market: an indictment unsealed in Manhattan federal court alleges that at least $2.5B of US-built AI servers (carrying B200s, H200s and more) flowed into China via Taiwan and Southeast Asian shell companies in 2024–2025 — over $510M in April–May 2025 alone, including one batch of 58 B200 servers worth about $17.4M. See our breakdown of the Supermicro smuggling indictment.
Supply uncertainty is also forcing substitution: DeepSeek, long a heavy H100/B200 user, was reported to have launched an in-house chip project codenamed "Aurora", running stockpiles, domestic chips and self-design in parallel.
The strongest head-on challenge comes from AMD: the MI355X, paired with SGLang and the MoRI communication library, hit $0.169 per million tokens in disaggregated DeepSeek-R1 inference — matching or beating the B200 (Dynamo + TRT-LLM) on TCO at key operating points, verified on SemiAnalysis' InferenceX platform. See the MI355X TCO report. Intel is playing the low-cost, low-barrier card with the air-cooled, LPDDR5-based Crescent Island.
The deeper variable is customers designing their own silicon: OpenAI's Broadcom-partnered Jalapeño inference chip joins Google, Amazon and SpaceX in the in-house camp. The B200 remains the de facto standard for training — but the era of "no second choice" is ending.
The B200 is the flagship AI accelerator GPU of NVIDIA's Blackwell architecture: 208 billion transistors (about 2.5x the H100), 5x FP8 training performance, and roughly 25% better energy efficiency. It is the de facto standard for large-model training clusters, with a Blackwell Ultra edition for AGI workloads unveiled at GTC 2026.
Demand is an industry-wide compute arms race — cloud providers, AI labs and sovereign buyers are all queuing. In April 2026, daily B200-shortage mentions on X jumped 625%, Reuters and Bloomberg confirmed the supply gap, and NVIDIA executives called demand "unprecedented". Supply is constrained by advanced-node and HBM memory capacity, which is why Wall Street treats HBM suppliers like Micron as the next beneficiaries.
Per vendor figures: about 2.5x the transistors, 5x FP8 training performance, +25% efficiency, and up to a claimed 30x inference in Blackwell Ultra configurations. Real-world gains depend heavily on precision formats (MXFP8/NVFP4), interconnect and the software stack — results vary widely by scenario, so benchmark your own workload before deciding.
Not legally. The B200 and H100 are both on the US export-control list for China, and controls tightened further in 2026. A Manhattan federal indictment alleges at least $2.5B of related servers were smuggled in via third countries during 2024–2025 — a gray channel now under active prosecution. Chinese firms are responding with stockpiles, domestic chips and in-house designs, such as DeepSeek's reported chip project.