Skip to main content
1M
NVDA$207.14+3.2%
MSFT$486.19+4.6%
GOOGL$375.19+5.4%
AAPL$303.05-1.9%
AMZN$285.86+5.3%
META$593.44+6.6%
TSLA$323.64+4.0%
IBM$226.69+1.4%
CRM$188.81+2.6%
AMD$480.57+0.9%
AVGO$388.72-0.1%
ARM$236.96-1.1%
TSM$403.29-0.2%
INTC$90.31+0.1%
QCOM$149.23+1.1%
MU$815.37-0.9%
CSCO$115.21-0.7%
ANET$181.33+0.5%
SMCI$28.45+0.2%
ORCL$137.58+5.9%
PLTR$125.14+1.7%
DELL$417.14+2.9%
HPE$49.03+2.4%
PSTG$77.17+2.8%
ALAB$317.11+1.9%
AMAT$511.65+0.8%
LRCX$290.96-0.7%
KLAC$179.81-1.6%
ASML$1,635.04+0.4%
TER$362.31-1.5%
GEV$997.86+0.8%
CEG$272.41+3.7%
UEC$9.82+2.3%
OKLO$41.30+6.4%
SMR$8.92+5.9%
BWXT$172.80+2.4%
VST$153.96+3.9%
D$68.80-0.5%
SO$94.22-0.3%
NEE$86.33-0.7%
CCJ$89.54+3.7%
LEU$183.35+3.6%
BE$217.12+5.5%
KMI$31.66-1.6%
EXC$45.68-0.3%
VRT$258.15+6.9%
ETN$432.93+4.3%
CAT$823.61+1.1%
PWR$674.29+1.0%
EME$814.20+2.1%
URI$1,103.76+2.3%
VMC$278.76+3.8%
J$138.17+2.4%
TT$461.02+1.3%
CARR$62.65+1.4%
JCI$146.68+0.0%
SIEGY$162.90-0.1%
ALB$119.04+1.2%
SQM$66.71-0.5%
LAC$2.97+3.5%
MP$43.73+5.7%
FCX$63.24+1.0%
GLW$145.80+5.5%
SCCO$185.35+1.4%
AA$44.55-1.6%
RIO$95.41-1.5%
VALE$14.57-3.3%
LYSDY$9.80-1.4%
EQIX$1,024.62+0.5%
DLR$191.93+1.8%
AMT$174.17+0.5%
LITE$761.92+6.7%
COHR$287.52+9.4%
CIEN$383.36+1.7%
IRM$124.97+2.2%
CCI$78.36+2.7%
MDB$357.81+6.0%
ADBE$253.37+1.2%
RIVN$15.58+2.4%
DDOG$275.55+2.8%
SNOW$312.84+6.7%
NOW$114.95+3.3%
PATH$13.02+2.0%
CRWD$197.10+3.3%
NVDA$207.14+3.2%
MSFT$486.19+4.6%
GOOGL$375.19+5.4%
AAPL$303.05-1.9%
AMZN$285.86+5.3%
META$593.44+6.6%
TSLA$323.64+4.0%
IBM$226.69+1.4%
CRM$188.81+2.6%
AMD$480.57+0.9%
AVGO$388.72-0.1%
ARM$236.96-1.1%
TSM$403.29-0.2%
INTC$90.31+0.1%
QCOM$149.23+1.1%
MU$815.37-0.9%
CSCO$115.21-0.7%
ANET$181.33+0.5%
SMCI$28.45+0.2%
ORCL$137.58+5.9%
PLTR$125.14+1.7%
DELL$417.14+2.9%
HPE$49.03+2.4%
PSTG$77.17+2.8%
ALAB$317.11+1.9%
AMAT$511.65+0.8%
LRCX$290.96-0.7%
KLAC$179.81-1.6%
ASML$1,635.04+0.4%
TER$362.31-1.5%
GEV$997.86+0.8%
CEG$272.41+3.7%
UEC$9.82+2.3%
OKLO$41.30+6.4%
SMR$8.92+5.9%
BWXT$172.80+2.4%
VST$153.96+3.9%
D$68.80-0.5%
SO$94.22-0.3%
NEE$86.33-0.7%
CCJ$89.54+3.7%
LEU$183.35+3.6%
BE$217.12+5.5%
KMI$31.66-1.6%
EXC$45.68-0.3%
VRT$258.15+6.9%
ETN$432.93+4.3%
CAT$823.61+1.1%
PWR$674.29+1.0%
EME$814.20+2.1%
URI$1,103.76+2.3%
VMC$278.76+3.8%
J$138.17+2.4%
TT$461.02+1.3%
CARR$62.65+1.4%
JCI$146.68+0.0%
SIEGY$162.90-0.1%
ALB$119.04+1.2%
SQM$66.71-0.5%
LAC$2.97+3.5%
MP$43.73+5.7%
FCX$63.24+1.0%
GLW$145.80+5.5%
SCCO$185.35+1.4%
AA$44.55-1.6%
RIO$95.41-1.5%
VALE$14.57-3.3%
LYSDY$9.80-1.4%
EQIX$1,024.62+0.5%
DLR$191.93+1.8%
AMT$174.17+0.5%
LITE$761.92+6.7%
COHR$287.52+9.4%
CIEN$383.36+1.7%
IRM$124.97+2.2%
CCI$78.36+2.7%
MDB$357.81+6.0%
ADBE$253.37+1.2%
RIVN$15.58+2.4%
DDOG$275.55+2.8%
SNOW$312.84+6.7%
NOW$114.95+3.3%
PATH$13.02+2.0%
CRWD$197.10+3.3%
NVDA$207.14+3.2%
MSFT$486.19+4.6%
GOOGL$375.19+5.4%
AAPL$303.05-1.9%
AMZN$285.86+5.3%
META$593.44+6.6%
TSLA$323.64+4.0%
IBM$226.69+1.4%
CRM$188.81+2.6%
AMD$480.57+0.9%
AVGO$388.72-0.1%
ARM$236.96-1.1%
TSM$403.29-0.2%
INTC$90.31+0.1%
QCOM$149.23+1.1%
MU$815.37-0.9%
CSCO$115.21-0.7%
ANET$181.33+0.5%
SMCI$28.45+0.2%
ORCL$137.58+5.9%
PLTR$125.14+1.7%
DELL$417.14+2.9%
HPE$49.03+2.4%
PSTG$77.17+2.8%
ALAB$317.11+1.9%
AMAT$511.65+0.8%
LRCX$290.96-0.7%
KLAC$179.81-1.6%
ASML$1,635.04+0.4%
TER$362.31-1.5%
GEV$997.86+0.8%
CEG$272.41+3.7%
UEC$9.82+2.3%
OKLO$41.30+6.4%
SMR$8.92+5.9%
BWXT$172.80+2.4%
VST$153.96+3.9%
D$68.80-0.5%
SO$94.22-0.3%
NEE$86.33-0.7%
CCJ$89.54+3.7%
LEU$183.35+3.6%
BE$217.12+5.5%
KMI$31.66-1.6%
EXC$45.68-0.3%
VRT$258.15+6.9%
ETN$432.93+4.3%
CAT$823.61+1.1%
PWR$674.29+1.0%
EME$814.20+2.1%
URI$1,103.76+2.3%
VMC$278.76+3.8%
J$138.17+2.4%
TT$461.02+1.3%
CARR$62.65+1.4%
JCI$146.68+0.0%
SIEGY$162.90-0.1%
ALB$119.04+1.2%
SQM$66.71-0.5%
LAC$2.97+3.5%
MP$43.73+5.7%
FCX$63.24+1.0%
GLW$145.80+5.5%
SCCO$185.35+1.4%
AA$44.55-1.6%
RIO$95.41-1.5%
VALE$14.57-3.3%
LYSDY$9.80-1.4%
EQIX$1,024.62+0.5%
DLR$191.93+1.8%
AMT$174.17+0.5%
LITE$761.92+6.7%
COHR$287.52+9.4%
CIEN$383.36+1.7%
IRM$124.97+2.2%
CCI$78.36+2.7%
MDB$357.81+6.0%
ADBE$253.37+1.2%
RIVN$15.58+2.4%
DDOG$275.55+2.8%
SNOW$312.84+6.7%
NOW$114.95+3.3%
PATH$13.02+2.0%
CRWD$197.10+3.3%
Modern GPU data center interior with rows of liquid-cooled server racks

The Rise of Neoclouds

How AI-specialized cloud providers became a $35 billion market in barely two years, and what happens next.

$35B
Neocloud market size (2026)
$100B+
Combined contracted revenue
65%
Cheaper than hyperscalers

A $35 Billion Market That Barely Existed Two Years Ago

A new category of cloud company grew from a niche idea into a $35 billion market in roughly two years. Not a gradual climb. A near-vertical one. As Nvidia's annual GTC conference wraps up in San Jose this week, that category is impossible to ignore. These companies have a name: neoclouds.

Nvidia used GTC to unveil its Vera Rubin platform with seven new chips. CoreWeave announced general availability of its most powerful server systems. Days before the conference, Meta finalized a $27 billion infrastructure deal with Nebius, and Nvidia disclosed a $2 billion investment in that same deal. The neocloud is no longer a footnote.

A new category of company quietly grew from a niche idea into a $35 billion market in roughly two years.

What Exactly Is a Neocloud?

Start with the cloud you already know. When a company says it runs on the cloud, it almost certainly means Amazon Web Services, Microsoft Azure, or Google Cloud. These hyperscalers operate at enormous, globe-spanning scale, offering hundreds of services: databases, file storage, email servers, video transcoding, website hosting. A one-stop shop for anything a modern business needs online.

A neocloud does one thing instead of hundreds. It rents access to GPUs (specialized chips originally designed for video games that turned out to be ideal hardware for training and running AI) and builds everything around that single purpose. No frills, no bundled extras, no legacy enterprise services. Just raw GPU power.

The term is relatively new. SemiAnalysis published a landmark report in late 2024 that gave the category a name and a framework. By March 2026, Reuters and Bloomberg use "neocloud" as a standard industry category. The vocabulary caught on because the business model had already proven itself.

The price difference tells the story. According to the Uptime Institute, an equivalent high-end GPU server instance costs roughly $98 per hour from a hyperscaler and around $34 per hour from a neocloud. Hyperscalers wrap their hardware in layers of virtualization software that let thousands of customers share the same physical machine. That adds overhead, complexity, and cost. Neoclouds skip most of it. Many offer bare-metal access, meaning customers rent the actual physical server rather than a virtual slice. Fewer layers, lower costs.

Think of hotel rooms versus furnished apartments. A big hotel chain (the hyperscaler) gives you a room divided and managed by a large staff operating hundreds of services. An apartment building (the neocloud) hands you the keys to a floor and gets out of the way. Less overhead, lower rent. You handle more yourself, but for AI companies burning through GPU compute at scale, that trade-off is almost always worth it.

Watch: GPUaaS Explained, Why CoreWeave and Others Are Fueling the Next Cloud Revolution (SiliconANGLE theCUBE, 14 min)

Three Companies Walk into a Cloud

The media treats "neocloud" as a single category. Three genuinely different businesses operate under that label, and mixing them up explains why one company prints money while another quietly bleeds out. The differences come down to position in the stack: who owns the hardware, who manages the software, and who connects buyers to sellers.

1 GPU Infrastructure Clouds

The real estate developers of AI

Think of a developer who buys land, builds apartments, and rents out units. Tier 1 neoclouds do the same with data centers and Nvidia server racks. They raise enormous capital, build or lease facilities, pack them with GPUs, and charge by the hour for raw computing power. You rent the actual physical machine (bare-metal access), not a virtual slice.

The economics are brutal and simple. A GPU cluster sitting idle earns nothing. Utilization (the percentage of GPUs actively running workloads at any given moment) needs to stay above 80% for the math to work after depreciation. CoreWeave carries a $66.8 billion revenue backlog, the largest in the sector. Nebius posts 479% revenue growth year-over-year as it expands across Europe.

CoreWeave Nebius Lambda Crusoe Applied Digital RunPod
2 Managed Inference and Model Platforms

The restaurant kitchen, not the farm

A farm grows food. A restaurant kitchen turns that ingredient into something a customer can use: fast, consistent, at scale. Tier 2 neoclouds do not own the GPUs. They run AI models on top of them as efficiently as possible. The product is inference (taking a trained AI model and running it to generate outputs, the process behind every ChatGPT response). Inference sounds simple but grows operationally complex at millions of requests per day.

These companies build software for autoscaling (spinning GPU capacity up when demand spikes and down when it drops, so customers never pay for idle machines), request routing, and latency optimization. Training a model is a one-time cost. Running it at commercial scale is the ongoing expense. Inference will represent two-thirds of all AI compute by end of 2026. Fireworks AI processes over 10 trillion tokens per day across 10,000+ companies. Baseten runs the AI infrastructure behind Cursor, Notion, and Mercor.

Baseten Fireworks AI Together AI

Training a model is a one-time cost. Running it millions of times per day is the ongoing expense. Inference will be two-thirds of all AI compute by end of 2026.

3 Control Planes and Abstraction Layers

The Kayak of AI compute

When you book a flight on Kayak, you are not flying on Kayak. You search across Delta, United, and Southwest simultaneously, and Kayak routes you to the best price and availability. Tier 3 neoclouds do this for GPU compute. They aggregate capacity from dozens of Tier 1 and Tier 2 providers, give developers a single API or dashboard, and handle routing in the background. The developer never needs to know which data center processed the request.

Hugging Face routes inference traffic across 15+ providers through a single endpoint; developers switch GPU clouds without changing a line of code. Shadeform aggregates 20+ GPU clouds into one console. Nvidia runs its own version through DGX Cloud Lepton, notable because Nvidia simultaneously supplies the hardware every other tier depends on. These businesses live and die by price arbitrage and availability, making them inherently competitive with the Tier 1 providers they route to.

Hugging Face Inference Providers Shadeform DGX Cloud Lepton

This taxonomy matters because the three tiers have different risk profiles, capital requirements, and moats. Tier 1 lives or dies by its ability to raise billions and keep utilization high. Tier 2 needs engineering talent and model-developer relationships. Tier 3 needs network effects and switching costs. When a headline says a neocloud is struggling or thriving, the first question: which tier?

The Deals That Defined Q1 2026

Start with a single data point: Nebius, a company that barely existed two years ago, has signed contracts worth more than $46 billion in the past six months. Nebius was spun out of Russian tech giant Yandex in 2024, an obscure corporate restructuring most people ignored. Today it is the fastest-growing infrastructure company on the planet.

In September 2025, Microsoft signed Nebius to a $19.4 billion deal. On March 11, 2026, Nvidia disclosed a $2 billion strategic investment for an 8.3% stake. Five days later, Meta signed a five-year, $27 billion AI infrastructure agreement: $12 billion in dedicated compute, with $15 billion more if demand runs hotter than expected. That deal builds on a $3 billion pilot Meta ran with Nebius since late 2025, a trial that went well enough to trigger one of the largest infrastructure commitments in tech history. Nebius generated $529.8 million in revenue in 2025 (+479% YoY), guides toward $7 to $9 billion annualized by end of 2026, and plans $16 to $20 billion in capex (capital expenditure: money spent building data centers and buying GPU clusters) to get there.

CoreWeave tells a different story: what happens when you grow fast enough to matter but borrow heavily to get there. The company went public on March 28, 2025, raising $1.5 billion at $40 per share in an IPO (selling shares to the public for the first time). The stock surged past $187, then reversed: Q4 2025 brought a $452 million net loss (nearly double expectations) and $388 million in interest expense as $14 billion in debt became impossible to ignore. The stock dropped 20%, and a securities class action followed. None of that changes the demand story. CoreWeave's backlog (signed contracts for future revenue not yet delivered) grew 342% to $66.8 billion, anchored by $22.4 billion from OpenAI and $14.2 billion from Meta. Full-year 2025 revenue hit $5.13 billion (+168%), with guidance of $12 to $13 billion for 2026. The question: whether cash from those contracts arrives fast enough to service the debt financing the infrastructure today.

Largest neocloud contracts (as of March 2026)
Deal value in billions USD
Nebius CoreWeave Applied Digital
Nebius (Meta)
$27B
CoreWeave (OpenAI)
$22.4B
Nebius (Microsoft)
$19.4B
CoreWeave (Meta)
$14.2B
Applied Digital
$11B+

Baseten raised $300 million at a $5 billion valuation in January 2026, with IVP and CapitalG co-leading and Nvidia contributing roughly $150 million. Baseten builds the systems software that makes AI models run reliably in production. Its customers (Cursor, Notion, Lovable) represent the current generation of AI-native products that need inference to be fast and predictable. Total funding: $585 million.

Fireworks AI launched on Microsoft Azure AI Foundry in public preview on March 11, embedding itself inside a hyperscaler rather than competing with one. It also acquired Hathora, a real-time server orchestration platform from the gaming industry, an odd fit until you hear the reasoning.

Gamers will tolerate lower graphics, but they will not tolerate lag. The same is true for AI inference users.

Lin Qiao, CEO of Fireworks AI

Fireworks processes 10 trillion+ tokens per day and generates $280 million in annualized revenue (up from $44 million in 2024) at a $4 billion valuation, all without building a single data center. Together AI made its GPU Clusters product generally available in February 2026, offering self-provisioned access to 8 to 100,000+ GPUs. At GTC, Together announced partnerships on Nvidia's Dynamo 1.0, FlashAttention-4, and Mamba-3. Total raised: $534 million at a $3.3 billion valuation.

The deals getting less attention may be the most instructive. RunPod: 500,000+ developers, $120 million ARR, no Series A, profitable. Lambda raised $1.5 billion (Series E, $4 billion valuation). Crusoe powers its data centers on stranded energy that would otherwise be flared, raised $1.38 billion, and sits at the center of OpenAI's $500 billion Stargate project. Applied Digital has $11 billion+ in contracted revenue. The number of companies at meaningful scale (not demos, but hundreds of millions in real revenue) separates this moment from every previous cloud buildout cycle.

Company Latest round / valuation Revenue or ARR Notable
Nebius$2B Nvidia investment (8.3% stake)$529.8M (2025, +479%)$46B+ in contracts signed
CoreWeaveIPO Mar 2025, $40/share$5.13B (2025, +168%)$66.8B backlog; $14B+ debt
Baseten$300M / $5B valuation (Jan 2026)UndisclosedNvidia contributed ~$150M
Fireworks AI$4B valuation$280M annualized (+536% YoY)10T+ tokens/day; Azure partnership
Together AI$534M raised / $3.3B valuationUndisclosedGPU Clusters GA; Nvidia GTC partnerships
Lambda$1.5B Series E / $4B valuationUndisclosedGPU cloud pioneer
Crusoe$1.38B raisedUndisclosedStranded energy; Stargate infrastructure
RunPodNo Series A (profitable)$120M ARR500K+ developers; bootstrapped
Applied DigitalPublic company$11B+ contracted revenueHPC and AI data centers

Three Forces That Made Neoclouds Inevitable

Force 1: Hyperscalers Cannot Build Fast Enough

Think about a fast-growing city. Demand for apartments shoots up, but new buildings take years: permitting, construction, utilities. AI infrastructure has the same problem, worse. Building a data center from scratch takes two to four years. Joining the interconnection queue (the line to hook up to the electrical grid) can take seven. Meanwhile, demand for GPU compute grows five times faster than new data center construction, according to KPMG.

Neoclouds solve this by retrofitting existing buildings. They lease warehouse space that already has power, add GPU infrastructure, and bring capacity online in months. McKinsey projects AI workloads will require 200 gigawatts of data center capacity by 2030, a 3.5x increase. No single company can build that fast alone. Microsoft signed more than $33 billion in external GPU capacity deals to bridge the gap. Hyperscalers choose neoclouds because the alternative is falling behind.

Force 2: The Shift from Training to Inference

Training and inference are fundamentally different operations. Training is like building a factory: assemble everything, run hard for days or weeks, done. Inference is operating that factory every day, around the clock, for every customer. Through 2024, the industry spent most of its GPU budget on training. That balance is shifting fast.

By 2030, inference is projected to account for 80 to 90 percent of all AI compute demand. Inference runs continuously, responds to individual requests in real time, and scales with user traffic rather than grinding through a fixed job. That environment rewards software optimization, latency management, and operational efficiency over raw GPU count. Neoclouds specializing in managed inference are building positions genuinely hard to replicate.

The inference economy rewards software optimization, not just raw GPU count.

Automated factory floor representing continuous AI inference operations
Inference is like running a factory 24/7 for every customer. The shift from training to inference is reshaping who wins in AI infrastructure.

Force 3: The Price Gap Is Still Enormous

Equivalent GPU instances on neoclouds cost 40 to 66 percent less than AWS or Azure. For startups training models or serving AI features, this determines whether a product is viable. Hyperscalers charge more partly because of bundling: hundreds of ancillary services whether you need them or not. Neoclouds strip that away and charge for compute alone.

Mordor Intelligence estimates the market will grow from $35.2 billion in 2026 to $236.5 billion by 2031, a 46% CAGR (compound annual growth rate). Synergy Research Group projects $180 billion by 2030 (69% annual growth). Both reflect the same reality: demand for affordable GPU compute is not slowing, and hyperscalers cannot fully capture it.

Projected neocloud market size
Revenue in billions USD. 46% CAGR projected through 2031. Sources: Mordor Intelligence, Synergy Research Group
2024
$8B
2025
$18B
2026
$35.2B
2027
$55B
2028
$82B
2029
$120B
2030
$180B
2031
$236.5B
Watch: Neoclouds vs Hyperscalers Explained (NEXTDC NEXTnow Podcast, 5 min)
GPU chip with green accent lighting on a dark reflective surface
Nvidia's chips power nearly every neocloud. That dependency is both an opportunity and a risk.

Why Nvidia Is Bankrolling the Competition's Landlords

Nvidia invested at least $4 billion directly into neoclouds: $2 billion into CoreWeave, $2 billion into Nebius, plus stakes in Lambda and Nscale (which closed a $2 billion Series C in March 2026). The reason: all three major hyperscalers are developing proprietary AI chips. Amazon has Trainium, Google has TPUs, Microsoft has Maia, Meta is building MTIA. If those chips work at scale, customers migrate off Nvidia hardware, eating directly into core revenue.

Neoclouds are Nvidia's insurance policy. Think of a supplier investing in retailers who sell exclusively its products. If big-box stores start stocking competing brands, dedicated retail partners become more valuable. Neoclouds run exclusively on Nvidia hardware by design, with no incentive to switch. CoreWeave and Nebius have each committed to deploying 5+ gigawatts of Nvidia capacity by 2030. The $4 billion buys committed buyers for hardware hyperscalers might otherwise replace.

Neoclouds are cloud providers that rely exclusively on Nvidia hardware. They are Nvidia's counterweight to the hyperscalers' chip ambitions.

The strategy runs deeper than hardware sales. Nvidia started with DGX Cloud, renting dedicated GPU instances at ~$37,000/month, then evolved it into DGX Cloud Lepton, a marketplace aggregating capacity across dozens of neocloud partners. Through this platform, Nvidia controls the software layer (NIM inference microservices, NeMo training framework) while neoclouds carry the capital expenditure of buying and housing hardware. Nvidia collects software revenue and ecosystem leverage without owning a single data center.

Hardware allocation makes the logic cleaner still. The latest Nvidia racks (GB200 and GB300 NVL72 systems, liquid-cooled units drawing ~120 kilowatts each at ~$3.1 million per rack) go to hyperscalers and the largest neoclouds first. The same priority applies to Vera Rubin, Blackwell's successor announced at GTC. A $2 billion investment buys preferred status: newest hardware before competitors. For neoclouds, priority hardware means the ability to sign contracts. For Nvidia, it creates cloud providers structurally dependent on its roadmap.

Glass office building at night with storm clouds and lightning
The growth is real. But so are the structural risks that come with building an industry on borrowed capital and borrowed time.

What Could Go Wrong

One chip to rule them all

Nearly every neocloud runs on Nvidia silicon. That is a structural dependency, not a coincidence. AMD remains a distant alternative because of CUDA (Nvidia's programming toolkit for its GPUs), which 4 million+ developers have built their workflows on over two decades. That ecosystem cannot be replicated in a product cycle or two. If Nvidia shifts allocation, raises prices, or competes more directly in cloud services, neoclouds have limited room to maneuver. They built on Nvidia's foundation because they had to. That foundation is also a ceiling.

Your biggest customer is also your biggest threat

CoreWeave derived 62% of its 2024 revenue from Microsoft alone. Nebius's backlog is dominated by Meta and Microsoft. These are existential relationships. Hyperscalers are simultaneously the neoclouds' largest customers and most capable competitors. Microsoft signed $33 billion+ in neocloud deals while investing $80 billion+ annually in its own infrastructure. When those numbers converge, and hyperscalers insource currently outsourced workloads, neocloud utilization rates will feel the shift immediately.

The margin math is brutal

Oracle internal data shows GPU rental gross margins (revenue remaining after direct costs) at just 14 to 16% after depreciation (equipment losing value over time), labor, and power. CoreWeave posted $1.6 billion in Q4 2025 revenue alongside a $452 million loss and $388 million in interest expense on $14 billion+ in debt, with $30 to $35 billion more in capex planned for 2026. McKinsey calls bare-metal GPU economics "almost no margin of safety," with returns flatlining below 80% utilization. GPU rental prices have fallen 60 to 75% from their 2023 peak, and the H100s and A100s neoclouds staked their businesses on depreciate on three-year cycles.

GPU-collateralized debt (loans backed by physical GPUs as assets) has never been tested through a sustained AI spending downturn. If enterprise budgets tighten or adoption slows, idle GPUs still incur depreciation, power costs, and financing charges. Unlike software companies with near-zero marginal costs, neoclouds have a hard floor on operating expenses that does not move.

There is a counterargument worth taking seriously. SemiAnalysis calls it the value cascade: as GPUs age out of frontier training, they move into inference workloads, then batch processing, extending their economic life to five or six years rather than the two or three that standard depreciation assumes. Google, Oracle, and Microsoft have all converged on six-year depreciation schedules. CoreWeave's CEO has said that H100s from expired contracts were rebooked at 95% of original price, and that A100s (announced in 2020) remain fully booked for inference in 2026. Deloitte projects inference will account for two thirds of all AI compute by the end of this year, and that growing demand structurally supports longer GPU lifespans.

The catch: this thesis holds best when supply is tight. As Nvidia ramps Blackwell production and more chips flood the market, the floor under older GPUs could soften. Michael Burry has argued that the real useful life is closer to two or three years, and H100 rental prices have already dropped over 60% from their peak. The value cascade is plausible, maybe even likely in the near term. But neoclouds betting their balance sheets on it are making a directional wager on sustained inference demand that has not yet been tested through a full cycle.

The Neocloud Paradox

To escape thin infrastructure margins, neoclouds need to move up the stack: build managed inference platforms, orchestration software, enterprise tooling, where software economics replace infrastructure economics. Software scales. Servers do not. But moving up the stack puts neoclouds in direct competition with customers who represent 50 to 60% of their revenue. Azure AI Foundry, AWS Bedrock, and Google Vertex AI already offer those managed AI services with more distribution, trust, and data than any neocloud has.

Imagine a restaurant that gets 60% of its business catering to one hotel chain. To grow, the restaurant needs to open its own dining room. But now it is competing with the hotel's restaurant, and the hotel controls the building.

That is the neocloud paradox. Standing still means accepting commodity margins that barely cover the debt service. Growing means picking a fight with the customers keeping the lights on. There is no clean path through it, only execution.

Data center building on the horizon at golden hour
The window that created neoclouds is real. Whether it stays open depends on what they build while it lasts.

The Window Is Real, But It Will Not Stay Open

Combined, the sector has locked in over $100 billion in contracted revenue. Market projections put GPU cloud infrastructure at $180 to $236 billion by 2030. These reflect real demand from companies that cannot build fast enough on their own.

But the conditions that created neoclouds are not permanent. They emerged from a convergence: explosive AI demand, constrained GPU supply, and hyperscaler capacity gaps. Grid capacity is expanding. Custom silicon at Microsoft, Google, and Amazon is maturing. GPU supply is normalizing. The structural gap is already narrowing, and the companies still standing in five years will be those that built something that does not depend on the gap remaining open.

What to watch: whether neoclouds develop software differentiation before their biggest customers scale their own infrastructure. Whether GPU-collateralized debt survives its first real stress test (the CoreWeave class action is an early signal). And whether any neocloud escapes the commodity layer without losing the enterprise relationships that got them there.

The next three years will answer the question this industry has deferred: are neoclouds a durable layer of cloud infrastructure, or a transitional chapter the hyperscalers eventually absorb? The contracted revenue says the former is possible. The margin structure says it will not happen automatically. The difference will be decided by companies most of us have not heard of yet.

Sources

Nebius-Meta deal: CNBC (Mar 16, 2026), Nebius press releases, Yahoo Finance/Reuters

Nvidia-Nebius investment: Bloomberg/BNN (Mar 11, 2026), Nvidia Newsroom, 24/7 Wall St.

Nvidia Vera Rubin / GTC 2026: Nvidia Newsroom (Mar 16, 2026), Nvidia blog

CoreWeave financials: Constellation Research, SEC filings (Q4 2025), CoreWeave investor relations

CoreWeave Q4 loss & lawsuit: GlobeNewsWire (Mar 11, 2026), Motley Fool (Mar 16, 2026)

Baseten funding: HPCwire, Business Wire, Bloomberg (Jan 23, 2026)

Fireworks AI: Microsoft Azure Blog (Mar 11, 2026), SiliconANGLE (Mar 9, 2026), GamesBeat, Hathora blog

Together AI: Data Center Dynamics (Feb 2026), Together AI blog, SiliconANGLE

Lambda, Crusoe, RunPod, Applied Digital: Data Center Dynamics, TechCrunch, Yahoo Finance

Market sizing: Mordor Intelligence, Synergy Research Group

Industry analysis: McKinsey ("The Evolution of Neoclouds and Their Next Moves"), KPMG, Equinix Blog

Neocloud economics: The Information (Oracle margin data), Data Center Dynamics

Nvidia strategy: MLQ Research, Futurum Group, Data Center Dynamics (Lepton/DGX Cloud)

SemiAnalysis taxonomy: "AI Neocloud Playbook and Anatomy" (2024)

Uptime Institute: GPU instance pricing comparison (hyperscaler vs. neocloud)

futuretechwithyanyan.com · March 20, 2026