The neocloud thesis was simple: the hyperscalers were too slow and too rationed to meet AI compute demand, so a provider that bought GPUs aggressively and rented them flexibly could name its price. For three years, that thesis printed money.
The sector now operates data centers measured in gigawatts and contracts measured in billions. It also carries debt against hardware that depreciates on a three-to-five-year curve while its successors arrive annually. That tension defines the next phase.
Why it matters
Neoclouds are no longer a sideshow: they train frontier models, serve a large share of startup inference, and act as the swing capacity of the whole AI economy. If their economics break, the effects propagate into model-lab costs, GPU order books and the data-center buildout itself.
They're also a live experiment in a question every infrastructure industry eventually faces: is compute a utility with utility margins, or a scarce asset with scarcity rents? The answer determines what the AI boom's infrastructure layer is actually worth.
How it works
The business model is arbitrage across time and customers: buy capacity at scale, finance it against contracts, and keep utilization high enough that revenue outruns depreciation plus debt service. The levers are contract length (long reserved deals de-risk, short on-demand deals pay more), hardware mix (training clusters versus inference-optimized fleets), and the secondary market for used GPUs, which sets the residual value that makes the financing math work.
The risks compound: a generation of GPUs that loses value faster than expected, or a softening in inference demand, hits utilization and residual value simultaneously — the two variables the model can least afford to move together.
Evidence
Public filings and reporting show the largest neoclouds carrying multi-billion-dollar revenue backlogs alongside substantial debt facilities secured against GPU fleets. Hyperscalers have responded with their own capacity expansion and with price cuts on older GPU instances — an early sign that scarcity pricing is eroding at the trailing edge even as leading-edge capacity stays tight.
Meanwhile, GPU rental price trackers show a widening spread: latest-generation accelerators command premiums, while two-generation-old hardware rents at a fraction of its 2024 rate. That spread is the depreciation problem made visible.
The competing read
Bulls argue demand elasticity will save the economics: every price cut has historically expanded AI usage faster than it cut revenue, and inference demand is still in its early innings. Bears see the airline pattern — capital-intensive capacity, commodity product, brutal cycles — and note that the largest customers are already building their own fleets, which caps the neoclouds' pricing power exactly when their debt comes due. The truth likely splits by operator: those with long contracts and cheap capital endure; those running a spot-market book do not.
What happens next
Watch utilization disclosures and used-GPU price indices — the two honest numbers in the sector. Also watch consolidation: the next downturn in GPU scarcity will sort the operators with real operations discipline from those who were, in effect, leveraged GPU traders with a data center attached.
