Author | Wei Rong Editor | Bao Yonggang
A year ago, Arkady Volozh, founder and CEO of AI cloud provider Nebius, was still answering a question: Who will buy so many GPUs? A year later, his judgment has shifted to the conclusion that the real bottleneck is how much compute capacity humanity can actually build.
In 2025, he primarily explained how a neocloud is established from three dimensions: technology, capital, and market. By 2026, the question is no longer just "how to build a neocloud," but rather how cloud infrastructure needs to be rebuilt in the AI era.
In Volozh's description, this new AI cloud extends downward to land, power, data centers, and racks, and upward to cover foundational cloud services, inference, and agent services. The scale of the infrastructure supporting it could be hundreds or even thousands of times larger than past construction scales, thus requiring "a completely new set of tools."
Leifeng.com compared Arkady Volozh's two interviews in Spotlight On in 2025 and 2026, revealing a shift in mindset: neoclouds are beginning to evolve from a GPU infrastructure supply model to a next-generation AI cloud foundation.
01. The Three Pillars of Neoclouds: Technology, Capital, and Market
For neoclouds, owning GPUs has never equated to owning a cloud.
Essentially, a GPU is a compute resource. To continuously deliver these resources to customers, data centers, racks, networking, software stacks, sufficient capital, and customer demand are also required. For early neocloud providers, competition was not just about acquiring GPUs, but also proving their ability to truly organize expensive compute resources into a cloud service.
In 2025, Arkady summarized this capability as "three pillars": technology and talent, capital, and market sales. The first pillar is technology and talent. Nebius originated from the former Dutch listed entity of Yandex. Following the Russia-Ukraine war, the former Yandex N.V. sold its Russian business in 2024. The retained team includes a group of engineers who have long been building large-scale machine learning infrastructure, as well as data centers located in Finland.
Nebius's ability to enter this market is largely attributed to the engineering team, data center construction experience, and capital base left by the former Yandex. Volozh emphasized at the time that the team can build data centers and racks independently, and build a software stack on top of the hardware, calling it a "full stack AI cloud."
The second pillar is capital. AI infrastructure requires massive upfront investments in chips, power, and data centers before revenue is realized. Technology determines whether the system can run, while capital determines whether it can scale. Nebius inherited the capital base of the former listed entity, holding over $2 billion in cash on its balance sheet with no debt; even so, Volozh stated they would continue to raise funds.
"You need good technology and people who understand it, as well as substantial capital to build AI infrastructure," Volozh said.
The third pillar is market sales. After the server rooms, racks, and software are in place, the company still needs to find real customers. Volozh judged at the time that the earliest batch of customers for AI infrastructure is mainly concentrated in the US: on one end are large tech companies, hyperscale cloud providers, and major ecosystem companies competing for frontier models; on the other end are AI-native startups developing models and AI applications.
However, he does not believe the market will remain confined to these two types of customers in the long term. As AI gradually integrates into daily applications, customers will continue to expand from AI-native companies to digital-native enterprises, and further penetrate the general enterprise market. Industries such as finance, pharmaceuticals, and robotics will progressively become users of AI infrastructure. Volozh even predicted at the time that a broader range of enterprise customers would become one of the primary markets in a few years.
These three pillars explain why "having GPUs" does not equal "having a cloud." A GPU is merely a compute resource, whereas a cloud must solve how to organize a massive number of GPUs into a stable service that customers can continuously invoke: where they are located, how they are powered and networked, how they are managed by software, and how they are delivered according to different customer requirements.
Therefore, the "full stack" concept in 2025 was first and foremost a proof of capability. The company needed to demonstrate to the market that Nebius not only possessed a batch of popular chips at the time, but also had the team to rebuild hardware and software, the capital to build the infrastructure, and the ability to find customers.
However, the interviews at that time already laid the groundwork for the next phase. Neoclouds face not only giants capable of purchasing massive amounts of compute capacity at once, but also a large number of AI-native startups, as well as enterprise customers continuously entering the market in the future.
For these customers, the question will ultimately not stop at "whether there are GPUs."
02. Large-Scale Bare Metal Contracts Are Turning Into Financing Tools
As AI customers begin to segment, neocloud products are also stratifying. For hyperscale tech companies like Microsoft and Meta, purchasing bare metal capacity from Nebius already meets their needs, as they possess their own software stacks and engineering capabilities and do not require cloud providers to supply a large number of tools further up the stack.
Those who truly rely on cloud providers' tools and platforms are a vast number of AI-native startups and enterprise customers: they lack the capability to build tens of thousands of GPUs on their own and lack complete software stacks, requiring cloud service providers to further encapsulate the underlying compute capacity into directly usable services. In his 2026 interview, Volozh directly referred to this segment of customers as Nebius's "primary market" and "core business."
"We do not treat these large bare metal contracts as our core business, but rather as one of our financing methods to build our own cloud and serve other customers in the market," Volozh said.
What Nebius aims to build is not a larger server room warehouse, but a cloud adapted to new users. This infrastructure system built entirely for the AI era demands greater scale and more tools, while also requiring the company to control a longer infrastructure chain.
This customer segmentation and positioning also introduces two distinct dimensions to AI cloud competition. The first is scale. AI infrastructure is increasingly resembling a heavy-asset construction race: hundreds of thousands, or potentially millions, of GPUs in the future require support from land, power, data centers, and continuous construction capabilities.
The other is product depth: the same batch of GPUs can be sold in bulk as bare metal, or foundational cloud services, platforms, and inference can be layered on top, ultimately transforming into token and agent services. Volozh summarizes this competition into two dimensions: "scale" and "product depth."
These two dimensions are intertwined: focusing solely on bare metal allows providers to serve large customers with mature software stacks, but makes it difficult to cover small and medium-sized customers who need comprehensive tools; focusing only on software and upper-layer services means that without stable underlying capacity, businesses like inference will be constrained by compute supply.
Therefore, AI cloud competition is beginning to extend in two directions simultaneously: upward to compete for higher-level software and services, and downward to continuously delve deeper into land, power, and data centers.
Volozh places the starting point on land: first acquiring land and industrial permits, then resolving power permits and access, followed by constructing data centers and internal electrical facilities; moving upward to racks, cluster software, and foundational cloud services, and finally to inference and agent products, ultimately building a complete and massive full-stack AI infrastructure from the bottom up.
Why must it start with land? His answer is almost uncompromising: "If it is outsourced, it can never be efficient." For every layer outsourced, the company must pay profits to data center builders, rack manufacturers, or software companies; more importantly, systems scattered across different suppliers are difficult to optimize as a whole. Vertical integration recaptures not just profits, but also construction speed, engineering efficiency, and expansion pace.
In Volozh's framework, bare metal is not even the endpoint of a typical neocloud business, but merely the starting point for Nebius's AI cloud.
However, such an architectural model is destined to make some neoclouds increasingly "heavy." In the early stages, to acquire capacity quickly, providers could entirely use third-party colocation data centers.
Nebius itself once adopted this approach, even though Volozh admitted that smaller colocation facilities are more expensive, because the most important thing at the time was to go online as quickly as possible. But as scale expands, the large-scale sites the company is building are reaching hundreds of megawatts or even gigawatt levels.
The industry logic behind this is not complex: colocation solves time-to-market, while self-building solves whether large-scale AI infrastructure can expand long-term according to its own cost and engineering requirements.
The other side of this model is that large customer contracts are also incorporated into the expansion cycle. Microsoft's bare metal contract of over $17 billion can be used to secure financing from banks; Meta's arrangements include approximately $12 billion in bare metal procurement and about $15 billion in offtake guarantees, providing a financing foundation for more cloud capacity.
Thus, Volozh describes a more complete AI cloud expansion model: large customers provide scale orders and credit, while small and medium-sized customers provide the market for higher-level cloud services; long-term contracts support financing, which is then transformed into new data centers, power, and GPU capacity.
The bare metal business, which neoclouds excel at, has not disappeared; it has merely shifted from the endpoint to one layer within this business model.
03. GPUs Move Behind the Scenes as AI Clouds Continue to Pioneer a Blue Ocean
As AI clouds continue to extend upward, what changes is not only the product form, but also the metering methods and users of compute capacity.
At the bare metal layer, customers are still purchasing compute resources such as GPUs, clusters, and usage time; whereas at the inference layer, compute capacity begins to adopt another metering method: it is no longer sold by GPU hours, but tokens are sold directly.
In the stratification constructed by Volozh, the long-term, large-scale sale of bare metal compute capacity is precisely the business of a "typical neocloud," but he believes Nebius does not intend to stop here. Instead, it will continue upward, building an Infrastructure as a Service (IaaS) layer on top of cluster software, and a platform to provide higher-level services to customers above the IaaS layer.
"What we are building is for a new generation of customers who have completely different needs," Volozh said.
Under this vision of serving the new generation, Volozh believes that the new generation of customers actually think about problem-solving in a completely different way from the customers of classic clouds a decade ago. The new generation of AI cloud users are beginning to use the cloud in a more "agent-centric" way; they may not even care as much about where the compute is specifically sent, nor do they need to piece together over a hundred different tools to use.
"Because these tasks are completed by agents, and the agents themselves will optimize efficiency," Volozh emphasized. What the new generation of users values more is a platform that is quantifiably more reliable, more efficient, and cheaper.
For users, GPUs are being hidden in the future; whereas for providers, to make tokens faster and cheaper, power, data centers, racks, and software must be highly aligned. The more the upper-layer delivery and metering methods tend toward standardization, the more the underlying engineering actually requires high integration.
This is also why Volozh believes there is still a window in the market. In his judgment, the AI cloud is not a red ocean where traditional providers and neoclouds fight each other to the death, but a blue ocean opened up by new users, new demands, and new services. "Facing these new demands and new users, traditional hyperscale cloud providers do not have the same protective moats as in their original businesses. We are not trying to steal their existing business, but to build new businesses together with them. This is also not a winner-takes-all consumer market, but a B2B market," Volozh stated.
Traditional cloud providers still have a massive customer base, capital, and software ecosystems. But Volozh is betting that when AI brings a batch of new customers with completely different usage methods, existing advantages may not automatically cover all newly added markets, thus giving new companies a window to re-enter the cloud computing table.
For neoclouds attempting to expand toward a full-stack AI cloud, expansion is happening simultaneously in two directions: downward to control land, power, and server rooms to reduce costs and expand construction scale; upward to enter token and agent services to adapt to new users and obtain higher-level revenue. For providers expanding upward, if they remain at the GPU and bare metal layers, they still need to bear heavy-asset construction and financing costs, while what they primarily sell to customers is still underlying compute resources.
However, moving upward does not mean the compute shortage has ended. Arkady expects that over the next year, compute capacity will remain the bottleneck, limiting how much infrastructure humanity can actually build to meet new demands. At a stage where demand is sufficiently high, new capacity can be sold in bulk as bare metal, or placed at a higher product level and sold in the form of tokens and AI services.
Starting from bare metal and GPU resources, some neoclouds are attempting to move toward another ultimate outcome: becoming a full-stack cloud rebuilt for a new generation of users, continuously outputting AI capabilities like infrastructure.