"In the era of Agent AI, CPU processing latency accounts for approximately 50% to 90% of the end-to-end latency of the entire Agent workload." This figure, presented by Xiao Jun, Director of Ecosystem Cooperation at Theo End Computing (Nanjing) Technology Co., Ltd., at the recently held ICDIA (Integrated Circuit Design and Application Exhibition), highlights the new challenges that AI Agent workloads pose to CPUs.

Xiao Jun, Director of Ecosystem Cooperation at Theo End Computing (Nanjing) Technology Co., Ltd.
Over the past two years, the pursuit of GPU computing power for large model training and inference has almost overshadowed another critical thread: the redefinition of CPU performance by AI Agent workloads. As Agentic AI moves from concept to implementation, in the closed loop from perception and cognition to decision-making and action, the CPU takes on functions such as environmental perception, task orchestration, and tool invocation, making it no longer just a "supporting role" in the system. Xiao Jun judges: "The requirements for CPUs—core performance, memory bandwidth, and I/O bandwidth—are higher than in the original cloud era."
This technological inflection point coincides precisely with a subtle moment in the CPU industry landscape. x86 has dominated the server market for over three decades, but between 2024 and 2026, ARM server CPUs represented by Amazon's Graviton 5 and Microsoft's Cobalt 200 have accelerated their penetration among top cloud providers. ARM's share in the server sector has leaped from a negligible fraction to 20% to 25%. Meanwhile, RISC-V is probing into higher computing power scenarios from the IoT (Internet of Things) and embedded fields, with an expected compound annual growth rate of over 30%. The two camps are launching an assault from the fringes into the core territory of x86.
Theo End Computing is a new player established exactly at this intersection. Founded in July 2022, this server CPU startup has chosen a rare strategic path in China: pursuing dual technology routes for ARM and RISC-V in parallel. The first-generation ARM product has been taped out, and the second generation is under development; the first-generation RISC-V product is advancing simultaneously. In Xiao Jun's view, this is not a dispersion of resources, but a pragmatic combination of "using ARM for rapid commercialization" and "using RISC-V to layout for the future." The 24-core BlueDot V1 chip targets the massive incremental market of edge AI inference servers.
Dual Paths in Parallel: ARM for Commercialization, RISC-V for the Future
Theo End Computing's most striking strategic choice is to simultaneously deploy both ARM and RISC-V technology routes.
"We are indeed pursuing two routes in parallel," Xiao Jun frankly stated in the interview. "ARM currently has a better ecosystem and performance. We chose ARM because it allows us to achieve commercialization as quickly as possible. At the same time, RISC-V has huge future potential due to its openness and flexibility, so we need to lay out in advance."
Xiao Jun summarizes this strategy as "three projects in parallel": the first-generation ARM product has been taped out, the second-generation ARM product is under R&D, and the first-generation RISC-V product is also being developed simultaneously. In terms of manpower investment, the ARM direction currently receives more investment, but as the company scales up, investment in RISC-V is also gradually increasing.
From a product planning perspective, Theo End Computing's first-generation ARM product, BlueDot V1, is built on the ARM v9 instruction set's Neoverse V3 core, focusing on enhancing vector computing capabilities, out-of-order pipeline efficiency, intelligent task scheduling, memory interconnect bandwidth, and energy efficiency ratio. Meanwhile, this chip supports the PCIe 6.0 high-speed interconnect protocol and DDR5-8000MT/s memory speed. The core adopted by Theo End Computing is the same ARM V3 core used by international mainstream players—Amazon's Graviton 5 and Microsoft's Cobalt 200. According to the company's disclosure, BlueDot V1 is currently being taped out, with plans to initiate small-batch trial commercialization by the end of 2026.
The Agent AI Era: CPU Latency Challenges and Architectural Transformation
The rise of large AI models and Agents is redefining CPU design requirements.
Xiao Jun pointed out that Agent AI introduces an autonomous loop of "perception, cognition, decision-making, and action." In this loop, the CPU is primarily responsible for environmental perception, task orchestration, and tool invocation. "CPU latency accounts for approximately 50% to 90% of the total latency, which is a relatively high proportion," Xiao Jun said. "Therefore, in the Agent AI era, the requirements for CPUs—core performance, memory bandwidth, and I/O bandwidth—are higher than in the original cloud era."
This judgment directly influences Theo End Computing's product definition. The high specifications of BlueDot V1 in memory speed and I/O interconnect—DDR5-8000MT/s and PCIe 6.0—are precisely designed to address the data throughput bottleneck in AI inference scenarios. Xiao Jun further explained the reliance of inference scenarios on memory bandwidth and I/O bandwidth: "Due to the large number of parameters in AI models, when the context length of an Agent is long, the GPU's VRAM capacity can hardly accommodate both the model weights and the context's KV Cache simultaneously. Therefore, during the inference process, the KV Cache needs to be transmitted back and forth between the GPU's VRAM and the CPU's memory. At this point, memory bandwidth and I/O bandwidth become extremely critical."
From an industry trend perspective, AI-oriented heterogeneous CPUs are emerging as a new development direction. In his speech, Xiao Jun predicted: "In the future, we will see AI-oriented heterogeneous CPUs rising prominently, featuring dedicated Tensor processing units or standalone NPU (Neural Processing Unit) accelerators within the CPU, which are relatively suitable for exerting efforts in AI inference scenarios." In Theo End Computing's future heterogeneous architecture solution, the big cores will adopt the ARM V3 core, while the small cores will use RISC-V for acceleration, with the number of small cores reaching the thousands. This design concept not only responds to the demand for parallel computing in AI inference but also reflects the deep integration of the dual technology routes at the product level.
Team Barriers: The "Experience Threshold" for Large Chips
In the domestic CPU track, there are not many manufacturers capable of truly delivering products. Xiao Jun is frank about this: "The number of manufacturers in the domestic server CPU field can be counted on the fingers of two hands. Why? Because the threshold for this industry is very high."
In his view, Theo End Computing's core advantage lies in its team. "Our core team mainly comes from the core teams of former large companies. In server CPU chips, every stage of the R&D process requires highly experienced individuals to take charge independently. Being experienced doesn't mean you gain experience after three or five years; it may take at least eight to ten years, or even more than ten years, to be fully capable of taking charge independently."
"Among so many stages, each stage requires such experienced people. A team like this is relatively hard to find in China," Xiao Jun said.
This judgment reveals the essential characteristics of server CPU R&D. Unlike consumer-grade chips, server-grade CPUs have extremely high requirements for reliability, stability, security, and performance. A mistake in any stage could lead to losses of tens of millions or even hundreds of millions. There are no shortcuts to accumulating experience; it can only be achieved through the precipitation of time and projects.
Ecosystem First: Layout Starting from the Product Definition Stage
As the Director of Ecosystem Cooperation, Xiao Jun repeatedly emphasized the concept of "ecosystem first" in the interview.
"The chip industry, especially the CPU chip industry, is an ecosystem-driven industry. The characteristic of an ecosystem-driven industry is: ecosystem first," Xiao Jun explained. "You cannot build the ecosystem after the product is released; it will be too late by then. We started our ecosystem layout during the chip product definition phase."
Specifically, Theo End Computing is advancing ecosystem construction simultaneously across multiple levels. At the tool level, the company began building migration and tuning tools during the product R&D process to ensure that the tools can be released synchronously when the product is launched. At the community level, the company has joined the openEuler community, the Loongson community, and multiple domestic alliances and organizations, participating in standard formulation and white paper writing. At the software adaptation level, the company has sorted out all mainstream basic open-source software in the industry—operating systems, databases, and middleware—and prepared test cases in advance. Regarding commercial software, the company has communicated with several large manufacturers, and both parties have a strong willingness to cooperate.
At the product validation level, Theo End Computing has already run major operating systems in the EMU (Emulation) and FPGA environments prior to tape-out. According to public information, the company is capable of running operating systems such as CentOS, openEuler, and Ubuntu in the EMU and FPGA stages.
Xiao Jun also specifically emphasized the issue of standardization in ecosystem construction. He cited the historical failure of the MIPS instruction set as an example: "MIPS technology was quite good, but it failed to control the development of its ecosystem. After licensing the instruction set to customers, they could add their own instructions to it, leading to ecosystem fragmentation." RISC-V similarly faces the risk of fragmentation brought by open source. Theo End Computing's response strategy is to "follow the relevant standards of the foundation, promote relevant standards within the foundation, integrate into the ecosystem of the foundation and domestic RISC-V, and avoid creating ecosystem fragmentation".
Market Positioning: Asymmetric Competition, Focusing on the Commercial Market
In terms of product positioning, Theo End Computing has chosen a differentiated route.
BlueDot V1 is designed with 24 cores, focusing on four major scenarios: general-purpose servers, edge AI inference servers, storage servers, and workstations. Compared with high-end products from international manufacturers like NVIDIA and ARM—such as ARM's 136 cores and x86's over 200 cores—Theo End Computing has explicitly chosen asymmetric competition.
"They are taking the high-end route, used in training or cloud data center scenarios. We are using it in edge inference scenarios, so there is a difference in scenarios first," Xiao Jun said. "In the high-end market, one CPU can be paired with 4 or 8 high-end inference devices, whereas one of our CPUs can be paired with 1, 2, or 4 inference devices. From a market perspective, we are staggered from them to avoid direct competition."
In terms of market regions, Theo End Computing is currently focusing on the domestic market. In terms of market types, the company is targeting the commercial market rather than the Xinchuang (IT application innovation) market. "The Xinchuang market has certain qualification requirements. To enter this market, you must first obtain these qualifications, which takes time and has a certain threshold. Therefore, at this stage, we are still primarily focusing on the commercial market."
In terms of business models, Theo End Computing also offers more flexible choices than other domestic manufacturers. "We can just sell chips. If channel-type manufacturers need complete machines, we can also provide complete machines. We have a relatively complete upstream and downstream ecosystem, which can help customers build complete machine products. After customers put their logos on them, they can sell them." According to public information, the company has initiated preliminary customer introduction work, covering typical scenarios such as intelligent edge servers, inference all-in-one machines, and file servers.
Opportunities and Challenges for RISC-V: Ecosystem is a Prerequisite
Regarding the future of RISC-V CPUs, Xiao Jun has both expectations and prudence.
From the history of instruction set development, he summarized the lessons from three cases: MIPS technology was good but failed due to ecosystem fragmentation; ALPHA technology was equally excellent but declined due to a closed business model. The reason ARM and x86 became mainstream is that they maintained ecosystem unity while allowing third-party participation.
Xiao Jun believes that for RISC-V CPUs to develop, "the ecosystem is one of the prerequisites." Specifically, it is necessary to "support the open-source ecosystem, prevent ecosystem fragmentation in the long term through unified standards and software/hardware platforms," "steadily advancing in terms of performance, reliability, stability, security features, and virtualization," and "start landing in some medium-to-low computing power scenarios and gradually expand to higher computing power levels."
From the perspective of the maturity of key components of RISC-V CPUs, Xiao Jun's judgment is: the maturity of the core is moderate, and single-core performance still lags behind ARM and x86; bus maturity is relatively high; MMU maturity is average, with some missing system features; PCIe and DDR controller maturity is high, but system stability still needs polishing. At the ecosystem level, "the ecosystem maturity of RISC-V CPUs is still insufficient, and there is still a large gap compared to x86 and ARM. It mainly involves the hardware component ecosystem, basic software ecosystem, and application software ecosystem, among which the most difficult is the commercial application software ecosystem."
"The technology stack of RISC-V CPUs is not exactly the same as that of ARM, so it cannot be completely reused," Xiao Jun said. "However, there are some partners in both circles who are the same. For example, some operating system manufacturers have both ARM and RISC-V versions of their products. Therefore, the same partner cooperating on both technology routes can be considered a form of reuse in this regard."
Since 2026, the RISC-V industry has been accelerating its development. The Chinese Academy of Sciences released the "Xiangshan" open-source high-performance RISC-V processor system and the "Ruyi" RISC-V native operating system; Alibaba DAMO Academy released the new generation flagship CPU Xuantie C950; Lenovo and Lanxin Computing jointly launched the world's first server complete machine product equipped with a RISC-V processor.
Against this backdrop, Theo End Computing's choice of a dual-path strategy for ARM and RISC-V is both a response to industry trends and a pragmatic risk hedge—using ARM's mature ecosystem to ensure commercial landing, and using RISC-V's openness to layout for the future.