In the first week of September 2026, the AI industry experienced an unprecedented "dense release week". On September 1, Anthropic took the lead in launching Claude Fable 5.1 and Claude Mythos 5.1, focusing on coding, knowledge work, and long-duration Agent tasks. The next day, Google released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, which specializes in cybersecurity—just three weeks after the release of the previous generation, Gemini 3.7 Flash, marking the third iteration of the Flash series within six weeks. Shortly after, Meta unveiled Muse Spark 1.3. In the early hours of September 4, OpenAI officially released the new-generation flagship model GPT-6 Astra. At the launch event, company president Greg Brockman declared, "Welcome to the AGI era." Astra delivered near-perfect scores across multiple benchmark tests—according to official OpenAI data, it scored 97.6% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and a perfect 100% on the ExploitBench cybersecurity test.
Parameters are expanding, context windows are lengthening, and reasoning capabilities are climbing. Almost every metric is telling people that the technical capabilities of large models are growing at an unprecedented speed.
Shifting the focus from overseas back to China, there is no shortage of major news. In July, Moonshot AI released Kimi K3, with a total parameter scale reaching 2.8 trillion. Adopting a sparse mixture-of-experts architecture, it natively supports visual understanding and a context window of up to 1 million tokens—making it the largest open-source model in the world by parameters at the time. Kimi K3 ranked third in the intelligence ranking by independent evaluation agency Artificial Analysis, and its release quickly attracted widespread attention from the global AI industry and research institutions.
However, when the focus shifts from laboratories and launch events to the real business world and daily life, another picture quietly unfolds.
01. The Temperature Gap Between Hype and Reality
Let's look at a set of numbers first. McKinsey's 2025 global survey shows that about 88% of enterprises have applied AI in at least one business function, but only about 39% have truly brought measurable profit contributions; a report released by MIT NANDA during the same period pointed out that about 95% of enterprise-level generative AI pilots failed to produce measurable operational returns in the short term. Gartner's survey is even more direct—over 90% of enterprises globally have launched generative AI pilots, but only about 41% of the projects have truly entered the production environment and formed scaled value. Focusing more specifically on Agentic AI, currently only 16% of enterprises have deployed it in production environments, an increase of only 7 percentage points from 9% in 2025.
These numbers point to the same conclusion: enterprise AI attempts are very common, but very few can actually run through and make money.
IDC predicts that by 2026, 50% of AI-driven digital application scenarios will fail to meet ROI targets. Another Gartner survey shows that only 11% of CFOs saw actual financial value brought by AI in 2025, and only 8% of Chinese enterprises have truly achieved revenue growth through AI.
Between the heat wave of technology and the coolness of business, there lies a huge temperature gap.
02. The Devoured Moat
This temperature gap does not come out of thin air. A thought-provoking phenomenon is: the stronger the capabilities of large models become, the more barren the application ecosystem built around them appears.
In the early stage when ChatGPT was just introduced, the market saw a large number of AI applications emerge with the business logic of "filling model defects"—legal tech companies reduced model hallucinations through professional databases, medical consultation applications improved reliability through secondary verification, and marketing copywriting tools generated content by relying on carefully designed prompt chains. Without exception, these applications seized the early capability shortcomings of large models, using engineering methods to build temporary supporting structures at the model's soft spots.
However, as model iterations accelerate, these moats are collapsing rapidly. GPT-4 leveled the precision gap in legal citation, and Claude's continuous optimization in reasoning capabilities has also significantly improved the reliability in professional scenarios such as medical consultations. After the context window expanded from several thousand tokens to the million level, those copywriting tools relying on external memory mechanisms found that their carefully maintained barriers became useless. The dilemma that application-layer value is difficult to scale is also reflected in macro data—a report released by an Oxford University research team shows that the total annual global AI investment has exceeded USD 400 billion, but only about 33% of enterprises have successfully pushed AI projects from pilots to scaled applications.
Even more worrying are the changes at the user level. The number of independent applications in app stores boasting "AI-driven" is still growing, but the head products with monthly active users exceeding one million remain a minority, and a large number of applications fall into user churn and revenue stagnation shortly after going online. The motivation for many users to switch their primary AI tools is mostly "heard that the new model works better", rather than "the new application solves things the old application couldn't do".
Users have no loyalty because all AI products are looking more and more alike—a uniform dialogue box plus an input box. When interaction methods converge and underlying capabilities come from the same batch of large models, where exactly is the independent value of the application layer?
03. Structural Dilemma
If the dilemma at the application layer is just the surface, then the deeper problem lies in the structural obstacles to enterprise AI implementation.
High costs bear the brunt. The computing resources required for large model training, and the power and energy expenditures behind data centers, will ultimately be borne by enterprises. Although companies like DeepSeek have been working hard to reduce training costs—in 2025, DeepSeek compressed the total cost of post-training such as reinforcement learning for the R1 series to about USD 294,000—inference costs remain a threshold for scaled applications. Taking Google's newly released Gemini 3.8 Flash as an example, although its current promotional price is only USD 0.75 per million input tokens and USD 3.75 for output (which will be raised to USD 1.5 and USD 7.5 starting in 2027), far lower than Claude Opus 5's USD 5 and USD 25, for small and medium-sized enterprises that need to call APIs on a large scale, this expenditure is still not to be underestimated. The shortage of high-end computing power supply and high application costs remain significant barriers.
Secondly, there is the dilemma of data quality. Gartner's survey shows that only 4% of enterprises rate themselves as having AI-ready data; the data of most enterprises is merely traditional business records, lacking structured semantics and adaptation to business scenarios. Data being "visible but unusable" has become a major bottleneck for enterprise AI implementation.
The deeper problem lies in the cognitive limitations of the models themselves. The Blue Book on the Technical System of Causal World Models released by Zhongshu Ruizhi points out that most current AI models stay at the association layer of the "ladder of causation", being good at answering "what is", but struggling to deduce "what will happen after intervention". In core industry scenarios with zero tolerance for errors such as oil and gas, electric power, and high-end manufacturing, AI outputs lacking causal logic and deductive capabilities cannot meet rigorous decision-making requirements. This is the core reason why a large number of AI projects are "usable but dare not be used, implemented but not deeply cultivated".
Sun Xin, Senior Principal Research Director at Gartner, has a concise summary of this: top AI models complete a round of capability iteration every 1.5 to 3 months, but the ability of enterprises to absorb AI and transform business value has not grown synchronously. There is a severe "impact gap" between technology supply and enterprise implementation needs.
04. Bubbles and Consolidation
In the first half of 2026, a batch of AI applications once hotly pursued by capital are exiting the stage one after another. OpenAI announced the discontinuation of the Sora video generator, which had been online for only half a year; Yupp.ai, an AI model evaluation platform that raised USD 33 million, announced its shutdown; Google began to shrink its internal AI application lines.
These exits are not entirely without value. They collectively expose the same problem: as the underlying models continue to upgrade, has the application layer formed a sufficiently thick independent value? Applications that are only supported by model dividends are losing the reason to continue existing independently.
But the bursting of the bubble does not mean the end of the story. In fact, this is more like a necessary clearance—clearing out those applications that "white-label" large model capabilities and rely solely on interface packaging, allowing truly valuable applications to surface.
In an industry observation at the end of 2025, the China Academy of Information and Communications Technology (CAICT) mentioned that continuous technological iteration has laid a solid foundation for the practical application of large models, and agents are becoming the main form of large model application implementation. GPT-6 Astra has already demonstrated an intuitive possibility—according to data published by OpenAI, in the offline evaluation of OSWorld 2.0, Astra's computer use performance reached 72.6%, taking about 40 minutes per task, achieving a significant improvement over GPT-5.6 Sol's 65.7% and 75 minutes. When models can independently complete a series of tasks such as filling out online forms, updating CRM records, organizing schedules, analyzing scientific data, and creating websites, the form of the application layer is bound to be redefined.
From "intelligent assistants" to "digital employees", from single-point tools to autonomously collaborating agent groups—this road is still very long, but the direction has begun to become clear.
Models are still getting stronger, this is a fact. In just the first four days of September 2026, four overseas giants have taken turns to step onto the stage. But the speed at which the scope of problems that AI can truly deliver in a closed loop expands is obviously slower than the growth of model capabilities. Narrowing this gap and truly transforming the sprint of technology into a leap in productivity is perhaps the most anticipated proposition in the second half of the AI game.
- End -