Written by | Wen Yehao Edited by | Wang Pan
While the World Robot Conference (WRC) still has residual warmth, the 2026 Bund Summit has once again heated up topics related to embodied intelligence.
Inside the venue, discussions on world models, data scaling, and commercial implementation never ceased. On September 11, during the 2026 Bund Summit, Zhu Xing, CEO of Robbyant, also held a dialogue with Photon Star and other media.
From data and world models to independent financing and commercialization, Zhu Xing repeatedly emphasized one issue: embodied intelligence is still in the cold start phase. The industry is eager to kickstart the data flywheel, but the capabilities of robots are not yet sufficient to support widespread applications.
At this moment, Robbyant is in the process of spinning off for independent financing.
Over the past year, the company has taken intensive actions: jointly releasing the first video generation foundation model for embodied AI, Robbyant-Video, with the industry, and training Robbyant-VA 2.0 based on it; Robbyant-VLA 2.0 adapted to robots from over a dozen mainstream brands during the pre-training phase. In addition, it has a self-developed "R-series" robot body, with deployment scenarios selected in pharmacies, warehousing, and industrial settings.
Just a day before the dialogue, on September 10, Robbyant-World 2.0 open-sourced its 1.3B Small version.
When asked whether large model vendors entering the embodied AI field would create a crushing dominance, Zhu Xing countered with an analogy: If OpenAI or Anthropic were to enter autonomous driving today, could they disrupt Tesla in a short time?
His judgment is that it would be very difficult. The advantage of large model vendors lies in training technology, but the more core barrier in embodied AI is data: what constitutes good data, the know-how of data, and the data pipeline.
At the same time, in Zhu Xing's view, collecting tens of thousands of hours of data does not directly indicate how much the model has learned. Zhu Xing observed that some models have three to five times the data volume of others, yet their overall performance remains poor, even though they excel in a few tasks. Behind this is often an imbalanced data distribution.
Robbyant's choice is to have the model team define the required scenarios, tasks, and quality standards, and then hand them over to suppliers for customized collection. Rather than promoting a larger collection number, they care more about how the data enters training and what the model still lacks.
Regarding implementation, he also maintains boundaries. In his view, over the next year or two, some productivity scenarios with controlled environments and short tasks are expected to gradually become economically viable. The team needs to restrain its pursuit of revenue figures, pass the test of paying customers, and accumulate technical and product capabilities.
The following is the main content of the dialogue (slightly edited for brevity without changing the original meaning):
Distribution Matters More Than "Quantity"
Q: Ego (egocentric) data currently lacks unified standards, and equipment specifications vary widely. Some studies suggest that scaling up collection in the future will cause a lot of waste. Can model vendors only designate a single high-spec data collection solution?
Zhu Xing: Ego is a general term and a good concept that can assist in the production and collection of human-centric real-world data. Building models pays great attention to data diversity, and behind diversity is nothing but the distribution of scenarios and tasks. Ego can significantly improve efficiency, being more portable and convenient for entering certain scenarios.
However, Ego is also subdivided into many methods, such as a headband with bare hands or a headband with gloves? The headband with bare hands currently has significant problems, with the core issue being accuracy: the hand trajectory is reconstructed, and this reconstruction has accuracy errors. When the error is too large, a large part of the data's inherent value is discounted. This should be the root cause behind the current issues of poor quality and low funnel rates.
In the future, combining high-precision, relatively portable tactile gloves should be the next important approach.
First, it retains the advantages of Ego; second, with the assistance of hardware, including algorithms, the accuracy problem can be solved; third, visual information and tactile information are naturally aligned across multiple modalities, as long as the error is not too large. Data that meets these characteristics is of great significance for the next step of model development.
Q: Some practitioners say that 90% of the data collected now is invalid. How does Robbyant view this?
Zhu Xing: I don't know how this number was calculated. Compared to the so-called single scale, we pay more attention to the quality distribution of the data itself.
Currently, regardless of the type of data source or collection method, there is a massive repetition of scenarios in the data. Data with excessively repeated scenarios is not very meaningful. Why do people always talk about corner cases? Because after data becomes abundant, the problem instead becomes how to get the specific part you want; quantity is simply not the issue.
What does Robbyant do? The demand for data itself is defined by model training, which is a very core know-how in model R&D. We need to define what type of data and what kind of distribution we need, and what scenarios and tasks this distribution covers. For some with higher collection difficulty, we conduct trial collections first, generate a good set of SOPs, and then push them to suppliers for scaling.
So going through the entire pipeline, Robbyant's core approach is customized collection. We stopped buying off-the-shelf data quite early last year. Off-the-shelf data indeed has problems: poor technical quality and massive repetition, which lead to a very low funnel rate. Quite early on, we increased the availability rate of the data funnel from raw data to data entering model training from about 15% to over 90%.
In addition, we pay special attention to one metric: from the start of data production, including the supplier's production, exactly how long it takes to enter model training—is it a week or three days. You will find that in our cooperation with all suppliers, data delivery is generally not done in batches, but in a streaming manner.
Q: Since the procured data is so unreliable, why not just build a self-owned data collection factory?
Zhu Xing: There is no need to view this as a black-and-white choice. Good data suppliers are very valuable to model vendors in at least two aspects.
First, they have relatively professional large-scale employment capabilities, managing a large number of production workers. How to assess and incentivize them, and maintain a certain elasticity in capacity, is a capability in itself. At least Robbyant is not good at this, and we don't want to make ourselves that complicated.
Second, in the future, data will need to enter more scenarios, and entering scenarios itself requires operations. Therefore, for Robbyant, it is definitely about achieving 1+1>2 with professional data service providers.
Q: Do you believe in Scaling Laws? How do you view the "quality" and "quantity" of data?
Zhu Xing: First, I also believe in Scaling Laws.
Second, the data behind Scaling Laws cannot be viewed so mechanically in terms of quantity. This phenomenon has already appeared in embodied models: some have tens of thousands of hours, which is 3 to 5 times the demand of another model, yet the effect is worse. However, it is found that they perform exceptionally well in certain types of scenario tasks. This is definitely caused by a severely imbalanced data distribution.
I can share a few of Robbyant's judgments on data.
First, quantity is important, but we cannot only talk about quantity; we must also talk about the underlying quality and distribution. Especially distribution, which is the most core know-how for model vendors. People always talk about "alchemy" (model training), and behind alchemy is the recipe.
Second, every different type of data has its value at different stages.
Is simulation data valuable? Yes, but more in mid-to-late training, especially for a specific task where I want to reinforce and pull up its success rate. But if you say simulation plays a huge value in pre-training today, how is that possible? If simulation can generate so much high-quality corpus in pre-training, it means your simulation capability has reached the level of controlling robots. Then you shouldn't bother with models anymore; it's too exhausting. Is Ego data valuable? It is also valuable.
Third, do not easily talk about the ratio between different data sources. People often ask me, do you use internet data, and how do you mix it? I say, what is there to mix? Millions of internet video manipulation data mixed with a tiny bit of real-robot data would just drown it out.
Q: Everyone is talking about the data flywheel. When do you think it can really start spinning?
Zhu Xing: Today is very early, a cold start, relying on forced feeding through various data collections. Taking a step forward, with a certain scale of application, at least in some scenarios, the flywheel for anomalous data can also start spinning. Actually, successful data is not of much significance.
What do I think is the real data flywheel for embodied intelligence? It is that everyone, in their daily life or work process, is producing data for robots. But if you give it to them now, they can't consume it; they can't digest "coarse grain."
From "Usable" to "Economically Viable"
Q: Why are the only implementations now in small convenience stores? Simply achieving a 100% success rate in P&P (Pick and Place) could theoretically unlock many scenarios. Why can't it be done?
Zhu Xing: Small convenience stores are actually hard to implement. Doing only a very specific scenario at a stage, without pursuing scaled replication, and directly throwing in a small model is not impossible. But it depends on whether the scaled scenario involves replication across scenarios, tasks, and locations; here lies the challenge of generalization.
Whether P&P can achieve 100% depends on the complexity of the operation sequence of the scenario, environment, and the task itself.
Relatively objectively speaking, to promote a certain scale of implementation at this stage, several dimensions need to be considered.
First, the environment cannot be too open, and generalization should not be too broad. Second, tasks cannot be too long-horizon, because long-horizon tasks have too many anomalies. Today's mainstream implementable products are basically still based on imitation learning, which fears anomalies the most because they are unseen. There is also the cost constraint: if the value you create is less than the original value, and it is relatively expensive, it won't work in the long run.
First is the capability issue; in some scenarios, the capability simply cannot reach the requirement. But we have seen that over the past year, the improvement in data and model capabilities has been quite fast.
Q: What is the relationship between the capability of the foundation model and the implementation success rate? Will there be a conflict?
Zhu Xing: No conflict. The great value of the foundation model is simply to make the few-shot performance in post-training better, achieving the same effect with less data, or achieving a better effect with the same data. This effect includes the success rate.
The improvement of the foundation model is like water rising and boats floating higher. This is the same principle as how the development of foundation model capabilities in large language models later drove the explosion of the entire ecological applications.
Q: What stage is this industry in now? When can we see quantitative growth?
Zhu Xing: There have been several positive changes in the industry over the past year.
First, people are increasingly concerned about how to better scale data; second, the entire industry cares more about implementation, and many industry players are starting to think more about how to combine embodied intelligence with their own businesses; third, the development of model routes is very exciting, including world models, VLA, etc., and we have seen some improvements in effects.
But no matter how good it is, it is still very early. The comprehensive complexity of embodied AI determines this. Autonomous driving has L1, L2, L3, L4, moving and stopping for so many years, passing through so many cycles. Embodied AI is the same; it is in its early stages, and everyone needs to be a bit patient. The good thing is that I see limited, real implementations starting to happen.
Implementation involves two things: the necessary condition is whether the capability can be reached—the scenario cannot be too open, the environment cannot be too complex, and the task cannot be too long-horizon. After all these are met, we also need to look at the cost. Cost is the sufficient condition; can the math work out? Facing households is another matter; I am talking about the productivity scenario.
In the visible next year or two, in places where task complexity is not that high, the math can gradually work out. I think from the second half of this year to the next year or two, it should enter a stage of rapid quantitative growth in productivity applications.
Robots Don't Need to Look Good, They Need to Be Reasonable
Q: Will world model companies in the digital world, those doing games, film and television, and entertainment, grow into embodied AI along the technical route?
Zhu Xing: Some concepts and terms are very broad now. What is the underlying interoperability between embodied world models and digital world models, including those for games? This type of world model based on the video generation route only uses the principle of video generation; it just grew out of the video generation school. The difference between the digital world and the physical world is quite large.
Let me take the most popular world models this year as examples: Kling, Wanxiang, and Sora. They are all excellent digital world models, but they corely satisfy the needs of the digital world: content needs to be rich, good-looking, innovative, with special effects and good image quality. Moreover, objectively speaking, as long as it is good, users can accept latency.
What do robots need? Robots don't need to look this good, nor do they need special effects. What robots need is "reasonableness"—whether it conforms to physical laws, because ultimately it needs to be translated into actions to control the robot.
For example, if I want to drink this bottle of mineral water, and the bottle dances in the sky and falls into my hand, it would definitely look good, but it doesn't conform to physical laws; it's meaningless. Second, robots need high performance and real-time capabilities because they are constantly inputting and outputting in the physical world. It's impossible to wait for you to generate a very good story for them.
On this point, Robbyant has a very firm persistence, including releasing the first video generation foundation model for embodied AI, Robbyant-Video, with the industry some time ago, and based on it, training the entire Robbyant-VA 2.0.
Q: Why not just fine-tune directly on general video generation models?
Zhu Xing: Essentially, during our 1.0 series, we also based it on general video generation models in the digital world, the most representative being Wan, to fine-tune and force it to adapt to robots.
But this route has two problems: first, the upper limit is not high; second, tuning it to the end will instead destroy the priors and knowledge of the previous model, and the side effect is reduced generalization.
When we reached the later stages, we really hit a brick wall, so we started more native R&D very early on. We added a large amount of real robot data when training our Video model.
People easily ask whether this world model is good or a certain model is good. You need to see what needs it satisfies. It does very well in meeting user needs, but forcing it to control a robot is asking too much, very much like forcing a man to give birth.
Q: On September 10, you open-sourced the 1.3B version of Robbyant-World 2.0, and previously Yu Jun said that edge deployment would limit computing power. Why still open a small model?
Zhu Xing: It's impossible to say that hallucinations are completely gone. With the model being this much smaller, the effect remaining completely consistent is also unrealistic.
First, we previously said we would open smaller versions in the future, hoping that those interested in world models, especially highly interactive world models, can play with them more. This is what we wanted to contribute to the technical community in the early stage.
Second, by opening this small model, we are open-sourcing some technologies. For example, everyone is quite concerned about how to train a unidirectional causal ability, rather than a bidirectional one, because a robot's time must be unidirectional. Online autoregressive prediction uses occurred observations in chronological order and cannot rely on future real frames. This training technology itself has a threshold, and we wanted to contribute a part of it to the community relatively early.
The "Battle of a Hundred Models" in Embodied AI Has Just Begun
Q: If companies like OpenAI enter the embodied AI field, how big is the threat? Is the window left for embodied AI only one or two years?
Zhu Xing: I can make an analogy. If OpenAI or Anthropic were to enter autonomous driving today, could they disrupt Tesla in a short time?
In what aspects do large model vendors have advantages? For example, training technology; the training capabilities of Anthropic and OpenAI are higher than Tesla's. But for the more core barrier of autonomous driving, to put it bluntly, it is still data. Embodied AI is the same.
Yesterday, the main opening forum of the Bund Summit was also discussing this issue, and my view is highly consistent with everyone's: this definitely has very good technical reference significance, including improving the efficiency of certain links in the embodied AI R&D process.
But objectively speaking, the difficult problems faced by embodied AI are still there. The challenge stuck on data remains, and it still needs to be solved by themselves.
Q: Does the market need so many embodied intelligence companies? Which companies with what traits can survive?
Zhu Xing: Whether Robbyant can survive in the long term is also a question, which I cannot answer. But is it a bad thing for an industry to receive excessive attention at a certain stage during its development? Actually, it's not a bad thing.
Returning to Robbyant, we still go steady and far, doing it for the long term. We will also choose some routes with larger investments and higher risks relatively early, because the upper limit of this route is higher. We hope Robbyant can cross several cycles.
From the perspective of the industrial landscape, there really don't need to be too many at the foundation model level. Look at large models, from the vigorous "Battle of a Hundred Models" to now being relatively converged, with four or five sitting at the main table. It's quite hard for domestic players to get a seat now. Embodied AI will also move towards a few main players. The future will definitely be a converged landscape, which is actually more efficient for the industry. But today, embodied AI might just be at the very beginning of the "Battle of a Hundred Models."
Q: Where exactly is Robbyant's advantage in self-research?
Zhu Xing: Regarding advantages, our proactive publicity is relatively low-key. But when overseas looks at domestic embodied intelligence, especially model capabilities, they might not know the four characters "Ant Robbyant," but they definitely know "Robbyant."
When more advanced overseas peers release some work, to some extent, they will include many of Robbyant's elements in baseline evaluations; very large global technical communities will also include Robbyant-related models in official recommended models or baselines.
Here, I can share two insights gained from the development stage of large language models.
The first is the issue of choice: after the rise of large language models, there were two choices in front of us. One is to fine-tune based on other foundation models, and the other is to resolutely invest in pre-training at greater risk. Later, those at the "main table"—none of those who chose the former made it out; those who chose the latter didn't necessarily make it out either, but all those who didn't choose the latter died. Today, embodied AI has also entered the early stage of the "Battle of a Hundred Models."
Second, why are only a few families always at this table? Behind it relies on the Infra capabilities delivered by excellent teams, large-scale training technology, and large-scale data pipelines. Robbyant's training scale is super large, and in many directions, it is a native foundation model in the first-mover field.
Q: What exactly is the route you persist in? How do you view companies like Physical Intelligence?
Zhu Xing: What we believe in most is that physical intelligence or embodied intelligence needs to be based on physical real-world data, training its own foundation model from 0 to 1. We do not believe that physical intelligence is merely a graft of digital intelligence.
The development of digital intelligence, including the development of multimodal and large language models, may provide a lot of assistance to embodied AI, but the core part still needs to be solved by embodied AI itself.
Physical Intelligence has made great contributions to the development of global embodied intelligence and has well promoted the development of the entire community in the past few years. But if you ask what exactly its model architecture has that others don't? A very large part of its core behind it is the know-how of data.
The model paradigm we talk about is not at the Transformer level; that is a large computing architecture, and everyone will switch if it changes. The model paradigm has not fully converged yet, including how to well utilize the capabilities of world models on top of it. The future learning paradigm for robots will also undergo major changes. These require continuous exploration. Data is very important now, and it will be even more important in the future.
Q: Some companies making the "brain" say they only self-develop the brain and the body is not important. Why do you still make the R-series body?
Zhu Xing: This doesn't need to be extreme. From the perspective of R&D, it is definitely integrated hardware and software, because models need data, and data is produced by hardware. If you don't do it yourself, you have to find others to customize it.
Second, Robbyant indeed has an "R-series" robot. One is to meet the needs of R&D; the other more important point is that the core business precipitation of Ant Group is all around life services. Our embodied intelligence also wants to explore life services earlier, so we need a suitable body.
One characteristic of our body is that the chassis is very small, equipped with a lot of perception devices. We have done a lot of hardware and software processing in many aspects such as force control. In life service scenarios, passing and obstacle avoidance are more complex, and what we consider is safety. This is not the current mainstream shipping scenario in the industry, and since we want to explore it earlier, we need a hardware platform to cooperate.
As for worrying about forming competition with partners, we do not touch shipping scenarios like industry and logistics; we only provide foundation models plus toolchains. The life services I talk about are strictly between To B and To C. I generally define it as "2B2C." It is an outpost before entering households, where some things can be explored in advance in terms of technology and business.
Q: If body manufacturers also start making models, what will your cooperative relationship become?
Zhu Xing: First, the entire Robbyant, including myself, believes more in open ecological cooperation. Especially in the relatively early stage of the industry, for the industry to go further, everyone needs to build an ecosystem together.
Second, returning to Robbyant itself, how could we have the virtue and ability to make the brain very leading and make everything very leading? I don't think God would favor Robbyant so much.
The core of what Robbyant needs to do is to output intelligence, outputting closed-source foundation models plus toolchains. The toolchains will help these manufacturers solve more efficient post-training, more efficient post-training data collection, including cleaning and annotation, higher-performance quantized deployment on the edge, and more efficient real-robot reinforcement learning.
We will give these jointly closed-source foundation models to manufacturers together, providing technical services in the process, exploring post-training together. To put it bluntly, it is about continuous integration, cultivating this ecosystem, and also hoping to build a data network to allow more anomalous data to flow back.
For me, this kind of cooperation is 1+1>2. Overall, robots are a non-standard product. From docking with customer needs to integrated delivery, including maintenance, these are very heavy tasks.
Q: How will the value of embodied AI be measured in the future? By task or by delivery?
Zhu Xing: Personally, I think the Token measurement method will also be suitable for embodied AI in the future, just with subtle differences in its approach, and the implementation method might be slightly different. Measuring by completed tasks or delivery is unlikely and not easy to measure.
Q: In the next year, what does Robbyant most want to accomplish? Where is the critical point of life and death for the industry?
Zhu Xing: From Robbyant's perspective, we care about two things.
First, we will not excessively push commercialization just for a better revenue figure. First, it is unsustainable; second, it will drain the team's energy at stages, truly picking up sesame seeds while losing the watermelon. Robbyant has chosen a very long-term route; otherwise, we wouldn't invest in large-scale training. In the process, we will maintain restraint on small things.
Second, no matter how restrained we are, we cannot do nothing. Only paid commercialization can meet more real tests. Commercialization is a very important core capability for a company to move forward. No matter how good the technology is, it is precipitated by continuously meeting the challenges of commercialization. Including cloud technology and large model technology, in the step-by-step commercialization process, the maturity of the entire technology, product, and capability will be better. We do limited commercialization to make the Robbyant team more capable in this regard.
Q: Why does Ant Group attach so much importance to embodied intelligence?
Zhu Xing: Why has Ant Group attached so much importance to embodied intelligence in the past, present, and future? A very important point is that in the next era of physical intelligence, in ten or twenty years, how to better serve people, and how to better serve people in households? The bridge that cannot be bypassed in the middle is embodied intelligence.
What Robbyant is doing is because of Ant's strategic positioning. When embodied intelligence continues to develop, Ant's long-term user accumulation, scenario accumulation, and data accumulation in life services can completely "join forces" with Robbyant more in the future.
WeChat ID | TMTweb
Official Account | Photon Star