Ant Group’s Robbyant Unveils LingBot-VA 2.0 for Real-World Robot Control
Event summary
- Robbyant, an Ant Group subsidiary, launched LingBot-VA 2.0 on July 10, 2026, the first embodied-native video-action world model designed specifically for physical environments.
- The model achieves real-time inference speeds of 150 Hz on a single GPU and can generalize to new tasks with as few as 20 demonstrations.
- LingBot-VA 2.0 integrates semantic visual-action tokenization, strict causal pre-training, mixture-of-experts architecture, and enhanced asynchronous inference for efficient robot control.
- The release is part of Robbyant’s full-stack launch week, introducing six models for embodied AI applications.
The big picture
Robbyant’s LingBot-VA 2.0 represents a strategic shift in robotics foundation models, moving from digital content generation to physical world control. This aligns with the broader industry trend of integrating AI with real-world applications, particularly in sectors like elderly care and medical assistance. The model’s architecture innovations address key challenges in execution efficiency and causal prediction, positioning Robbyant as a competitor in the embodied AI space.
What we're watching
- Execution Efficiency
- How the model’s real-time inference speed will impact deployment in industrial and real-world scenarios.
- Generalization Capability
- Whether LingBot-VA 2.0 can sustain its performance across diverse tasks with minimal demonstrations.
- Ecosystem Development
- The pace at which Robbyant expands its open technology and application ecosystem for broader robot adoption.
Related topics
