AGIBOT has announced the launch of Genie Envisioner 2.0 (GE 2 Sim), which represents a major advance in the progression of world models, moving from World Action Models to fully interactive World Simulators.
The new system introduces what AGIBOT calls a “physical evolution engine” for embodied AI. This creates a model based environment where robots can be trained, assessed, and improved at scale, without depending entirely on expensive real world experimentation.
Visit the project page here.
From Understanding the World to Learning Inside It
In 2025, AGIBOT released the first open source action driven world model platform, Genie Envisioner. It enabled robots to interpret the world by combining vision, language, and action within a single modelling framework.
Genie Envisioner 2.0 pushes this further. Instead of only helping robots understand the world, it allows them to learn within a world generated by models.
This shift mirrors a wider movement in embodied AI. The focus is moving from representing the world to simulating it. As world models become stable, high fidelity environments that respond to actions in physically consistent ways, they make it possible to train robots at scale in synthetic settings.
AGIBOT sees this as a key turning point on the path toward a genuine scaling law for embodied intelligence.
From World Action Model to World Simulator
This evolution builds on AGIBOT’s ongoing development of the World Action Model (WAM) framework, which expands traditional world models by treating actions as a core variable.
Instead of modelling only state, WAM captures the full cycle of: State → Action → State Evolution
This allows world models to act as a foundation for both policy learning and action generation.
AGIBOT has built several systems on top of this framework: • EnerVerse, which extends embodied environments into a computable 4D world model • Genie Envisioner Act (GE Act), which links world representation with action trajectory generation • Act2Goal, which enables long horizon, goal directed control
Although these systems supported policy learning, real world deployment revealed limitations such as dependence on physical environments, high evaluation costs, and constraints on data scalability.
This led to a core insight: the next major breakthrough requires turning world models into fully functional simulators rather than simply improving representation.
Making the World Runnable: Moving Toward Interactive Simulation
To support this transition, AGIBOT introduces new capabilities that move world models closer to interactive simulation: • EnerVerse AC, which adds action conditioned world modelling for future prediction • Genie Envisioner Sim (GE Sim), a neural simulator for closed loop policy evaluation • EWMBench, a benchmark for simulation fidelity, action accuracy, and semantic alignment
AGIBOT also introduces a new data and training approach: • Real2Edit2Real, which makes real world data editable and extensible, increasing scale and diversity
• Fidelity Aware Data Composition, which blends real and generated data to balance realism with generalisation
Together, these developments shift world models from representation tools to environment level infrastructure.
Genie Envisioner 2.0: A Physical Evolution Engine
Genie Envisioner 2.0 brings this evolution to its full expression. It is no longer only generative, but operational.
Key capabilities include:
Action driven world dynamics The system responds directly to robot actions, producing high fidelity environmental changes that follow physical and semantic rules. The world becomes an interactive process rather than a static depiction.
Long horizon temporal modelling The system supports stable simulation over minutes, enabling continuous generation of full task sequences instead of short, disconnected clips.
Embodied spatial consistency Multi view perception, cross view 3D consistency, and robot proprioception are unified into a single representation. Perception shifts from images to an interactive embodied world.
Built in evaluation and reward modelling A native General Reward Model enables self evaluation and optimisation based on textual feedback, supporting reinforcement learning within the world model without manually designed rewards.
Toward real time interaction Improved inference efficiency brings GE 2 Sim close to real time performance, enabling: • Evaluation within the world model • Reinforcement learning within the world model • Teleoperation within the world model
This marks the shift from offline world models to interactive system environments.
A Paradigm Shift: When Models Become Worlds
As these capabilities come together, embodied AI is undergoing a fundamental change.
The field is moving from using models to understand the world to learning and making decisions inside model generated worlds.
The combination of WAM and Vision Language Action (VLA) models supports a shift from reactive control to generative, predictive decision making.
At the same time, World Simulators allow robots to explore, iterate, and optimise at scale. They are no longer limited by real world data, but by the fidelity of the simulation itself.
When these two developments converge, robots can move beyond copying human demonstrations and begin to explore, adapt, and evolve within model generated environments.
Toward a New Foundation for Embodied Intelligence
AGIBOT envisions world models progressing from tools for understanding, to platforms for learning, and eventually to infrastructure that supports continuous evolution.
When models become worlds, reality is no longer the only place where training can occur. When worlds can be constructed, learning can be scaled. When evolution happens inside models, the limits of embodied AI can be redefined.
Source: AGIBOT





