Humanoid Robotics Technology
hrt awards banner
Home » AGIBOT Announces Genie Operator-2, a Next-Gen Embodied Foundation Model

AGIBOT Announces Genie Operator-2, a Next-Gen Embodied Foundation Model

genie operator

AGIBOT has introduced GO‑2, its next‑generation foundation model for embodied AI. GO‑2 is the first system to bridge the “last mile” between logical reasoning and precise execution within a unified architecture. Trained on tens of thousands of hours of interaction data, it has set new State‑of‑the‑Art records across multiple robotic benchmarks. It marks a shift from exploratory black‑box behaviour to a true unity of reasoning and action.

Building on its predecessor GO‑1, GO‑2 integrates logical reasoning and action execution within a single system. This allows robots not only to plan effectively but also to carry out those plans reliably in real‑world environments.

Core technical contributions from GO‑2 have been accepted to CVPR 2026 and ACL 2026, highlighting its impact across both computer vision and natural language processing.

Evolution of the GO Series: From Perception to Actuation

A year ago, AGIBOT released the Genie Operator‑1 (GO‑1) foundation model. Powered by the ViLLA architecture, it delivered unified modelling of vision, language and action. GO‑1 became a milestone in embodied AI, earning a Best Paper Nomination at IROS, acceptance in the robotics journal TRO, and the SAIL Star at WAIC. It now forms part of Genie Studio, AGIBOT’s end‑to‑end embodied development platform for deploying and validating models at scale.

GO‑1 enabled robots to understand instructions, interpret scenes and plan tasks. However, as deployments expanded into more complex environments, a key limitation emerged: robots could generate reasonable plans, but their actions did not always follow those plans precisely.

This was not a planning failure but a disconnect between reasoning and execution. The underlying issue is the long‑standing Semantic‑Actuation Gap, where high‑level reasoning signals and low‑level motor commands remain insufficiently aligned. As a result, control modules often bypass reasoning during execution, leading to accumulated errors in long‑horizon tasks.

GO‑2 is designed specifically to close this gap and to ensure that robots can reason about the world and act upon it with consistent stability.

Core Philosophy of GO‑2: Achieving a True Unity of Reasoning and Action

To unify reasoning and action, a system must solve two challenges at the same time:

  1. Generating action plans that are genuinely executable through deep spatial reasoning.
  2. Ensuring those plans can be executed stably in real‑world conditions.

GO‑2 addresses both through two major innovations.

1. Action Chain‑of‑Thought: Reasoning in the Action Space

GO‑2 performs reasoning directly within the Action Space using an Action Chain‑of‑Thought approach.
Instead of mapping instructions straight to motor commands, it produces a high‑level sequence of action intents that form a macro‑plan. This mirrors how a human mentally simulates a movement before performing it. GO‑2 explicitly plans a full behavioural path and executes it step by step, naturally breaking complex tasks into ordered stages.

This capability has been accepted by CVPR 2026 as a significant advancement in embodied AI.

2. Asynchronous Dual‑System: Low‑Frequency Planning and High‑Frequency Execution

High‑level reasoning alone cannot guarantee stable execution in noisy, unpredictable environments. GO‑2 introduces an Asynchronous Dual‑System architecture that translates reasoning into precise movement.

Semantic Planning Module (System 2):
Operates at a lower frequency and acts as a general commander. It generates structured high‑level action sequences using progressive refinement, ensuring that plans are inherently executable and provide stable geometric anchors.

Action Following Module (System 1):
Operates at a higher frequency and acts as an agile executor. It receives high‑level intents and combines them with real‑time observations to generate control signals, applying residual refinement to compensate for environmental disturbances.

Both systems are tightly aligned. GO‑2 uses a Teacher Forcing mechanism during training to ensure that execution remains robust even when reasoning is only approximately correct.

This architecture has been accepted by ACL 2026.

Performance: State‑of‑the‑Art Across Benchmarks

By unifying reasoning and action, GO‑2 delivers a major leap in behavioural performance and significantly outperforms models such as π0.5 and NVIDIA GR00T.

  • LIBERO Benchmark: GO‑2 ranks first across Spatial, Object, Goal and Long tasks with an average success rate of 98.5%.
  • LIBERO‑Plus: Achieves an 86.6% zero‑shot success rate in environments with disturbances.
  • VLABench: Scores an average of 47.4, outperforming existing methods in texture and category generalisation.
  • Genie Sim 3.0 (Sim‑to‑Real): Achieves an 82.9% real‑world success rate using only simulation‑trained data.

From Model to Deployment: Continuous Learning in the Real World

AGIBOT is extending GO‑2 into real‑world deployment through a pre‑training, post‑training and data‑feedback loop.

Integrated with Genie Studio, the system supports:

  • Continuous data collection across robot fleets
  • Cloud‑based collaborative training
  • Online post‑training in real environments

This infrastructure enables:

  • Distributed training across thousands of robots
  • Approximately ten‑fold improvements in training efficiency
  • Task start‑up times reduced to minutes
  • Minute‑level convergence in industrial tasks
  • Success rates improved by two to four times with more than 50% less data

GO‑2 becomes not just a model but a continuously evolving embodied system.

Toward Embodied Agents with Memory

AGIBOT is now exploring whether robots can develop long‑term memory and improve through experience. The OpenClaw Memory System (arXiv:2603.11558) introduces long‑term memory that allows robots to reuse reasoning traces from past interactions.

By combining action‑level reasoning, hierarchical execution and long‑term memory, AGIBOT is building a complete intelligent loop: perception, reasoning, action and memory.

From GO‑1 to GO‑2, AGIBOT has moved from enabling robots to understand the world to enabling them to act on it. GO‑2 represents the point at which embodied foundation models achieve a genuine unity of reasoning and action.

As the field progresses, AGIBOT aims to accelerate the transition from research breakthroughs to real‑world impact and to unlock the next phase of scalable, intelligent robotics.

Source: AGIBOT

HRT Badge

Humanoid Robotics Technology

  • Caohejing Kangqiao Business Oasis, No. 2555 Xiupu Road, Pudong New Area, Shanghai, Shanghai, CN
Visit Website View Profile
Join thousands of Humanoid Robotics Experts and get the latest updates straight to your inbox!