Dahyun Chung,
Siyoon Jin,
Hyunwook Choi,
Honggyu An,
Junyoung Seo,
Hyunsung Kim,
Seung Wook Kim†,
Seungryong Kim†
KAIST AI
† Co-corresponding authors.
- Project page: cvlab-kaist.github.io/ME-World
- arXiv preprint
- Inference code and model weights
- Training code
Egocentric world models predict first-person observations from an agent's actions, but most of them simulate a single agent. Real embodied settings involve several agents acting and interacting in one shared environment, where every interaction has to be visible from every agent's viewpoint and the resulting state changes have to appear in all observations at once.
ME-World formulates embodied multi-agent world modeling as synchronized ego-stream generation: it generates one first-person video per agent for agents interacting through fine-grained body and hand motion in a shared world. Three components keep the streams coupled:
- Joint multi-agent generation — all ego streams are denoised together in a single token sequence, so cross-stream information is exchanged at every layer.
- Shared action conditioning — every agent's body motion is projected into each agent's own camera: the wearer's own hands and arms, and the other agents' bodies, with a fixed palette per identity. Head motion enters as per-pixel rays in a shared canonical frame.
- Shared environment memory — all agents' observation history is pooled; the best-covering history frame is warped into each target view (stream-aligned geometric memory), and a greedily selected set of clean anchor frames is appended to the sequence for every stream to attend to (cross-stream anchor memory).
The model is trained on real two-person recordings and on a synthetic set rendered from retargeted human–human interactions, and is evaluated with shared-world consistency metrics for the environment (Senv), interaction-induced state updates (Supdate) and agent identity (Sid), alongside camera control, action control and video quality. ME-World improves all of them over multi-view video generation models, single-ego world models and general world models, and extends to three agents and to 221-frame autoregressive generation.
See the project page for videos: real and synthetic results, the architecture, explainers for shared action conditioning and shared environment memory, three-agent and long-horizon generation, comparisons and ablations.
@article{chung2026meworld,
title={{ME-World: Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction}},
author={Chung, Dahyun and Jin, Siyoon and Choi, Hyunwook and An, Honggyu and Seo, Junyoung and Kim, Hyunsung and Kim, Seung Wook and Kim, Seungryong},
journal={arXiv preprint},
year={2026}
}