← Back
XGEN-Labs

XGEN-Labs/XGEN-JING

View on GitHub ↗
Stars
266
Forks
24
Watchers
266
Open issues
0
Contributors
1
Language
Python
License
Other
Default branch
main
Created Sep 15, 2026Updated Sep 20, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

xgen-jing-demo.mp4

XGEN logo XGEN-JING: An Egocentric
Interactive Experience Model

XGEN Team

GitHub Hugging Face Gallery Blog arXiv: Coming Soon


We present XGEN-JING, an egocentric interactive experience model built on MiniMax-H3. Given actions, reference images, and observation history, JING generates first-person video and audio for navigation, object interaction, and conversation.

  • Camera control. Explore everyday places and imagined worlds through keyboard-controlled movement.
  • Interaction and dialogue. Guide object interactions and character conversations with text, with video and audio generated together.
  • Reference conditioning. Combine character, object, and scene images to compose an experience and explore different actions from the same starting point.

This release provides four-step bidirectional inference, example cases, and Prompt skills. The causal model and technical report are coming soon.

📋 Release Plan

  • JING-Flash-v1 — Four-step bidirectional model.
  • Inference code and examples — Camera controls, reference images, and joint audio/video generation.
  • Prompt skills — Generate validated inference cases from stories and reference images.
  • Causal model — Coming soon.
  • Technical report — Coming soon.

🚀 Quick Start

1. Installation

Use Python 3.12 and a compatible CUDA environment. The demo has been validated on six H100 GPUs: one for the text encoder, one for the video/audio VAEs, and four for the DiT with sequence parallelism. FlashAttention-4 is the default backend.

git clone https://github.com/XGEN-Labs/XGEN-JING.git
cd XGEN-JING
python3 -m pip install -r requirements.txt
SGLang runtime and validated CUDA versions

Install SGLang at commit 95140a7b0c9fc2f87a2a6cf6f6f0df8640a73174 separately. Keep the Diffusers revision pinned in requirements.txt; SGLang's diffusion extra pins a different version. The validated stack uses Torch 2.13.0+cu130, torchvision 0.28.0+cu130, Triton 3.7.1, FA4 4.0.0b26, and SGLang kernel 0.4.7+cu130 from the CUDA 13 wheel index.

2. Model weights

JING Flash is built upon MiniMax-H3 Ref2VA and FlashGen for a faster experience.

The demo loads the JING-Flash-v1 transformer from Hugging Face. The text encoder, tokenizer, processor, video/audio VAEs, and schedulers come from Diffusers-format MiniMax-H3.

Required model files are downloaded automatically on first use and reused from the Hugging Face cache. Set HF_HOME to choose the cache location. No manual download or weight directory is required. Repository IDs are configured in configs/base.yaml.

3. Inference

python3 demo_bidirection.py check_config=true
bash demo.sh cases=examples/bakery_greeting.json

Or:

bash demo.sh cases=examples/train_carriage_gaze.json

Results are saved to examples/outputs/.

Custom model paths and GPU selection
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5 bash demo.sh \
  model.h3=/path/to/MiniMax-H3 \
  model.transformer=/path/to/XGEN-JING \
  output.directory=/path/to/outputs
Write your own prompts and controls

Start from the bakery example. Each case combines reference images with a sequence of prompts and controls. A prompt chunk uses one control entry per repeated slice:

{
  "prompt": "First-person view: approach the counter and greet the baker.",
  "repeat": 3,
  "control": ["w", "w,a", ""]
}

w/s/a/d control forward/backward/left/right movement; i/k look up/down and j/l turn left/right. "w,a" or "w,j" combines keys in one slice. Opposite keys cancel; "" applies no keys. Keep control the same length as repeat. With the default layout, num_frames = 17 * sum(repeat) + 5. Reference images are ordered and addressed as <Picture 1>, <Picture 2>, and so on, up to five.

Prompt skills

Use Prompt skills to turn a story and reference images into a validated cases JSON file, ready to pass to demo.sh. It includes a standalone guide and supports a configurable Chat Completions API.

🤝 Acknowledgments

We thank the MiniMax-H3 team for opening their audio/video foundation model, and the FlashGen team for their four-step acceleration work and model release.

Our inference implementation builds on the open-source infrastructure provided by Diffusers, SGLang, and FlashAttention. See NOTICE for component attributions.

We also thank the WBench team for their open-source benchmark and evaluation tools for interactive video world models.

License

XGEN-JING code and model weights are released under the MiniMax H3 Community License Agreement. Third-party components retain their original licenses; see NOTICE.