Light-O1-Preview turns an instruction into whole-body humanoid action. It is pretrained on structured human action recovered from internet video, which gives it a transferable action prior, and then post-trained on purpose-collected data to adapt that prior to a target embodiment and to human intent.
Given a prompt, the model answers in two parts: it states in language what the instruction requires
of the body, then generates the action. The plan is readable, so you can see what the model
understood before it moved. Actions come out as a (frames, 138) unified human action
representation at 20 FPS, which renders a character directly or drives a behavior foundation model
on a robot.
Note
Light-O1-Preview releases the text-to-action capability only. We are working towards the full Light-O1, which will add vision and manipulation. Stay tuned — thanks for your interest!
The tech blog compares Light-O1-Preview with two public text-to-action models, HY-Motion-1.0 and Kimodo. On HY-Motion-Bench, a VLM judge decomposes each prompt into a checklist and scores the generated action question by question.
| HY-Motion-Bench, SSAE | Overall |
|---|---|
| Light-O1-Preview | 78.0 |
| HY-Motion-1.0 | 74.7 |
| Kimodo | 61.4 |
Protocols, baselines and caveats are in the tech blog.
| Model | Representation | Context | Download |
|---|---|---|---|
| Light-O1-Preview | human_action_138_v1 |
text | 🤗 Hugging Face |
The checkpoint ships the action decoder alongside the model weights:
checkpoint/
├── config.json
├── model.safetensors
├── tokenizer.json
├── tokenizer_config.json
├── chat_template.jinja
├── action_checkpoint_manifest.json
└── action_tokenizer/
├── decoder_config.json
├── manifest.json
└── action_decoder.safetensors
The checkpoint's architectures field must be Qwen3_5ActionForConditionalGeneration. Special tokens and
weight keys retain their trained spelling.
Try Light-O1-Preview without any local setup at Light-O1-Preview Playground — enter a prompt, watch the reasoning stream, and inspect the generated action in a 3D viewer.
There are three ways to run Light-O1-Preview locally, in increasing order of setup: a one-shot command line generation, the Python API, and a local web console with a 3D viewer.
Beyond that: running the Control Server and GPU API on separate hosts is covered in docs/deployment.md, the HTTP endpoints in docs/api.md, the optional MuJoCo simulation in docs/sonic.md, and tests and linting in docs/development.md.
| OS | Linux x86-64 |
| Python | 3.11 (the project pins >=3.11,<3.12) |
| GPU | NVIDIA, CUDA 13 compatible |
Note
macOS and Windows are not supported for inference. To try the model without a GPU, use the playground.
Download the weights from LightOriginsHQ/Light-O1-Preview. The action decoder ships inside the checkpoint — see 2. Model Downloads.
git clone https://github.com/lightorigins/Light-O1.git
cd Light-O1
uv sync --extra inferenceuv run --extra inference light-deploy \
--model /absolute/path/to/Light-O1-Preview \
--prompt "a person waves with the right hand" \
--thinking \
--output human_action.npyThis writes a float32 (frames, 138) array at 20 FPS. Pass --temperature, --top-p or --seed to
control sampling.
from light_deploy.generate import Inference
inference = Inference("/absolute/path/to/Light-O1-Preview", device="cuda:0")
output = inference.generate("a person raises the right arm and waves", enable_thinking=True)
output.action # (frames, 138) action array
output.reasoning # the Thinking textTwo interchangeable backends sit behind this API, selected per instance with backend=:
vllm(default) — the resident-GPU production runtime.transformers— torch-native, single-request, no vLLM. Works in lazy-init hosts such as ZeroGPU.
Both share the same masking FSM, prompt and output processors, and action decoder, so a token stream is masked and parsed identically whichever backend produced it.
uv run --extra inference light-deploy-server \
--model-path /absolute/path/to/Light-O1-Preview --port 8090Open http://127.0.0.1:8090 to enter prompts, watch the reasoning stream and inspect the generated action in
a 3D viewer. This also starts a managed GPU API on port 8030; change it with --inference-port. Model
loading is asynchronous, so generation becomes available only once that API reports ready. Without
--model-path the page starts without a model and lets you enter a checkpoint path.
Warning
This is a trusted-workstation service, not authenticated multi-user hosting — it permits local checkpoint selection and GPU work. Keep the localhost defaults, or put it behind an authenticated reverse proxy. Do not expose it directly to the Internet.
Our tech blog demonstrates Light-O1-Preview on both our LightBot robot and the Unitree G1. Light-O1-Preview outputs a unified action representation, which is connected to each robot's BFM for execution: our own BFM for LightBot, and GEAR-SONIC as the low-level motion controller for the Unitree G1. For the community, we provide a Sonic example that adapts this unified representation to GEAR-SONIC for closed-loop G1 simulation in MuJoCo.
The example produces a simulation video and rollout metrics. It requires a separate GEAR-SONIC low-latency policy checkpoint and runs in simulation only. See the setup guide to run it from the web console, or the example README for standalone usage and details of the action adapter.
The code in this repository is licensed under the Apache License 2.0.
The Light-O1-Preview weights are distributed under their own terms on the model page. Required model and asset notices are in THIRD_PARTY_NOTICES.md.
@misc{lightorigins2026lighto1,
title = {Light-O1: Scaling Whole-Body Intelligence with Human Action Pretraining},
author = {Light Origins Team},
year = {2026},
howpublished = {\url{https://www.lightorigins.com/en/blog/light-o1}}
}Join us on Discord, or scan the QR code to join the WeChat group:
Questions and bug reports are welcome as GitHub issues. For anything else, reach us through Discord or the WeChat group above.
