Video generation, references, motion controls and editing.
Paper🔥HOT · Introduction · Overview · Models · Quick start · ComfyUI · Tasks · Usage guides · Roadmap · Changelog · Organization · Acknowledgments · Citation
Important
🔥 ComfyUI supports Standard four-step and Flash three-step generation. Workflows, required
nodes and model layout are under comfyui/. For limited VRAM, use the matching
Standard Lite or Flash Lite workflow and checkpoint at its shipped step count; see
Recommended models for limited VRAM.
This repository is an early beta. Bugs, compatibility issues, unfinished features and inconsistent generation quality may remain.
For limited VRAM, we strongly recommend Standard Lite to preserve Standard quality, or
Flash Lite for a much smaller memory footprint and very fast generation. Start with the
matching Lite checkpoint and *_lite.json ComfyUI workflow when choosing a memory-saving setup.
- Standard Lite · quality first: BF16 Lite is 37.6 GiB; INT8 Lite is 20.4 GiB. At the shipped 4 steps, each preserves its corresponding original checkpoint's tested outputs; the BF16/INT8 switch remains available.
- Flash Lite · memory and speed first: the INT8 DiT is 16.7 GiB, paired with the Light VAE and shipped 3-step workflows. It retains Flash's fast generation and tested quality; Lite reduces the checkpoint footprint rather than promising extra speed over full Flash.
These are DiT file sizes, not total runtime VRAM requirements. See the ComfyUI Lite setup and checkpoints.
LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows
Technical report, September 2026.
lynnreal-demo-preview-9.4MB.mp4
Preview compressed to ~10 MB for this repository; for the full-quality demo, click the YouTube or Bilibili link below.
| YouTube | Bilibili |
|---|---|
| Watch on YouTube ↗ | Watch on Bilibili ↗ |
One model, unified across many video tasks. Built on a 32B shared multimodal diffusion transformer (following the MiniMax H3 architecture), LynnReal-Omni brings text-to-video, image-to-video, human- and hand-pose guided generation, structural control, omni-reference generation, style transfer, video editing, degraded-video restoration (we find that it also repairs videos affected by accumulated error) and streaming long-video generation into a single framework, all at four-step fast generation. It also accepts heterogeneous inputs such as appearance references, editable 3D renders and game recordings, so an Agent can compose visual conditions inside one model.
Flash, built for real-time rendering. We train a 27B LynnReal-Omni-Flash (three-step generation) that lowers inference cost through model and decoding acceleration and a lightweight VAE decoder. On a single H100, warm generation and decoding of a 22-frame 540p video takes 843 ms with the Standard model and 377 ms with Flash, laying the groundwork for real-time streaming video generation.
Data pipeline and MSAVP evaluation. We build a systematic multi-shot and omni-reference data pipeline covering video cleaning, subject association, multimodal annotation and alignment control. From the large corpus of collected videos we select a high-quality multi-shot audio-visual subset and extract a variety of omni-reference condition controls. We also propose MSAVP for evaluating multi-shot audio-visual generation.
| Generate | Control | Edit & repair |
|---|---|---|
| Text-to-video and first-frame conditioning | Single- and multiple-subject references | Instruction-guided image editing |
| Video continuation | Body- and hand-pose control | Video appearance editing |
| Speech and video generation | Game- and mesh-video rendering | General frame-by-frame video repair |
Standard · 4 steps / Flash · 3 steps / Optional lightweight VAE
Standard uses four denoiser forwards. Flash uses three forwards with 42 blocks, spatial token selection, INT8 projection weights and INT8 activations. Step counts refer to each generation call: Standard first-pass generation uses 4 steps; independent frame repair uses 4 steps per source frame. Long streaming can add 2 refinement steps per later chunk; see streaming generation.
The repositories below are the planned Hugging Face upload locations. Weights will become downloadable as their uploads are completed.
| Model | Sampling | Weights | Status |
|---|---|---|---|
| Standard · BF16 DiT | 4 steps | Hugging Face ↗ | Completed |
| Flash · W8A8 DiT | 3 steps | Hugging Face ↗ | Completed |
| Lightweight VAE | Optional codec | Hugging Face ↗ | Completed |
| Standard DiT INT8 | 4 steps | To be announced | Coming soon |
Place each downloaded bundle under weight/, preserving its configurations and
component subdirectories:
weight/
├── standard/ # Standard bundle: DiT, text encoder, codecs and configs
│ └── transformer/ # The single physical Standard DiT checkpoint
├── flash/ # Flash bundle
├── vae/ # Official VAE
└── light-vae/ # Optional lightweight VAE
Pass the complete weight/standard/ bundle to --weights. Its DiT lives in
weight/standard/transformer/; the other required components are loaded from the
bundle too. The lightweight VAE does not replace the DiT or text encoder.
Standard DiT INT8 is coming soon. Experimental Standard INT8 launchers are already in the codebase; the downloadable checkpoint has not been released. Flash W8A8 is a separate model variant.
Run the following commands from the repository root, with the required local weights in place.
From Conda base, install this checkout and activate the new environment:
pip install .
conda activate lynnrealThe bootstrap creates or reuses a Python 3.12 environment named lynnreal and
checks attention on the visible GPU. Use pip install . -v for detailed progress.
See Runtime notes for other environments, offline installation
and hardware-specific behavior.
# Standard: five seconds, native 768p, four denoiser forwards
bash script/sample/standard/bf16/t2v.sh --frame 5s --resolution 768p
# Flash: three denoiser forwards
bash script/sample/flash/int8/t2v.sh# First-frame conditioning
bash script/sample/standard/bf16/ti2v.sh --image test/assets/times_square.png
# Continue a video using the launcher's example input
bash script/sample/standard/bf16/v2v.sh
# Subject references: replace these paths with your own images
bash script/sample/standard/bf16/ref2v.sh --ref_image first.jpg,second.jpgEach launcher saves the video, prompt, full log and measured DiT/decoder times
under output/. --frame 120f requests exactly 120 frames; --frame 5s is
five seconds at 24 fps. The default output is 1344 × 768. Standard BF16 and Flash
W8A8 have separate entry points. The experimental script/sample/standard/int8/
launchers remain available for development; the downloadable Standard DiT INT8
checkpoint is coming soon.
For a repeatable run, use bash script/sample/standard/bf16/t2v.sh --seed 7 --frame 5s --name my_case.
Choose a new output name for each run.
Keep prompt, seed, geometry, references and attention backend fixed for comparisons.
Changing attention precision can change a four-step result. Standard BF16 defaults
to the native attention backend and official VAE.
We also ship a ComfyUI port — it is open, and you are very welcome to try it! 🎉
comfyui/ contains workflows for t2v, i2v, r2v, pose2v and v2v on the Standard
four-step checkpoints and for t2v, ti2v and ref2v on the Flash three-step checkpoint, the
small custom node pack they need (ComfyUI-LynnReal), the demo assets the workflows load, and
the exact model files each task needs. Those files are published in ComfyUI format alongside the
release weights on Hugging Face under
🤗 stdstu123/LynnReal-Onmi-beta-0.1 · comfyui/models.
Drop the workflows, the node pack and the models into your ComfyUI install and they run — no
launcher flags needed. Every task, its workflow and the weights it loads are listed in
comfyui/README.md.
The Flash checkpoints run in ComfyUI at their trained three steps: t2v, ti2v (one first
frame) and ref2v (reference pictures), W8A8 DiT plus the Light VAE! On a single H100 80 GB
at 1344×768, warm — model already loaded, the way a session runs — three measured runs per cell:
| Task | 5 s · generate | 5 s · click-to-video | 10 s · generate | 10 s · click-to-video |
|---|---|---|---|---|
| Text → video | 8.4 s | 12.2 s | 22.0 s | 29.1 s |
| First frame → video | 8.9 s | 13.0 s | 23.1 s | 30.1 s |
| References → video | 9.5 s | 13.2 s | 24.2 s | 31.1 s |
A five-second 1344×768 clip with native stereo audio — three denoiser steps and both decoders included — in about eight and a half seconds, on one GPU. Generate is the release's own Generate-wall convention (first denoiser forward to decoded frames); click-to-video is what you actually wait for, prompt encoding and muxing included, and repeat runs agree to ±0.02 s!
Warning
Videos longer than 11 seconds are not usable in the ComfyUI Flash path yet — the accelerated path for long clips is still being fixed.
Nothing is assumed about your card: the node pack picks the fastest attention it can find and verifies it numerically (FlashAttention-3 → FA2 → cuDNN SDPA → native), falls back to ComfyUI's own block math whenever a fused kernel is unavailable, and tunes the INT8 GEMMs for the GPU it actually runs on!
Warning
The ComfyUI port is experimental and under active construction 🚧
- It runs the same checkpoints and the same schedules, but the pipeline around them is
ComfyUI's, so performance numbers should be taken from the original scripts in
script/sample/— those are the reference implementation and the ones quoted in the paper. The port is measurably slower than the scripts today (fewer fused kernels, a different attention backend and no Hopper-specific INT8 grouping yet). - Same-seed output is not comparable between the two engines: the noise source and the decoder path differ. Compare quality, not pixel identity.
- The port covers the Standard four-step checkpoints (with an optional INT8 switch on the
canvas, off by default except
pose2v, backed bylynnreal_omni_standard_int8.safetensors) and the Flash three-step checkpoint. Videos longer than 11 seconds are not usable in the ComfyUI Flash path yet — the accelerated path for long clips is still being fixed. - We will keep improving it — speed, memory, more tasks and cleaner packaging are all on the list. Issues and pull requests are very welcome! 🙌
script/sample.py accepts repeated --reference paths. Image and video prompt
labels count separately from one. --aligned-reference indexes the complete
input list from zero: an appearance image followed by a pose video uses index 1.
| Task | Mode | Inputs |
|---|---|---|
| Text to video | t2v |
prompt |
| First-frame conditioning | ti2v --native-keyframes |
first image |
| Single or multiple subjects | reference |
reference images |
| Body/hand motion | reference |
appearance image and aligned pose video |
| Video editing | video-edit |
aligned source video, optional style image |
| Image editing | image-edit |
source image; output filename ends in .png |
Image editing saves both a selected image and its complete companion clip. See the prompt-writing guide for task-specific instructions and reference conventions.
frame_repair.sh is a general repair entry point. Supply your own video and
repair instruction for the scene and artifacts you want to address.
bash script/sample/standard/bf16/frame_repair.sh \
--video /path/to/input.mp4 \
--prompt-file /path/to/repair_prompt.txt \
--output output/my_frame_repair- Input: 24 fps; width and height must be divisible by 32.
- Process: edit each source frame independently, then assemble a silent video at 24 fps.
- Budget: four DiT forwards per source frame; 120 frames require 480 forwards.
- Preview: add
--indices 0,24,48 --keep-clipsto inspect a few frames before a full run. - Limitation: independent edits may introduce brightness or shape fluctuations across frames.
Understand frame selection
The default image-edit call generates an internal 22-frame clip and selects
index 11 (zero-based). Use --image-frame to choose another valid index.
One selected image is retained for each original source frame; the internal clip
does not extend the source timeline. No temporal interpolation or generated-frame
feedback is used.
A corrected still alone does not establish that the whole video is repaired. Inspect the complete assembled sequence for remaining artifacts and flicker.
script/sample/standard/bf16/stream.sh
generates an image-conditioned first chunk and continues its latent history.
Run these commands from the repository root:
# Default: 5 seconds, 768p, 24 fps, seed 7; four DiT forwards per chunk.
bash script/sample/standard/bf16/stream.sh --frame 5s
# 30 seconds: preserve the initial four-step prefix, then use 4+2-step refinement.
bash script/sample/standard/bf16/stream.sh --frame 30s
# 30 seconds with four steps throughout; disable refinement and context refresh.
bash script/sample/standard/bf16/stream.sh --frame 30s --no-refresh-context
# Generate only the native 22-frame first chunk.
bash script/sample/standard/bf16/stream.sh --first-chunk-onlyThe default five-second run uses 4 steps throughout. Above 136 requested
frames, the default enables context refresh: the bootstrap and first seven
continuation chunks use 4 Standard BF16 DiT forwards each; continuation
chunk eight onward uses 4 + 2 forwards, including the second pass.
--no-refresh-context keeps the fixed-context, four-step-only policy at any duration.
These modes use the same DiT in weight/standard/transformer/ and the official VAE.
| Configuration | Output frames | Total DiT forwards |
|---|---|---|
| First chunk only | 22 | 4 |
| Default 5 seconds | 120 | 32 |
| Default 30 seconds, with later refinement | 720 | 242 |
30 seconds with --no-refresh-context |
720 | 172 |
The totals include the bootstrap and every continuation chunk. They describe sampling work, rather than elapsed time or text-encoder/decoder computation. Long clips refresh image-aware text from the preceding chunk after the stable prefix. Local texture and shape drift can still occur.
Custom images, prompts and continuation captions
bash script/sample/standard/bf16/stream.sh \
--image inputs/first.png \
--prompt-file inputs/initial.txt \
--captions inputs/continuations.json \
--frame 5s --seed 7 \
--output output/my_stream--image: a 1344 × 768 first image; this entry currently supports only 768p.--prompt-file: the initial scene and action prompt. Use--prompt "..."for literal text instead.--captions: a JSON array of nonempty strings describing successive motion intervals. Provide at leastceil((frames - 17) / 17)captions: 7 for 5 seconds, 42 for 30 seconds.--frame:5,5sand120fall request 120 frames at 24 fps.--output: a new output directory; existing runs are preserved.
A custom image or initial prompt requires a matching continuation-caption file,
except with --first-chunk-only. Without custom inputs, the launcher selects the
bundled rainy-street image and five- or thirty-second caption plan.
Use --dry-run to validate inputs and inspect the planned commands and step counts
before loading the models. Set LYNNREAL_PYTHON if the runtime uses another interpreter.
Each run saves video.mp4, sample.log and run.json under
output/standard/bf16/stream/<timestamp>/, or the directory supplied with
--output. The first_chunk/ and continuation/ subdirectories retain their
prompts, latents and per-chunk timing records. run.json records the planned
step budget; per-chunk metadata records actual DiT forwards.
Long-stream context refresh stages the text encoder through CPU memory on a
single GPU, which increases end-to-end latency. An optional second-GPU
conditioning worker can be selected with --conditioner-service.
See streaming instructions for worker startup, native overlap,
appearance constraints and --audit-codec verification.
For continuation from an existing video, use
script/sample/standard/bf16/v2v.sh --continuation.
| What to explore | Documentation |
|---|---|
| Sampling launchers and arguments | Sampling guide |
| Timing and benchmark commands | Speed-test guide |
| Streaming generation and validation | Streaming reproduction |
| Prompt structure and reference conventions | Prompt-writing guide |
Installation options · Conda, existing environments and offline wheels
The local build hook creates/reuses the lynnreal Conda environment (Python 3.12),
installs the inference package and dependencies there, then probes attention on the
visible GPU. New environments use conda-forge without changing global channel settings.
Base receives only the small lynnreal-bootstrap receipt package.
Failures propagate to pip; installation does not activate the parent shell.
Use pip install . -v to see build/probe progress. The bootstrap uses PyPI for its
child installs; it does not change global pip settings. Weights remain in weight/.
For an existing complete local wheel cache, set LYNNREAL_WHEELHOUSE=/path/to/wheels
to install the same requirements offline inside the target environment.
Outside base, pip install . installs into the active Python 3.12/3.13 environment.
To explicitly install into the current environment, including base:
python script/setup_env.py --current-envThe standalone helper also creates the environment and shows progress directly:
python script/setup_env.py --pypi-only
conda activate lynnrealIt bootstraps checksum-verified Miniforge if Conda is absent. Existing incompatible
Python environments are not downgraded or deleted; choose a new --env-name.
--no-shell-init leaves shell startup files untouched. --skip-attention installs
only core dependencies. For ordinary wheel builds from base, set
LYNNREAL_INSTALL_CURRENT=1 to disable the local Conda bootstrap.
Attention and GPU compatibility · probes, memory and tested devices
Attention is tested in a fresh process, in this order: FA3 (Hopper with CUDA >=12.3),
FA2 (Ampere or newer), cuDNN SDPA, PyTorch Flash SDPA, then native SDPA. FA3 is pinned to commit
203b9b3dba39d5d08dffb49c09aa622984dff07d and built from source when installation
is needed. RTX 4090 cannot run Hopper FA3 and normally selects FA2. Legacy FA1 is
not installed over FA2 because it does not satisfy this H3 BF16 interface.
python script/setup_env.py --attention-only
# Strict FA3 acceptance on a supported Hopper GPU:
python script/setup_env.py --attention-only --attention _flash_3Auto mode prominently warns when FA3 is not active and records the selected backend
and numerical probes in output/setup/attention.json. An explicitly requested
backend must pass; otherwise installation exits nonzero. No working GPU means
attention remains unverified: rerun the probe on the sampling worker. Attention
backends can produce different four-step outputs; probe success alone does not
establish identical video quality. Sampling launchers retain explicit
--attention-backend selection for controlled comparisons. On GPUs below 64 GiB,
Flash launchers decode one tile at a time to bound temporary memory; spatial tile
geometry and blend order stay unchanged. This does not make the Standard model
fit into the same memory budget.
Blackwell GPUs select the same pinned PyTorch package versions with CUDA 12.8 wheels when the installed build is older. This follows PyTorch's Blackwell support requirements. H100 uses the pinned FA3 commit; its source build retains the dense FP16/BF16 forward paths used by H3 (head dimensions up to 128), omitting unused training/KV-cache kernels. FA3 is not selected on RTX 40/50-series, A100 or B200 merely because it is installed: other devices try FA2, cuDNN SDPA, PyTorch Flash SDPA, then native SDPA, with a numerical probe and a visible non-FA3 warning. An explicitly requested backend remains a strict requirement.
Actual GPU validation for this revision covers H100 80GB and the available RTX
4090 with 48GB memory. A100, RTX 50-series and B200 use compatibility selection
but have not been measured here. Kernel compatibility does not remove model
memory requirements; a standard 24GB RTX 4090 is not equivalent to the tested
48GB node. See output/setup/review/index.html for generated samples and
output/setup for installation/probe records.
On an uncached prompt, the text encoder retains the 50 layers actually consumed by H3 before moving to CUDA. When memory is limited, complete text layers are staged through CPU without changing their precision or execution order. The cached conditioning features remain compatible. This prevents the unused 14 text layers from causing a cold-prompt allocation failure on the tested 48GB RTX 4090; text encoding remains separate from DiT latency.
Kernel precompilation and persistent caches
Installation also precompiles representative inference kernels on the visible
GPU and writes output/setup/kernels.json. When local weights and GPU memory
permit, it also prepares both models at 540p/22 frames and 768p/5, 10, 15 seconds.
On the tested 48GB RTX 4090, full-model preparation covers Flash at 540p/22 frames.
Completed profiles are recorded under output/setup/profiles-*.json; repeated
installation reuses matching profiles. Profile signatures follow the local inference
imports, model/decoder configurations, and software versions; unrelated streaming
experiments do not invalidate them. Other shapes are prepared on demand.
Set LYNNREAL_PRECOMPILE=0 to skip full-model preparation. The installer and INT8 launchers share
persistent caches under output/.cache; GPU architecture, kernel source and
software versions distinguish cached selections. Unsupported optional compilation
emits a warning and sampling uses the compatible path; a numerical mismatch
fails installation instead of being accepted as a successful fallback.
Hopper INT8 inference additionally uses persistent TMA GEMM for measured long
projection shapes and a single-warp Q/K kernel that retains the reference reduction
order. Short GEMMs and non-Hopper devices keep their existing dispatch. Set
LYNNREAL_LONG_KERNELS=0 for the paired reference configuration. These changes
preserve exact latents, audio and RGB in the completed paired comparisons;
this statement does not cover changing attention, quantization or decoder settings.
Precompilation does not remove model loading or fresh-process tracing. The public
INT8 commands list preparation calls separately; their warm DiT+decoder times
are not shell-command elapsed times. Dense attention remains the main cost for
768p long clips, so short-clip speed does not extrapolate linearly with duration.
See output/speed_test/long_acceleration_review/index.html for measured pairs and videos.
Codecs and timing · official VAE, lightweight VAE and measurement scope
The core Python API and BF16 shell launchers retain the official VAE default.
INT8 shell launchers select the compiled lightweight decoder and adaptive tiles.
In the Python CLI, --light-vae weight/light-vae selects the 26-block distilled
decoder with the unchanged encoder and latent interface.
Both use indexed Safetensors and component configs. For reconstruction comparisons,
use the same encoded latent and report precision, tile geometry and hardware.
Timing records separate actual denoiser forwards and video decoder execution. BF16 refers to the DiT; the official video decoder uses its native FP16 autocast. Cold loading, conditioning, audio decoding and output encoding are outside this sum; end-to-end generation is recorded separately. CPU offload can substantially increase latency and must be disclosed in speed comparisons.
model/ inference, conditioning, codecs and kernels
script/ sampling and video continuation
skill/ prompt-writing guidance
weight/ standard/, flash/, vae/, light-vae/
test/ default prompts, reference media and continuation plans
output/ generated media, complete logs and timing records
tool/ helpers required by streaming and structure editing
All model components are local to weight/, including the text encoder,
tokenizer, image processor, audio codec and schedulers. The DiT bundles contain
complete inference parameters. No additional adapter is required. Weight paths
outside this directory are rejected. Model licenses and upstream attribution
are retained in each bundle.
Standard uses one DiT checkpoint in weight/standard/transformer/.
Standard clip launchers validate this directory and run four denoiser forwards
per generated clip. Streaming uses four per chunk; when context refresh is enabled,
continuation chunk eight onward adds two refinement forwards. The internal reference-task
component name transformer_ref maps to the same transformer/ directory through
modular_model_index.json; no transformer_ref/ weight directory is needed.
The sampler loads only the DiT component selected by the task layout. See
the model card included with the downloaded Standard weight bundle for details.
This repository will continue to receive updates. The current beta is still imperfect, and we appreciate your patience with its remaining defects.
Our planned releases and improvements include:
- The Standard DiT INT8 checkpoint.
- Improved model checkpoints with better generation quality and consistency.
- Training code to support further experimentation and model development.
- Selected datasets or dataset subsets, subject to their redistribution permissions.
- Continued fixes to installation, hardware compatibility, inference, documentation, and reproducible examples.
- ComfyUI: keep the Light VAE decoder compiled on DynamicVRAM machines — it currently falls back to eager decoding there, which roughly doubles decode time.
These are development plans, not a fixed release schedule. Availability and usage instructions will be updated here as each release is ready. Reproducible bug reports and feedback on failure cases are welcome.
This project is developed by Lynnreal Lab.
Leader: Xiaofeng Mao, Shaohao Rui, Weijie Ma
Core Contributors: Xiaofeng Mao, Peijia Lin, Shaohao Rui, Yibo Zhang, Haibin Wan, Weijie Ma
Welcome students with backgrounds in 3D reconstruction and interactive world models to apply for internships and collaborate! Please send your resume to hr@lynnreal.com.
We especially thank the MiniMax H3 team for the model that forms the foundation of LynnReal-Omni.
We also thank Qwen3-VL for the multimodal encoder, tokenizer and processor implementation used for conditioning.
The following community projects provided useful implementation references and technical discussions during our development and experiments:
| Project | Reference and contribution |
|---|---|
| ComfyUI-JZL-MiniMax-H3 | Ref2VA reference encoding, multimodal input organization and reference-scale handling. |
| ComfyUI-Minimax-H3-Prompt-Builder | Structured Ref2VA prompts, multi-segment continuity and second-pass sampling workflows. |
| DiffSynth-Studio | H3 reference/keyframe conditioning and training-loss implementations consulted during development. |
| MiniMax-H3-FineTuning | Community H3 fine-tuning guidance and numerical-correctness discussions. |
| ComfyUI-MiniMax-H3-Turbo | Community sampling implementations reviewed for audio/video time schedules and prediction conventions. |
We thank the authors of these projects for sharing H3 refinement implementations and workflows that informed our second-pass sampling experiments:
| Project | Reference and contribution |
|---|---|
| h3-latent-upscaler | Spatial latent upscaling between low-resolution and high-resolution sampling passes; see our sampling guide. |
| Comfyui_Minimax_h3_latent_Upscaler | Learned H3 latent upscaling and split-upscale workflows examined in our refinement experiments. |
These acknowledgments include development references and experimental comparisons; the released sampling paths and their validated settings are documented separately in this repository. We appreciate the authors' open-source contributions and the community's reproducible bug reports and feedback.
Please see LICENSE, NOTICE, and the respective upstream projects for their license terms and attribution notices.
If you use Yume for your research, please cite our paper:
@misc{mao2026lynnrealomninativemultimodalvideo,
title={LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows},
author={Xiaofeng Mao and Peijia Lin and Shaohao Rui and Yibo Zhang and Haibin Wan and Weijie Ma},
year={2026},
eprint={2609.15863},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.15863},
}
Thank you for trying LynnReal-Omni.
Better models, training code and selected datasets are planned.
Back to overview ↑