← Back
OmniJev

OmniJev/OneJev

🚀🚀 A multimodal System One decision model that gives calibrated answers to typed questions about screens, photos, video and text in one forward pass.

View on GitHub ↗https://omnijev.github.io/OneJev/ ↗
calibrationcomputer-usedecision-modelgui-agenthuggingfacejevllmmultimodalpytorchqwensystem-onetypesafevideo-understandingvision-language-modelvlm
Stars
103
Forks
12
Watchers
103
Open issues
0
Contributors
1
Language
Python
License
Apache License 2.0
Default branch
main
Created Sep 27, 2026Updated Sep 30, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

OneJev, a Multimodal System One Decision Model

English · 简体中文 · 日本語

Hugging Face Demo GitHub Website License

Data API Python PyTorch Transformers

OneJev is a multimodal System One decision model. It returns calibrated probabilities for typed questions about screenshots, photos, videos and text in a single forward pass. Available in 0.8B, 4B, 9B and 27B.

Results

Block bar charts: the four OneJev sizes against Jev 1.13, Jev-Omni 12B and Qwen3.8-27B thinking on the OneJev test set, DecisionBench hard, TypeSafe and MMStar

Results table, best score in each row in magenta, second best in light pink

Accuracy (%). The OneJev test set is held out from OneJev training. Jev 1.13 uses published text-only scores; Jev-Omni and Qwen3.8-27B thinking were evaluated by us.

Latency on one H200 for one 1280x720 screenshot: OneJev-0.8B 31 ms for 1 question, 51 ms for 10, 5.1 ms per question; OneJev-4B 64 ms for 1 question, 104 ms for 10, 10.4 ms per question; OneJev-9B 81 ms for 1 question, 131 ms for 10, 13.1 ms per question; OneJev-27B 189 ms for 1 question, 324 ms for 10, 32.4 ms per question

Quick start

Choose one backend, then run the Python example below.

Option A: PyTorch

For NVIDIA GPUs. Supports text, images and video.

pip install "qev[torch] @ git+https://github.com/OmniJev/OneJev.git"
qev serve --model OmniJev/OneJev-4B --port 8000

Option B: llama.cpp

For GGUF models. Supports text and images; use PyTorch for video. Install llama.cpp first (brew install llama.cpp on macOS).

pip install git+https://github.com/OmniJev/OneJev.git
qev serve --gguf mradermacher/OneJev-4B-GGUF:Q8_0 --port 8000

Send a request

Both backends serve the same API at http://localhost:8000. In another terminal, run this example with your own screenshot.png:

from qev import Client, Choice, Noul, Score
from qev.media import data_uri

r = Client("http://localhost:8000").system_one(
    state={"task": "Pay the open invoice from ACME", "screen": "<image:1>"},
    media=[{"type": "image", "data": data_uri("screenshot.png")}],
    questions={"done": Noul("The invoice has been paid"),
               "next": Choice("What should the agent do next?",
                              {"click": "click an element", "type": "type text", "scroll": "scroll", "stop": "stop"}),
               "progress": Score("How far along is the task?", ["not started", "halfway", "almost done", "done"])},
)
r.answers["done"].noul                  # probability of yes
r.answers["next"].probabilities         # one probability per option
r.answers["progress"].score             # expected level

The API is compatible with TypeSafe System One. More examples: video, curl, official TypeSafe SDK.

Training

git clone https://github.com/OmniJev/OneJev.git
cd OneJev
pip install -e ".[train]"

# Choose A or B. A is enabled below.
# A. Demo: 100 examples with images (23 MB)
hf download OmniJev/OneJev-Data sample/sample-100.parquet --repo-type dataset --local-dir data/onejev
python -m train.unpack data/onejev --sample

# B. Full dataset: 94,707 examples (17.7 GB). Uncomment these two lines instead of A.
# hf download OmniJev/OneJev-Data --repo-type dataset --include "data/*.parquet" --local-dir data/onejev
# python -m train.unpack data/onejev

Either option creates data/onejev/train.jsonl and extracts images and video frames into data/onejev/media/.

OneJev-Data releases 94,707 of the original 99,193 training questions; the remaining sources do not permit redistribution.

Fine-tune

The four models fine-tune Qwen3.5-0.8B, Qwen3.5-4B, Qwen3.5-9B and Qwen3.8-27B for one epoch with the vision tower frozen. The configs use a learning rate of 5e-6, a 16,384-token limit, and cross-entropy plus Brier loss over answer probabilities. To train the 4B model on four GPUs:

torchrun --nproc-per-node 4 -m train.sft --config train/configs/onejev_4b_full.yaml

Configs: 0.8B · 4B · 9B · 27B. Set --nproc-per-node to your GPU count; the 9B and 27B configs enable FSDP to shard model and optimizer state. For your own data, change train and media_root in the config.

The final 4B checkpoint is saved to train/runs/onejev_4b/final. Start it with:

qev serve --model train/runs/onejev_4b/final --port 8000

Citation

@misc{onejev2026,
  title        = {{OneJev}: A Multimodal System One Decision Model},
  author       = {{OmniJev Team}},
  year         = {2026},
  howpublished = {\url{https://github.com/OmniJev/OneJev}}
}

Friendly Links

  • LINUX DO
  • Jev
  • Awesome JEV
  • Awesome JEV Website
  • PlayJev

License

Apache 2.0