← Back
Yinsongxu

Yinsongxu/LLM2Jev

Turn local language models into Jev-style structured decision models. Get results from text and images with prefill alone—no token-by-token decoding required.

View on GitHub ↗
jevllmmllm
Stars
382
Forks
38
Watchers
382
Open issues
0
Contributors
2
Language
Python
License
Apache License 2.0
Default branch
main
Created Sep 19, 2026Updated Sep 26, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

LLM2Jev

🧠 LLM2Jev: Turn LLMs into Jev-Style Decision Models


Python License Jev API Backend

简体中文

Turn local language models into Jev-style structured decision models. Get results from text and images with prefill alone—no token-by-token decoding required.

LLM2Jev is an independent open-source project. It is not affiliated with or endorsed by Jev or TypeSafe.

📰 News

  • September 26 - JevBench evaluation: LLM2Jev accuracy and latency on 231 public items; P50 latency is below one tenth of the Jev official result.
  • September 23 - MLX backend: added text and image scoring, candidate batching, bounded prefix reuse, and a compatible System One HTTP service.
  • September 22 - Multimodal inputs: added text-and-image requests for SGLang, Transformers, and the System One HTTP API.
  • September 21 - Web and Snake demos: added interactive examples for composing mixed questions and model-driven decisions.
  • September 21 - Prefix reuse on cold requests: added staged candidate submission for reusing SGLang's Radix Cache, with architecture, usage, and benchmark documentation.
  • September 20 - SGLang and System One API: added the SGLang scoring backend and a compatible POST /v1/systemone endpoint.

✨ Key Features

  • Broad backend support: run local text and vision language models with SGLang, Transformers, or MLX on Apple Silicon through the Python API or HTTP service.
  • Prefill only: compute probabilities from logits during prefill and assemble results directly, without token-by-token decoding.
  • Multimodal inputs: combine text and images in state or instructions, with support for SGLang, Transformers and MLX-VLM.
  • Prefix reuse on cold requests: stage candidate submissions to reuse SGLang's Radix Cache within a single request, including a first request with no relevant cached prefix.

Candidates share state, and candidates for the same question also share its instructions. The SGLang backend first scores a real criteria candidate to establish the prefix cache, then submits candidates that can reuse it. Each candidate is scored once, reducing repeated computation for long inputs with many candidates. The MLX backend explicitly prefills shared prefixes before scoring candidate suffixes.

Staged candidate scoring reuses state and question instructions through SGLang Radix Cache.

Both backends preserve the same binary scoring interface. See the Usage guide for details.

Learn how it works: From Jev Request to LLM Request → Shared-prefix design.

🚀 Quick Start

On Linux with a supported NVIDIA GPU, run a local model through SGLang:

git clone https://github.com/Yinsongxu/LLM2Jev.git
cd LLM2Jev
uv sync --extra sglang
source .venv/bin/activate
python examples/sglang_inference.py --model-path /path/to/model

The example submits Choice, Score, and Noul questions and prints the response as JSON. Replace /path/to/model with a local Hugging Face-compatible causal language model directory.

For image serving, see Multimodal inputs.

📦 Installation

See Installation for environment requirements, SGLang, Transformers, and MLX dependencies, and uv or pip installation.

📖 Getting Started

See the Usage guide for complete examples:

  • Offline: Python API
  • MLX backend on Apple Silicon
  • Online: HTTP service
  • Choosing between staged and all
  • Multimodal inputs

🎮 Demos

LLM2Jev web demo LLM2Jev Snake demo
Web demo Snake demo
MuJoCo pick-and-place demo
MuJoCo pick-and-place demo

📊 Benchmarks

See Performance benchmarks for the Qwen3-1.7B / RTX 5090 measurements, test conditions, and comparison of staged and all across cold and warm caches. Gains depend on input length, candidate count, and cache state.

JevBench public accuracy

This test covers 231 public JevBench items. LLM2Jev with Qwen3.5-4B measured 76.2% accuracy.

JevBench public accuracy comparison

The recorded LLM2Jev latency is P50 48 ms / P95 346 ms, while the Jev 1.13.0 values are P50 652 ms / P95 722 ms. See the JevBench report for the data and reproduction commands.

🗺️ Roadmap

  • More benchmarks across model sizes, datasets, and workloads, covering decision quality, latency, and throughput.
  • An interactive web demo for submitting questions and inspecting probabilities.
  • Initial local-image support for Transformers and SGLang.
  • More multimodal tasks and demos.

🧪 Tests

python -m unittest discover -s tests -v

📄 License

This project is licensed under the Apache License 2.0.