Jiamu Zhang1 Tianze Yang1 Yucheng Shi2 Liang Wu1
1 Nokia, Sunnyvale, CA 2 Tencent Hunyuan
Qwen3-8B on a real BANKING77 item. Every number is a model output.
Tip
🆕 vLLM serves every level, L2 included. An embed server's pooler hands back the hidden
state a closed-form head reads, so a decision endpoint is a pooling server plus a few
kilobytes of head. python -m anyjev.pipeline <model> converts, serves and measures in one
command. Start here ↓
Three commands take a model off the Hub and put a calibrated decision endpoint in front of it.
pip install "anyjev[hf]"
# 1. keep the blocks a decision needs — usually about two thirds
python -m anyjev.truncate Qwen/Qwen2.5-7B-Instruct 18 ./qwen-b18
# 2. serve it. L2 reads a hidden state, so the pooler hands one back untouched
vllm serve ./qwen-b18 --task embed \
--override-pooler-config '{"pooling_type":"LAST","normalize":false,"softmax":false}'from anyjev import Decider, Question
from anyjev.backends.vllm import VLLMBackend
d = Decider(VLLMBackend("http://localhost:8000", "./qwen-b18"), level="L2")
route = Question.choice("Which team should handle this?",
["billing", "technical", "sales", "other"], name="route")
d.fit_head(route, states, labels, layers=[-1]) # 100–300 labels, one closed-form solve
d.decide(ticket, [route])["route"].distribution # {"billing": 0.81, "technical": 0.07, ...}Without labels, turn on the rotation budget — recommended for any K-option choice. L0 asks the
model once per option rotation so that no option is favoured by its position. Most decisions do not need
all K: read them one at a time, stop when the leader is far enough ahead, and the threshold can be
calibrated so the answer matches the full cycle's a stated fraction of the time — measured against
our own full-strength readout, so it needs no labels at all.
d = Decider(VLLMBackend("http://localhost:8000", "./qwen-b18"), adaptive_shifts=True)
d.calibrate_adaptive(route, unlabelled_tickets, target=0.01) # a few hundred states, no labels
d.decide_batch(tickets, route) # diagnostics: shifts_used, stop_threshold7.2 rotations instead of 18 at a certified 1% disagreement rate, 2.2× the decisions per second on
vLLM and 2.3×–2.7× on the transformers backend, accuracy unchanged
(docs/rotation_budget.md). It is opt-in in 0.2.0 only because the tables in
docs/ were measured before it existed; it becomes the default when they are regenerated.
An L2 deployment is a pooling server plus a few kilobytes of head. No logits, no parsing, no
patched engine, and nothing generated. raw / L0 / L1 run the same way against a --task generate
server. A head fit through transformers and served by vLLM answers the same as one fit and
served on either alone — 99.0% identical answers, mean |dp| 0.0011 on BANKING77-20.
Measure it on your own box instead of trusting ours:
python -m anyjev.pipeline Qwen/Qwen2.5-7B-Instruct --labels-from banking20That truncates, serves, fits a head, measures accuracy, ECE and ms per decision on held-out
states, shuts the server down, and repeats at full depth so there is something to compare
against. Timings are a median over --repeats passes with the spread printed next to them,
because on a shared machine a single pass can report the same configuration as both faster and
slower than the baseline.
Depth is usually a gain, not a trade. Cutting Qwen2.5-7B from 28 blocks to 18 left accuracy slightly higher and calibration better, and was faster: a middle block is a better feature space for a linear head than the last one, where the remaining blocks are busy turning the answer into tokens.
--quantization fp8is available and not recommended — it buys single-question latency and costs accuracy.
Ask any open LLM a typed question — a choice, a yes/no, a score — and get a decision with a probability you can threshold, read from one prefill of its next-token distribution. Nothing is generated and nothing is parsed. Raw logits change their answer when you reorder the options and their confidence cannot be trusted; AnyJev fixes the first with zero labels and the second with a few hundred.
| ⚪ raw logits | 🔵 L0 zero labels |
🟢 L1 + temperature |
|
|---|---|---|---|
| Labels required | none | none | 100–500 |
| Answer flips when options are reversed | 0.230 | 0.073 | 0.077 |
| Accuracy | 0.747 | 0.803 | 0.807 |
| Calibration error (ECE) | 0.240 | 0.184 | 0.095 |
| Auto-decidable at ≤5% error | 7.7% | 46.3% | 52.0% |
Qwen3-8B, BANKING77 20-way, 300 test items · every ablation
The last row is the point: accuracy moves six points, but the share of traffic you can safely automate goes 7.7% → 52.0%. With raw logits a "0.9" is not trustworthy enough to act on, so everything goes to a human. Once the probability means what it says, you can set a threshold.
A closed-form head per question, solved on 100–300 labels in seconds — no gradients, the model's weights untouched — and read from one prompt stopped partway down.
| model | L0, zero labels | L2 | block | cost vs one forward |
|---|---|---|---|---|
| Qwen3-1.7B | 0.494 | 0.730 | 18 / 28 | 0.70× |
| Qwen3-4B | 0.564 | 0.786 | 24 / 36 | 0.69× |
| Qwen3-8B | 0.647 | 0.771 | 24 / 36 | 0.68× |
| Qwen3-30B-A3B | 0.630 | 0.799 | 40 / 48 | — |
| Qwen3-32B | 0.700 | 0.798 | 52 / 64 | 0.84× |
LocalLLaMA/typed-decisions, 20 questions × 300 labels, 2,000 held-out decisions. Pooled ECE 0.03–0.05. Jev 0.727 and fine-tuned Laya 0.768 on the same set, as published by their authors · every cell
A 1.7B at 64% of its depth reaches the number Jev publishes; a 4B ties the fine-tuned 421M Laya.
100 labels already put the 8B head at 0.740. Heads for five Qwen3 models ship in
anyjev-heads/, 23 heads per model in one 1.8–4.4 MB file.
A head maintains itself. Only its feature mean and scale move afterwards, re-estimated from unlabelled traffic, so it follows its question across rewordings and option orders on its own — a reworded question drops the Qwen3-8B head from 0.77 to 0.65–0.70 and 30 unlabelled requests bring it back to 0.74–0.75, against 0.77 for a full relabelled refit. New labels are needed only for a new question. How the routing works →
| Level | Needs | Does | Does not |
|---|---|---|---|
raw |
nothing | restricted softmax over label tokens | anything about bias or calibration |
L0 |
nothing | averages position bias out over the K rotations, divides out the label prior | calibrate the uncertainty |
L1 |
100–500 labels per question | temperature scaling on top of L0 | change the ranking |
L2 |
100–300 labels per question | a closed-form head on the hidden state partway down, one prompt per state | transfer to another question or model |
Every Decision carries its level, and require="L1" makes downstream code refuse to act on a
weaker one. L0 costs K prefills for a K-option choice, or about 7 of 18 with the rotation budget on (adaptive_shifts=True, recommended — see Serve it). L2 costs less than one plain forward.
d.observe(q, state, label) collects labels as they arrive and solves the head by itself at 30,
re-solving at 60, 120, … so day 0 runs at L0 with nothing and L2 arrives when the loop has fed it.
The contract in full → · the method →
No GPU handy? python -m demo.jev_mode --backend fake runs the whole thing on a synthetic
model in under a second.
Jev mode · 2048 and Minesweeper · NanoJev maze · when L0 helps · small models · research log, negative results included
Every number is regenerated from committed JSON (bash scripts/regen_docs.sh); a second run
from a clean checkout reproduced every zero-label number bit for bit. Not affiliated with TypeSafe
AI or Jev; rows published by their authors were not rerun here.
-
choice,noulandscorefrom one prefill; L0 with zero labels; L1 artifacts - L2: a closed-form head per question, routing, label-free adaptation,
observe - Shipped heads for five Qwen3 models; a packaged demo
- 🚧 Speed optimization (ongoing): making every decision cheaper
- L2 on served engines: vLLM, through an embed server's pooler or a truncated checkpoint (
anyjev.pipeline,anyjev.truncate); SGLang not yet - Agent-loop evaluation: the same decisions inside a real agent, against the LLM they replace
- Heads on the Hugging Face Hub, an interactive Space, a technical report
- More models (Llama, Gemma, Mistral, DeepSeek), span readout beyond 26 options, conformal abstention
Dated plan and help-wanted files: ROADMAP.md.
- On typed-decisions, "accuracy" is agreement with a teacher LLM. The gold is the mean of three samples of one model; a fresh sample of that teacher agrees with it 0.735 of the time.
- L2 is per question and per model. Heads fit on other questions do not help a new one, and only Qwen3 heads ship. It needs hidden states, which transformers and a vLLM embed server both provide; other engines do not yet.
- Calibration cannot fix a model that cannot answer. On maze edges and Minesweeper no readout beats the trivial baseline.
- L0 is not a free win everywhere. The batch prior costs accuracy when one label dominates (when L0 helps).
Also: at most 26 options in the letter readout (a span readout is on the roadmap, not in the code); coverage at 5% risk is a high-variance estimate at n = 300; the headline tables are Qwen models; every decision here is scored in isolation, not inside an agent loop.
Backends and bench providers are one file each; several are help wanted (ROADMAP.md, CONTRIBUTING.md). Changes: CHANGELOG.md. Credits: CREDITS.md.
@software{anyjev2026,
title = {AnyJev: Turn any LLM into a Jev-style decision model},
author = {Zhang, Jiamu and Yang, Tianze and Shi, Yucheng and Wu, Liang},
year = {2026},
url = {https://github.com/nokia-applied-research/AnyJev}
}Apache-2.0, see LICENSE. Datasets keep their own licenses, see THIRD_PARTY.md.

