← Back
ollaya-dev

ollaya-dev/ollaya

Run open decision models locally: pull and serve Laya, decider, NLI and GLiClass behind a TypeSafe-compatible API. Ollama for decision models.

View on GitHub ↗https://ollaya.dev ↗
calibrationclassificationdecision-modelsgliclassjevlayallm-routinglocal-inferencenliollamaonnxonnxruntimerusttypesafezero-shot-classification
Stars
1.1K
Forks
57
Watchers
1.1K
Open issues
9
Contributors
5
Language
Rust
License
Apache License 2.0
Default branch
main
Created Sep 23, 2026Updated Oct 1, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

ollaya

Run open decision models locally, the way Ollama runs LLMs.

Website · Models · Results · Docs · Releases · Hugging Face

Created and maintained by Mert Cobanov (@mertcobanov).

A decision model reads a state (a message, an email, a ticket, any JSON) plus typed questions (choice, score, noul) and returns calibrated probabilities in a single forward pass, in milliseconds. It never generates text. Ollaya pulls these models by name, serves them from a local daemon, and speaks TypeSafe's /v1/systemone wire format, so existing Jev clients work by changing one environment variable.

curl -fsSL https://ollaya.dev/install.sh | sh
ollaya run winnow:e4b --preset triage "Third time this year you've double-charged me. Refund it today or I'm cancelling and moving to a competitor."
intent            refund                                ███████████████░ 0.91
is_urgent         yes                                   ███████████████░ 0.92
frustration       2.89 / 3  very angry or using stron…  ██████████████░░ 0.86
refund_requested  yes                                   ████████████████ 0.99
churn_risk        yes                                   ████████████████ 0.99

winnow:e4b is the recommended model: 0.722 accuracy on typed decisions (TypeSafe's Jev: 0.738) and 89 ms for these five questions on an RTX 4090. It is a 4B-class language model, so without an NVIDIA GPU start with laya, which answers in a fraction of a second on a CPU. All models and their numbers: ollaya.dev/search.

Features

  • One binary. ollaya serve runs the daemon; ollaya run, pull, list, ps, show, rm, cp, stop and create work the way they do in Ollama. If the daemon isn't running, the CLI starts it.
  • TypeSafe-compatible. POST /v1/systemone, /v1/decisions and GET /v1/models are wire-identical to TypeSafe. The official SDK works unchanged when you set TYPESAFE_BASE_URL=http://localhost:11435.
  • Native API. /api/decide adds routing information and timings. /api/pull streams NDJSON progress, and there are /api/tags, /api/show, /api/ps and more. See docs/api.md.
  • Weights come from their authors. Ollaya publishes only small ONNX graphs, about 3 MB each. These graphs read the original weight files (usually model.safetensors) from the author's Hugging Face repository, pinned to a commit and verified by sha256. Models whose authors publish GGUF files (winnow, jevk5, jeb) run that file itself on llama.cpp. Ollaya never re-hosts weights.
  • For agents. ollaya mcp serves the models to Claude Code, Claude Desktop, Cursor and other MCP clients (claude mcp add ollaya -- ollaya mcp), and the ollaya-decisions skill teaches agents when and how to use them (npx skills add ollaya-dev/ollaya --skill ollaya-decisions).
  • Routers. laya detects the script and language of each request, then answers with laya:en or laya:multilingual.
  • Modelfiles. You can bake a question set into your own model:
    FROM laya
    QUESTIONS ./triage.json
    PARAMETER precision fp32
    
    Then run ollaya create triage -f Modelfile and ollaya run triage "…".
  • Fast and exact.
    • Hardware: ONNX Runtime on CPU, and CUDA on NVIDIA GPUs. GGUF models run on llama.cpp: CPU, CUDA, and Metal on Apple silicon.
    • Precision: fp16 on GPU and fp32 on CPU, chosen when the model loads.
    • Accuracy: fp32 exports give the same decision as the PyTorch reference on 100% of 2,383 questions per checkpoint.

Models

Model What it is
winnow:e4b Recommended. EldanRing's Winnow-E4B, a Gemma 4 fine-tune run from the author's Q8_0 GGUF on llama.cpp: 0.722 on typed-decisions (Jev: 0.738), 89 ms for five questions on an RTX 4090
laya Router: picks laya:en or laya:multilingual by language
laya:en English decision model (ModernBERT-large, 421M). The fastest: 8–10 ms for five questions on an RTX 4090
laya:multilingual 100+ languages (mmBERT-base, 322M)
laya:typed-decisions Fine-tuned on the typed-decisions workflows
decider, decider:4b, decider:0.8b Mapika's Qwen3.5 decoders, 2B (the default), 4B and 0.8B: 0.680 on typed-decisions for 4B, 0.591 for 2B
decider:2b-vision Mapika's Qwen3.5-2B vision-language decider: questions about an image (--image, images on /api/decide) as well as the state
kev, kev:0.8b, kev:9b Jared Palmer's Kev: a LoRA and a pointer head on Qwen3.5 (4B by default, 0.8B, 9B), calibrated. kev:4b scores 0.669 on typed-decisions and kev:9b 0.722, as much as winnow:e4b
decision Decision 1.0 Eos by the vLLM Semantic Router contributors: a fine-tuned Qwen3.5-0.8B with an endpoint head, 16k-token rows
qwen3guard Qwen3Guard-Gen-0.6B safety guard with built-in questions: safe, controversial or unsafe, and the category
nli, nli:modernbert-large Moritz Laurer's zero-shot NLI classifiers (DeBERTa-v3-large, ModernBERT-large)
gliclass Knowledgator's instruction-following zero-shot classifier (DeBERTa-v3-large)
von Victor Hugo Panisa's Von 1.1 (ModernBERT-large): every option scored at its own marker, 8k-token context
winnow EldanRing's Winnow-12B, the larger sibling of winnow:e4b: 0.702 on typed-decisions
clm Contrastive-LM's CLM-v0.1-8B: the Qwen3-8B encoder and two projection heads score options by similarity, with questions and options cached. 0.357 on typed-decisions; built for agent, game and tool-calling states
jevk5 alibiserikbay's JevK5 v0.3, a Qwen3.5-4B fine-tune run from the author's Q8_0 GGUF on llama.cpp, up to 16 options
nimble Bespoke Labs' Nimble v2: a LoRA on Qwen3.5-9B that reads the whole request as a JSON schema and scores option codes, with the author's temperature: 0.665 on typed-decisions, up to 255 options, ~2.3 s for five questions on an RTX 4090
jeb, jeb:4b, jeb:27b Jason Brashear's Jebadiah (AINode): LoRAs merged into Qwen3.5 (9B by default, 4B) and Qwen3.8-27B, run from the authors' GGUF on llama.cpp with their per-type temperatures; jeb:9b answers five questions in 124 ms on an RTX 4090
jeeves PostHog's Jeeves-9B without its reasoning chain: Qwen3.5-9B (LoRA merged) and a pointer head: 0.680 on typed-decisions with an ECE of 0.031, 838 ms for five questions on an RTX 4090
cygnet blockbrain-ai's Cygnet: frozen Gemma 4 12B IT (Q8_0 GGUF) with a letter-readout prompt and temperature 3.4: 0.683 on typed-decisions, 202 ms for five questions on an RTX 4090

Browse them at ollaya.dev/search. Laya tags ending in -fp32 or -fp16 pin the precision. The derived files of every model are also published at huggingface.co/ollaya-dev.

Ollama 0.35 also serves decision models: Nimble and Tev1, through the same TypeSafe wire format. The FAQ compares the two projects.

Results

We measure every model on our own GPUs and CPUs and publish all of it, with the raw data, at ollaya.dev/results: accuracy and calibration on public benchmarks, speed on every machine, and parity with the authors' own code on each device.

Accuracy against latency on Bespoke Labs' public benchmark, RTX 5090: Ollaya's models and Ollama's

  • Accuracy. On Bespoke Labs' public benchmark (3,880 human-labeled questions from 13 datasets, scored with Bespoke's own code, one RTX 5090), winnow:12b scores 0.773 at 60 ms per question. Ollama's best, Nimble, scores 0.749 at 210 ms.
  • Calibration. On the same Nimble weights, the calibration error is 0.022 on Ollaya and 0.122 on Ollama: Ollaya applies each model's fitted temperature.
  • Speed. Every model on an RTX 5090, an RTX 4090 and two CPUs (Threadripper 3960X, i9-13900K), five questions per request through the HTTP API. The encoders (laya, nli, gliclass, von, qwen3guard) take 7 to 35 ms on a GPU and 0.3 to 2.4 s on a CPU; the decoders 0.1 to 0.9 s on a GPU and 1.3 to 25 s on a CPU; nimble:9b 1.8 to 2.3 s on a GPU.
  • Parity. Before a model ships, its runtime is checked question by question against the authors' code (or llama.cpp's own server, for GGUF models) on each device it runs on.

Median latency of five-question requests for every model on each GPU and CPU we measured

Install

  • Linux (x86_64 or arm64, glibc ≥ 2.38, e.g. Ubuntu 24.04+): curl -fsSL https://ollaya.dev/install.sh | sh. When an NVIDIA GPU is present (driver R525+), the installer adds the CUDA runtime: CUDA 13 for R580+, CUDA 12 for older drivers.
  • macOS (Apple silicon): the same command.
  • Windows (x64): irm https://ollaya.dev/install.ps1 | iex in PowerShell. When an NVIDIA GPU is present (driver R527+), the installer adds the CUDA runtime, as on Linux.
  • Desktop app for macOS, Windows and Linux: start and stop the server, download models and run them in one window. On macOS it lives in the menu bar. Get it from ollaya.dev/download.
  • Docker: docker run -d --gpus=all -p 11435:11435 ghcr.io/ollaya-dev/ollaya:cuda (:cuda12 for host drivers older than R580), or ghcr.io/ollaya-dev/ollaya for CPU only.

Configuration is through environment variables: OLLAYA_HOST, OLLAYA_MODELS, OLLAYA_KEEP_ALIVE, OLLAYA_DEVICE, OLLAYA_API_KEY and others, listed in docs/api.md §15.

Repository

Path What
crates/ollaya The binary: CLI, daemon, runner
crates/ollaya-server HTTP API, scheduler (one runner process per model), model resolution
crates/ollaya-api API types and client; the contract is docs/api.md
crates/ollaya-registry Model names, manifests, blob store, resumable pulls
crates/ollaya-decision Question schema, sequence layouts, calibration, answers
crates/ollaya-runner Inference engines (ONNX Runtime, and llama.cpp for GGUF models)
crates/ollaya-lang Script and language detection for routers
convert/ Build-time Python: ONNX export, parity checks, packaging
site/ The website and the static model registry host

Development

cargo test --workspace
cargo build --release -p ollaya --features cuda   # CUDA build (x86-64 Linux and Windows)

convert/ rebuilds models. It exports them, checks parity against the PyTorch reference, generates golden fixtures, and packages the result into registry/. See the module docstrings. cd convert && uv sync installs it: with CUDA 13 torch on Linux and Windows, and with the CPU and MPS build from PyPI on Apple silicon Macs, where exports and parity run on the CPU.

License

Apache-2.0. Each model keeps its own license: laya (Convai Innovations), decider (Mapika), kev (Jared Palmer, on Qwen3.5 by the Qwen team), decision (the vLLM Semantic Router contributors, on Qwen3.5), qwen3guard (Qwen team), gliclass (Knowledgator), von (Victor Hugo Panisa), winnow (EldanRing, on Gemma 4 by Google DeepMind), jevk5 (alibiserikbay, on Qwen3.5), nimble (Bespoke Labs, on Qwen3.5), jeeves (PostHog, on Qwen3.5), jeb (Jason Brashear, on Qwen3.5 and Qwen3.8), cygnet (Gemma 4 by Google DeepMind; the Cygnet recipe is MIT) and nli:modernbert-large are Apache-2.0, and nli:deberta-v3-large (Moritz Laurer) is MIT. llama.cpp, which Ollaya ships for GGUF models, is MIT.

Ollaya is an independent project. It is not affiliated with or endorsed by Ollama or TypeSafe.