← Back
pritornyi

pritornyi/giga-embeddings-10B-A1.8B_hybrid

GGUF port and imatrix-guided IQ4_XS/Q5_K quantization of Giga-Embeddings-instruct-10B-A1.8B, a 10B MoE embedding model for Russian, English and code, running on a single 16 GB GPU with stock llama.cpp. Pipeline, llama.cpp patches, and every measurement.

View on GitHub ↗
Stars
41
Forks
0
Watchers
41
Open issues
0
Contributors
1
Language
Python
License
MIT License
Default branch
main
Created Sep 25, 2026Updated Sep 25, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

giga-embeddings-10B-A1.8B_hybrid

A GGUF port and quantization of ai-sage/Giga-Embeddings-instruct-10B-A1.8B-0826, a 10B mixture-of-experts embedding model for Russian, English and code, built to run on a single 16 GB consumer GPU with stock llama.cpp.

This repository holds everything needed to produce the files and check them: the pipeline, three llama.cpp patches, the measurements, and the write-ups.

BF16 original Q8_0 IQ4_XS/Q5_K hybrid
File size 19.52 GiB 10.38 GiB 6.17 GiB
nDCG@10, 11 retrieval tasks 0.8434 0.8434 0.8440
Seconds per sequence, RTX 5080 16 GB, all layers on GPU 0.798 * 0.053 0.043

* BF16 does not fit into 16 GB of VRAM; the driver spills it into system RAM.

Start here

  • Want to use the model? Read the model card: which file to take, a quick start with stock llama.cpp, and the quality and speed numbers.
  • Want to know how it was done? Read the story: what broke during the port and the quantization, and how each problem was tracked down.
  • Want to reproduce it? See Reproducing below and docs/reproducibility.md.

What is in the repository

Path Contents
patches/ Three patches for llama.cpp (see below)
scripts/ PowerShell pipeline: bootstrap, build, and one stage per step
src/giga_quant/ Python: reference runs, parity analysis, corpus building, imatrix checks, quantization plan and audit, evaluation, benchmarks
configs/ Frozen protocols and thresholds for every stage, environment locks
tests/ Unit and integration tests
artifacts/ Measured results: reports, manifests, per-stage verdicts
docs/ The story, detailed notes per topic, and architecture decision records

Model weights, calibration data and evaluation data are not stored here. The pipeline downloads them from pinned revisions and verifies their hashes.

The patches

The finished GGUF files run on unmodified llama.cpp. The patches are needed to make them:

Patch What it does
0001-deepseek-v3-bidirectional-converter.patch Converts DeepseekV3BidirectionalModel to GGUF as the stock deepseek2 architecture with causal attention off, mean pooling, and the gpt-4o pre-tokenizer after checking that its regex matches the model's tokenizer. Python only.
0002-giga-taps-layer-dump-tool.patch llama-giga-taps, a tool that dumps hidden states, router decisions and matmul inputs per layer, for comparing llama.cpp with Transformers.
0003-imatrix-encoder-records.patch A records mode for llama-imatrix: one pre-tokenized record per pass with the embedding graph, instead of a concatenated causal-LM stream. Needed for any bidirectional encoder.

They apply to llama.cpp commit 3057bb66c86c46d5781e50e85462a760ba7d1feb.

How the work was checked

The pipeline is a chain of stages, each with a gate that decides on the artifacts the stage produced. Two rules held throughout:

  • Thresholds are fixed before the measurement. A result that misses its threshold stays in the history. The threshold is never lowered; if the comparison itself was wrong, the fix is a new protocol version with the reason written down, and the old record is kept.
  • An exit code proves nothing. Tensor types are read back from the finished file, token IDs are compared with the reference tokenizer, vectors with vectors.

The main checks:

Check Result
BF16 GGUF against the original model in float32, 21 texts worst cosine 0.99923672, threshold 0.999
Tokenization against the Hugging Face tokenizer 21 of 21 texts identical
Tensor types in the hybrid against the plan 413 of 413, no silent fallbacks
Retrieval, hybrid against BF16 +0.00062 nDCG@10, budget −0.01
Stock llama.cpp release b11160, both files, CPU and CUDA all smoke checks pass

The key numbers in the documentation are re-read from their artifacts by the documentation gate, so the text cannot quietly drift from the data.

Reproducing

The pipeline targets native Windows with PowerShell 5.1. It was run on Windows 11 with a Ryzen 7 9800X3D, 31.1 GiB of RAM and an RTX 5080 (16 GB, compute capability 12.0), using MSVC 14.44, CMake 4.4.3, Ninja 1.13.2 and CUDA 13.3. All tools and Python environments are installed inside the repository; nothing is installed system-wide.

.\scripts\preflight.ps1          # measure the host, write the resource budget
.\scripts\bootstrap-tools.ps1    # project-local CMake and Ninja
.\scripts\bootstrap-cuda.ps1     # project-local CUDA toolkit
.\scripts\pin-sources.ps1        # check out llama.cpp at the pinned commit
.\scripts\bootstrap.ps1          # project-local Python environments (uv)
.\scripts\build.ps1 -Backend CPU
.\scripts\build.ps1 -Backend CUDA

.\scripts\pipeline.ps1 -List     # stages and their gate status
.\scripts\pipeline.ps1 -Resume   # run the next stage whose prerequisites pass

.\scripts\pipeline.ps1 -Stage Release checks reproducibility end to end: it makes a clean clone, restores the environment from the lock file, runs the tests, applies the patches to the pinned llama.cpp commit, regenerates the quantization plan and a smoke importance matrix, and compares every hash the earlier stages pinned.

The heavy stages take real time on this hardware: the full importance matrix about 2 hours 40 minutes, the retrieval evaluation about 4 hours 20 minutes for three models, the benchmark about an hour. CI runs the tests and the documentation gate on a hosted Windows runner without a GPU; hardware results come from the checked-in artifacts.

Documentation

Topic Document
Model architecture and how it maps to llama.cpp architecture, converter, runtime
Port fidelity parity protocol, parity findings
Calibration corpus and importance matrix calibration corpus, imatrix
Quantization plan and audit quantization
Retrieval evaluation evaluation
Speed and memory benchmark
What was not tested limitations
Windows-specific notes and failures windows-native, troubleshooting
Pinned upstream sources sources
Why each non-obvious decision was made ADRs

Contribution rules are in CONTRIBUTING.md.

License

The code and documentation in this repository are MIT-licensed; see LICENSE. The llama.cpp patches are MIT, like llama.cpp itself.

The base model is labeled mit on its Hugging Face page, but its repository contains no license text; the GGUF files inherit whatever terms the base model carries. Calibration and evaluation datasets keep their own licenses and are not redistributed here.