A GGUF port and quantization of
ai-sage/Giga-Embeddings-instruct-10B-A1.8B-0826,
a 10B mixture-of-experts embedding model for Russian, English and code, built
to run on a single 16 GB consumer GPU with stock llama.cpp.
This repository holds everything needed to produce the files and check them: the pipeline, three llama.cpp patches, the measurements, and the write-ups.
| BF16 original | Q8_0 | IQ4_XS/Q5_K hybrid | |
|---|---|---|---|
| File size | 19.52 GiB | 10.38 GiB | 6.17 GiB |
| nDCG@10, 11 retrieval tasks | 0.8434 | 0.8434 | 0.8440 |
| Seconds per sequence, RTX 5080 16 GB, all layers on GPU | 0.798 * | 0.053 | 0.043 |
* BF16 does not fit into 16 GB of VRAM; the driver spills it into system RAM.
- Want to use the model? Read the model card: which file to take, a quick start with stock llama.cpp, and the quality and speed numbers.
- Want to know how it was done? Read the story: what broke during the port and the quantization, and how each problem was tracked down.
- Want to reproduce it? See Reproducing below and docs/reproducibility.md.
| Path | Contents |
|---|---|
patches/ |
Three patches for llama.cpp (see below) |
scripts/ |
PowerShell pipeline: bootstrap, build, and one stage per step |
src/giga_quant/ |
Python: reference runs, parity analysis, corpus building, imatrix checks, quantization plan and audit, evaluation, benchmarks |
configs/ |
Frozen protocols and thresholds for every stage, environment locks |
tests/ |
Unit and integration tests |
artifacts/ |
Measured results: reports, manifests, per-stage verdicts |
docs/ |
The story, detailed notes per topic, and architecture decision records |
Model weights, calibration data and evaluation data are not stored here. The pipeline downloads them from pinned revisions and verifies their hashes.
The finished GGUF files run on unmodified llama.cpp. The patches are needed to make them:
| Patch | What it does |
|---|---|
0001-deepseek-v3-bidirectional-converter.patch |
Converts DeepseekV3BidirectionalModel to GGUF as the stock deepseek2 architecture with causal attention off, mean pooling, and the gpt-4o pre-tokenizer after checking that its regex matches the model's tokenizer. Python only. |
0002-giga-taps-layer-dump-tool.patch |
llama-giga-taps, a tool that dumps hidden states, router decisions and matmul inputs per layer, for comparing llama.cpp with Transformers. |
0003-imatrix-encoder-records.patch |
A records mode for llama-imatrix: one pre-tokenized record per pass with the embedding graph, instead of a concatenated causal-LM stream. Needed for any bidirectional encoder. |
They apply to llama.cpp commit 3057bb66c86c46d5781e50e85462a760ba7d1feb.
The pipeline is a chain of stages, each with a gate that decides on the artifacts the stage produced. Two rules held throughout:
- Thresholds are fixed before the measurement. A result that misses its threshold stays in the history. The threshold is never lowered; if the comparison itself was wrong, the fix is a new protocol version with the reason written down, and the old record is kept.
- An exit code proves nothing. Tensor types are read back from the finished file, token IDs are compared with the reference tokenizer, vectors with vectors.
The main checks:
| Check | Result |
|---|---|
| BF16 GGUF against the original model in float32, 21 texts | worst cosine 0.99923672, threshold 0.999 |
| Tokenization against the Hugging Face tokenizer | 21 of 21 texts identical |
| Tensor types in the hybrid against the plan | 413 of 413, no silent fallbacks |
| Retrieval, hybrid against BF16 | +0.00062 nDCG@10, budget −0.01 |
Stock llama.cpp release b11160, both files, CPU and CUDA |
all smoke checks pass |
The key numbers in the documentation are re-read from their artifacts by the documentation gate, so the text cannot quietly drift from the data.
The pipeline targets native Windows with PowerShell 5.1. It was run on Windows 11 with a Ryzen 7 9800X3D, 31.1 GiB of RAM and an RTX 5080 (16 GB, compute capability 12.0), using MSVC 14.44, CMake 4.4.3, Ninja 1.13.2 and CUDA 13.3. All tools and Python environments are installed inside the repository; nothing is installed system-wide.
.\scripts\preflight.ps1 # measure the host, write the resource budget
.\scripts\bootstrap-tools.ps1 # project-local CMake and Ninja
.\scripts\bootstrap-cuda.ps1 # project-local CUDA toolkit
.\scripts\pin-sources.ps1 # check out llama.cpp at the pinned commit
.\scripts\bootstrap.ps1 # project-local Python environments (uv)
.\scripts\build.ps1 -Backend CPU
.\scripts\build.ps1 -Backend CUDA
.\scripts\pipeline.ps1 -List # stages and their gate status
.\scripts\pipeline.ps1 -Resume # run the next stage whose prerequisites pass.\scripts\pipeline.ps1 -Stage Release checks reproducibility end to end: it
makes a clean clone, restores the environment from the lock file, runs the
tests, applies the patches to the pinned llama.cpp commit, regenerates the
quantization plan and a smoke importance matrix, and compares every hash the
earlier stages pinned.
The heavy stages take real time on this hardware: the full importance matrix about 2 hours 40 minutes, the retrieval evaluation about 4 hours 20 minutes for three models, the benchmark about an hour. CI runs the tests and the documentation gate on a hosted Windows runner without a GPU; hardware results come from the checked-in artifacts.
| Topic | Document |
|---|---|
| Model architecture and how it maps to llama.cpp | architecture, converter, runtime |
| Port fidelity | parity protocol, parity findings |
| Calibration corpus and importance matrix | calibration corpus, imatrix |
| Quantization plan and audit | quantization |
| Retrieval evaluation | evaluation |
| Speed and memory | benchmark |
| What was not tested | limitations |
| Windows-specific notes and failures | windows-native, troubleshooting |
| Pinned upstream sources | sources |
| Why each non-obvious decision was made | ADRs |
Contribution rules are in CONTRIBUTING.md.
The code and documentation in this repository are MIT-licensed; see LICENSE. The llama.cpp patches are MIT, like llama.cpp itself.
The base model is labeled mit on its Hugging Face page, but its repository
contains no license text; the GGUF files inherit whatever terms the base model
carries. Calibration and evaluation datasets keep their own licenses and are
not redistributed here.