A simulator, a calibration method and a best recipe from Phase 1 of the AWS Trainium Frontier challenge.
Quick start · Results · Use it on your own runs · Findings · Limitations · Cite
| 0.9888 → 0.9614 | 1,206 | +0.0067 ± 0.0004 | 0.00057 bpb |
|---|---|---|---|
| official val_bpb, 24 Sep → 30 Sep (CDT) | Trainium run records the simulator is fitted on | rehearsal-to-official offset (mean ± sd, 15 uploads) | measured price of 1% of training steps |
- Chip time is the bottleneck.
ffsimpredicts a recipe's steps, step time and val_bpb in seconds on a laptop (numpy only), and tells you when a knob change is inside the noise so you can skip the run. - Official scores you can call in advance. A cold rehearsal of the exact upload file plus a near-constant offset has called every upload since K50: none missed its written projection by more than 0.0006, and both final uploads landed within 0.0001. You can test on your own chip instead of on the leaderboard.
- Speed has an exchange rate.
delta_bpb ~= 0.063 * ln(compute ratio)turns any throughput gain into score, so you can price a kernel before you write it.
flowchart LR
R["<b>Recipe</b><br/>train.py + FF_* knobs"] --> S{"<b>ffsim</b><br/>simulator"}
S -->|"Route 1: numpy surrogate"| P["Predicted steps, step time,<br/>val_bpb, P(beat best)"]
S -->|"Route 2: GPU proxy"| G["Equal-steps replay<br/>on a CUDA GPU"]
P -->|"only clear winners"| C["<b>Cold chip rehearsal</b><br/>exact upload file, seed 73,<br/>full charged budget"]
G -->|"only clear winners"| C
C -->|"+ offset ~ +0.0067"| O["<b>Official score,</b><br/>predicted before upload"]
official_bpb = quality(recipe, steps, seed) + offset(eval shard)
steps = the batch-schedule walk over the charged-time budget at the recipe's step time
Speed and quality are modelled separately because the challenge scores a fixed wall-clock budget.
| Route 1: the surrogate | Route 2: the GPU proxy | |
|---|---|---|
| What | step-time model per code lineage, a replay of train.py's charged-clock schedule, a ridge-regression quality model and a rehearsal-to-official offset model, run as a Monte Carlo with common random numbers |
a generated CUDA fork of the real train.py driven by a virtual charged clock, so every batch-phase, MTP-stage and cooldown switch fires at the same step as on the chip |
| Needs | Python 3.11 + numpy | an NVIDIA GPU, CUDA PyTorch, the organiser's prepare.py and tokenizer |
| Fitted / checked on | 1,206 run records (928 full runs, 278 screens); quality model on 131 full runs; 99 held-out same-seed pairs | Trainium pairs the campaign had already paid for |
| How close | K60 predicted 2,365 ± 9 steps and 0.9654 official, measured 2,361 and 0.9655; 85% sign agreement on the 99 pairs | one K60 run matched the chip to 4e-6 at equal steps (partly luck: default CUDA runs differ by 0.00056 between boxes); four paired effects within about 0.0002 |
Full details, calibration tables and caveats: docs/SIMULATOR.md. The rehearsal method:
docs/EXACT-ORACLE.md.
| upload | date (CDT) | official val_bpb | what it added |
|---|---|---|---|
| F2 | 15 Sep | 1.1463 | first upload |
| K40 | 21 Sep | 0.9953 | depth 10 x width 1024, Muon shard; first score under 1.0 |
| K44 | 24 Sep | 0.9888 | cooldown 0.60 |
| K50 | 26 Sep | 0.9760 | multi-token prediction, batch warm-up |
| K60 | 29 Sep | 0.9655 | attention-source reuse |
| K70a | 30 Sep | 0.9625 | row-pool data sampling |
| K73s4 | 30 Sep | 0.9620 | optimizer fusion + XU 48 + c_proj LR 0.8, shuffle salt 4 |
| K77a | 30 Sep | 0.9617 | + EMA blend 0.6 |
| K82s4 | 30 Sep | 0.96136 | + EMA prewarm (best scored); rehearsal 0.954472, projected 0.9614-0.9615 |
| K82s7 | 1 Oct | 0.96140 | K82s4 with shuffle salt 7; best rehearsal of the campaign (0.954379), projected about 0.9614 |
Every scored upload, with its rehearsal and offset, is in results/official-scores.csv;
the story behind them is in docs/FINDINGS.md. In the final leaderboard snapshot (about 7:40 AM
CDT on 1 Oct) rank #10 was 0.9554 and rank #1 was 0.9328, so we finished outside the top 10, 0.0060 behind #10.
git clone https://github.com/marker2601/trainium-simulator.git
cd trainium-simulator
pip install -r requirements.txt
python -m pytest -q tests/ # 255 pass, 5 skip (they need the full private chip harvest)Command line. Fit the three models from the shipped dataset (about 2 s), check them, then predict:
python -m ffsim fit --runs research/sim-data/runs.jsonl --pairs research/sim-data/validation-pairs.json \
--uploads research/sim-data/official-uploads.csv --out research/sim-data/models.pkl
python -m ffsim validate
# one recipe (omit --seed to average over seeds)
python -m ffsim simulate --recipe ffsim/examples/recipe-K60.json --n 2000 --seed 73
# every single and double knob change around K60, ranked, with a "confirm next" list
python -m ffsim search --space ffsim/examples/space-k60-local.json --n-sims 1000 --top 20 \
--out out/search-k60-localLocal app. A Gradio UI with three tabs: predict a recipe, a speed-to-score calculator, and an about page.
pip install -r space/requirements.txt
python space/app.py # then open http://127.0.0.1:7860From Python. Every button is an API endpoint (/predict, /predict_overrides, /speed_to_score,
/score_to_speed):
from gradio_client import Client
c = Client("http://127.0.0.1:7860/")
c.predict("K82s4", "FF_COOLDOWN_FRAC=0.6 FF_MATRIX_LR_SCALE=2.3", 73, "C", 2000,
api_name="/predict_overrides") # tweak our best recipe
c.predict(0.9614, 10, api_name="/speed_to_score") # what does +10% throughput buy?
c.predict(0.9614, 0.9554, api_name="/score_to_speed") # how much speed to reach a target?Argument details are in space/README.md. The GPU proxy has its own setup
(docs/SIMULATOR.md).
The shipped models know our recipe family. To make them know yours, give ffsim your own runs and refit.
-
Collect run directories. Lay them out as
<dir>/chipC/<run>/(chip letters A-D), each holding thetrain.log,eval.logandoverridesfiles your runs write. The three reference runs inresearch/sim-data/chipC/show the expected shape.ffsim/harvest.pycan copy them off Trainium instances over AWS SSM (read-only; configure it withFFSIM_AWS_PROFILEandFFSIM_CHIP_C_REGION/FFSIM_CHIP_C_INSTANCE). -
Build a dataset. Each run becomes one
RunRecordJSON line (schema inffsim/schema.py, contract inffsim/CONTRACT.md). You can also write those lines directly from your own logs.python -m ffsim build-dataset --chips-dir my-runs --experiments none --monitor none --out out/my-runs.jsonl
-
Add pairs and uploads. Same-seed treatment/control pairs score the quality model, and your rehearsal-vs-official table fits the offset. Copy the formats of
research/sim-data/validation-pairs.jsonandofficial-uploads.csv. -
Fit, validate, predict. Append your records to the shipped
runs.jsonlor fit on yours alone:python -m ffsim fit --runs out/my-runs.jsonl --pairs my-pairs.json --uploads my-uploads.csv --out out/my-models.pkl python -m ffsim validate --models out/my-models.pkl python -m ffsim simulate --recipe my-recipe.json --vs ffsim/examples/recipe-K60.json --models out/my-models.pkl
With only a handful of your own runs the quality model is mostly its priors (on the three reference runs alone its
sd is about 0.005 bpb), so append your records to the shipped runs.jsonl unless your recipe family is very different.
A recipe is a JSON file with a code version, a chip, a time target and its FF_* knobs (see
ffsim/examples/). Read results with the decision rule in ffsim/README.md:
confirm on a chip only candidates with support ok and a paired gain of 0.0003 or better; everything smaller is a
tie, and anything flagged EXTRAP, NEVER-VARIED or "cost unknown" needs a speed screen first.
Warning
models.pkl is a pickle. Build your own with fit; never load one from someone else.
- Row-pool data sampling was the best late lever: -0.0023 on chip, -0.0022 officially (0.9647 → 0.9625). It shuffles rows within a 256-batch pool, starting after the warm-up. Nearly free: 11 fewer steps (-0.5%) on its chip pair.
- EMA blend 0.6: about -0.0005 on one chip pair; officially K77a 0.9617 vs K73s4 0.9620.
- EMA prewarm moves the EMA kernel compile out of the charged clock. On chip G the EMA recipe ran 2,375 steps without it and 2,405 with it; officially K82s4 0.96136 vs K77a 0.9617, about -0.0003 (one pair).
- Steps were the main currency. A 30-minute run did 912 steps on 18 Sep and about 2,390 on 30 Sep at the same batch size, almost all of it from compiler-friendly code.
- Seeds are a lottery with a known shape. About 40% of seeds take a warm-up "spike" path in steps 3-10 that costs +0.003 to +0.0045 for good. Any change to steps 0-19 re-rolls it.
- What lost: bigger models (depth 10 about +0.004, 2 KV heads +0.0057), a smaller faster one (depth 8 +0.005), a late k=8 batch, split AdamW cooldown, AdamW LR retunes, NoPE, dropping MLPs. Scalar retunes won 0 of 61 knob keys.
- The gap to the top 10 was throughput, not recipe. From 0.9614, rank #10 in the final snapshot (0.9554) needed about +10% effective compute and rank #1 (0.9328) about +57%. We ran about 285k tokens/s, roughly 30% MFU.
More in docs/FINDINGS.md and docs/GAP-ANALYSIS.md.
Reproducing the recipe on Trainium
- Set up the organiser's kit (github.com/aws-neuron/trainium-frontier) on a trn2 instance as its README describes.
- Replace its
train.pywithrecipes/K82s4/train.py. Every setting is a default in the file: run it with no environment overrides, seed 73. - Run the kit's training and public-shard evaluation. A cold run on our chips gave val_bpb 0.954472 at 2,388 steps on the first 2M public tokens. Add about +0.0066 to +0.0070 for the private-shard score; the official score of this file was 0.96136.
recipes/K60/train.py is the 29 Sep recipe (official 0.9655); the GPU proxy is generated from it. Lineage and
checksums are in recipes/README.md.
ffsim/ the simulator: route 1 (numpy surrogate), gpu/ (route 2 proxy), cloud/, examples/
space/ the Gradio app (python space/app.py) and its API
recipes/ K82s4 and K60 train.py, byte-identical to the uploaded files (Apache-2.0)
docs/ FINDINGS, SIMULATOR, EXACT-ORACLE, GAP-ANALYSIS, campaign-time notes
docs/figures/ every figure (SVG, PNG, PDF) and the one script that builds them from repo data
paper/ OUTLINE.md: the plan, claims ledger and open experiments for the technical report
results/ official-scores.csv: every scored upload with rehearsal and offset
research/sim-data/ runs.jsonl, validation-pairs.json, official-uploads.csv, validation reports, searches
research/ experiments.csv (chips A/B table) and the K60 cold rehearsal log
tests/ 260 tests (255 pass; 5 need the private chip harvest) and fixture logs
- New mechanisms are outside what the surrogate can know. On 15 pairs that finished after the fit it scored 53% sign agreement and an MAE of 0.00104: a noise-floor detector, not an oracle. Every late win (row pool, EMA blend, prewarm) was a mechanism it had never seen.
- Differences under about 0.0003 are a coin flip. Read them as ties and confirm winners on a real chip.
- Model size is not modelled. Depth, width, head dim, KV heads and MLP width never varied in the fitted runs, so changing them gets a flat +3% "cost unknown" step-time charge and an unfitted quality prior.
- Two calibration targets are missed, on purpose and in the open. Step time is 0.545% off against a 0.5% target,
and the pair MAE (0.00039) sits at the noise floor above its 0.0003 target, so
ffsim validateprintspasses: False. - It is calibrated to one setting: single-chip trn2 runs of one nanoGPT-style recipe family and this challenge's evaluation shards. The offset is a property of that text, not a universal constant.
- Selected draws partly regress. K82s7 had the best rehearsal of the campaign (0.954379) but scored 0.96140, just behind K82s4 (0.96136); K73s6 gave back its whole local edge. Pick by rehearsal, but expect less.
If ffsim, the exact-oracle method or the findings help your work, please cite the software
(CITATION.cff; GitHub's "Cite this repository" button reads it). A paper describing the method is
in preparation; its outline and claims ledger are in paper/OUTLINE.md.
@software{frontierforge_2026,
author = {{FrontierForge team}},
title = {FrontierForge: a Trainium run simulator and exact-oracle score prediction},
year = {2026},
version = {1.0.0},
url = {https://github.com/marker2601/trainium-simulator},
license = {MIT}
}Thanks to the AWS Trainium Frontier organisers for the challenge, the hardware, the baseline and a leaderboard that kept us honest. Our recipes stand on the open-source nanoGPT and modded-nanogpt speedrun communities and the Muon optimizer community, whose public work made a 30-minute language model worth racing.
If ffsim saved you a chip-hour, a star helps other Trainium teams find it.
Our code, data and documents are under the MIT licence. recipes/*/train.py and ffsim/gpu/train_gpu.py
are modified versions of the organiser's Apache-2.0 baseline and stay under
Apache-2.0; see NOTICE. The organiser's kit, tokenizer and data are not
redistributed here.