← Back
usamahz

usamahz/cpu-performance-engineering

A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section.

View on GitHub ↗ai.usamah.me ↗
armawesome-listbenchmarkscomputer-architecturecpuedge-aiinferencelow-latencyoptimizationperformanceperformance-engineeringsimdsystems-programmingx86
Stars
288
Forks
34
Watchers
288
Open issues
2
Contributors
1
Language
C
License
MIT License
Default branch
main
Created Sep 17, 2026Updated Sep 27, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

CPU Performance Engineering

Links Quality Entries Benchmarks License Stars

Making a program fast on a modern CPU means knowing what the core does with each instruction, where the time actually goes, and how to prove a change helped. This is the reading that gets you there, in the order that makes the next piece legible.

Scope. x86 and Arm server parts, from one instruction through to serving a model on CPU. Not language runtimes, database internals, or anything above the socket.

Evidence. Primary sources only: the paper, the specification, the vendor manual, the repository, or a report by the person who did the work. Any number, anywhere in this repository, carries all seven fields set out in What earns a place, or it is not quoted.

Proof. Fourteen of the sections end in a benchmark under misc/benchmarks/: C source, the build line, the machine, the raw numbers and the analysis, all committed. Run them yourself.

Section 1 is a path through the rest; read it top to bottom before using the numbered sections as a reference.

Contents

  • 1. Start here
  • 2. One instruction, end to end
    • Fetch and decode
    • Rename and issue
    • Execute
    • Memory access and retire
  • 3. Microarchitecture
    • Limits of ILP and SMT
    • Branch prediction and speculation
    • Vendor estimates and measured tables
    • What the manuals leave out
  • 4. Memory hierarchy
    • Cache geometry, replacement and misses in flight
    • TLBs, page walks and prefetchers
    • Store buffers, ordering and cache-line contention
    • Struct layout, software prefetch and page size
  • 5. Measurement
    • Method and the whole-system view
    • Counters, events and precise sampling
    • CPU profilers and flame graphs
    • Microbenchmarks that lie
  • 6. Models
    • Roofline and the execution-cache-memory model
    • Top-down analysis
    • Scaling laws
    • Queueing
  • 7. Single-thread optimisation
    • Data layout and loop transforms
    • SIMD instruction sets
    • SIMD libraries and measured kernels
    • Branchless code and bit manipulation
  • 8. Compilers and codegen
    • Reading emitted code
    • Optimisation levels, inlining and link time
    • Target flags and auto-vectorisation
    • Profile-guided and post-link optimisation
  • 9. Concurrency
    • Memory models and atomics
    • Locks, contention and allocators
    • Lock-free structures and RCU
    • Thread pools and work stealing
  • 10. NUMA and multi-socket
    • NUMA and Linux memory placement
    • Topology and interconnects
    • Migration, balancing and measured effects
  • 11. OS and I/O
    • Syscalls and asynchronous I/O
    • Scheduling, affinity and isolation
    • Interrupts and kernel bypass
    • Cache and bandwidth partitioning
  • 12. Tail latency and production systems
    • Measuring the tail
    • Where jitter comes from
    • Load generation and production workloads
    • Mechanical sympathy
  • 13. Inference on CPU
    • GEMM and BLAS
    • Runtimes
    • Quantization
    • Matrix extensions
    • Threading for inference
    • When CPU beats GPU
  • 14. Hardware generations
    • Intel Xeon
    • AMD EPYC
    • Arm Neoverse server parts
    • Independent measurement across vendors
  • 15. Benchmarks
    • Standard suites
    • Microbenchmark suites
    • Methodology and what suites miss
  • 16. Watchlist
    • ISA extensions without a shipped server part
    • Parts without a public measurement
    • Memory and interconnect
    • Kernel paths and generated code
  • What earns a place

1. Start here

Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.

  1. Computer Architecture: A Quantitative Approach, 7th Edition - Its pipelining appendix and memory chapters define the hazard, speculation and cache vocabulary the list assumes.
  2. Optimizing software in C++ - Maps C++ onto pipeline mechanisms and shows why a loop-carried dependency chain, not instruction count, paces a loop.
  3. Intel Optimization Reference Manual - Its opening chapters show how a shipping x86 core implements the textbook pipeline, each rule tied to a mechanism.
  4. What Every Programmer Should Know About Memory - Measures the step in cost per access at each cache boundary and the gap a prefetcher hides.
  5. Memory Barriers: a Hardware View for Software Hackers - Explains why a second core makes loads and stores reorder and what a barrier drains.
  6. Systems Performance: Enterprise and the Cloud, 2nd Edition - Puts the method before the tools: what to measure, in what order, and how benchmarks mislead.
  7. Roofline: An Insightful Visual Performance Model for Multicore Architectures - Places a loop from a byte count and a datasheet bandwidth alone, before any counter is read.
  8. A Top-Down Method for Performance Analysis and Counters Architecture - Defines the split of pipeline slots into front end, bad speculation, back end and retiring, the tree profilers report.
  9. Performance Analysis and Tuning on Modern CPUs - Walks from a noisy timing to counters to a named bottleneck, applying roofline and top-down to whole programs.
  10. What Has My Compiler Done for Me Lately? Unbolting the Compiler's Lid - Shows how to read emitted assembly against its source, so each mechanism is checked in a listing, not assumed.

Work the exercises in Performance Ninja alongside them; reading alone will not build the instinct.

Reproduce it: misc/benchmarks/04-cache-latency, the cost per dependent load stepping up at each cache boundary, the curve the fourth entry measures.

2. One instruction, end to end

Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.

Fetch and decode

  • Fetch Directed Instruction Prefetching - The origin of the decoupled front end, where the predictor runs ahead of fetch and drives instruction prefetch.
  • Micro-Operation Cache: A Power Aware Frontend for Variable Instruction Length ISA - Introduces the decoded micro-op cache and its power case, the structure both x86 vendors later built.
  • Software Optimization Guide for the AMD Zen5 Microarchitecture - Where AMD states when the op cache feeds micro-ops, and the fusion and alignment rules for hot loops.
  • The microarchitecture of Intel, AMD, and VIA CPUs - Measures rather than quotes each x86 core's misprediction penalty, micro-op cache and loop buffer behaviour.
  • Intel Mitigations for Jump Conditional Code Erratum - States what a microcode fix evicts from the decoded cache and which counters show the fall back to legacy decode.

Reproduce it: misc/benchmarks/02-branch-misprediction, the cost of a mispredicted branch, sorted against unsorted against branchless.

Rename and issue

  • An Efficient Algorithm for Exploiting Multiple Arithmetic Units - The origin of tag-based renaming and reservation stations, from which every out-of-order issue queue descends.
  • Optimizing subroutines in assembly language - Shows which chains renaming cannot break, from partial registers to flags, and the zero-cost idioms that do.
  • Measuring Reorder Buffer Capacity - Sets the user-space method that measures the window and register files, and which idioms take no physical register.

Execute

  • Focusing Processor Policies via Critical-Path Prediction - Defines the dependence graph of an out-of-order core and the critical path through it that sets run time.
  • Instruction tables: latencies, throughputs and micro-operation breakdowns - Measures latency, reciprocal throughput and port assignment per instruction, the weights a dependence graph needs.
  • Arm Neoverse V2 Core Software Optimization Guide - States every stage of one Arm core from fetch to issue, then per-instruction latency, throughput and pipe assignment.
  • Entropy Decoding in Oodle Data: x86-64 3-Stream Huffman Decoders - Predicts a real decoder loop's cycle count from its carried chain and issue slots, then measures the prediction.
  • Faster zlib/DEFLATE decompression on the Apple M1 (and x86) - Predicts a decoder's refill chain from Arm core latencies, fits extra work under it, and measures the gain.

Memory access and retire

  • Memory Dependence Prediction using Store Sets - The origin of memory dependence prediction, which lets a load pass older stores and flushes on a wrong guess.
  • Store-to-Load Forwarding and Memory Disambiguation in x86 Processors - Measures which store and load size and offset pairs forward or stall, and which cores predict memory dependences.
  • Microarchitecture Optimizations for Exploiting Memory-Level Parallelism - Defines memory-level parallelism as the misses one window overlaps, and shows which core limits cap it.
  • Implementing Precise Interrupts in Pipelined Processors - Introduces the reorder buffer and defines a precise exception as in-order commit of out-of-order results.
  • Where Do Interrupts Happen? - Measures that an interrupt lands on the oldest unretired instruction, which decides what a sampling profiler blames.

3. Microarchitecture

Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.

Limits of ILP and SMT

  • The MIPS R10000 Superscalar Microprocessor - Shows rename, the active list and precise branch recovery fitted together in one shipped out-of-order core.
  • Limits of Instruction-Level Parallelism (WRL Research Report 93/6) - Separates perfect from realistic prediction, renaming and aliasing, and shows how little ILP survives the real ones.
  • Complexity-Effective Superscalar Processors - Puts circuit delay on wakeup, select and bypass, capping issue width and window size before ILP runs out.
  • Simultaneous Multithreading: Maximizing On-Chip Parallelism - The origin of SMT, where another thread fills one thread's idle issue slots at a cost to both.
  • A Mechanistic Performance Model for Superscalar Out-of-Order Processors - Turns window size, issue width and miss events into cycles lost, making an ILP limit a cycle count.

Branch prediction and speculation

  • A Case for (Partially) TAgged GEometric History Length Branch Prediction - Defines the tagged geometric-history predictor shipped designs converge on, and the limit of what it can learn.
  • A 64-Kbytes ITTAGE indirect branch predictor - Carries the tagged geometric scheme to indirect jumps and calls, where interpreters and virtual dispatch stall.
  • Characterizing the Branch Misprediction Penalty - Defines the penalty as pipeline refill plus window drain, so it exceeds pipeline depth and is not constant.
  • Spectre Attacks: Exploiting Speculative Execution - Establishes that predictor state is shared and trainable across contexts, the root of every mitigation and its cost.
  • Speculative Execution Side Channel Mitigations - Defines the indirect-branch and store-bypass controls (IBRS, STIBP, IBPB, SSBD) and states which carry a large cost.

Vendor estimates and measured tables

  • Software Optimization Guide for the AMD Zen5 Microarchitecture - Calls its own latency spreadsheet an estimate and lists its assumptions, the vendor claim the measured tables check.
  • The microarchitecture of Intel, AMD, and VIA CPUs - States where measured predictor, op cache and port findings disagree with the vendor's account of an x86 core.
  • Instruction tables: latencies, throughputs and micro-operation breakdowns - States why its measured figures differ from the vendor's and names which latencies cannot be measured accurately.
  • uops.info - Where each latency, throughput and port entry links to its microbenchmark, so any value can be re-run.
  • applecpu: Firestorm Overview - Measured per-instruction tables for an Apple AArch64 core, each entry linked to the counter experiment behind it.

What the manuals leave out

  • Performance Speed Limits - Sets the method for finding which hard bound binds a loop, testing its cycles against each in turn.
  • uiCA: Accurate Throughput Prediction of Basic Blocks on Recent Intel Microarchitectures - Models the predecoder, decoders, micro-op cache and loop buffer closely enough to predict front-end bound loops.
  • Gathering Intel on Intel AVX-512 Transitions - Measures what the vendor leaves untimed in an AVX-512 licence change, a throttled phase then a frequency step.
  • Hardware Store Elimination - Finds an undocumented optimisation by its counter signature and bandwidth, and sets the method for what manuals omit.

Reproduce it: misc/benchmarks/03-latency-vs-throughput, one dependency chain against eight independent accumulators.

4. Memory hierarchy

Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.

Cache geometry, replacement and misses in flight

  • What Every Programmer Should Know About Memory - One measured account of DRAM timing, cache geometry, TLBs and prefetchers that sets the padding and prefetch rules.
  • Measuring Cache and TLB Performance and Their Effect on Benchmark Run Times - The origin of the strided-loop method that recovers cache and TLB size, line size, associativity and miss cost.
  • Achieving Non-Inclusive Cache Performance with Inclusive Caches - Names inclusion victims as the cost of an inclusive last-level cache, the case for non-inclusive and victim designs.
  • Adaptive Insertion Policies for High Performance Caching - The origin of set duelling and bimodal insertion, the adaptive replacement that survives a streaming pass.
  • Lockup-Free Instruction Fetch/Prefetch Cache Organization - Origin of the lockup-free cache, whose miss-status registers keep misses in flight, so a stream beats a chase.

Reproduce it: misc/benchmarks/04-cache-latency, dependent-load latency from L1 to DRAM, with and without TLB pressure.

TLBs, page walks and prefetchers

  • Intel SDM Volume 3A: System Programming Guide, Part 1 - Fixes the paging walk, the page-walk caches, the TLB invalidation rules and what each memory type permits.
  • Translation Caching: Skip, Don't Walk (the Page Table) - Defines the page-walk cache design space and shows why cached partial translations let a walk skip page-table levels.
  • Improving Direct-Mapped Cache Performance - The origin of the stream buffer that every vendor stream prefetcher descends from, and of the victim cache.
  • Intel Optimization Reference Manual Volume 1 - Names each prefetcher and what trains it, and which stop at a page boundary and which cross it.
  • Arm Neoverse V2 Core Technical Reference Manual - Names an Arm server core's load-side and store-side prefetchers, its TLB levels and the bits that disable them.

Store buffers, ordering and cache-line contention

  • Memory Barriers: a Hardware View for Software Hackers - Derives store buffers and invalidate queues from the cost of coherence, the reason reordering exists at all.
  • x86-TSO: A Rigorous and Usable Programmer's Model for x86 Multiprocessors - The store-buffer model of Intel and AMD ordering, tested on hardware, fixing which reorderings a fence pays for.
  • Simplifying ARM Concurrency: Multicopy-Atomic Axiomatic and Operational Models for ARMv8 - The formal model Arm adopted into its architecture, and the record of why the non-multicopy-atomic option was dropped.
  • Everything You Always Wanted to Know About Synchronization but Were Afraid to Ask - Measures line-transfer cost between cores by coherence state and distance, the figure behind every contention rule.
  • C2C - False Sharing Detection in Linux Perf - The implementers' account of perf c2c, whose worked example reads contended lines, offsets and callers off the report.

Struct layout, software prefetch and page size

  • Cache-Conscious Structure Definition - Introduces and measures structure splitting and field reordering on real programs, the origin of hot-cold layout rules.
  • CppCon 2014: Data-Oriented Design and C++ - Argues from cache-line utilisation arithmetic that layout must follow the access pattern, the case behind structure of arrays.
  • dwarves (pahole) - Prints a struct's holes, padding and cache-line boundaries from DWARF, so a layout is seen rather than guessed.
  • When Prefetching Works, When It Doesn't, and Why - Sorts software prefetch into the cases where it helps and hurts, and shows how it mistrains hardware prefetchers.
  • Transparent Hugepage Support - The kernel's statement of THP: the enabled and defrag knobs, khugepaged, and the counters showing what was obtained.
  • Coordinated and Efficient Huge Page Management with Ingens - Measures the fault latency, memory bloat and unfairness that eager THP promotion causes, and shows what removes them.

5. Measurement

Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.

Method and the whole-system view

  • The USE Method - Sets the checklist that finds the saturated resource before any profiler is opened.
  • Performance Analysis and Tuning on Modern CPUs - Draws the line between counting, sampling, instrumentation and tracing, so a question is matched to its tool.
  • BPF Performance Tools - The reference for time a CPU sampler cannot see, off-CPU, scheduler and I/O waits, traced at bounded cost.
  • bpftrace - Makes a tracing hypothesis a one-line experiment over kprobes, uprobes, tracepoints and PMU events.
  • Google-Wide Profiling: A Continuous Profiling Infrastructure for Data Centers - The design continuous profilers descend from, always-on sampling across a fleet, cheap enough to leave running.

Counters, events and precise sampling

  • perf_event_open(2) - Defines the counting and sampling modes every Linux profiler uses, and the sample record fields, branch stack included.
  • Instruction-Based Sampling: A New Performance Analysis Technique - Defines skid, why a sample lands after the culprit, and how tagging one op through the pipeline removes it.
  • Intel Software Developer Manuals - Defines the architectural counters, what a PEBS sample captures, and why rdtsc counts time, not cycles, under DVFS.
  • Processor Programming Reference for AMD Family 1Ah Model 02h - Defines the event encodings and IBS registers for one Zen core, the tables perf's AMD events are derived from.
  • perf-arm-spe(1) - Defines Arm SPE as perf drives it, one sampled op in flight, the filters, and what a record holds.

CPU profilers and flame graphs

  • Linux perf wiki: Tutorial - The maintainers' walk from perf stat to perf record to perf annotate, with the sample fields each flag sets.
  • Intel VTune Profiler Documentation - Home of the user guide and cookbook, where each hardware analysis is defined by the events behind it.
  • AMD uProf User Guide - The vendor's reference for IBS-driven profiling on Zen, with metric presets defined per core generation.
  • The Flame Graph - Records the design decisions, width as sample share and alphabetical rather than time order, behind its reading rules.
  • Coz: Finding Code that Counts with Causal Profiling - Proves a hot function need not be worth optimising, by measuring what speeding up a line does to end-to-end time.

Microbenchmarks that lie

  • Producing Wrong Data Without Doing Anything Obviously Wrong! - The origin of measurement bias as a term, link order and environment size alone flipping a compiler flag comparison.
  • Non-Determinism and Overcount on Modern Hardware Performance Counter Implementations - Traces run-to-run variation in x86 retired-instruction counts to one extra count per interrupt and per fault.
  • clock_gettime(2) - Defines what each clock counts, NTP-slewed or raw monotonic time, or CPU time, and the resolution call.
  • Benchmarking tips (LLVM) - The compiler project's recipe for a quiet Linux host, governor, boost, SMT siblings, ASLR, a cpuset and tmpfs.
  • Google Benchmark User Guide - Documents the barriers that keep the optimiser from deleting the work under test, and repetition statistics.

Reproduce it: misc/benchmarks/05-measurement-pitfalls, dead-code elimination, run-to-run spread, cold against warm.

6. Models

A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code.

Roofline and the execution-cache-memory model

  • Roofline: An Insightful Visual Performance Model for Multicore Architectures - Defines operational intensity as traffic past the caches, and the ceilings that say which optimisation can pay.
  • Applying the Roofline Model - States the counter set and method that put a measured point on the plot in place of a hand-counted intensity.
  • LIKWID - Measures the roofs with likwid-bench and the point from likwid-perfctr groups on Intel, AMD and Arm server cores.
  • Cache-aware Roofline model: Upgrading the loft - Adds a roof per cache level against traffic at the core, so a kernel that hits in cache is not plotted as DRAM bound.
  • Performance bottlenecks of stencil computations using the Execution-Cache-Memory model - Times each level's transfer with overlap rules, predicting a core's rate and the core count where bandwidth saturates.

Reproduce it: misc/benchmarks/06-roofline, measured roofs and three kernels of rising arithmetic intensity.

Top-down analysis

  • A Top-Down Method for Performance Analysis and Counters Architecture - Defines the slot accounting that turns raw counters into a weighted tree of bottlenecks, and why the unit is a slot.
  • TMA_Metrics-full.xlsx (intel/perfmon) - The official home of every TMA formula, event, threshold and level per microarchitecture, from which toplev derives.
  • pmu-tools - Runs the TMA tree on Linux from the spreadsheet and states why multiplexed levels mislead on short or varied workloads.
  • pipeline.json (perf pmu-events, amdzen4) - States the AMD top-down formulas in events for both levels, the metric groups perf stat runs on Zen with no vendor tool.
  • Arm CPU Telemetry Solution Topdown Methodology Specification - Defines Arm's top-down as staged stall accounting, first locating the bottleneck and then measuring its resource.

Scaling laws

  • Validity of the single processor approach to achieving large scale computing capabilities - States the serial-fraction bound on speedup, the argument every later scaling law is written against.
  • Reevaluating Amdahl's Law - Defines scaled speedup, the bound that holds when the problem grows with the processor count instead of staying fixed.
  • Amdahl's Law in the Multicore Era - Extends the bound to chips of unequal cores under a fixed area budget, the arithmetic behind big and little cores.
  • A Simple Capacity Model of Massively Parallel Transaction Systems - The origin of the coherency term that makes throughput fall, not merely flatten, as processors are added.
  • Guerrilla Capacity Planning - Derives the universal scalability law and states the procedure that fits its coefficients to measured throughput.

Queueing

  • A Proof for the Queuing Formula: L = λW - Proves occupancy equals arrival rate times time in system with no assumption on arrivals or service.
  • Quantitative System Performance - Origin of the bound-and-bottleneck analysis roofline names as its ancestor, and of the asymptotic bounds on throughput.
  • Performance Modeling and Design of Computer Systems - Proves why open and closed systems answer a load question differently, and when scheduling, not capacity, sets latency.
  • Stochastic Processes Occurring in the Theory of Queues - Origin of the A/S/c notation, and of the embedded chain that solves a queue with non-memoryless arrivals or service.
  • The single server queue in heavy traffic - Derives the wait near saturation from utilisation and arrival and service variance, the formula behind the latency knee.

7. Single-thread optimisation

A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted.

Data layout and loop transforms

  • Intel Optimization Reference Manual - Where Intel states when a structure of arrays beats an array of structures, strided and hybrid cases included.
  • ispc: A SPMD Compiler for High-Performance CPU Programming - Measures one kernel in both layouts and traces the gain to the gathers the array-of-structures form forces.
  • A Data Locality Optimizing Algorithm - Defines interchange, skewing, reversal and tiling as one family and proves when each keeps a loop nest legal.
  • The Cache Performance and Optimizations of Blocked Algorithms - Traces the drops in a blocked loop's speed curve to self-interference misses, and shows when copying a tile pays.
  • Auto-Vectorization in LLVM - Where LLVM lists what its loop and SLP vectorisers accept: interleaving, reductions, if-conversion and runtime checks.

Reproduce it: misc/benchmarks/07-aos-vs-soa-simd, array of structs against structure of arrays, scalar against NEON.

SIMD instruction sets

  • Intel Intrinsics Guide - Maps each intrinsic to its instruction and CPUID flag, with the vendor's latency and throughput per microarchitecture.
  • Intel Software Developer Manuals - The normative semantics of every Intel vector instruction, with the masking, rounding and fault rules intrinsics hide.
  • Introduction to SVE - Where Arm explains vector-length-agnostic loops and predication, the model that removes remainder loops entirely.
  • Arm C Language Extensions - The specification the NEON and SVE intrinsics come from, so it settles what a compiler must accept and what is a bug.
  • Arm Architecture Reference Manual for A-profile architecture - The normative definition of NEON, SVE and SVE2 instructions and of the scalable vector and predicate register model.

SIMD libraries and measured kernels

  • Highway - Defines sizeless vector types with run-time dispatch, so one source serves SVE and every fixed-width ISA.
  • xsimd - Fixes a batch type per ISA and width, the compile-time vector length model, and still reaches NEON and SVE.
  • Faster Base64 Encoding and Decoding Using AVX2 Instructions - Where shuffle-based lookup replacing a byte loop is worked through step by step, with the measurement method spelt out.
  • Parsing Gigabytes of JSON per Second - Shows branch-free structural indexing with carry-less multiply and shuffles, measured against conventional parsers.
  • simdjson - Where that technique ships, with a kernel per ISA and the harness that keeps its published comparisons reproducible.
  • Hyperscan: A Fast Multi-pattern Regex Matcher for Modern CPUs - Decomposes regexes into string and automaton pieces so both run on SIMD, the design inside the matcher Snort embeds.

Branchless code and bit manipulation

  • Branch Prediction and the Performance of Interpreters - Counter evidence that current predictors absorb interpreter dispatch, so a branch removal must be measured, not assumed.
  • Hacker's Delight, 2nd Edition - Derives the branch-free integer tricks, division by a constant among them, with proofs rather than as a catalogue.
  • Faster sorted array unions by reducing branches - A worked branchless merge with code whose gain vanishes when the compiler emits no conditional move.
  • Array Layouts for Comparison-Based Searching - Measures branch-free search over sorted, Eytzinger and B-tree layouts, the winner changing with size and prefetch.
  • Faster Population Counts Using AVX2 Instructions - Shows a carry-save adder tree in vector registers beating the dedicated instruction, timed as the minimum of many runs.

8. Compilers and codegen

No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.

Reading emitted code

  • Compiler Explorer - Shows how a source change alters the emitted instructions across compilers, versions and flags, with nothing installed.
  • What Every C Programmer Should Know About Undefined Behavior - Explains how the signed-overflow and aliasing rules let a trip count be known and a store loop become memset.
  • llvm-objdump - Reads the binary that shipped, with source lines and symbolised branch targets, rather than a recompiled snippet.
  • llvm-mca - Predicts loop throughput and port pressure from the scheduling model, and states it models neither front end nor caches.
  • llvm-exegesis - Measures instruction latency and throughput with counters, so the model llvm-mca predicts from is checked, not trusted.

Optimisation levels, inlining and link time

  • Options That Control Optimization (GCC) - Lists what each -O level turns on, the inlining limits, and that -Ofast admits transforms invalid for conforming code.
  • There Are No Zero-cost Abstractions (CppCon 2019) - Shows with real codegen that an abstraction is free only when inlining and the ABI allow it, and the cost when either refuses.
  • Itanium C++ ABI - Fixes the rule that a non-trivial class goes by reference to a caller-made temporary, the cost a wrapped pointer pays.
  • How To Write Shared Libraries - States what PLT calls and interposition cost, and the visibility controls a library needs to inline its own exports.
  • LTO Overview (GCC Internals) - Defines whole-program LTO against partitioned WHOPR, and the LGEN, WPA and LTRANS stages that run -flto in parallel.
  • ThinLTO - Defines the thin link, summaries analysed whole-program then parallel backends, and the cache for incremental rebuilds.

Target flags and auto-vectorisation

  • x86 Options (GCC) - Defines -march against -mtune, the psABI levels, and -mprefer-vector-width, the switch for full-width AVX-512 code.
  • Function Multiversioning (GCC) - Defines target_clones, one function per ISA behind a resolver the dynamic linker runs, so a generic build ships AVX-512.
  • Controlling Floating Point Behavior (Clang) - Lists what -ffast-math implies, of which -fassociative-math alone frees a float reduction, and -ffp-contract for FMA.
  • Auto-Vectorization in LLVM - States what the vectorisers need, aliasing disproved or checked at run time, and where a float reduction stays in order.
  • Options to Emit Optimization Reports (Clang) - Defines the remarks that make the compiler say which loop it left scalar and why, so the fix targets the real blocker.

Reproduce it: misc/benchmarks/08-autovectorization-aliasing, the vectoriser with and without restrict.

Profile-guided and post-link optimisation

  • Profile Guided Optimization (Clang) - Defines the instrumented and sampled workflows, why their profiles cannot mix, and the cost of a wrong training input.
  • AutoFDO: Automatic Feedback-Directed Optimization for Warehouse-Scale Applications - Defines the address-to-source mapping with discriminators that lets a stale production profile still drive FDO.
  • BOLT: A Practical Binary Optimizer for Data Centers and Beyond - States why a profile applied to the final binary beats one mapped to source, and the layout passes accuracy enables.
  • BOLT (llvm-project/bolt) - States what full effect needs, relocations kept at link time and a branch-stack sample profile, neither on by default.
  • RFC: Propeller: A frame work for Post Link Optimizations - States the design, not a result: basic-block sections and a relink, no binary rewrite, so layout needs no disassembly.

9. Concurrency

Every cost below is a cache line moving between cores, so the ordering models and the measured line-transfer cost in the memory hierarchy section come first.

Memory models and atomics

  • Foundations of the C++ Concurrency Memory Model - Defines the data-race-free contract, sequential consistency for race-free programs and no meaning for a race.
  • Atomic operations, C++ working draft - The normative wording for every memory order, fence and read-modify-write, the text a compiler is checked against.
  • C/C++11 mappings to processors - The table that turns each memory order into x86 and Arm instructions, so what an order costs is read off the page.
  • Linux kernel memory-barriers.txt - States what the kernel assumes any CPU may reorder and what each barrier and access primitive guarantees.
  • herdtools7 - Where herd7, litmus7 and klitmus7 live, the tools that run a litmus test against the x86, Arm and kernel models.

Locks, contention and allocators

  • Is Parallel Programming Hard, And, If So, What Can You Do About It? - Derives counting, partitioning, locking and deferral with code that runs, the textbook the section assumes.
  • Algorithms for Scalable Synchronization on Shared-Memory Multiprocessors - The origin of the queue lock, each waiter spinning on its own line, measured against ticket and test-and-set locks.
  • Futexes Are Tricky - Derives a correct user-space mutex from futex and shows the lost wakeups and extra kernel entries naive versions pay.
  • Hoard: A Scalable Memory Allocator for Multithreaded Applications - Defines blowup and allocator-induced false sharing, which per-processor heaps under a bounded global heap avoid.
  • TCMalloc: Thread-Caching Malloc - The design statement for per-CPU caches built on restartable sequences and a hugepage-aware back end for TLB reach.

Reproduce it: misc/benchmarks/09-false-sharing, adjacent counters against padded counters across threads.

Lock-free structures and RCU

  • The Art of Multiprocessor Programming - States linearizability and the consensus hierarchy and builds both into working stacks, queues, lists and hash tables.
  • Simple, Fast, and Practical Non-Blocking and Blocking Concurrent Queue Algorithms - The lock-free queue later libraries copy, with the counted pointer against ABA and a two-lock queue beside it.
  • Hazard Pointers for C++26 - The standard-track form of safe reclamation, fixing when a retired node may be freed while a reader still holds it.
  • What is RCU? - The kernel's own statement of RCU as publish, wait for readers and keep old versions, with a free read side.
  • User-Level Implementations of Read-Copy Update - Defines liburcu's quiescent-state, signal-based and general RCU flavours and measures each read side against locks.

Thread pools and work stealing

  • Cilk: An Efficient Multithreaded Runtime System - Defines work and critical path, proves the work-stealing bound, and shows they alone predict a runtime's speedup.
  • The Implementation of the Cilk-5 Multithreaded Language - States the work-first principle, that overhead belongs on the rare steal path and not on every spawn.
  • Correct and Efficient Work-Stealing for Weak Memory Models - Gives the work-stealing deque a proven atomics form and derives which fence push, take and steal need on Arm and x86.
  • OpenMP Specifications - Fixes fork-join and tasking semantics, and the wait and binding controls deciding if idle workers spin, sleep or move.
  • oneTBB - The shipping work-stealing runtime, arenas and task groups over a deque, the home of grain size and spin-before-sleep.

10. NUMA and multi-socket

A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on.

NUMA and Linux memory placement

  • NUMA (Non-Uniform Memory Access): An Overview - The one account tying first touch, policy scope, zone reclaim and page movement together from the implementer's side.
  • What is NUMA? - Defines nodes, zonelists and the distance-ordered fallback that places an allocation once local memory runs out.
  • NUMA Memory Policy - The normative statement of policy scopes, every mode including weighted interleave, and the cpuset intersection rule.
  • Numa policy hit/miss statistics - Defines numa_hit, numa_miss and numa_foreign, the counters that show whether a policy put pages where it said.
  • numactl - Reference implementation of the policy API, prints the distance table and binds a binary that cannot be rebuilt.

Reproduce it: misc/benchmarks/10-first-touch, first touch of fresh pages against the second pass.

Topology and interconnects

  • NUMA Memory Performance - Explains the firmware-rated latency and bandwidth per initiator and target, and memory-side caches, that rank nodes.
  • Intel Xeon Processor Scalable Family Technical Overview - Where Intel names the mesh, UPI socket links, the directory-running home agent, and how SNC splits the cache.
  • Intel Xeon 6 with P-cores Configuration and Tuning Guide for HPC Applications - Defines SNC on current parts as one node per compute die, and fixes the numactl and numastat checks of placement.
  • BIOS and Workload Tuning Guide for AMD EPYC 9004 Series Processors - Discloses the I/O die, GMI and xGMI links, the NPS modes with their interleave widths, and the cache-as-NUMA override.
  • Arm Neoverse CMN-700 Coherent Mesh Network Technical Reference Manual - Defines the mesh, the home nodes holding the system cache and snoop filter, and the gateways joining sockets or CXL.

Migration, balancing and measured effects

  • move_pages(2) - Defines per-page migration of a running process, and a query reporting each page's node, the direct test of first touch.
  • sysctl kernel numa_balancing - Defines the hinting-fault sampling behind automatic balancing and tiering, and warns the overhead may not pay off.
  • Traffic Management: A Holistic Approach to Memory Placement on NUMA Systems - Proves against the kernel balancer that controller and link congestion, not remote latency, is what placement manages.
  • Intel Memory Latency Checker - Measures the node-to-node latency and bandwidth matrix and loaded latency on the x86 at hand, which no datasheet states.
  • High Performance Computing Tuning Guide for AMD EPYC 9004 Series Processors - Tabulates measured bandwidth by NPS mode, cores per die, boost and SMT, so the NPS trade-off is shown, not asserted.

11. OS and I/O

A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.

Syscalls and asynchronous I/O

  • vdso(7) - Defines the calls the kernel answers without a mode switch, and the clocksource condition for skipping the trap.
  • FlexSC: Flexible System Call Scheduling with Exception-Less System Calls - Separates a syscall's trap cost from its cache and TLB pollution, and shows the pollution can dominate.
  • An Analysis of Performance Evolution of Linux's Core Operations - Measures syscall and context switch cost across kernel releases and traces each slowdown to a named mitigation.
  • Efficient IO with io_uring - States the goals aio failed, and which io_uring features remove a syscall and which remove a copy.
  • Understanding Modern Storage APIs: A systematic study of libaio, SPDK, and io_uring - Measures io_uring's polling modes against libaio and SPDK, and shows the kernel poller needs its own core.

Reproduce it: misc/benchmarks/11-syscall-cost, the fixed cost of a kernel crossing across request sizes.

Scheduling, affinity and isolation

  • EEVDF Scheduler - Defines lag and virtual deadline, which the default class schedules by, and the slice request in sched_setattr.
  • The Linux Scheduler: a Decade of Wasted Cores - Proves cores sit idle while runnable threads queue, and gives the invariant checker that found the load-balancer bugs.
  • Control Group v2 - Defines cpu.max throttling, cpu.weight and the cpusets that bound affinity, the controls behind every container limit.
  • CPU Performance Scaling - Defines the governors, driver and boost switch that set a core's frequency, the sysfs state a measurement records.
  • CPU Isolation - Ties isolcpus, nohz_full, IRQ affinity, RCU offload and cpusets into one recipe, and lists the jitter it leaves.

Interrupts and kernel bypass

  • NAPI - Defines the polling, software coalescing, busy polling and IRQ suspension knobs that trade interrupts against latency.
  • DPDK Programmer's Guide - Defines the full bypass model, pinned poll-mode cores with no interrupts, that every kernel path is measured against.
  • The eXpress Data Path - Measures an in-kernel programmable path against DPDK and the stack per core, with the full configuration published.
  • Kernel vs. User-Level Networking: Don't Throw Out the Stack with the Interrupts - Separates direct and indirect NIC interrupt cost, measures the stack against bypass, and is where IRQ suspension began.
  • AF_XDP - Defines the socket and UMEM rings handing XDP frames to user space, and the zero-copy and need-wakeup modes.

Cache and bandwidth partitioning

  • Intel Resource Director Technology Architecture Specification - Defines classes of service, cache masks, bandwidth allocation and monitoring IDs, the model resctrl exposes.
  • User Interface for Resource Control feature (resctrl) - Defines the filesystem through which Linux exposes Intel, AMD and Arm partitioning, and the schemata format.
  • MPAM - Maps Arm's cache portion and bandwidth controls onto resctrl's schemata, and states which platform limits apply.
  • CPI2: CPU performance isolation for shared compute clusters - Shows at fleet scale that cycles per instruction alone finds an interfering neighbour and the one to throttle.
  • Heracles: Improving Resource Efficiency at Scale - Shows cache ways, cores, bandwidth and power must be partitioned together, or batch work reaches the tail.

12. Tail latency and production systems

A latency figure means nothing without its percentile, its load model and the way it was recorded.

Measuring the tail

  • The Tail at Scale - Shows why fan-out makes a rare slow server a common slow request, and names the techniques that tolerate variance.
  • Attack of the Killer Microseconds - Defines the stall band that out-of-order hardware cannot hide and a context switch cannot amortise.
  • How NOT to Measure Latency - Shows that a summary without a max discards the samples that define the tail, and closed-loop load never records them.
  • Coordinated Omission - The original definition of the recording error, with arithmetic for how far a reported percentile sits from the truth.
  • HdrHistogram - Keeps the whole distribution at fixed relative precision in constant time, so the far percentiles and max survive.

Reproduce it: misc/benchmarks/12-coordinated-omission, closed-loop against open-loop p99 under the same stalls.

Where jitter comes from

  • rt-tests - The reference wakeup-latency measurement for Linux, whose README states that an unloaded run proves nothing.
  • osnoise tracer - Counts the noise a spinning thread suffers and attributes each event to NMI, IRQ, softirq, thread or hardware.
  • Tales of the Tail - Derives the queueing-ideal tail and attributes the excess to scheduling, interrupt placement, power saving and NUMA.
  • Latency Implications of Virtual Memory - Measures with code the page-fault, TLB-shootdown and writeback stalls that memory mapping hides from the caller.
  • The KVM halt polling system - Defines the host-side polling after a vCPU halt that trades idle host CPU for guest wakeup time, unseen by the guest.

Load generation and production workloads

  • Open Versus Closed: A Cautionary Tale - Shows that open and closed load models disagree on response time and scheduling gains, with rules for choosing one.
  • wrk2 - Issues requests on a fixed schedule and times each from when it was due, so server stalls reach the percentiles.
  • Reconciling High Server Utilization and Sub-millisecond Quality-of-Service - Shows the tail, not throughput, caps a latency-critical server's utilisation, and how far co-located work lowers it.
  • TailBench - Pairs latency-critical services with an open-loop harness that records sojourn against service time per request.
  • Workload Analysis of a Large-Scale Key-Value Store - Measures the key, value and inter-arrival distributions of live key-value traffic, the shape load generators imitate.

Mechanical sympathy

  • Inter Thread Latency - Measures with code the floor for handing a cache line between cores, which every queue and lock is built on.
  • Single Writer Principle - States the design rule that removes write contention outright, using a contended increment's cost as the argument.
  • Optimizing a Ring Buffer for Throughput - Adds cached indices to a single-producer single-consumer ring and shows with counters the coherence traffic removed.
  • LMAX Disruptor - Applies the single writer rule and cache-line padding to a ring buffer, with the queue comparison that motivated it.
  • Aeron - Carries the single writer and batching rules through a whole transport, the reference beyond one in-process queue.

13. Inference on CPU

A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.

GEMM and BLAS

  • Anatomy of High-Performance Matrix Multiplication - Derives the cache blocking every fast CPU GEMM still uses by refining a memory model until the packed kernel falls out.
  • BLIS: A Framework for Rapidly Instantiating BLAS Functionality - Reduces a BLAS to one register-tile micro-kernel per architecture and states which loop owns which level of cache.
  • LLaMA Now Goes Faster on CPUs - Builds a register-tile kernel step by step and shows where outer-loop unrolling wins and where a vendor BLAS still does.
  • OpenBLAS - Sets run-time dispatch for a BLAS, one kernel table per microarchitecture chosen as a DYNAMIC_ARCH build loads.
  • oneDNN Matrix Multiplication Primitive - Fixes the data-type table, packed weight format and fused post-ops a CPU GEMM must expose to an inference graph.
  • Scaled Dot-Product Attention (oneDNN Graph) - Fixes the fused attention pattern, f32 accumulation under bf16 inputs and the shapes the fast CPU path accepts.

Reproduce it: misc/benchmarks/13-sgemm-naive-vs-blas, naive GEMM, a hand microkernel and the vendor BLAS, plus int8 against float32 dot products.

Runtimes

  • oneDNN - Generates VNNI and AMX kernels at run time under PyTorch, TensorFlow and OpenVINO, and calls Compute Library on Arm.
  • ggml - Defines the quantized block formats and the per-ISA dot-product kernels over them that every llama.cpp type rests on.
  • llama.cpp - Where new quantization types, kernels and thread pools land first, each with the perplexity and speed table behind it.
  • ONNX Runtime MLAS - Holds the CPU provider's GEMM, int8 and int4 MatMul kernels, dispatched per ISA at run time from SSE to AMX and SME.
  • OpenVINO CPU Device - States precision defaults per ISA, the int8 path through oneDNN and the streams model that turns cores into throughput.

Quantization

  • Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference - Defines the affine scale and zero-point scheme with integer accumulation that every int8 CPU runtime still implements.
  • Nuances of int8 Computations - States where int8 saturates on pre-VNNI x86, the compensation that makes signed by signed work, and what VNNI removes.
  • Quantize ONNX models - Defines the QDQ form against QOperator and dynamic against static, and which sign choices are safe on which CPUs.
  • k-quants - Defines super-block formats with quantized scales and the perplexity against size curve behind mixing types per tensor.
  • T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge - Replaces unpacking and multiplying low-bit weights with bit-sliced lookups, so a narrower weight costs less to apply.

Matrix extensions

  • Intel Software Developer Manuals - Defines AMX tiles, tile configuration and TMUL, the VNNI and bf16 dot products, and the XSAVE state a tile load needs.
  • Using XSTATE features in user space applications - Defines the arch_prctl permission Linux requires for AMX tile data and the trap a tile instruction takes until granted.
  • Add Intel Advanced Matrix Extensions (AMX) support to ggml - Holds the design report for the ggml AMX path, whose double buffering hides applying scales between tile products.
  • Arm Architecture Reference Manual for A-profile architecture - Defines streaming SVE mode, the ZA tile array and SME2 multi-vector instructions, with pseudocode a kernel must match.
  • SME Programmer's Guide - Walks through SME2 int8 and f32 matmul and gemv kernels and a lookup-table path, on a unit so far only in client parts.

Threading for inference

  • Thread management - Sets the physical-core default, the affinity it implies and the spin-wait controls that trade idle CPU for latency.
  • Performance Hints and Thread Scheduling - States the vendor defaults, one thread per core, SMT siblings off, core type by precision and one socket for latency.
  • Threadpool: take 2 - Defines the explicit thread pool with CPU masks, strict placement, priority and polling that ggml runs without OpenMP.
  • llama-bench - Defines the prompt and generation tests, repetitions and mean with deviation behind any comparable llama.cpp number.
  • Dual Epyc Genoa/Turin token generation performance bottlenecks - Traces poor decode scaling across sockets to remote NUMA access from weight placement, with numatop counts as evidence.

When CPU beats GPU

  • gpt-j example README (ggml) - Origin of ggml's bandwidth argument, NEON threads saturate a laptop's memory bus, so a GPU on that bus gains nothing.
  • llama.cpp Performance Testing - Holds the memory clock sweep with all else fixed, the token rate following it and flattening after a few threads.
  • SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs - Measures decode kernels on a Xeon as DRAM-bound, so AMX pays at batch one only once bytes per token shrink, with code.
  • MLPerf Inference v6.0 Results - Holds CPU-only entries whose logs show the batched Offline case lost to GPU entries at the same accuracy floor.

14. Hardware generations

The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.

Intel Xeon

  • Technical Overview of the 4th Gen Intel Xeon Scalable Processor Family - Where Intel states what Sapphire Rapids added: larger L2 and L3, DDR5, CXL, AMX and on-die accelerators.
  • Sapphire Rapids: Golden Cove Hits Servers - Measures L3 and memory latency across the tiled mesh and the slow clock ramp the vendor overview omits.
  • Emerald Rapids: 5th-Generation Intel Xeon Scalable Processors - The designers' statement of what changed: fewer, larger dies, a bigger shared L3, faster DDR5 and socket links.
  • A Look into Intel Xeon 6's Memory Subsystem - Measures per-die L3 under sub-NUMA clustering and the die-crossing cost on Granite Rapids beside Turin.
  • Benchmarking the Evolution of Performance and Energy Efficiency Across Recent Generations - Bandwidth-bound codes on Sapphire, Emerald and Granite Rapids and Sierra Forest with clocks, SMT and compiler stated.

AMD EPYC

  • AMD Next-Generation Zen 4 Core and 4th Gen AMD EPYC Server CPUs - The designers' account of the Zen 4 core and how it yields Genoa, Genoa-X, Bergamo and Siena.
  • Testing AMD's Bergamo: Zen 4c Spam - Tests the same-core claim for Zen 4c: cache latency, clock ceiling and core-to-core paths beside clock-matched Zen 4.
  • Software Optimization Guide for the AMD Zen5 Microarchitecture - Where AMD states what Zen 5 changed in front end, vector datapath and caches, the core Turin carries.
  • AMD EPYC 9005 Processor Architecture Overview - Defines Turin: Zen 5 or Zen 5c dies, which parts double die-to-IO links, NUMA modes and full-width AVX-512.
  • AMD's Turin: 5th Gen EPYC Launched - Measures what wider die-to-IO links and faster DDR5 do for Turin bandwidth, and where latency rose over Genoa.

Arm Neoverse server parts

  • AWS Graviton Getting Started - Where AWS states which Neoverse core, ISA revision, mesh, caches and compiler flag each Graviton generation carries.
  • Arm Neoverse V2 Core Software Optimization Guide - Sets the pipeline widths and instruction timings of the core Graviton 4, Grace and Axion share.
  • NVIDIA Grace Performance Tuning Guide - States the coherency fabric, LPDDR5X fit and MPAM cache and memory partitioning on Grace.
  • Arm Neoverse N2 Core Software Optimization Guide - States the timings and fusion rules of the narrower N line core in Cobalt 100 and Yitian.
  • Ampere Altra Rev A1 64-Bit Multi-Core Processor Datasheet - Where Ampere states the N1 part: private L2 per core, shared system cache, mesh and DDR4 fit.

Independent measurement across vendors

  • Microarchitectural Comparison and In-core Modeling of State-of-the-art CPUs - Measures Neoverse V2, Golden Cove and Zen 4 in-core at fixed clock, and each socket's clock under vector load.
  • On the Performance of Cloud-based ARM SVE for Zero-Knowledge Proving Systems - One SVE workload run on Graviton 3, Graviton 4, Yitian and Axion with compiler, runs and spread stated.
  • Arm's Neoverse V2, in AWS's Graviton 4 - Measures a sustained rename width below the stated one, cache latencies, mesh behaviour and cross-socket cost on V2.
  • ARM's Neoverse N2: Cortex A710 for Servers - Measures structure sizes, cache latencies and mesh behaviour of N2 on Yitian, the core Cobalt 100 carries.
  • AmpereOne at Hot Chips 2024: Maximizing Density - Puts vendor slides beside measurements of the predictor, small instruction cache, private L2 and long memory latency.

Reproduce it: misc/benchmarks/14-pcore-vs-ecore, the same three kernels on a performance core and an efficiency core.

15. Benchmarks

A score means what its suite's run rules say it means, so the rules come before the number.

Standard suites

  • SPEC CPU 2026 Run and Reporting Rules - Defines base against peak, rate against speed, the threading models a speed run may use and an Arm reference machine.
  • SPEC CPU: The Next Generation - Where the committee states how workloads were chosen and hardened, and defines the rolling round-robin rate.
  • SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison - Measures with counters what each workload stresses on x86 and Arm server parts, beside data-centre and inference suites.
  • MLPerf Inference Rules - Fixes model, accuracy floor and query pattern per scenario, so a Server score is throughput under a latency bound.
  • DCPerf: An Open-Source, Battle-Tested Performance Benchmark Suite for Datacenter Workloads - Shows standard suites misproject data-centre servers, and states the fleet-matching method the suite is built by.

Microbenchmark suites

  • Memory Bandwidth and Machine Balance in Current High Performance Computers - Defines sustainable bandwidth as what unit-stride loops get, not bus peak, and machine balance as flops per access.
  • STREAM Benchmark Reference Information - Sets the array size rule, timing over repeated trials, and counting bytes a loop asks for, not what the cache moved.
  • lmbench: Portable Tools for Performance Analysis - Origin of the one-mechanism-per-test method for memory, system call, pipe and socket latency, and what each leaves out.
  • uarch-bench - Isolates memory-level parallelism from load latency as separate tests, with DVFS held off before timing, x86 Linux only.
  • nanoBench: A Low-Overhead Tool for Running Microbenchmarks on x86 Systems - Shows why kernel mode with interrupts off matters, removes harness overhead, then recovers cache replacement policies.

Reproduce it: misc/benchmarks/15-stream-bandwidth, triad bandwidth by thread count against the vendor figure.

Methodology and what suites miss

  • How Not to Lie with Statistics: The Correct Way to Summarize Benchmark Results - Origin of the rule that normalised results take the geometric mean, which a SPEC ratio and the crimes list rest on.
  • Systems Benchmarking Crimes - Checklist of evaluation faults from sub-setting and improper baselines to arithmetic means of ratios, each with a fix.
  • Scientific Benchmarking of Parallel Computing Systems - Sets which mean fits costs, rates and ratios, when confidence intervals are owed, and the absolute base a speedup needs.
  • Rigorous Benchmarking in Reasonable Time - Decides how many builds, runs and iterations an experiment needs by measuring at which level the variation arises.
  • Profiling a warehouse-scale computer - Fleet counter profile showing services stall on instruction fetch and burn cycles in shared routines, which SPEC lacks.

16. Watchlist

Everything below is real but unproven: no item yet has all three of a written specification, a part you can buy, and a public measurement stating every one of the seven fields. Each line says what would promote it. Vendor multiples never qualify. Last checked 2026-09-15.

ISA extensions without a shipped server part

  • Intel AVX10.2 Architecture Specification - Folds the AVX-512 subsets into one versioned level and adds BF16 and FP8 forms, pending a shipped part and a run.
  • Intel Architecture Instruction Set Extensions Programming Reference - Ties APX, with doubled x86 registers, AVX10.2 and AMX FP8 tiles to Diamond Rapids, pending silicon and a public run.
  • SME in a Neoverse core, so far shipped only in client parts, pending a server core that carries it and a public run.
  • RISC-V Vector Extension, Version 1.0 - The ratified vector ISA, shipped only as IP and chiplets, pending a socketed server part and a run against Arm or x86.

Parts without a public measurement

  • 6th Gen AMD EPYC Server CPUs - The Zen 6 server family, so far a press release with no shipped part, pending shipment and a public run against Zen 5.
  • Intel Xeon 6+ Processors - The E-core-only sockets after Sierra Forest, shipped with vendor multiples footnoted off the page, pending a public run.
  • NVIDIA Vera CPU - Custom Arm cores with statically partitioned SMT and no architecture document, pending a specification and a public run.
  • Arm Neoverse V3 Core Software Optimization Guide - Vendor timing tables for the core shipped in Graviton 5 and previewed in Cobalt 200, pending a public run on the core.

Memory and interconnect

  • CXL Specification - Defines memory pooled across hosts on a coherent link, with latency so far estimated, pending a run on a shipped pool.
  • Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices - Measures expansion devices with frequency and SMT fixed, the nearest run to every field, pending compiler and flags.
  • DDR5MDB02 Multiplexed Rank Data Buffer, JESD82-552 - Defines the data buffer behind MRDIMMs, whose vendor bandwidth claims name no method, pending a run against RDIMMs.

Kernel paths and generated code

  • Extensible Scheduler Class - Lets a BPF program schedule at run time with safe fallback, once a run against the default scheduler states every field.
  • io_uring zero copy Rx - Lands payloads straight in user memory on header-splitting NICs, pending the implementer's epoll run naming every field.
  • T-MAC - Table-lookup kernels for low-bit weights, with a baseline stated but no frequency, compiler or flags, pending those.
  • Faster sorting algorithms discovered using deep reinforcement learning - Generated small sorts shipped in libc++, timed by CPU family with no model, compiler or flags stated, pending those.

What earns a place

Seven fields, and a number without all of them does not appear here:

1 CPU model and microarchitecture
2 core count used
3 frequency, with turbo and SMT state
4 compiler and flags
5 workload
6 baseline
7 measurement method

Miss one and the number is dropped; if the entry rests on that number, it moves to the watchlist or goes.

An entry itself has to be the thing, not writing about the thing: the paper that first described a mechanism, the specification or manual that defines it, the repository the implementation lives in, or a report from whoever did the work with code and reproducible measurements. Summaries, tutorials, surveys, marketing pages, mirrors and repackagings do not qualify. Every URL points at the live canonical copy, and misc/scripts/check_links.py and misc/scripts/check_format.py prove it on every push and again weekly.

CONTRIBUTING.md has the rules in full.

License

MIT. Maintained by @usamahz.

Format inspired by the GPU-side list at wafer-ai/gpu-perf-engineering-resources.