NEWS

Survey shows Rust's SIMD ecosystem more mature, but fragmented, in 2026

An independent survey on the state of SIMD in Rust in 2026 compares five vectorization libraries and shows what changes for those who need performance in databases, vector search, and AI inference.

An independent survey on the state of SIMD in Rust in 2026 compares five vectorization libraries and shows what changes for those who need performance in databases, vector search, and AI inference.

The bottleneck SIMD solves

Every modern processor carries far more arithmetic capacity than it can actually use, because decoding instructions is the bottleneck, not the arithmetic itself. The way out is to feed the processor a batch of numbers all at once, in a single instruction: instead of adding two numbers, you add two entire vectors.

It's this single instruction, multiple data technique that gives SIMD its name, and it's the engine behind vector databases, similarity search engines, and much of the AI model inference code running on CPU, according to Sergey "Shnatsel" Davidoff's survey on the state of SIMD in Rust in 2026 (https://shnatsel.github.io/state-of-simd-rust-2026/).

In theory, a recent x86 processor with 512-bit vectors can deliver up to an 8x gain in calculations with f64 or 64x with u8. In practice, according to the author, the result can run either slower or faster, depending on how the code is written and which hardware actually runs the binary.

The problem that only exists on x86

Each architecture gave its SIMD extension its own marketing name: ARM calls theirs NEON and makes it mandatory on every 64-bit CPU; WebAssembly has no marketing department and calls theirs the "WebAssembly 128-bit packed SIMD extension." x86_64, meanwhile, stacked extensions over the years: SSE2 came built in, SSE4.2 added operations, AVX and AVX2 brought 256-bit vectors, and AVX-512 went further with 512 bits.

The problem is that not every x86_64 CPU in production has AVX2 or AVX-512, so the Rust compiler, by default, can only assume SSE2. Teams that control their own hardware, such as those running their own cluster or using only a specific cloud, can force a minimum baseline:

RUSTFLAGS='-C target-cpu=x86-64-v3' cargo build --release

Those who distribute binaries to third parties don't have that option and need to resort to "function multiversioning": compiling the same function multiple times, one per CPU extension, and choosing the right version at runtime. ARM and WebAssembly don't suffer from this: NEON is guaranteed on every 64-bit ARM CPU, and WebAssembly solves it by compiling two binaries and letting JavaScript decide which one to load.

Three ways to vectorize in Rust

The survey organizes the ecosystem into three paths, from the simplest to the most granular:

  • Autovectorization: writing plain Rust and letting the compiler decide. Easier, but fragile: large or complex functions are rarely vectorized reliably, and the result changes between compiler versions.
  • Portable SIMD abstractions: types like i32x4 that add with plain +, hiding the intrinsics behind a generic API.
  • Platform-specific intrinsics: calling instructions like _mm256_add_ps (x86) or vaddq_u32 (ARM) directly, with full control and maximum verbosity.

A game-changing point for those working with f32/f64: Rust 1.98 stabilized algebraic operations like algebraic_add(), which let the compiler alter the result's precision in a controlled way, something similar to the -ffast-math flag, but safer. Before that, floating-point autovectorization simply didn't happen, because the compiler couldn't change the observable result of an addition.

The board of portable libraries

No single abstraction covers everything on its own today. The survey compares five portable SIMD crates used in production: std::simd (nightly), fearless_simd, wide, pulp, and macerator.

LibraryMultiversioningGeneric over vector widthTrigonometryMaturity
std::simdvia external crateyesweak (depends on a separate sleef crate)nightly-only
fearless_simdautomatic, even in small functionsyesnot yet availableversion 1.0
wideincompatible (except with cargo multivers)noavailable, but precision unspecifiedversion 1.0
pulpworks, but verboseyesweakused by faer
maceratorworksyesweakused by burn

std::simd remains restricted to nightly and suffers occasional API breakage, which weighs against those who need stable builds in production. wide, despite being mature and having good platform coverage, doesn't compose well with multiversioning outside of fixed hardware, which limits its use in software distributed to third parties.

fearless_simd: the crate that became the de facto standard

The survey's author became a maintainer of fearless_simd after contributing to the project over the past year, and the crate reached version 1.0, already with a formally documented security policy. Its central promise is solving multiversioning without boilerplate cost: just annotate the function with #[simd] and dispatch to the correct CPU extension happens automatically, even in small functions, where other approaches suffer from call overhead.

The crate also avoids using AVX-512 on CPUs where that extension only hurts performance, a behavior other libraries don't replicate by default. Its acknowledged weak point is trigonometry: there's still no port of something like SLEEF for fearless_simd's routines, so anyone needing vectorized sine and cosine with good precision still depends on other solutions.

Editor's note: the code snippets in this report reproduce examples from the original survey and were not run in our own environment. Validate before using in production.

For safe access to platform-specific intrinsics, Rust 1.87 started allowing platform intrinsics to be called without an unsafe block in the function signature, as long as the function is annotated with #[target_feature(enable = "avx2")]:

rust
#[target_feature(enable = "avx2")]
fn add_avx2(a: __m256, b: __m256) -> __m256 {
    _mm256_add_ps(a, b)
}

Even so, calling this function from outside still requires an unsafe block, unless a type token is used that proves at compile time that the CPU feature check has already been done. Crates like archmage and fearless_simd's own kernel! macro solve this problem in different ways.

The author was direct in justifying why he left out libraries whose development is mostly driven by generative AI:

I'm also excluding crates whose development is primarily AI-driven (e.g. magetypes, simdeez, thermite) because I cannot recommend them for production use, especially since the latter two are disconcertingly buggy.

Sergey "Shnatsel" Davidoff, maintainer of fearless_simd

What changes for those building in Brazil

Vector databases, semantic search engines, and local AI model inference pipelines on CPU are the most direct use cases for SIMD today, and Brazilian teams running this kind of workload on their own cloud or on dedicated servers now have more than one mature option besides writing intrinsics by hand. For those who control production hardware, common at companies with their own cluster or a closed contract with a single cloud provider, fixing -C target-cpu=x86-64-v3 in the build already guarantees AVX2 without any multiversioning.

For those distributing binaries to customers with varied hardware, a typical case for command-line tools or SDKs, fearless_simd reduces friction because multiversioning is configurable by whoever assembles the final binary, without requiring the library to be rewritten. Linear algebra libraries that already depend on pulp, such as faer, remain a solid option for those already in Rust's numerical computing ecosystem.

What's still missing

Vectorized trigonometry with good precision remains the weak point of nearly the entire ecosystem: the survey points to the sleef crate as the best available option, but it doesn't integrate with the multiversion crate, requiring anyone who needs both to manually fork and merge them. AVX-512 also remains tricky territory: libraries like wide, pulp, and macerator only check whether the extension exists, without verifying whether it actually pays off on the specific hardware, which can hurt performance instead of improving it.

This leaves the field open for anyone who wants good coverage of mathematical operations and safe automatic dispatch at the same time, without giving up either one.

Source

Sergey "Shnatsel" Davidoff, "The state of SIMD in Rust in 2026" (https://shnatsel.github.io/state-of-simd-rust-2026/).

Translated from the Brazilian Portuguese original · Read the original