Axle v0.14.1

SIMD and auto-vectorisation

Axle gives you SIMD two ways: vector types you drive lane by lane (f32x8, i32x4), and auto-vectorisation, where an ordinary array loop is rewritten into vector operations for you.

What this page covers: what SIMD is, in one minute · the two surfaces and which to pick · the manual surface — lane shapes, construction, operations, bitwise and shift, a complete example · reading scalars back out · the errors a vector mistake produces · what the automatic rewrite covers and what stays scalar · steering it with @vectorize and @unroll · what you can rely on, and what you cannot.

What SIMD is (start here)

SIMD stands for Single Instruction, Multiple Data. A normal (scalar) CPU instruction operates on one value at a time: add takes two numbers and produces one sum. A SIMD instruction does the same operation on a whole bundle of values at once — one add that produces four (or eight, or sixteen) sums in a single step.

The bundle lives in a vector register : a wide CPU register (128, 256, or 512 bits) sliced into equal lanes, one value per lane. A 256-bit register holds eight 32-bit integers, so it has eight lanes. A SIMD add on two such registers adds lane 0 to lane 0, lane 1 to lane 1, and so on — all eight additions happen in the time one scalar addition would take.

Scalar: four separate adds, one per step
   a0 + b0 → c0
   a1 + b1 → c1          four instructions,
   a2 + b2 → c2          four steps
   a3 + b3 → c3

Vector (SIMD): one add across four lanes, in one step
   ┌────┬────┬────┬────┐   ┌────┬────┬────┬────┐
   │ a0 │ a1 │ a2 │ a3 │ + │ b0 │ b1 │ b2 │ b3 │
   └────┴────┴────┴────┘   └────┴────┴────┴────┘
            │  one  vector  add  instruction  │
            ▼    ▼    ▼    ▼
   ┌────┬────┬────┬────┐
   │ c0 │ c1 │ c2 │ c3 │   one instruction, one step
   └────┴────┴────┴────┘

That is the entire idea : if a loop does the same arithmetic to every element of an array, doing four/eight/sixteen elements per step is a straight speed-up proportional to the lane count — the CPU already has the wide registers, SIMD just uses them. The lanes run in genuine parallel hardware, not a hidden loop.

Two things make SIMD non-automatic in general :

  • the lanes must be independent — lane i can’t need a result that lane i-1 produced in the same step (no loop-carried dependence) ;
  • the operation must exist as a machine instruction for that lane shape (every CPU adds vectors of integers; not every CPU fuses a multiply-add, which is why the runtime picks the best kernel per host — see CPU dispatch).

Axle’s two SIMD surfaces

Axle gives you SIMD two ways :

  1. Manual SIMD via the f32x8, i32x4, … vector types — you name the lane shape and call lane operations explicitly (FMA, masked blend, shuffle, horizontal sum). Use this when you’re writing the hot kernel (dot product, matmul block, FIR filter) and want exact control.
  2. Auto-vectorisation — the compiler recognises an ordinary array loop and rewrites it into vector form for you, no special types required. Use this when you have a normal for loop that should just go fast. The @vectorize / @unroll annotations let you nudge the loops the automatic rewrite doesn’t cover.

The rest of this page covers the manual surface first, then what the automatic rewrite does (and, honestly, does not do) for you.

Manual SIMD

Vector types

Fixed-lane shapes named <scalar>x<lanes>, spanning four register widths (64 / 128 / 256 / 512 bits) :

// 64-bit
i8x8    i16x4   i32x2   f32x2

// 128-bit (universal x86-64 baseline, ARM NEON)
i8x16   i16x8   i32x4   i64x2   f32x4   f64x2

// 256-bit (AVX2 / SVE)
i8x32   i16x16  i32x8   i64x4   f32x8   f64x4

// 512-bit (AVX-512)
i8x64   i16x32  i32x16  i64x8   f32x16  f64x8

The lane scalar follows the usual iN / uN / fN naming — every signed shape above has its unsigned twin (u8x16, u32x4, u64x2, …), at the same widths. Total bit-width must be 64, 128, 256, or 512 — anything else is rejected at compile time.

An unsigned vector is a distinct type from its signed twin, and selects the unsigned operation wherever the two differ — umax/umin on a reduction, udiv/urem on a division, a logical shift on >>. Its lanes are unsigned scalars, so building one takes u32s :

let all_ones : u32 = -1 as u32;
let v : u32x4 = u32x4(all_ones, 1, 1, 1);
let hi : u32 = v.max();          // 4294967295 — a signed vector answers 1

let s : i32x4 = i32x4(-1, 1, 1, 1);
let lo : i32 = s.max();          // 1

Unsigned lanes

Every integer shape also has an unsigned twin spelled with a u lane (u8x16, u16x8, u32x4, u64x2, … at the same widths). An unsigned shape is a distinct type from its signed counterpart and selects the unsigned machine op where the two differ — max / min lower to umax / umin, / % to udiv / urem, and >> to a logical shift. Operations that don’t depend on signedness (+, -, *, &, |, ^, ~, <<, eq, movemask, splat, lane access) behave identically.

An unsigned lane’s scalar is the unsigned scalar of its width — u8x16.splat(x) takes a u8, and v.lane(k) on a u32x4 returns a u32. Signed and unsigned shapes never mix implicitly: u32x4 + i32x4 is rejected (E0001).

Construction

Either broadcast a single value across every lane :

let zero : f64x4 = f64x4.splat(0.0);

Or list every lane explicitly :

let v : f64x4 = f64x4(1.0, 2.0, 3.0, 4.0);

The literal-list constructor requires exactly lanes arguments, each of the lane scalar type.

Feeding an f32 shape. A bare 0.0 literal is an f64, and there is no f32 literal suffix. In argument position nothing narrows it, so f32x8.splat(0.0) is rejected with E0716 (expects f32, got f64). An annotated let does narrow, so bind first :

let z : f32 = 0.0;
let zero : f32x8 = f32x8.splat(z);

Operations

The full method set. Two call shapes, by whether the op produces a vector or acts on one:

Producers — receiver is the vector type name (Tx<L>):

MethodSignatureDescription
Tx<L>.splat(v)(elem) → Tx<L>broadcast scalar to every lane
Tx<L>.load(buf, i)(elem[]\|elem[N], int) → Tx<L>typed <NxT> load at &buf[i], alignment 1 ; index is any integer width
Tx<L>(v0, …)(elem × L) → Tx<L>literal-list constructor — exactly lanes args
Tx<L>.select(mask, a, b)(iWxL, Tx<L>, Tx<L>) → Tx<L>masked blend — mask is the matching integer shape (f32xN→i32xN, f64xN→i64xN; an integer-lane vector is its own mask)
Tx<L>.convert(v)(UxL) → Tx<L>each lane to T, lane count kept: a wider lane extends by v’s signedness, a narrower one keeps the low bits ; integer lanes only
Tx<L>.saturate(v)(UxL) → Tx<L>each lane to the narrower T, clamped to T’s range instead of wrapped — the pack back to 8 or 16 bits ; integer lanes only
Tx<L>.reinterpret(v)(any shape of the same width) → Tx<L>the same bits read as Tx<L> — u16x8.reinterpret(i16x8), u8x16.reinterpret(i32x4)

Instance methods — receiver is the vector value (v.method(…)):

MethodSignatureDescription
v.store(buf, i)(elem[]\|elem[N], int) → voidtyped store at &buf[i], alignment 1 ; index is any integer width
a.fma(b, c)(Tx<L>, Tx<L>) → Tx<L>fused multiply-add a*b + c ; float lanes only
v.sum()() → elemhorizontal sum
v.max() / v.min()() → elemhorizontal max / min (smax/smin on signed lanes, umax/umin on unsigned)
v.shuffle(idx)(i32[N]) → TxNlanes of v picked by a compile-time index array, each entry in [0, lanes) ; N indices give N lanes, so this also splits a vector (v.shuffle([4, 5, 6, 7]) is the high half)
a.shuffle(b, idx)(Tx<L>, i32[N]) → TxNthe same over two vectors: entries [0, L) pick a’s lanes, [L, 2L) pick b’s — joins, interleaves ([0, L, 1, L+1, …]), any reorder across both
a.eq(b)(Tx<L>) → iWxLlane-wise equality — -1 (all-bits-set) in matching lanes, 0 elsewhere, in the matching integer-mask shape
a.lt(b) / a.gt(b)(Tx<L>) → iWxLlane-wise < / > — -1 where the relation holds, 0 elsewhere, same mask shape as eq. Unsigned lanes order unsigned (0xFFFFFFFF is the largest u32); a NaN lane answers 0
v.movemask()() → i32 or i64pack each lane’s MSB into the low lanes bits of an i32 — of an i64 for a 64-lane vector ; integer lanes only (E0718 on float lanes)
v.storeCompressed(buf, i, mask)(elem[]\|elem[N], int, iWxL) → voidwrite the lanes whose mask lane is non-zero one after another from &buf[i], nothing else ; mask is v’s integer-mask shape
a.laneMin(b) / a.laneMax(b)(Tx<L>) → Tx<L>the smaller / larger of each lane pair, by the lane’s signedness ; integer lanes
v.abs()() → Tx<L>absolute value per lane, the minimum staying the minimum ; signed integer lanes
a.addSaturating(b) / a.subSaturating(b)(Tx<L>) → Tx<L>lane sum / difference clamped to the lane’s range ; integer lanes
a.mulHigh(b)(Tx<L>) → Tx<L>the high half of each lane’s double-width product — a fixed-point or reciprocal multiply ; integer lanes
a.dotPairs(b)(Tx<L>) → WxL/2lanes 2i and 2i+1 multiplied at twice the width and summed (the sum wraps) — i16x16 gives i32x8 ; signed lanes of 8, 16 or 32 bits
v.leadingZeros() / v.popcount()() → Tx<L>leading zero bits (a zero lane answers its width) / set bits, per lane ; integer lanes
v[k] / v.lane(k)(const int) → elemextract lane k (extractelement) ; k is a literal const in [0, lanes)
v.withLane(k, x)(const int, elem) → Tx<L>non-destructive lane insert → new vector ; k const in [0, lanes), v unchanged

The split is the canonical surface: an op that acts on an existing vector is only the value-receiver form. The type-name spelling of an instance op (f32x8.sum(v), f32x8.store(buf, i, v), …) is rejected with E0714 — call it on the value instead.

shuffle takes its mask inline. LLVM’s shufflevector needs the permutation as a constant operand, so the index array is read straight off the call site: v.shuffle([3, 2, 1, 0]). A mask held in a named binding — even an i32[4] local initialised from literals — is rejected with E0716 (shuffle index array must contain compile-time constants), because the value is a runtime array by the time the call is checked. Write the literal list at the call, or, when the permutation genuinely varies at runtime, store to a buffer, reorder, and reload.

fn reversed(v : i32x4) : i32x4 {
    return v.shuffle([3, 2, 1, 0]);
}

lt and gt produce the same mask, so they feed select for a branch-free lane-wise choice and movemask for a bit set :

fn lanewiseMin(a : i32x8, b : i32x8) : i32x8 {
    return i32x8.select(a.lt(b), a, b);   // the smaller lane of each pair
}

fn lanesAbove(v : u32x8, limit : u32) : i32 {
    // bit k is set when lane k is greater than `limit`
    return v.gt(u32x8.splat(limit)).movemask();
}

The conversions and the lane-wise integer ops are what a fixed-point kernel is written in — here a 16-bit multiply-add packed back to bytes, the step a colour transform takes per pixel pair:

fn weighted(samples : i16x16, weights : i16x16) : u8x8 {
    let sums : i32x8 = samples.dotPairs(weights);       // one multiply-add per pair
    let rounded : i32x8 = (sums + i32x8.splat(1 << 13)) >> 14;
    return u8x8.saturate(rounded);                      // clamp to 0..255, no wrap
}

On x86 each of these is one instruction — dotPairs on i16 lanes is pmaddwd, mulHigh is pmulhw / pmulhuw, saturate a saturating pack — and a target without one computes the same value.

The eq + movemask pair is the canonical SwissTable probe shape :

// Locate any matching key in a 4-lane group with one branch.
fn probe(keys : i32[], base : i32, target : i32) : i32 {
    let group : i32x4 = i32x4.load(keys, base);
    let bits  : i32   = group.eq(i32x4.splat(target)).movemask();
    // `bits` is a 4-bit set of matching lanes — feed it into a
    // trailing-zero scan to read the matching slot.
    return bits;
}

Arithmetic operators +, -, *, /, % work directly between two vectors of the same shape :

fn ops(a : f32x8, b : f32x8, u : i32x4, v : i32x4, m : u32x4, n : u32x4) {
    let c : f32x8 = a + b;       // elementwise add
    let d : f32x8 = a * b;       // elementwise mul
    let r : f32x8 = a % b;       // elementwise IEEE remainder
    let q : i32x4 = u % v;       // signed integer remainder per lane
    let w : u32x4 = m % n;       // unsigned integer remainder per lane
}

Both operands must have identical (elem, lanes) and signedness — no implicit broadcasting, no signed/unsigned mixing. Use .splat(...) to lift a scalar.

Lane-wise / and % carry the same divide-by-zero protection as their scalar counterparts : if any lane’s divisor is zero (or, for signed integer lanes, a lane computes INT_MIN / -1), the program aborts with a runtime error before the divide — exactly as a scalar zero divisor would. Unsigned lanes are guarded on a zero divisor only (INT_MIN / -1 is not overflow for udiv/urem). A bad lane never silently yields ±Inf or a poison value :

let a : f64x4 = f64x4(1.0, 2.0, 3.0, 4.0);
let b : f64x4 = f64x4(2.0, 0.0, 2.0, 2.0);
let q : f64x4 = a / b;   // aborts: vector float division:
                         //         a lane divisor is zero

Bitwise and shift operators

&, |, ^, ~, <<, >> and >>> work lane by lane on integer-lane vectors, and so do their compound forms (&=, |=, ^=, <<=, >>=, >>>=) :

fn mix(a : i32x8, b : i32x8, u : u32x8) {
    let m : i32x8 = (a & b) ^ ~a;   // lane-wise and / xor / complement
    let s : i32x8 = a >> 31;        // arithmetic: each lane keeps its sign
    let z : i32x8 = a >>> 31;       // logical: zeros shifted in
    let h : u32x8 = u >> 28;        // unsigned lanes: logical
    let p : i32x8 = a << b;         // one count per lane
    u >>= 1;
}
  • The two operands of &, |, ^ must have the same shape and signedness, as for arithmetic — a & 1 is rejected; write a & i32x8.splat(1).
  • A shift count is either a vector of the same shape (one count per lane) or a single scalar applied to every lane. A scalar count has the lane’s own type: a literal adopts it (u >> 28 on u32x8), and a variable of another type needs a cast (u >> (n as u32)).
  • >> follows the lane type like its scalar counterpart: it keeps the sign on signed lanes and shifts in zeros on unsigned ones. >>> always shifts in zeros.
  • A literal count outside 0..width is a compile-time error (i32x8 << 32). A count computed at run time is taken modulo the lane’s width — n & 31 on 32-bit lanes — exactly as for a scalar shift.
  • Float lanes have no bitwise operators (E0001).

A complete example

Dot product over f64[8] arrays via two AVX-shaped FMA chunks plus a horizontal sum :

fn dot(a : f64[8], b : f64[8]) : f64 {
    let acc : f64x4 = f64x4.splat(0.0);

    let va0 : f64x4 = f64x4.load(a, 0);
    let vb0 : f64x4 = f64x4.load(b, 0);
    let acc1 : f64x4 = va0.fma(vb0, acc);

    let va1 : f64x4 = f64x4.load(a, 4);
    let vb1 : f64x4 = f64x4.load(b, 4);
    let acc2 : f64x4 = va1.fma(vb1, acc1);

    return acc2.sum();
}

This kernel compiles to two FMA instructions plus one horizontal reduce on any host. On a CPU with AVX2 + FMA (Haswell and newer), each FMA collapses to a single vfmadd231pd ; on older CPUs the compiler falls back to mulpd + addpd automatically. You don’t recompile or ship a different binary per CPU level — the dispatch happens at process startup based on the actual host capabilities.

Reading scalars back out

A vector is opaque ; there are four ways to pull values into locals, by intent :

You want…ApproachCost
an aggregate (sum / max / min)v.sum()one horizontal reduction
every lane in localsv.store(buf, 0) then buf[k]one store + N loads (often elided)
one lane known at compile timev[k] / v.lane(k)one extractelement
the match pattern of a comparisonm.movemask()one pack op
let v : f64x4 = f64x4(10.0, 20.0, 30.0, 40.0);
let total : f64 = v.sum();   // 100.0  — aggregate
let third : f64 = v[2];      //  30.0  — one lane, no buffer
let w : f64x4 = v.withLane(2, 99.0);   // <10, 20, 99, 40>, v unchanged

Errors you might see

CodeTrigger
E0001An operator on mismatched vector shapes (u32x4 + i32x4, a & b across lane counts), a vector beside a scalar where no broadcast is allowed (a & 1), a bitwise operator on float lanes, a shift count of the wrong scalar type, or a literal shift count outside 0..width.
E0714Unknown method on a vector type (f32x8.frobnicate()), or the type-name spelling of an instance op (f32x8.sum(v) instead of v.sum(), i32x4.lt(a, b) instead of a.lt(b)).
E0715Wrong number of arguments (a.fma(b) — fma takes two operands).
E0716Argument type doesn’t match the expected signature (including a.lt(b) / a.eq(b) on two different shapes, such as i32x8.gt(u32x8)), a lane index (v[k] / lane / withLane) is non-const or outside [0, lanes), a shuffle index is out of range or its count makes no vector type (i32x4 with three indices), or a conversion breaks its width rule (i32x4.convert(i16x8), i32x8.saturate(i16x8), i32x8.reinterpret(i16x8)).
E0717Literal-list constructor with the wrong lane count (f32x8(1.0, 2.0)).
E0718A float-only op (fma) on integer lanes; an integer-only op (movemask, convert, laneMin, mulHigh, …) on float lanes; abs on unsigned lanes; dotPairs on lanes with no signed pair sum.

Constraints

  • No implicit broadcasting. f64x4 + 1.0 is rejected ; use f64x4.splat(1.0) explicitly. The one exception is a shift count, which may be a scalar of the lane type.
  • No implicit lane-shape conversion. f64x4 + f64x2 is rejected ; change a lane type with convert / saturate / reinterpret, a lane count with shuffle.
  • f32 lane literals are awkward. Axle rejects implicit f64 → f32 narrowing and has no 0.0f32 literal — float SIMD work is easiest in f64xN.

Auto-vectorisation

Some ordinary array loops vectorise with no annotation and no vector types at all — you write a plain for and the compiler rewrites it into the vector-add-across-lanes shape from the diagram above :

fn axpy(out : f64[], a : f64[], b : f64[], x : f64, n : i32) {
    for i of 0..n {
        out[i] = a[i] * x + b[i];
    }
}

The compiler turns this into a vector main loop that processes one full register of elements per step, followed by a scalar tail that mops up the n % lanes leftover elements the vector loop couldn’t fill :

n = 10, lanes = 4

vector main loop        scalar tail
┌───────────┬───────────┐  ┌────┬────┐
│ 0 1 2 3   │ 4 5 6 7   │  │ 8  │ 9  │
└───────────┴───────────┘  └────┴────┘
   step 1       step 2       one at a time

This rewrite happens in the compiler front-end (sema), before LLVM, so the SIMD form is guaranteed — it does not depend on any optimiser cost model deciding it was worthwhile. That is the difference from the @vectorize hint further down, which only asks LLVM to try.

STATUS — the automatic rewrite is deliberately PARTIAL. It only fires on the narrow, provably-safe loop shape below. Anything outside it stays a scalar loop (and may still be vectorised by LLVM’s own pass, but without the guarantee). It is not a general vectoriser.

What the automatic rewrite covers

The body must be a single same-index store — d[i] = <expression> where every array read is s[i] at the same loop index i (never s[i-1], s[i+1], or s[j]), and i is used only as an index, never as a value. Same-index means each lane touches only its own column, so the lanes are independent and the rewrite is sound even when d aliases a source (d[i] = d[i] + s[i] is fine) — no alias analysis needed.

Inside that store, these build a vectorisable expression :

ShapeExampleNotes
map with a constant / invariantd[i] = s[i] * 2 + cloop-invariant scalars are broadcast (splatted) to every lane
zip-map of two arraysd[i] = a[i] + b[i]
AXPY (scale-and-add)d[i] = k * a[i] + b[i]on float lanes with an FMA-capable host this fuses into one fma
+ - *d[i] = a[i] - b[i]the core arithmetic
integer negationd[i] = -a[i]lowered to 0 - a[i] (exact for integers)
/ by a non-trapping literald[i] = a[i] / 4integer: divisor ≠ 0 and ≠ -1; float: divisor ≠ 0.0

The loop also needs its bounds provably in range (so the rewrite never drops a bounds check that the scalar loop would have run) and a trip count large enough to fill at least one vector — a constant count below 8 is left scalar as not worth the tail; a runtime count n is always rewritten (the main loop simply runs zero times when n is small).

What it does NOT do (stays scalar)

Not coveredWhy
Reductions — sum = sum + a[i]the store target isn’t d[i]; the result carries across iterations (a loop-carried dependence)
Stencils / shifted index — d[i] = a[i-1] + a[i+1]lane i would need a neighbour’s input — not independent
% (modulo)no safe vector form is emitted; kept scalar with its per-element divisor check
Runtime or trapping / — d[i] = a[i] / xa vector divide guard would abort on a whole chunk, changing which stores already happened
Float negation — d[i] = -f[i] on float0.0 - x differs from -x at x = +0.0 (IEEE keeps +0.0 and -0.0 distinct), and the HIR has no vector unary-negate node
Multi-statement bodies, conditional stores, computed indicesoutside the single-same-index-store shape

A reduction like for i of 0..n { sum = sum + a[i]; } is a common thing to want vectorised, and the automatic rewrite does not cover it. Write it with the manual surface (v.sum() over f64x4 accumulators) when it’s hot.

Steering the rest with annotations

For every loop the automatic rewrite leaves scalar, the two annotations below attach to for, while, or do-while and steer LLVM’s own vectoriser. Both are honoured at -O1 and up.

@vectorize

fn addInto(out : i32[], a : i32[], b : i32[], n : i32) {
    @vectorize
    for i of 0..n {
        out[i] = a[i] + b[i];
    }
}

Tells LLVM to enable the loop vectoriser on this loop. The vector width is picked by LLVM based on the host’s SIMD level and the body’s cost.

@vectorize(width: N)

fn mulInto(out : i32[], a : i32[], b : i32[], n : i32) {
    @vectorize(width: 8)
    for i of 0..n {
        out[i] = a[i] * b[i];
    }
}

Pins the vector width to N lanes. The compiler refuses widths the machine running the compiler can’t honour — a width above that machine’s maximum is rejected up front with E0292, never silently downgraded to a slower path. --target names the output and does not move the cap :

Host SIMD levelmax_lanes
Baseline (SSE2 on x86_64, NEON on aarch64)4
AVX28
AVX-512F16
No-SIMD embedded triple1

@vectorize(disable)

fn fold(table : i32[], n : i32) : i32 {
    let state : i32 = 0;
    @vectorize(disable)
    for i of 0..n {
        // Tight scalar dependency — vectorising would scatter through
        // memory more than it saves.
        state = state * 31 + table[i];
    }
    return state;
}

Forces the loop to stay scalar. Useful when the auto-vectoriser produces strided scatter / gather code that is slower than the plain form.

@unroll(N)

fn total(values : i32[], count : i32) : i32 {
    let accum : i32 = 0;
    @unroll(4)
    for i of 0..count {
        accum = accum + values[i];
    }
    return accum;
}

Pins the unroll factor to N. N must be a positive integer literal.

When to pick which

You want to …Use
Express a specific SIMD kernel (dot product, FFT butterfly, FIR tap)Manual SIMD (f64x4.fma, etc.)
Run a normal for faster, trust the compiler@vectorize
Avoid LLVM vectorising what you proved scalar@vectorize(disable)
Trade code size for ILP on a small constant-trip loop@unroll(N)

Both surfaces compose : you can write an auto-vectorised loop that calls a manual-SIMD helper, or use @vectorize(disable) on the outer loop of a manually-vectorised kernel so the inner SIMD path isn’t disrupted.

What you can rely on

  • Every vector method ships on every supported lane shape, except where its table row names a lane restriction: fma is float-lane only, movemask, the conversions and the lane-wise integer ops are integer-lane only, abs and dotPairs want signed lanes (each E0718 otherwise) ; and shuffle takes its index list inline at the call site (see the note under Operations).
  • The operators + - * / % and, on integer lanes, & | ^ ~ << >> >>> compute on each lane exactly what the scalar operator computes on that lane’s value, compound forms included.
  • a.eq(b), a.lt(b) and a.gt(b) set a lane to -1 exactly when the scalar ==, < or > on that lane’s values is true — ordered by the lane’s signedness, and false on a NaN lane — and to 0 otherwise.
  • A single binary adapts to the host CPU at startup — same .exe on a 2010 Westmere CPU, a 2015 Haswell, and a 2023 Zen 4.
  • Compile errors surface the wrong-shape problem at the call site (typed E0714-E0718 codes), never as a runtime crash.
  • The automatic loop rewrite covers same-index maps, zip-maps, and AXPY (+ - *, integer negation, / by a safe literal) — and guarantees the SIMD form for them whenever the HIR transforms run (-O1 and up); at -O0 the loop is left as written. Reductions, stencils, %, and float negation are not auto-vectorised; reach for the manual types there.
  • @vectorize / @unroll are honoured at every optimisation level above -O0.

Limitations

LimitWhy it holds
No implicit broadcasting — f64x4 + 1.0 is rejecteduse f64x4.splat(1.0); the one exception is a shift count, which may be a scalar of the lane type
No implicit lane-shape conversionchange a lane type with convert / saturate / reinterpret, a lane count with shuffle
A f32 lane literal cannot be written directlya bare 0.0 is an f64 and there is no f32 suffix; bind an annotated let first
A shuffle index array must sit inline at the callLLVM’s shufflevector needs the permutation as a constant operand
The automatic rewrite is deliberately partialreductions, stencils, %, float negation and multi-statement bodies stay scalar; reach for the manual surface there
A signed and an unsigned lane shape never mixu32x4 + i32x4 is rejected (E0001)

The per-construct tables above — Constraints, What it does NOT do — carry the detail behind each row.

See also

simdvectorisationperformancevectors