SIMD and auto-vectorisation
Axle gives you SIMD two ways: vector types you drive lane by lane
(f32x8, i32x4), and auto-vectorisation, where an ordinary array loop is
rewritten into vector operations for you.
What this page covers: what SIMD is, in one minute · the two surfaces and
which to pick · the manual surface — lane shapes, construction, operations,
bitwise and shift, a complete example · reading scalars back out · the errors a
vector mistake produces · what the automatic rewrite covers and what stays
scalar · steering it with @vectorize and @unroll · what you can rely on, and
what you cannot.
What SIMD is (start here)
SIMD stands for Single Instruction, Multiple Data. A normal
(scalar) CPU instruction operates on one value at a time: add takes two numbers and produces one sum. A SIMD instruction does the same operation on a whole bundle of values at once — one add that produces four (or eight, or sixteen) sums in a single step.
The bundle lives in a vector register : a wide CPU register
(128, 256, or 512 bits) sliced into equal lanes, one value per
lane. A 256-bit register holds eight 32-bit integers, so it has eight lanes. A SIMD add on two such registers adds lane 0 to
lane 0, lane 1 to lane 1, and so on — all eight additions happen
in the time one scalar addition would take.
Scalar: four separate adds, one per step
a0 + b0 → c0
a1 + b1 → c1 four instructions,
a2 + b2 → c2 four steps
a3 + b3 → c3
Vector (SIMD): one add across four lanes, in one step
┌────┬────┬────┬────┐ ┌────┬────┬────┬────┐
│ a0 │ a1 │ a2 │ a3 │ + │ b0 │ b1 │ b2 │ b3 │
└────┴────┴────┴────┘ └────┴────┴────┴────┘
│ one vector add instruction │
▼ ▼ ▼ ▼
┌────┬────┬────┬────┐
│ c0 │ c1 │ c2 │ c3 │ one instruction, one step
└────┴────┴────┴────┘ That is the entire idea : if a loop does the same arithmetic to every element of an array, doing four/eight/sixteen elements per step is a straight speed-up proportional to the lane count — the CPU already has the wide registers, SIMD just uses them. The lanes run in genuine parallel hardware, not a hidden loop.
Two things make SIMD non-automatic in general :
- the lanes must be independent — lane
ican’t need a result that lanei-1produced in the same step (no loop-carried dependence) ; - the operation must exist as a machine instruction for that lane shape (every CPU adds vectors of integers; not every CPU fuses a multiply-add, which is why the runtime picks the best kernel per host — see CPU dispatch).
Axle’s two SIMD surfaces
Axle gives you SIMD two ways :
- Manual SIMD via the
f32x8,i32x4, … vector types — you name the lane shape and call lane operations explicitly (FMA, masked blend, shuffle, horizontal sum). Use this when you’re writing the hot kernel (dot product, matmul block, FIR filter) and want exact control. - Auto-vectorisation — the compiler recognises an ordinary
array loop and rewrites it into vector form for you, no special
types required. Use this when you have a normal
forloop that should just go fast. The@vectorize/@unrollannotations let you nudge the loops the automatic rewrite doesn’t cover.
The rest of this page covers the manual surface first, then what the automatic rewrite does (and, honestly, does not do) for you.
Manual SIMD
Vector types
Fixed-lane shapes named <scalar>x<lanes>, spanning four register
widths (64 / 128 / 256 / 512 bits) :
// 64-bit
i8x8 i16x4 i32x2 f32x2
// 128-bit (universal x86-64 baseline, ARM NEON)
i8x16 i16x8 i32x4 i64x2 f32x4 f64x2
// 256-bit (AVX2 / SVE)
i8x32 i16x16 i32x8 i64x4 f32x8 f64x4
// 512-bit (AVX-512)
i8x64 i16x32 i32x16 i64x8 f32x16 f64x8 The lane scalar follows the usual iN / uN / fN naming — every signed
shape above has its unsigned twin (u8x16, u32x4, u64x2, …), at the same
widths. Total bit-width must be 64, 128, 256, or 512 — anything else is
rejected at compile time.
An unsigned vector is a distinct type from its signed twin, and selects the
unsigned operation wherever the two differ — umax/umin on a reduction, udiv/urem on a division, a logical shift on >>. Its lanes are unsigned
scalars, so building one takes u32s :
let all_ones : u32 = -1 as u32;
let v : u32x4 = u32x4(all_ones, 1, 1, 1);
let hi : u32 = v.max(); // 4294967295 — a signed vector answers 1
let s : i32x4 = i32x4(-1, 1, 1, 1);
let lo : i32 = s.max(); // 1 Unsigned lanes
Every integer shape also has an unsigned twin spelled with a u lane (u8x16, u16x8, u32x4, u64x2, … at the same widths). An
unsigned shape is a distinct type from its signed counterpart and
selects the unsigned machine op where the two differ — max / min lower to umax / umin, / % to udiv / urem, and >> to a
logical shift. Operations that don’t depend on signedness (+, -, *, &, |, ^, ~, <<, eq, movemask, splat, lane access)
behave identically.
An unsigned lane’s scalar is the unsigned scalar of its width — u8x16.splat(x) takes a u8, and v.lane(k) on a u32x4 returns a u32. Signed and unsigned shapes never mix implicitly: u32x4 + i32x4 is rejected (E0001).
Construction
Either broadcast a single value across every lane :
let zero : f64x4 = f64x4.splat(0.0); Or list every lane explicitly :
let v : f64x4 = f64x4(1.0, 2.0, 3.0, 4.0); The literal-list constructor requires exactly lanes arguments,
each of the lane scalar type.
Feeding an f32 shape. A bare 0.0 literal is an f64, and there
is no f32 literal suffix. In argument position nothing narrows it, so f32x8.splat(0.0) is rejected with E0716 (expects f32, got f64).
An annotated let does narrow, so bind first :
let z : f32 = 0.0;
let zero : f32x8 = f32x8.splat(z); Operations
The full method set. Two call shapes, by whether the op produces a vector or acts on one:
Producers — receiver is the vector type name (Tx<L>):
| Method | Signature | Description |
|---|---|---|
Tx<L>.splat(v) | (elem) → Tx<L> | broadcast scalar to every lane |
Tx<L>.load(buf, i) | (elem[]\|elem[N], int) → Tx<L> | typed <NxT> load at &buf[i], alignment 1 ; index is any integer width |
Tx<L>(v0, …) | (elem × L) → Tx<L> | literal-list constructor — exactly lanes args |
Tx<L>.select(mask, a, b) | (iWxL, Tx<L>, Tx<L>) → Tx<L> | masked blend — mask is the matching integer shape (f32xN→i32xN, f64xN→i64xN; an integer-lane vector is its own mask) |
Tx<L>.convert(v) | (UxL) → Tx<L> | each lane to T, lane count kept: a wider lane extends by v’s signedness, a narrower one keeps the low bits ; integer lanes only |
Tx<L>.saturate(v) | (UxL) → Tx<L> | each lane to the narrower T, clamped to T’s range instead of wrapped — the pack back to 8 or 16 bits ; integer lanes only |
Tx<L>.reinterpret(v) | (any shape of the same width) → Tx<L> | the same bits read as Tx<L> — u16x8.reinterpret(i16x8), u8x16.reinterpret(i32x4) |
Instance methods — receiver is the vector value (v.method(…)):
| Method | Signature | Description |
|---|---|---|
v.store(buf, i) | (elem[]\|elem[N], int) → void | typed store at &buf[i], alignment 1 ; index is any integer width |
a.fma(b, c) | (Tx<L>, Tx<L>) → Tx<L> | fused multiply-add a*b + c ; float lanes only |
v.sum() | () → elem | horizontal sum |
v.max() / v.min() | () → elem | horizontal max / min (smax/smin on signed lanes, umax/umin on unsigned) |
v.shuffle(idx) | (i32[N]) → TxN | lanes of v picked by a compile-time index array, each entry in [0, lanes) ; N indices give N lanes, so this also splits a vector (v.shuffle([4, 5, 6, 7]) is the high half) |
a.shuffle(b, idx) | (Tx<L>, i32[N]) → TxN | the same over two vectors: entries [0, L) pick a’s lanes, [L, 2L) pick b’s — joins, interleaves ([0, L, 1, L+1, …]), any reorder across both |
a.eq(b) | (Tx<L>) → iWxL | lane-wise equality — -1 (all-bits-set) in matching lanes, 0 elsewhere, in the matching integer-mask shape |
a.lt(b) / a.gt(b) | (Tx<L>) → iWxL | lane-wise < / > — -1 where the relation holds, 0 elsewhere, same mask shape as eq. Unsigned lanes order unsigned (0xFFFFFFFF is the largest u32); a NaN lane answers 0 |
v.movemask() | () → i32 or i64 | pack each lane’s MSB into the low lanes bits of an i32 — of an i64 for a 64-lane vector ; integer lanes only (E0718 on float lanes) |
v.storeCompressed(buf, i, mask) | (elem[]\|elem[N], int, iWxL) → void | write the lanes whose mask lane is non-zero one after another from &buf[i], nothing else ; mask is v’s integer-mask shape |
a.laneMin(b) / a.laneMax(b) | (Tx<L>) → Tx<L> | the smaller / larger of each lane pair, by the lane’s signedness ; integer lanes |
v.abs() | () → Tx<L> | absolute value per lane, the minimum staying the minimum ; signed integer lanes |
a.addSaturating(b) / a.subSaturating(b) | (Tx<L>) → Tx<L> | lane sum / difference clamped to the lane’s range ; integer lanes |
a.mulHigh(b) | (Tx<L>) → Tx<L> | the high half of each lane’s double-width product — a fixed-point or reciprocal multiply ; integer lanes |
a.dotPairs(b) | (Tx<L>) → WxL/2 | lanes 2i and 2i+1 multiplied at twice the width and summed (the sum wraps) — i16x16 gives i32x8 ; signed lanes of 8, 16 or 32 bits |
v.leadingZeros() / v.popcount() | () → Tx<L> | leading zero bits (a zero lane answers its width) / set bits, per lane ; integer lanes |
v[k] / v.lane(k) | (const int) → elem | extract lane k (extractelement) ; k is a literal const in [0, lanes) |
v.withLane(k, x) | (const int, elem) → Tx<L> | non-destructive lane insert → new vector ; k const in [0, lanes), v unchanged |
The split is the canonical surface: an op that acts on an existing
vector is only the value-receiver form. The type-name spelling of an
instance op (f32x8.sum(v), f32x8.store(buf, i, v), …) is rejected
with E0714 — call it on the value instead.
shuffletakes its mask inline. LLVM’sshufflevectorneeds the permutation as a constant operand, so the index array is read straight off the call site:v.shuffle([3, 2, 1, 0]). A mask held in a named binding — even ani32[4]local initialised from literals — is rejected withE0716(shuffle index array must contain compile-time constants), because the value is a runtime array by the time the call is checked. Write the literal list at the call, or, when the permutation genuinely varies at runtime,storeto a buffer, reorder, and reload.
fn reversed(v : i32x4) : i32x4 {
return v.shuffle([3, 2, 1, 0]);
} lt and gt produce the same mask, so they feed select for a
branch-free lane-wise choice and movemask for a bit set :
fn lanewiseMin(a : i32x8, b : i32x8) : i32x8 {
return i32x8.select(a.lt(b), a, b); // the smaller lane of each pair
}
fn lanesAbove(v : u32x8, limit : u32) : i32 {
// bit k is set when lane k is greater than `limit`
return v.gt(u32x8.splat(limit)).movemask();
} The conversions and the lane-wise integer ops are what a fixed-point kernel is written in — here a 16-bit multiply-add packed back to bytes, the step a colour transform takes per pixel pair:
fn weighted(samples : i16x16, weights : i16x16) : u8x8 {
let sums : i32x8 = samples.dotPairs(weights); // one multiply-add per pair
let rounded : i32x8 = (sums + i32x8.splat(1 << 13)) >> 14;
return u8x8.saturate(rounded); // clamp to 0..255, no wrap
} On x86 each of these is one instruction — dotPairs on i16 lanes is pmaddwd, mulHigh is pmulhw / pmulhuw, saturate a saturating pack —
and a target without one computes the same value.
The eq + movemask pair is the canonical SwissTable probe
shape :
// Locate any matching key in a 4-lane group with one branch.
fn probe(keys : i32[], base : i32, target : i32) : i32 {
let group : i32x4 = i32x4.load(keys, base);
let bits : i32 = group.eq(i32x4.splat(target)).movemask();
// `bits` is a 4-bit set of matching lanes — feed it into a
// trailing-zero scan to read the matching slot.
return bits;
} Arithmetic operators +, -, *, /, % work directly between
two vectors of the same shape :
fn ops(a : f32x8, b : f32x8, u : i32x4, v : i32x4, m : u32x4, n : u32x4) {
let c : f32x8 = a + b; // elementwise add
let d : f32x8 = a * b; // elementwise mul
let r : f32x8 = a % b; // elementwise IEEE remainder
let q : i32x4 = u % v; // signed integer remainder per lane
let w : u32x4 = m % n; // unsigned integer remainder per lane
} Both operands must have identical (elem, lanes) and signedness —
no implicit broadcasting, no signed/unsigned mixing. Use .splat(...) to lift a scalar.
Lane-wise / and % carry the same divide-by-zero protection as
their scalar counterparts : if any lane’s divisor is zero (or, for signed integer lanes, a lane computes INT_MIN / -1), the program
aborts with a runtime error before the divide — exactly as a scalar
zero divisor would. Unsigned lanes are guarded on a zero divisor only
(INT_MIN / -1 is not overflow for udiv/urem). A bad lane never
silently yields ±Inf or a poison value :
let a : f64x4 = f64x4(1.0, 2.0, 3.0, 4.0);
let b : f64x4 = f64x4(2.0, 0.0, 2.0, 2.0);
let q : f64x4 = a / b; // aborts: vector float division:
// a lane divisor is zero Bitwise and shift operators
&, |, ^, ~, <<, >> and >>> work lane by lane on integer-lane
vectors, and so do their compound forms (&=, |=, ^=, <<=, >>=, >>>=) :
fn mix(a : i32x8, b : i32x8, u : u32x8) {
let m : i32x8 = (a & b) ^ ~a; // lane-wise and / xor / complement
let s : i32x8 = a >> 31; // arithmetic: each lane keeps its sign
let z : i32x8 = a >>> 31; // logical: zeros shifted in
let h : u32x8 = u >> 28; // unsigned lanes: logical
let p : i32x8 = a << b; // one count per lane
u >>= 1;
} - The two operands of
&,|,^must have the same shape and signedness, as for arithmetic —a & 1is rejected; writea & i32x8.splat(1). - A shift count is either a vector of the same shape (one count per
lane) or a single scalar applied to every lane. A scalar count has the
lane’s own type: a literal adopts it (
u >> 28onu32x8), and a variable of another type needs a cast (u >> (n as u32)). >>follows the lane type like its scalar counterpart: it keeps the sign on signed lanes and shifts in zeros on unsigned ones.>>>always shifts in zeros.- A literal count outside
0..widthis a compile-time error (i32x8 << 32). A count computed at run time is taken modulo the lane’s width —n & 31on 32-bit lanes — exactly as for a scalar shift. - Float lanes have no bitwise operators (
E0001).
A complete example
Dot product over f64[8] arrays via two AVX-shaped FMA chunks plus
a horizontal sum :
fn dot(a : f64[8], b : f64[8]) : f64 {
let acc : f64x4 = f64x4.splat(0.0);
let va0 : f64x4 = f64x4.load(a, 0);
let vb0 : f64x4 = f64x4.load(b, 0);
let acc1 : f64x4 = va0.fma(vb0, acc);
let va1 : f64x4 = f64x4.load(a, 4);
let vb1 : f64x4 = f64x4.load(b, 4);
let acc2 : f64x4 = va1.fma(vb1, acc1);
return acc2.sum();
} This kernel compiles to two FMA instructions plus one horizontal
reduce on any host. On a CPU with AVX2 + FMA (Haswell and newer),
each FMA collapses to a single vfmadd231pd ; on older CPUs the
compiler falls back to mulpd + addpd automatically. You don’t
recompile or ship a different binary per CPU level — the dispatch
happens at process startup based on the actual host capabilities.
Reading scalars back out
A vector is opaque ; there are four ways to pull values into locals, by intent :
| You want… | Approach | Cost |
|---|---|---|
| an aggregate (sum / max / min) | v.sum() | one horizontal reduction |
| every lane in locals | v.store(buf, 0) then buf[k] | one store + N loads (often elided) |
| one lane known at compile time | v[k] / v.lane(k) | one extractelement |
| the match pattern of a comparison | m.movemask() | one pack op |
let v : f64x4 = f64x4(10.0, 20.0, 30.0, 40.0);
let total : f64 = v.sum(); // 100.0 — aggregate
let third : f64 = v[2]; // 30.0 — one lane, no buffer
let w : f64x4 = v.withLane(2, 99.0); // <10, 20, 99, 40>, v unchanged Errors you might see
| Code | Trigger |
|---|---|
| E0001 | An operator on mismatched vector shapes (u32x4 + i32x4, a & b across lane counts), a vector beside a scalar where no broadcast is allowed (a & 1), a bitwise operator on float lanes, a shift count of the wrong scalar type, or a literal shift count outside 0..width. |
| E0714 | Unknown method on a vector type (f32x8.frobnicate()), or the type-name spelling of an instance op (f32x8.sum(v) instead of v.sum(), i32x4.lt(a, b) instead of a.lt(b)). |
| E0715 | Wrong number of arguments (a.fma(b) — fma takes two operands). |
| E0716 | Argument type doesn’t match the expected signature (including a.lt(b) / a.eq(b) on two different shapes, such as i32x8.gt(u32x8)), a lane index (v[k] / lane / withLane) is non-const or outside [0, lanes), a shuffle index is out of range or its count makes no vector type (i32x4 with three indices), or a conversion breaks its width rule (i32x4.convert(i16x8), i32x8.saturate(i16x8), i32x8.reinterpret(i16x8)). |
| E0717 | Literal-list constructor with the wrong lane count (f32x8(1.0, 2.0)). |
| E0718 | A float-only op (fma) on integer lanes; an integer-only op (movemask, convert, laneMin, mulHigh, …) on float lanes; abs on unsigned lanes; dotPairs on lanes with no signed pair sum. |
Constraints
- No implicit broadcasting.
f64x4 + 1.0is rejected ; usef64x4.splat(1.0)explicitly. The one exception is a shift count, which may be a scalar of the lane type. - No implicit lane-shape conversion.
f64x4 + f64x2is rejected ; change a lane type withconvert/saturate/reinterpret, a lane count withshuffle. f32lane literals are awkward. Axle rejects implicitf64 → f32narrowing and has no0.0f32literal — float SIMD work is easiest inf64xN.
Auto-vectorisation
Some ordinary array loops vectorise with no annotation and no
vector types at all — you write a plain for and the compiler
rewrites it into the vector-add-across-lanes shape from the diagram
above :
fn axpy(out : f64[], a : f64[], b : f64[], x : f64, n : i32) {
for i of 0..n {
out[i] = a[i] * x + b[i];
}
} The compiler turns this into a vector main loop that processes
one full register of elements per step, followed by a scalar
tail that mops up the n % lanes leftover elements the vector
loop couldn’t fill :
n = 10, lanes = 4
vector main loop scalar tail
┌───────────┬───────────┐ ┌────┬────┐
│ 0 1 2 3 │ 4 5 6 7 │ │ 8 │ 9 │
└───────────┴───────────┘ └────┴────┘
step 1 step 2 one at a time This rewrite happens in the compiler front-end (sema), before
LLVM, so the SIMD form is guaranteed — it does not depend on
any optimiser cost model deciding it was worthwhile. That is the
difference from the @vectorize hint further down, which only asks LLVM to try.
STATUS — the automatic rewrite is deliberately PARTIAL. It only fires on the narrow, provably-safe loop shape below. Anything outside it stays a scalar loop (and may still be vectorised by LLVM’s own pass, but without the guarantee). It is not a general vectoriser.
What the automatic rewrite covers
The body must be a single same-index store — d[i] = <expression> where every array read is s[i] at the same loop index i (never s[i-1], s[i+1], or s[j]), and i is
used only as an index, never as a value. Same-index means each lane
touches only its own column, so the lanes are independent and the
rewrite is sound even when d aliases a source (d[i] = d[i] + s[i] is fine) — no alias analysis needed.
Inside that store, these build a vectorisable expression :
| Shape | Example | Notes |
|---|---|---|
| map with a constant / invariant | d[i] = s[i] * 2 + c | loop-invariant scalars are broadcast (splatted) to every lane |
| zip-map of two arrays | d[i] = a[i] + b[i] | |
| AXPY (scale-and-add) | d[i] = k * a[i] + b[i] | on float lanes with an FMA-capable host this fuses into one fma |
+ - * | d[i] = a[i] - b[i] | the core arithmetic |
| integer negation | d[i] = -a[i] | lowered to 0 - a[i] (exact for integers) |
/ by a non-trapping literal | d[i] = a[i] / 4 | integer: divisor ≠ 0 and ≠ -1; float: divisor ≠ 0.0 |
The loop also needs its bounds provably in range (so the rewrite
never drops a bounds check that the scalar loop would have run) and
a trip count large enough to fill at least one vector — a constant
count below 8 is left scalar as not worth the tail; a runtime count n is always rewritten (the main loop simply runs zero times when n is small).
What it does NOT do (stays scalar)
| Not covered | Why |
|---|---|
Reductions — sum = sum + a[i] | the store target isn’t d[i]; the result carries across iterations (a loop-carried dependence) |
Stencils / shifted index — d[i] = a[i-1] + a[i+1] | lane i would need a neighbour’s input — not independent |
% (modulo) | no safe vector form is emitted; kept scalar with its per-element divisor check |
Runtime or trapping / — d[i] = a[i] / x | a vector divide guard would abort on a whole chunk, changing which stores already happened |
Float negation — d[i] = -f[i] on float | 0.0 - x differs from -x at x = +0.0 (IEEE keeps +0.0 and -0.0 distinct), and the HIR has no vector unary-negate node |
| Multi-statement bodies, conditional stores, computed indices | outside the single-same-index-store shape |
A reduction like for i of 0..n { sum = sum + a[i]; } is a
common thing to want vectorised, and the automatic rewrite does not
cover it. Write it with the manual surface (v.sum() over f64x4 accumulators) when it’s hot.
Steering the rest with annotations
For every loop the automatic rewrite leaves scalar, the two
annotations below attach to for, while, or do-while and steer
LLVM’s own vectoriser. Both are honoured at -O1 and up.
@vectorize
fn addInto(out : i32[], a : i32[], b : i32[], n : i32) {
@vectorize
for i of 0..n {
out[i] = a[i] + b[i];
}
} Tells LLVM to enable the loop vectoriser on this loop. The vector width is picked by LLVM based on the host’s SIMD level and the body’s cost.
@vectorize(width: N)
fn mulInto(out : i32[], a : i32[], b : i32[], n : i32) {
@vectorize(width: 8)
for i of 0..n {
out[i] = a[i] * b[i];
}
} Pins the vector width to N lanes. The compiler refuses widths the
machine running the compiler can’t honour — a width above that machine’s
maximum is rejected up front with E0292, never silently downgraded to
a slower path. --target names the output and does not move the cap :
| Host SIMD level | max_lanes |
|---|---|
| Baseline (SSE2 on x86_64, NEON on aarch64) | 4 |
| AVX2 | 8 |
| AVX-512F | 16 |
| No-SIMD embedded triple | 1 |
@vectorize(disable)
fn fold(table : i32[], n : i32) : i32 {
let state : i32 = 0;
@vectorize(disable)
for i of 0..n {
// Tight scalar dependency — vectorising would scatter through
// memory more than it saves.
state = state * 31 + table[i];
}
return state;
} Forces the loop to stay scalar. Useful when the auto-vectoriser produces strided scatter / gather code that is slower than the plain form.
@unroll(N)
fn total(values : i32[], count : i32) : i32 {
let accum : i32 = 0;
@unroll(4)
for i of 0..count {
accum = accum + values[i];
}
return accum;
} Pins the unroll factor to N. N must be a positive integer
literal.
When to pick which
| You want to … | Use |
|---|---|
| Express a specific SIMD kernel (dot product, FFT butterfly, FIR tap) | Manual SIMD (f64x4.fma, etc.) |
Run a normal for faster, trust the compiler | @vectorize |
| Avoid LLVM vectorising what you proved scalar | @vectorize(disable) |
| Trade code size for ILP on a small constant-trip loop | @unroll(N) |
Both surfaces compose : you can write an auto-vectorised loop that
calls a manual-SIMD helper, or use @vectorize(disable) on the
outer loop of a manually-vectorised kernel so the inner SIMD path
isn’t disrupted.
What you can rely on
- Every vector method ships on every supported lane shape, except where
its table row names a lane restriction:
fmais float-lane only,movemask, the conversions and the lane-wise integer ops are integer-lane only,absanddotPairswant signed lanes (eachE0718otherwise) ; andshuffletakes its index list inline at the call site (see the note under Operations). - The operators
+ - * / %and, on integer lanes,& | ^ ~ << >> >>>compute on each lane exactly what the scalar operator computes on that lane’s value, compound forms included. a.eq(b),a.lt(b)anda.gt(b)set a lane to-1exactly when the scalar==,<or>on that lane’s values is true — ordered by the lane’s signedness, and false on a NaN lane — and to0otherwise.- A single binary adapts to the host CPU at startup — same
.exeon a 2010 Westmere CPU, a 2015 Haswell, and a 2023 Zen 4. - Compile errors surface the wrong-shape problem at the call site
(typed
E0714-E0718codes), never as a runtime crash. - The automatic loop rewrite covers same-index maps, zip-maps, and
AXPY (
+ - *, integer negation,/by a safe literal) — and guarantees the SIMD form for them whenever the HIR transforms run (-O1and up); at-O0the loop is left as written. Reductions, stencils,%, and float negation are not auto-vectorised; reach for the manual types there. @vectorize/@unrollare honoured at every optimisation level above-O0.
Limitations
| Limit | Why it holds |
|---|---|
No implicit broadcasting — f64x4 + 1.0 is rejected | use f64x4.splat(1.0); the one exception is a shift count, which may be a scalar of the lane type |
| No implicit lane-shape conversion | change a lane type with convert / saturate / reinterpret, a lane count with shuffle |
A f32 lane literal cannot be written directly | a bare 0.0 is an f64 and there is no f32 suffix; bind an annotated let first |
A shuffle index array must sit inline at the call | LLVM’s shufflevector needs the permutation as a constant operand |
| The automatic rewrite is deliberately partial | reductions, stencils, %, float negation and multi-statement bodies stay scalar; reach for the manual surface there |
| A signed and an unsigned lane shape never mix | u32x4 + i32x4 is rejected (E0001) |
The per-construct tables above — Constraints, What it does NOT do — carry the detail behind each row.
See also
- Optimisations overview — where auto-vectorisation sits in the full census of what the compiler optimises, and what the compiler tells LLVM about aliasing and ranges to unlock vector loops.
- Annotations reference — every
@…annotation the compiler recognises. - Memory model — the storage tiers vector buffers sit in.
- SIMD CPU dispatch — how the runtime picks AVX2+FMA vs SSE2 at startup, with no per-call overhead.
- Concept index — every SIMD operation on one page.