Axle v0.14.1

What the compiler tells LLVM

An optimiser can only do what it can prove. Two loads through two pointers can be reordered, hoisted out of a loop, or vectorised — but only if the optimiser knows the pointers don’t overlap. In C, it usually can’t know, so it assumes the worst and emits defensive code.

Axle has a structural advantage: its ownership system knows things C’s type system never records. An own T argument aliases nothing. A Shared<T> field is a live object of a known size. A dynamic array’s length header is written once and never changes. The compiler turns each of these facts into an LLVM attribute or a piece of metadata stamped onto the IR, so the backend can optimise as aggressively as if you’d hand-annotated everything.

None of this changes what your program does — it changes what the optimiser is allowed to assume, which is often the difference between a vectorised loop and a scalar one. This is a big part of why two look-alike programs can compile very differently. This page is the honest inventory: what gets stamped, and what does not.

No-alias: “these pointers don’t overlap”

The single most valuable fact for an optimiser is that two pointers can’t touch the same memory. Axle stamps it in three places where the ownership system guarantees it:

  • A freshly allocated array or object is based on no pre-existing pointer, so the allocator’s return slot carries noalias — LLVM knows a store through it can’t disturb anything the caller already holds.
  • A function that returns a freshly-minted new T() carries noalias on its result, the same malloc-style contract: the returned pointer is not reachable any other way. A function that merely returns an existing object (return self.field;) does not get it — that pointer might alias the caller’s own storage.
  • The lone pointer parameter of an own T function has nothing for it to alias, so it too is noalias.

You can also ask for it directly with @noalias on a parameter.

Parameter facts: what a function does to its pointers

For each pointer parameter, a compiler pass walks the body once and stamps what the function actually does with it — never guessing from the type or the ownership keyword (a plain borrowed self still writes its own fields, so the keyword alone would be a miscompile). The stamped attributes:

AttributeMeansLets LLVM…
readonlythe body never writes through this pointer, and never lets it escapekeep a load of it hoisted out of a loop
writeonlythe body writes through it but never reads its targetdrop a redundant load
captures(none)the pointer never escapes the callreuse the caller’s storage across the call
nonnullthe pointer is one of the non-null owning/borrowing shapesdelete the null-check on every downstream use
dereferenceable(N)it addresses at least N valid bytes — stamped only for an own / shared / borrow / arenaref of a resolvable classspeculate a load out of a branch or loop
dereferenceable_or_null(N)it addresses at least N valid bytes or is null — the array-allocation path, where the byte count is a compile-time constanthoist a load, keeping the null case explicit
align Nit has known alignmentpick the aligned load/store form

Each of these is a proof, not a hope: a wrong readonly would let -O2 delete a store you needed, so the analysis defaults to the conservative answer (assume written, assume captured) for anything it can’t model. A pointer that is not statically known to be non-null gets no dereferenceable / align: the bytes it might address are not known to exist.

Type-based alias analysis: disjoint fields

Beyond individual pointers, the compiler tells LLVM which kinds of access can’t overlap, via type-based alias metadata (!tbaa) on every typed load and store. Two accesses tagged with distinct types — or two distinct fields of the same object, at known offsets — provably don’t alias, so a loop that reads one field and writes another can be reordered, CSE’d, or vectorised.

This is where Axle’s richer type information pays off directly: it knows the exact byte offset and scalar type of every field, so it emits a struct-path descriptor that marks point.x and point.y as disjoint. On array-heavy code this can unlock a reordering or a vectorisation LLVM would otherwise leave on the table — most of all where removing the aliasing worry is the one thing standing between a scalar loop and a vector one. The size of the win is entirely workload-dependent; treat it as “this unblocks an optimisation”, not a fixed percentage.

Load facts: range, immutability, definedness

On individual loads, the compiler attaches what it has proven about the value they produce:

  • !range — a value whose storage constrains it. A bool load (i8) is [0, 2), a char (i32) is [0, 0x110000), and the array-length header load is [0, i64::MAX). LLVM then folds away redundant comparisons (if (b > 2) on a bool) and builds tighter jump tables. The bound is a property of how the language stores the value, so it holds on every execution of the load — which is what lets LLVM move the load speculatively. An enum tag is exactly such a storage-bounded integer. A runtime type-id is not: it is a hash of the class’s qualified name and reaches every value of its word, so no !range on it could hold — the compiler stamps none on either (see Limitations).
  • !invariant.load — a value that never changes after it’s first set: the runtime type-id of an object, and a dynamic array’s length header in a program where no buffer is ever grown in place. LLVM can hoist such a load clean out of a loop. A program that grows a buffer (the compiler turns a hand-written allocate-copy-free into one realloc) rewrites that header, so the claim is dropped program-wide and each bounds check reloads the length.
  • !noundef — every scalar load is defined, because Axle’s type system requires initialisation at declaration (every let has an initialiser; struct literals zero-fill). LLVM can then fold select / phi simplifications against the value without a “could-be-poison” escape hatch.
  • !nonnull / !align / !dereferenceable — the load-site mirrors of the parameter attributes above, for a pointer loaded out of a field.

Cold-path placement

The compiler also tells LLVM which branches are unlikely, via branch weights (!prof). The throw path after an exception poll, a divide-by-zero / bounds-check / out-of-memory guard, the last-reference drop of a Shared<T>, and a ?? fallback’s null branch are all tagged cold — so LLVM lays them off the hot instruction-cache line and prioritises the common path for register allocation. A user if / else arm that the compiler proves always throws or always calls a cold function is tagged too. A function is cold when every path through it throws, or when it is declared @heat(cold); a declared-cold function is also never inlined.

Floating-point fast-math

Float instructions carry a fast-math flag bundle driven by the function’s floating-point mode. The default mode is contract, which allows a multiply-add to fuse into a single fma instruction (the reason an a*x + y loop can become one FMA per element) but forbids reassociation — so results stay predictable. Full IEEE-relaxing fast-math (nnan, ninf, reassoc, reciprocal division, …) is a stronger mode that reorders and approximates; it is not applied by default, because it can change the last bit of a result, and a floating-point division is never silently turned into a reciprocal multiply. @float(fast) on a function opts it into the full fast-math bundle; @float(strict) pins it to bare IEEE-754, refusing even the default contraction. See Annotations reference.

Limitations

Everything above ships. What follows is what the compiler leaves conservative, so you don’t count on it:

  • Enum-tag and type-id !range. A tag’s bound is known — it indexes the variant list. A type-id’s is not: the tag is an FNV-1a hash of the class’s qualified name, spread over the whole u32. Only the bool / char / array-length loads are actually stamped, so a tag or type-id load carries no !range.
  • General-integer and loop-induction !range. The interval analysis computes these bounds, but announces them at the site that owns them (@llvm.assume, which is control-dependent) rather than publishing them as a per-load fact that could hoist out of the branch which established it. A general integer load therefore carries no !range even when its bounds are known.
  • !invariant.load on your own immutable fields. Proving a field is never written after construction needs an alias analysis the compiler does not perform — a field reachable through an aliased pointer stays writable — so only the runtime’s own always-immutable loads (array length, type-id) carry it.
  • noalias on disjoint-proved borrowed parameters. Two borrowed pointers that don’t overlap could in principle be noalias, but proving it needs a whole-program aliasing analysis that’s undecidable with indirect calls, so it isn’t stamped.

In every one of these cases the compiler emits the conservative IR — correct, just not as aggressively optimised as it could eventually be. None of them is a correctness gap; each is a proof the compiler does not attempt.

See also

optimisationllvmattributesmetadataalias-analysis