Axle v0.14.1

Optimisations

Axle is a statically-typed language that compiles to native code through LLVM, with compiler-inferred memory management — no garbage collector, and no lifetime annotations to write. A lot of the work that would otherwise be your job — deciding what goes on the stack, removing a bounds check the surrounding code already guarantees, turning a divide into a multiply, vectorising a numeric loop — the compiler does ahead of time and stamps onto the program.

This section is the complete census of that work: what the compiler optimises, how much it can prove, and — just as important — where it stops. Axle does not claim to make every program fast; it claims to remove the overheads it can prove are unnecessary, and to leave a runtime check in place the moment it cannot. A removed check that shouldn’t have been removed is a memory-safety bug, so every elision below is conservative: when in doubt, the check stays.

The full map

Every optimisation the compiler performs, grouped by what it acts on. Pages under this chapter are the optimisation-lens view; pages under Memory, Vectorisation, and Compiler internals are the deep dives — this chapter links to them rather than repeating them.

Where your objects live

What you getPage
The heap is the last resort. new T(...) that doesn’t escape becomes a stack slot (no malloc); a short-lived object graph shares an arena (bulk-freed at function exit); only a genuinely shared value pays a reference count; the raw heap is used only when nothing cheaper is provably safe.Memory placement
Escape analysis — the pass that proves which tier each allocation may use, and the eight routing strategies it stamps.Escape analysis and promotion
Shared<T> reference counting — the atomic-vs-plain refcount, the 16-byte header, and where inc/dec go non-atomic.Refcount and transitions

Where your indexes are checked

What you getPage
Bounds-check elimination — array and container index checks are removed when the compiler proves the index is in range, including across a chain of comparisons (i < n, n ≤ len), across a function call, and from how a container was built.Bounds-check elimination
Per-value range elision — the sibling interval analysis that proves a single value non-negative or in range, which also folds list.get(i) ?? fallback down to a bare load.Loop and idiom rewrites
Conditional-fault elision — when the compiler proves a container access is in range, the whole checked-access idiom folds away: recv.get(i) ?? d loses its nullable box and ??, and a recv.set(i, v) whose out-of-range case would throw loses the throw and its exception poll — a bare load or store remains. Rides the same bounds proof (including the loop-carried and cross-call chains).Bounds-check elimination

Your arithmetic

What you getPage
Strength reduction — constant divide / modulo become a multiply-shift (even at -O0); power-of-two / and % become a shift and a mask; a loop-invariant runtime divisor is turned into a hoisted multiply; x * 8 becomes x << 3; the divide-by-zero guard is dropped when the divisor is provably non-zero.Arithmetic strength reduction
Peepholes, CSE, abs-after-modulo, dead-code elimination — the small local rewrites that clean up arithmetic before the backend sees it.Loop and idiom rewrites

Your loops

What you getPage
Idiom recognition — a hand-written fill loop collapses to memset, a copy loop to memcpy; loop-invariant work is hoisted; a monotonic break flag exits early.Loop and idiom rewrites
Auto-vectorisation — an element-wise numeric loop is rewritten into SIMD form (guaranteed, before LLVM), with the @vectorize / @unroll annotations to steer the rest.SIMD and auto-vectorisation
CPU dispatch — the SIMD variant (AVX2+FMA vs SSE2 …) is chosen at process startup from the host’s feature set, one binary for every CPU.SIMD CPU dispatch

Your function calls

What you getPage
Inlining and tail calls — small bodies inline; self-tail-recursion becomes a constant-stack loop; the LLVM tail qualifier is stamped where it’s safe.Inlining, tail calls, and whole-program merge
Whole-program merge — the standard library and the hottest runtime leaves are compiled into your module before the optimiser runs, so a stdlib call isn’t a wall the optimiser stops at.Runtime and stdlib inlining

What the compiler tells LLVM

What you getPage
Stamped facts — Axle’s ownership types carry more than C’s, so the compiler stamps noalias / dereferenceable / align on fresh allocations, type-based alias metadata on field accesses, !range / !invariant.load on loads it can bound, and fast-math contraction on float ops. This is why two look-alike programs can compile very differently.What the compiler tells LLVM

The one idea behind all of it

Axle’s compiler is organised around a single contract: the analysis phase decides, and code generation only transcribes. Every shape decision — does this allocation escape? is this loop index in bounds? is this divisor non-negative? — is computed once by an analysis and recorded on the program. Code generation reads that recorded fact and emits the matching machine code; it never re-derives the decision. That is why two programs that look identical can compile very differently: what changes is not the syntax but what the analysis was able to prove about it.

Staying honest

A few things this section will not do:

  • No “faster than C” claims without a benchmark you can run. Where a number appears, it comes from the benchmark suite in the repository, at a stated optimisation level.
  • No pretending a proof is free. Some facts — a data-dependent index, a value read from an array, a non-local invariant across a whole data structure — a local compiler analysis genuinely cannot recover. Those keep their runtime check, and the pages below say so.
  • Absences are named. When a page reaches a capability the compiler does not have, it says so in as many words, and the mechanism that would have to exist for it to work is named too. Nothing here is sold as done when it is not.
  • Safety is never traded for speed. Every elimination here is gated on a proof; the fallback is always the checked path.

See also

optimisationperformancecompiler