Optimisations
Axle is a statically-typed language that compiles to native code through LLVM, with compiler-inferred memory management — no garbage collector, and no lifetime annotations to write. A lot of the work that would otherwise be your job — deciding what goes on the stack, removing a bounds check the surrounding code already guarantees, turning a divide into a multiply, vectorising a numeric loop — the compiler does ahead of time and stamps onto the program.
This section is the complete census of that work: what the compiler optimises, how much it can prove, and — just as important — where it stops. Axle does not claim to make every program fast; it claims to remove the overheads it can prove are unnecessary, and to leave a runtime check in place the moment it cannot. A removed check that shouldn’t have been removed is a memory-safety bug, so every elision below is conservative: when in doubt, the check stays.
The full map
Every optimisation the compiler performs, grouped by what it acts on. Pages under this chapter are the optimisation-lens view; pages under Memory, Vectorisation, and Compiler internals are the deep dives — this chapter links to them rather than repeating them.
Where your objects live
| What you get | Page |
|---|---|
The heap is the last resort. new T(...) that doesn’t escape becomes a stack slot (no malloc); a short-lived object graph shares an arena (bulk-freed at function exit); only a genuinely shared value pays a reference count; the raw heap is used only when nothing cheaper is provably safe. | Memory placement |
| Escape analysis — the pass that proves which tier each allocation may use, and the eight routing strategies it stamps. | Escape analysis and promotion |
Shared<T> reference counting — the atomic-vs-plain refcount, the 16-byte header, and where inc/dec go non-atomic. | Refcount and transitions |
Where your indexes are checked
| What you get | Page |
|---|---|
Bounds-check elimination — array and container index checks are removed when the compiler proves the index is in range, including across a chain of comparisons (i < n, n ≤ len), across a function call, and from how a container was built. | Bounds-check elimination |
Per-value range elision — the sibling interval analysis that proves a single value non-negative or in range, which also folds list.get(i) ?? fallback down to a bare load. | Loop and idiom rewrites |
Conditional-fault elision — when the compiler proves a container access is in range, the whole checked-access idiom folds away: recv.get(i) ?? d loses its nullable box and ??, and a recv.set(i, v) whose out-of-range case would throw loses the throw and its exception poll — a bare load or store remains. Rides the same bounds proof (including the loop-carried and cross-call chains). | Bounds-check elimination |
Your arithmetic
| What you get | Page |
|---|---|
Strength reduction — constant divide / modulo become a multiply-shift (even at -O0); power-of-two / and % become a shift and a mask; a loop-invariant runtime divisor is turned into a hoisted multiply; x * 8 becomes x << 3; the divide-by-zero guard is dropped when the divisor is provably non-zero. | Arithmetic strength reduction |
Peepholes, CSE, abs-after-modulo, dead-code elimination — the small local rewrites that clean up arithmetic before the backend sees it. | Loop and idiom rewrites |
Your loops
| What you get | Page |
|---|---|
Idiom recognition — a hand-written fill loop collapses to memset, a copy loop to memcpy; loop-invariant work is hoisted; a monotonic break flag exits early. | Loop and idiom rewrites |
Auto-vectorisation — an element-wise numeric loop is rewritten into SIMD form (guaranteed, before LLVM), with the @vectorize / @unroll annotations to steer the rest. | SIMD and auto-vectorisation |
| CPU dispatch — the SIMD variant (AVX2+FMA vs SSE2 …) is chosen at process startup from the host’s feature set, one binary for every CPU. | SIMD CPU dispatch |
Your function calls
| What you get | Page |
|---|---|
Inlining and tail calls — small bodies inline; self-tail-recursion becomes a constant-stack loop; the LLVM tail qualifier is stamped where it’s safe. | Inlining, tail calls, and whole-program merge |
| Whole-program merge — the standard library and the hottest runtime leaves are compiled into your module before the optimiser runs, so a stdlib call isn’t a wall the optimiser stops at. | Runtime and stdlib inlining |
What the compiler tells LLVM
| What you get | Page |
|---|---|
Stamped facts — Axle’s ownership types carry more than C’s, so the compiler stamps noalias / dereferenceable / align on fresh allocations, type-based alias metadata on field accesses, !range / !invariant.load on loads it can bound, and fast-math contraction on float ops. This is why two look-alike programs can compile very differently. | What the compiler tells LLVM |
The one idea behind all of it
Axle’s compiler is organised around a single contract: the analysis phase decides, and code generation only transcribes. Every shape decision — does this allocation escape? is this loop index in bounds? is this divisor non-negative? — is computed once by an analysis and recorded on the program. Code generation reads that recorded fact and emits the matching machine code; it never re-derives the decision. That is why two programs that look identical can compile very differently: what changes is not the syntax but what the analysis was able to prove about it.
Staying honest
A few things this section will not do:
- No “faster than C” claims without a benchmark you can run. Where a number appears, it comes from the benchmark suite in the repository, at a stated optimisation level.
- No pretending a proof is free. Some facts — a data-dependent index, a value read from an array, a non-local invariant across a whole data structure — a local compiler analysis genuinely cannot recover. Those keep their runtime check, and the pages below say so.
- Absences are named. When a page reaches a capability the compiler does not have, it says so in as many words, and the mechanism that would have to exist for it to work is named too. Nothing here is sold as done when it is not.
- Safety is never traded for speed. Every elimination here is gated on a proof; the fallback is always the checked path.
See also
- Memory model — the tier rules the memory optimisations enforce, from a user’s point of view.
- SIMD and auto-vectorisation — the user-facing surface of vectorisation.
- Compiler internals — the machinery behind these guarantees, for the curious.
- Concept index — every concept on one page.