Axle v0.14.1

Inlining, tail calls, and whole-program merge

Calling a function costs more than the work inside it: the CPU pushes the arguments and a return address, jumps away, runs the body, and jumps back. For a one-line accessor that overhead can dwarf the actual computation — and worse, a call is a wall the optimiser can’t see past, so everything on the far side has to be assumed to do anything.

Axle attacks both problems: it removes the call where it can, and where a call remains, it tries to make sure the optimiser can still see through it. Three mechanisms, all automatic.

Inlining small bodies

Inlining pastes a function’s body into the caller at the call site, so x = size(p) becomes x = p.len — no jump, and the surrounding optimiser now reasons about x as the field itself. Every function is classified at compile time into one of four inline hints:

  • inline fn keyword, or a small body → always inline. The keyword is an explicit request and is honoured whatever the body measures, even at -O0; a body within a small statement budget is folded on the compiler’s own judgement.
  • A larger body → hint. The compiler asks the optimiser to prefer inlining.
  • Anything larger → no hint. The body is left to LLVM’s cost model, which weighs the code-size cost of duplicating it.
  • A body that must stay out of line → no-inline. A function the compiler itself splits out, or one marked @heat(cold), is vetoed; a veto wins over the keyword.

The trade-off is code size: paste a large body into a thousand call sites and the binary balloons, so a body nobody marked is inlined eagerly only while it stays small. The full four-tier rule, and how the hint is stamped, is in Inlining and tail calls.

Self-tail-recursion becomes a loop

A tail call is the last thing a function does — return f(args) with nothing left to compute after f returns. When a function tail-calls itself, the recursion is really a loop in disguise, and Axle rewrites it into one before LLVM ever sees it — so a deep recursion runs at constant stack depth by construction, not by optimiser goodwill:

fn fact(n : i64, acc : i64) : i64 {
    if (n <= 1) { return acc; }
    return fact(n - 1, acc * n);
}

fn main() : i32 {
    return fact(5, 1) as i32 - 120;   // 120 - 120 = 0
}

Dump this with --emit=llvm and there is no recursive call left: fact is a single-block loop whose induction values are plain φ nodes, a compare, and a back-edge. The recursion became O(1) stack frames instead of O(depth) — a deep input that would have overflowed the stack now runs fine.

The rewrite fires only when it’s provably equivalent (every parameter is a by-value scalar or one the callee does not retain, the call at the function’s top level or inside plain if branches, never inside an inner loop or a try / synchronized / defer that has cleanup to run). The exact conditions, and the two-phase argument update that makes it sound, are in Inlining and tail calls.

The LLVM tail qualifier

For a return f(args) the loop rewrite doesn’t claim — a call to a different function, or mutual recursion — the compiler stamps the LLVM tail qualifier, which lets the optimiser reuse the caller’s dead frame for the jump. It’s applied only when every argument is a by-value scalar, because a pointer into the caller’s frame could be clobbered once that frame is reused. Details and the safety argument are in Inlining and tail calls.

The whole program is one module

Here is the part that sets Axle apart from most compiled languages. In a typical toolchain, a call into the standard library or the language runtime is a wall: the body lives in a separately-compiled archive, so the inliner, loop-invariant hoisting, and dead-store elimination all have to assume the worst about what happens across that call.

In Axle they don’t. The standard library — the function bodies themselves, shipped as LLVM bitcode — is merged into your module’s IR before the optimiser runs. There is then one IR unit (your code, the stdlib bodies, the folded runtime leaves) and one optimiser pass over all of it. mem2reg, the inliner, LICM, dead-store elimination, GVN, and the vectoriser all see straight through the stdlib boundary.

This is not ThinLTO: there’s no separate link-time step and no cross-module summary index — just a plain IR merge followed by the ordinary pass pipeline, at the same -O level you asked for. (At -O0, a standalone dead-global pass still runs after the merge so a debug binary doesn’t carry the whole standard library it never calls.)

On top of the merge, the runtime functions you hit on nearly every line of string-heavy or Shared<T>-heavy code are also compiled as self-contained, inlinable IR and folded in as available_externally definitions. At -O 2 and above on the host target, LLVM inlines them to a few instructions right where you call them:

$ axle build hot_strings.axle --emit=llvm -O 2 -o hot.ll
$ grep -c 'call.*@axle_string_byte_len' hot.ll
0                       # the length-header read inlined; no call left

Below -O 2, or when cross-compiling to a non-host target, those leaves stay ordinary calls — the fold is a release-build win, and debug builds keep clean stack traces and fast compiles. And because the folded copy is only inlining metadata, the real symbol still lives in the runtime archive: nothing changes when the fold doesn’t apply. The full mechanism, and the reference-count / arena work that was never a call to begin with, is in Runtime and stdlib inlining.

Where it stops — honestly

  • Large bodies aren’t force-inlined. Past the small-body threshold the decision is LLVM’s cost model, which can decline — inlining everything would bloat the binary and thrash the instruction cache.
  • The self-recursion rewrite only covers direct self-calls with scalar parameters. Mutual recursion and pointer-argument tail calls get the tail qualifier (a permission LLVM may or may not act on), not the guaranteed loop.
  • Leaf folding is a host-target, -O ≥ 2 feature. Cross-compiled or debug builds keep the plain cross-boundary call — correct, just not inlined.

See also

optimisationinliningtail-callsltoruntime