Inlining, tail calls, and whole-program merge
Calling a function costs more than the work inside it: the CPU pushes the arguments and a return address, jumps away, runs the body, and jumps back. For a one-line accessor that overhead can dwarf the actual computation — and worse, a call is a wall the optimiser can’t see past, so everything on the far side has to be assumed to do anything.
Axle attacks both problems: it removes the call where it can, and where a call remains, it tries to make sure the optimiser can still see through it. Three mechanisms, all automatic.
Inlining small bodies
Inlining pastes a function’s body into the caller at the call
site, so x = size(p) becomes x = p.len — no jump, and the
surrounding optimiser now reasons about x as the field itself.
Every function is classified at compile time into one of four inline
hints:
inline fnkeyword, or a small body → always inline. The keyword is an explicit request and is honoured whatever the body measures, even at-O0; a body within a small statement budget is folded on the compiler’s own judgement.- A larger body → hint. The compiler asks the optimiser to prefer inlining.
- Anything larger → no hint. The body is left to LLVM’s cost model, which weighs the code-size cost of duplicating it.
- A body that must stay out of line → no-inline. A function the
compiler itself splits out, or one marked
@heat(cold), is vetoed; a veto wins over the keyword.
The trade-off is code size: paste a large body into a thousand call sites and the binary balloons, so a body nobody marked is inlined eagerly only while it stays small. The full four-tier rule, and how the hint is stamped, is in Inlining and tail calls.
Self-tail-recursion becomes a loop
A tail call is the last thing a function does — return f(args) with nothing left to compute after f returns. When a function
tail-calls itself, the recursion is really a loop in disguise, and
Axle rewrites it into one before LLVM ever sees it — so a deep
recursion runs at constant stack depth by construction, not by
optimiser goodwill:
fn fact(n : i64, acc : i64) : i64 {
if (n <= 1) { return acc; }
return fact(n - 1, acc * n);
}
fn main() : i32 {
return fact(5, 1) as i32 - 120; // 120 - 120 = 0
} Dump this with --emit=llvm and there is no recursive call left: fact is a single-block loop whose induction values are plain φ
nodes, a compare, and a back-edge. The recursion became O(1) stack
frames instead of O(depth) — a deep input that would have overflowed
the stack now runs fine.
The rewrite fires only when it’s provably equivalent (every parameter is
a by-value scalar or one the callee does not retain, the call at the
function’s top level or inside plain if branches, never inside an inner
loop or a try / synchronized / defer that has cleanup to run). The exact
conditions, and the two-phase argument update that makes it sound,
are in Inlining and tail
calls.
The LLVM tail qualifier
For a return f(args) the loop rewrite doesn’t claim — a call to a different function, or mutual recursion — the compiler stamps the
LLVM tail qualifier, which lets the optimiser reuse the caller’s
dead frame for the jump. It’s applied only when every argument is a
by-value scalar, because a pointer into the caller’s frame could be
clobbered once that frame is reused. Details and the safety argument
are in Inlining and tail
calls.
The whole program is one module
Here is the part that sets Axle apart from most compiled languages. In a typical toolchain, a call into the standard library or the language runtime is a wall: the body lives in a separately-compiled archive, so the inliner, loop-invariant hoisting, and dead-store elimination all have to assume the worst about what happens across that call.
In Axle they don’t. The standard library — the function bodies
themselves, shipped as LLVM bitcode — is merged into your module’s IR
before the optimiser runs. There is then one IR unit (your code,
the stdlib bodies, the folded runtime leaves) and one optimiser pass
over all of it. mem2reg, the inliner, LICM, dead-store elimination,
GVN, and the vectoriser all see straight through the stdlib
boundary.
This is not ThinLTO: there’s no separate link-time step and no
cross-module summary index — just a plain IR merge followed by the
ordinary pass pipeline, at the same -O level you asked for. (At -O0, a standalone dead-global pass still runs after the merge so a
debug binary doesn’t carry the whole standard library it never
calls.)
On top of the merge, the runtime functions you hit on nearly every
line of string-heavy or Shared<T>-heavy code are also compiled as
self-contained, inlinable IR and folded in as available_externally definitions. At -O 2 and above on the host target, LLVM inlines them
to a few instructions right where you call them:
$ axle build hot_strings.axle --emit=llvm -O 2 -o hot.ll
$ grep -c 'call.*@axle_string_byte_len' hot.ll
0 # the length-header read inlined; no call left Below -O 2, or when cross-compiling to a non-host target, those
leaves stay ordinary calls — the fold is a release-build win, and debug
builds keep clean stack traces and fast compiles. And because the
folded copy is only inlining metadata, the real symbol still lives in
the runtime archive: nothing changes when the fold doesn’t apply. The
full mechanism, and the reference-count / arena work that was never a
call to begin with, is in Runtime and stdlib
inlining.
Where it stops — honestly
- Large bodies aren’t force-inlined. Past the small-body threshold the decision is LLVM’s cost model, which can decline — inlining everything would bloat the binary and thrash the instruction cache.
- The self-recursion rewrite only covers direct self-calls with
scalar parameters. Mutual recursion and pointer-argument tail calls
get the
tailqualifier (a permission LLVM may or may not act on), not the guaranteed loop. - Leaf folding is a host-target,
-O ≥ 2feature. Cross-compiled or debug builds keep the plain cross-boundary call — correct, just not inlined.
See also
- Inlining and tail calls — the inline-hint rule, the self-recursion rewrite, and the
tailqualifier in full, with the IR they produce. - Runtime and stdlib inlining —
the IR merge and the
available_externallyleaf folding. - Escape analysis and promotion — inlining sometimes exposes a fresh-then-discarded allocation the stack promoter can then catch.
- Optimisations overview — the complete census.