Runtime and stdlib inlining
What you’d notice
In most compiled languages a call into the standard library or the language runtime is a wall the optimizer stops at: the body lives in a separately-compiled archive, so the inliner, loop-invariant hoisting, and dead-store elimination all have to assume the worst about what happens on the other side of that call.
In Axle they don’t. The standard library and the hottest runtime leaves are compiled into your module, so the optimizer sees through them and treats them like code you wrote yourself. Two mechanisms produce that, and both are automatic.
One module, one optimizer pass
The standard library is shipped as LLVM bitcode. When your program is
compiled, that bitcode (and the
runtime’s inlinable leaves, below) is merged into your module’s IR
before LLVM’s optimizer runs. There is then a single IR unit — your
code, the stdlib bodies, the folded runtime leaves — and one default<O…> pipeline pass over all of it. mem2reg, the inliner,
LICM, DSE, GVN and the vectorizer all see straight through the stdlib
boundary.
This is not ThinLTO. There is no separate link-time optimization
step and no cross-module summary index — just a plain IR merge followed
by the ordinary pass pipeline, at the same -O level you asked for. At -O0, a standalone dead-global pass still runs after the merge so a
debug binary doesn’t carry the whole standard library it never calls.
Hot runtime leaves fold into the call site
On top of the merge, the runtime functions you hit on nearly every line
of string-heavy or Shared<T>-heavy code are compiled a second way:
as self-contained, inlinable LLVM IR, folded into your module as available_externally definitions. At -O≥2 on the host target, LLVM
inlines them to a handful of instructions right where you call them,
instead of a jump across the runtime boundary. Six entry points, from four
source modules, fold:
| Leaf | Inlines to |
|---|---|
axle_string_byte_len | the O(1) string-length header read — a load + a mask |
axle_string_byte_at | the bounded read of one byte of a string |
axle_string_free | the drop of an owned string |
axle_shared_free | the last-reference free of a Shared<T> |
axle_shared_downgrade | Shared<T> → Weak<T> |
axle_weak_upgrade | Weak<T> → Shared<T> \| null |
Because these are available_externally, the folded body is only a
copy for the inliner to consume: the real symbol still lives in the
runtime archive, so nothing changes when the fold doesn’t apply. And it
doesn’t always apply — at -O0, or when cross-compiling to a non-host
target, the leaves stay ordinary calls. That’s deliberate: debug builds
stay fast to compile and keep clean stack traces, and the inlining is a
release-build win. (One leaf, axle_weak_decrement, is also kept
out-of-line on purpose — it would otherwise collide with a sibling under
LLVM’s function-merging pass.)
Watch it happen
Emit the IR for a string-heavy program at -O2 and look for the length
leaf — it isn’t there as a call, because it inlined:
$ axle build hot_strings.axle --emit=llvm -O 2 -o hot.ll
$ grep -c 'call.*@axle_string_byte_len' hot.ll
0 Emit the same file at -O0 and the call reappears — the fold is gated
on optimization, exactly as described above.
Refcount and arena: inline from the start
Some runtime work was never a call to begin with. The Shared<T> reference-count increment and decrement are emitted directly as a
guarded atomicrmw — or as a plain load/add/store when the
compiler has proven the class never crosses a thread (see Refcount and transitions). The
arena bump-allocation fast path is emitted inline too, calling into the
runtime only on the rare slow path where a fresh chunk is needed. So the
reference-counting and arena bookkeeping that runs constantly is already
a branch and a couple of instructions, not a function call.
The result is that “the runtime” and “the standard library” are not a black box the optimizer has to route around. They are ordinary IR in your module, optimized together with your code.
See also
- Inlining, tail calls, and whole-program merge — the user-facing view of this merge, in the optimisations census.
- Refcount and transitions — the atomic-vs-plain reference count these leaves sit alongside.
- Memory model — where the arena and
Shared<T>tiers come from. - Inlining and tail calls — how the same inliner treats your function bodies.
- Loop and idiom rewrites — the sema-level rewrites that run before this pipeline.