Axle v0.14.1

Runtime and stdlib inlining

What you’d notice

In most compiled languages a call into the standard library or the language runtime is a wall the optimizer stops at: the body lives in a separately-compiled archive, so the inliner, loop-invariant hoisting, and dead-store elimination all have to assume the worst about what happens on the other side of that call.

In Axle they don’t. The standard library and the hottest runtime leaves are compiled into your module, so the optimizer sees through them and treats them like code you wrote yourself. Two mechanisms produce that, and both are automatic.

One module, one optimizer pass

The standard library is shipped as LLVM bitcode. When your program is compiled, that bitcode (and the runtime’s inlinable leaves, below) is merged into your module’s IR before LLVM’s optimizer runs. There is then a single IR unit — your code, the stdlib bodies, the folded runtime leaves — and one default<O…> pipeline pass over all of it. mem2reg, the inliner, LICM, DSE, GVN and the vectorizer all see straight through the stdlib boundary.

This is not ThinLTO. There is no separate link-time optimization step and no cross-module summary index — just a plain IR merge followed by the ordinary pass pipeline, at the same -O level you asked for. At -O0, a standalone dead-global pass still runs after the merge so a debug binary doesn’t carry the whole standard library it never calls.

Hot runtime leaves fold into the call site

On top of the merge, the runtime functions you hit on nearly every line of string-heavy or Shared<T>-heavy code are compiled a second way: as self-contained, inlinable LLVM IR, folded into your module as available_externally definitions. At -O≥2 on the host target, LLVM inlines them to a handful of instructions right where you call them, instead of a jump across the runtime boundary. Six entry points, from four source modules, fold:

LeafInlines to
axle_string_byte_lenthe O(1) string-length header read — a load + a mask
axle_string_byte_atthe bounded read of one byte of a string
axle_string_freethe drop of an owned string
axle_shared_freethe last-reference free of a Shared<T>
axle_shared_downgradeShared<T> → Weak<T>
axle_weak_upgradeWeak<T> → Shared<T> \| null

Because these are available_externally, the folded body is only a copy for the inliner to consume: the real symbol still lives in the runtime archive, so nothing changes when the fold doesn’t apply. And it doesn’t always apply — at -O0, or when cross-compiling to a non-host target, the leaves stay ordinary calls. That’s deliberate: debug builds stay fast to compile and keep clean stack traces, and the inlining is a release-build win. (One leaf, axle_weak_decrement, is also kept out-of-line on purpose — it would otherwise collide with a sibling under LLVM’s function-merging pass.)

Watch it happen

Emit the IR for a string-heavy program at -O2 and look for the length leaf — it isn’t there as a call, because it inlined:

$ axle build hot_strings.axle --emit=llvm -O 2 -o hot.ll
$ grep -c 'call.*@axle_string_byte_len' hot.ll
0

Emit the same file at -O0 and the call reappears — the fold is gated on optimization, exactly as described above.

Refcount and arena: inline from the start

Some runtime work was never a call to begin with. The Shared<T> reference-count increment and decrement are emitted directly as a guarded atomicrmw — or as a plain load/add/store when the compiler has proven the class never crosses a thread (see Refcount and transitions). The arena bump-allocation fast path is emitted inline too, calling into the runtime only on the rare slow path where a fresh chunk is needed. So the reference-counting and arena bookkeeping that runs constantly is already a branch and a couple of instructions, not a function call.

The result is that “the runtime” and “the standard library” are not a black box the optimizer has to route around. They are ordinary IR in your module, optimized together with your code.

See also

internalscompileroptimizationllvmruntimestdlib