Axle v0.14.1

Escape analysis and promotion

new T(...) always looks like a heap allocation in the source. The compiler decides where it actually lives — and most of the time, that’s not the heap.

What “escape” means

An object escapes a function when something outside that function can still reach it after the function returns. That’s the one question that decides whether the object can go on the cheap stack tier: a stack slot vanishes when the frame pops, so anything that must outlive the call cannot live there.

An object escapes when it is:

   returned            stored in an        sent to a
   from the fn         escaping owner      thread (spawn)
   ┌────────┐          ┌────────┐          ┌────────┐
   │  new T │──return─►│  new T │──field──►│  new T │──spawn──►
   └────────┘          └────────┘          └────────┘  ▲
        ▲                   ▲                           │
        │                   │                    another thread
   caller keeps it    outer object keeps    reads it after this
   after we return    it after we return    fn returns

If none of those can happen — the object is created, used, and dropped entirely inside the function — then it does not escape, and the compiler puts it on the stack. Escape analysis is the pass that answers this question for every allocation site, by following where each new T(...) flows.

   does the object escape its declaring function?
        │
        ├─ no, and never leaves its scope      → STACK  (alloca)
        ├─ no, but leaves its scope in-frame   → ARENA  (bump + reset)
        └─ yes (return / throw / field / spawn) → HEAP   (malloc / shared)

Escape analysis is conservative: when it cannot prove an object stays local, it assumes it escapes and keeps it on the heap. A false “it escapes” only costs a malloc that was avoidable; a false “it’s local” would free memory still in use — so the analysis always errs toward the safe, heap-resident answer.

Two-pass system

A dedicated analysis examines every allocation site and asks two questions :

  1. Does the value leave its declaring function ? If no → rewrite to a stack slot in the entry frame. The frame pop reclaims the storage ; no malloc ever runs. This pass — the stack promoter — runs first because it produces the smallest IR shape for every other pass downstream.
  2. Does it leave the request scope ? If no but it leaves the declaring fn → route to the per-function arena. A single reset at function exit releases every object allocated in that scope in O(chunks) time, regardless of how many allocations happened in between. The arena pass walks the alloc-dependency graph (constructor args, struct fields, array elements, field writes through a local receiver) so an alloc whose only escape paths terminate in scope-only positions (local field writes, transient temporaries) is still promoted — the dependency closure is what answers “does this transitively outlive the frame”.

Allocations that fail both checks (returned, thrown, stored into a field of an escaping object, passed to a function that retains them, spawned across a thread boundary) stay on the heap — plain malloc for owned bindings, the shared allocator (with its 16-byte refcount header) for Shared<T>.

What counts as “escape”

  • Returned from the function.
  • Stored into a field of a value that itself escapes.
  • Thrown from the function.
  • Assigned through an index, deref, or field write whose receiver escapes.
  • Passed to a callee whose parameter signature doesn’t promise non-retention (the callee captures into its own escape set).

The two-pass system is conservative — when in doubt, it stays on the heap. False positives (would have been safe to stack-promote, but went heap) are common at refactor boundaries ; false negatives (escape that the analysis missed) would be unsound, so the rules err strictly.

Watch it happen

One class, two call sites, two tiers — decided purely by whether the object escapes. Build with --emit=hir and read the keyword in front of the allocation (stack vs a bare new for heap):

class Box {
    v : i64;
    constructor(v : i64) { self.v = v; }
}

// Does NOT escape: b is read and dropped inside the function.
fn sum_local() : i64 {
    let b : Box = new Box(41);   // HIR: stack Box(41)   → alloca, no malloc
    return b.v + 1;
}

// ESCAPES via return: the value must survive the frame pop.
fn hand_out() : Box {
    let b : Box = new Box(41);   // HIR: new Box(41)      → heap malloc
    return b;                    // one owner; freed at the caller's drop
}

The bodies are identical except for the last line. Returning b is an escape, so hand_out cannot use the stack; sum_local keeps b entirely local and pays nothing. You never annotate this — the analysis reads the return and routes accordingly.

Restrictions

Classes with a user-declared destructor stay off the arena. The arena release is a bulk reset — running per-object destructors during reset would defeat the performance model, and a user destructor has observable side effects the reset would skip. Such classes are routed to heap or shared instead.

A class that only owns heap fields (and so gets a compiler- synthesised destructor that just frees those fields) follows one rule: a child’s tier follows its container’s. When a let p = new Container(...) provably doesn’t escape its frame, the compiler inlines the constructor at the allocation site, promotes the container to the stack, promotes each single-write owned child onto the stack with it, and drops the synthesised destructor for that instance — the frame pop reclaims the whole graph. So class Pair { a : Wrap; b : Wrap; } used locally puts the Pair and both Wraps on the stack with zero malloc; the optimiser then dissolves the entire graph into registers (there’s no allocation left to optimise). A child that itself owns heap, a reassigned field, or an escaping container keeps the ordinary heap path — and when the container escapes, its children stay on the heap with it, freed by the destructor.

   let p = new Pair(...)   provably local          p returned / stored → escapes
   ┌────────────────────────────────────┐          ┌──────────────────────────┐
   │ STACK frame                         │          │ HEAP                     │
   │   Pair ──owns──► Wrap a             │          │   Pair ──owns──► Wrap a  │
   │        └─owns──► Wrap b             │          │        └─owns──► Wrap b  │
   └────────────────────────────────────┘          └──────────────────────────┘
     child tier FOLLOWS the container:                the whole graph is heap;
     all three on the stack, zero malloc,             the synthesised destructor
     synthesised destructor dropped                   frees each owned child

Stdlib handle classes (registry-backed integer handles like ReentrantLock, BoundedChannel, ThreadPoolExecutor) also bypass the arena : the handle integer is registered with a separate thread-local table that owns the actual underlying state.

The eight array-routing strategies

Once the analysis has picked the tier, every array allocation in the program is tagged with one of these strategies. An object allocation (new T(...)) is routed by the same analysis but carries its tier on the expression itself — StackAlloc / ArenaAlloc / HeapAlloc / SharedAlloc — not through this enum, which answers for T[N](v) and malloc<T>(n) sites. The decision is made once at compile time; the code emitter mechanically picks the matching call shape and never re-inspects allocation sizes, types, or escape state.

StrategyWhen it’s pickedWhat gets emitted
StackA constant-count T[N] below the 32 KiB stack-allocation thresholdA single alloca [N x T] in the function’s entry block plus an optional zero-init memset
callocConstant-count, both count and sizeof(T) known at compile time, product below the huge-page threshold, and the buffer is read before being fully overwrittenlibc::calloc(count, size) — zero-fill in one call
Huge callocSame as calloc but the total allocation is ≥ 2 MiB__axle_runtime_huge_calloc(count, size) — mmap + MADV_HUGEPAGE + zero-fill
mallocDynamic-count malloc<T>(n), or constant-count above the stack threshold but below huge-page, or any allocation where the compiler proved the buffer is fully overwritten before the first read (so the zero-fill is skipped)libc::malloc(N) — uninitialised
Huge mallocAllocation ≥ 2 MiB with the zero-init proven unnecessary__axle_runtime_huge_alloc(size) — mmap + MADV_HUGEPAGE, no zero-fill
Arena bumpnew T(...) that escapes the local but not the per-function arena scope, and the class has no user-declared destructorInline bump of the per-thread @axle_arena_cursor, dropping to axle_arena_alloc_slow only when the active chunk is full
Shared allocatornew shared T(...) (or any Shared<T> binding)axle_shared_alloc(size) — 16-byte refcount header + 16-aligned payload
Stdlib FFI handlenew ClassName(args) on a stdlib handle class (ReentrantLock, BoundedChannel, ThreadPoolExecutor, …)__axle_stdlib_<module>_<Class>_new(args) — e.g. __axle_stdlib_io_File_new — returns an i64 handle into a thread-local registry; no malloc runs on the Axle side

The most conservative strategy (malloc) is the default — it’s always correct, never always fastest. Every other strategy is an upgrade applied when the analysis can prove the upgrade is safe.

Bookend behaviour you can rely on

  • Every let x = new T(...) whose receiver provably doesn’t escape becomes a stack slot. Tight loops allocate at zero runtime cost — the slot is materialised once at the function entry block and overwritten in place on each iteration.
  • Every arena-eligible allocation is freed deterministically at function return. Each arena-using frame snapshots the bump position it inherited with axle_arena_mark() in its entry block and rewinds to exactly that with axle_arena_reset_to_mark(mark) on every return path (including unwind / throw exits) — it never releases the whole pool, so a callee returning to a still-live caller leaves the caller’s arena objects intact (that nesting contract is what lets one arena function call another). Peak arena RSS for a loop body whose allocations don’t escape the iteration stays bounded by one iteration’s footprint — the compiler wraps the body with the same mark / rewind pair when it proves non-escape, so the per-iteration allocations are reclaimed at the loop back-edge rather than at function exit.
  • Heap allocations above ~2 MiB route through the runtime’s huge-page allocator (mmap + MADV_HUGEPAGE) instead of plain malloc. The compile-time decision and the runtime allocator share the same threshold constant, so they can’t drift.
  • The Shared<T> 16-byte refcount header is allocated contiguously with its payload (one allocation, not two), so a shared object behaves like a plain pointer through the rest of the pipeline — header offsets are baked into the access patterns at compile time.

See also

memoryescape-analysisstackoptimization