Memory placement — the heap is the last resort
The single biggest thing the Axle compiler does for your program’s
speed happens before a line of arithmetic runs: it decides where
each object lives. In most languages new T(...) means “call the
allocator” — a search through the heap on the way in, a free on the
way out, and pressure on the cache the whole time. In Axle, new T(...) means “put this object in the cheapest place its
lifetime allows”, and the heap is only one of four answers — the
one the compiler reaches for last.
You never annotate this. You write new T(...) everywhere, and an
analysis proves the tier. This page is the optimisation-lens view;
the full model — with the diagrams, the ownership keywords, and the
runtime guarantees — is in the Memory model chapter, and the analysis that drives it in Escape analysis and
promotion.
The ladder: stack > arena > refcount > heap
Every new T(...) lands in exactly one of four tiers. They form a
ladder from cheapest and least flexible to most flexible and most
expensive, and the compiler always picks the cheapest rung the
object’s lifetime allows, climbing only when it must:
cheapest ───────────────────────────────────────► most flexible
Stack Arena Refcount Heap (owned)
never leaves leaves the genuinely shared outlives the
the function scope, not / crosses a function, one
the function thread owner
───────── ───────── ───────── ─────────
frame pop bulk reset atomic count one malloc /
(free) (one add) (16-byte header) one free - Stack — the object provably never escapes the function that
created it. Allocating costs nothing: the storage is part of the
function’s frame, reclaimed the instant the function returns. No
malloc, nofree, no fragmentation. Even inside a tight loop the slot is materialised once and overwritten in place. - Arena — the object outlives its immediate block but not its function. It’s placed in a per-function bump allocator, and the whole arena is reset in one operation when the function returns.
- Refcount (
Shared<T>) — the one tier you opt into, for a value that is genuinely shared or crosses a thread. It pays a reference count per copy. - Heap (owned) — the object outlives its function with a single
owner. This is the classic
malloc/free, and Axle uses it only when the three cheaper tiers are provably unavailable.
The heap is correct for every case — which is exactly why it’s the default the analysis has to beat. Each cheaper tier is an upgrade applied only when the compiler can prove it safe.
What is an arena?
The arena (or bump allocator) is the tier most languages don’t have, and it fits a shape that shows up constantly: a burst of work allocates a lot of short-lived objects, then throws them all away at once — exactly what handling one request, or one loop iteration, looks like.
An arena is a block of memory with a single cursor. Allocating is “hand back the cursor, then move it forward by the object’s size” — one addition, no search. Freeing is the trick: you never free objects one by one. When the burst ends, you move the cursor back to where it started, and every object is reclaimed at once, in time proportional to the number of chunks — not the number of objects:
allocate A, B, C → reset (one assignment)
┌───┬───┬─────┬───────┐ ┌──────────────────────┐
│ A │ B │ C │ …free │ │ all free │
└───┴───┴─────┴───────┘ └──────────────────────┘
▲ cursor ▲ cursor back to start Axle uses one arena per function call, reset at return — and per loop iteration when it can prove the body’s allocations don’t outlive the iteration, so a long-running loop’s peak memory stays bounded by a single pass. Because the reset is bulk, arena objects never run destructors, which is why a class with a destructor (yours, or one synthesised for its owned fields) is kept off the arena.
Watch it happen
You don’t have to guess which tier you got. axle build file.axle --emit=hir prints each allocation with a tier keyword in front
of it. Here one class is used two ways:
class Point {
pub x : i64;
pub y : i64;
constructor(x : i64, y : i64) {
self.x = x;
self.y = y;
}
}
// p never leaves this function → STACK, no malloc ever runs.
fn local_use() : i64 {
let p : Point = new Point(3, 4);
return p.x + p.y;
}
fn main() : i32 {
return local_use() as i32 - 7; // 3 + 4 - 7 = 0
} Emitting the HIR shows the promotion in the keyword — a bare new would be the heap, stack is the promoted form:
$ axle build point.axle --emit=hir
let p : Point* = stack Point(3 as i64, 4 as i64); Return the object instead of consuming it locally, and the same allocation demotes to the heap — the value now has to survive the frame pop:
class Point {
pub x : i64;
pub y : i64;
constructor(x : i64, y : i64) {
self.x = x;
self.y = y;
}
}
fn hand_out() : Point {
let p : Point = new Point(3, 4); // HIR: new Point(3, 4) → heap malloc
return p; // one owner, freed at the caller's drop
} The bodies are identical except for the last line. That single difference — does the object escape? — is the whole decision.
Counting what actually ran
--emit=hir shows the tier a site got; it says nothing about how many times
that site ran, and a loop body that allocates once per iteration reads the same
as a one-off. axle build file.axle --alloc-stats answers the other half: the
built program counts every allocation it makes — the standard library’s and the
runtime’s included, not just yours — and prints the tally when it exits.
$ axle build word_count.axle --alloc-stats -o word_count && ./word_count
axle-alloc-stats: allocs=3000 bytes=32000 frees=3000 reallocs=0 mean_bytes=10
axle-alloc-stats: hist 8=2000 16=1000 The second line is a histogram of requested sizes by power-of-two band, so 8=2000 means two thousand requests between 8 and 15 bytes — usually the fastest
way to see that a program’s allocation profile is one shape repeated, which is
the shape the runtime recycles rather than re-allocates.
Two caveats. The counts are exact but the duration is not your program’s: each allocation pays a call and two atomics, so measure time with the flag off and shape with it on. And a build without the flag costs nothing at all — the counting is compiled in, not switched on at runtime.
Why no GC, and no lifetime annotations
Two well-known ways to manage the heap automatically, and why Axle takes neither:
- A garbage collector frees memory for you, but pauses your program at unpredictable moments to scan — unacceptable for a server that must answer within a few milliseconds every time. Axle has no GC and no pauses; cleanup is deterministic, tied to where the object’s lifetime ends.
- A borrow checker (as in Rust) proves safety at compile time
with no runtime cost, but makes you write lifetime annotations and
satisfy the checker. Axle asks for no lifetime annotations. The
only memory annotation you can write is
Shared<T>, for the rare genuinely-shared case.
The bet is that a compiler can infer the cheapest safe storage for most objects, and fall back to a small explicit opt-in only when sharing is real. You get stack-allocation speed where it’s provable, deterministic cleanup everywhere, and no annotations in the common case.
The cheapest tier is no allocation at all
The ladder answers where an object goes. There is a rung below it: an owned child that lives and dies with its parent does not need a home of its own — it can live inside the parent.
class Inner { value : i32 }
class Outer {
a : Inner;
constructor() { self.a = new Inner(); } // one allocation, not two
} Written this way, Outer is two objects and two mallocs: one for the parent,
one for the child, with a pointer between them. The compiler folds the child’s
storage into the parent’s, so there is one allocation, one free, and o.a.value becomes two constant offsets instead of an offset, a dependent load, and another
offset. Nothing about the source changes — o.a still names the child, and you
can still pass it to a function.
It applies when the compiler can prove the child never outlives its parent:
- every write to the field is a fresh
new— never a pointer handed in from elsewhere, and never a compound assignment; - the child is never stored anywhere else, returned, or thrown;
- the field is not
Shared<T>(a refcounted handle is shared by definition); - neither class is part of an inheritance chain, and neither is an
extern "C"layout (below).
Passing o.a to a function is fine as long as the callee only borrows it: a
parameter the compiler has already proven does not capture its argument and
frees nothing. That is a lending, not an escape.
Fields sit where alignment wants them
Independently of the fold, fields are ordered by descending alignment rather
than by the order you wrote them, so the padding between them disappears. A
struct of { i8, i64, i8 } occupies 24 bytes laid out as written and 16 laid
out by alignment.
This is invisible from Axle — p.flag names the field you declared, whichever
slot it landed in. It is not invisible across an FFI boundary, which is what extern "C" struct and extern "C" class are for: they declare that the C layout is
the contract, and the compiler then touches neither the field order nor the
fold. See FFI and interop.
Where it stops — honestly
The analysis is conservative: a false “it escapes” only costs a malloc that could have been avoided, but a false “it’s local” would
free memory still in use. So doubt always resolves to the safe, heap
answer. In particular:
- Escape through
return/throw/spawn/ a stored field demotes an object to the heap — it must outlive the frame. - A destructor-bearing class never goes on the arena (the bulk reset can’t run per-object destructors), even when escape analysis would otherwise allow it. It routes to heap or refcount instead.
- A value passed to a function that might retain it is assumed to escape unless the parameter’s contract promises otherwise.
Shared<T>is never silently downgraded to owned heap. Once you ask for shared ownership, an alias the compiler can’t see could outlive any single owner, so the refcount stays — dropping it would risk a use-after-free.
In every one of these cases the object is placed safely and freed deterministically; you lose a cheaper tier, never correctness.
See also
- Memory model — the full four-tier model,
ownership keywords (
mut/own/Shared<T>),defer, and the runtime guarantees. - Escape analysis and promotion — the analysis that proves each tier, and the eight allocator routing strategies it stamps.
Shared<T>reference counting — the one opt-in tier, in full.- Optimisations overview — the complete census.
- Concept index — every memory concept on one page.