Axle v0.14.1

Refcount and transitions

This is the deep-dive companion to Shared<T> reference counting — read that page first for what a reference count is and how to use Shared<T>. Here we open the box: what’s actually allocated, how the atomics are ordered, and how the explicit Shared::wrap / Shared::try_unwrap transitions lower to IR.

Both transitions are fully implemented: Shared::wrap emits the shared allocation and copies the owned payload in; Shared::try_unwrap emits the speculative atomic decrement and the unique-branch recovery.

Header layout

Every Shared<T> allocation is a single contiguous block :

┌──────────────────────────────────────────┬─────────────────────┐
│  16-byte refcount header                  │  payload : T        │
│  ┌───────────┬───────────┬──────────────┐ │                     │
│  │ AtomicU32 │ AtomicU32 │ u64          │ │                     │
│  │ strong    │ weak      │ payload size │ │                     │
│  └───────────┴───────────┴──────────────┘ │                     │
└──────────────────────────────────────────┴─────────────────────┘
^ the shared allocator returns this address
                                             ^ the compiler steps past
                                               the header to the payload

The payload size slot lets the matching free reconstruct the exact Layout Rust’s dealloc needs — no size threading required from the call site. The weak count drives the std::sync::Arc / Weak lifecycle: the strong references hold one implicit weak (so it starts at 1), the payload is dropped at strong 0, and the block is freed at weak 0.

Atomic ordering

The strong and weak counts are each an AtomicU32 (four billion references is unreachable before OOM, so carrying both still fits the 16-byte header). Every operation uses an explicit, deliberate ordering — never SeqCst; the initial store and the clone-side increments are the Relaxed ops, their happens-before carried by the drop side.

OpOrderWhy
Init (shared allocator returns)Relaxed store of 1The block is not yet published when the counts are written, so no ordering is owed — exactly Arc::new
inc (clone)Relaxed fetch_add(1)The caller already holds a live strong reference, so the block is valid; the cross-thread happens-before is carried by the Release dec and the last-drop Acquire fence — exactly Arc::clone’s discipline
dec (drop, count > 1)Release fetch_sub(1)Publishes the dropper’s writes ; pairs with the last-drop fence so the eventual destructor sees them
dec (last drop, count == 1)Acquire fence after the Release fetch_subThe destructor about to run must observe every write any prior holder made — even on cores that haven’t seen the dec yet

Both sides mirror Rust’s Arc<T> exactly : a Relaxed clone with an abort-on-overflow branch, a Release drop, and the Acquire fence on the last drop.

The Acquire fence on the last drop is load-bearing : it guarantees the destructor sees the fields’ final values, not stale cached copies on another core. Without it, a final write made by a holder on one core could still be in that core’s store buffer when the destructor on another core runs — UB by way of stale reads.

The implementation uses fetch_sub + branch-on-prev-value rather than a CAS loop : it’s cheaper on every architecture we target (one atomic op vs a retry loop), and the rarely-taken last-drop branch is colder than the common count > 1 path.

Atomic only when the object can cross a thread

The ordering table above describes the atomic variant. You do not always pay for it. Atomicity is elected per class, at compile time, by a thread-locality analysis in sema:

  • A Shared<T> class is thread-local unless one of its instances can reach a second thread. Sema seeds “reaches another thread” from every spawn (its arguments and the resolved target’s parameters / return / receiver) plus — conservatively — every extern "C" FFI boundary, then closes the set over class field graphs by fixpoint: if a thread-reachable class owns a Shared/Weak field, that field’s class escapes too.
  • For a class in the thread-local set, inc/dec lower to a plain load / add / store with no lock prefix and no Acquire fence on the last drop. For every other class, you get the atomic sequence above.

The decision is keyed on the class (its TypeId), never on a pointer or a binding — so it is automatically sound under aliasing and reassignment: two handles to the same object always share the class, hence the same atomicity. The header layout is identical either way (still 16 B), as is the overflow guard and the drop diamond; only the three instructions change.

You can watch it happen. A Shared<T> graph churned in a single thread compiles to zero atomic operations:

$ axle build shared_local.axle --emit=llvm | grep -c atomicrmw
0

Send the same graph through a spawn and the counter goes atomic again — the lock-prefixed atomicrmw and the last-drop fence reappear, exactly where correctness needs them and nowhere else. In other words: you pay for atomic reference counting only when the object actually crosses a thread.

Overflow

The inc path checks the pre-increment value and calls libc abort() when the 32-bit count is about to wrap from u32::MAX to 0. The alternative — silent wrap — would cause the next decrement to bring a still-live object to refcount zero and free it. A program with > 2³² live references to one object is already broken ; aborting is the safer failure mode. The check costs one compare + a never-taken branch in the steady state, both perfectly predicted.

Arc<T> in Rust’s stdlib makes the same trade-off — an abort-on-overflow branch on the clone — so the per-clone overhead is identical.

Shared::wrap lowering

class Cache {
    hits : i32;
    constructor() { self.hits = 0; }
}

fn promote() : Shared<Cache> {
    let owned = new Cache();                    // uniquely owned
    let s : Shared<Cache> = Shared::wrap(owned); // owned consumed
    return s;
}

The emission sequence :

  1. Call the shared allocator for sizeof Cache — fresh refcount header (count = 1) + uninitialised payload.
  2. memcpy from owned into the new payload slot.
  3. libc::free(owned) to release the original owned cell.
  4. Return the Shared<Cache> pointer.

The source binding is consumed at compile time (own semantics) — no scope-exit cleanup runs on it later. The class identity is preserved bit-for-bit ; no constructor is re-invoked.

Shared::try_unwrap lowering

class Cache {
    hits : i32;
    constructor() { self.hits = 0; }
}

fn recover(s : Shared<Cache>) : i32 {
    let owned = Shared::try_unwrap(s);   // unique value, or null
    if (owned == null) {
        return -1;
    }
    return owned.hits;
}

The emission speculatively decrements the refcount and branches on the pre-dec value :

  • old_count == 1 (last reference) → the dec hit zero. Emit a fresh owned allocation, memcpy the payload, free the shared cell, and return the owned pointer.
  • old_count > 1 → return null. There is no restore-increment : try_unwrap consumes s uniformly, so on the contended outcome your reference is genuinely released — the surviving holders keep the object alive and free it at their own scope exits.

The atomic dec is the same Release-ordered op as any drop ; the Acquire fence taken on the last drop already orders the memcpy against another thread’s final write through the shared handle. The nullable result lowers to a plain pointer (null = no value), so the two branches merge in a single phi. The value recovered on the unique branch is an ordinary owned heap object : its scope-exit cleanup frees it exactly once, unless you move it on (return it, store it, re-wrap it).

See also

memoryrefcountsharedatomics