Refcount and transitions
This is the deep-dive companion to Shared<T> reference
counting — read that page first for what a
reference count is and how to use Shared<T>. Here we open the
box: what’s actually allocated, how the atomics are ordered, and
how the explicit Shared::wrap / Shared::try_unwrap transitions
lower to IR.
Both transitions are fully implemented: Shared::wrap emits the
shared allocation and copies the owned payload in; Shared::try_unwrap emits the speculative atomic decrement and
the unique-branch recovery.
Header layout
Every Shared<T> allocation is a single contiguous block :
┌──────────────────────────────────────────┬─────────────────────┐
│ 16-byte refcount header │ payload : T │
│ ┌───────────┬───────────┬──────────────┐ │ │
│ │ AtomicU32 │ AtomicU32 │ u64 │ │ │
│ │ strong │ weak │ payload size │ │ │
│ └───────────┴───────────┴──────────────┘ │ │
└──────────────────────────────────────────┴─────────────────────┘
^ the shared allocator returns this address
^ the compiler steps past
the header to the payload The payload size slot lets the matching free reconstruct the
exact Layout Rust’s dealloc needs — no size threading
required from the call site. The weak count drives the std::sync::Arc / Weak lifecycle: the strong references hold one
implicit weak (so it starts at 1), the payload is dropped at
strong 0, and the block is freed at weak 0.
Atomic ordering
The strong and weak counts are each an AtomicU32 (four billion
references is unreachable before OOM, so carrying both still fits the
16-byte header). Every operation uses an explicit, deliberate
ordering — never SeqCst; the initial store and the clone-side
increments are the Relaxed ops, their happens-before carried
by the drop side.
| Op | Order | Why |
|---|---|---|
| Init (shared allocator returns) | Relaxed store of 1 | The block is not yet published when the counts are written, so no ordering is owed — exactly Arc::new |
inc (clone) | Relaxed fetch_add(1) | The caller already holds a live strong reference, so the block is valid; the cross-thread happens-before is carried by the Release dec and the last-drop Acquire fence — exactly Arc::clone’s discipline |
dec (drop, count > 1) | Release fetch_sub(1) | Publishes the dropper’s writes ; pairs with the last-drop fence so the eventual destructor sees them |
dec (last drop, count == 1) | Acquire fence after the Release fetch_sub | The destructor about to run must observe every write any prior holder made — even on cores that haven’t seen the dec yet |
Both sides mirror Rust’s Arc<T> exactly : a Relaxed clone
with an abort-on-overflow branch, a Release drop, and the Acquire fence on the last drop.
The Acquire fence on the last drop is load-bearing : it guarantees the destructor sees the fields’ final values, not stale cached copies on another core. Without it, a final write made by a holder on one core could still be in that core’s store buffer when the destructor on another core runs — UB by way of stale reads.
The implementation uses fetch_sub + branch-on-prev-value rather
than a CAS loop : it’s cheaper on every architecture we target
(one atomic op vs a retry loop), and the rarely-taken last-drop
branch is colder than the common count > 1 path.
Atomic only when the object can cross a thread
The ordering table above describes the atomic variant. You do not always pay for it. Atomicity is elected per class, at compile time, by a thread-locality analysis in sema:
- A
Shared<T>class is thread-local unless one of its instances can reach a second thread. Sema seeds “reaches another thread” from everyspawn(its arguments and the resolved target’s parameters / return / receiver) plus — conservatively — everyextern "C"FFI boundary, then closes the set over class field graphs by fixpoint: if a thread-reachable class owns aShared/Weakfield, that field’s class escapes too. - For a class in the thread-local set,
inc/declower to a plainload/add/storewith nolockprefix and noAcquirefence on the last drop. For every other class, you get the atomic sequence above.
The decision is keyed on the class (its TypeId), never on a pointer or a
binding — so it is automatically sound under aliasing and reassignment: two
handles to the same object always share the class, hence the same
atomicity. The header layout is identical either way (still 16 B), as is the
overflow guard and the drop diamond; only the three instructions change.
You can watch it happen. A Shared<T> graph churned in a single thread
compiles to zero atomic operations:
$ axle build shared_local.axle --emit=llvm | grep -c atomicrmw
0 Send the same graph through a spawn and the counter goes atomic again —
the lock-prefixed atomicrmw and the last-drop fence reappear, exactly
where correctness needs them and nowhere else. In other words: you pay for
atomic reference counting only when the object actually crosses a thread.
Overflow
The inc path checks the pre-increment value and calls
libc abort() when the 32-bit count is about to wrap
from u32::MAX to 0. The alternative — silent wrap — would
cause the next decrement to bring a still-live object to refcount
zero and free it. A program with > 2³² live references to one
object is already broken ; aborting is the safer failure mode.
The check costs one compare + a never-taken branch in the steady
state, both perfectly predicted.
Arc<T> in Rust’s stdlib makes the same trade-off — an abort-on-overflow
branch on the clone — so the per-clone overhead is identical.
Shared::wrap lowering
class Cache {
hits : i32;
constructor() { self.hits = 0; }
}
fn promote() : Shared<Cache> {
let owned = new Cache(); // uniquely owned
let s : Shared<Cache> = Shared::wrap(owned); // owned consumed
return s;
} The emission sequence :
- Call the shared allocator for
sizeof Cache— fresh refcount header (count = 1) + uninitialised payload. memcpyfromownedinto the new payload slot.libc::free(owned)to release the original owned cell.- Return the
Shared<Cache>pointer.
The source binding is consumed at compile time (own semantics) —
no scope-exit cleanup runs on it later. The class identity is
preserved bit-for-bit ; no constructor is re-invoked.
Shared::try_unwrap lowering
class Cache {
hits : i32;
constructor() { self.hits = 0; }
}
fn recover(s : Shared<Cache>) : i32 {
let owned = Shared::try_unwrap(s); // unique value, or null
if (owned == null) {
return -1;
}
return owned.hits;
} The emission speculatively decrements the refcount and branches on the pre-dec value :
old_count == 1(last reference) → the dec hit zero. Emit a fresh owned allocation,memcpythe payload, free the shared cell, and return the owned pointer.old_count > 1→ returnnull. There is no restore-increment :try_unwrapconsumessuniformly, so on the contended outcome your reference is genuinely released — the surviving holders keep the object alive and free it at their own scope exits.
The atomic dec is the same Release-ordered op as any drop ; the Acquire fence taken on the last drop already orders the memcpy against another thread’s final write through the shared
handle. The nullable result lowers to a plain pointer (null =
no value), so the two branches merge in a single phi. The value
recovered on the unique branch is an ordinary owned heap object :
its scope-exit cleanup frees it exactly once, unless you move it
on (return it, store it, re-wrap it).
See also
Shared<T>reference counting — user-facing API and lifecycle.- Memory model — when the compiler picks shared vs heap vs arena.
- Escape analysis and promotion — the other deep-dive page in this section.
- Compiler internals overview — index of all internals pages.
- Concept index — every shared-memory concept on one page.