Axle v0.14.1

Multithreading and async

Axle has two distinct concurrency worlds. They are not different APIs for the same thing — they have different execution models, different costs, and different sweet spots.

Before the APIs, the one distinction everything else follows from:

  • Parallelism — two things happening at the same instant, on two physical CPU cores. This needs OS threads: only the operating system can place two streams of instructions on two cores. That is the thread world.
  • Concurrency — many things in flight but only one running at any instant, interleaved on a single core by handing control back and forth. No extra cores, no OS threads — just one thread that switches between tasks whenever one would otherwise sit idle waiting. That is the async world.

A CPU-bound job (it never waits, it just computes) only goes faster with real parallelism — more cores. An I/O-bound job (it spends its life waiting for a socket, a disk, a timer) doesn’t need more cores; it needs a way to keep thousands of waits in flight cheaply. The two worlds are Axle’s answer to those two shapes.

The two worlds

Thread world — spawn / Task<T>

spawn launches a function call on a fresh OS thread and returns a Task<T> handle. The worker runs in true parallel — it gets a real CPU core, can block freely, and does not cooperate with anything. The calling thread continues independently; .join() collects the result when you’re ready.

  main thread                      worker thread
  ───────────────────────          ─────────────────────────
  let t = spawn encode(data)  →    fn encode(data) { … }
  do other work                    (runs on its own core,
  let result = t.join()  ←──────   can block, uses real CPU)

Use it for: CPU-bound computation (encoding, compression, simulation, ray-tracing), coarse-grained parallelism where you want N cores working simultaneously, or blocking libraries you can’t restructure as async.

Cost: each spawn creates an OS thread + stack (~1–8 MiB). A hundred threads is fine; ten thousand is not.

Async world — async fn / Future<T>

An async fn body is compiled into a coroutine state machine. Calling it does nothing — it returns a Future<T> handle. The body starts running only when await enters the event loop: a scheduler on the calling thread that drives every live machine in round-robin. When a machine hits an await, it suspends — saves its locals into a heap frame, yields to the scheduler, which picks up the next ready machine. When the awaited thing resolves (a socket, a timer, a nested future), the scheduler resumes the machine from where it left off.

  single thread — the event loop
  ───────────────────────────────────────────────────────────
  machine A: connect to API → suspends (waiting for TCP)
    machine B: read from socket → suspends (waiting for data)
      machine C: 50ms timer → suspends
      machine C: timer fires  → resumes → done
    machine B: data arrives   → resumes → done
  machine A: TCP ready        → resumes → done

Use it for: I/O-bound concurrency — thousands of simultaneous connections, HTTP requests, file reads — anywhere the work is waiting, not computing. A thousand coroutines sleep on one thread while a reactor wakes them as data arrives.

Cost: each suspended coroutine is just a heap frame (~hundreds of bytes for typical locals). No OS threads involved. Ten thousand concurrent coroutines is routine.

The two worlds, side by side

DimensionThread world (spawn / Task<T>)Async world (async fn / Future<T>)
UnitOS threadcoroutine state machine
Parallelismtrue — N cores run at oncenone within one loop — interleaved on its thread
Schedulingpre-emptive (the OS)cooperative (yields at await)
Cost per unita stack, ~1–8 MiBa heap frame, ~hundreds of bytes
Practical ceiling~hundredstens of thousands
Blocking isfine — only that worker stallspoison — stalls every coroutine on the thread
Best atCPU-bound crunchingI/O-bound waiting
Consume with.join() (blocks) / .isAlive() (peek)await / .timeout(ms)

The dividing line is blocking. A thread may block freely — the OS just runs something else on that core. A coroutine must never block the thread: the whole point is that one thread hosts thousands of them, so a single blocking call freezes them all. That is why CPU crunching belongs in the thread world and waiting-on-I/O belongs in the async world.

Quick pick

SituationWorld
CPU-bound: image compression, simulationspawn / Task<T>
I/O-bound: HTTP, sockets, file readsasync fn / Future<T>
Blocking third-party libraryspawn / Task<T>
Thousands of concurrent connectionsasync fn / Future<T>
CPU crunch inside an async pipelineasync fn that spawns a Task

The worlds compose, in both directions:

  • An async fn can spawn a Task, launch real parallel work, and .join() it back across a suspension. The coroutine sleeps while the thread crunches; the event loop drives other machines in the meantime.
  • spawn can target an async fn. The worker thread gets its own event loop — its own task table, its own reactor — drives that coroutine to completion there, and hands the result back through the Task<T>. That is how you put more than one core to work on async code: N spawned threads, N independent loops, no shared scheduler to contend on.

A future handle belongs to the loop that minted it, though: pass a Future<T> or Task<T> itself to another thread and the handle is unknown there. The compiler already blocks the compile-time cases — a handle is not a portable spawn / await payload (E0729).


Thread world in practice

Parallel image processing

fn jpegCompress(pixels : i32[], width : i32, height : i32) : string {
    return "jpeg:" + pixels.length + ":" + width + "x" + height;
}

fn loadFrames() : i32[][4] {
    let f : i32[][4];
    for (i of 0..4) { f[i] = malloc<i32>(1920 * 1080); }
    return f;
}

fn writeFrame(encoded : string, index : i32) : void {
    println(index + ": " + encoded);
}

fn encodeFrame(pixels : i32[], width : i32, height : i32) : string {
    // CPU-bound: runs on its own core, blocks freely
    return jpegCompress(pixels, width, height);
}

fn main() : i32 {
    let frames : i32[][4] = loadFrames();

    let workers : Task<string>[4];
    for (i of 0..4) {
        workers[i] = spawn encodeFrame(frames[i], 1920, 1080);
    }

    let i : i32 = 0;
    while (i < 4) {
        let encoded : string = workers[i].join();
        writeFrame(encoded, i);
        i = i + 1;
    }
    return 0;
}

Each frame compresses on its own core. Total time is max(frame_times) instead of sum(frame_times).

Check whether a task is still running

.isAlive() is a non-consuming peek — it does not .join() and does not block.

use std::net::HttpClient;

fn processLocalData(tick : i32) : void {
    println("tick " + tick);
}

fn save(body : string) : void {
    println("saved " + body.length + " bytes");
}

fn longDownload(url : string) : string ! IOException {
    return HttpClient::get(url).body();   // blocks the worker thread
}

// `spawn`ing a throwing function carries its error set to the caller,
// so `main` re-declares `! IOException`.
fn main() : i32 ! IOException {
    let t : Task<string> = spawn longDownload("https://files.example.com/data.zip");

    let ticks : i32 = 0;
    while (t.isAlive()) {
        processLocalData(ticks);
        ticks = ticks + 1;
    }

    let data : string = t.join();    // instant — already done
    save(data);
    return ticks;
}

.isAlive() returns false once the thread finishes or the handle is already consumed. Use it for progress loops and timeout counters.

Three-stage pipeline

use std::concurrent::BoundedChannel;
use std::io::BufferedReader;
use std::io::FileInputStream;

fn readLines(path : string, out : BoundedChannel) : i32
    ! InterruptedException, IOException, FileNotFoundException
{
    let reader : BufferedReader = new BufferedReader(new FileInputStream(path));
    defer reader.close();
    let line : string = reader.readLine();
    while (!line.isEmpty()) {
        out.send(line.length);            // a cheap per-line value to forward
        line = reader.readLine();
    }
    out.send(-1);       // shutdown sentinel
    return 0;
}

fn hashLines(input : BoundedChannel, out : BoundedChannel) : i32
    ! InterruptedException
{
    let v : i32 = input.receive();
    while (v != -1) {
        out.send(v * 2654435761);
        v = input.receive();
    }
    out.send(-1);
    return 0;
}

fn main() : i32 ! InterruptedException, IOException, FileNotFoundException {
    let raw    : BoundedChannel = new BoundedChannel(64);
    let hashed : BoundedChannel = new BoundedChannel(64);

    let reader : Task<i32> = spawn readLines("/var/log/app.log", raw);
    let hasher : Task<i32> = spawn hashLines(raw, hashed);

    let count : i32 = 0;
    let v : i32 = hashed.receive();
    while (v != -1) { count = count + 1; v = hashed.receive(); }

    reader.join();
    hasher.join();
    raw.close();          // a channel is Closeable — close it on every path
    hashed.close();
    return count;
}

Three stages run in parallel — reader, hasher, and the main thread collecting — each on its own core.


Async world in practice

How a coroutine works

Calling an async fn allocates a frame and returns immediately. The body starts only when await enters the event loop.

async fn greet(name : string) : string {
    let msg : string = "Hello, " + name;
    await Async::sleep(100);          // suspension point
    return msg + "!";
}

fn main() : i32 {
    let f : Future<string> = greet("world");   // nothing ran yet
    let s : string = await f;                  // NOW it runs
    println(s);
    return 0;
}

At Async::sleep the machine saves msg into its heap frame and yields. The event loop parks for 100 ms, then resumes the machine exactly where it left off — msg is still there.

The compiler splits greet at its one await into a small state machine. The locals live in a heap frame that outlives every suspension; a hidden __state word records where to resume:

  heap frame (survives the suspension)
  ┌─────────────────────────────────────────────┐
  │  __state : 1        msg : "Hello, world"     │
  └─────────────────────────────────────────────┘
         ▲                       │
         │ resume dispatches     │ locals restored from the frame
         │ on __state            ▼
  ── state 0 ─────────────────────────────────────────────────
     msg = "Hello, " + name
     arm the 100 ms timer;  __state = 1;  yield ──► event loop
  ── event loop drives other machines for 100 ms ─────────────
  ── state 1  (timer fired → scheduler resumes this machine) ──
     return msg + "!"   ──►  Done("Hello, world!")

A plain function keeps its locals on the call stack and unwinds when it returns; a coroutine can’t, because it must pause and let the stack be used by other machines. Moving the locals into a heap frame is what lets it freeze and thaw — and is why a suspended coroutine costs only its frame (a few hundred bytes), not a whole thread stack.

A thousand concurrent timers on one thread

When you await f, the loop drives every registered machine, not just f. Concurrent machines make progress automatically:

async fn tick(id : i32) : i32 {
    await Async::sleep(30);          // all 1000 sleep simultaneously
    return id;
}

fn main() : i32 {
    let futures : Future<i32>[1000];
    for (i of 0..1000) {
        futures[i] = tick(i);
    }

    let sum : i32 = 0;
    for (j of 0..1000) {
        sum = sum + (await futures[j]);
    }
    return 0;                       // done in ~30 ms — not 30 000 ms
}

When await futures[0] starts the loop, all 1000 machines hit Async::sleep in the first few rounds. The loop parks until the timer fires — and all 1000 complete together.

Joining a set of futures — await all

Awaiting futures one by one already runs them concurrently, but you often want the whole set as a single value. await all (sugar for Async::all(…)) joins a Future<T>[] into one Future<T[]>:

async fn fetch(id : i32) : i32 {
    await Async::sleep(1);
    return id * 2;
}

fn main() : i32 {
    let futures : Future<i32>[] = [ fetch(1), fetch(2), fetch(3) ];
    let results : i32[] = await all futures;
    return results[0] + results[1] + results[2] - 12;   // 2 + 4 + 6
}

The results come back in array order, not completion order. Two properties are worth knowing:

  • The joined handle is an ordinary Future<T[]>. Async::all(futures) mints it without awaiting anything, so you can store it in a variable, pass it to another function, return it, and await it later — exactly like the future of any async fn.
  • If several children throw, the one with the lowest array index wins, even if another failed sooner. Every child is still driven to a terminal state so its resources are released; the losing outcomes are discarded.

await all 5 or await all over an i32[] is E0738 — the operand must be an array of futures. Calling Async::all with anything other than that one argument is E0739.

Concurrent HTTP health checks

use std::net::HttpClient;

async fn check(url : string) : bool {
    try {
        let body : string = HttpClient::getAsync(url).join();
        return body.length > 0;
    } catch e : IOException {
        return false;
    }
}

fn main() : i32 {
    let endpoints : string[] = [
        "https://api.service-a.com/health",
        "https://api.service-b.com/health",
        "https://api.service-c.com/health",
        "https://api.service-d.com/health",
    ];

    let checks : Future<bool>[4];
    for (i of 0..4) { checks[i] = check(endpoints[i]); }

    let healthy : i32 = 0;
    for (j of 0..4) {
        if (await checks[j]) { healthy = healthy + 1; }
    }
    println("Healthy: " + healthy + "/4");
    return 0;
}

All four HTTP requests are in-flight simultaneously on a single thread. Total time is the slowest response, not the sum.

Files run on threads, not the reactor

Files have no fd the event loop can poll, so there is no File coroutine verb — file work belongs in the thread world. Hand each file to a spawned Task and .join() the results; the OS schedules the blocking reads across cores:

use std::io::BufferedReader;
use std::io::FileInputStream;

fn lineCount(path : string) : i32 ! IOException, FileNotFoundException {
    let reader : BufferedReader = new BufferedReader(new FileInputStream(path));
    defer reader.close();
    let n : i32 = 0;
    let line : string = reader.readLine();
    while (!line.isEmpty()) { n = n + 1; line = reader.readLine(); }
    return n;
}

// `spawn`ing a throwing function propagates its error set to the caller,
// so `main` re-declares `! IOException, FileNotFoundException`.
fn main() : i32 ! IOException, FileNotFoundException {
    let files : string[4] = ["a.txt", "b.txt", "c.txt", "d.txt"];
    let jobs  : Task<i32>[4];
    for (i of 0..4) { jobs[i] = spawn lineCount(files[i]); }

    let total : i32 = 0;
    for (j of 0..4) { total = total + jobs[j].join(); }
    return total;
}

Each lineCount runs on its own worker thread, so all four files are read concurrently. The async world (async fn / reactor) is for sockets, DNS and HTTP — see Async I/O.

DNS + async connect

use std::dns::dnsResolveAsync;
use std::net::Socket;

async fn connectTo(host : string, port : i32) : Socket
    ! IOException
{
    let ip   : string = await dnsResolveAsync(host);
    let sock : Socket = new Socket(ip, port);
    await sock.connectAsync();
    return sock;
}

fn main() : i32 ! IOException {
    let sock : Socket = await connectTo("db.internal", 5432);
    await sock.writeAsync("SELECT 1");
    let reply : string = await sock.readAsync(256);
    sock.close();
    return 0;
}

DNS lookup, TCP handshake, write, and read each suspend the coroutine and let the event loop do other work. No thread is blocked.

Timeouts

.timeout(ms) runs the loop with a deadline. If the future resolves in time you get its value; if the deadline hits, the future is cancelled and TimeoutException is thrown:

use std::net::HttpClient;

fn fetchWithFallback(url : string) : string {
    try {
        // A request carries its own deadline, so nothing here waits on a
        // `Task` to bound it: `timeoutMillis` belongs to this request and
        // the verb closes it.
        return HttpClient::request(url)
            .timeoutMillis(2000)
            .get()
            .body();
    } catch e : IOException {
        return "";
    }
}

Catching errors across an await

A bare try/catch around an await works exactly like around a regular call:

use std::net::HttpClient;

async fn fetchSafe(url : string) : string {
    try {
        return HttpClient::getAsync(url).join();   // join() suspends
    } catch e : IOException {       // getAsync declares `! IOException`
        return "";
    }
}

Composing threads and async

An async fn can spawn a Task for CPU work and .join() it across a suspension. The coroutine sleeps while the real thread runs:

use std::net::HttpClient;

fn lz4Compress(data : string) : string {
    return "lz4:" + data.length;
}

fn compress(data : string) : string {
    return lz4Compress(data);        // CPU-bound, runs on its own core
}

async fn processRequest(raw : string) : string ! IOException {
    let extra : string = HttpClient::getAsync("https://config/schema").join();
    let worker : Task<string> = spawn compress(raw + extra);
    return worker.join();            // suspends — loop keeps running
}

The other direction works too: spawn an async fn and the worker thread drives that coroutine on its own event loop. This is how async work spreads over several cores — one loop per thread, nothing shared between them:

async fn crunch(n : i32) : i32 {
    await Async::sleep(0);
    return n * n;
}

fn main() : i32 {
    let jobs : Task<i32>[4];
    for (i of 0..4) {
        jobs[i] = spawn crunch(i);   // 4 threads, 4 independent loops
    }
    let total : i32 = 0;
    for (j of 0..4) {
        total = total + jobs[j].join();
    }
    return total - 14;               // 0 + 1 + 4 + 9
}

Each loop’s futures are private to its thread, so don’t hand a Future<T> or Task<T> handle itself to another thread — the compiler rejects the static cases as a non-portable handle payload (E0729).


Restrictions

These all concern the body of an async fn. A plain fn (including main) may await / .timeout() freely, anywhere — those calls just drive the event loop inline.

ErrorConstruct (inside an async fn body)Workaround
E0721await or .timeout(ms) inside a loopMake the loop body a recursive async fn
E0722await or .timeout(ms) inside defer / try-finally / synchronizedRestructure; bare try/catch is fine
E0723async fn mainmain is the event-loop driver — it can’t itself be a coroutine
E0725Arena allocation inside async fnUse heap allocation (new) instead
E0726async fn with no suspension pointAdd an await, or drop async

These apply anywhere, not only inside an async fn:

ErrorConstructWorkaround
E0719await on a Task<T> (or .join() on a Future<T>)Match the verb to the world: .join() a task, await a future
E0728await of something that is not a Future<T>Drop the await — .timeout(ms) already resolves the future
E0510Double .join() or double awaitA handle is one-shot
E0734A capturing closure passed to spawn, or held live across an awaitPass the captured values as parameters instead
E0729spawn / await of a function whose result can’t cross the handleReturn a portable result (scalar, string, struct/tuple/fixed-array, class instance / Shared<T>, stdlib handle); box anything else behind new
E0738Async::all(x) / await all x where x is not a Future<T>[]Pass an array of futures
E0739Async::all with more or fewer than one argumentAsync::all(futures) takes exactly the array

spawn of an async fn is allowed — the worker thread runs the coroutine on its own event loop.

A result handed back from a spawned Task or an async fn’s Future travels across the thread/coroutine boundary through a uniform handle, so it must be a portable value. Integer and float scalars, bool / char, strings, struct / tuple / fixed-array values, class instances and Shared<T> references, and stdlib handles all round-trip cleanly. A few shapes can’t ride the handle — a raw ptr<T>, a SIMD vector, a function value, an unresolved type parameter, a Range, a nullable wrapper, or another Task / Future. If a return type is one of those, the compiler reports E0729 at the spawn / await site; wrap the value in a small class and return that (a new-allocated instance is portable) instead:

class Report {
    pub rows  : i32;
    pub bytes : i64;
    constructor(r : i32, b : i64) {
        self.rows  = r;
        self.bytes = b;
    }
}

fn countRows(path : string) : i32 {
    return path.length;
}

fn fileSize(path : string) : i64 {
    return path.length as i64;
}

// Return a shared reference (portable) rather than a non-portable result.
// A `Shared<T>` is built with `new shared T(…)` — the `shared` keyword puts
// the instance under reference counting so it can outlive the worker thread.
fn analyse(path : string) : Shared<Report> {
    return new shared Report(countRows(path), fileSize(path));
}

fn main() : i32 {
    let job : Task<Shared<Report>> = spawn analyse("/var/log/app.log");
    let r   : Shared<Report> = job.join();
    return r.rows;
}

Under the hood

Thread world

spawn fn(args) emits a private thunk — an extern "C" fn(*mut u8) -> i64 that unpacks the argument block, calls the target, frees the block, and widens the result to i64. Blocks come from a per-thread recycling pool only while the process has spawned no thread at all — the pool is keyed by the freeing thread, so it stays coherent only when the freeing thread is the allocating one, which is the single-threaded case (an event-loop thread minting and releasing coroutine frames of one shape). Once any thread has been spawned, spawn argument blocks are allocated and freed directly through the global allocator, so a tight spawn/join loop pays one allocator call per spawn. The i64 is narrowed back on .join() (truncated, bitcast, or pointer-cast as needed).

Async world

An async fn is rewritten into a flat state machine before codegen. Its locals move into a heap frame (word 0 is __state); each suspension point becomes a numbered state; the resume function dispatches on __state and returns 0 (“still running”) or the widened result (“done”). Each OS thread that drives coroutines keeps its own task registry and its own reactor, so two loops never take a lock on each other’s behalf. Exceptions thrown inside a coroutine are captured (class + message) in a Done entry and re-thrown on the consuming thread at the await, so a try/catch around the await intercepts them naturally. When the coroutine finishes (by any path), the runtime frees any string fields in the frame before releasing it — no leaks from strings built inside an async fn.


See also

concurrencythreadsasyncparallelism