Multithreading and async
Axle has two distinct concurrency worlds. They are not different APIs for the same thing — they have different execution models, different costs, and different sweet spots.
Before the APIs, the one distinction everything else follows from:
- Parallelism — two things happening at the same instant, on two physical CPU cores. This needs OS threads: only the operating system can place two streams of instructions on two cores. That is the thread world.
- Concurrency — many things in flight but only one running at any instant, interleaved on a single core by handing control back and forth. No extra cores, no OS threads — just one thread that switches between tasks whenever one would otherwise sit idle waiting. That is the async world.
A CPU-bound job (it never waits, it just computes) only goes faster with real parallelism — more cores. An I/O-bound job (it spends its life waiting for a socket, a disk, a timer) doesn’t need more cores; it needs a way to keep thousands of waits in flight cheaply. The two worlds are Axle’s answer to those two shapes.
The two worlds
Thread world — spawn / Task<T>
spawn launches a function call on a fresh OS thread and returns a Task<T> handle. The worker runs in true parallel — it gets a
real CPU core, can block freely, and does not cooperate with anything.
The calling thread continues independently; .join() collects the
result when you’re ready.
main thread worker thread
─────────────────────── ─────────────────────────
let t = spawn encode(data) → fn encode(data) { … }
do other work (runs on its own core,
let result = t.join() ←────── can block, uses real CPU) Use it for: CPU-bound computation (encoding, compression, simulation, ray-tracing), coarse-grained parallelism where you want N cores working simultaneously, or blocking libraries you can’t restructure as async.
Cost: each spawn creates an OS thread + stack (~1–8 MiB). A
hundred threads is fine; ten thousand is not.
Async world — async fn / Future<T>
An async fn body is compiled into a coroutine state machine.
Calling it does nothing — it returns a Future<T> handle. The
body starts running only when await enters the event loop: a
scheduler on the calling thread that drives every live machine in
round-robin. When a machine hits an await, it suspends — saves
its locals into a heap frame, yields to the scheduler, which picks
up the next ready machine. When the awaited thing resolves (a
socket, a timer, a nested future), the scheduler resumes the machine
from where it left off.
single thread — the event loop
───────────────────────────────────────────────────────────
machine A: connect to API → suspends (waiting for TCP)
machine B: read from socket → suspends (waiting for data)
machine C: 50ms timer → suspends
machine C: timer fires → resumes → done
machine B: data arrives → resumes → done
machine A: TCP ready → resumes → done Use it for: I/O-bound concurrency — thousands of simultaneous connections, HTTP requests, file reads — anywhere the work is waiting, not computing. A thousand coroutines sleep on one thread while a reactor wakes them as data arrives.
Cost: each suspended coroutine is just a heap frame (~hundreds of bytes for typical locals). No OS threads involved. Ten thousand concurrent coroutines is routine.
The two worlds, side by side
| Dimension | Thread world (spawn / Task<T>) | Async world (async fn / Future<T>) |
|---|---|---|
| Unit | OS thread | coroutine state machine |
| Parallelism | true — N cores run at once | none within one loop — interleaved on its thread |
| Scheduling | pre-emptive (the OS) | cooperative (yields at await) |
| Cost per unit | a stack, ~1–8 MiB | a heap frame, ~hundreds of bytes |
| Practical ceiling | ~hundreds | tens of thousands |
| Blocking is | fine — only that worker stalls | poison — stalls every coroutine on the thread |
| Best at | CPU-bound crunching | I/O-bound waiting |
| Consume with | .join() (blocks) / .isAlive() (peek) | await / .timeout(ms) |
The dividing line is blocking. A thread may block freely — the OS just runs something else on that core. A coroutine must never block the thread: the whole point is that one thread hosts thousands of them, so a single blocking call freezes them all. That is why CPU crunching belongs in the thread world and waiting-on-I/O belongs in the async world.
Quick pick
| Situation | World |
|---|---|
| CPU-bound: image compression, simulation | spawn / Task<T> |
| I/O-bound: HTTP, sockets, file reads | async fn / Future<T> |
| Blocking third-party library | spawn / Task<T> |
| Thousands of concurrent connections | async fn / Future<T> |
| CPU crunch inside an async pipeline | async fn that spawns a Task |
The worlds compose, in both directions:
- An
async fncanspawnaTask, launch real parallel work, and.join()it back across a suspension. The coroutine sleeps while the thread crunches; the event loop drives other machines in the meantime. spawncan target anasync fn. The worker thread gets its own event loop — its own task table, its own reactor — drives that coroutine to completion there, and hands the result back through theTask<T>. That is how you put more than one core to work on async code: N spawned threads, N independent loops, no shared scheduler to contend on.
A future handle belongs to the loop that minted it, though: pass a Future<T> or Task<T> itself to another thread and the handle is
unknown there. The compiler already blocks the compile-time cases — a
handle is not a portable spawn / await payload (E0729).
Thread world in practice
Parallel image processing
fn jpegCompress(pixels : i32[], width : i32, height : i32) : string {
return "jpeg:" + pixels.length + ":" + width + "x" + height;
}
fn loadFrames() : i32[][4] {
let f : i32[][4];
for (i of 0..4) { f[i] = malloc<i32>(1920 * 1080); }
return f;
}
fn writeFrame(encoded : string, index : i32) : void {
println(index + ": " + encoded);
}
fn encodeFrame(pixels : i32[], width : i32, height : i32) : string {
// CPU-bound: runs on its own core, blocks freely
return jpegCompress(pixels, width, height);
}
fn main() : i32 {
let frames : i32[][4] = loadFrames();
let workers : Task<string>[4];
for (i of 0..4) {
workers[i] = spawn encodeFrame(frames[i], 1920, 1080);
}
let i : i32 = 0;
while (i < 4) {
let encoded : string = workers[i].join();
writeFrame(encoded, i);
i = i + 1;
}
return 0;
} Each frame compresses on its own core. Total time is max(frame_times) instead of sum(frame_times).
Check whether a task is still running
.isAlive() is a non-consuming peek — it does not .join() and
does not block.
use std::net::HttpClient;
fn processLocalData(tick : i32) : void {
println("tick " + tick);
}
fn save(body : string) : void {
println("saved " + body.length + " bytes");
}
fn longDownload(url : string) : string ! IOException {
return HttpClient::get(url).body(); // blocks the worker thread
}
// `spawn`ing a throwing function carries its error set to the caller,
// so `main` re-declares `! IOException`.
fn main() : i32 ! IOException {
let t : Task<string> = spawn longDownload("https://files.example.com/data.zip");
let ticks : i32 = 0;
while (t.isAlive()) {
processLocalData(ticks);
ticks = ticks + 1;
}
let data : string = t.join(); // instant — already done
save(data);
return ticks;
} .isAlive() returns false once the thread finishes or the handle
is already consumed. Use it for progress loops and timeout counters.
Three-stage pipeline
use std::concurrent::BoundedChannel;
use std::io::BufferedReader;
use std::io::FileInputStream;
fn readLines(path : string, out : BoundedChannel) : i32
! InterruptedException, IOException, FileNotFoundException
{
let reader : BufferedReader = new BufferedReader(new FileInputStream(path));
defer reader.close();
let line : string = reader.readLine();
while (!line.isEmpty()) {
out.send(line.length); // a cheap per-line value to forward
line = reader.readLine();
}
out.send(-1); // shutdown sentinel
return 0;
}
fn hashLines(input : BoundedChannel, out : BoundedChannel) : i32
! InterruptedException
{
let v : i32 = input.receive();
while (v != -1) {
out.send(v * 2654435761);
v = input.receive();
}
out.send(-1);
return 0;
}
fn main() : i32 ! InterruptedException, IOException, FileNotFoundException {
let raw : BoundedChannel = new BoundedChannel(64);
let hashed : BoundedChannel = new BoundedChannel(64);
let reader : Task<i32> = spawn readLines("/var/log/app.log", raw);
let hasher : Task<i32> = spawn hashLines(raw, hashed);
let count : i32 = 0;
let v : i32 = hashed.receive();
while (v != -1) { count = count + 1; v = hashed.receive(); }
reader.join();
hasher.join();
raw.close(); // a channel is Closeable — close it on every path
hashed.close();
return count;
} Three stages run in parallel — reader, hasher, and the main thread collecting — each on its own core.
Async world in practice
How a coroutine works
Calling an async fn allocates a frame and returns immediately.
The body starts only when await enters the event loop.
async fn greet(name : string) : string {
let msg : string = "Hello, " + name;
await Async::sleep(100); // suspension point
return msg + "!";
}
fn main() : i32 {
let f : Future<string> = greet("world"); // nothing ran yet
let s : string = await f; // NOW it runs
println(s);
return 0;
} At Async::sleep the machine saves msg into its heap frame and
yields. The event loop parks for 100 ms, then resumes the machine
exactly where it left off — msg is still there.
The compiler splits greet at its one await into a small state
machine. The locals live in a heap frame that outlives every
suspension; a hidden __state word records where to resume:
heap frame (survives the suspension)
┌─────────────────────────────────────────────┐
│ __state : 1 msg : "Hello, world" │
└─────────────────────────────────────────────┘
▲ │
│ resume dispatches │ locals restored from the frame
│ on __state ▼
── state 0 ─────────────────────────────────────────────────
msg = "Hello, " + name
arm the 100 ms timer; __state = 1; yield ──► event loop
── event loop drives other machines for 100 ms ─────────────
── state 1 (timer fired → scheduler resumes this machine) ──
return msg + "!" ──► Done("Hello, world!") A plain function keeps its locals on the call stack and unwinds when it returns; a coroutine can’t, because it must pause and let the stack be used by other machines. Moving the locals into a heap frame is what lets it freeze and thaw — and is why a suspended coroutine costs only its frame (a few hundred bytes), not a whole thread stack.
A thousand concurrent timers on one thread
When you await f, the loop drives every registered machine, not
just f. Concurrent machines make progress automatically:
async fn tick(id : i32) : i32 {
await Async::sleep(30); // all 1000 sleep simultaneously
return id;
}
fn main() : i32 {
let futures : Future<i32>[1000];
for (i of 0..1000) {
futures[i] = tick(i);
}
let sum : i32 = 0;
for (j of 0..1000) {
sum = sum + (await futures[j]);
}
return 0; // done in ~30 ms — not 30 000 ms
} When await futures[0] starts the loop, all 1000 machines hit Async::sleep in the first few rounds. The loop parks until the timer
fires — and all 1000 complete together.
Joining a set of futures — await all
Awaiting futures one by one already runs them concurrently, but you often
want the whole set as a single value. await all (sugar for Async::all(…)) joins a Future<T>[] into one Future<T[]>:
async fn fetch(id : i32) : i32 {
await Async::sleep(1);
return id * 2;
}
fn main() : i32 {
let futures : Future<i32>[] = [ fetch(1), fetch(2), fetch(3) ];
let results : i32[] = await all futures;
return results[0] + results[1] + results[2] - 12; // 2 + 4 + 6
} The results come back in array order, not completion order. Two properties are worth knowing:
- The joined handle is an ordinary
Future<T[]>.Async::all(futures)mints it without awaiting anything, so you can store it in a variable, pass it to another function, return it, andawaitit later — exactly like the future of anyasync fn. - If several children throw, the one with the lowest array index wins, even if another failed sooner. Every child is still driven to a terminal state so its resources are released; the losing outcomes are discarded.
await all 5 or await all over an i32[] is E0738 — the operand
must be an array of futures. Calling Async::all with anything other than
that one argument is E0739.
Concurrent HTTP health checks
use std::net::HttpClient;
async fn check(url : string) : bool {
try {
let body : string = HttpClient::getAsync(url).join();
return body.length > 0;
} catch e : IOException {
return false;
}
}
fn main() : i32 {
let endpoints : string[] = [
"https://api.service-a.com/health",
"https://api.service-b.com/health",
"https://api.service-c.com/health",
"https://api.service-d.com/health",
];
let checks : Future<bool>[4];
for (i of 0..4) { checks[i] = check(endpoints[i]); }
let healthy : i32 = 0;
for (j of 0..4) {
if (await checks[j]) { healthy = healthy + 1; }
}
println("Healthy: " + healthy + "/4");
return 0;
} All four HTTP requests are in-flight simultaneously on a single thread. Total time is the slowest response, not the sum.
Files run on threads, not the reactor
Files have no fd the event loop can poll, so there is no File coroutine verb — file work belongs in the thread world. Hand each
file to a spawned Task and .join() the results; the OS schedules
the blocking reads across cores:
use std::io::BufferedReader;
use std::io::FileInputStream;
fn lineCount(path : string) : i32 ! IOException, FileNotFoundException {
let reader : BufferedReader = new BufferedReader(new FileInputStream(path));
defer reader.close();
let n : i32 = 0;
let line : string = reader.readLine();
while (!line.isEmpty()) { n = n + 1; line = reader.readLine(); }
return n;
}
// `spawn`ing a throwing function propagates its error set to the caller,
// so `main` re-declares `! IOException, FileNotFoundException`.
fn main() : i32 ! IOException, FileNotFoundException {
let files : string[4] = ["a.txt", "b.txt", "c.txt", "d.txt"];
let jobs : Task<i32>[4];
for (i of 0..4) { jobs[i] = spawn lineCount(files[i]); }
let total : i32 = 0;
for (j of 0..4) { total = total + jobs[j].join(); }
return total;
} Each lineCount runs on its own worker thread, so all four files are
read concurrently. The async world (async fn / reactor) is for
sockets, DNS and HTTP — see Async I/O.
DNS + async connect
use std::dns::dnsResolveAsync;
use std::net::Socket;
async fn connectTo(host : string, port : i32) : Socket
! IOException
{
let ip : string = await dnsResolveAsync(host);
let sock : Socket = new Socket(ip, port);
await sock.connectAsync();
return sock;
}
fn main() : i32 ! IOException {
let sock : Socket = await connectTo("db.internal", 5432);
await sock.writeAsync("SELECT 1");
let reply : string = await sock.readAsync(256);
sock.close();
return 0;
} DNS lookup, TCP handshake, write, and read each suspend the coroutine and let the event loop do other work. No thread is blocked.
Timeouts
.timeout(ms) runs the loop with a deadline. If the future resolves
in time you get its value; if the deadline hits, the future is
cancelled and TimeoutException is thrown:
use std::net::HttpClient;
fn fetchWithFallback(url : string) : string {
try {
// A request carries its own deadline, so nothing here waits on a
// `Task` to bound it: `timeoutMillis` belongs to this request and
// the verb closes it.
return HttpClient::request(url)
.timeoutMillis(2000)
.get()
.body();
} catch e : IOException {
return "";
}
} Catching errors across an await
A bare try/catch around an await works exactly like around a
regular call:
use std::net::HttpClient;
async fn fetchSafe(url : string) : string {
try {
return HttpClient::getAsync(url).join(); // join() suspends
} catch e : IOException { // getAsync declares `! IOException`
return "";
}
} Composing threads and async
An async fn can spawn a Task for CPU work and .join() it
across a suspension. The coroutine sleeps while the real thread runs:
use std::net::HttpClient;
fn lz4Compress(data : string) : string {
return "lz4:" + data.length;
}
fn compress(data : string) : string {
return lz4Compress(data); // CPU-bound, runs on its own core
}
async fn processRequest(raw : string) : string ! IOException {
let extra : string = HttpClient::getAsync("https://config/schema").join();
let worker : Task<string> = spawn compress(raw + extra);
return worker.join(); // suspends — loop keeps running
} The other direction works too: spawn an async fn and the worker thread
drives that coroutine on its own event loop. This is how async work spreads
over several cores — one loop per thread, nothing shared between them:
async fn crunch(n : i32) : i32 {
await Async::sleep(0);
return n * n;
}
fn main() : i32 {
let jobs : Task<i32>[4];
for (i of 0..4) {
jobs[i] = spawn crunch(i); // 4 threads, 4 independent loops
}
let total : i32 = 0;
for (j of 0..4) {
total = total + jobs[j].join();
}
return total - 14; // 0 + 1 + 4 + 9
} Each loop’s futures are private to its thread, so don’t hand a Future<T> or Task<T> handle itself to another thread — the compiler rejects the
static cases as a non-portable handle payload (E0729).
Restrictions
These all concern the body of an async fn. A plain fn (including main) may await / .timeout() freely, anywhere — those calls just
drive the event loop inline.
| Error | Construct (inside an async fn body) | Workaround |
|---|---|---|
| E0721 | await or .timeout(ms) inside a loop | Make the loop body a recursive async fn |
| E0722 | await or .timeout(ms) inside defer / try-finally / synchronized | Restructure; bare try/catch is fine |
| E0723 | async fn main | main is the event-loop driver — it can’t itself be a coroutine |
| E0725 | Arena allocation inside async fn | Use heap allocation (new) instead |
| E0726 | async fn with no suspension point | Add an await, or drop async |
These apply anywhere, not only inside an async fn:
| Error | Construct | Workaround |
|---|---|---|
| E0719 | await on a Task<T> (or .join() on a Future<T>) | Match the verb to the world: .join() a task, await a future |
| E0728 | await of something that is not a Future<T> | Drop the await — .timeout(ms) already resolves the future |
| E0510 | Double .join() or double await | A handle is one-shot |
| E0734 | A capturing closure passed to spawn, or held live across an await | Pass the captured values as parameters instead |
| E0729 | spawn / await of a function whose result can’t cross the handle | Return a portable result (scalar, string, struct/tuple/fixed-array, class instance / Shared<T>, stdlib handle); box anything else behind new |
| E0738 | Async::all(x) / await all x where x is not a Future<T>[] | Pass an array of futures |
| E0739 | Async::all with more or fewer than one argument | Async::all(futures) takes exactly the array |
spawn of an async fn is allowed — the worker thread runs the
coroutine on its own event loop.
A result handed back from a spawned Task or an async fn’s Future travels across the thread/coroutine boundary through a uniform
handle, so it must be a portable value. Integer and float scalars, bool / char, strings, struct / tuple / fixed-array values, class
instances and Shared<T> references, and stdlib handles all round-trip
cleanly. A few shapes can’t ride the handle — a raw ptr<T>, a SIMD
vector, a function value, an unresolved type parameter, a Range, a
nullable wrapper, or another Task / Future. If
a return type is one of those, the compiler reports E0729 at the spawn / await site; wrap the value in a small class and return
that (a new-allocated instance is portable) instead:
class Report {
pub rows : i32;
pub bytes : i64;
constructor(r : i32, b : i64) {
self.rows = r;
self.bytes = b;
}
}
fn countRows(path : string) : i32 {
return path.length;
}
fn fileSize(path : string) : i64 {
return path.length as i64;
}
// Return a shared reference (portable) rather than a non-portable result.
// A `Shared<T>` is built with `new shared T(…)` — the `shared` keyword puts
// the instance under reference counting so it can outlive the worker thread.
fn analyse(path : string) : Shared<Report> {
return new shared Report(countRows(path), fileSize(path));
}
fn main() : i32 {
let job : Task<Shared<Report>> = spawn analyse("/var/log/app.log");
let r : Shared<Report> = job.join();
return r.rows;
} Under the hood
Thread world
spawn fn(args) emits a private thunk — an extern "C" fn(*mut u8) -> i64 that unpacks the argument block,
calls the target, frees the block, and widens the result to i64.
Blocks come from a per-thread recycling pool only while the process
has spawned no thread at all — the pool is keyed by the freeing
thread, so it stays coherent only when the freeing thread is the
allocating one, which is the single-threaded case (an event-loop
thread minting and releasing coroutine frames of one shape). Once any
thread has been spawned, spawn argument blocks are allocated and
freed directly through the global allocator, so a tight spawn/join
loop pays one allocator call per spawn. The i64 is narrowed back on .join() (truncated, bitcast, or pointer-cast as needed).
Async world
An async fn is rewritten into a flat state machine before codegen.
Its locals move into a heap frame (word 0 is __state); each
suspension point becomes a numbered state; the resume function
dispatches on __state and returns 0 (“still running”) or the
widened result (“done”). Each OS thread that drives coroutines keeps
its own task registry and its own reactor, so two loops never take a
lock on each other’s behalf. Exceptions thrown inside a coroutine are
captured (class + message) in a Done entry and re-thrown on the
consuming thread at the await, so a try/catch around the await intercepts them naturally. When the coroutine finishes (by
any path), the runtime frees any string fields in the frame before
releasing it — no leaks from strings built inside an async fn.
See also
- Concurrency primitives — locks, channels, atomics
- Async I/O — sockets, files, DNS and HTTP in depth
Shared<T>— cross-thread object lifetimesstd/concurrentreference