Axle v0.14.1

Strings and text

string is a built-in immutable UTF-8 type. The full operator and method surface lives in std/text — the most-used members are auto-imported, so basic operations need no use.

Equality and ordering

let a : string = "hello";
let b : string = "Hello";

if (a.equals(b))                  { /* false */ }
if (a.equalsIgnoreCase(b))        { /* true  */ }

let cmp : i32 = a.compareTo(b);   // <0, 0, >0 — lexicographic, byte-order

== and != on strings compare content, the same answer equals gives — two strings built independently are equal when their bytes match. Reach for equals when you want the comparison to read as one at the call site, and equalsIgnoreCase when case must not matter. It is classes that == compares by identity, so an @derive(Eq) class is where == will not do what you meant.

Length and emptiness

let s : string = "héllo";

let n : i32 = s.length;                         // 5 — one slot per character
if (s.isEmpty()) { /* false */ }
if ("   ".isBlank())   { /* true — only whitespace */ }

.length returns the codepoint (character) count, not the byte count — é counts as one character even though it takes two bytes in UTF-8. .byteLength returns the byte count, and it is the bound a byte walk needs: .byteAt(i) reads octets, so length would stop short by the number of multi-byte characters. For codepoint iteration, see “Character access” below.

let s : string = "héllo";
let chars : i32 = s.length;        // 5
let bytes : i64 = s.byteLength;    // 6

let i : i64 = 0;
while (i < s.byteLength) {
    let b : i32 = s.byteAt(i);
    i = i + 1;
}

Slicing

fn main() : i32 ! IndexOutOfBoundsException {
    let s : string = "hello world";

    let hello : string = s.substring(0, 5);      // "hello" — [start, end)
    let rest  : string = s.substringFrom(6);     // "world"
    println(hello + " " + rest);
    return 0;
}

Both throw IndexOutOfBoundsException on out-of-range indices.

Trimming and case

fn main() : i32 {
    println("  hello  ".trim());                 // "hello"
    println("  hello  ".trimStart());            // "hello  "
    println("  hello  ".trimEnd());              // "  hello"

    println("AxLe".toLowerCase());               // "axle"
    println("AxLe".toUpperCase());               // "AXLE"
    return 0;
}
fn main() : i32 {
    let path : string = "/usr/local/bin/axle";

    println(path.startsWith("/usr"));            // 1     (true)
    println(path.endsWith("/axle"));             // 1     (true)
    println(path.contains("local"));             // 1     (true)
    println(path.indexOf("/"));                  // 0     (first match)
    println(path.lastIndexOf("/"));              // 14    (last match)
    println(path.indexOf("nope"));               // -1    (sentinel)
    return 0;
}

indexOf / lastIndexOf return -1 when the substring is absent.

println renders a bool as 1 / 0; call bool.toString() to read true / false.

Transformation

fn main() : i32 {
    println("hello, world".replace(",", ";"));   // "hello; world"
    println("ab".repeat(3));                     // "ababab"
    println("hello".reverse());                  // "olleh"  (codepoint-aware)
    return 0;
}

replace substitutes every occurrence. There’s no “first-only” variant — use indexOf + slicing if needed.

Character access

charAt(i) returns the codepoint at character (codepoint) index i, or -1 if i is past the end. It indexes codepoints, not bytes — for a multi-byte character charAt(1) is the second character, whatever its byte offset:

let s : string = "AB";
let a : i32 = s.charAt(0);                       // 65
let b : i32 = s.charAt(1);                       // 66
let z : i32 = s.charAt(2);                       // -1

To iterate codepoints (not bytes) safely, walk via charAt and your own index — UTF-8 boundary handling is up to you.

Classifying a codepoint

A codepoint is just an i32, so a match with range patterns reads the character classes directly — name the categories with an enum and let the arms map each ASCII range to a variant:

enum CharClass {
    Digit,
    Upper,
    Lower,
    Other,
}

fn classify(c : i32) : CharClass {
    return match c {
        48..=57  => CharClass::Digit,    // '0'..'9'
        65..=90  => CharClass::Upper,    // 'A'..'Z'
        97..=122 => CharClass::Lower,    // 'a'..'z'
        _        => CharClass::Other,
    };
}

match is an expression and the enum match is exhaustive thanks to the _ arm; the named variants make a downstream match cls { … } read far better than juggling raw code ranges at every call site.

Padding

fn main() : i32 {
    println("42".padStart(5, 48));               // "00042" (48 = '0')
    println("42".padEnd(5, 32));                 // "42   " (32 = space)
    return 0;
}

The pad character is a Unicode codepoint (i32).

Parsing

use std::text;

fn readNumbers() : i32 ! ParseException {
    let n   : i32  = text::parseI32("42");
    let big : i64  = text::parseI64("1234567890123");
    let f   : f64  = text::parseF64("3.14");
    let b   : bool = text::parseBool("true");
    return n;
}

parseXxx(s) throws ParseException if s doesn’t represent a value of the target type.

Conversion to string

let n : i32     = 42;
let f : f64     = 3.14;
let b : bool = true;

let s1 : string = n.toString();                   // "42"
let s2 : string = f.toString();                   // "3.14"
let s3 : string = b.toString();                   // "true"

toString() is defined on every primitive ; the output format matches the usual forms ("42", "3.14", "true").

Concatenation

+ on strings is left-to-right concat:

let name : string = "Alice";
let id   : i32    = 7;
let line : string = "user " + name + " id=" + id;
// "user Alice id=7"

+ with a mixed left-side primitive does work : the left operand is stringified when one operand is string.

let n  : i32    = 42;
let s  : string = "n=" + n;                       // "n=42"
let s2 : string = n + " is the answer";           // "42 is the answer"

Building strings — StringBuilder

For long iterative concatenation, reach for StringBuilder — each + copies the whole growing left operand, so a loop of + is O(n²).

use std::collections::ArrayList;
use std::text::StringBuilder;

fn buildCsv(rows : ArrayList<i32>) : string {
    let sb : StringBuilder = new StringBuilder();
    let i : i32 = 0;
    while (i < rows.size()) {
        if (i > 0) { sb.append(","); }
        sb.appendI32(rows.get(i) ?? 0);     // in range — never null
        i = i + 1;
    }
    return sb.toString();
}

API surface:

MethodEffect
append(s)append a string
appendChar(c)append a codepoint (i32)
appendI32(v) / appendI64(v)append an int
appendF64(v)append a float
appendBool(v)append "true" / "false"
length() / byteLength() / isEmpty()character count / byte count / empty test
toString()freeze into a fresh owned string ; the builder stays usable
clear()reset the buffer for reuse
reverse()in-place reverse

All append* methods return self so you can chain:

use std::text::StringBuilder;

fn main() : i32 {
    let s : string = new StringBuilder()
        .append("[")
        .appendI32(42)
        .append(", ")
        .appendF64(3.14)
        .append("]")
        .toString();
    println(s);
    return 0;
}

StringBuilder is O(n) for n total characters ; iterative + on string is O(n²). Use the builder for loops with more than a handful of iterations.

toString() allocates a fresh owned string every call — the builder keeps its buffer and stays usable, but the result is a new heap string whose ownership moves to you. Bind it or return it, and let scope release it; a toString() whose result is dropped leaks that allocation. A builder reused across a loop must therefore call toString() once, at the end (as buildCsv and join above do) — calling it per iteration allocates a string each pass. Use clear() to reset the buffer for reuse without freezing.

A string carries its byte length in a header just before the bytes, so byteLength and byteAt are O(1) (a header read, not a scan) — byteLength compiles to the header read itself, with no call. length counts codepoints and scans. The O(n²) of iterative + is therefore purely the repeated copying of the growing left operand — not a length recomputation — which is exactly what StringBuilder amortises away.

JSON

use std::text::StringBuilder;
use std::text::json;

fn buildPayload(name : string, age : i32) : string {
    let sb : StringBuilder = new StringBuilder();
    sb.append("{\"name\":");
    sb.append(json::jsonQuoteString(name));     // produces `"Alice"` with quotes
    sb.append(",\"age\":");
    sb.appendI32(age);
    sb.append("}");
    return sb.toString();
}

fn isValid(payload : string) : bool {
    return json::jsonValidate(payload);
}

std/text/json exposes:

  • jsonValidate(text) — true iff text is valid JSON.
  • jsonQuoteString(s) — wraps s in JSON-escaped quotes ("hi" → "hi" with quotes ; embedded " becomes \").
  • jsonEscapeString(s) — escapes " / \ / control chars without adding outer quotes.
  • jsonI64ToString(v) / jsonF64ToString(v) — JSON-compliant numeric formatting (NaN / Inf → "null").

std/text/json validates and escapes ; it does not parse into a document tree. For value extraction, run a Pattern (regex) over the raw text, or keep the payload flat enough to slice by hand.

CSV rows

std/text/csv is an RFC 4180 single-row parser written in pure Axle. It reads one row at a time by index — no array is materialised — so you loop over the fields with csvFieldCount + csvField:

use std::text::csv;

fn main() : i32 ! IndexOutOfBoundsException {
    let row : string = "alice,\"Smith, Jr.\",42";
    let n : i32 = csv::csvFieldCount(row);       // 3 — quoting-aware
    let i : i32 = 0;
    while (i < n) {
        println(csv::csvField(row, i));  // alice / Smith, Jr. / 42
        i = i + 1;
    }
    return 0;
}

A quoted field ("…") may contain commas, and a doubled "" inside it is a literal ". csvField(row, i) throws IndexOutOfBoundsException past the last field — the ! clause above propagates it.

To write a row, csvEncode wraps a field in quotes only when it has to (the field contains a comma, quote, or newline); csvJoin2 encodes two fields and joins them with a comma:

use std::text::csv;

fn main() : i32 {
    println(csv::csvEncode("plain"));          // "plain"        — no quoting needed
    println(csv::csvEncode("has,comma"));      // "\"has,comma\"" — quoted
    return 0;
}

csvJoin2 joins two already-encoded fields with a comma — encode each field first, then join (it does no quoting of its own):

use std::text::csv;

fn main() : i32 {
    println(csv::csvJoin2(csv::csvEncode("a,b"), csv::csvEncode("c")));  // "\"a,b\",c"
    return 0;
}

Wider rows compose: csvJoin2(csvJoin2(a, b), c), or just build the row with a StringBuilder and csvEncode each cell as you go.

FunctionEffect
csvFieldCount(row)number of fields in one row (quoting-aware)
csvField(row, i)the i-th field, unquoted — throws IndexOutOfBoundsException
csvEncode(field)quote a field iff it contains , " or newline
csvJoin2(a, b)join two already-encoded fields with a comma

Multi-line records (a " straddling a newline) are not supported — feed each physical line separately. Split a multi-row document on "\n" first (see Patterns → Splitting below), then parse each line.

Regex

std::text::regex (alias std::regex) is a Pattern / Matcher API : compile a Pattern once, reuse it, and close() it when done (it holds a registry slot, so pair it with defer).

use std::text::regex::Pattern;

fn extractDigits(s : string) : string ! IllegalArgumentException, IOException {
    let p : Pattern = new Pattern("[^0-9]");
    defer p.close();
    return p.replaceAll(s, "");          // strip everything non-digit
}
MethodEffect
matches(s)whole-string test (anchored at both ends)
find(s)true if the pattern occurs anywhere in s
replaceAll(s, repl) / replaceFirst(s, repl)rewrite matches
split(s)pieces of s between matches, joined by \n
matcher(s)a Matcher cursor for iterating matches and reading capture groups

The constructor throws IllegalArgumentException on a bad pattern. Syntax is the Rust regex crate’s — no backreferences or lookaround.

Encoding helpers

use std::text::encoding;

fn roundTrip() : string ! ParseException {
    let bytes : string = encoding::base64Encode("Hello");   // "SGVsbG8="
    let plain : string = encoding::base64Decode(bytes);     // "Hello"

    let hex : string = encoding::hexEncode("AB");           // "4142"
    let back : string = encoding::hexDecode(hex);           // "AB"
    return plain;
}

The encode direction never fails ; the decode direction throws ParseException on malformed input (a bad base64 symbol, an odd-length hex string, bytes that aren’t valid UTF-8), so base64Decode / hexDecode must sit in a function that lists ! ParseException or wraps the call in a catch.

URLs

use std::text::url;
use std::lang::ParseException;

fn main() : i32 ! ParseException {
    let escaped : string = url::urlEncode("hello world&foo=bar");
    // "hello%20world%26foo%3Dbar"

    let raw : string = url::urlDecode("a%20b");
    println(escaped + " " + raw);
    return 0;
}

Hashing

use std::text::hashing;

let h : i64 = hashing::hashString("hello");      // FNV-1a, deterministic across runs
let c : i64 = hashing::hashCombine(h, hashing::hashI64(42));

These are fast non-cryptographic hashes. For integrity checksums (CRC-32, Adler-32), see std/text/checksum — neither module substitutes for a cryptographic hash:

use std::text::checksum;

let sum : i64 = checksum::crc32("hello");        // 907060870
let adl : i64 = checksum::adler32("hello");

crc32Update(seed, data) / adler32Update(seed, data) fold more bytes into a running checksum — pass the previous result as seed to hash a stream in chunks.

UUIDs

std/text/uuid generates and validates version-4 (random) UUIDs:

use std::text::uuid;

fn main() : i32 ! ParseException {
    let id : string = uuid::uuidV4();            // "f47ac10b-58cc-4372-a567-0e02b2c3d479"
    let hex : string = uuid::uuidV4Hex();        // same, 32 hex chars, no dashes

    if (uuid::uuidIsValid(id)) { println("ok"); }

    let canon : string = uuid::uuidNormalize("F47AC10B-58CC-4372-A567-0E02B2C3D479");
    // → "f47ac10b58cc4372a5670e02b2c3d479" — dashes stripped, lowercased;
    //   throws ParseException on a non-UUID
    return 0;
}
FunctionEffect
uuidV4()fresh random UUID, canonical 36-char dashed form
uuidV4Hex()fresh random UUID, 32 hex chars, no dashes
uuidIsValid(s)true iff s is a well-formed UUID, dashed or not (never throws)
uuidNormalize(s)the 32 hex digits, dashes stripped and lowercased — throws ParseException

uuidV4 draws from a cryptographically-seeded generator, so the values are suitable as unguessable identifiers.

Semantic versions

std/text/version parses and compares semver strings — the common “is this release newer” question, done right (so 1.10.0 sorts after 1.2.0, which a plain string compare gets wrong):

use std::text::version;

fn main() : i32 ! ParseException {
    let cmp : i32 = version::versionCompare("1.2.0", "1.10.0");   // -1 (older)
    if (version::versionIsValid("2.0.1-rc.1")) { println("valid"); }

    let major : i32 = version::versionMajor("2.0.1");            // 2
    let pre : string = version::versionPrerelease("2.0.1-rc.1"); // "rc.1"
    return cmp;
}
FunctionEffect
versionCompare(a, b)-1 / 0 / 1 — precedence-correct (prerelease < release)
versionIsValid(v)true iff v parses as semver (never throws)
versionMajor/Minor/Patch(v)the numeric components — throw ParseException
versionPrerelease(v) / versionBuild(v)the -… / +… tags, or ""

versionCompare orders by major, minor, patch, then prerelease per the semver spec — a release always outranks its own prereleases.

Patterns

Splitting

Pattern.split cuts a string on every match and returns the pieces joined by '\n' — feed that to a line-by-line consumer:

use std::text::regex::Pattern;

fn splitCsvLine(line : string) : string ! IllegalArgumentException, IOException {
    let comma : Pattern = new Pattern(",");
    defer comma.close();
    return comma.split(line);            // "a,b,c" → "a\nb\nc"
}

For single-character separators, an indexOf + substring loop also works — collect the pieces into an ArrayList<string>.

Joining

use std::collections::ArrayList;
use std::text::StringBuilder;

fn join(parts : ArrayList<string>, sep : string) : string {
    let sb : StringBuilder = new StringBuilder();
    let i : i32 = 0;
    while (i < parts.size()) {
        if (i > 0) { sb.append(sep); }
        sb.append(parts.get(i) ?? "");      // in range — never null
        i = i + 1;
    }
    return sb.toString();
}

Formatting numbers with padding

fn formatI32Padded(v : i32, width : i32) : string {
    return v.toString().padStart(width, 48);     // 48 = '0'
}

Common pitfalls

== on a class

class Point {
    x : i32;
    y : i32;
    constructor(x : i32, y : i32) { self.x = x; self.y = y; }
}

let p : Point = new Point(1, 2);
let q : Point = new Point(1, 2);
if (p == q) { /* false — same fields, different objects */ }

== compares a reference against a reference: two Points with identical fields are still different objects. Give the class an equals — @derive(Eq) writes one comparing the fields, or declare it by hand. A string is the exception: == on two strings compares their contents, because that is the only identity a string has.

.length counts characters, not bytes

let s : string = "café";
let n : i32 = s.length;                          // 4 — `é` is one character

.length is the codepoint (character) count: café is 4 characters even though é occupies two bytes in UTF-8. Walk the characters with charAt (it returns -1 once you pass the end) when you need per-character access; reach a raw byte with .byteAt(i), bounded by .byteLength (5 here, not 4).

Iterative + in a loop

use std::text::StringBuilder;

fn slowJoin() : string {
    // BAD — quadratic.
    let s : string = "";
    for (i of 0..1000) { s = s + " " + i; }
    return s;
}

fn fastJoin() : string {
    // GOOD — linear.
    let sb : StringBuilder = new StringBuilder();
    for (i of 0..1000) { sb.append(" ").appendI32(i); }
    return sb.toString();
}

See also

stringstextutf-8formatting