Strings and text
string is a built-in immutable UTF-8 type. The full operator and
method surface lives in std/text — the most-used members are
auto-imported, so basic operations need no use.
Equality and ordering
let a : string = "hello";
let b : string = "Hello";
if (a.equals(b)) { /* false */ }
if (a.equalsIgnoreCase(b)) { /* true */ }
let cmp : i32 = a.compareTo(b); // <0, 0, >0 — lexicographic, byte-order == and != on strings compare content, the same answer equals gives — two strings built independently are equal when their
bytes match. Reach for equals when you want the comparison to read
as one at the call site, and equalsIgnoreCase when case must not
matter. It is classes that == compares by identity, so an @derive(Eq) class is where == will not do what you meant.
Length and emptiness
let s : string = "héllo";
let n : i32 = s.length; // 5 — one slot per character
if (s.isEmpty()) { /* false */ }
if (" ".isBlank()) { /* true — only whitespace */ } .length returns the codepoint (character) count, not the byte
count — é counts as one character even though it takes two bytes in
UTF-8. .byteLength returns the byte count, and it is the bound a byte
walk needs: .byteAt(i) reads octets, so length would stop short by
the number of multi-byte characters. For codepoint iteration, see
“Character access” below.
let s : string = "héllo";
let chars : i32 = s.length; // 5
let bytes : i64 = s.byteLength; // 6
let i : i64 = 0;
while (i < s.byteLength) {
let b : i32 = s.byteAt(i);
i = i + 1;
} Slicing
fn main() : i32 ! IndexOutOfBoundsException {
let s : string = "hello world";
let hello : string = s.substring(0, 5); // "hello" — [start, end)
let rest : string = s.substringFrom(6); // "world"
println(hello + " " + rest);
return 0;
} Both throw IndexOutOfBoundsException on out-of-range indices.
Trimming and case
fn main() : i32 {
println(" hello ".trim()); // "hello"
println(" hello ".trimStart()); // "hello "
println(" hello ".trimEnd()); // " hello"
println("AxLe".toLowerCase()); // "axle"
println("AxLe".toUpperCase()); // "AXLE"
return 0;
} Search
fn main() : i32 {
let path : string = "/usr/local/bin/axle";
println(path.startsWith("/usr")); // 1 (true)
println(path.endsWith("/axle")); // 1 (true)
println(path.contains("local")); // 1 (true)
println(path.indexOf("/")); // 0 (first match)
println(path.lastIndexOf("/")); // 14 (last match)
println(path.indexOf("nope")); // -1 (sentinel)
return 0;
} indexOf / lastIndexOf return -1 when the substring is absent.
println renders a bool as 1 / 0; call bool.toString() to read true / false.
Transformation
fn main() : i32 {
println("hello, world".replace(",", ";")); // "hello; world"
println("ab".repeat(3)); // "ababab"
println("hello".reverse()); // "olleh" (codepoint-aware)
return 0;
} replace substitutes every occurrence. There’s no
“first-only” variant — use indexOf + slicing if needed.
Character access
charAt(i) returns the codepoint at character (codepoint) index i,
or -1 if i is past the end. It indexes codepoints, not bytes — for a
multi-byte character charAt(1) is the second character, whatever its
byte offset:
let s : string = "AB";
let a : i32 = s.charAt(0); // 65
let b : i32 = s.charAt(1); // 66
let z : i32 = s.charAt(2); // -1 To iterate codepoints (not bytes) safely, walk via charAt and your own index — UTF-8 boundary handling is up to you.
Classifying a codepoint
A codepoint is just an i32, so a match with range patterns reads
the character classes directly — name the categories with an enum and let the arms map each ASCII range to a variant:
enum CharClass {
Digit,
Upper,
Lower,
Other,
}
fn classify(c : i32) : CharClass {
return match c {
48..=57 => CharClass::Digit, // '0'..'9'
65..=90 => CharClass::Upper, // 'A'..'Z'
97..=122 => CharClass::Lower, // 'a'..'z'
_ => CharClass::Other,
};
} match is an expression and the enum match is exhaustive thanks to
the _ arm; the named variants make a downstream match cls { … } read far better than juggling raw code ranges at every call site.
Padding
fn main() : i32 {
println("42".padStart(5, 48)); // "00042" (48 = '0')
println("42".padEnd(5, 32)); // "42 " (32 = space)
return 0;
} The pad character is a Unicode codepoint (i32).
Parsing
use std::text;
fn readNumbers() : i32 ! ParseException {
let n : i32 = text::parseI32("42");
let big : i64 = text::parseI64("1234567890123");
let f : f64 = text::parseF64("3.14");
let b : bool = text::parseBool("true");
return n;
} parseXxx(s) throws ParseException if s doesn’t represent a
value of the target type.
Conversion to string
let n : i32 = 42;
let f : f64 = 3.14;
let b : bool = true;
let s1 : string = n.toString(); // "42"
let s2 : string = f.toString(); // "3.14"
let s3 : string = b.toString(); // "true" toString() is defined on every primitive ; the output format
matches the usual forms ("42", "3.14", "true").
Concatenation
+ on strings is left-to-right concat:
let name : string = "Alice";
let id : i32 = 7;
let line : string = "user " + name + " id=" + id;
// "user Alice id=7" + with a mixed left-side primitive does work : the left
operand is stringified when one operand is string.
let n : i32 = 42;
let s : string = "n=" + n; // "n=42"
let s2 : string = n + " is the answer"; // "42 is the answer" Building strings — StringBuilder
For long iterative concatenation, reach for StringBuilder — each + copies the whole growing left operand, so a loop of + is O(n²).
use std::collections::ArrayList;
use std::text::StringBuilder;
fn buildCsv(rows : ArrayList<i32>) : string {
let sb : StringBuilder = new StringBuilder();
let i : i32 = 0;
while (i < rows.size()) {
if (i > 0) { sb.append(","); }
sb.appendI32(rows.get(i) ?? 0); // in range — never null
i = i + 1;
}
return sb.toString();
} API surface:
| Method | Effect |
|---|---|
append(s) | append a string |
appendChar(c) | append a codepoint (i32) |
appendI32(v) / appendI64(v) | append an int |
appendF64(v) | append a float |
appendBool(v) | append "true" / "false" |
length() / byteLength() / isEmpty() | character count / byte count / empty test |
toString() | freeze into a fresh owned string ; the builder stays usable |
clear() | reset the buffer for reuse |
reverse() | in-place reverse |
All append* methods return self so you can chain:
use std::text::StringBuilder;
fn main() : i32 {
let s : string = new StringBuilder()
.append("[")
.appendI32(42)
.append(", ")
.appendF64(3.14)
.append("]")
.toString();
println(s);
return 0;
} StringBuilder is O(n) for n total characters ; iterative + on string is O(n²). Use the builder for loops with more than a
handful of iterations.
toString() allocates a fresh owned string every call — the builder
keeps its buffer and stays usable, but the result is a new heap string
whose ownership moves to you. Bind it or return it, and let scope release
it; a toString() whose result is dropped leaks that allocation. A
builder reused across a loop must therefore call toString() once, at
the end (as buildCsv and join above do) — calling it per iteration
allocates a string each pass. Use clear() to reset the buffer for reuse
without freezing.
A string carries its byte length in a header just before the
bytes, so byteLength and byteAt are O(1) (a header read, not a
scan) — byteLength compiles to the header read itself, with no call. length counts codepoints and scans. The O(n²) of iterative + is
therefore purely the repeated copying of the growing left operand — not
a length recomputation — which is exactly what StringBuilder amortises away.
JSON
use std::text::StringBuilder;
use std::text::json;
fn buildPayload(name : string, age : i32) : string {
let sb : StringBuilder = new StringBuilder();
sb.append("{\"name\":");
sb.append(json::jsonQuoteString(name)); // produces `"Alice"` with quotes
sb.append(",\"age\":");
sb.appendI32(age);
sb.append("}");
return sb.toString();
}
fn isValid(payload : string) : bool {
return json::jsonValidate(payload);
} std/text/json exposes:
jsonValidate(text)—trueifftextis valid JSON.jsonQuoteString(s)— wrapssin JSON-escaped quotes ("hi"→"hi"with quotes ; embedded"becomes\").jsonEscapeString(s)— escapes"/\/ control chars without adding outer quotes.jsonI64ToString(v)/jsonF64ToString(v)— JSON-compliant numeric formatting (NaN/Inf→"null").
std/text/json validates and escapes ; it does not parse into a
document tree. For value extraction, run a Pattern (regex) over
the raw text, or keep the payload flat enough to slice by hand.
CSV rows
std/text/csv is an RFC 4180 single-row parser written in pure Axle. It
reads one row at a time by index — no array is materialised — so you loop
over the fields with csvFieldCount + csvField:
use std::text::csv;
fn main() : i32 ! IndexOutOfBoundsException {
let row : string = "alice,\"Smith, Jr.\",42";
let n : i32 = csv::csvFieldCount(row); // 3 — quoting-aware
let i : i32 = 0;
while (i < n) {
println(csv::csvField(row, i)); // alice / Smith, Jr. / 42
i = i + 1;
}
return 0;
} A quoted field ("…") may contain commas, and a doubled "" inside it
is a literal ". csvField(row, i) throws IndexOutOfBoundsException past the last field — the ! clause above propagates it.
To write a row, csvEncode wraps a field in quotes only when it has to
(the field contains a comma, quote, or newline); csvJoin2 encodes two
fields and joins them with a comma:
use std::text::csv;
fn main() : i32 {
println(csv::csvEncode("plain")); // "plain" — no quoting needed
println(csv::csvEncode("has,comma")); // "\"has,comma\"" — quoted
return 0;
} csvJoin2 joins two already-encoded fields with a comma — encode
each field first, then join (it does no quoting of its own):
use std::text::csv;
fn main() : i32 {
println(csv::csvJoin2(csv::csvEncode("a,b"), csv::csvEncode("c"))); // "\"a,b\",c"
return 0;
} Wider rows compose: csvJoin2(csvJoin2(a, b), c), or just build the row
with a StringBuilder and csvEncode each cell as you go.
| Function | Effect |
|---|---|
csvFieldCount(row) | number of fields in one row (quoting-aware) |
csvField(row, i) | the i-th field, unquoted — throws IndexOutOfBoundsException |
csvEncode(field) | quote a field iff it contains , " or newline |
csvJoin2(a, b) | join two already-encoded fields with a comma |
Multi-line records (a " straddling a newline) are not supported — feed
each physical line separately. Split a multi-row document on "\n" first
(see Patterns → Splitting below), then parse each line.
Regex
std::text::regex (alias std::regex) is a Pattern / Matcher API : compile a Pattern once, reuse it, and close() it when
done (it holds a registry slot, so pair it with defer).
use std::text::regex::Pattern;
fn extractDigits(s : string) : string ! IllegalArgumentException, IOException {
let p : Pattern = new Pattern("[^0-9]");
defer p.close();
return p.replaceAll(s, ""); // strip everything non-digit
} | Method | Effect |
|---|---|
matches(s) | whole-string test (anchored at both ends) |
find(s) | true if the pattern occurs anywhere in s |
replaceAll(s, repl) / replaceFirst(s, repl) | rewrite matches |
split(s) | pieces of s between matches, joined by \n |
matcher(s) | a Matcher cursor for iterating matches and reading capture groups |
The constructor throws IllegalArgumentException on a bad
pattern. Syntax is the Rust regex crate’s — no backreferences or
lookaround.
Encoding helpers
use std::text::encoding;
fn roundTrip() : string ! ParseException {
let bytes : string = encoding::base64Encode("Hello"); // "SGVsbG8="
let plain : string = encoding::base64Decode(bytes); // "Hello"
let hex : string = encoding::hexEncode("AB"); // "4142"
let back : string = encoding::hexDecode(hex); // "AB"
return plain;
} The encode direction never fails ; the decode direction throws ParseException on malformed input (a bad base64 symbol, an odd-length
hex string, bytes that aren’t valid UTF-8), so base64Decode / hexDecode must sit in a function that lists ! ParseException or
wraps the call in a catch.
URLs
use std::text::url;
use std::lang::ParseException;
fn main() : i32 ! ParseException {
let escaped : string = url::urlEncode("hello world&foo=bar");
// "hello%20world%26foo%3Dbar"
let raw : string = url::urlDecode("a%20b");
println(escaped + " " + raw);
return 0;
} Hashing
use std::text::hashing;
let h : i64 = hashing::hashString("hello"); // FNV-1a, deterministic across runs
let c : i64 = hashing::hashCombine(h, hashing::hashI64(42)); These are fast non-cryptographic hashes. For integrity
checksums (CRC-32, Adler-32), see std/text/checksum — neither
module substitutes for a cryptographic hash:
use std::text::checksum;
let sum : i64 = checksum::crc32("hello"); // 907060870
let adl : i64 = checksum::adler32("hello"); crc32Update(seed, data) / adler32Update(seed, data) fold more bytes
into a running checksum — pass the previous result as seed to hash a
stream in chunks.
UUIDs
std/text/uuid generates and validates version-4 (random) UUIDs:
use std::text::uuid;
fn main() : i32 ! ParseException {
let id : string = uuid::uuidV4(); // "f47ac10b-58cc-4372-a567-0e02b2c3d479"
let hex : string = uuid::uuidV4Hex(); // same, 32 hex chars, no dashes
if (uuid::uuidIsValid(id)) { println("ok"); }
let canon : string = uuid::uuidNormalize("F47AC10B-58CC-4372-A567-0E02B2C3D479");
// → "f47ac10b58cc4372a5670e02b2c3d479" — dashes stripped, lowercased;
// throws ParseException on a non-UUID
return 0;
} | Function | Effect |
|---|---|
uuidV4() | fresh random UUID, canonical 36-char dashed form |
uuidV4Hex() | fresh random UUID, 32 hex chars, no dashes |
uuidIsValid(s) | true iff s is a well-formed UUID, dashed or not (never throws) |
uuidNormalize(s) | the 32 hex digits, dashes stripped and lowercased — throws ParseException |
uuidV4 draws from a cryptographically-seeded generator, so the values
are suitable as unguessable identifiers.
Semantic versions
std/text/version parses and compares semver strings — the common “is this release newer” question, done right (so 1.10.0 sorts after 1.2.0, which a plain string compare gets wrong):
use std::text::version;
fn main() : i32 ! ParseException {
let cmp : i32 = version::versionCompare("1.2.0", "1.10.0"); // -1 (older)
if (version::versionIsValid("2.0.1-rc.1")) { println("valid"); }
let major : i32 = version::versionMajor("2.0.1"); // 2
let pre : string = version::versionPrerelease("2.0.1-rc.1"); // "rc.1"
return cmp;
} | Function | Effect |
|---|---|
versionCompare(a, b) | -1 / 0 / 1 — precedence-correct (prerelease < release) |
versionIsValid(v) | true iff v parses as semver (never throws) |
versionMajor/Minor/Patch(v) | the numeric components — throw ParseException |
versionPrerelease(v) / versionBuild(v) | the -… / +… tags, or "" |
versionCompare orders by major, minor, patch, then prerelease per the
semver spec — a release always outranks its own prereleases.
Patterns
Splitting
Pattern.split cuts a string on every match and returns the
pieces joined by '\n' — feed that to a line-by-line consumer:
use std::text::regex::Pattern;
fn splitCsvLine(line : string) : string ! IllegalArgumentException, IOException {
let comma : Pattern = new Pattern(",");
defer comma.close();
return comma.split(line); // "a,b,c" → "a\nb\nc"
} For single-character separators, an indexOf + substring loop
also works — collect the pieces into an ArrayList<string>.
Joining
use std::collections::ArrayList;
use std::text::StringBuilder;
fn join(parts : ArrayList<string>, sep : string) : string {
let sb : StringBuilder = new StringBuilder();
let i : i32 = 0;
while (i < parts.size()) {
if (i > 0) { sb.append(sep); }
sb.append(parts.get(i) ?? ""); // in range — never null
i = i + 1;
}
return sb.toString();
} Formatting numbers with padding
fn formatI32Padded(v : i32, width : i32) : string {
return v.toString().padStart(width, 48); // 48 = '0'
} Common pitfalls
== on a class
class Point {
x : i32;
y : i32;
constructor(x : i32, y : i32) { self.x = x; self.y = y; }
}
let p : Point = new Point(1, 2);
let q : Point = new Point(1, 2);
if (p == q) { /* false — same fields, different objects */ } == compares a reference against a reference: two Points with
identical fields are still different objects. Give the class an equals — @derive(Eq) writes one comparing the fields, or declare
it by hand. A string is the exception: == on two strings compares
their contents, because that is the only identity a string has.
.length counts characters, not bytes
let s : string = "café";
let n : i32 = s.length; // 4 — `é` is one character .length is the codepoint (character) count: café is 4 characters
even though é occupies two bytes in UTF-8. Walk the characters with charAt (it returns -1 once you pass the end) when you need
per-character access; reach a raw byte with .byteAt(i), bounded by .byteLength (5 here, not 4).
Iterative + in a loop
use std::text::StringBuilder;
fn slowJoin() : string {
// BAD — quadratic.
let s : string = "";
for (i of 0..1000) { s = s + " " + i; }
return s;
}
fn fastJoin() : string {
// GOOD — linear.
let sb : StringBuilder = new StringBuilder();
for (i of 0..1000) { sb.append(" ").appendI32(i); }
return sb.toString();
} See also
std/textreference — full string surfacestd/text/jsonreference — JSON helpersstd/text/regexreference — regex API- Collections → HashMap word counter — practical string keys
- Concept index — every string / text concept cross-linked