std/text
std/text — string operations and parsing.
Free functions over the primitive string type, backing thestring members (the .length property, .equals(), .byteAt(), …)
the compiler lowers onto them.
The byte-level operations (search, comparison, slicing by character
index, formatting, the polynomial hash) are pure Axle, built on thebyteAt primitive (strByteAt) and a strByteLen O(1) header read —
both @link(symbol = …) bindings straight to the runtime, which owns the
string layout. What stays native (a Rust #[axle_native] shim) is
what cannot be reproduced byte-for-byte without it:stringFromBytes / __string_concat (the raw bridges),
Unicode whitespace (isBlank / trim*) and case folding
(toLowerCase / toUpperCase), and IEEE-754 float parse / format
(parseF64 / __f64ToString); the integer / bool parsers trim
Unicode whitespace before parsing, so they stay native too.
Free functions
| Type | Method and description |
|---|---|
| stringFromBytes(buf : ptr, len : i64) : string Copy len bytes from a raw libc buffer into an owned string(UTF-8, lossy). The read-side counterpart to passing a stringstraight to a char* libc binding. The result owns its allocation,so the source buffer may be freed immediately after. |
| bytesToString(buf : i8[], len : i64) : string Build an owned string from the first len bytes of a byte array.The array → ptr decay required by the native bridge is confinedto this function's unsafe block. |
| bufferPtr(buf : i8[]) : ptr<i8> Decay a byte array to the typed pointer the scratch-buffer routines ( std::core::bytes) index. The array → ptr decay is confined here.The result stays a pointer the whole way down: an i64 address wouldforce each access to rebuild a pointer with inttoptr, which coststwo extra instructions and, because the result has no provenance, stops alias analysis from proving anything about the buffer. |
| bufferAddr(buf : i8[]) : i64 The same base address as an i64, for the few sites that hand itstraight to a libc binding taking an opaque ptr (fread / fwrite/ realpath) rather than indexing it.Prefer bufferPtr for anything that reads or writes bytes: theinteger form loses provenance, and every access through it has to rebuild a pointer. This exists only so the decay itself is still written once. |
| spanFits(buf : i8[], offset : i64, len : i64) : bool Does the window buf[offset .. offset+len) lie inside buf?The one place a caller-supplied (offset, len) is measured against thebuffer it addresses. Every site that reaches a byte through [ bufferAddr] has already left the array behind — the FFI call it handsthe address to sees a bare pointer and a count, and no bound at all. So the measurement has to happen here, while buf is still an array andstill knows its own length. It refuses rather than clamping. A short read is a legitimate answer from a stream; a request reaching past the buffer is a caller that has miscounted, and silently serving a prefix of it hides the miscount until the data is wrong somewhere else. The addition is written as a subtraction because Axle's + wraps: offset + len on two large valuesis a small number that passes. |
| addrAsPtr(addr : i64) : ptr Reinterpret a buffer address as the opaque ptr an FFI binding takes —the read side of bufferAddr, for a call that needs an offset into abuffer rather than its base. Here, and only here, so the i64 -> ptr decay has one site the waybufferPtr gives the array -> ptr one its site. Prefer bufferPtrwhenever the result is indexed: an address that has been through an integer carries no provenance, and every access through it has to rebuild a pointer. |
| __stringByteAt(s : string, i : i64) : i32 Raw byte 0..255 at byte-index i, or -1 if out of range. Theprimitive every hand-rolled scanner builds on; backed directly by the runtime's bounds-checked accessor ( strByteAt) since the runtime ownsthe string layout. |
| __string_concat(a : string, b : string) : string Runtime helper backing a + b when either operand is a string. Theresult is a fresh owned string; the operands are left untouched. |
| __stringRepeat(s : string, n : i32) : string s repeated n times ("" when n <= 0). Builds the result in asingle allocation filled in place — no scratch buffer, no re-copy. |
| __stringConcat(a : string, b : string) : string Backs the .concat(other) method — distinct from the + operator, whichlowers to the native __string_concat. Joins a and b into a freshowned string ("" when both are empty). |
| __stringIsBlank(s : string) : bool True when s is empty or contains only Unicode whitespace. |
| __stringTrim(s : string) : string Drop leading and trailing Unicode whitespace, returning an owned string. |
| __stringTrimStart(s : string) : string Drop leading Unicode whitespace, returning an owned string. |
| __stringTrimEnd(s : string) : string Drop trailing Unicode whitespace, returning an owned string. |
| __stringToLowerCase(s : string) : string Full Unicode lowercase folding (locale-independent), as an owned string. |
| __stringToUpperCase(s : string) : string Full Unicode uppercase folding (locale-independent), as an owned string. |
| parseI32(s : string) : i32 Parse a base-10 i32, trimming surrounding Unicode whitespace first. |
| parseI64(s : string) : i64 Parse a base-10 i64, trimming surrounding Unicode whitespace first. |
| parseF64(s : string) : f64 Parse an IEEE-754 f64, trimming surrounding Unicode whitespace first. |
| parseBool(s : string) : bool Parse "true" / "false" (after trimming Unicode whitespace). |
| __stringToI32(s : string) : i32 s.toI32() — @throws ParseException when the text is not a 32-bit integer. |
| __stringToI64(s : string) : i64 s.toI64() — @throws ParseException when the text is not a 64-bit integer. |
| __stringToF64(s : string) : f64 s.toF64() — @throws ParseException when the text is not a float literal. |
| __stringToBool(s : string) : bool s.toBool() — @throws ParseException for anything but "true" / "false". |
| __f64ToString(v : f64) : string IEEE-754 shortest round-trip text of v (the value parseF64 inverts). |
| __txtSubstr(s : string, start : i64, end : i64) : string Copy bytes [start, end) of s into an owned string. |
| __charCountPrefix(s : string, byteEnd : i64) : i32 Count of UTF-8 characters in [0, byteEnd) — every byte that is nota continuation byte ( 0b10xxxxxx) starts a new character. Backed bythe runtime's O(1) header read + one pass over the byte slice ( strCharCountPrefix), not a per-byte byteAt FFI loop (which madelength / indexOf / substring O(n) function calls on the length). |
| __charByteIndex(s : string, srcLen : i64, charIdx : i32) : i64 Byte offset of the charIdx-th character in s (whose byte lengthis srcLen): the string end when charIdx equals the charactercount, or -1 when it exceeds it. |
| __charWidth(lead : i32) : i64 UTF-8 width (1..4 bytes) of the character whose lead byte is lead.A stray continuation byte counts as width 1 to keep scans advancing. |
| __decodeCodepoint(s : string, byteOffset : i64) : i32 Decode the Unicode codepoint of the character starting atbyteOffset (assumed valid UTF-8). |
| __txtPutUtf8(buffer : ptr<i8>, offset : i64, codepoint : i32) : i64 UTF-8 encode codepoint (invalid scalars folded to U+FFFD, matchingchar::from_u32(..).unwrap_or(..)) into buffer at offset,returning the offset past the bytes written. |
| __byteIndexOf(s : string, sourceLen : i64, sub : string, subLen : i64, from : i64) : i64 Byte offset of the first occurrence of sub in s at or afterfrom, or -1. An empty needle matches at from (mirrorsstr::find("")). |
| __txtFormatDecimal(v : i64) : string Decimal text of a signed 64-bit value. Builds the result in a single right-sized allocation, digits formatted in place (handles the minimum value without negating). |
| __stringLength(s : string) : i32 Number of UTF-8 characters in s. |
| __stringIsEmpty(s : string) : bool True when s holds no bytes. |
| __stringSubstring(s : string, start : i32, end : i32) : string Substring of the characters [start, end). RaisesIndexOutOfBoundsException when the range is negative, inverted, orpast the character count. |
| __stringSubstringFrom(s : string, start : i32) : string Substring from character start to the end. RaisesIndexOutOfBoundsException when start is negative or past thecharacter count. |
| __stringStartsWith(s : string, prefix : string) : bool True when s begins with prefix (byte comparison; empty prefix is always true). |
| __stringEndsWith(s : string, suffix : string) : bool True when s ends with suffix (byte comparison; empty suffix is always true). |
| __stringContains(s : string, sub : string) : bool True when sub occurs anywhere in s (empty sub always matches). |
| __stringIndexOf(s : string, sub : string) : i32 Character index of the first occurrence of sub, or -1. |
| __stringLastIndexOf(s : string, sub : string) : i32 Character index of the last occurrence of sub, or -1. |
| __stringReplace(s : string, from : string, to : string) : string Replace every non-overlapping occurrence of from with to. Anempty from inserts to around every character (mirroringstr::replace). |
| __stringReverse(s : string) : string Reverse by character (not by byte) so multi-byte UTF-8 stays valid. |
| __stringCharAt(s : string, i : i32) : i32 Unicode codepoint at character index i, or -1 if out of range. |
| __stringPadStart(s : string, len : i32, c : i32) : string Pad s on the left with character c until it spans lencharacters (no-op when already at least that long). |
| __stringPadEnd(s : string, len : i32, c : i32) : string Pad s on the right with character c until it spans lencharacters (no-op when already at least that long). |
| __stringHashCode(s : string) : i32 31-polynomial hash over the codepoints: h = 31*h + c, wrapping ini32 (the truncation each step matches the native wrapping_*).Reads length through the __string_byte_len compiler builtin like thecomparison family below it (see the note above __stringEquals), notthe strByteLen FFI shim. |
| __stringEquals(a : string, b : string) : bool Byte-exact equality (no case folding, no Unicode normalisation). Every == between two strings lands here, so it is the hottest loop astring-keyed table has. Once the lengths are known equal the comparison is a memory compare of a known extent, and it is done eight bytes at a time: one i64 load per side per word instead of one byte load per side per byte.The tail is a single overlapping word — re-reading bytes already found equal costs nothing and removes the loop. It does not delegate to bytesMatchAt, which is shared with indexOf /startsWith / endsWith / replace and must keep answering for an arbitraryoffset and a needle shorter than its haystack. |
| __stringEqualsIgnoreCase(a : string, b : string) : bool ASCII case-insensitive equality (only A-Z ↔ a-z are folded,matching eq_ignore_ascii_case). |
| __stringCompare(a : string, b : string) : i32 Lexicographic byte comparison: <0, 0, >0. A proper prefix sorts before the longer string. |
| __i32ToString(v : i32) : string Decimal text of a signed 32-bit value. |
| __i64ToString(v : i64) : string Decimal text of a signed 64-bit value. |
| __boolToString(v : bool) : string "true" or "false". |
| i64ToString(v : i64) : string Decimal text of a signed 64-bit value — the public entry point onto the module's shared decimal formatter, so a caller outside this module reuses one implementation instead of recreating it. n.toString() isthe spelling most code reaches for; this is the same formatter under a name a module can use. |
| f64ToString(v : f64) : string IEEE-754 round-trip text of v — the public entry point onto thenative __f64ToString shortest-round-trip formatter, so a calleroutside this module can format a float without touching the reserved __-prefixed primitive, which E0002 forbids it to name. |
Method detail
#stringFromBytes
Copy len bytes from a raw libc buffer into an owned string
(UTF-8, lossy). The read-side counterpart to passing a string
straight to a char* libc binding. The result owns its allocation,
so the source buffer may be freed immediately after.
buf raw source buffer to copy fromlen number of bytes to copy out of buf#bytesToString
Build an owned string from the first len bytes of a byte array.
The array → ptr decay required by the native bridge is confined
to this function's unsafe block.
buf byte array whose contents are copiedlen number of bytes to read from buf#bufferPtr
Decay a byte array to the typed pointer the scratch-buffer routines
(std::core::bytes) index. The array → ptr decay is confined here.
The result stays a pointer the whole way down: an i64 address would
force each access to rebuild a pointer with inttoptr, which costs
two extra instructions and, because the result has no provenance,
stops alias analysis from proving anything about the buffer.
buf byte array whose base address is taken#bufferAddr
The same base address as an i64, for the few sites that hand it
straight to a libc binding taking an opaque ptr (fread / fwrite
/ realpath) rather than indexing it.
Prefer bufferPtr for anything that reads or writes bytes: the
integer form loses provenance, and every access through it has to
rebuild a pointer. This exists only so the decay itself is still
written once.
buf byte array whose base address is taken#spanFits
Does the window buf[offset .. offset+len) lie inside buf?
The one place a caller-supplied (offset, len) is measured against the
buffer it addresses. Every site that reaches a byte through
[bufferAddr] has already left the array behind — the FFI call it hands
the address to sees a bare pointer and a count, and no bound at all. So
the measurement has to happen here, while buf is still an array and
still knows its own length.
It refuses rather than clamping. A short read is a legitimate answer
from a stream; a request reaching past the buffer is a caller that has
miscounted, and silently serving a prefix of it hides the miscount until
the data is wrong somewhere else. The addition is written as a
subtraction because Axle's + wraps: offset + len on two large values
is a small number that passes.
buf the array the window is taken fromoffset first index of the windowlen window length in bytes#addrAsPtr
Reinterpret a buffer address as the opaque ptr an FFI binding takes —
the read side of bufferAddr, for a call that needs an offset into a
buffer rather than its base.
Here, and only here, so the i64 -> ptr decay has one site the waybufferPtr gives the array -> ptr one its site. Prefer bufferPtr
whenever the result is indexed: an address that has been through an
integer carries no provenance, and every access through it has to
rebuild a pointer.
addr base address, from bufferAddr, plus any byte offset#__stringByteAt
Raw byte 0..255 at byte-index i, or -1 if out of range. The
primitive every hand-rolled scanner builds on; backed directly by the
runtime's bounds-checked accessor (strByteAt) since the runtime owns
the string layout.
s string to readi byte offset into s, 0-based#__string_concat
Runtime helper backing a + b when either operand is a string. The
result is a fresh owned string; the operands are left untouched.
a left operand, placed first in the resultb right operand, appended after a#__stringRepeat
s repeated n times ("" when n <= 0). Builds the result in a
single allocation filled in place — no scratch buffer, no re-copy.
s string to repeatn number of repetitions#__stringConcat
Backs the .concat(other) method — distinct from the + operator, which
lowers to the native __string_concat. Joins a and b into a fresh
owned string ("" when both are empty).
a left operand, placed first in the resultb right operand, appended after a#__stringIsBlank
True when s is empty or contains only Unicode whitespace.
s string to test for blankness#__stringTrim
Drop leading and trailing Unicode whitespace, returning an owned string.
s string to trim on both ends#__stringTrimStart
Drop leading Unicode whitespace, returning an owned string.
s string to trim on the leading end#__stringTrimEnd
Drop trailing Unicode whitespace, returning an owned string.
s string to trim on the trailing end#__stringToLowerCase
Full Unicode lowercase folding (locale-independent), as an owned string.
s string to lowercase#__stringToUpperCase
Full Unicode uppercase folding (locale-independent), as an owned string.
s string to uppercase#parseI32
Parse a base-10 i32, trimming surrounding Unicode whitespace first.
s decimal text to parseParseException when the text is empty or not a valid 32-bit integer.#parseI64
Parse a base-10 i64, trimming surrounding Unicode whitespace first.
s decimal text to parseParseException when the text is empty or not a valid 64-bit integer.#parseF64
Parse an IEEE-754 f64, trimming surrounding Unicode whitespace first.
s floating-point text to parseParseException when the text is not a valid float literal.#parseBool
Parse "true" / "false" (after trimming Unicode whitespace).
s bool text to parseParseException for any other text.#__stringToI32
s.toI32() — @throws ParseException when the text is not a 32-bit integer.
#__stringToI64
s.toI64() — @throws ParseException when the text is not a 64-bit integer.
#__stringToF64
s.toF64() — @throws ParseException when the text is not a float literal.
#__stringToBool
s.toBool() — @throws ParseException for anything but "true" / "false".
#__f64ToString
IEEE-754 shortest round-trip text of v (the value parseF64 inverts).
v f64 to render#__txtSubstr
Copy bytes [start, end) of s into an owned string.
s source stringstart inclusive start byte offset into s, 0-basedend exclusive end byte offset into s#__charCountPrefix
Count of UTF-8 characters in [0, byteEnd) — every byte that is not
a continuation byte (0b10xxxxxx) starts a new character. Backed by
the runtime's O(1) header read + one pass over the byte slice
(strCharCountPrefix), not a per-byte byteAt FFI loop (which madelength / indexOf / substring O(n) function calls on the length).
s source stringbyteEnd exclusive end byte offset of the prefix to count#__charByteIndex
Byte offset of the charIdx-th character in s (whose byte length
is srcLen): the string end when charIdx equals the character
count, or -1 when it exceeds it.
s source stringsrcLen byte length of scharIdx character index to locate, 0-based#__charWidth
UTF-8 width (1..4 bytes) of the character whose lead byte is lead.
A stray continuation byte counts as width 1 to keep scans advancing.
lead lead byte of a UTF-8 sequence#__decodeCodepoint
Decode the Unicode codepoint of the character starting atbyteOffset (assumed valid UTF-8).
s source stringbyteOffset byte offset of the character's lead byte#__txtPutUtf8
UTF-8 encode codepoint (invalid scalars folded to U+FFFD, matchingchar::from_u32(..).unwrap_or(..)) into buffer at offset,
returning the offset past the bytes written.
buffer destination buffer base addressoffset byte offset at which to write the encoded bytescodepoint Unicode scalar to encode#__byteIndexOf
Byte offset of the first occurrence of sub in s at or afterfrom, or -1. An empty needle matches at from (mirrorsstr::find("")).
s haystack stringsourceLen byte length of ssub needle substring to search forsubLen byte length of subfrom byte offset at which the search starts#__txtFormatDecimal
Decimal text of a signed 64-bit value. Builds the result in a single
right-sized allocation, digits formatted in place (handles the minimum
value without negating).
v signed integer to render in base 10#__stringLength
Number of UTF-8 characters in s.
s string to measure#__stringIsEmpty
True when s holds no bytes.
s string to test for emptiness#__stringSubstring
Substring of the characters [start, end). RaisesIndexOutOfBoundsException when the range is negative, inverted, or
past the character count.
s source stringstart inclusive start character index, 0-basedend exclusive end character index#__stringSubstringFrom
Substring from character start to the end. RaisesIndexOutOfBoundsException when start is negative or past the
character count.
s source stringstart inclusive start character index, 0-based#__stringStartsWith
True when s begins with prefix (byte comparison; empty prefix is always true).
s string to testprefix candidate leading substring#__stringEndsWith
True when s ends with suffix (byte comparison; empty suffix is always true).
s string to testsuffix candidate trailing substring#__stringContains
True when sub occurs anywhere in s (empty sub always matches).
s haystack stringsub needle substring to search for#__stringIndexOf
Character index of the first occurrence of sub, or -1.
s haystack stringsub needle substring to search for#__stringLastIndexOf
Character index of the last occurrence of sub, or -1.
s haystack stringsub needle substring to search for#__stringReplace
Replace every non-overlapping occurrence of from with to. An
empty from inserts to around every character (mirroringstr::replace).
s source stringfrom substring to search for and replaceto replacement substring#__stringReverse
Reverse by character (not by byte) so multi-byte UTF-8 stays valid.
s string to reverse#__stringCharAt
Unicode codepoint at character index i, or -1 if out of range.
s source stringi character index, 0-based#__stringPadStart
Pad s on the left with character c until it spans len
characters (no-op when already at least that long).
s string to padlen target length in charactersc Unicode codepoint of the pad character#__stringPadEnd
Pad s on the right with character c until it spans len
characters (no-op when already at least that long).
s string to padlen target length in charactersc Unicode codepoint of the pad character#__stringHashCode
31-polynomial hash over the codepoints: h = 31*h + c, wrapping ini32 (the truncation each step matches the native wrapping_*).
Reads length through the __string_byte_len compiler builtin like the
comparison family below it (see the note above __stringEquals), not
the strByteLen FFI shim.
s string to hash#__stringEquals
Byte-exact equality (no case folding, no Unicode normalisation).
Every == between two strings lands here, so it is the hottest loop a
string-keyed table has. Once the lengths are known equal the comparison is a
memory compare of a known extent, and it is done eight bytes at a time:
one i64 load per side per word instead of one byte load per side per byte.
The tail is a single overlapping word — re-reading bytes already found equal
costs nothing and removes the loop.
It does not delegate to bytesMatchAt, which is shared with indexOf /startsWith / endsWith / replace and must keep answering for an arbitrary
offset and a needle shorter than its haystack.
a first string to compareb second string to compare#__stringEqualsIgnoreCase
ASCII case-insensitive equality (only A-Z ↔ a-z are folded,
matching eq_ignore_ascii_case).
a first string to compareb second string to compare#__stringCompare
Lexicographic byte comparison: <0, 0, >0. A proper prefix sorts
before the longer string.
a left operand of the comparisonb right operand of the comparison#__i32ToString
Decimal text of a signed 32-bit value.
v i32 to render in base 10#__i64ToString
Decimal text of a signed 64-bit value.
v i64 to render in base 10#__boolToString
"true" or "false".
v bool to render#i64ToString
Decimal text of a signed 64-bit value — the public entry point onto
the module's shared decimal formatter, so a caller outside this module
reuses one implementation instead of recreating it. n.toString() is
the spelling most code reaches for; this is the same formatter under a
name a module can use.
v i64 to render in base 10#f64ToString
IEEE-754 round-trip text of v — the public entry point onto the
native __f64ToString shortest-round-trip formatter, so a caller
outside this module can format a float without touching the reserved__-prefixed primitive, which E0002 forbids it to name.
v f64 to renderC class StringBuilder
Growable UTF-8 byte buffer. data is an owned i8[] — the compiler
frees it through the synthesised destructor at scope exit, so no
manual release is needed. Mutating methods return self so chains
stay idiomatic. Typed appends delegate to the same-module
converters; toString() materialises an owned string.
Constructors
| Type | Method and description |
|---|---|
| constructor() Create an empty builder with a 16-byte backing buffer. |
Method detail
#constructor
Create an empty builder with a 16-byte backing buffer.
Methods
| Type | Method and description |
|---|---|
| _reserve(mut self, extra : i64) : void |
| _pushByte(mut self, b : i32) : void |
| append(mut self, s : string) : StringBuilder |
| appendChar(mut self, c : i32) : StringBuilder |
| _reverseRange(mut self, lo : i64, hi : i64) : void |
| _writeI64(mut self, v : i64) : void |
| appendI32(mut self, v : i32) : StringBuilder |
| appendI64(mut self, v : i64) : StringBuilder |
| appendF64(self, v : f64) : StringBuilder |
| appendBool(self, v : bool) : StringBuilder |
| length(self) : i32 |
| isEmpty(self) : bool |
| byteLength(self) : i64 |
| toString(self) : string |
| clear(mut self) : StringBuilder |
| reverse(mut self) : StringBuilder |
| dispose(mut self) : void |