Axle v0.14.1
Package

std/core/bytes

std/core/bytes — shared byte-level primitives for the pure-Axle
stdlib.

The low-level toolbox every hand-rolled scanner / encoder / parser
reuses: raw scratch-buffer access, byte search over a string,
UTF-8 encoding of a scalar, strict UTF-8 validation, and range
comparison. Declared once here and shared by use std::core::bytes::…
so the same loop isn't rewritten in version / url / escape /
csv / json / encoding / … . Single-character classification and
case / hex-digit conversion live in std/core/ascii.

Buffers are libc_malloc allocations addressed as a ptr<i8>; reads
and writes go through getByte / putByte, which index it. A
string's bytes are read with the byteAt builtin. Nothing here
allocates or builds a string (that needs stringFromBytes, which
would couple this module to std/text); callers materialise results
themselves.

The buffer travels as a pointer rather than an i64 so that every
access keeps the provenance of the allocation it came from. Carrying
the address as an integer forces (buffer + offset) as ptr<i8> at
each access — three instructions where a getelementptr is one, and
an inttoptr whose result alias analysis must treat as pointing
anywhere. text::bufferPtr is the one place an i8[] decays.

Free functions

TypeMethod and description
ptr<i8>
bytePtrFromAddr(addr : i64) : ptr<i8> View a raw address as the byte pointer the accessors below index.

For the handful of routines whose scratch space is a stack array
(i32[N], byte-addressed) rather than a heap i8[]: a static array
does not decay to a pointer, so &scratch is the only way to name
its address and it arrives as a managed pointer. This is the single
confined unsafe that turns one into the raw form; the heap path
(text::bufferPtr) never goes through an integer at all.
void
putByte(buffer : ptr<i8>, offset : i64, value : i32) : void Store the low 8 bits of value at buffer + offset.
i32
getByte(buffer : ptr<i8>, offset : i64) : i32 Load the unsigned byte 0..255 at buffer + offset.
i64
putStr(buffer : ptr<i8>, offset : i64, lit : string) : i64 Copy every byte of the literal lit into buffer at offset,
returning the offset just past the last byte written.
i64
putStrLen(buffer : ptr<i8>, offset : i64, lit : string, litLen : i64) : i64 Like putStr but accepts a pre-computed byte length — avoids a
redundant strByteLen call when the caller already knows it.
i64
findByte(s : string, start : i64, end : i64, target : i32) : i64 Index of the first target byte in s within [start, end), or
-1 if absent.
i64
rfindByte(s : string, start : i64, end : i64, target : i32) : i64 Index of the last target byte in s within [start, end), or -1
if absent.
i64
findFirstOf3( s : string, start : i64, end : i64, targetA : i32, targetB : i32, targetC : i32 ) : i64 Index of the first byte in [start, end) equal to a, b, or c,
or end when none occurs. Pass an impossible value (e.g. -1) for an
unused target.
bool
bytesMatchAt(s : string, sub : string, subl : i64, at : i64) : bool true if the subl bytes of sub occur in s at byte offset at.

The iteration bound subl is already validated by the caller, so
the __string_byte_at builtin's inline gep i8 + load i8 path is
safe — it lowers to a direct memory read with no FFI hop and no -1
sentinel, unlike the codepoint-aware byteAt shim this loop
formerly used. The two byteAt calls per iteration were opaque to
the vectoriser; the builtin form is fully inlinable.
i32
byteRangeCompare( a : string, aStart : i64, aEnd : i64, b : string, bStart : i64, bEnd : i64 ) : i32 Lexicographic byte comparison of a[aStart, aEnd) against
b[bStart, bEnd): <0, 0, >0. A proper prefix sorts before the
longer range.
i64
putUtf8(buffer : ptr<i8>, offset : i64, codepoint : i32) : i64 UTF-8 encode the scalar value codepoint into buffer at offset,
returning the offset past the encoded bytes. The caller must pass a
valid scalar (non-negative, ≤ U+10FFFF, not a surrogate); passing a
byte value 0..255 yields its U+00xx form (1 byte under 128, else the
2-byte sequence).
bool
utf8Valid(buffer : ptr<i8>, len : i64) : bool Strict UTF-8 well-formedness check over the len bytes at buffer,
matching the definition str::from_utf8 enforces: rejects overlong
encodings (C0/C1, E0 80..9F, F0 80..8F), UTF-16 surrogates
(ED A0..BF), and codepoints above U+10FFFF (F4 90..BF, F5..FF).
A decoder calls this before materialising a string, so malformed
bytes raise rather than corrupt.

Method detail

#bytePtrFromAddr

bytePtrFromAddr(addr : i64) : ptr<i8>

View a raw address as the byte pointer the accessors below index.

For the handful of routines whose scratch space is a stack array
(i32[N], byte-addressed) rather than a heap i8[]: a static array
does not decay to a pointer, so &scratch is the only way to name
its address and it arrives as a managed pointer. This is the single
confined unsafe that turns one into the raw form; the heap path
(text::bufferPtr) never goes through an integer at all.

Parameters
addr address of a buffer that outlives every access made through it

#putByte

putByte(buffer : ptr<i8>, offset : i64, value : i32) : void

Store the low 8 bits of value at buffer + offset.

Parameters
buffer base address of the destination scratch buffer
offset byte offset from buffer of the slot to write
value source word; only its low 8 bits are stored

#getByte

getByte(buffer : ptr<i8>, offset : i64) : i32

Load the unsigned byte 0..255 at buffer + offset.

Parameters
buffer base address of the source scratch buffer
offset byte offset from buffer of the slot to read

#putStr

putStr(buffer : ptr<i8>, offset : i64, lit : string) : i64

Copy every byte of the literal lit into buffer at offset,
returning the offset just past the last byte written.

Parameters
buffer base address of the destination buffer
offset byte offset where the first byte of lit is written
lit source string; its byte-length bytes are copied verbatim

#putStrLen

putStrLen(buffer : ptr<i8>, offset : i64, lit : string, litLen : i64) : i64

Like putStr but accepts a pre-computed byte length — avoids a
redundant strByteLen call when the caller already knows it.

#findByte

findByte(s : string, start : i64, end : i64, target : i32) : i64

Index of the first target byte in s within [start, end), or
-1 if absent.

Parameters
s string whose bytes are scanned
start inclusive lower byte index of the search window
end exclusive upper byte index of the search window
target byte value 0..255 to match (compared masked with 0xFF)
Returns byte index of the first matching byte, or -1 if none in the window

#rfindByte

rfindByte(s : string, start : i64, end : i64, target : i32) : i64

Index of the last target byte in s within [start, end), or -1
if absent.

Parameters
s string whose bytes are scanned backwards
start inclusive lower byte index bounding the scan
end exclusive upper byte index; scanning begins at end - 1
target byte value 0..255 to match (compared masked with 0xFF)
Returns byte index of the last matching byte, or -1 if none in the window

#findFirstOf3

findFirstOf3( s : string, start : i64, end : i64, targetA : i32, targetB : i32, targetC : i32 ) : i64

Index of the first byte in [start, end) equal to a, b, or c,
or end when none occurs. Pass an impossible value (e.g. -1) for an
unused target.

Parameters
s string whose bytes are scanned
start inclusive lower byte index of the search window
end exclusive upper byte index; also the not-found sentinel
targetA first candidate byte 0..255 (or -1 to disable)
targetB second candidate byte 0..255 (or -1 to disable)
targetC third candidate byte 0..255 (or -1 to disable)
Returns byte index of the first matching byte, or end if none matches

#bytesMatchAt

bytesMatchAt(s : string, sub : string, subl : i64, at : i64) : bool

true if the subl bytes of sub occur in s at byte offset at.

The iteration bound subl is already validated by the caller, so
the __string_byte_at builtin's inline gep i8 + load i8 path is
safe — it lowers to a direct memory read with no FFI hop and no -1
sentinel, unlike the codepoint-aware byteAt shim this loop
formerly used. The two byteAt calls per iteration were opaque to
the vectoriser; the builtin form is fully inlinable.

Parameters
s haystack string compared against
sub needle string whose leading bytes are matched
subl number of bytes of sub to compare
at byte offset in s where the comparison starts

#byteRangeCompare

byteRangeCompare( a : string, aStart : i64, aEnd : i64, b : string, bStart : i64, bEnd : i64 ) : i32

Lexicographic byte comparison of a[aStart, aEnd) against
b[bStart, bEnd): <0, 0, >0. A proper prefix sorts before the
longer range.

Parameters
a left-hand string
aStart inclusive lower byte index of the left range
aEnd exclusive upper byte index of the left range
b right-hand string
bStart inclusive lower byte index of the right range
bEnd exclusive upper byte index of the right range

#putUtf8

putUtf8(buffer : ptr<i8>, offset : i64, codepoint : i32) : i64

UTF-8 encode the scalar value codepoint into buffer at offset,
returning the offset past the encoded bytes. The caller must pass a
valid scalar (non-negative, ≤ U+10FFFF, not a surrogate); passing a
byte value 0..255 yields its U+00xx form (1 byte under 128, else the
2-byte sequence).

Parameters
buffer base address of the destination buffer
offset byte offset where the encoded sequence begins
codepoint Unicode scalar to encode; must be ≤ U+10FFFF, non-surrogate

#utf8Valid

utf8Valid(buffer : ptr<i8>, len : i64) : bool

Strict UTF-8 well-formedness check over the len bytes at buffer,
matching the definition str::from_utf8 enforces: rejects overlong
encodings (C0/C1, E0 80..9F, F0 80..8F), UTF-16 surrogates
(ED A0..BF), and codepoints above U+10FFFF (F4 90..BF, F5..FF).
A decoder calls this before materialising a string, so malformed
bytes raise rather than corrupt.

Parameters
buffer base address of the byte run to validate
len number of bytes from buffer to check