std/core/bytes
std/core/bytes — shared byte-level primitives for the pure-Axle
stdlib.
The low-level toolbox every hand-rolled scanner / encoder / parser
reuses: raw scratch-buffer access, byte search over a string,
UTF-8 encoding of a scalar, strict UTF-8 validation, and range
comparison. Declared once here and shared by use std::core::bytes::…
so the same loop isn't rewritten in version / url / escape /csv / json / encoding / … . Single-character classification and
case / hex-digit conversion live in std/core/ascii.
Buffers are libc_malloc allocations addressed as a ptr<i8>; reads
and writes go through getByte / putByte, which index it. Astring's bytes are read with the byteAt builtin. Nothing here
allocates or builds a string (that needs stringFromBytes, which
would couple this module to std/text); callers materialise results
themselves.
The buffer travels as a pointer rather than an i64 so that every
access keeps the provenance of the allocation it came from. Carrying
the address as an integer forces (buffer + offset) as ptr<i8> at
each access — three instructions where a getelementptr is one, and
an inttoptr whose result alias analysis must treat as pointing
anywhere. text::bufferPtr is the one place an i8[] decays.
Free functions
| Type | Method and description |
|---|---|
| bytePtrFromAddr(addr : i64) : ptr<i8> View a raw address as the byte pointer the accessors below index. For the handful of routines whose scratch space is a stack array ( i32[N], byte-addressed) rather than a heap i8[]: a static arraydoes not decay to a pointer, so &scratch is the only way to nameits address and it arrives as a managed pointer. This is the single confined unsafe that turns one into the raw form; the heap path( text::bufferPtr) never goes through an integer at all. |
| putByte(buffer : ptr<i8>, offset : i64, value : i32) : void Store the low 8 bits of value at buffer + offset. |
| getByte(buffer : ptr<i8>, offset : i64) : i32 Load the unsigned byte 0..255 at buffer + offset. |
| putStr(buffer : ptr<i8>, offset : i64, lit : string) : i64 Copy every byte of the literal lit into buffer at offset,returning the offset just past the last byte written. |
| putStrLen(buffer : ptr<i8>, offset : i64, lit : string, litLen : i64) : i64 Like putStr but accepts a pre-computed byte length — avoids aredundant strByteLen call when the caller already knows it. |
| findByte(s : string, start : i64, end : i64, target : i32) : i64 Index of the first target byte in s within [start, end), or-1 if absent. |
| rfindByte(s : string, start : i64, end : i64, target : i32) : i64 Index of the last target byte in s within [start, end), or -1if absent. |
| findFirstOf3(
s : string,
start : i64,
end : i64,
targetA : i32,
targetB : i32,
targetC : i32
) : i64 Index of the first byte in [start, end) equal to a, b, or c,or end when none occurs. Pass an impossible value (e.g. -1) for anunused target. |
| bytesMatchAt(s : string, sub : string, subl : i64, at : i64) : bool true if the subl bytes of sub occur in s at byte offset at.The iteration bound subl is already validated by the caller, sothe __string_byte_at builtin's inline gep i8 + load i8 path issafe — it lowers to a direct memory read with no FFI hop and no -1 sentinel, unlike the codepoint-aware byteAt shim this loopformerly used. The two byteAt calls per iteration were opaque tothe vectoriser; the builtin form is fully inlinable. |
| byteRangeCompare(
a : string,
aStart : i64,
aEnd : i64,
b : string,
bStart : i64,
bEnd : i64
) : i32 Lexicographic byte comparison of a[aStart, aEnd) againstb[bStart, bEnd): <0, 0, >0. A proper prefix sorts before thelonger range. |
| putUtf8(buffer : ptr<i8>, offset : i64, codepoint : i32) : i64 UTF-8 encode the scalar value codepoint into buffer at offset,returning the offset past the encoded bytes. The caller must pass a valid scalar (non-negative, ≤ U+10FFFF, not a surrogate); passing a byte value 0..255 yields its U+00xx form (1 byte under 128, else the 2-byte sequence). |
| utf8Valid(buffer : ptr<i8>, len : i64) : bool Strict UTF-8 well-formedness check over the len bytes at buffer,matching the definition str::from_utf8 enforces: rejects overlongencodings ( C0/C1, E0 80..9F, F0 80..8F), UTF-16 surrogates( ED A0..BF), and codepoints above U+10FFFF (F4 90..BF, F5..FF).A decoder calls this before materialising a string, so malformedbytes raise rather than corrupt. |
Method detail
#bytePtrFromAddr
View a raw address as the byte pointer the accessors below index.
For the handful of routines whose scratch space is a stack array
(i32[N], byte-addressed) rather than a heap i8[]: a static array
does not decay to a pointer, so &scratch is the only way to name
its address and it arrives as a managed pointer. This is the single
confined unsafe that turns one into the raw form; the heap path
(text::bufferPtr) never goes through an integer at all.
addr address of a buffer that outlives every access made through it#putByte
Store the low 8 bits of value at buffer + offset.
buffer base address of the destination scratch bufferoffset byte offset from buffer of the slot to writevalue source word; only its low 8 bits are stored#getByte
Load the unsigned byte 0..255 at buffer + offset.
buffer base address of the source scratch bufferoffset byte offset from buffer of the slot to read#putStr
Copy every byte of the literal lit into buffer at offset,
returning the offset just past the last byte written.
buffer base address of the destination bufferoffset byte offset where the first byte of lit is writtenlit source string; its byte-length bytes are copied verbatim#putStrLen
Like putStr but accepts a pre-computed byte length — avoids a
redundant strByteLen call when the caller already knows it.
#findByte
Index of the first target byte in s within [start, end), or
-1 if absent.
s string whose bytes are scannedstart inclusive lower byte index of the search windowend exclusive upper byte index of the search windowtarget byte value 0..255 to match (compared masked with 0xFF)#rfindByte
Index of the last target byte in s within [start, end), or -1
if absent.
s string whose bytes are scanned backwardsstart inclusive lower byte index bounding the scanend exclusive upper byte index; scanning begins at end - 1target byte value 0..255 to match (compared masked with 0xFF)#findFirstOf3
Index of the first byte in [start, end) equal to a, b, or c,
or end when none occurs. Pass an impossible value (e.g. -1) for an
unused target.
s string whose bytes are scannedstart inclusive lower byte index of the search windowend exclusive upper byte index; also the not-found sentineltargetA first candidate byte 0..255 (or -1 to disable)targetB second candidate byte 0..255 (or -1 to disable)targetC third candidate byte 0..255 (or -1 to disable)end if none matches#bytesMatchAt
true if the subl bytes of sub occur in s at byte offset at.
The iteration bound subl is already validated by the caller, so
the __string_byte_at builtin's inline gep i8 + load i8 path is
safe — it lowers to a direct memory read with no FFI hop and no -1
sentinel, unlike the codepoint-aware byteAt shim this loop
formerly used. The two byteAt calls per iteration were opaque to
the vectoriser; the builtin form is fully inlinable.
s haystack string compared againstsub needle string whose leading bytes are matchedsubl number of bytes of sub to compareat byte offset in s where the comparison starts#byteRangeCompare
Lexicographic byte comparison of a[aStart, aEnd) againstb[bStart, bEnd): <0, 0, >0. A proper prefix sorts before the
longer range.
a left-hand stringaStart inclusive lower byte index of the left rangeaEnd exclusive upper byte index of the left rangeb right-hand stringbStart inclusive lower byte index of the right rangebEnd exclusive upper byte index of the right range#putUtf8
UTF-8 encode the scalar value codepoint into buffer at offset,
returning the offset past the encoded bytes. The caller must pass a
valid scalar (non-negative, ≤ U+10FFFF, not a surrogate); passing a
byte value 0..255 yields its U+00xx form (1 byte under 128, else the
2-byte sequence).
buffer base address of the destination bufferoffset byte offset where the encoded sequence beginscodepoint Unicode scalar to encode; must be ≤ U+10FFFF, non-surrogate#utf8Valid
Strict UTF-8 well-formedness check over the len bytes at buffer,
matching the definition str::from_utf8 enforces: rejects overlong
encodings (C0/C1, E0 80..9F, F0 80..8F), UTF-16 surrogates
(ED A0..BF), and codepoints above U+10FFFF (F4 90..BF, F5..FF).
A decoder calls this before materialising a string, so malformed
bytes raise rather than corrupt.
buffer base address of the byte run to validatelen number of bytes from buffer to check