core/text
Shared: available on every enabled target.
Byte cursors, validated UTF-8 views, byte search and bounded decimal conversion.
Cursors advance by bytes. The language provides separate UTF-8 indexing and traversal rules; this module does not turn byte offsets into validated character boundaries. UTF-8 equality, hashing and ordering use encoded bytes, with no normalization or locale collation. Views borrow their input.
Items
- position type
- invalid_number atom
- invalid_text atom
- no_space atom
- number_overflow atom
- past_end atom
- nowhere value
- advance function
- at function
- at_end function
- byte function
- bytes function
- contains function
- copied function
- eq function
- first function
- from_bytes function
- from_c function
- hash function
- less function
- next function
- ordinal function
- slice function
- to_u32 function
- write_byte function
- write_u32 function
- written function
Executable example
This complete program is maintained in the repository runtime tests. View source.
import core/text exercise: () -> (ok: bool) ! text.past_end = ok = false source: [1]u8 = [42] first := text.first(source[0..<1]) found := try text.byte(source[0..<1], first) return when found <> 42 _ = text.byte(source[0..<1], text.next(first)) else (problem) ok = problem == text.past_end return end end exercise public main: () -> (code: i32) = code = 1 ok := exercise() else false if ok then code = 42 end if end main
position type
public position: type = position_valueOpaque byte offset for scanning. It does not retain or validate a source; when used to index UTF-8 it must identify a codepoint boundary.
invalid_number atom
public invalid_number: atomThe text is empty or contains a non-decimal digit where an unsigned integer is required.
invalid_text atom
public invalid_text: atomThe supplied bytes are not well-formed UTF-8.
no_space atom
public no_space: atomThe destination does not have room for the complete requested output.
number_overflow atom
public number_overflow: atomThe parsed unsigned decimal value exceeds u32.
past_end atom
public past_end: atomA byte read was requested at or beyond the end of its source.
nowhere value
public nowhere: position = (offset: 0)The zero-offset position sentinel. It has the same offset as first; it is not a separate invalid-position tag.
advance function
public advance: (inout cursor: position) -> noneAdvance the cursor by one byte in place. Does not validate the source or skip a whole UTF-8 scalar.
at function
public at: (source: []u8, offset: usize) -> (result: position)Construct a cursor with the given byte offset. Does not check bounds or UTF-8 boundaries; later operations enforce their own requirements.
Reify an ordinal as an opaque position for byte-oriented scanners. The source argument makes the intended coordinate space explicit even though The compact representation stores only the offset.
at_end function
public at_end: (source: []u8, cursor: position) -> (yes: bool)Return whether the byte offset is at or beyond the source length.
byte function
public byte: (source: []u8, cursor: position)
-> (value: u8) ! past_endRead one source byte at the cursor. Reports past_end before reading outside the source.
bytes function
public bytes: (source: utf8) -> (result: []u8 from source)Lend the encoded bytes of a UTF-8 view without copying or allocating.
contains function
public contains: (source: utf8, sought: utf8) -> (yes: bool)Test whether the sought UTF-8 byte sequence occurs in the source. Empty text matches; no normalization or allocation is performed.
Two-Way search uses a critical split of the sought bytes. Its two lexicographic passes and the search each take linear time; only scalar counters are kept, so neither a table nor a caller allocator is needed.
copied function
public copied: (cursor: position) -> (result: position)Copy a cursor value without retaining the aggregate from which it was read.
Copy a position out of aggregate storage without borrowing that storage. This matters to scanners that retain the token start while advancing the cursor field from which it was read.
eq function
public eq: (left: utf8, right: utf8) -> (same: bool)Compare UTF-8 values by their encoded bytes. Does not normalize Unicode spelling.
Equality and containment are exact encoded-byte operations. Valid UTF-8 has one shortest-form encoding per scalar sequence, so neither operation needs normalization or locale authority.
first function
public first: (source: []u8) -> (result: position)Return a cursor at byte offset zero, including for an empty source.
from_bytes function
public from_bytes: (source: []u8)
-> (result: utf8 from source) ! invalid_textValidate a byte slice and lend it as UTF-8 without allocation. Reports invalid_text; the caller keeps the bytes alive and valid while the view is used.
The checked adapters report foreign or file-data defects before the ordinary conversion establishes utf8's invariant. Their result is a source-derived view; no allocation, copy or canonical empty carrier is introduced.
from_c function
public from_c: (source: cstring)
-> (result: utf8 from source) ! invalid_textValidate a terminated C string and lend it as UTF-8. The pointer must remain readable through its terminator; invalid UTF-8 reports invalid_text.
hash function
public hash: (value: utf8) -> (result: u64)Hash UTF-8 encoded bytes consistently with eq and the byte-slice hash. No Unicode normalization or cryptographic guarantee is provided.
utf8 is a map key and sortable as it stands: equality, hash and order are over the encoded bytes, which valid UTF-8 makes unique per scalar sequence. The hash is 64-bit FNV-1a, the same as core/map's for []u8.
less function
public less: (left: utf8, right: utf8) -> (yes: bool)Compare encoded bytes lexicographically, with a shorter equal prefix sorting first. This is not locale-aware collation.
Byte-wise lexicographic order, which for valid UTF-8 is scalar order.
next function
public next: (cursor: position) -> (result: position)Return a cursor one byte later. This is byte traversal, not Unicode-scalar traversal; offset overflow follows ordinary checked arithmetic.
ordinal function
public ordinal: (cursor: position) -> (offset: usize)Return the byte offset stored in a position.
slice function
public slice: (source: []u8, begins: position, ends: position)
-> (result: []u8 from source)Lend bytes between two cursor offsets, with an exclusive end. Requires an ordered interval within the source; invalid ranges follow the language's checked range behavior.
to_u32 function
public to_u32: (source: utf8)
-> (result: u32) ! invalid_number | number_overflowParse nonempty ASCII decimal digits into u32. Rejects signs, whitespace and non-digits with invalid_number, and excessive values with number_overflow.
write_byte function
public write_byte: (inout into: []mut u8, used: usize, value: u8)
-> (next: usize) ! no_spaceAppend one byte at used and return the new used count. Checks capacity before writing and reports no_space without partial output.
Both writers preflight capacity before their first mutation. A no_space refusal therefore preserves every caller-owned byte and the written count, including the multi-byte decimal operation.
write_u32 function
public write_u32: (inout into: []mut u8, used: usize, value: u32)
-> (next: usize) ! no_spaceAppend an unsigned decimal representation and return the new used count. Checks space for the whole number before writing; no_space leaves the buffer unchanged.
written function
public written: (source: []u8, used: usize)
-> (result: []u8 from source) ! no_spaceLend the byte prefix of length used. Requires used within the source length; it does not validate UTF-8.