Landin library reference source

core/text

Shared: available on every enabled target.

Byte cursors, validated UTF-8 views, byte search and bounded decimal conversion.

Cursors advance by bytes. The language provides separate UTF-8 indexing and traversal rules; this module does not turn byte offsets into validated character boundaries. UTF-8 equality, hashing and ordering use encoded bytes, with no normalization or locale collation. Views borrow their input.

Items

Executable example

This complete program is maintained in the repository runtime tests. View source.

import core/text

exercise: () -> (ok: bool) ! text.past_end =
    ok = false
    source: [1]u8 = [42]
    first := text.first(source[0..<1])
    found := try text.byte(source[0..<1], first)
    return when found <> 42
    _ = text.byte(source[0..<1], text.next(first)) else (problem)
        ok = problem == text.past_end
        return
    end
end exercise

public main: () -> (code: i32) =
    code = 1
    ok := exercise() else false
    if ok then
        code = 42
    end if
end main

position type

public position: type = position_value

core/text/text.ldn:17

Opaque byte offset for scanning. It does not retain or validate a source; when used to index UTF-8 it must identify a codepoint boundary.

invalid_number atom

public invalid_number: atom

core/text/text.ldn:24

The text is empty or contains a non-decimal digit where an unsigned integer is required.

invalid_text atom

public invalid_text: atom

core/text/text.ldn:21

The supplied bytes are not well-formed UTF-8.

no_space atom

public no_space: atom

core/text/text.ldn:28

The destination does not have room for the complete requested output.

number_overflow atom

public number_overflow: atom

core/text/text.ldn:26

The parsed unsigned decimal value exceeds u32.

past_end atom

public past_end: atom

core/text/text.ldn:19

A byte read was requested at or beyond the end of its source.

nowhere value

public nowhere: position = (offset: 0)

core/text/text.ldn:34

The zero-offset position sentinel. It has the same offset as first; it is not a separate invalid-position tag.

advance function

public advance: (inout cursor: position) -> none

core/text/text.ldn:81

Advance the cursor by one byte in place. Does not validate the source or skip a whole UTF-8 scalar.

at function

public at: (source: []u8, offset: usize) -> (result: position)

core/text/text.ldn:68

Construct a cursor with the given byte offset. Does not check bounds or UTF-8 boundaries; later operations enforce their own requirements.

Reify an ordinal as an opaque position for byte-oriented scanners. The source argument makes the intended coordinate space explicit even though The compact representation stores only the offset.

at_end function

public at_end: (source: []u8, cursor: position) -> (yes: bool)

core/text/text.ldn:43

Return whether the byte offset is at or beyond the source length.

byte function

public byte: (source: []u8, cursor: position)
             -> (value: u8) ! past_end

core/text/text.ldn:87

Read one source byte at the cursor. Reports past_end before reading outside the source.

bytes function

public bytes: (source: utf8) -> (result: []u8 from source)

core/text/text.ldn:199

Lend the encoded bytes of a UTF-8 view without copying or allocating.

contains function

public contains: (source: utf8, sought: utf8) -> (yes: bool)

core/text/text.ldn:233

Test whether the sought UTF-8 byte sequence occurs in the source. Empty text matches; no normalization or allocation is performed.

Two-Way search uses a critical split of the sought bytes. Its two lexicographic passes and the search each take linear time; only scalar counters are kept, so neither a table nor a caller allocator is needed.

copied function

public copied: (cursor: position) -> (result: position)

core/text/text.ldn:58

Copy a cursor value without retaining the aggregate from which it was read.

Copy a position out of aggregate storage without borrowing that storage. This matters to scanners that retain the token start while advancing the cursor field from which it was read.

eq function

public eq: (left: utf8, right: utf8) -> (same: bool)

core/text/text.ldn:209

Compare UTF-8 values by their encoded bytes. Does not normalize Unicode spelling.

Equality and containment are exact encoded-byte operations. Valid UTF-8 has one shortest-form encoding per scalar sequence, so neither operation needs normalization or locale authority.

first function

public first: (source: []u8) -> (result: position)

core/text/text.ldn:37

Return a cursor at byte offset zero, including for an empty source.

from_bytes function

public from_bytes: (source: []u8)
                   -> (result: utf8 from source) ! invalid_text

core/text/text.ldn:183

Validate a byte slice and lend it as UTF-8 without allocation. Reports invalid_text; the caller keeps the bytes alive and valid while the view is used.

The checked adapters report foreign or file-data defects before the ordinary conversion establishes utf8's invariant. Their result is a source-derived view; no allocation, copy or canonical empty carrier is introduced.

from_c function

public from_c: (source: cstring)
               -> (result: utf8 from source) ! invalid_text

core/text/text.ldn:192

Validate a terminated C string and lend it as UTF-8. The pointer must remain readable through its terminator; invalid UTF-8 reports invalid_text.

hash function

public hash: (value: utf8) -> (result: u64)

core/text/text.ldn:416

Hash UTF-8 encoded bytes consistently with eq and the byte-slice hash. No Unicode normalization or cryptographic guarantee is provided.

utf8 is a map key and sortable as it stands: equality, hash and order are over the encoded bytes, which valid UTF-8 makes unique per scalar sequence. The hash is 64-bit FNV-1a, the same as core/map's for []u8.

less function

public less: (left: utf8, right: utf8) -> (yes: bool)

core/text/text.ldn:431

Compare encoded bytes lexicographically, with a shorter equal prefix sorting first. This is not locale-aware collation.

Byte-wise lexicographic order, which for valid UTF-8 is scalar order.

next function

public next: (cursor: position) -> (result: position)

core/text/text.ldn:75

Return a cursor one byte later. This is byte traversal, not Unicode-scalar traversal; offset overflow follows ordinary checked arithmetic.

ordinal function

public ordinal: (cursor: position) -> (offset: usize)

core/text/text.ldn:48

Return the byte offset stored in a position.

slice function

public slice: (source: []u8, begins: position, ends: position)
              -> (result: []u8 from source)

core/text/text.ldn:96

Lend bytes between two cursor offsets, with an exclusive end. Requires an ordered interval within the source; invalid ranges follow the language's checked range behavior.

to_u32 function

public to_u32: (source: utf8)
               -> (result: u32) ! invalid_number | number_overflow

core/text/text.ldn:343

Parse nonempty ASCII decimal digits into u32. Rejects signs, whitespace and non-digits with invalid_number, and excessive values with number_overflow.

write_byte function

public write_byte: (inout into: []mut u8, used: usize, value: u8)
                   -> (next: usize) ! no_space

core/text/text.ldn:367

Append one byte at used and return the new used count. Checks capacity before writing and reports no_space without partial output.

Both writers preflight capacity before their first mutation. A no_space refusal therefore preserves every caller-owned byte and the written count, including the multi-byte decimal operation.

write_u32 function

public write_u32: (inout into: []mut u8, used: usize, value: u32)
                  -> (next: usize) ! no_space

core/text/text.ldn:377

Append an unsigned decimal representation and return the new used count. Checks space for the whole number before writing; no_space leaves the buffer unchanged.

written function

public written: (source: []u8, used: usize)
                -> (result: []u8 from source) ! no_space

core/text/text.ldn:404

Lend the byte prefix of length used. Requires used within the source length; it does not validate UTF-8.