Landin the fixtures source

Shared fixtures

Fixtures live here rather than under compiler/ada/ because they describe Landin, not the Ada implementation that currently checks them. When a stage is eventually rewritten, these must still be the tests it has to pass.

Layout

compiler/tests/
  fixtures/<class>/<name>/fixture.meta   the fixture and its metadata
  harness-cases/malformed/               trees that must be rejected; see
                                         harness-cases/README.md
  fuzz/                                  a seeded corpus mutator that drives the
                                         server, and what it found; see fuzz/README.md
  server/<name>/session.lsp              a scripted language-server session, and its
                                         workspace; see "Language-server sessions"
  registers.md                           source: the four evidence registers the matrices are generated from
  constructs.matrix                      generated: every [NNNN], its evidence and inventory
  diagnostics.catalogue                  generated: every code and its rule
  diagnostics.matrix                     generated: code contracts, emitters and owners
  guarantees.matrix                      generated: classified semantic boundaries
  conformances.matrix                    generated: conformance/evidence mechanisms
  prototypes.matrix                      generated: derivations, their oracles and target results
  targets.matrix                         generated: applicability of every fixture
  lexical.tokens                         generated: the scanned corpus
  layout.targets                         recorded: what each target measures
  lowering.ir                            recorded: every positive fixture, lowered

Generated and recorded are not the same word here. check.py writes the generated files and refuses each when it is stale; the last two are written by ./scripts/test.sh --record, because producing them means running compiler stages and asking the target model, which check.py cannot do. It will not tell you those two are stale — the harness and the gate will. constructs.matrix began as the list of every construct either document defines against what the corpus says about it, and is now the construct inventory: each row also gives the strongest claim per product target, read from fixture metadata together with darwin/parity.json, cortex-m/corpus.json and driver/fixture.json; every named refusal with whether its note says the form is a recorded boundary, withdrawn or transferred to a named successor; and the state, applicable targets, gaps and open owner that the construct inventory records, whose disposition there explains the row. A full check.py refuses a missing, stale, unowned or unexplained row, and scripts/tests/test_construct_inventory.py proves each refusal fires. Regenerate it with python3 check.py --matrix.

Fixture classes, and the directory each uses:

classdirectorywhat it covers
unitunita note of one behaviour an implementation-side case covers
positivepositivea program that must be accepted
negativenegativea program that must be rejected, with the codes it must produce and, where a code alone could hide a wrong refusal, the exact report
runtimeruntimea program whose behaviour when run is the assertion
ABIabiemitted Landin assembly compiled with ordered C11 companions, then executed
end-to-endend-to-endthe toolchain from source to result

The scripted source-debugger programs live in debugging/ and run through scripts/debug.sh, independently of the Ada fixture harness. The debugger metadata class has no directory; these sessions use debugger assertions rather than the harness's process-output fixture contract.

Focused developer runs

The harness can select one suite, one case, or one recorded fixture by exact name. The developer wrapper combines that with checksum-based minimum recompilation:

./scripts/dev-test.sh --suite='fixture execution'
./scripts/dev-test.sh --case='harness/filters select exact cases'
./scripts/dev-test.sh --fixture=positive/variant-match-exhaustive
./scripts/dev-test.sh --fixture=negative/variant-match-duplicate
./scripts/dev-test.sh --fixture=runtime/variant-match-selects-tag
./scripts/dev-test.sh --fixture=abi/r440-smoke

A fixture selector accepts positive, negative, runtime, abi, or any other discovered class whose fixture has a recorded expect. It invokes the real scanner-through-backend path appropriate to that class, including assembling, linking and executing a runtime or ABI fixture. Every selected transcript begins with FILTERED, and an unknown selection fails: focused feedback cannot look like the complete suite by accident. On macOS use --host, optionally with one --suite or --case; native workload cases remain excluded and cannot be selected through that combination. --fixture and recording cannot combine with --host. Run Linux workloads on a native Linux host and Darwin workloads natively; a filtered run is never a verdict on the complete suite.

./scripts/dev-test.sh --host runs every compiler case except the two native target-workload emission/execution cases. Its HOST-ONLY banner identifies the scope; compiler units, diagnostics and the complete IR golden remain included.

Lanes

A run executes one product target's corpus, its lane: the fixtures whose targets: name it are emitted, compiled, linked and run, and every other is left to its own lane. The lane is the compiler's own host (D257) unless --target=NAME names another, with refine's own name for it, so the same command is the Linux x86-64 lane on an x86-64 runner and the Linux arm64 lane on an arm64 one, and nothing in the harness asks the machine which. Either option may accompany any selection, and the counts a lane is held to are its own fixtures'. A recorded fixture whose args: name a target is that target's verdict whatever the host, and every lane runs it.

A lane that is not the host's is a cross run. Its transcript begins LANE CROSS NAME, every compilation names the target, and an executable runs through --runner=PROGRAM and links with --toolchain=DRIVER:

./scripts/dev-test.sh --target=linux-arm64 --runner=qemu-aarch64 \
    --toolchain=aarch64-unknown-linux-gnu-gcc --suite='fixture execution'

That is QEMU's user emulation with nixpkgs' cross driver, which is evidence that the emitted code is right and none about the pinned toolchain on its own host; the gate's ubuntu-24.04-arm jobs are that. A level above the default is not asked of the host on a cross run, since the emulator offers every level its lane has. The retired routine acceptance policy used this scope for the debug compiler and ran the complete runtime/ABI matrix with the release compiler; milestone acceptance ran the complete suite in both modes. Those exact-revision acceptance runs are retired. The current .github/workflows/gate.yml has separate Linux compiler and release jobs; each runs unfiltered ./scripts/test.sh, so the complete suite runs in both debug and release modes. The gate is a safety net, not an exact-revision acceptance verdict. See docs/process.md.

D213's r490-distinct-* fixtures cover exact base construction/extraction, opaque identity, no inherited operators or conformances, ordinary identifiers, generic/fixed identity keys, erased dispatch, origins and module image cycles. The two generic runtime fixtures run all six profiles; the module-image fixture checks scalar/float/atom images, arrays, nested records, slices, text and callback relocations. abi/r490-distinct-c-roundtrip crosses the native C boundary in both directions for integer, float, pointer and mixed C-record bases. The IR atom-image unit case accepts a member identity and rejects a stored identity absent from the field's set.

The construction regressions distinguish runtime field/payload/fill arguments from static type arguments. negative/r491-construction-type-arguments pins value diagnostics in module and local contexts. negative/r491-construction-type-fills pins expression diagnostics for type-only trailing fills and preserves recovery through later declarations. negative/r491-construction-static-address preserves the existing static-image address exclusion. positive/r491-construction-static-arguments retains generic type arguments, local addresses, value fills and a callback body that takes a local address. runtime/r491-variant-array-construction exercises small root, nested and wrapped arrays, repetition and later variant replacement. The checker and lowering seam cases separately assert case identities and verified storage paths; driver cases assert that refused source writes no output and invokes no tool.

The checker case nested calls retain flow effects uses paired small sources to pin labelled-call assignment checks, nested sink and try effects, descriptor reads and assignment-destination order. Slice descriptor reads preserve independent element liveness. It preserves unevaluated fixed-array measurements and separately checked anonymous bodies. The driver refusal case requires the corresponding invalid sources to produce one L0302 before any output or tool invocation, for both assembly and executable requests.

The checker case joined destinations keep escape obligations covers body-local joins of frame, parameter, module and untracked destinations, with accepted same-origin and escaping controls. Two driver refusal sources require one L0314 and no output or tool effects. assigned children cover element descendants checks whole-child initialization and branch containment in both orders, preserving independent siblings, indices and consumed leaves. These are small compiler checks; they neither assemble nor execute Landin programs.

try failures check reference cleanups checks origins on propagated failure after call arguments run. Its 14 small sources preserve escaping/static arguments, recovered calls, early transfers and success-path isolation while refusing frame/non-escaping references retained by failure cleanups. The driver refusal case also requires L0314 before any output or tool call.

negative/r491-operand-diagnostic-cascades pins one type diagnostic for each invalid float remainder or shift while retaining an independent integer division-by-zero diagnostic inside a refused operator.

runtime/r491-function-final-values covers statement prefixes followed by scalar, aggregate, array, pointer, slice, function and multiple-result values. It includes generic and anonymous functions, fallible calls, named assignments followed by none-returning calls or controls, and cleanup ordering. negative/r491-function-final-value-refusals preserves assignment, origin and complete-value checks. negative/r491-function-final-value-prefix preserves the grammar's exclusion of unconditional exits and unchecked regions from a final expression's statement prefix.

The platform cases native timeout stops descendants and native arguments and capture are preserved exercise the real host process adapter. The timeout witness in tool_process_probe.py forks one short-lived child and checks that it cannot write a delayed marker after the runner stops the group. These tests invoke no compiler, assembler or linker. Existing native cases retain ordinary exit, signal and default-capture coverage.

The readiness regressions cover static concept labels in both instantiation orders and a separately renamed receiver; computed callees retain unhandled/undeclared error diagnostics, inference, propagation and recovery. The driver exercises those refusals and over-limit binary chains with writes and tools disabled. The parser checks the binary-chain boundary and following-declaration recovery. A 4,000-link ordinary alias chain checks iterative settlement, and runtime/r491-cleanup-verifier-storage preserves pop-before-run behavior for 12 guarded cleanups across the standard profiles. Those fixed controls do not claim unlimited compilation size or linear cleanup expansion.

Optimization profiles and object quality

Every runtime and ABI fixture runs separately under none/off, size/off, size/auto and speed/auto (objective/specialization). Every runtime and ABI fixture explicitly records profiles: standard for those four, or profiles: specialization to add none/all and speed/all. The latter covers the generic, erased-dispatch and complete hosted workloads that need forced-specialization evidence. Policy travels with the metadata when a fixture is renamed; no part of its directory name selects coverage. Missing, duplicate, unknown or misplaced profile policy is a metadata error. A selected runtime/ABI fixture uses the same matrix. Artifact names and assertion labels include the profile. Each profile independently checks the original exact status, trap and output oracle; agreement with another profile alone is never success. Timeout never satisfies traps: yes.

The harness adds compiler controls directly, not through a fixture's args or run_args; profiles selects the harness matrix separately from the program's original request and oracle. --build-mode is a separate source-configuration axis, and building the Ada compiler in debug/release does not select an emitted-code objective either.

After the compiler build, ./scripts/quality.sh runs the Linux x86-64 numeric acceptance in quality/check.py. It requires LANDIN_GNAT_HOME and takes gcc, objdump and size from that checksum-pinned installation, not an arbitrary host PATH. It compiles and repeats each request, parses factual JSON, assembles the same output to ELF objects, measures sections, function bytes, prologue frames and disassembled instruction sites, and executes that same assembly against exact exit/stdout/stderr oracles. Its probes cover scalar chains and loops, a tiny leaf, compact large-array arithmetic with real stack-page touches, explicit optimal layout, single-instance evidence specialization, a source-level two-instance size/speed threshold and final private-body folding. The specialization and folding probes also run none/all and speed/all. The complete derived-parser client runs all six quality profiles with its original input path, three ordered diagnostics and status 42. Repeated requests must produce byte-identical assembly and build reports, and the measured object must execute that original oracle. Its source inventory reaches the real parser and lexer; object measurements add no size or timing threshold. The complete derived-containers client runs all six profiles, preserving its status-42 and empty-output oracle. Its reports must identify real core container instances and factual specialization actions with retained evidence ABIs; forced specialization must actually select a container entry. Its object measurements are observations, not a new size budget or timing claim. The complete derived-hosted-memory application also runs all six quality profiles with its status-42 and empty-output oracle. Source inventories must reach its real application module; factual specialization reports and emitted indirect machine calls must preserve runtime provider dispatch even with forced specialization. Measurements remain object observations without a new size or timing threshold. The quality and debugger runners give these large rooted workloads a separate 900-second compiler-subprocess limit: its full-debug compilation already exceeds the ordinary 120-second limit on a native development host. Executable and debugger timeouts remain 120 seconds; slow compilation does not excuse a hung program or debugger. Existing insertion-sort and sieve-of-eratosthenes sources retain their original status-42 oracles; the other probes return zero. Scalar acceptance requires substantial stack-site, frame and instruction reductions, with no tiny-leaf growth or gratuitous callee-save overhead; existing workload text has a bounded regression allowance. Source threshold decisions are checked against both measured cost inputs and actual direct/indirect machine sites, including shared fallback bodies. The runner retains disassembly, symbol and size output and compiler/assembly/object hashes in JSON after all acceptance checks pass. The numeric bounds live in quality/check.py itself. The script writes actual observations to the selected build tree's quality/measurements.json; it never updates an acceptance bound or recorded fixture. Acceptance is the current command's zero exit, not the presence of this file: a failed rerun leaves an earlier successful observation untouched. This is structural/object smoke evidence, not timing or competitive benchmark evidence. A non-Linux host fails rather than claiming a skip as a pass.

Source debugger acceptance

After building the compiler, ./scripts/debug.sh runs the Linux source-debugger acceptance in debugging/check.py for this host's own Linux target, x86-64 or arm64; --target=linux-arm64 with --qemu and --driver selects the other under QEMU's GDB stub, as cross evidence. It uses GDB to test the emitted program's line information, breakpoints, stepping, stack frames and selected parameters and locals. Routine debugger risk runs the full release matrix; milestones run it with both debug and release Ada compilers. The script fails if its tools or debugger operations are unavailable; missing debugger evidence is not a pass.

On native macOS arm64, use ./scripts/debug.sh --target=darwin-arm64 --output DIR with a fresh output directory. debugging/darwin.py checks the shared selected Linux source-debugging fixture under none/off, size/auto and size/all, plus all thirteen scalar types in darwin-scalars.ldn. LLDB's SB API asserts source lines, stepping, nested caller values, represented aggregate/variant members, generic instances, source aliases and unavailable locals. It requires the packaged dSYM and a normal status-42 inferior exit; missing assertions or transport is a failure. --profile selects development feedback and cannot satisfy acceptance.

The runner retains source bytes, compiler/tool hashes, commands and exits, assembly, objects, executables, dSYMs, maps, DWARF verification/unwind dumps, UUIDs and LLDB transcripts. It checks mismatch refusal, comment-only identity changes, default/explicit none equivalence and stripped deployment without filenames or source breakpoint locations. Malformed identity controls run without a native debugger in scripts/tests/test_macho_identity.py. When debugging is selected, committed Mac acceptance requires all profiles and verifies their artifact identities; routine uses release and milestones use both compiler modes. --parity also runs the complete P2/P3/P4 LLDB workloads before these selected checks; see the native parity guide.

The complete derived-parser program runs in the same runner using none/off, size/auto and size/all. It receives the fixture's original input path and must produce its exact ordered diagnostics and status 42, including after stripping. GDB inspects initialized parser state, recursive source frames, recovery, nesting depth and the final success/failure flags. Its whole reached source inventory receives the same source-map, line-table and build-identity checks. check.py holds the complete P2/P3/P4 workload schedule to all three profiles.

The complete derived-containers program is another workload in that same runner, using none/off, size/auto and size/all. Source markers in its real containers_run path locate the sorted list and completed composition. Two concrete evidence_less calls use the same high-bit operand as signed and unsigned values, requiring opposite comparison results. GDB must identify each selected provider's frame and source line, verify the two unwind steps back to evidence_less and containers_run, and inspect the caller's completed result binding. The stack includes the application and entry point; acceptance does not depend on GDB automatically printing return values. The runner retains source maps, build reports and transcripts; it checks source hashes, line tables, stripping and exact executable identity for the whole reached library closure, not a fixture-only copy of it. The complete derived-hosted-memory application is another workload, using none/off, size/auto and size/all. GDB stops inside the runtime-selected sample_keep and text_emit providers, identifies their actual source lines, inspects the sampling state before and after its increment and the destination delivery cursor, and requires the caller stack to include process, run_logged, run and the fixture entry point. The text destination also retains its emit_retry caller frame. The whole memory-world application then completes its status-42 oracle. The same full source inventory, assembly hashes, executable identity, line-table, stripping and source-map checks apply to this application and its reached library closure. These sessions exercise ordinary any dispatch in the application; the real hosted I/O fixtures separately assert native process behavior. python3 compiler/tests/debugging/test_check.py exercises transcript refusals without GDB; scripts/debug.sh runs those regressions before the real sessions.

The current Linux transport is native GDB on the native runner. The historical translated troubleshooting path uses --runner=qemu; --qemu=PATH selects the emulator explicitly. It stayed outside the development and acceptance loop of the macOS and Cortex-M work. This uses QEMU's GDB remote stub with the same assertions and reports its transport. It never turns a failed native session into a pass by automatically falling back to emulation.

Native report identity and build inventory

Cortex source debugging is a separate mandatory embedded lane under environments/cortex-m/evidence.py, invoked by run.py on native Linux. It uses --debug=lines, source/function/ordinary-frame assertions and fail-closed artifact selection, with no advertised variable/type interface. Complete-driver device execution, QEMU CPU/startup and independent exception/stack controls stay distinct. See the embedded evidence guide and target contract.

python3 compiler/tests/test_native_report_identity.py --refine ABSOLUTE_PATH checks real destination identities without substituting the fake filesystem. It retains the original collision/refusal and successful-output oracles for source and artifact links, absent leaves, actual symlink parents, case rules, dangling links and indeterminate identities. --directory optionally chooses another filesystem for the temporary cases. The gate runs it with both compiler build modes on Linux and on macOS. python3 scripts/tests/test_build_inventory.py exercises the production developer-build inventory decision, including C/header addition, removal and renaming, without invoking a builder; the gate's scripts job runs it with every other scripts/tests module.

On an unknown filesystem, two absent ASCII names that still differ after ASCII case folding are distinct, because no filesystem equates them; the platform suite holds that rule at the C boundary. Other absent names there remain conservatively indeterminate: a case-folded match, a non-ASCII byte, a ~ or :, or a trailing dot or space. In particular, a Linux container cannot infer the host volume's case rules from its virtiofs mount. For local object-quality measurements, put the runner's --output on the container's own filesystem (for example /tmp), then retain its measurements.json in the host build tree. This changes no source, profile, execution oracle or acceptance threshold; the native gate's ordinary scripts/quality.sh output already resides on its Linux filesystem.

Language-server sessions

Each directory under server/ is one scripted session of refine lsp: a session.lsp transcript and, if the session reads files the editor has not opened, a workspace/ tree, served under file:///workspace. The transcript is lines: -> and one JSON message the editor sends, <- and one message the server must send, byte for byte, pause where the editor waits and the server catches up, chunk: N to feed the input N bytes at a time, raw: and bytes for a frame that is wrong on purpose, # for a comment, and exit: N last, the status the session must end with. Every message the server sends must be in the transcript, in order.

The server suite's every session runs as written case runs each transcript in the test program against the fake channel and filesystem. python3 compiler/tests/server/native_session.py --refine ABSOLUTE_PATH runs the same transcripts through the executable over pipes, with each workspace copied into a temporary directory and its URI written in place of file:///workspace; the gate runs it beside the native report identity, with both build modes on Linux and on macOS. ./scripts/test.sh --record rewrites every <- line and exit status with what the server sends now, as it does lowering.ir; read the difference before keeping it, because a recording is not a verdict.

Complete programs to try

The runtime fixtures include small, complete programs rather than only single-construct probes. Eleven of them are collected in examples.md and use the language and the hosted library:

  • sensors polls two kinds of sensor through any, grows a list of readings in an arena its caller lends, and skips or propagates declared failures;
  • FizzBuzz traverses one through 100, prints the traditional lines and tallies their atom classifications;
  • greatest common divisor implements Euclid's remainder reduction;
  • insertion sort sorts caller-owned storage through inout and a writable slice;
  • binary search searches a read-only slice and returns a found-or-missing variant;
  • the sieve of Eratosthenes marks composites in caller-owned fixed storage;
  • run-length encoding transforms a read-only slice into caller-owned structured output;
  • merge sort divides recursively and loops over caller-owned storage and a local work array;
  • fannkuch-redux enumerates all permutations of seven values and reports the official checksum and maximum flip count;
  • Mandelbrot plots the official 200-by-200 correctness image as a binary portable bitmap;
  • FASTA emits the official 1,000-unit repeated and weighted-random DNA sequences.

The examples use loops for ordinary traversal, reserve recursion for merge sort's divide-and-conquer step, and verify their results through status 42; FizzBuzz and the three Benchmark Game programs additionally have exact output oracles. Together they exercise aggregate parameters, fixed arrays, slices, inout, atoms, variants, pattern matching, computed indexing, valued loop exits, text literals, floating-point arithmetic, concepts, any dispatch, lent allocators, binary output and hosted I/O. The Benchmark Game ports use its published algorithms and small correctness inputs, not its performance inputs; they are correctness and compiler-pressure workloads, not competitive benchmark targets. Their generated oracles were compared byte for byte with the official fannkuch-redux, Mandelbrot, and FASTA outputs.

The narrower runtime fixtures retain the single-construct and composition coverage behind those examples. On Linux x86-64, compile one from the repository root with:

refine --root=. --target=linux-x86-64 --emit=exe \
  -o /tmp/landin-insertion-sort \
  compiler/tests/fixtures/runtime/insertion-sort
/tmp/landin-insertion-sort
test $? -eq 42

Each program returns 42 when its result is the expected one. They are runtime fixtures as well as examples. .github/workflows/gate.yml compiles, runs and checks all eleven on each full gate run. AGENTS.md describes the verified explanatory-only reuse exception.

The test program validates its complete suite-name inventory before any selected or complete run. It rejects missing and unlisted suites separately from individual case selection. check.py also compares every suite source with its registration call and expected name, so dropping both a call and its expected name cannot silently omit a source package.

The metadata-only fixtures/repository fixtures are clean case compares Ada discovery with every fixture identity and target list in targets.matrix, which check.py generates independently. Missing, additional and duplicate rows fail. Corpus runners separately compare their attempted programs, recorded outputs and runtime/ABI profiles with metadata-derived obligations, and require the positive, negative, runtime, ABI and recorded-output coverage categories they own to remain present. Malformed metadata stops those runners before they invoke a compiler. These are inventory and selection checks; individual verdicts and output oracles still decide success. Intentionally removing a fixture and regenerating the inventory still needs review of the semantic coverage that fixture provided.

Live Landin fences in the tour and prototypes also pass the independent lexical scanner, including the lowercase identifier rule. Historical findings and the tour's dropped-design section retain their original bytes and are excluded. This does not make the prototypes standalone programs: their documented omissions and future constructs remain. The tour's initialized-object and three-buffer acquisition examples are copied into positive/r491-tour-allocator-examples; check.py holds those two excerpts to the compiled witness. Its count is two bytes per buffer and validation only emits assembly text.

Metadata

fixture.meta is key: value lines, with # comments and blank lines.

keyrequiredmeaning
classyesmust match the directory the fixture sits in
summaryyesone line, what the fixture proves
profilesyes for runtime and ABI onlystandard for the four baseline compiler profiles, or specialization for all six; independent of fixture name
levelsno, and runtime onlycomma-separated CPU feature levels (D255), each some product target's, at which the program is also built and executed under every profile; each lane runs its own family's (Linux x86-64 the x86-64 ones, Linux arm64 and Darwin the arm64 ones, the Cortex-M corpus the M-profile ones), and the default level always runs. A level runs only once the processor that runs the lane confirms every feature of it: a Linux lane reads /proc/cpuinfo, a native one its host's and a cross one as its runner presents it, through a C program the lane's driver links; Darwin asks sysctl. An unconfirmed level fails as UNVERIFIED rather than being skipped or inferred from another lane. On Linux arm64, which has no ISA note, the image built at a level is held to its LSE instructions and to no exclusive-monitor loop, and the default's image of the same fixture to no LSE instruction
programyes for runtime, ABI, and a rooted positive or negativethe .ldn program the fixture runs or uses as its compile-only corpus file
withnothe rest of the Landin module, when one file is not enough; never a C source
rootnoa positive, negative, runtime or ABI fixture's import root, relative to its directory; the directory itself becomes the entry module
c-sourcesyes for ABIcomma-separated, ordered C companion sources relative to the fixture directory
c-argsnowhitespace-separated C compiler and linker arguments for an ABI fixture
expectnothe file holding the expected bytes
argsnothe arguments refine is run with
run_argsnothe arguments handed to a compiled runtime or ABI program
run_expectnothe file holding a runtime or ABI program's expected bytes on the selected stream
statusnothe exit status refine or a compiled program must produce (default 1 for a negative, 0 otherwise)
trapsnoyes if a runtime or ABI program must end without returning a status
streamnooutput (the bytes must be on standard output, and standard error must be empty) or merged (default)
lexnothe exact complaint the scanner must produce, for a fixture whose fault is lexical
codesyes for a negative with a programthe diagnostic codes the report must carry, in order, errors and warnings alike; on a positive, runtime or ABI fixture, the warnings its accepted program reports
constructsyes for a fixture with a programthe [NNNN] ids, without brackets, this fixture is evidence about
targetsyescomma-separated targets the fixture applies to
fixednofor a negative with a program and codes: golden results after the first fix of every diagnostic that offers one is applied; a bare name is the program's result and source -> result another source's

codes also says which stage refused the fixture, and that is what decides whether the grammar must derive its program. The frontend refuses what the grammar cannot derive, so a fixture whose first code the scan or the parse can raise must not derive; a later stage refuses source that parsed, so a fixture whose first code belongs to one of those must derive exactly as a positive fixture does. One parser rule is not the grammar's: end may repeat only the name it closes, which a context-free "end" identifier? cannot say, so a fixture refused for nothing but L0109 is derivable and must derive. check.py reads which codes the frontend raises out of Landin.Diagnostics.Lexical and Landin.Diagnostics.Syntactic rather than out of the number, because the catalogue's own header forbids reading a stage off a code — L0010 began in lexical refusal and is now raised only by the parser.

A warning is pinned exactly as an error is. An accepted program's report is held to its codes: a positive fixture's emission, and a runtime or ABI fixture's compilation, must report exactly those codes in that order, and without the key must report nothing, so a warning nobody pinned fails there. A negative fixture whose report is only warnings says status: 0, since a warning never refuses a program; negative/mut-never-written and negative/unused-pure-local both exercise this case.

codes is an ordered list and not a set. Two refused constructs in one file are two reports, and a regression that doubles a count is invisible to a set, so a fixture that contains two refused uses names its code twice in source order. Spaces around comma boundaries are insignificant; the harness canonicalizes them without sorting the codes or removing duplicates. Both source-only and recorded negative executions compare that ordered list and the declared or default exit status. Recorded negatives also compare their exact report bytes, including those with args but no program. float-literal-not-enabled names one L0301: its one literal is a float in an integer context, refused by the checker, so the grammar must derive it. check.py holds every name in codes to the catalogue, and refuses a negative fixture with a program that names none; the parser suite scans and parses the program and holds the report to the exact sequence.

with is how a fixture is more than one file. [1840] says the module scope is "every file compiled together", so a claim about it cannot be made by a fixture that can only name one; program stays the file the fixture is named for and with is handed to refine after it, in the order written. Naming the rest of a module with no program to be the rest of is a reported fault, and check.py holds every file either key names to being there — a name pointing at nothing would compile one file while claiming to have compiled two. Every .ldn in a fixture directory is held to the grammar already, so the extra files are derived like any other. C companions never belong in with; an ABI fixture names them only with c-sources.

fixed records the golden bytes for the first-fix pass. The fixes suite compiles the fixture as the parser suite does, requires its report to carry the ordered codes and at least one diagnostic to offer a fix, then applies the first fix from every diagnostic that offers one. Every edited source must have a recorded result, and its bytes must match; the resulting module must compile with an empty report.

The suite separately tries every later fix. For each alternative it keeps the first fixes of the other diagnostics, requires every edit of the selected fix to name a loaded source, applies its edits to all sources they name, and requires the resulting module to compile with an empty report. An alternative has no separate fixed golden. Multi-source edits are checked across the whole loaded module: no selected edit may be dropped, edits may not clash, and all affected sources are compiled together. A one-file or with fixture runs on a fake filesystem holding its sources; a rooted one reads its module closure from the repository, as its negative case does, with only the edited sources replaced. check.py holds every file fixed names to being there and every result to the grammar, however the original was refused, and the diagnostic matrix holds every code whose catalogue row admits a fix to at least one fixture that applies one. A result is named *.fixed rather than *.ldn, because an entry module is every .ldn in its directory and a rooted fixture's result would otherwise be compiled beside the program it replaces.

root is the directory-module counterpart for a positive, negative, runtime or ABI program. It is relative to the fixture directory; when present, the fixture directory is passed to refine as the entry module and the root is passed first with --root. It cannot be combined with with, because rooted discovery owns the source membership. A rooted positive or negative fixture must also name a nonempty program: that is the corpus file whose compile-only acceptance or refusal is counted. This is how compile-only fixtures and executable fixtures alike import repository-owned modules such as core/* without keeping fixture-only copies. The parser's exact-code case retains its fake filesystem for ordinary negative fixtures, but reads the real module closure for a rooted negative because its pinned report depends on those imports resolving.

An ABI fixture requires program, c-sources, constructs, and targets. c-sources is a comma-separated ordered list. Every entry must be a portable, slash-separated relative path ending in .c, must contain no backslash or colon, must remain below the fixture directory (no absolute, empty, . or .. component), and must name a file that exists. The list is passed to the target driver in the order written. c-args, when present, is a whitespace-separated argument-vector suffix; no shell interprets it. Both keys belong only to ABI fixtures, duplicate keys or C source paths are faults, and C source files in with are faults.

The ABI harness runs refine for Linux x86-64 with the Landin inputs followed by --target=linux-x86-64 --emit=asm -o <assembly>. It then selects Driver_For (Linux_X86_64, "") and invokes that driver once with arguments in this exact order:

<assembly> <c-sources in metadata order>
-std=c11 -Wall -Wextra -Werror -no-pie
<c-args in metadata order> -o <executable>

The resulting executable may get main from Landin or from a C companion. The harness passes run_args, compares run_expect with the selected stream when it is present, and then compares either the exit status or the non-returning traps: yes verdict. Thus a C main can drive exported Landin routines, and a Landin main can drive imported C routines, without adding C inputs to the product compiler. Runtime and ABI fixtures honor stream: output by capturing stdout and stderr separately: stdout must match run_expect when present, and stderr must be empty even without an output file. The default merged policy preserves combined bytes. Recorded compiler expectations use the same stream contract.

constructs is what the construct matrix indexes the corpus by, and it is a written list rather than a reading of the summary. A citation in prose is prose: it is there to explain the fixture to a person, it may name a paragraph the fixture merely mentions, and a heuristic over English is how a check ends up believing 114 lines of it were code. check.py holds every id to a paragraph tour.md or spec.md actually defines, and the harness holds it to being four digits — the two halves of the question, asked by the side that can answer each.

The decision register's Pinned by paragraphs are different: those paths are the current evidence they promise a reader, so check.py holds every named fixture directory to existing. Historical prose may still name a retired fixture when the retirement is the point; a live evidence list may not.

Name a construct when the fixture's *passing* would change if that construct were implemented wrong, and not when the construct merely appears in the text. Every runtime program contains literals, so naming [1770] everywhere would make that row read "covered" while saying nothing; a fixture whose asserted values come from literals earns it. The failure this rule prevents is the one a matrix is most prone to: a full column that means nothing. When a claim turns out not to be earned, the honest repairs are to drop it or to make it true — runtime/statements-run-as-they-read claimed [1840] before it declared anything inside an arm, and grew a function that does. The construct audit dropped [1550] from four fixtures whose passing another backend would not change, and [1730] from two that could not observe an elided check, and added [1570] to the firmware driver, whose interrupt handlers execute on Cortex-M. A claim that would fit every fixture discriminates none.

A fixture with a program must name at least one, because a .ldn program is written in the language and is therefore evidence about some construct of it. A fixture without one is about the tool rather than the language — an unknown option, the identity text, an implementation-side note — and names none for the same reason. A construct the kernel does not enable yet is perfectly good: negative/convention-not-enabled names [1830] for the refusal and [0900] for the thing being refused, and [0900] is a paragraph about a construct no fixture can yet use.

targets is required, and check.py holds each name to the targets the target applicability register allows: the seven product targets linux-x86-64, linux-arm64, macos-arm64, freebsd-x86-64, freebsd-arm64, linux-rv64 and cortex-m, and synthetic-32, the 32-bit model that preceded the Cortex-M backend. A name outside that list is refused, because that is how a fixture quietly stops applying to anything.

linux-arm64 is the Linux arm64 lane's own name and its driver name alike, so a fixture whose args: select --target=linux-arm64 names it, and one that names it is run by that lane with the target its args choose, or the lane's own when they choose none. A runtime or ABI fixture the Linux x86-64 lane runs names linux-arm64 too, or has a record in linux-arm64/parity.json giving the reason it cannot and the counterpart that runs on arm64 instead; a record for a fixture the lane runs, or a counterpart it does not run, is refused. The hand-ordered file is edited by hand, as darwin/parity.json and cortex-m/corpus.json are.

expect and args come as a pair. An expectation with no way to produce it is dead data that looks like coverage, and arguments with nothing to compare against are a command nobody checks, so either one alone is a reported fault.

A fixture carrying both is executed: the harness runs refine with those arguments through the real tool adapter, and compares the captured bytes and the exit status with what the fixture claims. Standard output and standard error are captured together, in the order the process wrote them.

Discovery is strict. An unknown key, a repeated key, a missing required key, a class that disagrees with its directory, a line that is not a pair, a fixture directory without metadata, and a plain file where a fixture belongs are all reported, and the fixture is not accepted. A fixture that is half-accepted is a fixture whose fault stops being visible.

Names beginning with . are skipped: host clutter is not a fixture and not a fault.

Ordering is by class, then by name, so a run reports the same sequence everywhere.

What the harness does with them

classtoday
unita note of what an implementation-side case covers; the case itself lives in compiler/ada/tests
negative, end-to-endexecuted: recorded cases compare refine output/status with expect/status; program-only negative cases require the expected failure status and ordered codes: sequence
runtimeexecuted: refine compiles and links program, the result is run, and its own exit status is compared with status — or, with traps: yes, it is held to having ended without returning one
ABIexecuted in the existing runtime fixtures execute case: refine emits assembly, the selected Linux x86-64 C driver compiles it with c-sources, and the result's output/status/trap verdict is checked
positiveexecuted: the grammar must derive the program, refine must accept it through checking, lowering and verification, and the Linux x86-64 backend must emit assembly for it
debuggersessions in debugging/ run separately through scripts/debug.sh

Each producing attempt first removes its expected output through the platform filesystem interface. An output that cannot be removed stops that attempt; a zero exit status then requires a newly created file. This applies to positive assembly, runtime executables and both ABI production steps. Fake-command cases pin missing, stale, repeated, failed, timed-out and unavailable producers without invoking a compiler or toolchain. A positive fixture may legitimately contain only types or external declarations and emit no function body: emission is stage coverage, while instruction semantics need IR or runtime oracles.

A class with no fixtures is the normal state early in the roadmap, and an empty class directory is not a fault. A fixture that records an expectation nobody runs is.

That last sentence decides what a runtime or ABI fixture does on a host that cannot finish the target, and the answer is that the run fails. A macOS host with no ELF toolchain reports the existing runtime fixtures execute case as one failing case: runtime fixtures carry refine's own L0500 report, and ABI fixtures name the triplet-selected C driver that could not be run. Skipping would be the quiet non-run the sentence refuses, and it would also hide the gate losing its toolchain. This is the same rule scripts/env.sh already applies one level up: a machine without the pinned GNAT is told so and stops, rather than quietly building nothing.

A runtime or ABI fixture carries program and either an exit status (zero by default) or traps: yes, and neither expect nor args, because nothing compares refine's own output — what is asserted is what the compiled program did. One without a program is a reported fault, for the same reason expect without args is: a status nobody produces is dead data. ABI additionally requires c-sources; c-args, run_args, and run_expect remain optional.

Accepted, emitted and executed are three claims and not one, which is why three classes make them. A positive fixture is a program the compiler must accept, and asking only that was how four of [1810]'s statement forms reached the first Linux backend's audit without ever having been handed to it: every stage accepted them and no case asked for a byte of assembly. So the positive class now emits as well, and a construct that reaches a compiler defect on the way to .s fails there rather than waiting for a runtime fixture to happen to use it. It is still not executed — most of the corpus is a fragment with no entry point to run, and a claim about a machine belongs to the runtime class.

traps: yes replaces status rather than joining it. spec.md [1960] says a trap is synchronous and non-returning and that its operating-system signal or status is not stable program behaviour, so a fixture may assert that the program ended without returning a status and may not assert which signal ended it. Naming both is a reported fault: a program that trapped has no status, and a fixture claiming one is claiming an answer nobody can observe. Nothing in the format or in Landin.Platform carries a signal number, deliberately.

What that can and cannot tell apart is worth knowing before writing one. runtime/checked-overflow-traps adds one to a 255u8 the compiler cannot read: without the backend's own check the instruction keeps the low byte and the program returns 42, so the fixture fails when the trap edge is removed — which is measured, not assumed. runtime/a-zero-divisor-traps cannot make that distinction, because x86-64 faults on a zero divisor whether or not the compiler guarded it; it proves [1950]'s obligation is met and not which of the two stopped the program. D11 is where the choice to emit a deliberate ud2 rather than inherit the incidental fault is recorded, and deterministic assembly is what pins it.

The defer and undo no-unwind fixtures register exit(99) as cleanup. Reaching that observer, whether before or after the fault, produces an ordinary exit and fails traps: yes. Merely changing the divisor in cleanup would not expose cleanup reached after the fault. A fake-outcome control pins the distinction; the fixture still cannot identify which instruction generated a signal.

The grammar corpus

A .ldn file under positive/ must be derivable from the enabled grammar in spec.md. A negative program a later stage refuses must derive too; one whose first diagnostic comes from the scanner or parser must not. check.py enforces those stage-sensitive verdicts on every full run, and it enforces that every construct in the grammar section is named by at least one fixture, so a production nothing pins is a reported fault rather than a quiet one.

The corpus made the specification and its examples check each other before a compiler existed. The parser now has to agree with the same corpus, and a disagreement between the parser and the grammar is a defect in one of them rather than a matter of opinion.

Two rounds of reading the grammar by hand found sixty-eight defects between them and still missed that a lone _ parsed as a name. The corpus found that in a second.

A negative fixture may add lex: <complaint> to pin the independent Python grammar scanner's refusal wording. check.py compares that field; the Ada fixture reader accepts the metadata but does not use it as an Ada diagnostic oracle. Ada diagnostics are pinned by the ordered codes: list and, when present, the exact expect report. Lexer unit cases separately inspect tokens, spans and complaints. Neither a Python lex: match nor codes alone proves the Ada report's wording or span is correct.

Programs are read as bytes. Text mode would turn CR LF and a lone CR into LF, so the terminator rule [1750] states could not be tested however many fixtures were written for it; positive/line-ends-crlf carries a CR byte and the checker asserts it is still there when read.

What the corpus cannot see, recorded so nobody assumes otherwise: it cannot tell CR LF read as one terminator from CR and LF read as two, because both produce the same tokens. The distinction belongs to the line map, and Landin.Source's own case for it is what holds that.

lexical.tokens

compiler/tests/lexical.tokens is generated: python3 check.py --tokens writes it from check.py's tokeniser, one line per token as first last spelling. The Ada harness reads it and compares every token with what Landin.Tokens.Lexer produced, and check.py regenerates it on every full run and fails if the committed copy is stale.

Kinds are deliberately not in it. The two implementations have different kind vocabularies, and what a disagreement actually looks like is a boundary in a different place.

That is two independent implementations of one grammar, held to each other over every program in the corpus. The first thing it caught was real: the Ada scanner appended each file's tokens to the previous file's, because a limited out parameter is passed by reference and Lex had not cleared it.

Diagnostic and semantic coverage registers

python3 check.py --catalogue writes both diagnostics.catalogue and diagnostics.matrix. The compact catalogue comes from Landin.Diagnostics.Catalogue, which is the only place in the compiler where a code is written. The matrix crosses every row's source/span/label/note contract with its emitter and fixture or fake-host test owner. A live code with no emitter or owner, a retired code still emitted, and a source diagnostic with no negative-program owner are gate failures; L0111's parser nesting limit and L0325's routine and struct size limits have their implementation-side unit owners instead, because a program that meets either is too large to keep in the corpus.

python3 check.py --coverage writes guarantees.matrix, conformances.matrix, prototypes.matrix and targets.matrix. Their source registers are D148 in spec.md and the prototype registers in registers.md. The checker closes the guarantee rows over every construct the independent construct matrix says is accepted or emitted, validates every fixture, diagnostic, decision and prototype finding they cite, requires every fixture to name applicable targets, and recovers prototype finding line numbers from the prototype sources. The copies are generated for reading; editing one cannot change its source.

prototypes.matrix carries three derived columns: inputs, outputs and results, never asserted beside the row. inputs and outputs come from the fixture's own record — its program, root, arguments, C peer, status, ordered codes and a digest of every golden it cites — so an edited expected output or a changed code list moves the column and a stale matrix fails the gate. results is the verdict each product target's retained record reaches for that one derivation, read through the same per-fixture reader the construct inventory folds over constructs: executed, compiled or refused, strongest first. targets stays the applicability claim it always was, and the two are deliberately separate — the Cortex corpus records verdicts for fixtures whose metadata names only the hosted targets, and a recorded Cortex refusal of a hosted derivative is a result rather than a silence. synthetic-32 never appears under results: it is the model that preceded the Cortex-M backend and applies to no construct, so a verdict under it would be a verdict about nothing. compiler/tests/driver/fixture.json carries the firmware driver's oracle, the assertion groups its runner defines; scripts/tests/test_prototype_coverage.py proves each of those refusals fires, including a renamed oracle, an absent golden and a target a row claims that no record places.

Every full check.py run fails if any generated copy is stale, and it refuses a code literal written anywhere else under compiler/ada/src.

The catalogue check earned itself immediately: the driver had held L0001 to L0004 as literals since it was first written, and moving them into the catalogue was the first thing it demanded.

lowering.ir and layout.targets

compiler/tests/lowering.ir is generated: ./scripts/test.sh --record writes it by lowering every positive fixture and rendering the Unit with Landin.IR.Dump. compiler/tests/layout.targets is written by the same command, and records what Landin.Targets says scalar and aggregate shapes measure, align to and offset their fields by on each described target. Its D74 rows also work the tag-first variant part and one containing aggregate on each description. Both are recorded artefacts check.py does not touch, and that difference matters enough to state. check.py generates the other two because it owns their sources — its own tokeniser, and the catalogue's Ada text. It owns nothing here: producing these files means running compiler stages and asking the target model, so the Ada harness is what can produce them and python3 check.py will not tell you either is stale. ./scripts/test.sh will, and so will the gate.

layout.targets exists for an ordering reason: target-parametric layout came before any 32-bit backend, a description is the only thing a compiler with no such machine can be held to, and the synthetic 32-bit target has no backend. The Cortex-M0 target separately instantiates its layout and preserves these original goldens. Recording both targets rather than that one is deliberate — what a reader needs is not "the 32-bit model says four" but the two columns beside each other, because the defect being guarded against is a description quietly inheriting the development host's answers. A usize that read eight in both would be exactly that, and it is a one-line change away at any time.

Recording runs no case, and no case ever writes. Two disjoint modes in one binary, chosen only by an argument a human typed: there is no environment variable, nothing writes the file when it is missing, and nothing rewrites it when it does not match. A golden that repairs itself on a mismatch records the defect instead of reporting it. The loop is closed by hand — record, then run the suite again with no argument — or by ./scripts/test.sh --record-and-run, whose shell wrapper makes both explicit binary invocations after one build.

What the file is for is narrow, and landin-ir-dump.ads says it: it proves the lowering has not changed its mind. That the corpus derives from the grammar is check.py's, and what the instructions mean is Landin.IR.Verifier's. No origin is printed, deliberately, so a comment edit above an instruction does not rewrite the artefact.

Derived programs

The four prototype text files in the repository root remain design-stress sketches, including their omissions and historical findings. Complete derived .ldn programs are separate artefacts and arrive with the roadmap work that can compile them.

runtime/derived-parser hosts the complete prototype-2-derived lexer and recovering configuration parser in examples/config_parser. Its DERIVATION.md maps the executable behavior and negative controls to the prototype.

runtime/derived-containers hosts examples/derived_containers/workload as the complete prototype-3-derived workload. Its DERIVATION.md maps every prototype section and Z finding to the ordinary core modules, executable paths and negative corpus. The runtime entry requires every path to succeed before returning 42, with no stdout or stderr: list growth and sorting, direct fixed array mutation, nested-array field ranges, initialized raw storage, vector and small-vector failure/retry, map collisions/reuse/enumeration, each of the three map acquisition failures, tree traversal and heterogeneous dispatch. Providers remain explicit capabilities and cleanup is observable; no fixture-private replacement container library or implicit resource ownership stands in for the prototype. This same workload is mandatory in the six-profile runtime matrix, object-quality measurements and source-debugger acceptance above.

runtime/derived-hosted-memory hosts the complete prototype-4-derived application in examples/derived_hosted/app. The runnable hosted entry and its derivation map live in examples/derived_hosted. Runtime configuration constructs heterogeneous filters and either a counting or text destination; the processing loop calls their ordinary any evidence entries. The reader retains partial lines across chunks and emits a final unterminated line. Copied arguments and message storage survive helper returns, and delivery retains its committed byte cursor across an explicit retry. The memory world makes input fragmentation, output contents and failure cleanup deterministic; the hosted fixtures use the same application with actual native Linux I/O. The memory composition is also mandatory in all six quality profiles and the three debugger workload profiles described above.

runtime/r480-generic-provider-entry pins the bound entry required when a concept provider itself has constrained generic parameters. Two nested providers execute through direct evidence and any calls, forwarding a by-value aggregate, an aggregate result, an inout array and a declared failure. The table entry supplies the provider's concrete evidence arguments to the ordinary generic body; its source calling convention and the two-word any representation remain unchanged.

Hosted parity asks for more than a populated construct column. Every hosted row must cite a Linux runtime or ABI program, except the explicitly registered compile-time rules whose exact acceptance/refusal is the oracle. check.py enforces this boundary; it does not infer semantic adequacy from metadata. The audit adds distinct and inline nominal types, exact once-evaluated field fills, atom comparisons across structural sets, atom-bearing arrays/fields/payloads, parameterized union aliases, conformance-key controls and recovered-error generic deduction. Review probes include static distinct Boolean/pointer images, type-name misuse, exact union application diagnostics and unused symbolic pointer obligations. Original refusal sources promoted to enabled grammar remain byte-for-byte positive fixtures, and all unchanged derived-application recovery and provider oracles remain required. D212's ordinary allocator authority and all existing workload profiles remain required.

Native Darwin lowering

compiler/tests/darwin/cases.json selects the first Darwin lowering and ABI corpus, run by compiler/tests/darwin/check.py with none/off, size/auto and speed/all profiles. It reuses shared source/verdicts and has separate Apple varargs, packed-stack/HFA/indirect-result and errno probes. Each run retains command arguments, outputs, hashes, assembly and native artifacts. Every emitted routine's frame record and reserved-register exclusion are checked; a C peer also checks a live Landin parent frame. --case/--profile are visibly filtered development feedback. The binding runner regenerates and executes the complete adapter-category corpus with the pinned Apple target and tests exact archives.

These scripts require native arm64 macOS. check.py --parity selects every shared runtime/ABI fixture and its four/six profiles; diagnostics.py checks every applicable Darwin source verdict. Explicit native differences and replacements are recorded in darwin/parity.json, with strict executable evidence. Linux golden records remain unchanged. The complete derived parser now also uses all six runtime profiles. The parity guide describes full coverage, debugger oracles and retained platform limitations. The retired native acceptance and matching-revision approval are documented in the Mac guide.

The cortex ABI compiler-host suite compares scalar, nested/variant, evidence/any and call-plan results with cortex-m.contract; it also checks real lowered source and target/budget/refusal boundaries. The independent Cortex-M probes compare every original synthetic-32 golden with GCC measurements and run C/assembly ABI controls in QEMU on the supported native Linux host. The retired native acceptance's documents job retained their evidence. These are executable ABI contract controls. Generated M0 execution is separately identified; these earlier controls retain their independent role.

r630-memory-scalars runs all supported widths, old-value and compare-exchange outcomes, wrapping addition, volatile access and barriers. abi/r630-native-memory runs Landin-generated operations under independent pthread scheduling and compares increments with a C atomic control; publication and mixed-width aliases have exact value assertions. Alignment fixtures require traps even inside unchecked regions. Negative fixtures pin ordered diagnostic codes for illegal types, permissions, orderings, arity and target operations. The ir opt/memory events case checks event preservation under every policy and malformed memory metadata refusals. Embedded controls and abstract cache models are separate evidence described in the Cortex-M guide.

The memory-model ABI counter passes through a concept provider and inferred generic call, so specialization profiles exercise real evidence dispatch. Its bounded SC-fence store-buffer litmus rejects 0/0, but the absence of that observation is not proof of a memory model. Every new negative pins exact diagnostic text; M0 RMW and eight-byte load requests retain their explicit target selection.

The development case targets/packed image algebra and access plans checks the compiler library with an independent bit-by-bit reference over 257024 combinations, enum holes, 64-bit boundaries, indexed fields and the complete bounded access-mode table. It also compares four target descriptions. That case is compiler-host evidence. Source fixtures named r640-* separately cover packed geometry, enum maps, raw copies, validated extraction, fit/index traps, exact diagnostics, register access modes, reserved policies, and generic and erased evidence paths across the six specialization profiles. Independent C ABI controls enumerate all byte images and observe backing storage after pre-store traps. The ordinary-slice DMA control retains D227's barrier contract. The native debugger source checks the raw carrier and its size; named packed bitfield/array display is not claimed.

packed.py executes independent M0 C firmware. packed_native.py executes compiler-generated hosted instructions against the encoding model through a C line transport, with literal independent image/trace assertions. Both are mandatory in the embedded probe/export path, alongside the profile, layout/ABI and memory controls. The Cortex-M guide records their results, assumptions and limits.

cortex-m/corpus.json inventories every shared runtime/ABI fixture. The mandatory embedded runner executes every selected optimization/specialization profile, checks precise source refusals, and retains physical limits separately. cortex-m/counterparts.json pins exact 32-bit/architecture differences without editing the hosted originals. Independent numeric layout expectations stay literal, including pointer/slice carriers and nested variant placement. The embedded guide describes generated execution, independent ABI/instruction controls, peripheral traces, helper provenance and the external startup/linker boundary. GDB observes machine instructions and frame records; this is not Landin source debugging.

D229's positive/r660-machine-directives pins target-fixed parsing of typed machine functions and placement without enabling hosted machine semantics. negative/r660-materialization pins the L0505 pre-GC budget. The cortex ABI suite exercises firmware request failures, conventions/conversions, naked body restrictions, placement conflicts and unsupported hosted assembly. Actual compiler-generated firmware and independent machine controls live in the embedded execution lane; those QEMU results are not hosted fixture passes or source-debugging acceptance. Existing runtime/ABI fixtures and their Cortex dispositions remain unchanged.

D230's positive/r670-scalar-assembly pins the target-fixed syntax without enabling hosted assembly. cortex ABI/assembly IR lowers the shorthand and the named form to one instruction and independently corrupts an input, an output slot, a register and a width, including a register Cortex-M0's table never answers for. The machine-directive cases check source operands, registers, arity, naked restrictions and hosted refusals. The freestanding library lane executes CPU, allocation, pool, vector and ordinary-slice DMA consumers through compiler-owned firmware. It retains the original shared fixture oracles and the reviewed raw-storage 32-bit counterpart. This is additional target evidence; it neither removes a shared verdict nor claims a complete freestanding core.

D232 adds abi/r670-panic on both native hosts, with independently pinned kinds and source-byte sites, no later actions/cleanup, recursion and hosted-root controls across specialization profiles. negative/r670-panic-handler pins L0506 and the driver suite checks malformed hooks on all targets. Cortex uses core-panic.ldn and its default-handler derivative in the mandatory freestanding lane, including real interrupt delivery and separately labelled debugger fault injection. Optional map mismatch tests run under full check.py. These remain additional evidence and preserve the inherited fixture verdicts and lane boundaries.

Generated-device source fixtures live under devices/, outside the shared runtime corpus. devices/test.py verifies offline provenance and fresh regeneration; devices/check_sources.py records seven exact diagnostic-code refusals. The mandatory Cortex device lane compiles all six generated modules and executes five consumers at six profiles, with its own retained independent oracles. No shared fixture is removed or reclassified by this addition.

The complete Cortex driver lives in driver/, outside the hosted fixture runner. Its fixture.json supplies the explicitly named firmware/derived-driver prototype-matrix row; check.py checks the source/ mapping paths and mandatory runner connection. Native Linux documents/tooling acceptance runs its QEMU evidence through compiler-owned firmware. The shared runtime/r690-recovery-loop-context fixture separately checks inferred recovery break/continue and cleanup on both hosted targets and Cortex.

The FreeBSD lanes execute their metadata-selected runtime and C ABI fixtures in separate x86-64 and arm64 VMs, after Linux emission. Their runtime lanes also check every applicable positive and negative source verdict with the FreeBSD target selected. Their LLDB sessions run natively in the guests; a local Mac-hosted VM is developer feedback, and the recurring Linux-runner VM jobs supply the gate evidence.

The physical RV64 lanes cross-emit checked payloads on Linux and report native runtime, GDB, independent nonempty C ABI and ISA-extension verdicts separately. The payload is tied to the source revision, complete fixture/profile selection and source and binary checksums. ISA execution confirms the actual instruction on hardware before running extended code; source checks and emission alone do not establish any execution verdict.