Development and validation environments
Canonical hosting is GitHub; git.sr.ht is a mirror. The exact-revision native acceptance described below approved every revision through 0.2.0 and no longer runs. .github/workflows/gate.yml is what runs now, on every target unless a push meets the verified explanatory-only reuse policy below. Historical SourceHut gate results below keep their original meaning. The acceptance operations described here are historical; scripts/ci/ has been removed.
Native macOS arm64 remains a development and gate environment. Use ./scripts/dev-test.sh --host for compiler-host feedback; the gate runs darwin-host, darwin-parity and lldb on macos-26. The former Darwin acceptance used python3 scripts/ci/darwin.py accept COMMIT on the Mac. Its verified bundle had to match the Linux bundle with compatible committed scope at approval (--darwin DARWIN_BUNDLE), binding both source and execution identities. See the native Mac guide for current development commands and the retired acceptance record.
The verified explanatory-only remote reuse policy in AGENTS.md permits qualifying edits to reuse a full successful main gate from the preceding seven days. Document and script checks always run.
Embedded firmware and freestanding library consumers run on Linux x86-64 through environments/cortex-m/run.py, in the gate's cortex-m job on each full gate run. They keep separate QEMU CPU and synthetic device evidence. Mac --host checks compiler behavior for Cortex; they do not execute embedded workloads.
FreeBSD x86-64 and arm64 run in system VMs on Linux. Each has separate recurring runtime, native guest LLDB and C ABI verdicts. Linux x86-64 emits and inspects hashed payloads against a checksum-locked FreeBSD 14.4 sysroot. The self-hosted freebsd-amd64 and freebsd-arm64 Linux controllers execute them in their FreeBSD 15.1 KVM guests; the evidence records the actual guest version and CPU probes.
Environments
| environment | role | status |
|---|---|---|
| native macOS arm64 | compiler-host development; darwin-host, darwin-parity and lldb gate lanes | working |
| exact-revision Darwin acceptance | verified Mac evidence matched Linux before approval | retired |
Apple Container, linux/amd64 under Rosetta | retained environment troubleshooting, outside the development workflow | available |
Apple Container, linux/arm64 on Apple silicon | Linux arm64 work before a push: the pinned aarch64 toolchain runs natively, with no emulation, in a Debian container | working |
| QEMU user emulation on Linux x86-64 | a cross lane for another Linux architecture, with a cross driver: evidence about the emitted code, none about the pinned toolchain | working |
| FreeBSD amd64 and arm64 system VMs on Linux | independent runtime, native guest LLDB and C ABI gate lanes with checked CPU levels | working |
| native Linux x86-64 runner | explicit exact-revision acceptance | retired with the SourceHut gate |
| builds.sr.ht | retired: the repository submits no build manifest | retired |
| GitHub Actions | gate.yml runs every target unless the verified explanatory-only reuse policy above applies: both compiler modes on Linux x86-64 and on Linux arm64 with GDB, the host suite, parity and LLDB on macOS arm64, and every Cortex-M lane and the called determinism.yml comparison; links.yml, release.yml and pages.yml run separately | working |
The retired acceptance controller ran the committed scripts/ci/policy.json scope against one committed archive. Routine promotion ran debug compiler-host checks, the complete release suite and native identity, release quality, bindings, and document/tooling checks. GDB ran only for substantial debugging regression risk or a major milestone. Milestones restored the full matrix in both compiler modes. Actual versions, binary hashes, commands, logs and artifacts were retained; verified export, annotated approval and atomic main promotion were required. See the retired native acceptance operations and validation workflow for that historical process.
No build manifest is submitted. GitHub is canonical and git.sr.ht is a mirror, kept in step by a second push URL on the same remote; the Pages and GitHub-mirror manifests are retired with the SourceHut gate, and scripts/site.sh renders and packages without publishing. .github/workflows/pages.yml is the only publisher. The former automatic Nix manifest was retired earlier, in favor of explicit supplemental native Nix checks when the shell's inputs change.
Native macOS arm64 began as a *development* loop and became a validated target of its own, with its own compiler build, platform tools and debugger gate; a result produced here is not Linux evidence, and a Linux container is never Darwin evidence.
./scripts/macos.sh --output .scratch/macos-validation-1 captures the native environment, assemble/link/run an arm64 smoke program, exercise LLDB, build and run compiler-host checks in both modes, and sample the inherited resource limits. The Apple tool policy, reproduction commands and the historical full-harness option are documented in environments/macos-arm64/README.md. This established the compiler host; Darwin lowering and source debugging followed. The retired acceptance selected compatible native policies with full release hosted coverage for routine changes and full release GDB/LLDB for debugger risk. Historical schema 3 remains both-mode milestone evidence.
QEMU full-system x86 is supplemental. It is not the daily loop and it is not the gate.
Commands
The same commands run in every environment:
export LANDIN_GNAT_HOME=... # the pinned GNAT for this host export LANDIN_GPRBUILD_HOME=... # the pinned GPRbuild for this host ./scripts/toolchain.sh ./scripts/clean.sh ./scripts/build.sh ./scripts/test.sh
On a Mac use ./scripts/dev-test.sh --host --target=linux-x86-64, optionally with an exact --suite or --case. Every selected check must pass. The first native Mac run's unfiltered missing-Linux-driver result is a historical environment observation, not a current success rule.
Linux work during the macOS and Cortex-M work ran on the native Linux runner, using focused development slots or committed acceptance. The container commands below are retained for explicit environment troubleshooting; they are not part of the Mac development or delivery loop:
./scripts/linux-loop.sh # build and run the suite in linux/amd64 ./scripts/linux-loop.sh sh -c '...' # anything else, in the same environment
Artifact-writing checks need a guest filesystem with known name rules. The shared virtiofs mount cannot prove the host volume's rules for absent output names, so the collision guard can refuse distinct-looking output paths there. For runtime and quality checks, copy the sources and place the executable and its output directory on the container's own /tmp filesystem. The baseline code generation work's local evidence met the same requirement. Building the bootstrap and comparing its recorded IR can still use the shared mount.
For the edit/test loop, two developer wrappers retain checksum-safe staleness checking while avoiding a clean rebuild for every Ada edit:
./scripts/dev-build.sh ./scripts/dev-test.sh --suite='fixture execution' ./scripts/dev-test.sh --case='harness/filters select exact cases' ./scripts/dev-test.sh --fixture=runtime/variant-match-selects-tag ./scripts/dev-test.sh --host --target=linux-x86-64 --suite=checking
The selectors are exact, accept one selection at a time, and print FILTERED in the transcript. They are fast feedback, not validation evidence. The developer build asks the pinned GPRbuild for checksum-based Ada recompilation; a changed source inventory, project file or selected GPRbuild executable still makes it clean. The ordinary build.sh and test.sh remain complete native Linux commands; --host selects Mac compiler checks. Nothing approves a revision now. The container command is retained troubleshooting, not a routine gate.
LANDIN_BUILD_MODE accepts only debug or release, before any build path is used. Builds and tests hold an OS lock for their host tag and mode through the entire command, including the build nested in test.sh. Different modes and tags can run concurrently. Host cleanup waits for both modes; clean.sh --all waits for every tag. The permanent lock files live in compiler/ada/.build-locks/, outside the directories cleanup removes. Python's standard-library fcntl supplies these locks on the supported macOS and Linux hosts; the kernel releases them when the last using process exits.
A session profile taken while aggregates and value layout were implemented had enough timestamps to set the priority: clean debug and release builds and whole-fixture runs occupied almost all measured time, while all six measured container lifecycle phases rounded to zero seconds and the complete one-shot startup stayed below one second. Keeping a persistent container would add state without attacking the bottleneck. The default Linux loop instead stopped calling build.sh before test.sh, since test.sh already owns its build, and test.sh --record-and-run now records and runs after one build when both are deliberately requested.
The loop asks for 4 GiB, and LANDIN_LINUX_MEMORY overrides it. That is not a preference. A release build is -O2 with -gnatn, so gprbuild's -j0 runs one gnat1 per core doing cross-unit inlining, and in a default-sized VM the kernel kills one of them:
gcc: fatal error: Killed signal terminated program gnat1 compilation terminated. compilation of landin-ir.adb failed
The unit named there is whichever was unlucky, not a unit with anything wrong in it — the same source builds in release natively and on the x86-64 gate. This is written down because the message reads exactly like a compiler defect in one file and cost an investigation once already.
environments/linux-amd64/Containerfile pins its base image by digest and verifies both Ada toolchain archives against the checksums in environments/pins.sh before unpacking either of them. It also installs the versioned Debian stable clang-19 package beside libc6-dev for C header extraction and generated-adapter tests. The frontend is deliberately separate from the pinned GNAT that builds refine: Clang supplies an external JSON AST, not a product backend. Its package comes from the container's existing Debian channel rather than a third download authority. The C boundary work refreshed that one base pin from Debian 12 to the official Debian 13 trixie-20260824 image index, which gives the local loop Clang 19.1.7. The gate installs Ubuntu 24.04's clang-19 at the exact version environments/pins.sh names; the binding generator's tests pass on both. GNAT and GPRbuild retain their existing versions and archive checksums.
GDB comes from that same Debian channel for scripts/debug.sh in the local container, and the Linux nix shell provides GDB from its locked package set. The gate runs the script with the release compiler and the GDB the pinned GNAT bundles. It checks source debugging of emitted programs, which is separate from debugging the Ada compiler itself.
Native GDB cannot read registers through Rosetta's ptrace interface: even /bin/true reports Cannot PTRACE_GETREGS and Couldn't get CS register. The local image therefore also supplies qemu-user. Its qemu-x86_64 GDB remote stub supports the same breakpoint, register and stack operations without Rosetta's ptrace path. This explicit local transport remains emulated evidence; the native gate uses GDB directly and has no fallback.
environments/pins.sh remains the one place an independently downloaded Ada toolchain version or checksum is written; check.py holds the recipe, compiler/ada/TOOLCHAIN.md and the nix shell to those same values, and checks the current native Mac tool policy against TOOLCHAIN.md. Objects are kept apart per host by LANDIN_BUILD_TAG, which scripts/env.sh defaults to os-arch: one checkout is built by two hosts, and .ali files from both in one directory is a build that fails confusingly.
scripts/toolchain.sh prints the host, the build mode and the exact compiler and builder versions, and build.sh and test.sh print it before doing anything. A captured log therefore names its own toolchain, which is what recorded evidence requires.
Recorded results
| date | environment | toolchain | result |
|---|---|---|---|
| 2026-08-20 | macOS arm64 (Darwin 25.5.0, Apple M1 Pro) | GNAT 16.1.0, GPRbuild 26.0.0 (aarch64-apple-darwin) | clean build; debug and release |
| 2026-08-20 | Apple Container 1.2.2, linux/amd64 under Rosetta, Linux 6.18.15 | GNAT 16.1.0, GPRbuild 26.0.0 (x86_64-pc-linux-gnu) | build from an empty build directory; debug |
| 2026-08-20 | builds.sr.ht debian/stable, Linux 6.12.94 x86-64 hardware, job 1867022 | GNAT 16.1.0, GPRbuild 26.0.0 (x86_64-pc-linux-gnu) | both archives verified against their checksums; clean build; debug and release; check.py clean; 47 seconds |
| 2026-10-02 | Apple Container 1.2.2, linux/arm64 Debian trixie on an Apple M1 Pro, 8 CPUs, 8 GiB | GNAT 16.1.0, GPRbuild 26.0.0 (aarch64-linux-gnu) | both archives verified against their checksums; clean build; the whole test program on the Linux arm64 lane, 800 cases, in 339 seconds; the GDB sessions with the bundled GDB 17.2 |
The retired builds.sr.ht gate job also printed refine --identify, so "no release version is assigned" appeared in the log of every run rather than only inside a test; the GitHub gate does not.
Case and check counts move as the suite grows, so they are not recorded here; the run itself is the record, and scripts/toolchain.sh output heads every one. What is recorded is that each environment built from clean and finished with no failures, in the modes named.
The two transcripts are byte-identical, which is the property worth having: the same cases in the same order with the same counts, on two hosts whose toolchains were built for different architectures.
One caveat, recorded because it was seen: a single early run ended in an unhandled exception and a traceback, and it has not reproduced in thirty subsequent runs including four from a clean checkout. The exception was not captured, so there is nothing to diagnose from. What changed as a result is that Landin.Testing.Run now catches an exception from a case, reports it as that case's failure, and keeps running the rest; a defect that used to take the whole transcript with it now costs one line of it.
A nix shell, for convenience
flake.nix provides nix develop with the pinned toolchain, contributed by ZAZPRO. It installs the same archives the container recipe and the GitHub workflows install, verified against the same checksums, because it reads environments/pins.sh rather than naming a nixpkgs attribute — at the time of writing nixpkgs carries GNAT 16.2.0 and GPRbuild 25.0.0, and the pin is GNAT 16.1.0 with GPRbuild 26.0.0.
It is defined for the three systems the pins carry checksums for, x86_64-linux, aarch64-linux and aarch64-darwin. The prebuilt .#refine-bin exists only for a system a release has published an asset for, with its hash recorded in the flake, so aarch64-linux has the shell and the source build .#refine and gains the prebuilt compiler with the first release that publishes it. The aarch64-linux shell is evaluated on x86-64, not yet built on arm64 hardware; Nix CI stays deferred.
The shell sets LANDIN_BUILD_TAG=nix-${system}: nix-x86_64-linux, nix-aarch64-linux or nix-aarch64-darwin. Shells sharing a checkout therefore use separate object trees and locks, apart from the other environments' trees. The nix build derivation still uses nix in its isolated build directory. python3 comes with it, so check.py and scripts/site.sh work in that shell too. On Linux it also selects llvmPackages."19".clang and glibc.dev from the package set fixed by flake.lock, matching the Debian environments' Clang major without following nixpkgs' default. The Darwin shell does not pretend that its SDK is a Linux sysroot; run binding-generator tests through scripts/linux-loop.sh there.
Its Linux behaviour is not settled by the local container, and that is not a formality. The pinned gprbuild dies with a segmentation fault when argv[0] names something other than the executable that is running — which is exactly what nixpkgs' makeWrapper arranges — and under Rosetta the same store paths ran without complaint. Evidence about this shell therefore comes from nix on x86-64 hardware, running the same scripts as every other environment.
Measured since, and the reason that is not merely a preference: the local loop cannot build this shell at all. Rosetta's binfmt handler replaces a custom argv[0] with the interpreted binary's own path — a program exec'd as I-AM-ARGV0 reads its own path back out of /proc/self/cmdline — so every nixpkgs wrapper built with --inherit-argv0 loses the name it was given. auto-patchelf is one of them: it runs under a bare interpreter that cannot import elftools, autoPatchelfHook fails, and the pinned archives are never patched. That is the paragraph above read from the other side — the same sensitivity, met by the translation rather than by the toolchain — and it is why this shell is answered for on hardware.
The first native Linux path needed one more thing of it, and only running programs found it. nixpkgs wraps a compiler under the names gcc, cc, g++ and cpp, and prefixes a driver's name with a GNU triplet only when it is cross-compiling; refine names its driver by triplet, so it reached the pinned archive's own x86_64-pc-linux-gnu-gcc rather than the wrapper, and linked with a compiler holding no libc — cannot find Scrt1.o. gprbuild was unaffected, because it links through the wrapper, which is why the compiler built in that shell and the programs it emitted did not. The flake now points every triplet-prefixed driver the archive ships at the wrapper.
Both of those were found by hand. The former automatic Nix manifest checked the shell at builds.sr.ht as a non-gate only when shell inputs changed. That automatic manifest is now retired. Run the same scripts/toolchain.sh and scripts/test.sh explicitly inside nix develop on native Linux after shell-input changes; this remains supplemental environment validation. Its first run settles both halves. The shell built on x86-64 nix with the pinned GNAT 16.1.0 and GPRbuild 26.0.0, and the suite passed 133 cases with no failures — runtime fixtures execute among them, which is the case that reported cannot find Scrt1.o and the reason any of this was looked at. That is a result for this shell and for nothing else: the table above is still what the three environments say.
The static-library directive's fixture also needs nixpkgs' separate static glibc output. The Linux shell includes it as a library input so an explicit linker.library("m") request can find libm.a. The shared glibc output precedes the archive output so the driver's own -lc stays dynamic; the shell check verifies that dependency on the static-library fixture. Job 1883411 passes the complete 418-case suite with 12,425 checks and verifies the dynamic libc dependency.
It is a convenience for editing on a nix machine and carries no authority of its own: the table above is unchanged by it, and a result produced in it is not evidence for any of the three environments. flake.lock pins the nixpkgs revision the shell is built from, but what it installs is decided by environments/pins.sh rather than by that revision.
What each environment is authority for
What the container does and does not settle is worth being exact about. It runs a real Linux kernel and a real x86-64 userspace, so it catches everything that depends on the operating system, the C library, the linker and the 64-bit-little-endian layout of the target: the whole suite passing there is real evidence, and it is why the Linux checksums in compiler/ada/TOOLCHAIN.md are now verified rather than transcribed. What it does not do is execute x86-64 instructions on x86-64 hardware — Rosetta translates them — so now that refine produces runnable executables, instruction-level and timing-sensitive results from this loop are not authority. That distinction is why the roadmap named the native gate before there was any code to run in it, and it is why the gate now exists: refine emits instructions. The original SourceHut gate ran them on their target hardware, the native acceptance did so through 0.2.0, and the gate does now, on every target.
Hosting is canonical GitHub, mirrored to git.sr.ht. GitHub Actions runs the gate, publishes the pages and checks host-independent emission; scripts/ci/ was removed with the acceptance, and no revision is accepted. The underlying compiler and test commands remain ordinary repository scripts.
Native LLDB runs through scripts/debug.sh --target=darwin-arm64 and the committed Mac policy. It validates emitted Landin programs with the pinned Apple debugger, dsymutil and dwarfdump, retaining dSYM/UUID and source-map matching evidence. The shared DWARF change selected routine Linux release GDB, and full hosted parity ran its dual-native milestone matrix in both compiler modes. See the native parity guide for the complete coverage and explicit physical-image limitation.
Embedded environment probes
Native Linux x86-64 with glibc 2.38 or later is the host for the pinned QEMU execution profile; its lock carries the Debian 13 runtime those tools load, and the gate runs it on Ubuntu 24.04. The retired native Linux acceptance's documents job executed and retained these small C/assembly probes from its exact archive. The Mac continues native compiler-host and Darwin workload/LLDB validation; it does not run Linux containers for embedded tests. QEMU owns CPU/startup evidence, and a repository-owned harness on its debugger stub runs the explicit synthetic peripheral lane on the same QEMU. Neither is physical hardware fidelity or Landin backend proof.
The gate's QEMU lane also runs with independently compiled C layouts and C/assembly ABI witnesses, compared with the compiler's Cortex-M planner and original synthetic-32 goldens. The retired acceptance export retained those artifacts beside the CPU and peripheral evidence. No embedded tools run on the Mac; the former native hosted acceptance policies covered their own work.
M0 scalar atomic, barrier and nested interrupt controls and bounded memory/cache models run on this same Linux path. Native hosted fixtures execute Landin atomics, generic evidence calls, SC fences and an escaping ordinary DMA slice on each real host. Cacheless emulator results never stand in for cached DMA maintenance. The retired dual-native archive bound all three evidence classes; the memory probe guide records limits.
Compiler-generated packed image execution runs on both native hosts and the Linux line-transport lane, including unnamed-encoding traps at all six optimization/specialization profiles. Independent M0 C controls and literal access traces remain separate from native Landin execution. The retired acceptance documents job built the archived compiler and exported these lanes beside previous embedded evidence; that lane on its own was not a Cortex-M compiler backend.
The same pinned tools support compiler-generated Cortex-M0 execution of the inventoried shared runtime corpus and direct synthetic-device packed/byte/ordinary-slice DMA transactions. The external startup and linker test harness does not enable Landin firmware entry or vectors. The retired export retained objects, ELF/map/disassembly, literal ABI and device oracles, helper identities and bounded image/stack observations; complete measured firmware and Landin source debugging belong to the freestanding evidence lane. Native GDB/LLDB coverage checks each hosted backend separately.
The compiler-owned Cortex firmware lane runs only in the existing native Linux embedded environment. It uses the same pinned tool inventory and produces generated startup/vector/linker inputs and fresh-image comparisons. The retired acceptance export included them. The external backend harness, hosted line transport and independent C/assembly probes remain separate evidence. See the firmware execution contract. The Mac remains the native compiler-host/Darwin/LLDB lane; no local Linux container or Nix CI path is introduced.
The generated-device lane also runs in the native Linux embedded entry, after the inherited probes. The retired export retained separate devices artifacts; vendor-input and regeneration checks run offline on both hosts. The selected vendor is provenance, not a replacement QEMU board.
The complete derived driver runs in the native Linux embedded entry. Its QEMU boot, separate synthetic protocol model and independent layout control retain their distinct evidence roles. Native GDB/LLDB routine risk coverage validated the checking and linker repairs it needed.
The evidence.py lane adds actual Cortex line/function GDB sessions and complete-application resource scenarios to that Linux embedded entry. It preserves the fixed board map and separate CPU/peripheral/control identities. The retired export retained source/debug matching, ELF load accounting, SP/paint/frame observations and bounded results. The retired dual-native milestone policies retained both compiler modes and native GDB/LLDB; no Cortex debugger session replaces a native one. See the evidence contract.