A uniform, two-tier test environment for SIMPLE¶
Date: 2026-09-22
Status (2026-09-25): completed and archived. Phases 0 to 4, 6 and 7 are
complete; Phase 5 is Ruben's and continues in its own handover (below). Every compile_*.sh --compile-tests build
first checks that the CTest registrations, the test UI and the test routers
agree, then runs the fast gate: thirteen fast area suites, about 5 s real
on the reference Mac in Debug, under the 30 s budget and the process-count
ratchet (SIMPLE_CTEST_BUDGET = 25: 13 fast, 5 library, 6 workflow,
1 platform). The review of everything else (Phase 3) gave every identity a
verdict and built the five nightly library suites (lib_reconstruction,
lib_cart_align3D, lib_heterogeneity, lib_single, lib_stream) as it
went; simple_test_exec is the only test executable. The tests written by
the review found and fixed the production defects recorded in section 9.7
and removed the dead routines recorded there. What remains is Phase 5, the
simulation-truth gates of the workflow entries and the nightly runner, and the
checks nobody has observed yet (section 16, criterion 10); both are carried by
doc/refactoring_notes/phase5_workflow_gates_and_nightly_runner_handover.md,
the live document from here on. The policy for day-to-day work (what a test
is, where it goes, how to write it, what runs when, and where every old test
went) is doc/policies/test_environment_policy.md; the review's verdicts and
retired tests are doc/refactoring_notes/completed/test_review_record.md. This was a large
project with four workstreams (section 1.1), delivered in slices that were
each useful on their own.
Validation level: static source inspection of SIMPLE, and of X's test system
(production/CMakeLists.txt, AGENTS.md "Test admission" and "Test tiering",
docs/design_notes/completed/test_suite_refactoring.md) as the reference
design. No SIMPLE source was compiled and no test executable was run while
preparing this plan; the timing inventory in Phase 0 is the first thing that
will change that.
This record is closed. It keeps the design, the migration and the batch records (section 9.7) as they happened; it is not updated any more. New work on the test environment follows the policy, and Phase 5 is recorded in its handover.
1. Objective¶
SIMPLE gets two test gates with one front door.
The fast gate runs as part of the build whenever --compile-tests is on:
compile_*.sh --compile-tests builds, installs, then runs ctest on the
fast label, and the whole label finishes within 30 s of wall time on the
reference Mac (24 cores; CI runners get a timeout multiplier, not a bigger
budget). Every test in it can fail on its own assertions, runs in-process
inside a fused per-area suite, needs no network, no download and no external
data, and pins itself to one OpenMP thread. This is the everyday loop.
Its content is decided (owner decision, 2026-09-22): the fast gate is the
unit suites that simple_test_units runs today, 39 sub-suites (the union of
its two routes) over the *_tester modules and the test_* procedures of
core types, split into a handful of area suites. Nothing else in the tree is
a fast-tier candidate.
Another test joins only by passing the admission rules of section 5.1 through
a review verdict, never by default.
The extensive gate runs overnight on a dedicated machine and has two
kinds of entries. Library suites group the tests of one coherent part of
the library (Fourier transforms, geometry, masks, numerics, optimisation,
statistics, I/O, projects, search, ...) into one process each, built from
the standalone tests that survive the review; they can take minutes, use
real-sized fixtures, and must fail on their own assertions. Workflow
gates are Ruben's self-contained simulated workflows (simulated_workflow
on the embedded 6VXX/1JYX systems, single_workflow, mini_stream, the
stream suite, pcg_recon, rec3D_backends, ...), which
generate their own data from known atomic models. Because the truth is known,
these workflows can be gated on it (section 5.2) without the real-data
validation registry and blessed baselines that X's validate needs; that
machinery is too costly for SIMPLE and is a non-goal (section 17).
One front door serves both:
simple_test_exec test=<suite-or-test> [arguments]
The standalone program simple_test_* sources under production/tests are
converted into callable procedures in grouped test modules and dispatched
through area test commanders. The test executable and every test-only source
stay out of production builds (BUILD_TESTS=OFF, the default of every
compile_*.sh since 2026-09-21).
Uniformity alone does not give a build-time gate. Three facts about the
current tree shape the plan: nothing today bounds what a test run costs;
every test is its own process, which is the direction X had to reverse after
reaching 172 CTest processes; and more than half of the existing tests cannot
fail, because they contain no assertion, no error stop, no THROW_HARD and
no tests_failed check (section 4.4). The tiering therefore comes first, the
fast tier has an admission rule, and quality and performance of the fast-tier
tests is a workstream of its own.
1.1 Workstreams¶
| Workstream | Depends on | Value on its own | |
|---|---|---|---|
| A | Fast gate scaffolding: ctest in the compile scripts, labels, working directories, thread pinning, timeouts, the budget check, the review dossier script, registration of the tests that already qualify |
nothing | yes: a build-time gate over today's assertion-bearing tests |
| B | Review and unification: per-area review verdicts (section 9), then standalone programs into grouped modules and fused suites behind simple_test_exec; overlaps merged, dead tests deleted; the glob removed |
A for the registration shape and the dossiers | yes, per area |
| C | Fast-tier performance and hermeticity: split units into area suites, reconcile its two routes, time each sub-suite, fit the 30 s budget, confine or move the sub-suites that use sockets, HTTP or child processes |
Phase 0 timing of units |
yes: the gate gets faster and more local with each step |
| D | Extensive tier: the library suites assembled from the review's survivors, simulation-truth gates on the simulated workflows, the overnight runner, result archiving | B for each area | yes, per library suite and per gated workflow |
Phase 0 and A come first. B, C and D proceed by test area and can interleave.
2. What SIMPLE takes from X, and what it does not¶
X's xlms_test is, by its own comment, "SIMPLE's simple_test_exec pattern":
every unit test a commander, dispatched by name. On top of that X added, and
SIMPLE adopts:
- Fused suites. The fast checks run as about ten in-process
bounded-context suites (
xlms_test unit_io), each accumulating ordinary failures and stopping once at the process boundary. Focused selectors stay available for debugging. Only genuine process-death contracts, the binary smokes and resource-hungry cases keep their own process. - A ratcheted process budget.
XLMS_CTEST_BUDGETmust equal the registered count or configuration fails; the number only goes down (172 to 22). "Checks are free; processes are budgeted." - A time budget on the label. The regression label stays under 15 s;
the everyday loop is
ctest, run bycompile_debug.shandrecompile.shright after the build. (X also asks any entry above 0.5 s in Release to say why; SIMPLE does not adopt that one, section 5.1 rule 4.) - Hermetic tests. No outcome depends on free memory, core count, load, wall clock, network or another test's files; every entry owns its working directory; ordinary work runs on one OpenMP thread.
- Abort tests only where dying is the claim, accepted on the dedicated hard-error status, never on "any non-zero exit".
SIMPLE does not adopt X's validate (a registry of real datasets and arms run
through the production paths, compared against platform-keyed blessed
baselines with a report directory and gated blessing). The datasets, the
curation and the blessing discipline are more than SIMPLE can carry. The
extensive tier uses simulated data with known truth instead (section 5.2).
3. Preservation contract¶
The refactor must preserve:
- Every supported test assertion, fixture, argument, deterministic seed, expected failure, output artifact and success/failure exit status. A test that has no failure path has no behaviour to preserve in this sense: it either gains assertions (workstream C) or is reclassified as extensive or manual; it is not admitted to the fast gate as it stands.
simple_test_exec test=listas the discoverable catalogue of suites and tests.- The ability to execute one named test independently, in its own process, for debugging.
- Process isolation where an existing mother suite deliberately launches
child cases, where a test's contract is that the process dies, or where a
platform launcher (
cafrun,mpirun) requires a separate process. - Conditional availability for coarray, MPI, OpenMP offload, OpenACC, CUDA, network and external-data tests.
- Existing production-library behaviour. This is a test-architecture change, not a numerical, scientific or workflow refactor. Adding an assertion to a test is in scope; changing what the code under test does is not.
BUILD_TESTS=OFFexcludingsimple_test_execand every test-only source from production builds (in place since 2026-09-21).
The migration must establish one authoritative implementation for each test. Temporary wrappers may exist within a migration phase, but duplicate program and commander bodies must not remain as the completed state.
4. Current-state audit¶
Four terms are used with fixed meanings throughout:
- Canonical test identity: one
test=<name>value. There are 150. An identity may be implemented on one or both routes. - Route implementation: one implementation of an identity, either a
standalone
program simple_test_<name>underproduction/testsor a commander procedure reached bysimple_test_exec test=<name>. There are 203 (107 standalone, 96 commander); 53 identities have both. - Unit sub-suite: one
begin_test_suite/end_test_suitegroup insideunits(39 in the union of its two routes). Sub-suites are not identities and are not registered individually. - Registered CTest process: one
add_testentry. This is what the process budget counts; it is neither an identity nor a sub-suite.
All headline counts below are regenerated by scripts/test_review_dossier.py
from the inventory and are per identity unless stated per route.
At the 2026-09-22 baseline, static source inspection finds:
| Test representation | Count | Current ownership |
|---|---|---|
Standalone route implementations (production/tests/simple_test_*.f90) |
107 | Auto-globbed into separate executables and CTest entries by production/CMakeLists.txt |
Commander route implementations (simple_test_exec test= dispatch cases) |
96 | Fourteen topic routers under src/main/exec; 95 exec_test_* procedures plus exec_volume_shape_descriptors |
| Canonical test identities | 150 | 107 + 96 − 53 on both routes |
| Identities with two route implementations | 53 | Two independently maintained implementations or wrappers |
Reusable *_tester.f90 modules |
24 | Scattered alongside the production domains they test |
The existing simple_test_exec route is:
production/simple_test_exec.f90
-> test UI registry
-> topic execution router
-> test commander type
-> test body or reusable tester procedure
The standalone route is:
production/tests/simple_test_<name>.f90
-> program body
-> optional suite-specific helper modules or *_tester modules
The standalone programs are discovered with a filename glob. Each is compiled,
linked, installed and registered separately with CTest, with no TIMEOUT, no
WORKING_DIRECTORY, no labels and no thread pinning. CI runs 30 of them by
name plus two simple_test_exec test=... calls, all sequentially.
4.1 Duplication and drift¶
Exact-name overlap does not guarantee equivalent behaviour. The standalone
simple_test_units and the test=units implementation in
simple_commanders_test_class maintain separate import lists and suite
schedules and have already diverged: the standalone program includes suites
the commander version does not, while the commander version has its own
additions and terminal message. Several other pairs contain copied program
bodies. A correction made in only one path leaves the other stale.
4.2 Registration overhead¶
The 96 simple_test_exec cases are represented in three parallel layers: UI
program construction under src/main/ui/simple_test, commander types and
implementations under src/main/commanders/test, and topic-specific
select case routing under src/main/exec. The UI and execution registration
remain useful because tests accept typed SIMPLE arguments and are discoverable
through test=list. The problematic part is treating each small test as its
own commander object when a coherent area commander could call a test-module
procedure directly.
4.3 Tests are not all the same kind¶
The current sources include small deterministic unit tests; numerical and
file-format regressions; multi-stage workflow tests; mother suites that launch
isolated child cases; performance and I/O benchmarks; coarray, MPI,
accelerator, SIMD and OpenMP platform tests; socket client/server roles; tests
requiring an external reference volume (11 sources read vol1=, stk= or
projfile=); and tests that download data or need other external services.
Uniform entry does not mean forcing these into one process or pretending they
have identical requirements. The tiers in section 5 make the requirements
explicit.
4.4 Failure paths¶
Only 36 sources use simple_test_utils (assert_*, begin_test_suite,
report_summary), and only 20 turn tests_failed into a non-zero exit. The
rest fail by error stop (31 standalone programs) or THROW_HARD (which ends
in error stop 1), or not at all. Per route, 61 of the 107 standalone
programs and 57 of the 96 commander procedures contain no assertion, no
error stop, no THROW_HARD and no tests_failed check; per identity, 84
of the 150 have no failure path on any route. They print, and they pass
whenever they do not crash. Examples: angres, corrs2weights, eigh, ft_expanded,
gencorrs_fft, io, io_parallel, lbfgsb, mask, neigh, otsu,
pca_all, serialize, sym, uniform_euler, uniform_rot. A build-time
gate made of these would be green without meaning anything. Workstream C
exists for them.
simple_end is not a terminator: it prints a banner and optionally touches a
file, then returns. The terminators that matter for a fused suite are
error stop, THROW_HARD and the bare stop on the test=list path of
simple_test_exec.
4.5 Timing¶
Unknown. No per-test run time has been recorded. Phase 0 records it.
4.6 units, the fast gate in embryo¶
simple_test_units (and its commander twin, test=units in
simple_commanders_test_class) already has the shape the fast gate needs:
one process, seed_rnd, its own working directory, and its sub-suites (38
on the standalone route, 36 on the commander route, 39 in the union) each
run inside begin_test_suite / end_test_suite over the *_tester modules
(string, syslib, fileio, the hashes, linked list, cmdline, ori,
oris, rec_list, starfile, project merge, motion gain, ...) and the
test_* procedures of core types (image, imghead, ftiter,
ftexp_shsrch, online_var, bspline_smoother, aff_prop, hclust,
srchspace_map2D_io), plus validate_ui_json, then report_summary and
error stop 1 if anything failed. The two routes have diverged: the
standalone runs class compatibility, particle sieve and
2D search-space map I/O, which the commander does not; the commander runs
atoms, which the standalone does not. Four sub-suites (IPC TCP socket,
HTTP POST, forked process, persistent worker server) use sockets or
child processes and need checking against the hermeticity rule. It is one
CTest entry today (units, called by CI as simple_test_units), so a
failure anywhere fails the whole thing and ctest --parallel cannot spread
it.
Measured 2026-09-22 (Debug, one OpenMP thread, reference Mac, all sub-suites
passing): 21.6 s standalone, 21.5 s through simple_test_exec, so the
front door costs nothing measurable and the two routes are equivalent apart
from the four sub-suites they do not share. (An earlier run the same day
gave 47 s and 71 s; the machine was loaded, and the difference between the
routes it suggested was noise.) simple_test_utils now times each sub-suite
(commit 045c6ba39); the per-sub-suite times are:
| sub-suite | s | note |
|---|---|---|
| forked process | 12.85 | 28 assertions; the time is c_usleep polling at FORK_POLL_TIME (100 ms) around real child processes |
| Fourier shift search | 4.37 | test_ftexp_shsrch: 240-pixel box, 50 trials, 20 noisy images, at -O0; 2 hard checks |
| image | 0.87 | |
| orientation data | 0.73 | |
| HTTP POST | 0.56 | localhost |
| orientation collection | 0.34 | 155 checks |
| particle sieve | 0.32 | standalone route only |
| class compatibility | 0.20 | standalone route only |
| UI JSON | 0.18 | |
| every other sub-suite | < 0.15 | 30 sub-suites, 1.2 s together |
Two sub-suites were 80% of the time. Without forked process the whole of
units is 8.7 s single-threaded; without both it is 4.3 s. The 30 s budget
is therefore met already on one thread, and the area split is about failure
locality and parallel spread rather than about fitting the budget. Both
decisions were taken the same day (owner):
forked processis not part of the build. It spawns real children and polls them on a clock, which is what the hermeticity rule excludes. It keeps its assertions and moves to theplatformlabel in the nightly run (section 5.3); in Phase 1 it is still inside the provisionalunitsentry, and Phase 2 takes it out.test_ftexp_shsrchruns on a 128-pixel box (was 240,SQRADscaled 60 to 32;TRS,HP,LPunchanged). The change also removed two things the test did without checking anything: 20 noisy images generated and never used, andprofile_corrs, a benchmark that filled and FFT'd two 4096-pixel images to print CPU times. Committed 2026-09-22 and measured the same day: 4.37 s before, 0.30 s after, both routes, checks passing.
5. Tiers and admission rules¶
Every test identity in the inventory (section 8) gets exactly one tier. The tier decides how the test is registered, where it runs and what it may cost.
5.1 Fast tier (fast label, part of the build)¶
Admission rules, all of them:
- It can fail. The test makes at least one assertion through
simple_test_utils(or an equivalent typed check) whose failure reaches the process exit status. "Completed without crashing" is admitted only for a closed set of binary smokes registered with aPASS_REGULAR_EXPRESSION(X has four; SIMPLE should need no more than a handful: onesimple_execprogram with a pass pattern, onesimple_private_exec, onesingle_exec, onesimple_streaminvocation that starts and stops). - It is hermetic. No network, no download, no user-supplied file, no dependence on core count, load, wall clock or another test's files. Fixtures are generated in the test from a fixed seed, or committed to the repository and small.
- It runs in-process inside its area suite (section 6) on one OpenMP
thread, and returns on success. A sub-suite whose subject is a threaded
path opens its own small team with
num_threads(the discrete stack reader ofunit_coreand the masks ofunit_imageuse three); the entry'sOMP_NUM_THREADSstays 1. It does notstop, does not change the working directory without restoring it, and does not leave global state (random-number generator, module variables, open units) that the next test in the suite can see. - It is cheap. The whole
fastlabel finishes in 30 s of real time on the reference Mac withctest --parallel, in every build type. The 30 s is a ratchet: it does not go up, and additions that would exceed it are paid for by shrinking fixtures or moving something out. There is no per-entry time rule: X's "an entry over 0.5 s says why" guarded a process count that SIMPLE'sSIMPLE_CTEST_BUDGETalready guards, and it would measure the front door rather than the tests (asimple_test_execprocess pays about 1 s at-O0for the test UI registry before its first check; 2026-09-22:unit_numerics, whose sub-suites take 0.1 s, runs in 1.07 s). Insteadctest_budget.pywrites the per-entry table beside the log on every build, so a suite that grows is seen when it grows. - It has a
TIMEOUT(default 60 s) and its ownWORKING_DIRECTORYunderbuild/test_runs/<suite>.
Members: the sub-suites of units (section 4.6), regrouped into area suites
along these lines, to be settled by the Phase 0 timing of each sub-suite:
| suite | sub-suites of units |
Debug, 1 thread (2026-09-22) |
|---|---|---|
unit_core |
string (with comma-separated integer lists and ANSI formatting since the utils review), syslib, fileio, stack I/O (with the discrete reader, three threads, since the singles review), character hash, hash, value-reference hash, linked list, record list, command line (with a full processing line since the utils review) | 0.2 s |
unit_ori |
orientation, orientation collection, symmetry, Euler shift (orientation data retired 2026-09-25) | 1.2 s (3.6 s with symmetry) |
unit_image |
image, image header, Fourier iterator, B-spline smoother (one sub-suite, 2D and 3D, since 2026-09-25), masks (with the threaded path on a team of three since the wrap-up), binary image, segmentation, trailing-reconstruction blend, CTF, image serialisation (utils review) | 1.3 s (before the shift search moved out and the mask suites moved in) |
unit_numerics |
online variance, random draws (shuffles and the multinomial draw, asserting since the singles review), affinity propagation, statistics (weights), shift search (one sub-suite since 2026-09-25; hierarchical clustering retired with hclust), cavg quality relations, diffusion-map graphs — the ft_expanded shift search is a motion-correction optimiser, not an image test (Hans, 2026-09-22) |
0.1 s before the additions |
unit_project |
STAR file, STAR project (with the RELION phase-shift contract), project merge, class compatibility, particle sieve (with the collector's hard-gate rejection since the stream review), motion gain (2D search-space map I/O retired with the module 2026-09-25) (atoms moved to unit_single) |
0.7 s |
unit_ui |
UI JSON, GUI metadata, GUI assembler, UI hash, UI visibility | 0.2 s |
unit_ipc |
IPC TCP socket, HTTP POST, persistent worker server, persistent worker message — localhost only, bounded; forked process is excluded by decision and goes to platform |
0.6 s |
unit_reconstruction |
rec3D backend, observation noise, class-average accumulator — added by the reconstruction review (2026-09-23, section 9.7); pcg_recon joins once its one-thread time is known |
0.9 s |
unit_pftc_align2D3D |
polar correlation (gen_objfun_vals and calc_frc on generated images, since the wrap-up and the open items), continuous in-plane, refine3D in-plane state, 2D probability table I/O, sigma2 state, class-average registration (utils review) — added by the inplane review (2026-09-23, section 9.7); registration on the polar Fourier transform, shared by the 2D and 3D searches | 0.4 s |
unit_cart_align3D |
Cartesian Fourier, pose refiner, pose adapter — added by the pose review (2026-09-23, section 9.7); the Cartesian (continuous) 3D registration; its nightly counterpart lib_cart_align3D holds the 1JYX recovery gate |
0.1 s |
unit_heterogeneity |
flex PCA (deconvolution of 4 000 particles), flex PCG operator (box 32, baseline solve) — added by the heterogeneity review (2026-09-23, section 9.7); its nightly counterpart lib_heterogeneity runs the deconvolution on 20 000 particles, the operator at box 64 and the solve sweep; flex_gpu is the CUDA platform entry |
4.5 s (Mac; 47.3 s in the first build, cut down, section 9.7) |
unit_parallel |
qsys control, qsys environment — added by the parallel review (2026-09-23, section 9.7); distributed execution, scripts only, nothing submitted | to be measured |
unit_single |
atoms, C-alpha finder — added by the single review (2026-09-23, section 9.7); SINGLE (nanoparticles, atomic models); its nightly counterpart lib_single holds Ruben's pipelines and pdb2mrc of the built-in models |
to be measured |
Measured through the gate on 2026-09-22 (Debug, ctest -j12, one thread
per entry): all seven pass, 3.3 s real, 12.9 processor-seconds;
unit_ori 3.3 s, unit_image 2.6 s, unit_ipc 1.8 s, unit_project
1.7 s, unit_ui 1.25 s, unit_core 1.2 s, unit_numerics 1.1 s. About
1 s of every entry is the front door (the test UI registry at -O0), which
was paid once when units was one process and is paid seven times now;
the parallel spread more than covers it. The budget is comfortable, and the
process count for the ratchet is seven.
No other test identity in the tree is a fast-tier candidate by default. The
remaining 149 identities are reviewed for the extensive tier, manual use or
deletion (section 9); one of them joins the fast gate only through a keep
or modify verdict that states which admission rules it meets and what it
measured in Phase 0.
5.2 Extensive tier (library and workflow labels, overnight)¶
The extensive tier runs nightly on the dedicated machine as
ctest -L "library|workflow". It has two kinds of entries.
5.2.1 Library suites (library label)¶
A library suite is the tests of one coherent part of the library, run in one
process through simple_test_exec test=lib_<area>, one CTest entry per
suite. It is the extensive-tier counterpart of the fast area suites: the same
fused shape (each member inside begin_test_suite / end_test_suite,
report_summary at the end, non-zero exit on any failure, a focused selector
for one member), without the 30 s budget. Members are the standalone tests
the review keeps (demote to lib_<area>), so a numerical test that takes a
minute on a realistic box has a home and runs every night instead of never.
Admission rules for a library suite member:
- it can fail: at least one assertion whose failure reaches the process
exit status (a
demoteof a print-only test carries the assertion it must gain, section 9.4); - it runs unattended: no user-supplied file, no download, fixtures generated from a seed or committed; a member may take minutes and use full-sized boxes, but it must finish;
- it is deterministic across runs on one machine (declared seed) and
restores the working directory and any module state it changes, since it
shares a process with its suite; CTest sets
SIMPLE_SEEDfor every entry, which fixes the seedseed_rnddraws from (parameters%newcalls it, so every commander does), and the runner reseeds before every sub-suite (section 9.7, stream); - it is registered with a
TIMEOUTand the suite's total is recorded in the nightly summary, so growth is visible.
Suites as built (2026-09-24). The provisional list drawn from the router
areas before the review (lib_fft, lib_geometry, lib_masks,
lib_numerics, lib_optimize, lib_stats, lib_io, lib_project,
lib_search, lib_parallel) is superseded: the review sent the members of
those areas to the fast area suites, to the suites below or to a program, or
retired them (section 9.7). The last three, gencorrs_fft, angres and
msk_routines, became the polar correlation sub-suite of
unit_pftc_align2D3D, the program measure_projspace_angres and the
threaded path of masks in unit_image (section 9.7, the wrap-up).
| suite | sub-suites |
|---|---|
lib_reconstruction |
PCG half-set |
lib_cart_align3D |
pose 1JYX recovery |
lib_heterogeneity |
flex PCA deconvolution 20k, flex PCG operator 64, flex PCG solve sweep |
lib_single |
nanoparticle atoms, C-alpha molecules, pdb2mrc |
lib_stream |
optics assignment, picking references, pick and extract |
Their runtimes are recorded by the first night of the runner (Phase 5).
Registration: one entry per suite, LABELS library, its own working
directory, OMP_NUM_THREADS set explicitly (a suite may use a team, since
the library suites run before the serial workflow gates and can be scheduled
against each other), a TIMEOUT sized to the suite. A library suite is not
a workflow: it does not start simple_exec or distributed workers; a test
that needs them is a workflow gate.
5.2.2 Workflow gates (workflow label)¶
The workflow gates are the simulated workflows, gated on the truth they were
simulated from. Today simulated_workflow, single_workflow, mini_stream
and stream_preproc check that the pipeline completes: files exist,
counts match, the heartbeat is well-formed, abinitio2D produced classes.
They do not compare the result with the model that generated the data. Since
the data comes from embedded atomic coordinates (6VXX, 1JYX) with known
orientations, defocus and B-factor, the comparison is available at no data
cost. Each workflow declares floors such as:
- resolution: FSC=0.143 between the reconstructed map and the ground-truth map simulated from the same coordinates, at or better than a declared value for that system, box and particle count;
- poses: fraction of particles whose recovered orientation is within a declared angular distance of the simulated one, after symmetry and hand alignment; shift error likewise;
- workflow structure: particle, class and state counts, heartbeat completeness, project-file consistency (the checks that exist today);
- stream: the same, per stage, over the simulated movie set.
Floors are declared in the test, versioned with it, and loosened only with a written justification in the commit. There are no blessed baselines to maintain, no platform keys and no report registry; the run archives its summary (commit, host, compiler, per-workflow metrics and pass/fail) into a directory on the dedicated machine so a regression can be dated.
Registration: one CTest entry per workflow, LABELS workflow,
RUN_SERIAL TRUE (each owns the machine's OpenMP team and may start
distributed workers), a long TIMEOUT, its own working directory. Members
as registered (2026-09-24): simulated_workflow_6vxx,
simulated_workflow_1jxy, single_workflow, pcg_recon,
simulate_particles (which absorbed reproject) and stream_preproc (the
in-process stream stages are in lib_stream). The review settled the rest
of the original list: mini_stream, pcg_frac_update and rec3D_backends
need user data and are manual; the nano workflows are sub-suites of
lib_single (nanoparticle atoms, C-alpha molecules); the second picker
of simulated_workflow is open. The truth gates and their floors are
Phase 5, assigned to Ruben
(doc/refactoring_notes/phase5_workflow_gates_and_nightly_runner_handover.md).
Its map gate is specified there: the abinitio3D map is neither docked to
the truth map nor necessarily of the right hand, so the gate makes the truth
map with pdb2mrc on the same grid, docks the map to it with dock_vols at
a low-pass of 15 to 20 A, keeps the hand (the map or its mirror('x')) that
docks with the higher correlation, and takes the masked FSC at 0.143 against
a declared floor; poses are compared after composing each recovered
orientation with the docking rotation (and the mirror).
5.2.3 The nightly run¶
ctest -L "library|workflow" after a clean --compile-tests build, started
by cron or a scheduler on the dedicated machine. The library suites run
first (in parallel, each pinned), then the workflow gates serially. The
summary (commit, host, compiler, per-suite and per-workflow times, metrics
against floors, pass/fail) is archived into a dated directory on that
machine and mailed or written where the team looks. The whole run must fit
the night; a suite or workflow that grows past its share is reported by the
summary, and trimming it is a reviewed change, as for the fast gate.
Ruben designs and writes the runner (2026-09-24, the Phase 5 handover,
part B): a locked, scheduled run of ./compile_clean.sh --compile-tests
on a known commit, then ctest -L library, ctest -L workflow, and
ctest -L platform where the machine has the capability; a dated summary
directory outside build/ (which the compile scripts delete) with the
status, times and metrics.tsv of every entry; a cumulative history file;
and a short notice where the team looks.
5.3 Isolated and platform tests (platform label)¶
Coarray (cafrun -np 2), MPI, OpenMP offload, OpenACC, CUDA, socket
client/server pairs, the forked process suite (real child processes,
clock-based polling) and expect-abort tests keep their own process. They are
registered only when CMake has confirmed the capability and launcher, carry
the platform label, and are excluded from the fast gate unless a specific
entry fits the budget on the reference Mac and is hermetic (the coarray test
currently runs in its own CI job and stays there). Socket tests must be
bounded and self-terminating or be driven by an orchestrating test that
starts both roles and checks both exit statuses. Expect-abort tests are
accepted on the SIMPLE hard-error status (error stop 1 from
simple_exception), never on "any non-zero exit", and exist only where dying
is the contract.
5.4 Not registered¶
Benchmarks (io_parallel, simd, openmp as timing tools), tests that
download data, and tests that need a user-supplied volume or stack remain
runnable by name through simple_test_exec and are documented as manual.
They are not CTest entries.
6. Target architecture¶
simple_test_exec
-> test UI metadata and argument parsing
-> topic execution router
-> area test commander (suite: runs every test of its area in-process,
accumulates failures, stops once at the process boundary;
focused selector: runs one named test)
-> grouped callable test module
|-- focused test procedures
|-- suite-owned fixtures, built once per process and handed over
`-- reusable domain tester modules
-> production SIMPLE APIs
| Layer | Owns | Does not own |
|---|---|---|
simple_test_exec |
Process front door, command parsing, timing, logging, memory monitoring, final process status | Individual test algorithms |
| Test UI | Suite and test names, arguments, help, required inputs, defaults | Test execution |
| Topic router | Routing a registered suite or test to one area commander | Test bodies |
| Area test commander | Running its area's tests in one process, shared fixtures and setup, failure accumulation, workflow orchestration in the extensive tier | Low-level assertions or copied test algorithms |
| Grouped test module | Callable test bodies and closely related suite helpers | CLI front-door behaviour or production orchestration policy |
Existing *_tester or domain test APIs |
Reusable focused checks close to the type or subsystem under test | Executable lifecycle |
| CMake/CTest | Build gating, launcher selection, labels, timeouts, working directories, thread pinning, the budgets | A second implementation of test behaviour |
6.1 Grouped callable modules¶
Related tests are grouped into modules named for their domain. A provisional grouping, to be settled by the Phase 0 inventory:
simple_test_cases_core
simple_test_cases_fft
simple_test_cases_geometry
simple_test_cases_io
simple_test_cases_masks
simple_test_cases_numerics
simple_test_cases_optimization
simple_test_cases_parallel
simple_test_cases_project
simple_test_cases_reconstruction
simple_test_cases_stream
simple_test_cases_workflows
simple_test_cases_utils
Large numerical suites that already have a mother module and several focused helper modules keep that cohesion rather than being pasted into a monolithic topic file. A grouped module is private by default and exports only its callable entries:
module simple_test_cases_io
use simple_cmdline, only: cmdline
implicit none
private
public :: test_imgfile
public :: test_mrc_validate
public :: test_stack_io
contains
subroutine test_stack_io(cline)
class(cmdline), intent(inout) :: cline
! Migrated standalone test behaviour, returning on success.
end subroutine test_stack_io
end module simple_test_cases_io
A common dummy argument is not imposed where it adds nothing: a test with no
inputs stays a no-argument procedure; tests that need CLI values take
cmdline at the pre-parse boundary or, preferably, typed parameters after
their area commander has initialised them; tests that exercise one production
object may receive it directly.
6.2 Area test commanders and suites¶
One commander per coherent test area, not one per test:
type, extends(commander_base) :: commander_test_io
contains
procedure :: execute => exec_test_io
end type commander_test_io
test=unit_core runs its sub-suites in one process: it builds any shared
fixture once, calls each callable procedure inside a begin_test_suite /
end_test_suite pair, and after the last one calls report_summary and exits
non-zero if anything failed, exactly as units did for all of them at once.
(Implemented in simple_commanders_test_class: each area is a table of
unit_suite entries, a name and a no-argument procedure, built by
suites_<area> and run by run_unit_suites; procedures that take arguments
are wrapped.) The area suites registered with CTest are the authoritative
gate.
test=units remains as a developer convenience that runs every fast suite in
sequence in one process; it is not authoritative, and a failure it shows that
the CTest gate does not is a state leak between suites, to be fixed in the
suite that leaks. A focused selector runs one sub-suite for debugging:
test=unit_core suite=hash, where suite is an optional string input
registered on each area program in the test UI (its help lists that area's
sub-suite names, lowercase with underscores) and read from cline in the
area commander, in the way simulated_workflow reads system. Like
system, it needs a field in simple_parameters (suite), because the
command-line parser accepts only keys that the generated argument list
knows.
The area commander is what CTest registers; the focused names are what a
developer types.
For the extensive tier the area commanders are the library-suite commanders
(test=lib_reconstruction, ..., section 5.2.1), one per suite and the same
fused shape as the fast ones (their tables are in
simple_commanders_test_class beside the fast ones), and the workflow
commanders (simulated_workflow, single_workflow, pcg_recon,
simulate_particles, preproc), one CTest entry each (section 5.2.2).
Per-test commander types are removed as their bodies migrate. The fourteen topic areas are a starting point; the inventory may split, merge or rename them by real test ownership.
6.3 Test procedure lifecycle¶
Callable test procedures must:
- return normally on success;
- report checks through
simple_test_utils(or an equivalent that reachestests_failed); - signal failure through the accumulating path, reserving
THROW_HARDfor conditions that make continuing the suite meaningless (a missing fixture); - leave process-wide timing, Git-version output, log closure and memory
monitoring to
simple_test_exec; - not parse raw process arguments;
- not call
stopor any success-path terminator; - restore the working directory and release resources they own, including on early-failure paths, and reset any module-level state they changed;
- read the fixture handed to them by the suite rather than rebuilding it, and build their own only when run alone.
Contained procedures of a standalone program move into the owning grouped module or an existing suite-specific helper module, never into a commander.
6.4 Process isolation¶
Ordinary fast-tier tests share a process by design (section 5.1). Isolation is kept for the cases in section 5.3 and for mother suites that already launch child cases; those launch the same executable with another selector:
simple_test_exec test=pose_cont_refinement case=<case-name>
An eventual test=all convenience runs the suites in-process one after
another and delegates only the isolated cases to subprocesses.
6.5 Specialized launchers¶
cafrun -np 2 simple_test_exec test=coarrays
mpirun -np N simple_test_exec test=<mpi-case>
simple_test_exec test=<gpu-case>
CMake and CI select the launcher and register the test only when the required capability is available.
7. Build and CTest design¶
- Gating.
BUILD_TESTS(SIMPLE's option, default ON in CMake, OFF in everycompile_*.shwithout--compile-tests) gatessimple_test_exec, the test-only library sources and every CTest registration. Nothing test- related is built otherwise. - The fast gate runs from the compile scripts. With
--compile-tests, eachcompile_*.shruns, betweenmakeandmake install(the X order: build, test, install; a failed gate is a failed build and nothing is installed, so the install banner is the last thing on a green build):
bash
ctest --test-dir build -L fast --output-on-failure --parallel "$NJOBS" --timeout 120 \
2>&1 | tee build/test_runs/ctest_fast.log
scripts/ctest_budget.py build/test_runs/ctest_fast.log --budget 30 --quiet
Until Phase 2 declares the fast gate, the label expression is
"fast|provisional" and ctest_budget.py runs with --no-budget, so the
provisional units entry runs and is timed on every build without a
budget it cannot yet meet. This is scripts/run_fast_gate.sh, called by
every compile_*.sh --compile-tests between build and install and by
make check; GATE_DECLARED inside it is the Phase 2 switch. The fast
suites run in-process from the build tree and need nothing from the
install tree. Before ctest, the script runs the registry-consistency
check (item 6).
NJOBS is the core count divided by two (tests are pinned to one thread
but do I/O). ctest_budget.py reads the ctest output, fails if the
run's real time exceeded 30 s or if any entry failed, and writes the
per-entry table sorted by time beside the log so the numbers are kept.
What the developer sees is ctest's own report, as in X (2026-09-22, by
decision): with --quiet the checker prints nothing on a green run
within budget and prints the table and the problems only when there is
something to fix. The check target runs the same thing by hand.
3. Registration by suite. One add_test per fast area suite
(simple_test_exec test=unit_core), per library suite
(test=lib_stream), per workflow gate, per platform case and per binary
smoke. Every registration sets LABELS (fast, library, workflow,
platform), TIMEOUT, WORKING_DIRECTORY under build/test_runs/<name>
(created at configure time) and an explicit OMP_NUM_THREADS (1 for the
fast tier). Workflow entries set RUN_SERIAL TRUE.
4. Process budget ratchet. SIMPLE_CTEST_BUDGET in
production/CMakeLists.txt must equal the registered count or
configuration fails, as in X. It is armed at the end of Phase 2, once the
area suites exist, at the count registered then; before that the count is
allowed to change (Phase 1 registers one units process, Phase 2 replaces
it with about seven). From then on it only goes down; raising it is an
owner decision recorded here.
5. No more standalone executables. The simple_test_*.f90 glob, its
per-program add_executable, install and add_test are removed at the
end of workstream B (Phase 7). Done 2026-09-24 by the utils review
(section 9.7): production/tests is gone and simple_test_exec is the
only test executable. Until then the glob keeps building the
programs that have not migrated, but none of them is registered with
CTest once Phase 1 lands (they are runnable by name and through
test_timing_run.sh); CI keeps calling the not-yet-migrated ones by
name until Phase 7 switches it to labels. CI runs by label since
2026-09-24 (section 9.7, the wrap-up).
6. Registry consistency. A configure-time or check-time script compares
the CTest registrations with simple_test_exec test=list and fails on a
registered selector that is not dispatchable, or on a selector the test UI
marks as registrable (area suites, library suites, workflow gates,
platform cases) that is not registered. Manual tools, focused sub-suite
selectors and platform cases whose capability is absent on this machine
are dispatchable without being registered, by design. A generated common
registry is a possible later improvement, not a prerequisite.
Done 2026-09-24 as a static check of the source tree rather than of
test=list, so it needs no binary and runs before the gate:
scripts/check_test_registry.py, called by run_fast_gate.sh before
ctest (exit status 1 fails the gate). It fails when a test= selector in
production/CMakeLists.txt (with the foreach lists expanded) is not a
program of the test UI; when a program of the test UI has no router case
in src/main/exec/simple_test_exec_*.f90, or more than one; when a
router case is not a program of the test UI; when a program is defined
twice; and when an area or library suite (unit_*, lib_*, except the
umbrella units) is not registered. Workflow gates and platform cases
follow no naming convention, so their registration is not enforced;
their selectors are still checked for being dispatchable.
7. Install. A --compile-tests install contains simple_test_exec and
no standalone test binaries.
NICE's Python tests remain under their Python runner. Python validators, shell compatibility checks and external oracle packages are not converted into Fortran modules. The uniformity goal applies to SIMPLE's Fortran tests.
8. Test inventory¶
Before moving code, create the migration inventory with one row per current test identity (both routes), with at least:
| Field | Purpose |
|---|---|
| Canonical test ID | Final test=<name> value |
| Current sources | Standalone program, commander procedure, tester module, helpers |
| Authoritative behaviour | Which implementation or merged behaviour is retained |
| Tier | fast, extensive, platform, or not registered (section 5) |
| Wall time / run state | From Phase 0, Debug and Release, single-threaded: measured <s> (passed or failed on its own), timed out, crashed, missing fixture (refused for lack of arguments or data), unsupported capability (coarray, MPI, GPU or a launcher this machine lacks), manual (persistent server, interactive). A runtime is required only for measured rows; every other state is itself the Phase 0 result for that row |
| Failure path | assertion / error stop / THROW_HARD / none |
| Performance action | none, shrink fixture, share fixture, drop I/O, split, move tier |
| Target grouped module and area commander | Owners after migration |
| Arguments | Existing CLI keys, defaults, parsing path |
| Launcher/capability | Serial, child process, coarray, MPI, GPU, network |
| Fixtures | Generated, committed, downloaded, user-supplied |
| Working files | Products, cleanup, retained evidence |
| Baseline result | Exit status, success marker, assertions, tolerances |
| Verdict | keep, modify, merge into, demote, delete, retire, investigate (section 9.4), with reviewer and date |
| Verdict note | The required note for the verdict |
| Migration status | Unreviewed, baselined, callable, routed, old executable removed, validated |
The inventory is also the review record (section 9): every row carries a verdict before its test is touched, and a retired tests table at the end of the inventory lists every identity removed, with date, reason and replacement.
The inventory lives at doc/code_overview/test_inventory.md (until 2026-09-25
doc/refactoring_notes/), one table
per area, generated by scripts/test_review_dossier.py. Since 2026-09-25 it is
generated on every build (target generate_test_inventory) and not
committed; the hand-written part, the verdict and note of each identity and
the retired-tests table, is doc/refactoring_notes/completed/test_review_record.md,
which the generator reads and which is updated in the same commit as the
verdicts it records.
scripts/test_review_dossier.py --coverage-after doc/refactoring_notes/completed/test_review_record.md
is the coverage accounting of section 9.5.
9. Test review¶
Unification and tiering move tests; they do not decide whether a test deserves to exist. That decision is a review, and it is where most of the overlap and dead weight will be found. The review is part of the plan, not a side activity: no test is migrated (workstream B), given assertions (workstream C) or gated (workstream D) before it has a recorded verdict.
9.1 What the review faces¶
Static inspection of the 150 identities (203 route implementations) gives the starting picture:
- Two route implementations: 53 identities. Two implementations claiming
the same test; the routes of
unitsare the known divergent pair. - Overlap by what they exercise: 14 clusters covering 30 identities.
Comparing the production calls each identity makes (both routes pooled),
17 pairs share at least half their footprint or one is a subset of the
other:
ori,oris,eigh,maxnloc,starfile,class_sampleandbounds_from_mask3Dagainst their_testtwins;ioinsideio_parallel;lbfgsbandlbfgsb_cosine;pca_allandpca_imgvar;cc_connectivityandimage_bin;gen_pickrefsandpreproc;reprojectandsimulate_particles; and thecontinuous_inplane_*gradient tests witheval_polarftcc. - Too thin to compare: 57 identities make fewer than three production calls. These are the print-only smokes and demos of section 4.4, and the prime candidates for deletion or merging.
-
No usable age signal. Every test source was touched in 2026 (the test tree was reorganised this year), so "last modified" says nothing about whether a test is still meaningful. Staleness has to be judged from what the test exercises and whether that code path still exists.
-
The question is not "fast or extensive". With the fast gate fixed as the
unitssub-suites, the review of the other 149 identities asks: which library suite does this belong to (it has, or can be given, a verifiable claim about a production path and runs unattended), is it a workflow gate, is it a manual tool worth keeping under its name, or does it go? The proposal column in the inventory is pre-filled on that basis.
The footprint comparison is a screen, not a verdict: it finds candidates, and a person decides.
9.2 Unit, reviewer, order¶
The unit of review is the test identity. Review proceeds by area, in the
order the areas are migrated, so verdicts are fresh when the migration
happens. The reviewer is the owner of the production subsystem the area
tests (by authorship today: Ruben for io, parallel, single, stream,
utils and the workflow tests; Cyril for masks; Hans for the rest), with
Hans as the arbiter for disagreements and for every deletion. A reviewer may
review their own tests; a second person reads the batch summary.
9.3 The dossier¶
scripts/test_review_dossier.py generates one dossier per test identity
from the tree and the Phase 0 timing runs (scripts/test_timing_run.sh), so
the reviewer does not have to assemble it:
- name, route(s), source file(s), line count, author history;
- production modules imported and type-bound procedures called (the footprint), and which of those no other test touches (unique coverage);
- failure path (assertion,
error stop,THROW_HARD, none) and what the test would report if it ran on wrong results; - fixtures (generated, committed, downloaded, user-supplied), arguments, launcher, working files;
- measured wall time, Debug and Release;
- callers: CI, scripts, documentation, other tests;
- overlap candidates: same-name twin, footprint cluster members, subset relations;
- a proposed tier (section 5) from the rules, to be confirmed or overruled.
The dossiers are regenerated for each batch and are inputs, not records; the record is the verdict.
9.4 Verdicts¶
Each test identity gets exactly one verdict, recorded in the inventory with reviewer and date:
| Verdict | Meaning | Required note |
|---|---|---|
keep |
Migrate as is; tier assigned | tier |
modify |
Migrate with a stated change: add an assertion, shrink or share a fixture, split, pin threads, remove a workflow run | what changes, and why the test is worth it |
merge into <id> |
Its unique coverage is folded into the named test; this identity is deleted | which checks move |
demote |
Not fast-tier material; goes to a named library suite (lib_<area>), the workflow gates, platform or manual |
destination, the assertion it must gain if it has none, and what it would take to promote it |
delete |
Removed without replacement | one of the reasons below |
retire |
Removed because the production path it tests is itself being removed | the production change |
investigate |
Parked: the verdict needs more than the dossier gives | owner, the open question, and the condition or date on which it is revisited; an investigate older than one migration phase is escalated to Hans |
A test is deleted when at least one of these holds and the note says which: it duplicates another test's coverage entirely; it exercises code that no longer exists or is deprecated for removal; it is a demo or benchmark with no verifiable claim and no plausible assertion; it cannot be made hermetic and has no extensive-tier value; or it has been broken with no caller for long enough that nobody noticed. A test is kept when it has unique coverage of a production path that matters, or when it is an isolation or launcher case (section 5.3). "It might be useful someday" is not a reason to keep; the history has it.
Deleted and retired identities go into a retired tests table in the inventory (name, date, reason, replacement if any), so the same test is not rewritten by accident and a reader of an old note or CI log can find out what became of it.
9.5 Coverage accounting¶
Deleting tests must not silently drop coverage. For each batch the dossier
script reports the union of production modules and type-bound procedures
exercised by the area's tests before and after the verdicts are applied, and
lists every production module that loses all test coverage. That list is
part of the batch summary; each entry is either accepted with a reason
("was only exercised by a print-only smoke") or answered by a modify or
merge. This is a call-footprint proxy, not line coverage; if it proves too
coarse, an instrumented (--coverage) build of the extensive tier is the
stronger tool and can be added later.
The review also reads in the other direction: a test that only exercises
dead production code has found dead production code. Whether that code goes
too is an owner decision, taken separately from the test's verdict; when the
answer is yes, test and code go in one commit so neither outlives the other.
First instance, 2026-09-22: subproject_distr and
ptcls_ppca_subproject_distr were the only callers of the subproject
scheduling framework in simple_qsys_ctrl/simple_qsys_env (added with
them on 2026-04-06); both tests and the framework were removed together.
The call-footprint proxy counts, for each unit_<area> identity, the
tester modules and local sub-suites that suites_<area>() registers, so a
merge into unit_<area> is credited to the tester that received the checks
rather than reported as lost coverage.
9.6 Mechanics and pace¶
The review of an area is one commit series: the verdicts (inventory rows
and the retired table) first, as reviewable text; then one commit per test
for the action taken (extract, assert, shrink, merge, delete). A verdict
takes about ten minutes with the dossier in hand; a test that needs longer is
marked investigate with a note and parked, not allowed to block the batch.
At that pace an area of fifteen tests is an afternoon, and the whole tree is
a few weeks of reviewer time spread across the migration.
Verdicts are revisited only through the same process: a keep that later
fails the budget goes back through review with its timing, not straight to
deletion.
9.7 Batch record¶
Geometry (2026-09-22, Hans). Nine identities on both routes: angres,
ori/ori_test, oris/oris_test, sym/sym_test, uniform_euler,
uniform_rot. The dossiers showed three print-only smokes with a handful of
THROW_HARDs, one print-only sweep, and two sampling demos. The question
"is what they touch exactly covered elsewhere?" was answered by comparing
their call footprint with the ori/oris tester modules procedure by
procedure: twelve ori methods (ori_from_rotmat, get_axis_angle,
reject, append_ori, delete_entry, get_keys, ori_strlen_trim,
ori2chash/chash2ori, ori2json, get_ctfvars, print_ori) and three
oris behaviours (reallocate, write/read round-trip, rnd_oris bounds)
were touched only by the old tests, so the verdicts are merge into
unit_ori, with asserting tests added to simple_ori_tester and
simple_oris_tester first (print_ori stays print-only). sym had no
assertions at all; it is replaced by a new simple_sym_tester module
(sub-suite symmetry of unit_ori) that pins the order and classification
of every group, the Euler limits, the subgroup tables, the group axioms of
the operator set (identity first, proper rotations, distinct, closed under
composition), apply consistency, rnd_euler limits, rot_to_asym,
symrandomize and build_refspiral; the print-only sym_tester routine
left simple_sym. angres is modify: the sweep lives in
simple_test_exec test=angres (tier lib_geometry) with assertions
against the recorded resolution ladder. uniform_euler and uniform_rot
are delete. Coverage accounting: the retired tests made 31 distinct
production calls and imported two production modules; every one is still
made by a remaining test, none lost. Two findings, both acted on by the
owner the same day: ori%get_axis_angle had no production caller and fed
Euler angles in degrees straight into cos/sin, so it was removed
rather than tested (second instance of a test finding dead production
code, section 9.5); ori_strlen_trim over-counted by one for a particle
with neither hash nor chash entries (it always added the pparms separator)
and now counts one separator between non-empty parts, which the tester
pins for every combination of parts. First build of the batch: 712 of 714
new assertions passed; the two failures were the spiral redundancy check
for d7 and i, one pair each. That pair is the jittered north pole and
its mirror mate: for d/o/i the mirror of the pole is symmetry-equivalent
to the pole, build_refspiral nudges it by at most 0.5 degrees so the
two are not identical, and the mate can land within a few thousandths of
a degree of it. The test now allows exactly that pair (mirror partners,
at most once) and still forbids every other near-coincidence; whether the
spiral should instead replace the degenerate mate with another
asymmetric-unit direction is an open owner question. Second build: gate
green, 7/7, 3.6 s real on the reference Mac in Debug (unit_ori 3.6 s
with the three ori/oris/sym sub-suites at 714 assertions).
fft (2026-09-22, Hans). Eight identities, six on both routes, none
with a failure path, none timed. corrs2weights_test (and its standalone
twin corrs2weights) and rank_weights printed or plotted weight curves
of production code that nothing else tested (corrs2weights drives the
motion-correction frame weights under every wcrit; the rank kernels are
its sum|cen|exp|inv modes): merge into unit_numerics through a new
simple_stat_tester (sub-suite statistics) that pins sums, signs,
monotonicity, closed-form spot values and the single-/all-zero edge cases.
ft_expanded ran test_ftexp_shsrch (already in the gate) and
test_ftexp_shsrch2 (never run): both are now sub-suites of
unit_numerics, moved out of unit_image because the expanded-Fourier
shift search is a motion-correction optimiser, not an image test.
order_corr (PASSED on an array size), phasecorr (a convention demo
with gnuplot windows), rotate_ref (a benchmark of two local copies of
what is now polarft_calc%rotate_ref_8) and eval_polarftcc
(user-supplied volume, timing printout) are delete; the oris tester
gained the corr-descending assertion order_corr never made.
gencorrs_fft is modify: its four unique calls were the image-to-polar
path, so it is now hermetic (three low-passed noise images from a fixed
seed) and asserts that gen_objfun_vals peaks at rotation 1 with
correlation 1 for an image against itself, at the applied step for a
copy rotated with rtsq (either angular convention; the sign is not what
is tested), and below 0.5 for an unrelated image. Coverage accounting:
33 calls, 2 not made by name any more (image%polarize and
set_ptcl_pft), both accepted: the new test reaches them through the
pftc's own polarize_ref_pft/polarize_ptcl_pft, which is the
production path. Findings: polarft_calc%rotate_ref_8 has no unit test
(the deleted benchmark validated a copy of it, not it); the dossier's
fixture detection missed defined('vol1'), so continuous_inplane_*
were listed as hermetic when they need a volume (fixed, the inventory
now says user-supplied). First run of the new gencorrs_fft failed two
checks through its own fixture: uniform noise (mean 0.5) under the soft
mask gave every image the same disc term, which correlates under any
rotation; zero-mean Gaussian noise, a Gaussian low-pass and a sixth-of-a-
turn probe fixed it (8/8, 0.03 s). The run also pinned the convention:
a real-space rtsq by +60 degrees peaks at polar index 5·nrots/6 + 1
(300 degrees), i.e. the pftc's rotation index runs opposite to rtsq's
angle, with cc 0.998. The IEEE_DIVIDE_BY_ZERO note seen in the first
run came from image%bp(0., lp), whose get_find(1, 0.) divides by
zero; production calls bp(0., lp) in imgops and resolest, harmless
because the flag is raised, not trapped.
masks (2026-09-22, Hans). Nine router identities plus three
standalone-only twins (bounds_from_mask3D, otsu, cc_connectivity);
two with a failure path. bounds_from_mask3D(_test), graphene_mask,
mask, image_bin and cc_connectivity are merge into unit_image
through a new simple_image_msk_tester with two sub-suites: masks
(bounds against a brute-force scan, the three-shells-per-band graphene
rule, disc/transfer2bimg/cos_edge with the edge pinned at 1, 0.5 and 0,
and the hard/soft/softavg mask semantics in 2D and 3D: 1 inside
mskrad-COSMSKHALFWIDTH, 0 beyond mskrad+COSMSKHALFWIDTH, monotone
cosine between, softavg filling with the outside average) and binary
image (the old image_bin examples with their answers, and the
26-connectivity contract that only the standalone cc_connectivity
enforced). otsu(_test) is merge into unit_numerics (statistics): a
two-Gaussian mixture, the threshold between the modes, class sizes, the
three overloads agreeing. msk_routines is modify: the single-thread
semantics live in masks; the exec case keeps what needs threads and
asserts parallel == serial for all six routines with the coordinates
memoised once outside the region (tier lib_masks, nthr=8 in CI).
nano_mask and score_volume_shape are demote to manual;
vol_shape_descr/calc_3D_shape_descriptors stay for the latter.
ptcl_center is delete (an RCSB download and a centering experiment);
its gap is recorded: it was the only test naming masscen, roavg,
window_center, shift2Dserial, power_spectrum, fproject and
get_nyq, image-area basics for the image review. Coverage accounting:
46 calls, none lost. Findings: calc_graphene_mask excludes the three
shells nearest each band unconditionally, so at a pixel size where a
band lies beyond Nyquist it silently drops the highest shells instead
(the test uses 0.358 A, where both bands are inside); otsu on a
constant sample divides by zero in its range scaling (not tested for that
reason). First build: 46 of 53 masks checks passed; the seven failures
were all the code, not the test. (1) image%disc (the npix form)
applied its threshold to the whole rmat including the two Fourier
padding columns, which cendist leaves with a partial distance, so the
padding was set to 1 and npix over-counted by two discs' worth (18671
against the 17077 voxels a 48-box sphere of radius 16 actually has);
fixed to the logical dimensions, as the lmsk form already did (its one
production caller, opt_filter, does not read npix). (2) The memoised
mask routines (mask2D_soft/softavg/hard, mask3D_*) compute the edge
weight at pixel i and apply it to the mirror pixel n+1-i as well, but
the memoised coordinate of pixel i is -n/2 + (i-1) (origin at pixel
n/2+1, the convention of cendist and of the per-pixel routines), so
the mirror of coordinate -(r+1) is applied to coordinate +r: every
mask is one pixel tighter on the positive side of each axis than on the
negative side (a hard mask of radius R keeps -R..R-1; the soft mask
reads 0.368 instead of 0.5 on the radius at +x, 0.5 at -x). The
fix is to mirror about the origin pixel (ir = n+2-i, with the
-n/2 row having no partner and the origin row applied once). This
changes results by one pixel on the positive side in every mask
consumer (40 files); Hans decided to fix it the same day. The four
mirrored routines now mirror about the origin pixel (softavg loops
over every pixel and was never affected), the loop structure was
checked against a direct per-pixel evaluation for boxes 6 to 64 (every
pixel touched exactly once, identical result), and the masks sub-suite
gained the assertion that would have caught it: each mask reads the same
at +r and -r along every axis and the same along x, y (and z). msk_routines passed its first run (7/7, 0.23 s with 24 threads); it
reads the thread count from the OpenMP environment, not from nthr=, so
CI sets OMP_NUM_THREADS=8 and CTest will pass it the same way. Second build of the batch: gate green, 7/7, 3.5 s real (unit_image
2.1 s with the two new sub-suites, unit_numerics 1.3 s).
segmentation (2026-09-22, Hans). A category rather than an area:
Otsu is a thresholding method, not a statistic, so test_otsu moved out
of statistics into a new simple_segmentation_tester (sub-suite
segmentation of unit_image, beside binary image, which keeps the
connected-component contract). peak_thres_fdr, an assertion-bearing
exec case for detect_peak_thres_fdr that its router had filed under
utils, is merge into unit_image there and its exec case is gone. What
the category is for: simple_segmentation and image_bin hold about
twenty production routines with callers and no test (otsu_img 8
callers, binarize and masscen_cc 5, erode, grow_bins,
diameter_cc, cc2bin 4, canny, sobel, sauvola 3,
detect_peak_thres_sortmeans, otsu_robust_fast, elim_ccs,
order_ccs, set_edgecc2background, feret_minmax 1 to 2); they are
to be pinned here on generated fixtures with known answers (two-level
images for the thresholds, a disc whose edge is a one-pixel ring for the
edge detectors, erode/grow round trips, a placed blob for masscen_cc
and diameter_cc) as a scheduled slot of its own. Owner list for the
section 9.5 decision, routines with no caller at all: hough_line,
polish_ccs, diameter_bin, elim_largestcc, detect_peak_thres_sortmeans
(border_mask, listed at first, is what erode uses). Built and green the same day: gate 7/7, 3.5 s, unit_image 2.2 s with
eight sub-suites.
The scheduled slot followed the same day (Hans: "pin the seg routines").
segmentation now pins, on generated fixtures with known answers:
binarize in all three forms (a ramp image, one pixel per value);
otsu_img plain, positive, tight and tighter (two-, three- and
four-level discs with 1 % noise, the binarised image equal to the object
to the pixel); otsu_robust_fast (salt-and-pepper on a disc, every
flipped pixel away from the edge repaired); sauvola (local standard
deviations equal to brute-force window statistics, the binarisation
following the Sauvola formula pixel by pixel); calc_gradient (a unit
ramp has gradient exactly 1 inside) and sobel (a ring along a square's
edge, nothing elsewhere); canny with explicit thresholds (a thin edge
around the square, nothing elsewhere, the input untouched);
detect_peak_thres in both forms, detect_peak_thres_for_npeaks and
refine_peak_thres_sortmeans (200 background scores and 20 peaks). The
binary image sub-suite pins the image_bin morphology and
bookkeeping: erode/dilate (a 10x10 square loses and regains its
outer layer exactly), grow_bins (cross template: no corners; the
13-pixel digital disc of radius 2), size_ccs, masscen_cc,
diameter_cc, cc2bin, elim_ccs, order_ccs on three placed blobs,
set_edgecc2background (a square ring is filled) and feret_minmax
(a 5x21 bar: 5 and the 21.4 diagonal). Three defects found while
deriving the expected answers, all fixed: (1) otsu returned the centre
of the last background bin instead of its upper edge, so the upper half
of that bin (up to 1/512 of the range) was classified as foreground; a
few background pixels per image, in a third of the emulated runs
(thresh = T + 0.5 in bin units); (2) otsu never assigned thresh
for a two-valued input, because the first bin is already the optimal
split and only strictly better splits assign (initialised to the first
bin); (3) image%binarize(npix) kept npix+2 pixels (forsort(n-npix-1)
with a >= comparison; now n-npix+1). Its one caller is the
binarize commander's npix option. detect_peak_thres_sortmeans is
referenced only from a comment and prints debug lines; it joins the
dead-code list for the owner. Built green first time: gate 7/7, 3.4 s, unit_image 2.1 s.
io (2026-09-23, Hans). Nine router identities plus the four binoris
identities from the utils and unassigned tables (two of them empty exec
stubs). imgfile is delete: SPIDER/MRC squares and cubes converted both
ways and compared by correlation is a strict subset of test_image part
20, already in the gate. io and io_parallel are delete: 40 GB
throughput benchmarks with no assertion. star_export is delete: it
timed two writers that STAR file asserts. mrc2jpeg and mrc_validate
are demote to manual (a filetab-to-JPEG converter and a read/write-back
of a user volume). stack_io is modify: the exec case copied a committed
stack and asserted nothing, the standalone was the real test; its hermetic
part now lives in a new simple_stack_io_tester (sub-suite stack I/O of
unit_core) and gained what it lacked: open/close state, same_stk,
buffer sizes of 2, 3, the whole stack and more than the stack (a partial
last window in each case), forward skipping reads through the refill loop,
get_image from the current buffer, the float32 header beside the
float16 one, and a position-dependent pattern in the 1025-box float16
stack so a pixel displaced across the converter's 1 M-element buffer
flush would be seen (the old test wrote a constant). The benchmarks are
gone. inside_write, binoris and binoris_io are merge into
unit_project through a new simple_binoris_tester (sub-suite binoris):
header-only files; a hash-backed segment and a fixed-width particle
segment round-tripped with every header field checked against the file
size; partial and sub-range particle reads landing at absolute indices
(what merge_algndocs relies on); a 40-value particle record written by
hand under a narrower header, read back with zeros in the twelve newer
slots (the legacy-project path of read_particle_record, which no test
had touched); write_segment_inside growing and then shrinking the
middle of three segments with the neighbours byte-identical, in both its
oris and string-array forms; the sp_project front door rewriting stk
in place (the old inside_write case, now asserted) and falling back to
a full write when the file is missing; and the four binoris_io
dispatchers on .txt and .simple, including the ctf/state/eo merge that
keeps the keys the file does not carry. starfile_test/starfile are
modify: the wrapper demo became assertions in simple_starfile_tester
(table names, comment, string, doubles pinned to the %12.6f/%12.6e
formatting the C++ writer uses, absent labels, first/next iteration), and
run_all_starproject_tests, a 23-test suite that only this exec case ran,
is registered as sub-suite STAR project of unit_project. Two things in
it were incompatible with a shared process and were removed: it called
report_summary and error stop on the process-wide failure counter
(so a failure in any earlier sub-suite would have aborted the run), and
it set the OpenMP thread count to 4 for good; it now restores the count
it found. Its tier is provisional: the per-entry timing table decides
whether the 20 000-row export stays in the gate or moves to lib_project
in Phase 4. Coverage accounting: 44 calls, 4 not made by name any more
(image%corr, made by test_image part 16; image%ran, a random fill;
rslices/wmrcslices, the imgfile layer under stack_io%read/write),
all accepted; simple_imgfile is no longer imported by a test directly.
The dossier script now counts a bare call obj%meth (no argument list) as
a type-bound call; it had missed write_header and update_byte_ranges.
First build: unit_core (with stack I/O) green in 1.1 s, STAR file
67/67, STAR project 67/67 in 0.41 s (so it stays in the gate), and
unit_project crashed with SIGSEGV in the binoris tester's fifth test. The
file left behind showed the production path was right (the grown segment
and the moved particle segment byte-correct); the fault was the tester's:
a helper's optional dummy named nmics hid the module constant NMICS
(Fortran is case-insensitive), so the helper read the absent optional.
Renamed, with the same trap removed from verify_stack in the stack_io
tester (bufsz), and a comment at each. Second build: gate green, 7/7,
3.5 s real (unit_project 2.0 s with the three new sub-suites,
unit_core 1.1 s).
Findings, not acted on: stack_io%read loops for ever on a backward read
(the refill loop only advances), where a THROW_HARD would name the
contract; binoris%open on a file that does not exist yet leaves
fname unset, so the error messages of a first write name an empty file;
discrete_stack_io (standalone only, assertion-bearing, unassigned) tests
dstack_io and the float16 encoder boundaries and is the natural next
addition to stack I/O.
numerics (2026-09-23, Hans). Five router identities plus the
standalone twins eigh and maxnloc; the theme is not what was deleted
but what had no tester: simple_linalg, simple_kbinterpol,
simple_srch_sort_loc and the symmetric neighbour searches. eigh_test
(prints, then eigh of a random 15000x15000 matrix thrown away) is
merge into unit_numerics through a new simple_linalg_tester
(sub-suite linear algebra): the LAPACK example matrix with numpy's
eigenvalues and inverse as the reference, eigh largest and smallest
with orthonormal eigenvectors and the residual, sparse_eigh against
eigh (the standalone's one check), svdcmp reconstruction and
singular values, matinv plus the singular flag, jacobi/eigsrt,
svdfit/svd_multifit on exact and noisy polynomials (chi-squared
pinned to numpy's least squares), fit_straight_line (its corr is r
squared, now said so), the plane fits, the vector helpers and gemm_tn;
test_eigh left simple_linalg. kbinterpol_fast is modify into a
new simple_kbinterpol_tester (Kaiser-Bessel kernel): the printed
outer-product comparison became assertions at fixed sub-pixel positions,
and the tester adds what nothing checked: apod against the closed form
(I0 series in double precision), the fast polynomial's coefficients as
(beta^2/4)^k/(k!)^2, apod_fast_value_deriv against central
differences, the three device forms bit for bit, apod_mat_3d_fast_grad
against finite differences of apod_mat_3d_fast with the switch
margin, and instr; box 16, no timing loops. maxnloc_test is merge
into a new simple_srch_sort_loc_tester (search, sort, locate) that
pins every routine of the module against brute force. neigh is merge
into unit_ori: symmetry now checks find_closest_proj,
nearest_proj_neighbors in both forms, sym_dists and find_angres on
a 200-direction spiral for c1, c2 and d2 against a brute-force scan over
the symmetry-expanded distances, and the c1 forms against the oris
forms. trail_rec_blend is modify: moved as it is into
simple_accum_blend_tester (trailing-reconstruction blend of
unit_image). With nothing left in the category, the numerics
commander module, router and UI module are gone (three files, three
call sites); CI lost simple_test_neigh. Coverage accounting: 28 calls,
5 not made by name any more (ran_tabu%shuffle, progress_gfortran,
rotmat2D, all scaffolding; test_eigh, removed; calc_stats, now
pinned in statistics), none a loss. Two production defects found
while deriving the expected answers, both fixed: (1) selec (the
Numerical Recipes selection behind median, median_nocopy, the
nu-filter evidence thresholds and the image edge median) tested
ir-1 == 1 where ir-l == 1 was meant, so a two-element final
partition away from the array start was left unsorted; emulated in
Python, the median of a random array was wrong in 14 % of cases for
n >= 10 (an adjacent order statistic), and the selec-for-every-k
check and a median test on two arrays the typo gets wrong now pin it;
(2) reverse on an even-length double-precision array kept element 1
in place and reversed the rest (the body of reverse_f, the
Fourier-origin-preserving variant, pasted into reverse_drarr); no
caller passes double arrays today, fixed and pinned. Also seen and
fixed the same day (Hans: "fix now"): oris%nearest_proj_neighbors
(count form) recomputed and sorted the distance table n times over an
outer loop whose index was unused, O(n^2 log n) for an O(n log n) job;
the loop is gone, the result is the same and symmetry pins it. Dead
code, removed the same day under the section 9.5 rule as Hans stated it
("if they are not used they go"): nine public simple_linalg routines
with no caller anywhere (hermitian_eigh, hermitian_invert,
hermitian_solve, svd_solve, normal_solve, svdvar, outerprod,
l1dist, same_energy_euclid, 295 lines, with the LAPACK interface
declarations only they used: zheev, zposv, dposv, dgelss,
sgelsy), and the five from the segmentation batch's owner list
(hough_line and detect_peak_thres_sortmeans in simple_segmentation,
polish_ccs, diameter_bin and elim_largestcc in image_bin, 310
lines; the one commented-out call in simple_pickref went with them).
The section 9.5 owner list is empty. First build: unit_image,
unit_core, unit_project green with the new sub-suites; Kaiser-Bessel
kernel 93/93, search, sort, locate 75/75, statistics 74/74;
linear algebra 117/120 and symmetry 427/428. Of the four failures,
one was the tests and three were real: (1) norm_2([3,4]) returned 0.
norm_2_sp and vabs_sp called BLAS snrm2; Apple's Accelerate
returns single-precision function results (snrm2, sdot, sasum) in
the f2c/g77 convention, as a double, so a gfortran caller reading a
float gets 0. On macOS, since the switch to external BLAS on 2026-06-10,
norm_2 (the gradient-norm convergence tests of simple_opt_helpers,
the BFGS2 and steepest-descent optimisers, two nanoparticle radius
checks) and vabs (the chi-squared of svdfit/svd_multifit, which
the noisy-fit test also caught as exactly 0) had returned 0; Linux with
OpenBLAS was unaffected. Both now accumulate in double precision without
BLAS; dnrm2 stays for the double versions, which the convention does
not touch. (2) jacobi is an ssyev wrapper that reports nrot = 0;
the test now pins that instead of expecting rotations. (3) euldist of a
direction with its own copy is acos(1 - eps), about 5e-4 rad in single
precision; the c1 representative check allows that. The two
near-coincident spiral directions printed for symmetry are the known
pole/mirror pair (section 9.7, geometry). Ninth production defect.
Third build: gate green, 7/7, 4.9 s real (unit_ori 4.9 s with the
neighbour searches on three 200-direction spirals plus a 400-direction
one per group, unit_numerics 1.3 s with the three new sub-suites).
stats (2026-09-23, Hans). Nine router identities plus the unassigned
standalone twins class_sample and ctf; nothing in the category
asserted except the two standalones (ctf, the 12-check phase-shift
policy test, and sp_project, of which the exec case was an older
subset), and nothing in the gate covered the CTF, the PCA classes, the
decay schedules or the class-sampling file. clustering (one call to
test_aff_prop, already the affinity propagation sub-suite) and
multinomal_test (prints; multinomial random draw is in the gate) are
retire. eo_diff is retire: it needed refine3D half-volumes in cwd
and asserted nothing, and ran_phases_below_noise_power, its only
production call, had no production caller (removed). class_sample_test
and its twin are merge into unit_core through
simple_class_sample_io_tester (class sample I/O): the ragged
round trip field by field, an empty class (unallocated pinds, as
get_class_sample_stats leaves it) coming back with pop 0 and
zero-sized arrays, replacement of a previously allocated array.
ctf_test and ctf are merge into unit_image through
simple_ctf_tester (CTF): the policy checks moved as they are, now
through eval_canonical and canonical_phshift, plus the 300/200 kV
wavelength, the ctfvars unit conversions, apply_convention, the
closed-form CTF at three frequencies (chi of 10, 50 and 160 radians,
tolerances accordingly) and along/across the astigmatism axis,
nextrema at three frequencies and at the first three zeros (computed
from the quadratic in double precision), ctf2img against the closed
form, ft2img placement in three modes and the gen_fplane4rec
restoration contracts. extr_frac is merge into unit_numerics through
simple_decay_funs_tester (decay schedules): every public schedule
pinned on endpoints, quarter points, monotonicity and the mirror
symmetry of cos_decay/inv_cos_decay, calc_nsampl_fromto on both
branches, the update fractions, extremal_decay and extremal_decay2D
with their clamps. pca_all and pca_imgvar are merge into
unit_numerics through simple_pca_tester (PCA), on Hans's
instruction that the PCA suite gets real unit tests: pca_svd on both
branches (D >= N and the transposed D < N) against numpy's SVD;
ppca on a rank-2-plus-noise 5x16 fixture against the Tipping-Bishop
maximum-likelihood solution (EM converges to it from a random start,
the tolerances follow the slow direction at the built-in stopping
thresholds), reconstruct_external, calc_bic and suggest_rank;
kpca_svd on a two-cluster 4x12 fixture against a double-precision
emulation of the same pipeline (kpca_ref.py, kpca_nys.py): kernel
eigenvalues, features and pre-images for exact/cosine, exact/RBF,
Nystroem with every point a landmark for both kernels, a 6-landmark
Nystroem run with local support (ordered spectrum, pre-images inside the
data's bounding box and on their own cluster's side) and
suggest_kpca_nystrom_neigs. sp_project is merge into unit_project
through simple_sp_project_tester (project records): the phase-shift
checks moved as they are, the write/read round trip at 7 mics / 300
particles, the read probe from another cwd, the three-document merge,
and, on Hans's rule that print_segment_json (the GUI's segment view,
called by print_project_field) can only go if unused, that routine
parsed back from a diverted logfhandle with json-fortran: the data
window and indices_pre/indices_post, ascending and descending sorts,
the histogram and plot blocks, a particle segment. With nothing left in
the category, the stats commander module, router and UI module are gone
(three files, three call sites); CI lost simple_test_ctf,
simple_test_extr_frac and simple_test_multinomal; the phase-shift
policy document (sections 9 and 10.1) and the staged-refactor note now
name the sub-suites. Coverage accounting: 41 calls, 6 not made by name
any more: three removed routines, test_aff_prop (in the gate as a
sub-suite; the dossier script does not see a bare procedure argument),
and get_res/subtr (accepted, trivial image arithmetic). Tenth
production defect, fixed: print_segment_json sized indices_post
from the caller's window (fromto(2)) instead of the remapped one
(ffromto(2)) when sorting in descending order, so the array was too
long for the section assigned to it; the descending-window test pins
the sizes. Simplified, behaviour preserved: calc_update_frac clamped
to half the particles and to the minimum before overriding both with
the maximum; it now computes the maximum over the particle count
directly. Dead code removed under the section 9.5 rule (nine routines):
ran_phases_below_noise_power, nsampl_decay, write_segment2txt
(85 lines, only the old test called it), print_class_sample,
class_samples_same (compared only the integer part of a record) and
the private unserialize_class_sample, ctf%eval in both forms (the
six-argument one said its angle was in radians and passed it to init,
which converts degrees) and eval_sign; eval_canonical is the CTF
evaluator (7 production callers). spafreqsqatnthzero was removed too
and the first build failed on simple_ctf_estimate_fit, which calls it
as SpaFreqSqAtNthZero: the caller search had been case-sensitive.
Restored, with the CTF fit's use (the fitting ranges between the first
zeros) pinned in CTF; caller searches are grep -i from now on.
Second build: everything green except five assertions, all the tests':
three in CTF because the apply_convention locals were named dfx,
dfy, angast and hid the module constants of the same names (the
trap of the io batch, a third time; renamed), two in project records
that expected add_single_movie to store a defocus (a movie record
carries optics only until CTF estimation; the test now says so).
PCA and decay schedules passed first time, unit_numerics 1.35 s.
Third build: green except suggest_rank, which returned 3 for the
rank-2 fixture where the second build had returned 2. Its BIC is
residual-based, D N log(rss/(D N)) + (D Q + 1) log(D N), so an extra
component lowers it whenever the eigenvalue it explains outweighs the
D log(D N) penalty, which the largest remaining eigenvalue always does
here (rank 3 has rss 1.6 against 3.3); which rank wins after the
ten-iteration cap depends on the random start, and the test had pinned
luck. It now pins what the routine guarantees (rank 1 loses by ~200,
ranks 2 and 3 are close, sigma^2 falls with the rank, duplicates are
skipped) and the finding goes to Hans: the auto-neigs of ppca classes
(cluster2D, PPCA_AUTO_CAND up to 16, 15 iterations) is decided by the
iteration cap and BIC_TOL rather than by the data; the PPCA marginal
likelihood (Tipping & Bishop, sigma^2 the mean of the discarded
eigenvalues) would stop at the rank where the spectrum flattens.
Hans: "make the ppca change". ppca%calc_bic is now -2 ln L + p ln N
with ln L = -(N/2)[D ln 2pi + sum_k ln lambda_k + (D-Q) ln sigma^2 + D]
with ln|C| from the fitted retained eigenvalues and sigma^2 and
tr(C^-1 S) evaluated at the fitted W (the stationary-point form, which
drops that trace as D, was tried first and let a ten-iteration rank-3
fit score above its own optimum and win; the exact likelihood cannot),
p = D Q - Q(Q-1)/2 + 1. On the fixture the converged values are 249.8,
171.5, 176.7, 180.5 for ranks 1 to 4. The scan's hard cap of ten EM
iterations then failed the Linux build: a ten-iteration rank-2 fit
scores 174-176 against a rank-3 fit that may reach 177, inside
BIC_TOL, and which one wins depends on the random start. The cap is
gone: suggest_rank honours the caller's maxpcaits (cluster2D
passes 15) and EM stops on its own tolerances before that; the tester
scans with 500 and pins the converged BICs and sigma^2.
The tester pins the rank-2 BIC and the scan again; cluster2D's
auto-neigs for ppca classes changes behaviour accordingly.
Thirteenth. Seen on the build in between: suggest_rank skipped a
repeated candidate only when it equalled the previous slot, so a
third copy (or [1,1,1], or candidates clamped to the same rank) was
fitted again; it now compares with the previous fitted rank.
Four more findings, decided by Hans the same day (1 and 3) or left to
the reviewer (2 and 4), all acted on: (1) the Nystroem kPCA backend
returned unit-norm eigenvectors as features and weighted its projected
kernel column by the eigenvalues, where the exact backend (and
Schoelkopf's projection) return sqrt(lambda_k) v_k and weight the
column by v_k(i) v_k(j); with every point a landmark the two backends
agreed on the spectrum but not on the features or the RBF pre-images
(Nystroem's collapsed each cluster onto its centroid on the fixture).
Fixed ("fix"): master_nystrom now stores sqrt(lambda_k) v_k as the
features and normalises the eigenvectors as the exact backend does for
the projected column; with every point a landmark the two backends now
agree to 1e-15 in the emulation, and the tester pins Nystroem against
the exact constants. cls_split with pca_mode=kpca (non-default)
sees differently scaled embedding coordinates from now on. Eleventh
production defect. (2) The exact cosine pre-image iteration did not
converge when the projected kernel column mixed signs within the
point's own cluster (the L1-normalised update flipped direction); on
the fixture one cluster ran the full 500 iterations and landed in the
other cluster. Fixed (reviewer's call): the weights are now
max(0, projected column) x max(0, cosine), a convex combination as the
RBF rule and the Nystroem cosine rule already were; every point
converges in three iterations and moves towards its cluster centre
(clipping the product alone, tried first, let anti-aligned points of
the other cluster in with positive weight). Twelfth. (3)
master_nystrom had PROFILE = .true. as a parameter and printed
twenty-odd timing lines on every call; off ("turn off profile"). The
per-percent pre-image progress lines are not under PROFILE and remain
(the tester diverts logfhandle around the Nystroem calls). (4)
print_segment_json for ptcl2D printed every record with
os_ptcl2D%print(iori) as it went (the loop index, not even the
selected record) - a debug leftover, removed; its histogram and plot
blocks dereferenced the optional sort_key whenever hist or
plot_key was passed, which the one caller always does together -
now guarded, and the tester asks for both without a sort key. Left as
is: in a descending window indices_pre/indices_post refer to the
ascending order (swapped relative to the displayed order); the GUI's
reading of them is not known here.
optimize (2026-09-23, Hans). Five router identities plus the
unassigned standalone lpstages. lbfgsb and lbfgsb_cosine (twins
on both routes, print-only) are merge into unit_numerics through
simple_opt_tester (optimisers): nothing in the gate had touched
the optimiser framework production runs on (lbfgsb at five sites,
de in CTF estimation, simplex in the volume symmetry search). The
tester goes through opt_factory/opt_spec as production does:
bookkeeping of the specification, L-BFGS-B on the 1D quadratic free
and against an active bound (converged flag, one gradient per cost
evaluation), on the direction problem (as 1 - cos, since acos has an
infinite derivative at the optimum the old test aimed at) and on
Rosenbrock, DE and the restarted simplex on a 2D quadratic from preset
and random starts. Hans's (a): the bfgs2, bforce and stde
optimisers, with no production caller, are gone, and with them
simple_opt_helpers (only they used it), the line searches, the
hill-climbing selector and the limit corrector of simple_opt_subs
(now amoeba alone, 465 to 135 lines), their factory cases, and the
opt parameter whose bfgs default nothing read: 1195 lines.
lplims (prints, was in CI), lpstages_test (prints) and the
unassigned lpstages (five THROW_HARDs) are merge into unit_numerics
through simple_lpstages_tester (low-pass stages): the three clamp
regimes of mskdiam2lplimits, lpstages with one stage, from a
falling FRC (the crossings, thresholds, crop boxes through the magic
boxes, cropped sampling and shift limits, pinned on a double-precision
emulation), the particle and class-average threshold floors on an FRC
that separates them, the linear fallback of a flat FRC (what the old
exec case ran without noticing), lpstages_fast with its floor and
force_lpstart, lpstages_setlims including the no-crop case, and
the Butterworth kernel in all four forms against 1/sqrt(1+(s/fc)^16);
lpstages_fast and lpstages_setlims gained the verbose optional
lpstages already had (default unchanged). opt_lp is retire: a
manual experiment that downloads 1JYX from RCSB, reprojects it and
prints Butterworth-band residuals; create_hist_vector, which it
called, stays because otsu uses it. With nothing left in the
category the optimize commander module, router and UI module are gone
(three files, three call sites); CI lost simple_test_lplims.
Coverage accounting: 20 calls, 3 not made by name any more
(apply_filter, exercised by the image self-test and inside
butterworth_filter; avg_sdev; create_hist_vector, inside otsu),
none a loss. Nothing found wrong in what was reviewed. First build:
green except two assertions of the tester's own making (DE reaches
the quadratic minimum to about 1e-2 in cost at its population
tolerance, not 1e-3; the notch value at s = 40 is 0.0996, the
assertion had said below 0.01).
The ipc gate and its sleeps (2026-09-23, Hans: "we cannot spend time
sleeping in unit tests that are part of the build process").
unit_ipc took 24 s on Linux and, on a later run, 16 s on the Mac
(1.4 s before): the four ipc sub-suites do not sleep, the production
code they call does. persistent_worker_server%kill slept a fixed
2 s "to allow workers to receive TERMINATE" on every call, and the
server tester kills eight started servers: 16 s. The TCP client's
send_recv_msg paused 1 s between its five retries, and the tester's
no-listener test pays four of them: 4 s. The socket server's
start_listener polled the ready flag every 100 ms. Fixed in
production, not in the tests: kill() now waits until the listener has
sent TERMINATE to every worker that was registered when kill() was
called (a mutex-protected counter the listener increments when it
replies TERMINATE), polling every 10 ms up to the old 2 s cap, so a
server with no workers stops at once and a server with workers stops
as soon as they have been told; the client's retry pause is a
component with a setter (set_retry_backoff_ms, default unchanged at
1 s) that the failure-path test sets to 0; the listener start polls
every 1 ms under the same 2 s cap. The server tester asserts that
kill() with no workers returns within half a second.
class and highlevel (2026-09-23, Hans). Two areas in one batch, because
class has one live item. Class: ui_hash_test called test_ui_hash, a
print-only PASS/FAIL routine embedded in the production module
simple_ui_hash; it is merge into unit_ui as the UI hash sub-suite
(simple_ui_hash_tester: set by character key and get by string key with
pointer identity and reference semantics, absent key and wrong dynamic type
as typed misses with a null pointer, overwrite retargeting the pointer to the
new object, key trimming, found optional). Reviewing the module showed that
production uses exactly two of its eight accessors (add_ui_program sets by
character key, simple_ui gets by string key): the set_ui_param/
get_ui_param family and the other two overloads had no caller and are
deleted with the embedded test (Hans: "(a)"; six routines, the generic
interfaces collapsed to two type-bound procedures). forked_process
(platform, section 4.6), the seven unit_<area> suites (fast) and units
(unregistered umbrella) are keep and recorded so the area is closed.
Highlevel: mini_stream is demote to manual (user movies, gain reference,
cluster); its standalone route was byte-identical and is deleted, and the
commander route lost a copy-paste relic (command_argument_count() /
parse_oldschool on a cline that arrives parsed) and a delete('nran') on
an empty cline. pcg_frac_update and rec3D_backends are demote to
manual: both need a refine3D project and are the equivalence and
backend-comparison gates of doc/policies/3D/reconstruct3D_pcg_policy.md;
the dossier's "delete candidate, no failure path" was an artefact, since
their THROW_HARDs live in validate_rec3D_pcg_fractional_updates (production)
and in run_rec3D_backends_single/gate_fail (the commander module's
helper), outside the exec_test_* body the script inspects. pcg_recon is
keep, workflow (registered, in CI): box 24, fixed seed, fourteen gated
stages; whether it belongs in the fast gate instead depends on its one-thread
time, which is to be measured (a move would mean rewriting its
THROW_HARD/all_ok checks on simple_test_utils). reproject is merge
into simulate_particles: one embedded 6VXX volume (centred), reproject with
nspace=100 and simulate_particles with nptcls=200 and CTF, each checked for
stack presence, image count, square box, smpd and one orientation record
per image; nthr comes from the command line instead of the hard-coded 16.
It is registered under workflow (simulate_particles, nthr=8) as the only
nightly run of either commander (simulated_workflow uses simulate_movie),
so SIMPLE_CTEST_BUDGET goes 19 -> 20 with that reason. simulated_workflow
is keep (registered twice). Coverage accounting: the retired routes made
no production call that a remaining test does not make (test_ui_hash is
deleted with its module's dead accessors; reproject's three calls are a
subset of simulate_particles'; the mini_stream commander route stays), so
nothing is lost. Noted, not done: the simulation checks are still
bookkeeping (counts, box, smpd); a simulation-truth floor (a zero-Euler
reprojection against the volume's z-sum, say) is Phase 5 material.
reconstruction (2026-09-23, Hans). The first family of the unassigned
standalone-only area, and the batch that creates the reconstruction area
in both tiers (Hans: "B"). continuous_3D_pcg_reconstruction (twelve
files, 1 645 lines, 2026-08-27, never in CI) was a driver re-executing
itself as a child process per case, with three cases: a self-test of its
own phantom builder, the gauran/add_gauran noise contracts, and a
half-set study (24 and 48 views per half, iteration trajectories with and
without support, a thirteen-value lambda sweep at forty iterations, a
gridding control, FSC between halves) that wrote twelve MRC volumes per
run. Verdict modify: the noise contracts become the observation noise
sub-suite of the new fast area suite unit_reconstruction
(simple_gauran_tester: N(mean, sdev^2) moments, the SNR definition
var_noise = var_signal/snr realised within 6 %, zero-mean and
signal-uncorrelated noise, replay after reseeding, independence of
consecutive draws; statistical tolerances, since the stream behind
random_number is compiler-specific), and the half-set study becomes the
PCG half-set sub-suite of the new nightly library suite
lib_reconstruction (simple_pcg_halfset_tester, next to
simple_reconstructor_pcg): in-process, no volume output, the same
observations (simulate_particles projection path, seeded noise) and the
same solves, asserting half ownership, realised SNR and noise
independence, exact iteration counts, reproduction of a solve by a fresh
operator (relative L2 < 1e-5 rather than bit identity, since the operator
reduces under OpenMP), finite gridding and PCG half maps that differ, the
FSC contract (bounded, low-shell mean > 0.4, decay of at least 0.1 towards
Nyquist), noiseless recovery of the supported truth (corr > 0.85; the first
nightly run, 2026-09-28, found the 24-view solve under-determined, 0.82 at
40 iterations against 0.98 at 48 views, so the FSC test runs 96 views per
half and the bar is 0.97; the lambda sweep is cut to six values around the
optimum, 0.1 to 100), and on
the 48-view matrix a noisy raw-L2 lambda optimum interior to the sweep
that beats the gridding control. The phantom fingerprints stay as one
small test of the fixture. rec3D_backend (in CI) is merge into
unit_reconstruction as rec3D backend (simple_rec3D_strategy_tester):
the defaults (rec_backend=gridding, maxits_pcg=2, rtol<=0), name
resolution (case-sensitive, blanks trimmed, empty invalid), wiring, and
the factory's dynamic type on all six branches (the old test pinned two).
flex_pcg is deferred to the heterogeneity family (Hans). Registration:
unit_reconstruction is the eighth fast entry and lib_reconstruction
the first library entry (3600 s, 8 threads, RUN_SERIAL), the template
for the other Phase 4 library suites; SIMPLE_CTEST_BUDGET 20 -> 22. The
CI line simple_test_rec3D_backend is gone. Coverage accounting: the
retired programs' production calls are all made by the two new testers
or by pcg_recon/the image and ori testers; nothing is lost by name.
HTTP POST left the network (2026-09-23, Hans). The eight-suite gate
passed on the Mac but took 58.9 s: unit_ipc 58.9 s, of it HTTP POST
58 s. The four tests posted to https://jsonplaceholder.typicode.com and
pinned FNV hashes of that site's responses — a live external service in
the build gate, against the admission rule (localhost only) and the
unit_ipc row of section 5.1, and on that day 14 s per request. The
tester now runs its own loopback HTTP/1.1 server on a listener thread of
ipc_tcp_socket_server (accept, read until the header block and
Content-Length bytes of body are in, answer, close; the server's kill
sentinel ends the loop): a request with a body is echoed back with 201,
a body-less request gets a canned document with 200, and the assertions
are on the exact content and content type rather than hashes, plus the
reset of the response between three requests on one object and the fast
failure against a closed port. The request tests are skipped on the
platforms where the ipc listener-thread tests are skipped (_WIN32,
__FreeBSD__, which the Mac build defines), like persistent worker
server; the lifecycle test runs everywhere. Finding, not changed: a
body-less http_post%request sets no POST fields, so libcurl issues a
GET; every production caller passes a body, and the test pins the
behaviour as it is.
inplane (2026-09-23, Hans). The second family of the unassigned area:
seven identities, fourteen files, about 2 500 lines from the continuous
in-plane rotation project of August, none in CI, all standalone. Three
things were under test: (A) the polar continuous-angle evaluators of
simple_polarft_calc (gen_raw_euclid/corr/hybrid_grad_at_angle against
the discrete gen_raw_euclid_grad_for_rot_8 and gen_objfun_vals) and
the joint route of simple_pftc_shsrch_grad; (B) the refine3D search
state (strategy3D_srch storage routes, seed_continuous_inplane_candidate,
resolve_inplane_e3, joint_evaluation_invalid) and the inpl_cont
policy; (C) post-run project metadata scans and a TSV baseline. Nothing in
the gate touched polarft_calc before. Hans's area decision: the 2D and
3D searches share the machinery, so the area is pftc_align2D3D
for everything on the polar Fourier transform, with cart_align3D
reserved for the Cartesian continuous implementations (the pose family).
Verdicts: inplane_cc_grad, inplane_hybrid_grad, inplane_rot2D_stage1,
the numeric and route-construction halves of inplane_rot2D_routes and
refine3D's synthetic_recovery are merge into unit_pftc_align2D3D
as the continuous in-plane sub-suite (simple_pftc_inplane_tester,
beside simple_polarft_calc). All five needed vol1=; the fixture is now
hermetic — a four-blob phantom (box 64, 1.3 A, 60 A mask) written to the
run directory for the suite and removed at its end; the checks are
properties of the evaluators, not of a molecule. Kept as asserted: the
score = exp(-raw loss) and scalar-route identities, parity of the
continuous evaluator with the discrete reference at grid angles (loss and
x/y gradient), analytic-vs-central-difference gradients at twelve probe
poses per evaluator (1 % + 3e-3 floor), the non-negative loss series and
[0,1] scores under near-noiseless shell-dependent sigma2 with stale and
re-memoised square sums (scan thinned from 5x5x81 to 3x3x41), the cc
penalty for an undefined denominator (finite, > 1, zero gradient), the
hybrid capability flags, the joint route's seed parity with the legacy
callback, and recovery: the joint solve from the grid seed does not worsen
the objective, improves the angle over the grid and lands within 0.25 px
rms of the known shift; plus the full-band hard-edge fixture at 359.375
degrees and the zero-angle fixture (first grid angle selected, vanishing
loss), and the strategy2D route flags under inpl_cont=no|yes and the
probabilistic mode. Dropped: the hybrid-objective rejection (the
constructor error stops; the old test spawned itself to observe it) and
the aliasing experiment (printed, never asserted). refine3D's
search_state, joint_state, direct_route, metadata_state and
policy are merge into unit_pftc_align2D3D as refine3D
in-plane state (simple_strategy3D_inplane_tester): hermetic, no
fixture, error stop became assertions, resolve_inplane_e3 gained the
first-grid-angle case. inplane_rot2D (a driver: GNU find -printf for a
1JYX project made by a simple_test_1jyx_abinitio that no longer exists,
two simple_exec prg=abinitio2D runs, the other programs as subprocesses)
is delete. inplane_rot2D_meta, refine3D metadata_project and
baseline (all need a finished project, the baseline a TSV snapshot too)
are delete: the e3/inpl consistency they scan for is what
resolve_inplane_e3 guarantees, and a project-level assertion after a
real run belongs to the simulated workflow (Phase 5). Registration:
unit_pftc_align2D3D is the ninth fast entry, SIMPLE_CTEST_BUDGET
22 -> 23. Coverage accounting: every production call of the fourteen files
is made by the two new testers except the abinitio2D workflow run and the
project scans, which were the deleted drivers' own business; nothing is
lost by name. Build notes: vol_pad2ref_pfts fills nspace references
and checks the bank (a 2026-09-10 change the August tests predate), so the
fixture asks for nspace=2 (the smallest even count build_refspiral
accepts) and a two-reference bank; and the Linux bounds-checked build
showed that the old tests projected from the unpadded volume while the
pftc's polar coordinates live on the padded lattice, reading past the
expanded bounds at the full band (silent on the Mac at lp = 8): the
fixture now builds vol_pad exactly as simple_matcher_refvol_utils does
(pad_fft by OSMPL_PAD_FAC, expand_cmat) and projects from it.
pose (2026-09-23, Hans). The third family of the unassigned area:
pose_cont_refinement (numerics, solver, helpers), pose_cont_refine3D_adapter
(adapter contracts and the opt-in 1JYX experiment) and, because it is the
Cartesian Fourier layer under the pose refiner and under PCG,
cartesian_fourier — eleven files, about 2 600 lines, all hermetic, none
in CI, written mid-September on the simple_cartesian_pose_refiner,
simple_pose_cont_refine3D_adapter and simple_cartesian_fourier modules.
Hans's area decision: cart_align3D for the Cartesian continuous
implementations, so this batch creates unit_cart_align3D (tenth
fast entry) and lib_cart_align3D (second library entry).
Verdicts: pose_cont_refinement is merge into unit_cart_align3D
as pose refiner (simple_cartesian_pose_refiner_tester, 669 lines):
prepared-particle validity and shell capping, exact matches giving zero
objective and gradient without CTF, with CTF and shell whitening and with
phase flip, the inverse-envelope constructor applying its correction once,
the Fourier shift phase sign on the native pixel scale, the 1-NCC formula
against an independent evaluation and its invariance to particle gain,
five-parameter gradients of both objectives against central differences,
the Cartesian gather against the PFTC projector kernel at a matched
boundary, orthogonality under the right rotation increment; then the LM
solvers: shift-only recovery within the step bound and the exact shift
retained, joint recovery of a known five-parameter pose within its bounds,
active-parameter masks, the cumulative guard, the NCC solver on a
gain-scaled particle, and invalid or unobservable inputs leaving the pose
untouched — everything the old test asserted, on simple_test_utils.
The adapter case of pose_cont_refine3D_adapter is merge as pose
adapter (simple_pose_cont_refine3D_adapter_tester): reference-artifact
workspace lifecycle, the observation adapter against the established
particle path, native/crop shift conversion, the inpl_cont-to-pose_cont
seed handoff and round trip, the transaction contracts of both routes and
the seed validity. Its 1jyx_reconstruction case is demote to
lib_cart_align3D as pose 1JYX recovery
(simple_pose_cont_1jyx_tester): 1JYX from the embedded coordinates at
box 144, 5 000 simulated particles with varying CTF and noise, every pose
perturbed by 15 degrees and 2 px, LM on each, three reconstructions with
FSC; the aggregate objective, rotation and shift errors must fall and the
refined reconstruction correlate better with the truth than the perturbed
one; the MRC volumes and TSVs it writes are the reviewable record and stay
in the nightly's run directory. cartesian_fourier is merge as
Cartesian Fourier (simple_cartesian_fourier_tester): the fast KB
polynomial and derivative against the ideal Bessel window, normalised
stencil derivatives and partition of unity, the stencil-switch jump, the
packed/Friedel gather derivative against finite differences, and the
parity of the extracted neutral operations with the pre-extraction
oracles retained in the tester; the self-re-executing driver is gone.
SIMPLE_CTEST_BUDGET 23 -> 25. Coverage accounting: every production call
of the eleven files is made by the four testers; nothing is lost by name.
Naming (Hans, same day): the areas are pftc_align2D3D and cart_align3D
("registration" was too long for a suite name); the first was renamed from
pftc_registration2D3D, under which it was committed in 675fdcd8e.
heterogeneity (2026-09-23, Hans: "go"). The flex trio: three
standalone drivers (52 lines) over self-tests that lived inside the
production modules — about 290 lines in simple_flex_pca_model,
_weights and _deconv, 412 in simple_flex_pca_pcg, about 1 300 in
simple_flex_gpu — all failing by THROW_HARD, none in CI. flex_pca
is merge into unit_heterogeneity as flex PCA
(simple_flex_pca_tester): the embedding-cache round trip (bit-exact,
seven payloads), the derived settings against the validated IgG and
Ribosembly scales, state placement with a population floor, kernel
weights at bandwidth with bounded widening, covariance state weights on a
bimodal embedding, and latent deconvolution (noise-scale calibration,
held-out K = 2, posterior means closer to the truth); the six routines
left the production modules (the ui_hash precedent) and
write_embedding_cache, read_embedding_cache,
place_states_with_population_floor, FLEX_AUTO_K_START/MIN became
public for the tester (Hans agreed). Two of them were not reproducible:
the population floor seeded from /dev/urandom through seed_rnd and
the deconvolution called a bare random_seed(); both now draw from a
fixed seed. Tidy found on the way: COV_MAX_BW_GROW was defined three
times (model, weights, util); the model's copy was unused and is gone,
the weights module now imports the util one, which is public. flex_pcg
is modify: test_flex_pcg_operator is white-box (private components of
flex_pcg_t, sixteen private scatter/fold kernels) and stays in the
module; it gained an optional passes(4) out-argument, a fixed seed in
place of seed_rnd, and a FAIL line that also reports (D). The
simple_flex_pcg_tester asserts (A) operator vs exact Gram, (B) rhs
deposit vs exact adjoint, (C) the CG solve, (D) band lists vs dense fold
by name: box 32 with 200 samples as flex PCG operator in
unit_heterogeneity, box 64 with 400 samples in lib_heterogeneity; the
three single-sample debug runs whose results the driver ignored are
dropped. flex_gpu is keep, platform: test=flex_gpu through
simple_test_exec, the five CPU-vs-CUDA routines as sub-suites (they skip
themselves without a CUDA build or device and stay in simple_flex_gpu),
registered only with USE_FLEX_CUDA and counted like coarrays.
SIMPLE_CTEST_BUDGET 25 -> 27. Coverage accounting: every production call
of the self-tests is made by the new testers or the wrapped routines;
nothing is lost.
The first gate with unit_heterogeneity passed 11/11 in 47.3 s against
the 30 s budget: flex PCA 36.2 s, all of it the deconvolution
(20 000 particles, noise variance 2..20 per axis, K ladder to 4), and
flex PCG operator 10.1 s. Two fixes and a split. (1) Production:
xd_fit factored T_ik = R_i Sigma_k R_i^T + N_i twice per particle and
component and iteration — once for the responsibility, then again with an
explicit inverse for the conditional moments — and xd_posterior did the
same. One Cholesky factor now serves both (component_factor,
component_moments: W = L^-1 R Sigma, b = mu + W^T y,
B = Sigma - W^T W); component_terms is gone. Mathematically identical,
about half the E-step arithmetic by operation count, in the loop behind
the 1 178 s K ladder once measured on 105k particles in 17 dimensions
(the reason for the XD_CV_MAX subsample). (2) The PCG
self-test's sample_exact summed the box^3 volume in a triple loop for
each of the 38 256 central-slice samples of (C) at box 32 (1.25e9 terms);
it now does the x sum for all columns as one real (2,N) x (N,N^2) library
matmul, then y and z. (3) The deconvolution runs on 4 000 particles at
noise variance 0.5..5 in the gate (the ladder stops at n/2000 = 2; an
independent numpy emulation of the EM over nine seeds gave a held-out
margin of 54-123 nats for K = 2 over K = 1 and an mse ratio of 0.33-0.35
against the asserted 0.6) and on 20 000 at 2..20 nightly as
flex PCA deconvolution 20k (the emulation gave K = 3 about 5-6 nats
below K = 2, as the Fortran run did). The PCG solve (C) runs only the
clean baseline in the gate; the twelve-setting sweep is flex PCG solve
sweep in lib_heterogeneity, where every clean solve (six of them, err
inner 0.004-0.009 in the first run) is now held to the baseline criterion
instead of only the first; the noisy ones stay unasserted, they show what
the Tikhonov term is for. The unused loc_fixed argument went with the
debug runs that used it. Found on the way, not changed: the XD EM runs to
XD_MAXIT = 150 without meeting XD_TOL in most K >= 2 fits of the
emulation (slow EM convergence when the noise dominates the component
widths), so the ladder's cost is iteration-capped rather than
tolerance-limited.
Differential evolution and the harness seed (2026-09-23, found by the
heterogeneity rebuild). The second gate was 4.9 s (unit_heterogeneity
4.5 s) but unit_numerics failed: test_de_quadratic stopped at
(1.38, -2.47). Nothing in the batch touched it; the test drew from a
generator that run_unit_suites and test_multinomal had seeded from
/dev/urandom (seed_rnd), so it had been passing by chance. Two
production defects in de_minimize, both clear: (1) the worst member was
tracked wrongly: a rejected trial set worst = X (whose cost had not
changed; when X was the best member the spread became zero), and an
accepted trial on the worst member left worst pointing at a member that
had just improved. The relative population-tolerance stop then fired on a
spread that was not the population's; a numpy replay of the routine on
this quadratic over 500 seeds stopped early in all 500 (median 270 of 3000
trials) and missed the asserted optimum in 121. Now worst is recomputed
when the worst member improves and a rejection changes nothing; the replay
runs to maxits in all 500 and reaches cost below 1e-18 in every one.
(2) The component that must always be mutated was chosen as i == X,
comparing a dimension index with a population index; it is now a random
dimension jrand. The unused nworse counter went. DE is production's
CTF search (ctf_estimate_cost, maxits 400): it no longer stops on a
spurious spread, so it can spend all 400 trials, and its fits change (not
measured on data). Reproducibility: simple_test_utils
gained set_fixed_seed (the formula every tester copy used; the copies in
the gauran and flex PCA testers are gone); run_unit_suites seeds with it
instead of seed_rnd, test_multinomal too, and the DE and simplex tests
seed themselves. Production code that calls seed_rnd (for example
simple_ftexp_shsrch, simple_parameters_phases) still re-seeds from
/dev/urandom for whatever runs after it.
Linux follow-up (after db50811a6). The bounds-checked Linux gate
stopped flex PCG operator in (C) at rho_t index -5: the test's 1x
gridding density is laid out on the h >= 0 half (rho_lb(1) =
-(iwinsz+1)), but the rotated central-slice samples cover both halves, so
on the Mac, without bounds checks, every sample with h < 0 was added in
front of the array. A defect of the test since it was written; a sample
with h < 0 now counts at its Friedel mate, as the reconstructor's rho_exp
does. The same run printed -fcheck=array-temps warnings for the strided
row z(i,:) handed to component_factor in xd_fit, xd_loglik and
xd_posterior; the row is now copied once per particle.
singles I (2026-09-23, Hans: verdicts 1-11). Thirteen of the
twenty-four unassigned standalones, to the areas of their machinery; no
new area (Hans: "no new area"). Three were duplicate drivers of gate
sub-suites (gui_assembler, gui_metadata, project_merge) and are
deleted; multinomal's standalone had gone in the stats batch and its row
is closed. rnd_shuffle and the multinomial draw form
simple_rnd_tester (random draws, unit_numerics, replacing
multinomial random draw, which printed frequencies): the shuffle
invariants and the draw asserted at a fixed seed within four binomial
standard deviations (1000 draws: +-0.05 at p = 0.8, +-0.04 at p = 0.1),
and an entry of probability zero is never drawn. phshift_star is
test_relion_phase_shift in the STAR project tester. ui_visibility
(about 170 assertions in a program body) is simple_ui_visibility_tester
(UI visibility, unit_ui, six tests), and phshift_policy's error stops
are its test_phshift_contract; the registrations it pins were
spot-checked against the current UI (every named program exists, the
category descriptors and the three category counts match).
discrete_stack_io joined stack I/O (unit_core): the concurrent
dstack_io reads of twelve float32 and twelve int16 stacks with three
threads (option a: the concurrency is the point, milliseconds), and the
float16 rounding (half to even), bit patterns on disk, image-layer mode
inheritance and statistics update, and subnormal/signed-zero boundaries
that the tester did not cover yet; its THROW_HARDs are assertions.
sigma2_state is simple_sigma2_state_tester (sigma2 state) and
eul_prob_tab2D_io is simple_eul_prob_tab2D_tester
(2D probability table I/O), both in unit_pftc_align2D3D (option a for
sigma2: the likelihood objective consumes it). projdir_accumulator is
simple_classaverager_tester (class-average accumulator,
unit_reconstruction: class averaging is 2D reconstruction).
cavg_quality_relations was an 8-line driver over a self-test inside
simple_cavg_quality_relations; the test left the module for
simple_cavg_quality_relations_tester (cavg quality relations,
unit_numerics) and calculate_promoted_feature became public for it.
Twelve standalone files deleted; the phase-shift policy (section 9 and
10.1) and the GUI onboarding note name the sub-suites. Coverage
accounting: every assertion of the twelve programs is carried over, the
error stops and THROW_HARDs as assertions; nothing is lost. The first
build stopped unit_ui in UI visibility: make_ui had already run in
UI JSON, and a second build aborts because add_ui_program refuses a
key that is already registered ("Key: export_relion already in ui_hash").
The standalone had been the only builder in its process; units has
the same double build (UI JSON, then refine3D in-plane state) and
could not have run past it. make_ui and make_test_ui now build
their table once per process and return on a second call; the duplicate
guard in add_ui_program stays, it catches two programs registering one
key. The other ten fast suites passed in that build (gate 5.0 s).
singles II (2026-09-23, Hans: verdicts 1-9). The remaining eleven
unassigned standalones. nu_envmask, nu_filter and phase_rand_fsc
are "not needed": simple_exec prg=nu_filt3D exercises the NU filter
and the fsc commander phase_rand_fsc; deleted. Every routine the two
NU programs called stays used in production (the evidence-state
accessors from the sharpening step, the margin from the envelope), so
nothing follows them out. diff_map_graphs is
simple_diff_map_graphs_tester (diffusion-map graphs, unit_numerics):
its thirteen checks were stop 'message', which exits with status 0, so
none could ever fail; they are assertions now (gated neighbours stay in
their projection bin, block-row assembly equals the whole build,
occupancy weights, the Perron vector of the view-balanced operator,
Nystrom coefficients equal to the eigenfunctions, an uncapped spectral
scan). create_gain and search_gain_flips were movie-driven runners
without assertions; stream preprocessing does both in production and
motion gain tests the summing and the analyser; deleted, and
normalized_inverse_average_intensity, which nothing tested, gained a
closed-form test in motion gain (per-pixel mean 2 with one pixel at 4
and one at 0: global mean 2, gain 1, 0.5 and 0). atomfit needed a PDB
in the working directory and asserted nothing; the routine it called,
atoms%fit_bfactors (133 lines), had no caller and is deleted with it.
stream_initial_analysis ran the p03 commander on a missing folder;
deleted. openmp_offload is Cyril's and moved as it is (Hans): the
standalone becomes a platform CTest entry registered with
USE_OPENMP_OFFLOAD (nthr=8 device=0, OMP_TARGET_OFFLOAD=MANDATORY),
counted in the process budget only when registered; its failed checks
stop with status 0, so the entry fails on Fatal error|FAIL: in the
output instead of by exit code (the one-word fix, error stop, is left
to him). qsys_ctrl and qsys_env wait for the parallel-area batch
(option a). Eight standalones deleted; the NU envelope algorithm note,
the motion-gain policy and the code map follow.
parallel (2026-09-23, Hans: "agreed", and a unit_parallel). Four
identities on both routes plus the two deferred qsys programs. coarrays
is two tests under one name and both stay: the standalone sync check is
the coarrays platform CTest entry, the exec case is what the coarray CI
job runs (SIMPLE_QSYS=coarray: simulated noise, check_nptcls over the
partitions through the coarray backend); its closing line claimed
"coarray Euler shift checks passed", it now says what ran. openacc (an
exec body that was entirely commented out, a standalone saxpy timing on
10^9 elements), openmp (the nowait race of the OpenMP runtime, printed
passed/failed) and simd (a timing) asserted nothing and touched no
SIMPLE code: deleted on both routes, simple_test_openmp and
simple_test_simd dropped from the CI workflow, and, as nothing in src
uses OpenACC, the USE_OPENACC option and its CMake block with them.
qsys_ctrl and qsys_env are the twelfth fast suite, unit_parallel
(Hans: useful for more later): simple_qsys_ctrl_tester
(qsys control: the controller over the local backend, four partitions,
two computing units, scripts only; the thirteen flag checks are
assertions, the fresh-status check used .and. where it needed .or.,
and the partition script's range, the restored job description and the
multi-job script's contents are now checked) and simple_qsys_env_tester
(qsys environment: the project-stored installation path is dropped and
ignored, the executable resolves from the local SIMPLE_PATH of the CTest
environment, an empty project gets 0-0:1:40). SIMPLE_CTEST_BUDGET
27 -> 28. Also noted: CI runs simple_test_exec test=units, which stopped
at the double make_ui from the in-plane batch until 5f3c705ce.
The first build failed one check: the standalone's "the queue
description keeps no simple_path" contradicts the policy of 9a87c12a8
(executables come from the environment), under which qsys_env%new
writes the runtime SIMPLE_PATH into the queue description on purpose
and builds the executable path from it; the program had never run. The
test now pins that the description carries the local SIMPLE_PATH and
not the project's. The same run printed "SIMPLE_QSYS_PARTITION is not
defined" from production: an optional variable read like a required one,
in every run without a partition. simple_getenv gained silent= and
the optional reads (SIMPLE_QSYS_PARTITION, SIMPLE_EMAIL, which has a
default) in qsys_env and update_compenv use it.
network (2026-09-23, Hans: "agreed"). socket_client,
socket_server, socket_io (both routes) and socket_comm_distr (exec)
were role programs to run by hand, asserting nothing: an endless accept
loop, a server thread that never ends with five 10 s sleeps, a 10 s
sleep. What they drove was dead: simple_distr_comm (71 lines) had no
caller and simple_socket_comm (298 lines) was used only by it; the TCP
transport production uses is simple_ipc_tcp_socket_*, tested by the
bounded localhost IPC TCP socket sub-suite of unit_ipc. Deleted: the
four tests on both routes, the network test category
(simple_commanders_test_network, simple_test_exec_network,
simple_test_ui_network and their hooks in simple_test_exec_api,
simple_test_exec.f90, simple_ui_test_group), the two modules, and
production/tests/test_socket_comm_distr.f90, which had never been
built (without the simple_ prefix the CMake glob does not see it).
network is no longer a simple_test_exec category. nice waits for
the utils batch.
single (2026-09-23, Hans: "Ruben's code: transfer it, and instruct him").
The six SINGLE cases are Ruben's; they moved as they were and
doc/refactoring_notes/single_area_tests_handover.md tells him, test by
test, what they have to assert. detect_calpha (the only one with a
failure path) is simple_calpha_finder_tester, sub-suite C-alpha finder
of the new thirteenth fast suite unit_single, which also takes the
atoms sub-suite from unit_project (the atoms module is SINGLE's).
simulate_nanoparticle and detect_atoms were prefixes of atoms_stats
and are retired; the pipeline runs nightly as nanoparticle atoms
(Pt, smpd 0.358) and detect_calpha_molecules as C-alpha molecules,
both in the new fourth library suite lib_single, through their
unchanged commanders; neither asserts anything yet (the report shows
them "completed"). Two defects fixed on the way: the CTest entry
single_workflow passed only element=Pt, so the required smpd was
missing, simple_cmdline%parse printed the usage and stopped with status
0, and the entry "passed" without running (no other registered entry has
required keys); it now passes smpd=0.358, for Ruben to confirm. And
every stage of atoms_stats and single_workflow ran with nthr=40
(module constant), whatever the entry's 8 threads; they take
params%nthr. SIMPLE_CTEST_BUDGET 28 -> 30.
stream (2026-09-24, Hans: "go" on the eleven proposals). The seven
stream cases are Ruben's (2026-08-25/26) and were seven nightly workflow
entries. Unlike the SINGLE cases they checked real things through
THROW_HARD; the review moved each to the area of what it tests and turned
its checks into assertions (same conditions, same messages; reads that
depend on a missing file are skipped; the fixture directory, which every
run used to leave behind, is removed when every check passed).
doc/refactoring_notes/stream_area_tests_handover.md tells Ruben what
they should pin beyond counts and files.
sieve_cavgs tested ptcl_sieve%collect_and_reject: it is
test_collect_and_reject_hard_gates of simple_ptcl_sieve_tester, so in
the fast gate (unit_project, particle sieve): two class averages of
64², one kept, one blank and rejected, exact expectations.
assign_optics (the stream's p02 stage, exact truth: two beam-shift
clusters, populations 2 and 3, centroids to 0.01), gen_pickrefs
(make_pickrefs: counts and diameter metadata) and pick_extract (three
copies of a reference picked and extracted) are the sub-suites optics
assignment, picking references and pick and extract of the new fifth
library suite lib_stream (simple_stream_tester, one thread). Optics
assignment takes a minute: the production watcher imports a project only
once it is LONGTIME = 60 s old. The inventory's "manual (needs
user-supplied)" for it and for preproc was a dossier artefact
(dir_target, dir_movies); both generate their fixtures. pick_extract
sets nboxes_max=3, the number it asserts, so over-picking cannot fail it
and the positions are never compared (handover).
master never started the stream master: it tested
gui_assembler%assemble_stream_heartbeat over seven live forked children,
running and finished only. It is run_stream_heartbeat_tests of
simple_gui_assembler_tester (whose header had said the heartbeat was
untestable there), sub-suite stream heartbeat of the forked_process
platform entry beside the forked-process lifecycle tests.
preproc submits its jobs to the local queue, so it stays the workflow
entry stream_preproc; simple_commanders_test_stream keeps only it. It deletes simulate_movie_params.txt and the optimal average, the
truth it could be compared with (handover).
abinitio2D_stream is retired: it ran abinitio2D for one iteration on 24
noise-free particles with a hand-written command line that differs from the
one the stream's chunk code builds (cls_init, rank_cavgs, chunk,
objfun, refine), checked counts and files but not that the two particle
families separate, and abinitio2D runs nightly in both
simulated_workflow systems.
Seeds. parameters%new calls seed_rnd, which read /dev/urandom, so
every test that runs a commander drew unseeded numbers from its first
commander on (the movie simulator's noise and positions, class
initialisation; lib_single and the workflow entries alike). seed_rnd
now honours the environment variable SIMPLE_SEED: set, the seed is that
integer advanced by 7919 per call since the last fixed seed, so successive
commanders in one process draw different but reproducible numbers, and
distributed workers inherit it; unset or empty, production is unchanged;
not an integer, it stops. seed_rnd_fixed in simple_rnd holds the
fixed-seed formula (set_fixed_seed of simple_test_utils and the flex
PCG self-test's private copy call it) and restarts the count. CTest sets
SIMPLE_SEED=20260923 for every entry, and run_unit_suites reseeds
before every sub-suite instead of once per process, so a full run, a
focused run (suite=<name>) and a reversed run draw the same numbers in
each sub-suite. The project-records tests called seed_rnd themselves (a
non-reproducible fast test); they take a fixed seed.
Tidy: the unit_project UI text still listed atoms; the router comment
said twelve fast suites; the distributed-execution comment had landed on
suites_single. SIMPLE_CTEST_BUDGET 30 -> 25 (13 fast + 1 platform +
6 workflow + 5 library).
First build: 12/13 in 5.0 s; unit_numerics stopped in straight-line
fit (THROW_HARD, so the rest of the suite did not run). The reseeding
moved the draws it starts from, and the test was flaky by construction:
10 000 random exact lines, slope 5·U with a random sign, and r squared
required at or above 0.9999. For a near-flat line r squared is 0/0 in
single precision; a float32 emulation of the fit puts r squared below
0.9999 for |slope| under about 5e-6·|intercept|, about one draw in
200 000, so about one seed in twenty failed. The closed forms of
fit_straight_line (exact line, perturbed line with its analytic r
squared) were already in linear algebra; the sub-suite is gone, and
linear algebra gained test_fit_straight_line_recovery: 35 exact lines,
slopes from -5 to 5 through near-flat and flat, intercepts from -10 to 10,
slope and intercept within 1e-5 (the emulation gives 6e-8 at worst). r
squared is not asserted there; the one production caller,
guinier_bfac, uses only the slope. The Linux build also warned that
write_merged_coordinates in simple_flex_pca_model (the coordinates
table of the paired-merge delivery, added 2026-09-17) is defined but never
called; it is removed.
utils and the standalone programs (2026-09-24, Hans: "we still have a
number of simple_test* executables other than simple_test_exec. They need to
go and the tests they execute need to become part of the new
environment"; verdicts: as proposed, apart from cif2mrc/cif2pdb "programs
rather than tests", nice deleted, offload 11a). Ten standalone programs
were left in production/tests, eight of them with a simple_test_exec
twin in the utils area, so the batch retires both. production/tests, the
simple_test_*.f90 glob with its per-program executables and installs, and
the utils test category (commander, router and UI modules, and their
hooks in simple_test_exec, the exec API and the test UI group) are gone;
simple_test_exec is the only test executable. No CTest process is added
or removed (budget 25).
ansi_colors printed seven coloured words: test_ansi_format_str in the
string tester (string, unit_core) asserts the escape sequences of
format_str and the ANSI code table, which found C_MARKED_WHITE = 46,
cyan's background; it is 47. stringmatch printed list_of_ints2arr of a
list: test_list_of_ints2arr (same sub-suite) pins blanks, a single number,
and trailing and doubled commas; for 1,2, it sized the array for two
entries and wrote a third past its end, and the routine is rewritten to
skip empty entries (it reads the state and class selections of two
commanders). cmdline mostly duplicated command line; its full
processing line (typed, name and path values, checkvar/check, delete) is
test_read_line_typed_values there, and cmdline%writeline, which only
the test called, is removed. serialize displayed a masked square and
checked nothing, while production (PCA denoising, stack operations, the
nanoparticle tools) had no test of it: simple_image_serialize_tester
(image serialisation, unit_image) pins the masked and full round trips,
the column-major order and the zeros outside the mask.
cavg_registration was a self-test inside simple_strategy2D_utils; it
moved to simple_cavg_registration_tester (class-average registration,
unit_pftc_align2D3D; match_imgs and match_imgs2ref are exported for it)
with its correlation checks as assertions, plus the applied rotation
within one polar step (either sign of e3) and the shift length within
0.5 px. pdb2mrc is simple_pdb2mrc_tester, sub-suite pdb2mrc of the
nightly lib_single (6VXX and 1JYX, default and explicit file names, the
exec twin's checks as assertions).
cif2mrc and cif2pdb downloaded 6VXX from RCSB (curl's status ignored)
and ran the production programs of those names, checking nothing: runs of
programs, not tests, and deleted. nice posted to an unresolvable
"testserver" and slept 20 s: deleted. install ran simple_test_units,
which no longer exists, so it checked nothing although CI ran it twice;
deleted, and doc/installation.md now points to simple_test_exec
test=units. The standalone coarrays only synchronised two images; the
coarrays platform entry now runs simple_test_exec test=coarrays, the
production coarray path end to end (qsys=coarray images of the
coarray-linked simple_private_exec under cafrun). openmp_offload
(Cyril's) is simple_openmp_offload_tester in src/utils, run as
simple_test_exec test=openmp_offload nthr=8 device=0 by the same
platform entry (USE_OPENMP_OFFLOAD, OMP_TARGET_OFFLOAD=MANDATORY); its 65
stops, the text ones exiting with status 0, are THROW_HARD with the same
text (the numeric CUDA/cuFFT codes in the message), and the entry keeps
its failure-message pattern as a second net; built without offload it
reports that it was skipped. Its source joins the
conditional-branch sources that suppress unused-variable warnings. The
offload branch compiles only in an offload build: Cyril should build it
once.
CI's test step lost its calls of simple_test_install (twice),
simple_test_ansi_colors, simple_test_serialize and
simple_test_stringmatch; the UI visibility test asserts angres,
openmp_offload and that the retired cavg_registration program is gone;
test_timing_run.sh and test_review_dossier.py note that the standalone
route is empty. The unit_image UI list of sub-suites was stale (it named
the moved shift search and missed five); it is complete.
The wrap-up (2026-09-24, Hans: "We can finish up the rest"; on the truth
gates: the abinitio3D maps are neither docked to the reference volume nor
necessarily of the right hand, so the gate involves docking and mirroring,
which SIMPLE has but which must be specified; on angres: "we need
something that actually measures the angular resolution of the projection
directions used in the search", a standard program rather than a test,
measure_projspace_angres with nspace given, with the tabulated values
kept where they fit; the rest as proposed). Three test programs were left
in the fft, geometry and masks categories of simple_test_exec, all
modify verdicts of 2026-09-22 waiting for library suites of the
provisional list (section 5.2.1) that the review never built.
gencorrs_fft is simple_polarft_corr_tester, sub-suite polar
correlation of unit_pftc_align2D3D: three smooth zero-mean images of box
64 go through the production polarisation path, and gen_objfun_vals
(objfun cc) peaks at rotation 1 with correlation 1 for an image against
itself, at the applied rotation (within one step, either sense) above 0.9
for a rotated copy, and stays below 0.5 at every rotation for an unrelated
image. Its seed was set before parameters%new, which reseeds through
seed_rnd, so without SIMPLE_SEED the images came from /dev/urandom;
it is set_fixed_seed(20260922) after parameters%new now, and its
private copy of the seed formula is gone.
msk_routines is test_masks_parallel_equals_serial in the mask tester
(masks, unit_image): eight noise images, box 128 in 2D and 48 in 3D,
are masked serially and then inside an OpenMP loop on an explicit team of
three (num_threads), with the mask coordinates memoised once outside the
region; the soft, soft-average and hard routines in 2D and 3D give the
serial result. The program needed OMP_NUM_THREADS; the explicit team
makes the check independent of the entry's one thread, as for the discrete
stack reader of unit_core (admission rule 3 now says so).
angres printed, and since 2026-09-22 asserted, oris%find_angres of a
spiral of 500 to 20 000 directions. It is the program simple_exec
prg=measure_projspace_angres nspace=<n> [pgrp=<pg>] [moldiam=<A>]
(orientation processing, advanced visibility). It builds the reference
directions as the 3D search does (sym%build_refspiral with nspace and
the point group) and reports the largest angle from a direction to its
third-nearest neighbour among the symmetry copies of the other directions
(for C1 the measure of find_angres), the mean of that angle and, with
moldiam, the resolution at the rim of the molecule (resang). It costs
nspace^2 times the number of symmetry operations dot products on the
threads given (4e8 at 20 000 directions in C1). The ladder of the original
program (spiral, 500 to 20 000 in steps of 500) is a comment above
find_angres in simple_oris_dists. The dead nspace_commander of
simple_commanders_refine3D (no program, no caller) is removed.
The fft and geometry categories of simple_test_exec are gone (commander,
router and UI modules and their hooks, as utils before them); the masks
category keeps the manual nano_mask and score_volume_shape. The UI
visibility test asserts nano_mask in the masks category and
measure_projspace_angres in orientation processing in place of the
retired angres. CI runs by label and name: the test step runs
ctest -L platform and ctest -R '^pcg_recon$' after the build step's
fast gate, and the coarray job ctest -R '^coarrays$', each with
--no-tests=error. The registry-consistency check (section 7, item 6) is
scripts/check_test_registry.py, run by run_fast_gate.sh before ctest.
On this tree it finds 27 registered selectors, 37 test programs and 37
router cases, consistent; on a scratch copy with an unregistered area
suite, a renamed router case and so a program without one, it reports the
three problems and exits with status 1. The code map is regenerated with
scripts/generate_codeoverview.pl, which also restores two files the
hand-kept map had missed and removes a duplicate entry. No CTest process is
added or removed (budget 25).
Phase 5 is Ruben's
(doc/refactoring_notes/phase5_workflow_gates_and_nightly_runner_handover.md):
the truth gates of each workflow entry, with the map gate specified as in
section 5.2.2 (a same-grid pdb2mrc truth map, dock_vols at 15 to 20 A,
the hand chosen by the docking correlation of the map and its mirror, the
masked FSC at 0.143 against a declared floor, poses composed with the
docking rotation and the mirror before sym_dists), a fix for the
NTHR = 4 constant of simulated_workflow, and the design and code of the
nightly runner. Section 12 and section 16 record where every phase and
criterion stands.
The open items (2026-09-25, Hans: "Fix the remaining smaller open
items"; on the pole/mirror pair of build_refspiral and the XD_MAXIT
iteration cap: leave them). The findings recorded as not acted on by the
batches above, and the test gaps they left.
Production defects, each pinned by a test. calc_graphene_mask excluded the
three shells nearest each graphene band unconditionally, so a band beyond
Nyquist cost the three highest shells; SINGLE's graphene subtraction had its
own copy with a third band (calc_3bands_mask) and the same flaw. One routine
now takes the bands (calc_graphene_mask(box, smpd, bands), GRAPHENE_BAND3
joined the other two in simple_defs), skips a band finer than 2*smpd, and the
mask tester checks two bands, three bands and two bands beyond Nyquist;
image%pspec_graphene_mask, its only other caller, had no caller and is gone.
image%bp converted both limits with get_find before looking at them, so
the common bp(0., lp) divided by zero (the flag, not a trap); only the
limits in use are converted, and the image tester asserts that bp(0., lp)
and bp(hp, 0.) raise no division by zero. stack_io%read only ever
advanced its buffer window, so a backward read looped for ever; the window is
now set to the block holding the image (the same windows a forward walk
loads), and the stack I/O tester reads backward and out of order.
binoris%open returned early on a new file before storing its name, so the
errors of a first write named an empty file; the name is stored first (no
test: it is private and only shows in messages). In a descending
print_segment_json window indices_pre and indices_post were the
ascending head and tail; NICE's micrograph panel reads them as the records
above and below the window as displayed (selectBelow, selectAbove in
panelmicrographs.js), so a descending view selected the wrong side. They
are swapped and listed in display order, and the project tester checks their
contents. otsu on a constant sample divided by zero in its range scaling;
it now returns the value (everything background). atoms%atom_validate and
map_validate cut the per-atom window one voxel off the atom (ang2vox is
1-based and window_slim adds one to the corner); in a numpy emulation of
convolve and the window mask, an atom correlated with its own simulated
density at 0.35-0.44 before and 0.99 after, which the atoms tester now pins
(above 0.95). Ruben's SINGLE handover has a note: his per-atom scores rise.
Test gaps. The fast gate's image sub-suite was test_image inside
simple_image: THROW_HARD checks, print-only filter, mask, rotation and
binarisation parts (and a rotational average that ran only with more than two
threads), and six image files left behind. It is simple_image_tester:
construction and access, dimension checks and foreground statistics, the FFT
round trip, get_nyq, the band-pass edges, apply_filter, the power spectrum
of a plane wave, shift against both shift2Dserial forms and a circular
shift, the autocorrelation's centre and shift invariance, rtsq (identity,
quarter turn, a turn and back), roavg of a Gaussian and a square, masscen
of pixels and voxels with and without a mask, corr of two Gaussians
(closed form 0.967), bit-exact SPIDER and MRC round trips of stacks and a
volume, and fproject against fproject_serial and across orientations of
an isotropic volume, centred. That covers the image basics only the deleted
ptcl_center named, except window_center, which had no caller and is gone.
First build: 54 of 58 image checks; the failures were the test's. A rotation
and its inverse and the rotational average of a Gaussian were compared over
the whole box, where the circular closure of rtsq (the corners take in the
opposite side) and the sharp disc edge dominate (a numpy emulation of rtsq
gives 0.978 and 0.022); they are compared inside radius 20 and 30 (0.9999 and
3e-5). corr of the two Gaussians was taken in a box of 64, where leaving out
the indices with |h|**2 < 2 gives 0.959; it uses the box of 100 of the original
test_image again (0.9672 emulated, closed form 0.9665). The run also fixed the
conventions, now pinned: shift(s) gives out(x) = in(x + s), and rtsq by 90
degrees gives out(i,j) = in(2c-j,i) about the centre c.
polarft_calc%rotate_ref_8 had no test: calc_frc, its production caller,
now has to peak where the FFT path of gen_objfun_vals does for probes of
+60, -60 and 180 degrees, which put the peaks in each branch of the rotation
(the identity, both halves of the in-plane range, the half turn);
calc_corr_rot_shift, a benchmark copy with no caller, is gone. The
atoms sub-suite was test_atoms inside simple_atoms, with private
assertions that stopped at the first failure; it is simple_atoms_tester
(access, geometry, a PDB round trip, the ANISOU columns, density simulation,
map_validate and atom_validate), and what only the self-test called is
removed: cc_res (its sum was never initialised), find_masscen (a copy of
get_geom_center), rotate, geometry_analysis_pdb, get_num,
does_exist, print_atom, get_atom_corr, plus the equally unused
does_exist of dstack_io and stream_watcher.
Seeds. pose_cont_1jyx and pcg_halfset carried private copies of the
fixed-seed formula; they call set_fixed_seed, which draws the same numbers.
pcg_recon put an all-42 seed; it is set_fixed_seed(42), which draws
different numbers, so it needs one run by hand (section 16). The polar
correlation tester seeded before parameters%new (see the wrap-up); no
private seeding is left in the tree.
Headers. The 41 source files without a !@descr: line (and one with an empty
tag) have one; check_descr.py now also rejects an empty tag. The code map is
regenerated.
Left open by decision: the jittered pole and its mirror mate in
build_refspiral for d, o and i (the symmetry tester tolerates exactly that
pair); the XD EM running to XD_MAXIT (needs measurements on real flex data
before choosing acceleration or a looser tolerance).
The policy (2026-09-25, Hans: a policy document for the new environment:
what a test is and is not, where it goes and whether it is part of the build,
how to implement one, why not standalone programs, the library support, what
was deleted and where each old test is now, and what the nightly suite holds
and how it is designed). doc/policies/test_environment_policy.md. It also
answers the question developers ask about scratch work: a local git worktree,
a developer program, or a test; the test area is not a scratch area. Hans's
emphasis: SIMPLE practises extreme programming, everyone works on and pushes
to master, and a branch in the online repository needs a strong reason;
scratch work lives in a git worktree on the developer's computer and is
merged into master and pushed when it is good. Writing it found
three things. suite_id turned the blank padding of the character(32)
sub-suite names into underscores, so suite= matched no sub-suite at all
since Phase 2; it now stops at the last non-blank, and drops commas and
slashes (search_sort_locate, stack_io), still accepting the old spelling.
The suite= lists in the test UI had drifted from the suite tables for
unit_core, unit_ori, unit_numerics and unit_project (a third of the
sub-suites missing), and flex_gpu had no list; they are complete, and
check_test_registry.py now fails the gate when a list and its table differ
(the old UI file gives six problems). Ruben fixed the same padding defect
independently (0e4d441fe, trim and underscores only); the merge keeps the rule
above, which also drops commas and slashes, with his adjustl, and the check
caught his three new or renamed sub-suites (ANSI formatting, cif2mrc,
cavg registration) missing from the UI lists, now added. The repository's agent skills
(.github/skills) still described the production/tests glob; they, the
onboarding slides and three policy and algorithm notes point to the policy
now.
The SPIDER header (2026-09-25, first build of the policy batch: unit_image
stopped in test_file_roundtrip with an index 33 above the bound 32 of the
header array in simple_imghead::read). getLabbyt returned the record
length lenbyt (4*nx bytes) instead of the header length labbyt (labrec
records of lenbyt, at least 1024 bytes). The reader allocated nx words and
took 43 fields from them, so a SPIDER file with a box below 43 read past the
array (a bounds-checked build stops, a release build reads heap memory into
the fields from nx+1 on, below a box of 13 labrec itself); the writer wrote
nx words, so a box below 43 lost the fields from nx+1 on, the pixel size among
them. getLabbyt returns labbyt; the reader takes the 43 fields in one
statement (a header is never shorter), the writer puts the whole header, the
fields and then zeros, in one statement instead of one statement per word, and
both check iostat. get_spifile_info read the header three times with the
same unit settings, labelled native, big and little endian (the convert= had
gone), and returned the label to two callers that dropped it; it reads once,
and conv is gone, as are the unused print_entire of read and pos and
print_entire of read_tiff. starproject%check_stk_params (the box of an
imported stack) opened MRC and SPIDER stacks with an unset status, could not
make a SPIDER header (new without dimensions throws) and left TIFF handles
open; it calls find_ldim_nptcls. The in-module test_imghead (one box of
120, dimensions only, a file left behind) is replaced by
simple_imghead_tester (sub-suite image header of unit_image): the SPIDER
header geometry for boxes 32, 120 and 300, round trips of two images and a
volume with the file length, the pixel size, iform and maxim, a box-32 header
reading a box-120 file, find_ldim_nptcls and find_img_smpd on SPIDER, and
an MRC round trip. Budget unchanged (25).
The in-module self-tests (2026-09-25, Hans: is anything left of one test
per class? Then do it now; srchspace_map2D can go; hclust and the jpg
wrapper by the recommendations). Eight production modules still held
self-tests registered as fast-gate sub-suites; none used the assertion
library, three (online_var, aff_prop, hclust) only printed their verdict
and could not fail, the others stopped the whole area process at the first
failure. Each is now a tester next to its module, pinning what the old test
claimed, with independent expected values:
simple_online_var_tester (closed-form means and sample variances, one and no
samples, a 1e4 offset, the two-pass moment); simple_aff_prop_tester (the
exemplars are the cluster medoids 7, 19, 31, found by brute force in the test
and by a float32 numpy emulation for preferences -1 to -200; the true
partition; simsum recomputed; preference 0 makes every point an exemplar and
-1000 one cluster; restarts reproduce exactly; the input is untouched);
simple_ftiter_tester (loop limits for even, odd and non-square boxes; the
logical half maps one to one onto the half-complex array, comp_addr_logi
inverts it, a negative h addresses its Friedel mate, the three forms agree,
for six 2D and 3D boxes checked in a Python emulation; Nyquist and Angstrom
conversions; the low-pass clamps); simple_ftexp_shsrch_tester (through the
public ospec callbacks: the cost is lowest at the applied integer shift, the
correlation there is 1, the fdf callback agrees, the gradient vanishes at the
peak and matches central differences off it; minimize returns the applied
sub-pixel shift within 0.05 px, which the old test never looked at);
simple_bspline_smoother_tester (a cosine comes out scaled by the closed-form
transfer function B2/(B2 + lambda R) of the quadratic B-spline and its
gradient Gram kernel, 2D square and non-square and 3D at 32**3, which a numpy
emulation of the code matches to 2e-7; Fourier input). test_oris repeated
the oris tester with broken checks (a verdict never read, an assignment loop
that compared x only, five files left behind); it goes with corr_oris, its
only caller, and the oris tester checks deep-copy assignment. Sub-suites:
orientation data, hierarchical clustering and 2D search-space map I/O
retired; B-spline smoother 2D/3D and shift search, correlator/optimiser
are one sub-suite each.
Defect: ftiter%loop_lims with a low-pass limit ran the third dimension of a
volume from lhps(3) = 0, so image%corr(..., lp_dyn) on volumes
(symanalyzer, volcluster) saw half of the half-space and
image%mul(..., lp) left half of it unmultiplied; it is symmetric now, like the
second dimension. hclust (unused since April) cut its tree at the first N-k
merges in the order the chain found them, not the N-k lowest (0, 5, 100, 100.1,
100.2 into three gave {0,5},{100},{100.1,100.2}); a restore from git has to fix
that and keep the chain across merges (O(N**2)).
Removed, no callers: hclust and the unreachable hclust, hybrid and
refine branches of cluster_dmat; srchspace_map2D_io, its call in
cls_split (the class-to-cluster map is also in the project's cls2D/cls3D
cluster field) and SRCHSPACE_MAP_FNAME; online_var add_2, reset_mean,
serialize, unserialize; aff_prop work arrays Y, Y2, I, I2, tmp,
dA; ftiter set_hp, get_llp, comp_addr_phys_orig, the physical mode of
loop_lims (its upper bound was ldim(1), not the half-complex extent) and
the fields nothing read; bspline_smoother new(img) and four unread fields;
the dead test_CPlot2D and test_jpg_export (ImageMagick and a GUI resource
path); the jpg wrapper keeps its writers (writeJpg for images and volumes,
write_rgb_jpeg), its loaders, getters, setters, montage, the integer
writer and the unused C interfaces (one bound to the wrong C name) go. The
only self-tests left in production modules are test_flex_pcg_operator
(white-box model) and the five flex_gpu tests of the GPU platform entry.
Budget unchanged (25).
Build-time records (2026-09-25, Hans: the code map and the test inventory
are generated at build time, both gitignored). The first version (Codex,
in 84e05b43f) wrote both into doc/ as custom-command outputs. The code map,
no longer committed, first appeared during the build inside the inventory's
CONFIGURE_DEPENDS glob of doc/, so the make install after the fast gate
reported a GLOB mismatch, reconfigured and rescanned the tree (the build
seemed to start again). The inventory, still committed, was rewritten on
builds, and its writer escaped every pipe again, one more backslash per run.
Now both are written in the build tree and copied into doc/ only when their
content differs (a stamp marks the run), both are made at configure time
when missing, the inventory target waits for the map, and both generators
run with --quiet, so the build log shows only CMake's own "Regenerating"
lines. The hand-written
part of the inventory, the 26 verdicts and notes and the 136 retired rows,
moved to the committed test_review_record.md, which the generator reads (in
doc/refactoring_notes/completed/ since the review closed);
both live in doc/code_overview/ beside the code map, so the build does not
depend on the refactoring notes;
the inventory it generates is identical to the committed one apart from its
header. Checked in a scratch CMake project with the same block: no mismatch
on a fresh clone, nothing reruns on an unchanged tree.
10. Fast-tier performance¶
The 30 s budget will not be met by classification alone; the fast candidates have to be made cheap. After the Phase 0 timing run, for every fast candidate in descending order of wall time:
- Shrink the fixture. Box size, particle count, iteration count and number of orientations are chosen to exercise the code path, not to be realistic. A 64-pixel box and a few dozen particles test an FFT, a CTF or a correlation as well as a 256-pixel box with thousands.
- Build once, share. The suite builds one fixture per process and hands it to its tests; a test copies the immutable parts and writes only derived files. No fixture is cached in the build tree across runs.
- Stay in memory. Prefer in-memory images and stacks over writing and re-reading MRC files unless file I/O is what the test is about.
- No workflow in a unit test. A test that runs
cluster2Dorrefine3Dto check a utility is an extensive-tier test wearing a fast label; replace it with a direct check of the utility and let the workflow tests cover the integration. - No sleeps, polls or real timers. Test scheduling and watchdog policies as state machines with injected time.
- One thread. OpenMP teams inside a
ctest --parallelrun oversubscribe the machine and make timings meaningless; parallel correctness is tested inside the sub-suite of the code in question with an explicit small team (num_threads), as the discrete stack reader and the masks do. - Measure again.
ctest_budget.pykeeps the per-entry table beside the log on every build; a suite that grows is seen when it grows, and the 30 s label total is what fails the gate.
The same pass gives the print-only tests their assertions (section 4.4): the person who shrinks a fixture is looking at what the test computes and can state what the right answer is.
11. Command and naming contract¶
test= is the canonical selector:
simple_test_exec test=list
simple_test_exec test=unit_core
simple_test_exec test=unit_core suite=hash
simple_test_exec test=lib_stream
simple_test_exec test=simulated_workflow system=6vxx
Suites are named unit_<area> for the fast tier; extensive workflows keep
their descriptive names. Documentation, CTest, CI and child-suite launchers
use this form. Historical prg= examples are normalized during migration;
whether a temporary prg= alias is retained is a compatibility decision, and
it must not remain the documented interface.
Test identifiers describe the behaviour under test and do not encode the former executable shape. Existing stable IDs are retained unless misleading or colliding. Renames require a documented alias or a coordinated update of CI, scripts, implementation notes and user instructions.
12. Staged migration¶
| Phase | Workstream | Change | Exit gate |
|---|---|---|---|
| 0 | all | Timing and failure-path inventory. Build with --compile-tests in Debug and Release; run scripts/test_timing_run.sh in each (every standalone binary and every simple_test_exec case, each in its own directory under a timeout, single-threaded). Run scripts/test_review_dossier.py --timing ... to generate the dossiers and the inventory (section 8): proposed tier, failure path, run state and time, overlap candidates, fixtures, launchers, callers. Propose the grouped-module and commander map and the area review order. |
Every identity has a dossier, a run state (a measured time where it could run; otherwise timed out, crashed, missing fixture, unsupported capability or manual) and a proposed tier. |
| 1 | A | Scaffolding and a provisional gate. ctest after install in every compile_*.sh --compile-tests; ctest_budget.py; labels, timeouts, working directories, thread pinning. Register units as it is under the label provisional, not fast: it runs on every --compile-tests build and reports its time, but the budget is not enforced and nothing carries the fast label yet, because units still contains the socket, HTTP and child-process sub-suites that the fast admission rules exclude. Register the simulated workflows under workflow with their current checks. SIMPLE_CTEST_BUDGET is not yet set. |
compile_debug.sh --compile-tests builds and runs units green; its per-sub-suite times are known; CI still passes. Met 2026-09-22: commit 8b7dfd4d7; units 17.5 s through the gate (Debug, 1 thread). |
| 2 | C | Split units into hermetic area suites, then declare the fast gate. Reconcile the two routes into one implementation (the union of their sub-suites), move the sub-suite lists into the grouped modules of section 6.1, move forked process out to its own platform entry (decided, section 4.6) and confirm the remaining unit_ipc sub-suites are localhost-only and bounded. Register one entry per area suite (section 5.1 table); when every registered suite meets the admission rules, relabel them fast, drop the provisional entry, set SIMPLE_CTEST_BUDGET to the registered count, and turn on the 30 s check in ctest_budget.py. Shrink what is over budget. Remove the standalone simple_test_units and its CI call. |
Every area suite runs in one process and meets the admission rules; the fast label is under 30 s with ctest --parallel; the budget ratchet is armed; a failure names its suite. Landed 2026-09-22 (simple_commanders_test_class rewritten as area tables over a unit_suite type, suite= input, SIMPLE_UNIT_ORDER=reverse, forked_process under platform, SIMPLE_CTEST_BUDGET=19, GATE_DECLARED=yes). Met 2026-09-22: the --compile-tests build passed 7/7 in 3.3 s real; every suite also passed with SIMPLE_UNIT_ORDER=reverse, so no sub-suite leaks state into its neighbours in either direction. |
| 3 | B | Review everything else. The other 149 identities, area by area (section 9): demote to a named library suite or the workflow gates, keep as manual, merge, delete or retire; the 53 two-route identities and 14 footprint clusters resolved to one implementation each; deletions applied with their retired-tests rows and coverage accounting. |
Every identity has a verdict naming its destination; no pair or cluster retains two implementations of the same coverage. Met 2026-09-24: all 92 identities of the inventory's area tables have a verdict with reviewer and date, and the retired-tests table has 136 rows (section 9.7, from the geometry batch to the wrap-up). |
| 4 | B + D | Build the library suites. Area by area: the survivors move into the grouped module, gain the assertions their verdicts require, and are registered as one lib_<area> entry under library; standalone binaries removed as each suite completes. The first suite (lib_fft or lib_geometry) is the pilot for the fused extensive shape. |
Each library suite runs in one process nightly with a recorded time; its members' binaries are gone. Done within Phase 3, 2026-09-24: the review built each library suite as it reviewed the area, so Phase 4 had no pass of its own: lib_reconstruction, lib_cart_align3D, lib_heterogeneity, lib_single, lib_stream (section 5.2.1). The pilot named here never existed: the last members of lib_fft, lib_geometry and lib_masks became fast sub-suites and a program. Every standalone binary is gone. The recorded nightly time of each suite comes with the first night of the runner (Phase 5). |
| 5 | D | Simulation-truth gates and the nightly runner. simulated_workflow, single_workflow and mini_stream compare against the generating model (FSC to the truth map, pose agreement) with declared floors; the nightly ctest -L "library|workflow" run and its archive on the dedicated machine. |
The nightly run completes unattended and reports per-suite times and per-workflow metrics against floors. Assigned to Ruben 2026-09-24 (Hans): the truth gates and the design and code of the nightly runner, doc/refactoring_notes/phase5_workflow_gates_and_nightly_runner_handover.md. |
| 6 | B | Mother suites, platform and socket cases with explicit isolation and launcher policy. | Parent suites launch simple_test_exec children with full accounting; platform cases skip or register predictably and cannot hang the fast gate. Met by the review, 2026-09-24: no mother suite is left (those that launched child cases were merged into in-process sub-suites or deleted, section 9.7); the platform entries (forked_process, and coarrays, flex_gpu and openmp_offload when CMake has the capability) carry their own label and timeout and are outside the fast gate; the socket role programs are deleted and the IPC socket tests in unit_ipc are bound to localhost. |
| 7 | B | Retire the glob. Remove the standalone executable glob, switch CI to ctest -L fast plus the platform jobs, normalize documentation, delete stranded per-test commander types, routers and UI entries. |
A clean --compile-tests build produces only simple_test_exec; registry consistency passes; CI uses no standalone Fortran test executable. Met 2026-09-24: the glob and production/tests went with the utils review; the stranded fft and geometry test categories and the dead nspace_commander went with the wrap-up; scripts/check_test_registry.py passes and runs on every --compile-tests build; CI runs ctest by label and name (section 9.7). |
Complete a vertical area slice (baseline, extract, assert, shrink, route, switch callers, delete duplicate) before starting the next. Do not copy every program into a module and leave both systems standing.
13. Migration rules for individual tests¶
For each test, in this order:
- Confirm it has a review verdict (section 9). A
deleteorretireends here: remove the sources, the UI entry, the router case and the callers in one commit, and add the retired-tests row. Amergemoves its unique checks into the target before this identity is removed. - Record its unchanged baseline and its Phase 0 wall time before editing.
- Identify all source and script callers by test executable name.
- Compare any existing
simple_test_execimplementation with the standalone program line by line and select or merge the authoritative behaviour. - If it is kept (fast or extensive) and has no failure path, write the assertion it was implicitly making (what would have been wrong if the printed numbers were wrong). A manual tool needs none.
- Move the body into its grouped module; keep suite-specific helper modules that already express ownership.
- Remove raw
get_command_argumentorparse_oldschoolownership from the body; register required keys in the test UI and consume the parsed command state through the established lifecycle. - Make the procedure return on success and accumulate failures; keep process
lifecycle in
simple_test_exec. - Apply the performance actions from the inventory row; re-measure.
- Route it through its area suite without adding a one-test commander type.
- Change CTest, CI, scripts and documentation to the suite or
test=<id>. - Compare exit status, assertions, products, tolerances and expected-failure behaviour with the baseline.
- Delete the standalone program and any duplicate commander body in the same completed slice.
Mechanical extraction, adding an assertion, and shrinking a fixture are three separate commits per test, so each can be reviewed and reverted alone. None of them changes what the production code does.
14. Risks and mitigations¶
| Risk | Level | Mitigation |
|---|---|---|
| The fast gate is green but hollow because print-only tests were admitted. | High | Admission rule 1; the failure-path column in the inventory; Phase 1 registers only assertion-bearing tests. |
| The 30 s budget is missed and quietly raised. | High | The budget is a ratchet checked by ctest_budget.py on every --compile-tests build; raising it is an owner decision recorded here. |
| Fused suites leak state between tests (RNG, module variables, cwd, open units) and produce order-dependent results. | Medium-high | Lifecycle rule 7; during Phase 2 run each area suite in a few fixed alternate orders (its table order, reversed, and with any I/O or IPC sub-suites first) and require identical results, since the sub-suite procedures have different signatures and are not trivially shuffled; keep the focused selector so a failing test can be reproduced alone. |
| Duplicate implementations have diverged; selecting one loses coverage. | High | Review verdicts with the dossier's unique-coverage list; merge missing assertions before deleting either path (units is the known example); coverage accounting per batch. |
| The review deletes tests that were the only coverage of something that matters. | Medium-high | Section 9.5: every module that loses all coverage is listed in the batch summary and accepted with a reason or answered by a modify/merge; deletions go through Hans. |
| The review stalls on hard cases and blocks migration. | Medium | The investigate verdict parks a test without blocking its batch; ten-minute pace with the dossier. |
| Conversion from program to subroutine changes initialisation, working directory, logging or teardown. | Medium-high | Baselines; explicit lifecycle ownership; a file-producing test in the pilot. |
A callable test calls stop, error stop or THROW_HARD on a path that should accumulate, ending the suite early. |
Medium | Audit terminal calls in extraction; THROW_HARD only for missing fixtures; expect-abort tests isolated. |
| Mother suites lose crash isolation when child programs disappear. | High | Launch the same simple_test_exec binary as a subprocess for each isolated case. |
| Coarray/MPI/GPU tests compile but run under the wrong launcher or build. | High | Launcher and capability checks in CMake/CI; compile guards retained. |
| Simulation-truth floors are set from one run and become flaky. | Medium | Set floors with margin from several seeded runs; declare the seed; treat a floor change as a reviewed commit. |
| The overnight machine drifts (compiler, libraries) and the extensive tier fails for reasons unrelated to SIMPLE. | Medium | The archive records compiler and host; the runner does a clean build each night. |
| Grouped modules become dumping grounds. | Medium | Group by domain; keep cohesive suite helpers; split by responsibility when a file becomes hard to review. |
CTest and test=list drift apart. |
Medium | The registry-consistency check; stable canonical IDs. |
ctest --parallel oversubscribes the machine and the budget is missed for scheduling reasons. |
Medium | OMP_NUM_THREADS=1 on fast entries; NJOBS at half the cores; explicit teams on the library suites and RUN_SERIAL on the workflow gates. |
| Removing many executables breaks scripts and historical validation packages. | Medium | Search all callers, update active ones atomically, document the command translation; historical evidence is preserved, not pretended runnable. |
15. Validation plan¶
Compilation and runtime validation are user-run unless separately authorised.
15.1 Static¶
- Every inventory row carries a verdict with reviewer and date, and resolves to exactly one final callable implementation or a retired-tests row.
- Every test UI entry has one reachable commander/module path.
- Every CTest registration refers to a valid suite or
test=<id>, and every registration has a label, a timeout and a working directory. - Every fast registration is assertion-bearing (grep for
assert_,tests_failedor an equivalent typed check in its module). - No active invocation of a removed
simple_test_*executable remains. - Test implementation modules contain no program units and no
stop. - No migrated test has a new dedicated commander type whose only purpose is to call one procedure.
BUILD_TESTS=OFFsource filtering excludes all test-only modules (in part, by design: section 16, criterion 8).git diff --checkand non-compiling syntax diagnostics pass for edited files.
15.2 Per-test behavioural¶
For every migrated test compare before and after: process exit status; assertion and failure counts; expected normal-stop or failure marker; required inputs and defaults; output files and retained artifacts; numerical values at the pre-existing tolerance; cleanup and working-directory behaviour; launcher, process count, threads and device requirements; skip behaviour when a capability or fixture is absent; and wall time against the inventory.
15.3 Integration¶
compile_debug.sh --compile-testsandcompile_clean.sh --compile-testsrun the fast gate green in under 30 s on the reference Mac, in Debug and Release, andctest_budget.pyreports no violation.simple_test_exec test=listcontains every suite and canonical test ID once.- Each area suite passes in one process and each of its tests passes alone.
- A deliberately failing assertion in one test fails its suite, fails the gate, and is reported by name.
- Mother suites survive an intentionally failing child and return a failing aggregate status.
- Coarray and MPI invocations use the required launcher and process count.
- The nightly
library|workflowrun completes on the dedicated machine and its archive holds per-suite times and per-workflow metrics against declared floors. - A
BUILD_TESTS=OFFbuild contains no test executable or test-only object (in part, by design: section 16, criterion 8); a--compile-testsinstall containssimple_test_execand no standalonesimple_test_*binaries.
16. Acceptance criteria¶
The project is complete when:
compile_*.sh --compile-testsbuilds and then runs a fast gate that finishes under 30 s on the reference Mac, in which every entry can fail.- The extensive gate runs nightly unattended: the library suites, one process per coherent part of the library, and the simulated workflows reporting resolution and pose-recovery metrics against declared floors.
simple_test_execis the sole public Fortran test executable, and all supported former standalone programs are callable procedures in grouped modules run by area suites.- Every test has one authoritative implementation and a recorded review verdict; every removed test has a retired-tests row.
- CTest, CI, documentation and child suites use suites or
test=<id>. - Process isolation and special launchers are preserved where section 5.3 requires them, and nowhere else.
- The
production/tests/simple_test_*.f90glob is gone. - Test-only code is absent when
BUILD_TESTS=OFF. - The process budget and the time budget are enforced on every build.
- Compilation and runtime checks not actually observed are listed as outstanding rather than claimed as passing.
Where the criteria stand (2026-09-24).
- Met: 1 (thirteen fast suites, about 5 s real on the reference Mac in
Debug, every one assertion-bearing); 3; 4 (section 12, Phase 3); 5 (CI by
label and name, documentation and handovers by
test=<id>, no mother suite left); 6; 7; 9 (SIMPLE_CTEST_BUDGETat configure time,ctest_budget.pyand the registry check on every--compile-testsbuild). - Open: 2, the nightly run with the truth gates (Phase 5, Ruben).
- Not fully attainable, by design: 8.
BUILD_TESTS=OFFdropssimple_test_execwith its commanders, routers and API module, every*_testermodule (src/CMakeLists.txt) and every CTest registration (production/CMakeLists.txt). Three kinds of test code stay in the library: the test UI registry (src/main/ui/simple_test,simple_ui_test_group), becausesimple_uibuilds the test-program table beside the production ones;simple_test_utilsinsrc/utils; and the white-box self-tests that stay inside production modules because they read private components (for exampletest_flex_pcg_operatorinsimple_flex_pca_pcg, section 9.7, heterogeneity). Separating them would take a conditional-compilation layer aroundsimple_uiand the self-tests for little gain; no production executable has a test entry point. - Outstanding checks (criterion 10), carried to the Phase 5 handover when
the plan was archived (2026-09-25): a green
--compile-testsbuild of the tree with the in-module self-test batch and the build-time records (Hans); one run ofsimple_test_exec test=pcg_recon, whose seed changed on 2026-09-25 (Hans; it also gives its runtime); a build of this tree without--compile-tests(the compile scripts then configureBUILD_TESTS=OFF), to confirm that it links without the test-only sources (Hans); the offload branch ofsimple_openmp_offload_testerin an offload build (Cyril); the first night of the runner, with the runtimes of the library and workflow entries (Ruben).
17. Non-goals¶
- An X-style
validatetier: a registry of real datasets and arms, blessed platform-keyed baselines, a report directory and gated blessing. Too costly for SIMPLE; the extensive tier gates on simulation truth instead. - Adding test-specific commander types.
- Rewriting numerical algorithms while moving their tests.
- Forcing all tests to share one procedure signature when their dependencies differ.
- Running the extensive tier, or any subprocess-isolated case, inside the fast gate.
- Converting NICE/Python tests or external oracle analyzers into Fortran.
- Making network-dependent tests part of any registered tier.
- Replacing CTest, the SIMPLE command-line system or
simple_test_utilsas a prerequisite for consolidation. - A generated universal test manifest before executable unification shows it is needed.
18. Effort and delivery shape¶
This is a large project, because two of the four workstreams are not mechanical:
- Phase 0 and workstream A (the timing inventory and the fast-gate scaffolding over today's qualifying tests): days, and immediately useful.
- Workstream B (review of 149 identities, then unification of what survives; 53 two-route identities and 14 footprint clusters to reconcile): three to six weeks of area-by-area work, dominated by the review and by preserving the meaning and launch behaviour of what is kept. The review itself is a few weeks of owner and reviewer time, an afternoon per area, spread across the migration.
- Workstream C (splitting
unitsinto area suites, reconciling its routes, timing and trimming to 30 s, settling the socket and child-process sub-suites): about a week, and it delivers the whole fast gate. Assertions for print-only tests are needed only for the ones the review keeps for the extensive tier, and are written as those tests are migrated in Phase 4. - Workstream D (the library suites as their areas are reviewed, the simulation-truth gates, the nightly runner): a few days per library suite once its review is done, one to two weeks for the first three workflow gates; more as workflows are added.
The order of value is A, then C (the fast gate is complete after it), then the review and migration of the rest area by area as those areas are touched in ordinary development, so the migration rides on work that is happening anyway.