Skip to content

The PCG reconstruction backend: what it is today

Status: current as of 2026-09-10. Present tense only. This page describes the backend as it runs on master; it does not argue for it. The binding contracts are the policy documents (doc/policies/3D/reconstruct3D_pcg_policy.md, the PCG sections of refine3D_policy.md and refine3D_auto_policy.md, automasking_policy.md); when they and this page disagree, the policy is right and this page needs a fix. How the backend got here, decision by decision, is pcg_decision_log.md; the full experiment record is pcg_priors_history.md. New decisions are appended to the log, and the affected paragraph here is rewritten, never annotated.

1. What it is

rec_backend=pcg replaces the gridding density quotient with the weighted least-squares reconstruction: it solves the normal equations of the CTF/sigma-weighted Fourier-slice projection operator by preconditioned conjugate gradients, on a real-space support, with the optional ML prior inside the operator. The particle pass is done once per iteration into raw accumulators; the operator is then applied in kernelized (Toeplitz) form, so the cost of a CG iteration is a handful of FFTs and independent of the particle count. pcgop=kernel is required in production; the matrix-free operator is the exact reference and exists for tests.

Both backends produce the same kinds of products (even/odd/merged state volumes, an _unfil base pair under ml_reg=yes, FSC/cFAR and the resolution text), at the same amplitude convention (the data-quotient convention; deapodization is inside the solver), and feed the same assembly-owned nonuniform (NU) competition afterwards. The PCG backend carries no prior of its own beyond the FSC/SSNR precision P_tau; every NU-derived prior that was tried inside the solve has been removed (decision log, 2026-08-27, 08-29, 09-06).

2. The two solves

Per state and per half, from one particle accumulation:

  • Base solve (H + lambda_0 I) x = b, lambda_0 = PCG_LAMBDA = 1e-3 (absolute Tikhonov ridge). Its half pair is the _unfil pair: the FSC, cFAR and resolution authority, and the seed of the NU candidate bank.
  • Regularized pair (2026-09-14): the closed-form optimum of the diagonal model of (H + P_tau + lambda_0 I) x = b, derived from the base solution on the replayed operator (same raw accumulators, the FSC/SSNR shell-diagonal precision P_tau from tau, hp): each Fourier coefficient of the base map on the padded lattice scaled by (rho+floor)/(rho+floor+P_tau), the ratio of the regularized and base preconditioners, voxelwise (shrink_by_ml_prior); maxits_ml (default 0) optional coupled PCG iterations from that start (solve_regularized_half, 2026-09-16: refinement-invisible on bgal, kept as a knob). Its pair is the shipped map and the finest NU bank member. It replaced the two-iteration replay solve, which was prior-dominated over most of Fourier space (prior/data 400-600 wherever the FSC reaches zero inside Nyquist), started from a shell-isotropic FSC-shrunk base that was rejected as worse than zero on every anisotropically sampled dataset, and reported an L2 residual above 1 that was mostly noise the prior refuses to fit -- in the preconditioned norm both solves had always converged alike (decision log).

Starts. The base solve starts from zero. There are no cross-iteration warm starts: nothing from a previous iteration's half maps enters a solve (2026-09-10), and no production solve starts nonzero (2026-09-14).

Budget. maxits_pcg=2, rtol=0 (exactly two iterations) in refinement; the original-sampling final reconstruction uses at least FINAL_PCG_MAXITS_FLOOR = 5. Two iterations from a state-free start is the calibrated regime (simulated data: two iterations beat gridding, beyond five the residual moves but nothing interpretable in the map does). A solve from a nonzero start that loses positive-definiteness is retried once from zero; a solve from zero that loses it is fatal.

Preconditioner. The sampling-density diagonal 1/(rho + floor + P_tau) with a shell-relative floor (RHO_FLOOR_FRAC). With the solvent prior the mean of its real-space ridge over the solve domain is folded in as a constant (fold_solvent_ridge_into_precond, 2026-09-22): the ridge lambda_s (1-w(r)) has no Fourier-diagonal representation, but its mean is one, and without it CG on the ridge system ran with a preconditioner built for the prior-free operator and was left at RESID 0.2-0.3 (bgal, lambda_rel 1.2, 2 iterations) against 0.04-0.06 prior-free.

3. The support

The solve is constrained by a real-space support P: the system solved is the projected P H P u = P b on the hard domain window > 0, and the shipped map is window * u (PCG_HARD_SOLVE_SUPPORT). Which window:

  • automsk=no: the soft spherical support at msk_crop, the same mask3D_soft the gridding restoration applies after deapodization. Base and replay both.
  • automsk=yes: the conservative density envelope -- automask3D of the lag-one reference volume at envmsklp (20 A), Otsu core, largest component, spherical dilation by binwidth layers, outward cosine skirt of edge voxels -- constrains BOTH the base and the replay once a prior reconstruction exists. The first reconstruction bootstraps the base on the sphere and builds the replay support from its own pair.
  • pcg_mskfile=<vol>: an explicit [0,1] volume, development mode, constrains every solve regardless of automsk and is reported as the state support.

The dilation has a shared physical minimum, ENVMSKWIDTH_A_MIN = 7.5 A (binwidth = max(binwidth, ceiling(7.5/smpd_crop)) whenever the envelope is in use; an explicit binwidth wins in either direction). The support is the one place a mask enters the estimate; no PCG product is multiplied by any mask after the solve (postprocess, _pproc, _mirr included), and the support-provenance sidecar <vol>_pcg_support.txt (solve_support=density|sphere, solve_kind=base|regularized|mixed|gridding|gridding_regularized) records what the shipped pair carries.

Opt-in soft solvent prior (pcg_solvent=yes, default no). None of the supports above sees solvent finer than the scale they were drawn at (the envelope's dilation ring and skirt, cavities and gaps below ~20 A). With pcg_solvent=yes the BASE solve is done twice per half on the same accumulators: first prior-free, then again from zero with the same budget after a real-space, position-dependent ridge has been installed, (H + lambda I + Lambda_s) x = b, Lambda_s = lambda_s (1 - w(r)), lambda_s = pcg_solvent_lambda x data_scale (the same reference the relative ridge uses; not given = estimated per state and iteration, see below): a Gaussian prior with position-dependent variance, the real-space twin of P_tau. The prior-free pair's FSC=0.143 sets the smoothing scale and each half's own prior-free map yields its protein weight w(r) in [0,1] (simple_pcg_solvent_sidecar): smoothed absolute density at max(8 A, 2 x FSC0.143), Otsu threshold inside the production support, logistic of the statistic around the threshold with the solvent class's spread as width. Nothing is zeroed or masked: where the data term is strong the prior is irrelevant, where it is weak solvent is pulled toward zero, and a misassigned voxel is over-regularized rather than deleted. Estimate on the base pair, apply to the prior'd pair (2026-09-21, Hans): the prior-free pair stays the base pair, i.e. the FSC (the cap, the band and the resolution claim), the NU competition and its solvent-calibrated whitening and evidence null, the _unfil pair and everything postprocess reads; the prior'd pair, written beside it as _even_solvent/_odd_solvent, is the base of the closed-form replay (P_tau from the base pair's FSC) and the pair the NU label field is applied to when the _nu_filt references are composed, so every reference voxel carries the prior. An FSC of the prior'd pair over-estimates (the per-half weights are 97-99.6% correlated, a shared soft envelope with no phase randomization; exp_gate 2026-09-21: 3.64 against 3.88 A prior-free, the stage-8 crossing pinned at the crop Nyquist) and is not computed; the whitening of the NU unary is the MAD of even minus odd per radial shell over the support, which the prior would collapse. The strength (2026-09-22): when pcg_solvent_lambda is not given it is chosen by cross-validation with the NU objective over the production support, in closed form on the prior-free pair (estimate_solvent_prior_lambda, simple_pcg_solvent_sidecar): each half shrunk voxelwise by h/(h + lambda data_scale (1-w)), h the real-space diagonal of the data operator (get_realspace_diagonal, the mean of D over all native shells), scored by the whitened Huber cross-half prediction error against the prior-free other half (image%nu_objective), on the grid 0.1-20 with one parabolic step in log lambda; the table, the chosen value and edge/flat flags are logged (PCG SOLVENT PRIOR LAMBDA; flat means the weight map is the limit), and the one real re-solve runs at that strength. pcg_solvent_check=yes adds the same grid by real re-solves, with their residuals (shared-memory and distributed paths), to validate the closed form; the approximation to watch is that the real ridge acts more on high frequencies of a solvent voxel than the scalar shrink. An explicit pcg_solvent_lambda bypasses the estimate (provenance auto|set). maxits_ml stays at its default 0: the replay is the Wiener shrink of the prior'd base map (coupled replay iterations, when requested, carry the same ridge). The support is untouched. It runs under any automsk setting; in abinitio3D the key rides with the PCG stages only. Validation: one PCG SOLVENT PRIOR line per half (smoothing scale, threshold, width, solvent fraction, mean weight, coefficient), the prior-free FSC, the even/odd weight correlation and solvent-fraction gap, the KIND=pre (prior-free) and KIND=base (prior'd) solve summaries, and the weight volumes pcg_solvent_weight_state01_even.mrc, _odd (one pair per state, overwritten every iteration). With pcg_solvent=no no prior code runs and the strategy is bit-identical to the version without it.

4. The FSC and the envelope

automsk=yes implies envfsc=yes on both backends (derived in parameter validation, logged when it overrides a no). The FSC is then obtained in one of three ways, and the way used is REPORTED on every evaluation as >>> FSC MODE in the log and in the resolution text:

  • spherical support from the reconstruction, no envelope, no correction (envfsc=no);
  • density envelope applied post hoc to the halves with the phase-randomized solvent correction (gridding, envfsc=yes);
  • halves estimated on a support envelope, no post-hoc mask and no correction (PCG under automsk=yes, or pcg_mskfile). The envelope contribution to the FSC is deliberately NOT removed; a common window on both halves can contribute correlated power, and the policy is to say so rather than hide it.

The automask3D_stateNN.mrc artifact is written on the envfsc path for its other consumers (postprocess of non-PCG products, the abinitio final rec).

5. NU filtering and the evidence envelope

The NU competition is assembly-owned and identical on both backends (simple_nu_state_filter, 2026-09-16, the nu_refine shell walk retired; the generated ladder withdrawn 2026-09-18): the static ladder [20,15,12,10,8,6,5,4] A capped at fsc/1.5 of the base pair, with the ML-regularized pair as one more member beside the finest retained rung, competing with it at zero prior cost (ml_reg=yes) -- the ed36eb4c abinitio3D machinery, the only NU mechanism since 2026-09-18. The matching low-pass handoff is the finest member of the bank: the finest rung under the fsc/1.5 cut, or the regularized pair once its FSC=0.143 is at or beyond the ladder's finest rung (2026-09-19; the pair joins the bank only then). The NU objective always runs on the spherical mskdiam support.

Under automsk=yes the NU evidence envelope (nu_envmask3D_stateNN.mrc, regenerated every competition from the live evidence) is the mask that controls the filtering: the filter-field background is its complement, and background voxels take the coarsest bank candidate, so excluded density (detergent, disordered belt) reaches the matching references heavily low-pass filtered rather than removed. References are never multiplied by an envelope. The evidence envelope's null model has two regimes, keyed on how the base pair was solved:

  • spherical base pair (gridding; PCG bootstrap): robust median + nu_msk_sig MAD of the margin over the observed support; valid while solvent holds the majority;
  • envelope-constrained base pair (PCG automsk=yes, pcg_mskfile): the estimator has removed the far solvent, so the null is designated by Euclidean geometry -- median/MAD on the density envelope's dilation ring at full weight of the base support (set_nu_evidence_null_shell); labels are free on the observed density envelope and fixed solvent outside it, so the evidence envelope is nested inside the density envelope; valid while the shell is populated.

If the null is invalid or the envelope empty, the density envelope itself is armed as the background (EVIDENCE FALLBACK), and the provenance string records which envelope ran.

6. Where it runs

  • refine3D / refine3D_auto / refine3D_states: rec_backend=pcg selects the distributed master (execute_rec3D_pcg_distributed_master): workers accumulate and write raw statistics, the master reduces, solves both halves concurrently, runs the FSC, the replay and the NU competition, and ships. Trailing reconstruction (trail_rec) keeps accumulator chains; the bootstrap iteration blends with the previous pair at the update weight exactly as the gridding bootstrap.
  • abinitio3D: the PCG backend engages from PCG_REC_START_STAGE = 3; stages 1-2 are gridding. Stage policy sets envfsc/automsk at the last stage (AUTOMSK_STAGE = ENVFSC_STAGE = NSTAGES). In the NU stages the NU handoff sets the matching low-pass, capped at the ladder's hard fine bound LPSTOP_BOUNDS(1) (4.5 A; a coarser explicit lpstop is retained), with the stage limit promoted to the previous stage's FSC=0.5 crossing when that is finer; NU stages run their full iteration budget (minits = maxits).
  • Final reconstruction (bootstrap_rec3D, both workflows): image-power sigma seed, a gridding ML bootstrap map carrying the workflow's filt_mode/automsk (the residual sigmas depend on the regularization of the reference they are scored against), one residual sigma pass, then the shipped PCG map at the native box: cold, at least five iterations, filt_mode=none, automsk inherited (so its support matches the refinement's), postprocess=yes (the regularized map B-sharpened and Butterworth-filtered at FSC=0.143 with no second FSC weighting, 2026-09-26), with the gridding bootstrap map passed as vol<state> so the density envelope constrains the base pair too and the reported FSC is estimator-constrained (2026-09-11; previously the final base pair bootstrapped on the sphere). refine3D_auto's startup reconstruction likewise receives the initial volume as vol1.
  • Shared-memory (nparts=1) runs the same worker and master in one process (since 2026-09-27; the separate execute_rec3D_pcg_shared route is retired), trailing included.

7. Reading a run

Per half and solve, one summary line:

>>> PCG DISTRIBUTED | STATE= 1 | HALF=even | KIND=base | N=  8416 | ITS= 2 | INIT= 1.000E+00 | RESID= 6.100E-02 | TIME=   1.9 s | STOP=fixed_iterations

INIT is the L2 relative residual of the start (1.0 from zero), RESID the final one, MRES the final residual in the preconditioned norm -- the one CG drives, and the one on which the two kinds are comparable (2026-09-14; the L2 residual of a prior-dominated system is mostly noise the prior refuses to fit). Healthy base values on PfCRT at box 140-160: 1 -> 0.06-0.09 (L2), MRES 0.2-0.3 after two iterations. The KIND=ml line is the closed form (maxits_ml=0, default): ITS=0, STOP=closed_form, RESID/MRES its residuals against the replay system, a diagnostic of the support coupling it leaves out, not a convergence measure. With maxits_ml>0 INIT is the closed form's L2 residual, ITS=maxits_ml, MRES the final one, and a CF MRES a -> b line gives the FSC between the start and the solved map inside the band (2026-09-16).

Other lines to grep: PCG SOLVE SUPPORT (which support, and why), FSC MODE, NU ENVELOPE OCCUPANCY, NU DILATION RING OCCUPANCY (how much of the density envelope's dilation ring the evidence labels signal -- the number to consult before tightening binwidth), NU NULL SHELL GEOMETRY (envelope/support Dice, ring retained at full weight), NU BACKGROUND (evidence envelope or fallback), NU BANK CAP (the static ladder's FSC cap and retained count), NU LOW-PASS ASSIGNMENTS, NU filter promoted matching low-pass, PCG BEYOND-BAND EXCESS (post-band RMS >= 10x the band-edge shell; the regression signal for solver defects), RECONSTRUCTION MASTER PHASE (wall time, thread-seconds, peak RSS). Per-half solve diagnostics (residual history, data scale, effective lambda, prior statistics) go to reconstruct3D_pcg_stateNN_<half>_<kind>*.txt sidecars.

The stage-6 take-off is the thing to check on an abinitio3D run: at the first NU iteration the base pair should support a few percent of the mask at ~8 A, the matching low-pass should be promoted below 6 A, and FSC-0.143 should drop from ~9 to ~4.5 A within the stage. If the matching low-pass sits at the stage ladder value, the base pair is not carrying fine-shell evidence.

8. Cost

At PfCRT scale (16.8k particles, box 140-160, 10 parts x 8 threads) the PCG master phase is ~16 s per refinement iteration against gridding's ~4 s, and the refine3D worker step is ~12 s against ~8 s because the PCG workers write the raw accumulators. On a 133-iteration abinitio3D that is 7460 s against 4820 s. Iterations of CG themselves are cheap (~1 s per half per iteration under the kernel operator); the cost is the accumulation and the master's reduction, not the solve.

9. Tests

simple_test_exec test=pcg_recon (operator identities, support contract, window band regression), test=pcg_frac_update (trailing accumulator arithmetic), test=rec3D_backends (gridding vs PCG on a project with poses, with optional truth). doc/implementation_notes/automsk_yes_code_review.md lists the nine-case validation matrix for the envelope/NU changes of 2026-09-09.

10. What is deliberately not there

No solvent prior, no Wilson prior, no NU-evidence prior in the solve, no auto-lambda; no cross-iteration warm starts; no post-hoc mask on any PCG product; no envelope multiplication of matching references; no AMSK_FREQ regeneration cadence. Each was implemented, measured and removed; the decision log has the date and the evidence for every one.