Importance Sampling and Fractional Update Policy¶
This document records durable workflow contracts for sampled particle updates,
probabilistic candidate sampling, fractional class-average restoration, and
trailing reconstruction in abinitio2D, cluster2D, abinitio3D, and
refine3D. It is policy, not a line-by-line implementation map.
1. Core Model¶
SIMPLE has two sampling layers that must remain separate:
- outer fractional-update sampling chooses which particles participate in the current iteration
- inner importance sampling chooses which reference, orientation, or in-plane candidates are explored for those participating particles
The outer subset is recorded in the project through sampled and updatecnt.
Downstream restoration and reconstruction consume that recorded state. They must
not infer participation from the nominal command-line update_frac alone.
Probabilistic pre-alignment is a sample-once-and-reuse path: the pre-alignment commander chooses the outer subset, probability-table workers reuse it, and the matcher reuses it again for the hard particle update.
2. Ownership¶
simple_commanders_abinitio2D.f90 owns abinitio2D orchestration: defaults,
stage execution, final fill-in, and final class-average generation.
simple_abinitio2D_controller.f90 owns the 2D stage policy: NSAMPLE_DEFAULT_2D,
nsample override handling, stage-local update_frac, search-mode transitions,
and the rule that stage 1 may sample particles without fractionally restoring
previous class averages.
simple_commanders_abinitio.f90 and simple_abinitio_controller.f90 own 3D
stage scheduling: dynamic update_frac, fillin, frac_best, balance,
trail_rec, and transitions between early prob_neigh modes, prob, and
late prob_neigh.
simple_matcher_smpl_and_lplims.f90 owns the shared outer subset-selection
helpers for 2D and 3D. This is where full update, random sampling,
update-count-biased sampling, class-balanced sampling, fill-in sampling, and
subset reproduction are dispatched.
simple_oris_sampling.f90 and simple_oris_getters.f90 own the bookkeeping:
sampled, updatecnt, exact subset reproduction, global realized update
fraction, and class-local realized update fractions.
simple_commanders_prob.f90 owns probabilistic pre-alignment orchestration:
sampling the outer subset once, writing it to the project, running table
generation, aggregating probability-table outputs, and writing the assignment
artifact.
simple_eul_prob_tab*.f90 owns inner candidate importance sampling. These
modules may sample references, orientations, neighbors, or in-plane candidates
inside the active particle subset, but they must not choose a new particle
subset.
simple_strategy2D_matcher.f90 and simple_strategy3D_matcher.f90 own
particle-domain search on the active subset, assignment consumption, pose or
class updates, sigma updates during search, and writing partition-local
reconstruction or class-average inputs.
The classaverager modules own 2D class-average restoration and assembly.
commander_volassemble owns 3D volume assembly and trailing reconstruction.
These layers consume sampled-update state; they do not own particle selection.
3. Bookkeeping Contracts¶
sampled marks the current sampling round. All particles with the latest
sampled value belong to the current active subset.
updatecnt tracks cumulative update history. Count-biased and fill-in paths use
it to prefer under-updated or never-updated active particles.
sample4update_reprod is the only correct way to reuse a previously selected
probabilistic subset. A probability-table worker or downstream matcher must not
silently resample when a probabilistic pre-step has already sampled the subset.
get_state_update_fracs returns state-local realized update fractions for 3D
trailing from the current sampled round, state labels, and particles with
updatecnt > 0. get_update_frac remains available for callers that need the
legacy global summary.
get_class_update_fracs returns per-class realized update fractions for 2D
class-average carry-over. It uses active particles, current class assignments,
the latest sampled round, and updatecnt > 0.
The nominal update_frac is a target used by sampling. The realized fraction in
simple_oris is the downstream restoration and trailing contract.
In-plane representation during probabilistic search¶
inpl_cont does not change either sampling layer or the assignment-file
schema. With inpl_cont=yes, callback-style local angle/shift profiling is
replaced by joint (sx,sy,rotind_frac) optimization, but inner importance
sampling still carries a canonical rounded in-plane index. Its stored shift is
expressed in that rounded-index frame, and no fractional angle is persisted in
the probability table.
When a sampled class or state/projection candidate is chosen for joint
profiling, its sampled inpl is not the continuous seed. The selected
class/state/projection remains fixed, the shift is converted to the native
particle frame, and one all-angle discrete evaluation at that exact shift
selects the profiling seed. The continuous route does not search a 5-by-5 grid
of alternative shifts.
The hard-assignment matcher owns the durable continuous result. In 3D, the
rounded probability-table assignment is authoritative: the matcher recovers
its native shift and reruns the joint optimizer locally within plus or minus two
cells of that in-plane index, without another global all-angle selection. It
persists fractional e3, integer inpl, shift, and score for the same final
pose. Valid non-improving work retains the incoming discrete pose with a
consistent re-scored objective; invalid work leaves the assignment untouched.
The 2D durable path retains its global all-angle seed selection. Neither path
may enter the legacy callback route. These policies do not alter sampled,
updatecnt, top-K support, assignment probabilities, or fractional-update
weighting.
4. Abinitio2D and Cluster2D¶
abinitio2D uses a fixed run-local target sample size:
- default:
NSAMPLE_DEFAULT_2D = 200000 - override:
nsample=<integer>
The stage controller converts that target into:
update_frac_2D = min(1.0, real(min(nptcls_eff, nsample_target_2D)) / real(nptcls_eff))
where nptcls_eff is the number of active particles with state > 0. If the
target covers almost all active particles, the stage command omits
update_frac and naturally becomes a full update.
Current stage policy:
- stage 1 uses the sampled-update machinery when needed, but fractional class-average carry-over is disabled
- while
startit == 1,sample_ptcls4update2Dkeeps the initial subset sticky by reproducing it after the first random draw - later non-probabilistic iterations use
sample4update_cnt, which is stochastic but biased toward particles with lowerupdatecnt - probabilistic stages use
prob_align2Dto sample once, thenprob_tab2Dandcluster2D_execreproduce the same subset - staged
fillin=yescurrently acts as a full-assignment coverage guard. It requires active particles to have assignments before convergence, while particle selection still follows the normal sampled-update path - staged
abinitio2Drefinement uses sampled SNHC (refine=snhc_smpl) for stages 1-2. From stage 3 onward,refine=probuses dense probabilistic assignment;refine=prob_snhcuses sparse probabilistic SNHC until the final staged invocation, which uses denserefine=prob - when staged updates were sampled,
abinitio2Dthen runs a separate terminal dense greedy all-particle pass withupdate_fracandfillindisabled, refreshing class, in-plane, and shift parameters before final class-average generation
Fractional 2D restoration is class-local. cavger_init_online reads or centers
previous partial sums when fractional update is active, obtains per-class
realized fractions through get_class_update_fracs, and weights previous
even/odd class sums and CTF-squared sums independently for each class. This is
the 2D analogue of respecting independently updated objects in 3D.
Distributed cleanup must preserve class-average partial sums while fractional restoration still needs them as carry-over input. Assignment and distance artifacts are per-iteration handoffs and may be removed before the next iteration writes replacements.
5. Abinitio3D and Refine3D¶
The 3D controller derives the abinitio3D outer update policy from nsample.
The resulting update fraction is capped by UPDATE_FRAC_MAX.
Current high-level ab initio stage policy:
- stages 1 and 2 use
prob_neighwithprob_neigh_mode=shc - stages 3-5 use
prob - final neighborhood stages use
prob_neigh - stage 1 uses
nspace=500; stages 2-4 usenspace=1000 - every stage gets its low-pass and crop information independently from the normal schedule
- final active stages may switch to
fillin, except where the multi-state policy disables it
For abinitio3D multivol_mode=independent, the default policy is an
inspection-first multi-state run: nstages=5 and lpstop=6.0 A unless the
user overrides them. This stops after the prob phase and before
prob_neigh, staged NU filtering, independent-mode trailing reconstruction,
and staged automasking. The workflow still runs the final reconstruction step
at the configured last stage so it writes inspectable final state volumes. To
increase the chance that all active particles receive assignments before that
exit, independent mode starts stochastic balanced sampling at stage 4: the child
refine3D stages use greedy_sampling=no with frac_best=1.0 from stage 4
onward. This keeps class-balanced quotas but draws from the whole class, not a
top-ranked fraction. The outer particle target remains the fixed
nsample-derived update fraction at every stage.
abinitio3D multivol_mode=docked has an explicit split/update epoch policy.
Stages before the split run as one state. The default split stage is 6, so the
split occurs after stage 5. Docked early stops before the split are rejected.
Ordinary pre-split stages use the single-state target
min(UPDATE_FRAC_MAX, nsample / active_particles). The stage immediately before
the split increases that target to
min(UPDATE_FRAC_MAX, nstates * nsample / active_particles). This deliberately
broadens the one-state pose coverage before state labels are introduced.
Immediately before the cohort-forming assignment pass, the commander clears
sampled and updatecnt. It then always runs one class-balanced refine=prob
pass with frac_best=1.0, fillin=no, trail_rec=no, volrec=no, and
sticky_class_sampling=no. In the fractional regime its nominal target is:
min(active_particles,
max(effective_post_split_target,
min(NSAMPLE_HET_SPLIT_CAP, round(2.5 * nstates * nsample))))
NSAMPLE_HET_SPLIT_CAP is 100000 consistently with refine3D_states. The
corresponding fraction remains capped by UPDATE_FRAC_MAX, so the effective
target cannot exceed 90 percent of the active particles. The cap limits cohort
expansion but cannot make the cohort smaller than the effective post-split update.
In the global full-sampling regime, the pass instead targets all active particles.
At the split, the commander restores the requested state count, recomputes the
fixed post-split target
min(UPDATE_FRAC_MAX, nstates * nsample / active_particles), and randomizes
active particles into balanced uniform state labels without clearing the cohort
markers. It then selects one post-split-sized class-balanced subset restricted
to sampled > 0 and reconstructs state-specific starting volumes and halfmaps
from exactly that latest sampled round without trailing. The reconstruction
reproduces the subset.
Throughout ordinary post-split docked refinement, sample4update_class applies
the same sampled > 0 eligibility restriction when set_cline_refine3D calls
abinitio_docked_cohort_active and emits the
sticky_class_sampling=yes child flag. This flag is consumed only by the
class-balanced sample4update_class path, where it enables sampled_only; it
does not change unbalanced or full particle sampling. It is emitted only after
the fractional cohort pass succeeds and the requested state count is restored;
the matcher does not infer it from nstates or multivol_mode. The
cohort-forming reset ensures that eligibility means membership in the pre-split
cohort. The latest round remains sampled == max(sampled), and updatecnt
rotates selection toward less-updated cohort members. Particles outside the
cohort remain excluded until the separate terminal missing-assignment policy is
invoked, for which sticky_class_sampling=no.
The split-stage refinement uses refine=prob_state, keeps the post-split
fractional target, and keeps fillin=no. From TRAILREC_STAGE_SINGLE onward,
trailing reconstruction is enabled in the fractional regime and consumes
realized state-local update fractions. This includes the default split stage 6;
an earlier custom split stage does not enable trailing before that threshold.
The previous artifacts at the default split are the state-specific split
halfmaps reconstructed from the initial post-split-sized subset, not the
pre-split one-state halfmaps. With a full 2.5-times cohort, the first realized
fraction is approximately 0.4. Later docked stages keep the same fractional target;
neighborhood stages use prob_neigh_mode=geom, which selects the geometric
neighborhood containing each particle's previous best projection and evaluates
that same neighborhood for every state.
The global full-sampling switch remains authoritative. When
nsample / active_particles > 0.9, emitted docked stage commands omit
update_frac, nsample, and fillin, and trailing remains disabled, including
at the split stage. Because the cohort pass resets the counters before selecting
its members, updatecnt after the split is cohort-local selection history. Final docked reconstruction still requires every active particle to have a multi-state
assignment and may therefore invoke the separate terminal missing-update pass.
sample_ptcls4update3D applies the normal 3D subset policy:
- if fractional update is off, select all active particles
- if
balance=yes, use class-balanced sampling. If the sampling had been setup withpartition=yes, the class-balancing is based of the clustering of the underlying classes as materialized bycluster_cavgs - otherwise use update-count-biased sampling
sample_ptcls4fillin is a separate late-stage coverage policy. Its purpose is
to update particles with insufficient history, not to preserve the normal
balanced or count-biased exploration distribution.
The 3D matcher writes partial reconstructions from the active subset. Volume
assembly then restores volumes, calculates FSCs, postprocesses references, and
applies trailing reconstruction when requested. Trailing uses an explicit
ufrac_trec only when the parsed params%l_ufrac_trec_defined flag is true
and the run is single-state; otherwise it consumes realized per-state fractions
from get_state_update_fracs. The numeric params%ufrac_trec field has a
default value and must not be interpreted as an active override by itself.
Multi-state convergence reporting records those effective state-local fractions
as TRAIL_REC_UPDATE_FRAC_STATE01, TRAIL_REC_UPDATE_FRAC_STATE02, and so on.
6. Probabilistic Pre-Alignment¶
Probabilistic pre-alignment is not a second outer sampler.
The workflow is:
- choose the outer subset through the normal 2D or 3D sampling helper
- write the sampled project state
- run probability-table generation only for that subset
- aggregate table outputs into one assignment artifact
- reproduce the same subset in the matcher
- perform the hard particle update
simple_eul_prob_tab.f90, simple_eul_prob_tab_neigh.f90, and
simple_eul_prob_tab2D.f90 perform candidate-level importance sampling inside
that subset. They may use score-derived candidate distributions,
angle_sampling, greedy_sampling, or neighborhood sampling, but the selected
particle set is already fixed before they run.
7. Restoration and Assembly¶
2D class-average restoration consumes class-local realized update fractions:
- previous class contribution:
1 - rho(class) - current class contribution: the new partial sums for that class
Classes with no active updated particles keep a zero realized fraction. Classes with full sampled participation replace previous sums.
3D volume assembly performs trailing in the accumulator domain, mirroring the
2D scheme. The persistent per-state chain (trailrec_stateNN_{even,odd} plus
rho files and a trailrec_stateNN.txt manifest) holds blended, unregularized
e/o Fourier sums and sampling densities at full-dataset sampling mass. Two
fractions govern the blend:
f— the realized state-local fraction that produced the current partials (get_state_update_fracs); always computedu— the applied map-update weight; equalsfunless a single-stateufrac_trecoverride is provided
The recurrence keeps the chain at full mass D and makes u the restored
current-map coefficient, preserving the historical ufrac_trec meaning:
- current contribution: partial sums and rho scaled by
u / f(mass(u/f) * f * D = u * D) - previous chain contribution: sums and rho scaled by
1 - u - a single sampling-density correction after the blend restores the trailed halves, so each Fourier component is weighted by its accumulated sampling density; the FSC is estimated post-blend and describes the on-disk artifact
The chain is written before restoration (regularization mutates rho in place).
When the chain does not exist yet, volassemble bootstraps: it uses the legacy
previous-halfmap volume-domain blend for that iteration's outputs and seeds the
chain with the current partials scaled by 1/f, so the stored chain carries
full-dataset mass and the next iteration's effective update weight is the
requested fraction (an unnormalized fractional seed would make a 10 percent
request act like a ~53 percent update). Stage-boundary full reconstructions
seed the chain at full-dataset weight through the internal trail_seed
handshake, but only when the consuming stage actually trails.
The four accumulator files plus manifest form one artifact set. The manifest is deleted before and rewritten after the data files with per-component byte sizes, generation counter, and provenance (box, sampling, particle population, state layout), so interrupted writes never validate. Readers accept a chain only when the manifest parses, provenance matches the current project, every component size matches, and the grid is not larger than the current one with the same physical extent; smaller grids are zero-padded on read (downsampling ramp). Any validation failure discards the complete set and re-seeds. Cross-directory continuation carries chains over as complete sets only, manifest last.
Neither class-average restoration nor volume assembly should make new particle
sampling decisions. If a restoration or assembly change requires a different
subset policy, that policy belongs in the commander/controller/sampling-helper
layer and must be reflected in sampled and updatecnt.
Online matcher restoration/reconstruction paths must read active particle images from disk once per batch and reuse those batch images for both matching and restoration/reconstruction. Do not introduce a trailing full reconstruction or class-average restoration pass that re-reads image stacks as a memory optimization unless the single-read performance contract is explicitly changed. Probabilistic table-generation programs and explicit offline assembly commands are separate workflow stages and may perform their own reads.
8. Invariants¶
- Outer particle sampling happens before probabilistic table generation.
- Probabilistic table workers and downstream matchers reproduce the same subset.
- Candidate importance sampling never changes the particle subset.
sampledremains the current-round marker.updatecntremains cumulative update history.- Downstream restoration uses realized update state, not only nominal
update_frac. - Stage 1 of
abinitio2Dmay be sampled but must not fractionally carry over previous class-average sums. - The
abinitio3Ddocked split starts a new multi-statesampled/updatecntepoch. - The docked cohort pass resets sampling history before selecting its persistent
pre-split cohort; its
sampled > 0markers survive state relabeling. - Ordinary post-split docked class sampling is restricted to that sticky cohort, while reconstruction and downstream probabilistic work reproduce the latest sampled round exactly.
- Docked split-stage
prob_stateremains fractional unless the global full-sampling switch is active. - Independent multi-state
abinitio3Ddefaults to a five-stage,lpstop=6.0 Ainspection run, starts stochastic balanced sampling at stage 4, and still writes final reconstruction outputs. - In fractional docked mode, trailing starts at
TRAILREC_STAGE_SINGLE; for the default split stage this means split-stage refinement uses the freshly reconstructed state-specific split artifacts plus realized state-local update fractions. - 2D fractional class-average restoration remains class-local.
- Staged
abinitio2Dfillin=yesremains a full-assignment coverage guard unless the implementation is deliberately changed to missing-only assignment. - Sampled
abinitio2Druns a terminal dense greedy all-particle refresh before final class-average generation. volassembleand the classaverager remain consumers of sampled-update state, not producers of particle-selection policy.- Online matcher restoration/reconstruction reuses the particle images already read for the current batch.
9. Review Checklist¶
For sampling, probabilistic alignment, class-average restoration, or volume assembly changes, check:
- Does the outer subset get selected exactly once for a probabilistic pre-alignment iteration?
- Do table workers and matchers reuse the recorded subset through
sample4update_reprod? - Is candidate-level importance sampling kept separate from particle-level subset selection?
- With
inpl_cont=yes, does candidate profiling retain only rounded in-plane metadata and defer durable fractionale3to the final hard assignment? - Does candidate profiling perform one all-angle selection at the supplied shift without a coarse shift scan?
- Does final 3D refinement retain the authoritative rounded assignment and avoid a second global angle selection?
- Can any joint no-improvement or invalid-result path accidentally enter the legacy callback?
- Are
sampledandupdatecntupdated consistently before downstream restoration or trailing consumes them? - Does 2D restoration use class-local realized fractions?
- Does 3D trailing consume the realized or explicit trailing fraction?
- Does the online matcher path preserve one image-stack read per particle batch?
- Are shared-memory and distributed paths preserving the same scientific workflow and artifact contracts?