====== GMM-based focused refinement for single particle analysis (2026) ======
* Most programs are available in EMAN2 builds after 2026-06, but newer versions are typically better.
* The tutorial is only tested on Linux with Nvidia GPU and CUDA.
* For reference of the method, see the [[https://arxiv.org/abs/2605.30518 | preprint]].
* A separate tutorial about the same method applied to CryoET data can be found [[https://blake.bcm.edu/emanwiki/doku.php?id=eman2:e2tomo_atpsyn | here]].
In this tutorial, we use a refinement of TRPV1 from a CryoSPARC [[https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/case-study-dktx-bound-trpv1-empiar-10059 | tutorial]]. In theory you can also do everything in EMAN2 from scratch, but this makes a clearer comparison with other software.
===== Import existing refinement =====
Start from the consensus refinement from CryoSPARC, and import the particles with:
e2convertrelion.py J35_008_particles.star --onestack particles/all_particles.hdf
Depending on the masking, the resolution should be around 2.7 to 3 Å. After the import finishes, you should find relavent files in the ''r3d_00'' folder.
===== Global orientation refinement =====
First build the GMM representation from the 3D volume.
e2gmm_guess_n.py r3d_00/threed_00.hdf --evenodd --nopdb --thr 4.5 --maxres 3
Here the ''--evenodd'' option will build GMM for the two half-set separately. Pick a threshold for ''--thr'' so all protein density of interest are visible. ''--nopdb'' force it to write points as txt instead of pdb, since the pdb format limits the maximum number of atoms. This should produce ''r3d_00/threed_00_gmm_even/odd.txt''.
Run the global refinement next.
e2gmm_refine_iter.py r3d_00/threed_00.hdf --startres 2.8 --initpts r3d_00/threed_00_gmm.txt --maskpp r3d_00/mask.hdf --jax --niter 3 --parallel thread:32 --sym c4
It is ok to use ''r3d_00/threed_00_gmm.txt'' even though the file does not exist, since the program will look for the even/odd versions automatically. Make sure to add ''--jax'' to call the JAX backend, as we are retiring the Tensorflow backend gradually. Without ''--maskpp'', the program will use auto-generated masks. Here we keep the same masking for better comparison.
After the refinement, the global resolution should reach Nyquist, at 2.5 Å, according to the "gold-standard" FSC.
{{:eman2:gmm_trpv_global.png|global refinement}}
===== Focused refinement =====
Now we run focused refinement on the CTD of one asymmetrical unit. First expand the c4 symmetry. Note that ''e2proclst.py --sym'' operates in place, so it is safer to backup.
e2proclst.py gmm_00/ptcls_03_even.lst --sym c4
e2proclst.py gmm_00/ptcls_03_odd.lst --sym c4
Next, make a mask that covers the target domain. This can be done using ''FilterTool'' through the e2display browser. I used ''mask.soft'' and ''mask.auto3d.thresh'', but many other combinations also work. Save the processed map and rename it ''mask_unit1.hdf''.
{{:eman2:gmm_trpv_mask.png|focus mask}}
Finally, run the acutal GMM-based focused refinement.
e2gmm_refine_iter.py gmm_00/threed_03.hdf --startres 2.5 --initpts r3d_00/threed_00_gmm.txt --maskpp r3d_00/mask.hdf --jax --niter 4 --parallel thread:32 --sym c1 --mask mask_unit1.hdf
Compared to the global refinement command, here we start from the existing GMM refinement ''gmm_00'', and provide the focus mask through ''--mask''. The program will run one iteration of alignment with neural network and the rest of iterations with gradient descent. When there are multiple flexible domains for focused refinement, simply make multiple masks and combine them into a 3D volume stack. Then use the volume stack file for the ''--mask'' option.
In the output folder, ''fsc_foc_XX.txt'' reports the FSC curve of the target region. Here although the global resolution reached Nyquist already, the local resolution at CTD is still at 2.7 Å after the GMM-based global refinement. After the focused refinement, the resolution of the domain also reaches Nyquist.
After each iteration, the program align and reconstruct every domain of focus (''threed_XX_YY_even/odd.hdf'', where XX is the iteration number and YY is the index of the focused region), and generate composite maps for even/odd half-sets separetly (''threed_XX_even/odd.hdf''). The ''fsc_masked_xx.txt'' are computed using the composite maps.
{{:eman2:gmm_trpv_focus.png|focus refinement}}
Here the subunit of focus is on the right. The second image is a composite map colored by local resolution (blue is better).
===== Heterogeneity analysis =====
Since the first iteration alignment is done by a neural network, it leaves its latent layer output in the same folder, and we can visualize the dynamic of the system from it. However since we keep strict "gold-standard" data separation, the even/odd half-sets are processed with two different networks independently, and there is currently no way to combine the latent space of two networks. Therefore, we can only visualize the dynamics one half-set at a time. Replace the ''even'' in the following command with ''odd'' to visualize the other half.
e2gmm_eval.py --pts gmm_01/mid_01_even.txt --pcaout gmm_01/mid_pca_even.txt --ptclsin gmm_01/ptcls_00_even.lst --ptclsout gmm_01/ptcls_cls_01_even.lst --mode regress --ncls 10 --nptcl 20000 --axis 0 --parallel thread:96
To visualize dynamics of the the combined dataset and without the rigid body constraint, you can also run
e2gmm_heter_refine.py gmm_01/threed_00.hdf --mask mask_unit1a.hdf --nogoldstandard --model gmm_01/model_03_even.txt --maxres 7
Here we use ''gmm_01/threed_00.hdf'' as input so it uses the orientation assignment from GMM-based global refinement.
{{:eman2:gmm_trpv_mov0.gif|movement - large}}
Finally, we can visualize the more subtle movement within the CTD, after the large scale rigid-body movement has been corrected by the focused refinement. This time run the same ''e2gmm_heter_refine.py'' command, but using the focused alignment orientation from the last iteration.
e2gmm_heter_refine.py gmm_01/threed_04.hdf --mask mask_unit1a.hdf --nogoldstandard --model gmm_01/model_03_even.txt --maxres 7
{{:eman2:gmm_trpv_mov1.gif|movement - fine}}