Abstract
Abstract This paper describes outcomes of the 2019 Cryo-EM Model Challenge. The goals were to (1) assess the quality of models that can be produced from cryogenic electron microscopy (cryo-EM) maps using current modeling software, (2) evaluate reproducibility of modeling results from different software developers and users and (3) compare performance of current metrics used for model evaluation, particularly Fit-to-Map metrics, with focus on near-atomic resolution. Our findings demonstrate the relatively high accuracy and reproducibility of cryo-EM models derived by 13 participating teams from four benchmark maps, including three forming a resolution series (1.8 to 3.1 Å). The results permit specific recommendations to be made about validating near-atomic cryo-EM structures both in the context of individual experiments and structure data archives such as the Protein Data Bank. We recommend the adoption of multiple scoring parameters to provide full and objective annotation and assessment of the model, reflective of the observed cryo-EM map density.
🔬 Techniques
💻 Software
🧪 Sample Preparation
💻 Software Details
💾 Data Repositories
🏛️ Research Organizations (ROR)
Affiliated research institutions:
📋 Methods
Challenge process and organization Informed by previous challenges 2 , 6 , 11 , the 2019 Model Challenge process was substantially streamlined in this round. In March, a panel of advisors with expertise in cryo-EM methods, modeling and/or model assessment was recruited. The panel worked with EMDR team members to develop the challenge guidelines, identify suitable map targets from EMDB and reference models from the PDB and recommend the metrics to be calculated for each submitted model. The challenge rules and guidance were as follows: (1) ab initio modeling is encouraged but not required. For optimization studies, any publicly available coordinate set can be used as the starting model. (2) Regardless of the modeling method used, submitted models should be as complete and as accurate as possible (that is, equivalent to publication-ready). (3) For each target, a separate modeling process should be used. (4) Fitting to either the unsharpened/unmasked map or one of the half-maps is strongly encouraged. (5) Submission in mmCIF format is strongly encouraged. Members of cryo-EM and modeling communities were invited to participate in mid-April 2019 and details were posted on the challenges website (challenges.emdataresource.org). Models were submitted by participant teams between 1 and 28 May 2019. For APOF targets, coordinate models were submitted as single subunits at the position of a provided segmented density consisting of a single subunit. ADH models were submitted as dimers. For each submitted model, metadata describing the full modeling workflow were collected via a Drupal webform, and coordinates were uploaded and converted to PDBx/mmCIF format using PDBextract 51 . Model coordinates were then processed for atom/residue ordering and nomenclature consistency using PDB annotation software (Feng Z., https://sw-tools.rcsb.org/apps/MAXIT ) and additionally checked for sequence consistency and correct position relative to the designated target map. Models were then evaluated as described below ( Model evaluation system ). In early June, models, workflows and initial calculated scores were made available to all participants for evaluation, blinded to modeler team identity and software used. A 2.5-day workshop was held in mid-June at Stanford/SLAC to review the results, with panel members attending in person. All modeling participants were invited to attend remotely and present overviews of their modeling processes and/or assessment strategies. Recommendations were made for additional evaluations of the submitted models as well as for future challenges. Modeler teams and software were unblinded at the end of the workshop. In September, a virtual follow-up meeting with all participants provided an overview of the final evaluation system after implementation of recommended updates.
Show full methods section
Challenge process and organization Informed by previous challenges 2 , 6 , 11 , the 2019 Model Challenge process was substantially streamlined in this round. In March, a panel of advisors with expertise in cryo-EM methods, modeling and/or model assessment was recruited. The panel worked with EMDR team members to develop the challenge guidelines, identify suitable map targets from EMDB and reference models from the PDB and recommend the metrics to be calculated for each submitted model. The challenge rules and guidance were as follows: (1) ab initio modeling is encouraged but not required. For optimization studies, any publicly available coordinate set can be used as the starting model. (2) Regardless of the modeling method used, submitted models should be as complete and as accurate as possible (that is, equivalent to publication-ready). (3) For each target, a separate modeling process should be used. (4) Fitting to either the unsharpened/unmasked map or one of the half-maps is strongly encouraged. (5) Submission in mmCIF format is strongly encouraged. Members of cryo-EM and modeling communities were invited to participate in mid-April 2019 and details were posted on the challenges website (challenges.emdataresource.org). Models were submitted by participant teams between 1 and 28 May 2019. For APOF targets, coordinate models were submitted as single subunits at the position of a provided segmented density consisting of a single subunit. ADH models were submitted as dimers. For each submitted model, metadata describing the full modeling workflow were collected via a Drupal webform, and coordinates were uploaded and converted to PDBx/mmCIF format using PDBextract 51 . Model coordinates were then processed for atom/residue ordering and nomenclature consistency using PDB annotation software (Feng Z., https://sw-tools.rcsb.org/apps/MAXIT ) and additionally checked for sequence consistency and correct position relative to the designated target map. Models were then evaluated as described below ( Model evaluation system ). In early June, models, workflows and initial calculated scores were made available to all participants for evaluation, blinded to modeler team identity and software used. A 2.5-day workshop was held in mid-June at Stanford/SLAC to review the results, with panel members attending in person. All modeling participants were invited to attend remotely and present overviews of their modeling processes and/or assessment strategies. Recommendations were made for additional evaluations of the submitted models as well as for future challenges. Modeler teams and software were unblinded at the end of the workshop. In September, a virtual follow-up meeting with all participants provided an overview of the final evaluation system after implementation of recommended updates.
Coordinate sources and modeling software
Modeling teams created ab initio models or optimized previously known models available from the PDB. Models optimized against APOF maps used PDB entries 2fha , 5n26 or 3ajo as starting models. Models optimized against ADH used PDB entries 1axe , 2jhf or 6nbb . Ab initio software included ARP/wARP 41 , Buccaneer 37 , Cascaded-CNN 46 , Mainmast 45 , Pathwalker 49 and Rosetta 40 . Optimization software included CDMD 39 , CNS 52 , DireX 43 , Phenix 21 , REFMAC 13 , MELD 48 , MDFF 44 and reMDFF 47 . Participants made use of VMD 53 , Chimera 38 , COOT 29 and PyMol for visual evaluation and/or manual model improvement of map-model fit. See Table 1 for software used by each modeling team. Modeling software versions/websites are listed in the Nature Research Reporting Summary .
Model evaluation system
The evaluation system for 2019 challenge (model-compare.emdataresource.org) was built on the basis of the 2016/2017 Model Challenge system 11 , updated with several additional evaluation measures and analysis tools. Submitted models were evaluated for >70 individual metrics in four tracks: Fit-to-Map, Coordinates-only, Comparison-to-Reference and Comparison-among-Models. A detailed description of the updated infrastructure and each calculated metric is provided as a help document on the model evaluation system website. Result data are archived at Zenodo 54 . Analysis software versions/websites are listed in the Nature Research Reporting Summary . For brevity, a representative subset of metrics from the evaluation website are discussed in this paper. The selected metrics are listed in Table 2 and are further described below. All scores were calculated according to package instructions using default parameters. Fit-to-Map The evaluated metrics included several ways to measure the correlation between map and model density as implemented in TEMPy 16 – 18 v.1.1 (CCC, CCC_OV, SMOC, LAP, MI, MI_OV) and the Phenix 21 v.1.15.2 map_model_cc module 19 (CCbox, CCpeaks, CCmask). These methods compare the experimental map with a model map produced on the same voxel grid, integrated either over the full map or over selected masked regions. The model-derived map is generated to a specified resolution limit by inverting Fourier terms calculated from coordinates, B factors and atomic scattering factors. Some measures compare density-derived functions instead of density (MI, LAP 16 ). The Q -score (MAPQ v.1.2 (ref. 8 ) plugin for UCSF Chimera 38 v.1.11) uses a real-space correlation approach to assess the resolvability of each model atom in the map. Experimental map density is compared to a Gaussian placed at each atom position, omitting regions that overlap with other atoms. The score is calibrated by the reference Gaussian, which is formulated so that a highest score of 1 would be given to a well-resolved atom in a map at an approximately 1.5 Å resolution. Lower scores (down to −1) are given to atoms as their resolvability and the resolution of the map decreases. The overall Q -score is the average value for all model atoms. Measures based on Map-Model FSC curve, Atom Inclusion and protein sidechain rotamers were also compared. Phenix Map-Model FSC is calculated using a soft mask and is evaluated at FSC = 0.5 (ref. 19 ). REFMAC FSCavg 13 (module of CCPEM 42 ) integrates the area under the Map-Model FSC curve to a specified resolution limit 13 . EMDB Atom Inclusion determines the percentage of atoms inside the map at a specified density threshold 14 . TEMPy ENV is also threshold-based and penalizes unmodeled regions 16 . EMRinger (module of Phenix) evaluates backbone positioning by measuring the peak positions of unbranched protein C γ atom positions versus map density in ring paths around C ɑ –C β bonds 15 . Coordinates-only Standard measures assessed local configuration (bonds, bond angles, chirality, planarity, dihedral angles; Phenix model statistics module), protein backbone (MolProbity Ramachandran outliers 20 ; Phenix molprobity module) and sidechain conformations, and clashes (MolProbity rotamers outliers and Clashscore 20 ; Phenix molprobity module). New in this challenge round is CaBLAM 22 (part of MolProbity and as Phenix cablam module), which uses two procedures to evaluate protein backbone conformation. In both cases, virtual dihedral pairs are evaluated for each protein residue i using C ɑ positions i − 2 to i + 2. To define CaBLAM outliers, the third virtual dihedral is between the CO groups flanking residue i . To define Calpha-geometry outliers, the third parameter is the C ɑ virtual angle at i . The residue is then scored according to virtual triplet frequency in a large set of high-quality models from PDB 22 . Comparison-to-Reference and Comparison-among-Models Assessing the similarity of the model to a reference structure and similarity among submitted models, we used metrics based on atom superposition (LGA GDT-TS, GDC and GDC-SC scores 23 v.04.2019), interatomic distances (LDDT score 24 v.1.2), and contact area differences (CAD 26 v.1646). HBPLUS 50 was used to calculate nonlocal hydrogen bond precision, defined as the fraction of correctly placed hydrogen bonds with more than six separations in sequence (HBPR > 6). DAVIS-QA determines for each model the average of pairwise GDT-TS scores among all other models 27 . Local (per residue) scores Residue-level visualization tools for comparing the submitted models were also provided for the following metrics: Fit-to-Map, Phenix CCbox, TEMPy SMOC, Q -score, EMRinger and EMDB Atom Inclusion; Comparison-to-Reference, LGA and LDDT; and Comparison-among-Models, DAVIS-QA. Metric score pairwise correlations and distributions For pairwise comparisons of metrics, Pearson correlation coefficients ( P ) were calculated for all model scores and targets ( n = 63). For average per-target pairwise comparisons of metrics, P values were determined for each target and then averaged. Metrics were clustered according to the similarity score (1 − | P |) using a hierarchical algorithm with complete linkage. At the beginning, each metric was placed into a cluster of its own. Clusters were then sequentially combined into larger clusters, with the optimal number of clusters determined by manual inspection. In the Fit-to-Map evaluation track, the procedure was stopped after three divergent score clusters were formed for the all-model correlation data (Fig. 4a ), and after two divergent clusters were formed for the average per-target clustering (Fig. 4b ). Controlling for model systematic differences As initially calculated, some Fit-to-Map scores had unexpected distributions, owing to differences in modeling practices among participating teams. For models submitted with all atom occupancies set to zero, occupancies were reset to one and rescored. In addition, model submissions were split approximately 50/50 for each of the following practices: (1) inclusion of hydrogen atom positions and (2) inclusion of refined B factors. For affected fit-to-map metrics, modified scores were produced excluding hydrogen atoms and/or setting B factors to zero. Both original and modified scores are provided at the web interface. Only modified scores were used in the comparisons described here.
Evaluation of group performance
Rating of group performance was done using the group ranks and model ranks (per target) tools on the challenge evaluation website. These tools permit users, either by group or for a specified target and for all or a subcategory of models (for example, ab initio), to calculate composite Z -scores using any combination of evaluated metrics with any desired relative weightings. The Z -scores for each metric are calculated from all submitted models for that target ( n = 63). The metrics (weights) used to generate composite Z -scores were as follows. Coordinates-only CaBLAM outliers (0.5), Calpha-geometry outliers (0.3) and Clashscore (0.2). CaBLAM outliers and Calpha-geometry outliers had the best correlation with Comparison-to-Reference parameters (Fig. 4f ), and Clashscore is an orthogonal measure. Ramachandran and rotamer criteria were excluded since they are often restrained in refinement and are zero for many models. Fit-to-Map EMRinger (0.3), Q -score (0.3), Atom Inclusion (0.2) and SMOC (0.2). EMRinger and Q -score were among the most promising model-to-map metrics, and the other two provide distinct measures. Comparison-to-Reference LDDT (0.9), GDC_all (0.9) and HBPR >6 (0.2). LDDT is superposition-independent and local, while GDC_all requires superposition; H-bonding is distinct. Metrics in this category are weighted higher, because although the reference models are not perfect, they are a reasonable estimate of the right answer. Composite Z -scores by metric category (Extended Data Fig. 6a ) used the Group Ranks tool. For ab initio rankings (Extended Data Fig. 6b ), Z -scores were averaged across each participant group on a given target, and further averaged across T1 + T2 and across T3 + T4 to yield overall Z -scores for high and low resolutions group 54 models were rated separately because they used different methods. Group 73’s second model on target T4 was not rated because the metrics are not set up to meaningfully evaluate an ensemble. Other choices of metric weighting schemes were tried, with very little effect on clustering. Molecular graphics Molecular graphics images were generated using UCSF Chimera 38 (Fig. 2 and Extended Data Fig. 3 ) and KiNG 55 (Extended Data Figs. 1 , 2 and 4 ). Reporting Summary Further information on research design is available in the Nature Research Reporting Summary linked to this article.
Online content Any methods, additional references, Nature Research reporting summaries, source data, extended data, supplementary information, acknowledgements, peer review information; details of author contributions and competing interests; and statements of data and code availability are available at 10.1038/s41592-020-01051-w.
Supplementary information Supplementary Information Supplementary Fig. 1. Reporting Summary
Supplementary information The online version contains supplementary material available at 10.1038/s41592-020-01051-w.
📊 Figures
Fig. 1
Single particle cryo-EM models in the Protein Data Bank.
a , Plot of reported resolution versus PDB release year. Models derived from single particle cryo-EM maps have increased dramatically since the u2018resolution revolutionu2019 circa 2014. Higher-resol...
Fig. 2
Challenge targets: cryo-EM maps at near-atomic resolution.
Shown from left to right are u0251-helix-rich APOF at 1.8, 2.3 and 3.1u2009u00c5 (EMDB entries EMD- 20026 , EMD- 20027 and EMD- 20028 ) and ADH at 2.9u2009u00c5 (EMDB entry EMD- 0406 ). a , Full maps ...
Fig. 3
Challenge pipeline.
a u2013 c , Overview of the challenge setup ( a ), submissions ( b ) and evaluation ( c ) strategy. d , Scores comparison. Multiple interactive tabular and graphical displays enable comparative evalua...
Fig. 4
Evaluation of metrics.
Model metrics (Table 2 ) were compared with each other to assess how similarly they performed in scoring the challenge models. a u2013 d , Fit-to-Map metrics analyses. a , Pairwise correlations of sco...
Extended Data Fig. 1
Evaluation of peptide bond geometry.
All 63 Challenge models were evaluated using MolProbity. APOF and ADH each have one cis peptide bond per subunit before a proline residue. ( a ) Counts of peptide bonds with each of the following conf...
Extended Data Fig. 2
Classic CaBLAM outlier with no Ramachandran outlier.
a , Mis-modeled peptide (identified by red ball at carbonyl oxgen position) is flagged by two successive CaBLAM outliers (magenta dihedrals), a bad clash (hot-pink spikes), and a bond-angle outlier (n...
Extended Data Fig. 3
Evaluation of a short sequence misalignment within a helix.
Local Fit-to-Map and Coordinates-only scores are compared for a 3-residue sequence misalignment inside an u0251-helix in an ab initio model submitted to the Challenge (APOF 2.3u2009u00c5 54_1). a , Mo...
Extended Data Fig. 4
Modeling errors around omitted Zinc ligand in ADH.
Target 4 (ADH) density map with examples of modeling errors caused by omission of Zinc ligand. a , Reference structure with Zinc metal ion (gray ball) coordinated by 4 Cysteine residues (blue sidechai...
Extended Data Fig. 5
Fit-to-Map Scores with and without refined B-factors (ADP).
Two representative metrics are shown: a , CCmask correlation, b , FSC05 resolution u22121 . Each plotted point indicates the calculated score for atom positions with B-factors included (horizontal axi...
Figure images are served from the NIH/NLM PubMed Central Open Access Subset or Europe PMC; copyright remains with the publishers and authors.
💬 Discussion
0 commentsNo comments yet. Be the first to start a discussion!
Leave a Comment