Abstract
The reference-free averaging of three-dimensional electron microscopy (3D-EM) reconstructions with empty regions in Fourier space represents a pressing problem in electron tomography and single-particle analysis. We present a maximum likelihood algorithm for the simultaneous alignment and classification of subtomograms or random conical tilt (RCT) reconstructions, where the Fourier components in the missing data regions are treated as hidden variables. The behavior of this algorithm was explored using tests on simulated data, while application to experimental data was shown to yield unsupervised class averages for subtomograms of groEL/groES complexes and RCT reconstructions of p53. The latter application served to obtain a reliable de novo structure for p53 that may resolve uncertainties about its quaternary structure.
🔬 Techniques
🔭 Microscopes
🧬 Organisms
💻 Software
✨ Fluorophores
🧪 Sample Preparation
🏭 Microscope Brands
📷 Detectors
💻 Software Details
🏛️ Research Organizations (ROR)
Affiliated research institutions:
📋 Methods
The Data Model We model our data set to consist of N 3D reconstructions (particles) X i , with 3D Fourier transforms X i , which are described according to the following model in Fourier space: (1) X i = R Φ i A k i + G i for i = 1 , 2 , … , N where: X i ∈ℜ J is a J -dimensional vector of complex values ( X i ) j or X ij in short, which can be divided into a vector X i o of observed components X ij o and a vector X i m of missing components X ij m . Each data vector X i may have a unique pattern of missing components, but the pattern is known for every X i . We formalize the notion of a missing data generation mechanism by defining a J -dimensional missing data indicator vector w i , such that w ij = 1 for observed elements, and w ij = 0 for missing elements. A k i is one of K unknown 3D structures in Fourier space A 1 , A 2 ,…, A K ∈ℜ J . These are the objects we wish to estimate from the data. k i is an unknown, random integer with possible values 1, 2,…, K , indicating which of the unknown structures corresponds to particle X i . R Φ i is a transformation matrix, which maps the unknown structure A k i on particle X i . The actual transformation Φ i for particle X i is unknown. We parameterize this transformation by ϕ i , a 6D vector (corresponding to three Euler angles q rot , q tilt , q psi , and three real-space translation coordinates, q x , q y , q z ). G i ∈ℜ J is a J -dimensional vector of unknown, independent Gaussian noise with mean zero and unknown standard deviation σ. In this model, the data space comprises the entire J -dimensional vector X i , including the part that remained unobserved, which contrasts with alternative approaches where the data space is restricted to only the observed part of X i , (e.g., Forster and Hegerl, 2007 ). We note here that also in the latter case one may devise a maximum-likelihood algorithm, and investigations into alternative ways to deal with preferred orientations within such a data model are currently underway (unpublished results). Also, the assumption of white Gaussian noise in this model may only be an approximate description of experimental noise in 3D-EM. Still, similar assumptions in maximum-likelihood (ML) classification approaches for single-particle data have been shown to be highly effective ( Scheres et al., 2005 , 2007a ). Moreover, the mathematical framework presented here could easily be adapted to describe colored noise by introducing a resolution-dependent estimate for the standard deviation of the noise. Again, for single-particle analysis such an approach has previously been shown to be effective ( Scheres et al., 2007b ), but we chose not to implement a similar approach here because of its increased computational complexity.
Show full methods section
The Data Model We model our data set to consist of N 3D reconstructions (particles) X i , with 3D Fourier transforms X i , which are described according to the following model in Fourier space: (1) X i = R Φ i A k i + G i for i = 1 , 2 , … , N where: X i ∈ℜ J is a J -dimensional vector of complex values ( X i ) j or X ij in short, which can be divided into a vector X i o of observed components X ij o and a vector X i m of missing components X ij m . Each data vector X i may have a unique pattern of missing components, but the pattern is known for every X i . We formalize the notion of a missing data generation mechanism by defining a J -dimensional missing data indicator vector w i , such that w ij = 1 for observed elements, and w ij = 0 for missing elements. A k i is one of K unknown 3D structures in Fourier space A 1 , A 2 ,…, A K ∈ℜ J . These are the objects we wish to estimate from the data. k i is an unknown, random integer with possible values 1, 2,…, K , indicating which of the unknown structures corresponds to particle X i . R Φ i is a transformation matrix, which maps the unknown structure A k i on particle X i . The actual transformation Φ i for particle X i is unknown. We parameterize this transformation by ϕ i , a 6D vector (corresponding to three Euler angles q rot , q tilt , q psi , and three real-space translation coordinates, q x , q y , q z ). G i ∈ℜ J is a J -dimensional vector of unknown, independent Gaussian noise with mean zero and unknown standard deviation σ. In this model, the data space comprises the entire J -dimensional vector X i , including the part that remained unobserved, which contrasts with alternative approaches where the data space is restricted to only the observed part of X i , (e.g., Forster and Hegerl, 2007 ). We note here that also in the latter case one may devise a maximum-likelihood algorithm, and investigations into alternative ways to deal with preferred orientations within such a data model are currently underway (unpublished results). Also, the assumption of white Gaussian noise in this model may only be an approximate description of experimental noise in 3D-EM. Still, similar assumptions in maximum-likelihood (ML) classification approaches for single-particle data have been shown to be highly effective ( Scheres et al., 2005 , 2007a ). Moreover, the mathematical framework presented here could easily be adapted to describe colored noise by introducing a resolution-dependent estimate for the standard deviation of the noise. Again, for single-particle analysis such an approach has previously been shown to be effective ( Scheres et al., 2007b ), but we chose not to implement a similar approach here because of its increased computational complexity.
EXPERIMENTAL PROCEDURES GroEL/groES Subtomogram Data The tomographic data on groEL/groES mixtures used in this work was published previously ( Forster et al., 2008 ). These authors collected single-axis tilt series on purified protein samples of GroEL and GroELGroES complexes mixed with a 10 nm BSA-colloidal gold suspension. Carbon-coated grids were vitrified by plunge freezing, and low-dose images were collected using a Tecnai G2 Polara microscope (FEI, Eindhoven, The Netherlands) equipped with a 2 k × 2 k CCD camera. The tilt series ranged from −65 to 65 with a 2–2.5 increment, defocus values between 4 and 7 μm were used, and the pixel size at specimen level was 0.6 nm. Manual particle selection yielded several thousand subtomograms, which were downscaled to 32 × 32 × 32 voxels with a final voxel size of 1.2 nm. From this data set, 793 subtomograms were selected based on a cross-correlation criterion with the average of all subtomograms aligned against a low-pass-filtered crystal structure of the GroELGroES complex ( Xu et al., 1997 ).
P53 Electron Microscopy
A deletion mutant of human p53, lacking its 33 C-terminal residues (NTC, p53 1-360 ) and with four stabilizing mutations in the core domain (M133L/V203A/N239Y/N268D) ( Nikolova et al., 1998 ), was expressed in E. coli , and purified as described elsewhere ( Veprintsev et al., 2006 ). Complexes with a double-stranded DNA probe containing the TGFA response element were prepared at a high concentration (10 μM) to avoid free p53. These samples were simultaneously separated on molecular weight and cross-linked with glutaraldehyde using the Grafix method ( Kastner et al., 2008 ). For all fractions, the interaction of the DNA probe with p53 was confirmed by electrophoretic mobility shift assay, and the size of the complex was confirmed using SDS-PAGE. Only fractions containing p53 tetramers in complex with DNA were selected for EM. Negative staining was performed with uranyl formiate in carbon sandwich technique over Quantifoil grids. Micrographs were taken in tilted pairs, at tilt angles of 0° and 45°, under low-dose conditions on a JEOL JEM-2200FS electron microscope. Images were recorded on a 4k × 4k CCD camera at a magnification of 64,305 ×, corresponding to a 2.1 Å pixel size. A total of 2521 pairs of particles were manually selected, downscaled to a pixel size of 8.4 Å, and processed using standardized RCT protocols in Xmipp ( Scheres et al., 2008 ) to generate 40 reconstructions of 30 × 30 × 30 voxels. Similarly, a second data set consisting of 7600 untilted particles was obtained using the same protein sample, but incubated with a dsDNA probe containing the GADD45 response element. Implementation The ML algorithm outlined above was implemented in the xmipp_ml_tomo program of the Xmipp package ( Sorzano et al., 2004 ). The same program also implements a standard multireference refinement scheme based on weighted averaging and maximum constrained cross-correlations (CC). In addition, although we did not derive the corresponding formulas here, the program implements a modified maximum-likelihood algorithm that employs weighted averaging to calculate the reference structures. The latter two implementations depend on an additional free parameter in the form of a threshold value to prevent divisions by zero, which was set to 0.1 in all calculations. Although the program reads 3D maps in single-file SPIDER format only, the Xmipp package provides utilities to convert to and from other common formats like MRC ( Crowther et al., 1996 ) or EM ( Hegerl, 1996 ). Missing data regions may be defined as cones, wedges, or pyramids, and the program may also be used to align 3D maps without missing regions. The 6D integrations over all orientational and translational assignments are performed exhaustively, although the orientational integrations may optionally be limited to a given angular distance from a set of initial orientations for all particles. The integrations over all orientations are replaced by discrete Riemann sums over approximately even distributions of the three Euler angles. It is noteworthy that these distributions are oriented differently according to a small random perturbation at every iteration to prevent the optimization from getting stuck in local minima. The program has been parallelized with near-linear efficiency for the numbers of processors employed in this study, using a hybrid scheme of distributed and shared memory parallelization by the combined use of the message passing interface (MPI) and threads. The computational cost of the implementation scales linearly with the number of particles, the number of classes and the number of sampled rotations. In addition, it scales approximately linearly with the number of voxels in the 3D maps. The runs with the model data took approximately 7–8 hr wall clock time each, using 20 3GHz Intel Xeon processors in parallel. The groEL runs took in total 17 hr on 128 2.3GHz IBM Power PC processors, and the alignment of the p53 maps took 20 min on 20 3GHz Intel Xeon processors.
Supplementary Material supplement
📊 Figures
Figure 1
Model Calculations
(A) Class purities (see below) for the ML (black) and CC (gray) approach to unsupervised multireference refinement of simulated data sets with varying amounts of noise. Averages and standard deviation...
Figure 2
GroEL/GroES Subtomogram Averaging
(A) Central slices through XY and YZ of the three class averages obtained for the data set of subtomograms of groEL and groEL/groES complexes. (B) Symmetrized average subtomogram for class 2 with fitt...
Figure images are served from the NIH/NLM PubMed Central Open Access Subset or Europe PMC; copyright remains with the publishers and authors.
💬 Discussion
0 commentsNo comments yet. Be the first to start a discussion!
Leave a Comment