Abstract
Electron tomography is a technique for three-dimensional reconstruction, that is widely used for imaging macromolecules, macromolecular assemblies or whole cells. Combined with cryo-electron microscopy, it is capable of visualizing structural detail in a state close to in vivo conditions in the cell. In electron tomography, micrographs are taken while tilting the specimen to different angles about a fixed axis. Due to mechanical constraints, the angular tilt range is limited. As a consequence, the reconstruction of a 3D image is missing data, which for a single axis tilt series is called the "missing wedge", a region in reciprocal space where Fourier coefficients cannot be obtained experimentally. Tomographic data is analyzed by extracting subvolumes from the raw tomograms, by alignment of the extracted subvolumes, multivariate data analysis, classification, and class-averaging, which results in an increased signal-to-noise ratio and substantial data reduction. Subvolume analysis is a valuable tool to discriminate heterogeneous populations of macromolecules, or conformations of a macromolecule or macromolecular assembly as well as to characterize interactions between macromolecules. However, this analysis is hampered by the lack of data in the original tomograms caused by the missing wedge. Here, we report enhancements of our subvolume processing protocols in which the problem of the missing data in reciprocal space is addressed by using constrained correlation and weighted averaging in reciprocal space. These procedures are applied to the analysis of myosin V and simian immunodeficiency virus (SIV) envelope spikes. We also investigate the effect of the missing wedge on image classification and establish limits of reliability by model calculations with generated phantoms.
🔬 Techniques
💻 Software
✨ Fluorophores
🧪 Sample Preparation
🏭 Microscope Brands
📷 Detectors
💻 Software Details
💾 Data Repositories
🏛️ Research Organizations (ROR)
Affiliated research institutions:
📋 Methods
Tomographic data
The tomographic data sets used in this study and the results from subvolume averaging were published previously: the structure of myosin V in the inhibited state ( Liu et al., 2006a ), and SIV virus envelope spikes ( Zhu et al., 2006 ). In these previous studies, the processing methodology did not include any missing wedge compensation. In the following, we summarize the data collection and image processing techniques used in the previous studies that produced the tomograms which are the basis of this work. The SIV image data of the virus cryo-samples was collected on a Philips CM300 FEG electron microscope equipped with a goniometer and a Tietz TemCam F224 CCD camera (2048×2048 pixels), at 300 kV and a magnification of 43,200 under low-dose conditions. Defocus values were in the range of 4 – 6 μm. Three tilt series with 70 – 80 images were collected that covered an angular tilt range up to 70°. Variable increments were chosen according to the Saxton rule ( Saxton et al., 1984 ), starting with a 2° step at 0° tilt for two series, and a 3° step for the other series. The pixel size at the specimen level was 0.56 nm. The tilt series were processed with the “protomo” software package ( Winkler and Taylor, 2006 ) using marker-free alignment and the final maps were computed with weighted backprojection. A total of 2,004 subvolumes were selected by visual inspection from the original 6,175 subvolumes of the earlier study. Selection criteria were the overall appearance of the subvolumes and the location within the tomogram, i. e. subvolumes originating from regions of the tomograms with apparent lower resolution or poorer structural preservation were rejected, as well as subvolumes near the edges of the tomograms. Cryo-electron microscopy and tomography of the second specimen, myosin V in the inhibited state adsorbed on lipid monolayers ( Liu et al., 2006a ), were carried out under the same conditions as described above for the SIV envelope spikes. Eight tilt series were collected with a starting angle increment of 2°, and the tilt range covered angles up to 70°. The pixel size at the specimen level was 0.56 nm. The defocus range of the eight collected tilt series was between 5 and 12 μm. For the last tilt series alignment cycle and the computation of the final maps, defocus corrected micrographs were used. Images taken at tilt angles greater than 30° were defocus gradient corrected ( Winkler and Taylor, 2003 ), while the other images were simply corrected with a Wiener filter. The subvolume positions were derived from the location of the “flower motifs” in the myosin V arrays, whereby the parameters for each of the six “petal motifs” (consisting of two lever arms and a cargo-binding domain of the myosin V molecule) were calculated by applying a shift and a rotation, resulting in a total of 11,112 subvolumes containing the petal motifs. Subvolumes from each data set were reanalyzed with the volumetric data processing package described in ( Winkler, 2007 ). Unlike the previously published results, the alignment of the subvolumes now included a compensation for the missing wedge, which is based on the principle of constrained correlation ( Frangakis et al., 2002 ). Multivariate data analysis and hierarchical ascendant classification were applied to analyze the structural heterogeneity. Class averages were computed by averaging Fourier coefficients, so that missing regions are taken into account explicitly. The processed 2,004 SIV subvolumes had a size of 48×48×48 pixels, whereas subvolumes of 72×60×36 pixels were used for the 11,112 petal motifs of myosin V.
Show full methods section
Tomographic data
The tomographic data sets used in this study and the results from subvolume averaging were published previously: the structure of myosin V in the inhibited state ( Liu et al., 2006a ), and SIV virus envelope spikes ( Zhu et al., 2006 ). In these previous studies, the processing methodology did not include any missing wedge compensation. In the following, we summarize the data collection and image processing techniques used in the previous studies that produced the tomograms which are the basis of this work. The SIV image data of the virus cryo-samples was collected on a Philips CM300 FEG electron microscope equipped with a goniometer and a Tietz TemCam F224 CCD camera (2048×2048 pixels), at 300 kV and a magnification of 43,200 under low-dose conditions. Defocus values were in the range of 4 – 6 μm. Three tilt series with 70 – 80 images were collected that covered an angular tilt range up to 70°. Variable increments were chosen according to the Saxton rule ( Saxton et al., 1984 ), starting with a 2° step at 0° tilt for two series, and a 3° step for the other series. The pixel size at the specimen level was 0.56 nm. The tilt series were processed with the “protomo” software package ( Winkler and Taylor, 2006 ) using marker-free alignment and the final maps were computed with weighted backprojection. A total of 2,004 subvolumes were selected by visual inspection from the original 6,175 subvolumes of the earlier study. Selection criteria were the overall appearance of the subvolumes and the location within the tomogram, i. e. subvolumes originating from regions of the tomograms with apparent lower resolution or poorer structural preservation were rejected, as well as subvolumes near the edges of the tomograms. Cryo-electron microscopy and tomography of the second specimen, myosin V in the inhibited state adsorbed on lipid monolayers ( Liu et al., 2006a ), were carried out under the same conditions as described above for the SIV envelope spikes. Eight tilt series were collected with a starting angle increment of 2°, and the tilt range covered angles up to 70°. The pixel size at the specimen level was 0.56 nm. The defocus range of the eight collected tilt series was between 5 and 12 μm. For the last tilt series alignment cycle and the computation of the final maps, defocus corrected micrographs were used. Images taken at tilt angles greater than 30° were defocus gradient corrected ( Winkler and Taylor, 2003 ), while the other images were simply corrected with a Wiener filter. The subvolume positions were derived from the location of the “flower motifs” in the myosin V arrays, whereby the parameters for each of the six “petal motifs” (consisting of two lever arms and a cargo-binding domain of the myosin V molecule) were calculated by applying a shift and a rotation, resulting in a total of 11,112 subvolumes containing the petal motifs. Subvolumes from each data set were reanalyzed with the volumetric data processing package described in ( Winkler, 2007 ). Unlike the previously published results, the alignment of the subvolumes now included a compensation for the missing wedge, which is based on the principle of constrained correlation ( Frangakis et al., 2002 ). Multivariate data analysis and hierarchical ascendant classification were applied to analyze the structural heterogeneity. Class averages were computed by averaging Fourier coefficients, so that missing regions are taken into account explicitly. The processed 2,004 SIV subvolumes had a size of 48×48×48 pixels, whereas subvolumes of 72×60×36 pixels were used for the 11,112 petal motifs of myosin V.
Subvolume processing
Two strategies of aligning and classifying subvolumes have been applied. The first one follows established procedures used in single particle analysis. An alignment step brings the raw subvolumes into register and is followed by a classification and averaging step. This process is iterative: class averages obtained in one cycle are used as alignment references in the subsequent cycle. For cross-correlation alignment, one or more references were created to which the raw subvolumes were aligned. Except for the initial alignment, multiple references were always used, since the investigated structures are expected to be variable to a certain extent. The second strategy is called “alignment by classification” ( Dube et al., 1993 ). With this variant, differences in orientation are separated by classification. Subsequently, class averages are aligned with respect to each other, rather than raw images. The resulting alignment transformations are then applied to the raw subvolumes, at which point a new cycle of the iterative procedure can be started. For the SIV envelope spikes, the orientation was initially determined by fitting an ellipsoidal surface to the picked spike positions and the spike axes were approximated by the calculated surface normals at the picked positions ( Winkler, 2007 ). This procedure cannot determine the spike rotation about the spike axis, so that only two of the three Euler angles are obtained at this stage. The initial orientation parameters (3 spike origin coordinates and 2 Euler angles) were refined by a cross-correlation alignment to a global spike average that had been rotationally averaged about the spike axis. The choice of a rotationally averaged reference reduces the rotational grid search to two parameters, thus speeding up the otherwise time-consuming procedure in the initial stages, when the spike orientations are still inaccurate and larger grid search ranges are required. The third Euler angle was included in the refinement of the orientation parameters in later processing cycles, after a classification produced appropriate references that discriminated the varying orientations. For the more or less planar arrangement of the myosin V petal motifs, positions (2 origin coordinates) and in-plane rotations (1 rotation angle) were derived from the flower motifs, each of which consists of six petals. The flower motifs are arranged in an imperfect hexagonal lattice, which simplified locating the motifs automatically by cross-correlation methods and determining the in-plane rotations. This was all carried out with projections of the myosin V maps to speed up the initial processing, and the first cycle of processing used projections rather than volumes also. In subsequent cycles, the alignment was switched to the volume data and the search range was limited to a few degrees in order to take into account the wrinkling of the monolayers. At the beginning of each cycle of the first processing variant (alignment of raw subvolumes), multiple references were selected from class averages produced by the preceding cycle. Care has been taken in the selection process that the references represented the whole spectrum of variance obtained by the classification. For the SIV specimen, 10 – 25 references were selected, and, for myosin V, where the larger number of subvolumes allowed us to produce more classes with a similar number of class members, up to 60 classes were selected. The references were first windowed to exclude regions outside a volume slightly larger than the observed molecular density. In addition, the membrane of the SIV virion was also partially windowed in order to focus the alignment on the spike structure rather than the membrane density, which appeared to be relatively strong in some classes. Finally, the references were band-pass filtered, the low-pass limit was chosen to be equal or less than the spatial frequency of the first zero of the CTF, and the high-pass limit chosen to remove low-frequency density variations. Each raw subvolume was extracted from the maps given the geometric parameters (position and orientation) from the previous cycle and the subvolume was cross-correlated with each reference. The constrained cross-correlation coefficient served as a similarity measure. The new parameters were obtained with a grid search over the rotation parameters from the reference which produced the highest co-efficient, and the geometric parameters were updated accordingly. Aligned subvolumes were not stored, since it is sufficient to record the updated parameters. Instead, the aligned subvolumes were re-interpolated on the fly with the updated alignment parameters, when needed. This approach saved large amounts of storage space for intermediate results during the processing cycles, and avoided multiple re-interpolations of the image data. The rotational alignment was carried out with a grid search by modifying the orientation with which the subvolumes were extracted. An exhaustive search over all possible orientations in space was not needed, since an approximate orientation of the motifs was known. Thus, for the SIV spikes the polar angle was only searched in 2° steps on cones with half-widths of 2° and 4°, the half-width subsequently being reduced to 1° for the final cycle. Rotation about the spike axis was initially tested in 7.5° steps and gradually reduced to 2° steps for later cycles. The myosin V orientation was refined after the initial projection alignment to correct for deviations caused by wrinkling and to account for lattice distortions of the para-crystalline array. In the final cycles, the orientation was searched on cones with half-widths of 1°, the in-plane rotational search was restricted to ±0.5°. Principal component analysis using the modulation metric ( Borland and van Heel, 1990 ) was applied after the subvolume alignment. Relevant voxels of the re-extracted aligned subvolumes were selected by specifying a binary mask. The mask was generated in such a way that it contained mostly voxels of the molecular volume, and excluded the surrounding vitrified ice. At some stages of the SIV analysis, cylindrical masks containing a spike were also used that necessarily included a slightly larger volume in the stalk region, but were easier to generate. In any case, the masks always excluded the virion membrane from the analysis. Hierarchical ascendant classification ( van Heel, 1989 ) was carried out to find groups of similar subvolumes. For the SIV specimen, 50 classes were normally generated so that the number of averaged subvolumes per class was around 30 – 60. Whereas in the iterative processing stages, a large number of classes was generated to obtain as many potential variations of the spike conformations, this number was reduced to 8 in the final cycle for visualization. In the case of myosin V, where five times as many subvolumes were available as compared to SIV, 20 – 80 classes were generated resulting in averages with an accordingly higher number of subvolumes. Since most averages did not differ substantially from each other, only 5 classes were chosen in the final classification. Averaging was carried out in Fourier space, so that Fourier coefficients falling in the region of the missing wedge could be excluded from the summation. One cycle of the first variant of subvolume processing can be summarized as follows ( Fig. 1a ): Multivariate data analysis: The subvolumes were re-interpolated based on the stored alignment parameters, then masked so that only voxels representing the molecular volume were active in the analysis. An eigenvector/eigenvalue decomposition was performed and factorial coordinates of each subvolume were computed. Classification: Factorial coordinates of the subvolumes, corresponding to the most significant eigenvectors were clustered with a hierarchical ascendant algorithm. Subvolumes of the desired size were extracted from the tomograms and were averaged separately for each cluster. Reference selection: Multiple references were selected from the class averages. The references were band-pass filtered and an apodized real space mask was applied. Multi-reference alignment: Subvolumes were re-extracted and cross-correlated with each reference. Rotational alignment was achieved by an orientation grid search, translational alignment by locating the position of the correlation peak. The new alignment parameters of the subvolumes were selected from the cross-correlation with the reference that produced the highest correlation peak, and stored for the next cycle. The second variant of subvolume processing was applied to the SIV data sets only. This variant is based on a reference-free alignment method dubbed “alignment by classification” ( Dube et al., 1993 ). In the cited study, the method was applied to two-dimensional images of a portal protein of a bacteriophage, and in order not to introduce any symmetry bias, the authors first aligned the noisy images translation-ally to a rotationally averaged reference, then utilized classification procedures to find similar images in similar rotational orientations. In our three-dimensional case, we aligned the raw subvolumes to a single, rotationally averaged reference only at the very beginning. To avoid any reference bias in subsequent alignment cycles, class averages were aligned, which have a higher SNR than raw subvolumes. The alignment transformations obtained from mutual alignment of class averages were then applied to the constituent members of each class. At this point, we continued processing with classification to split the subvolumes into as many classes of different structure and orientation as possible. When the distribution of orientations is continuous, a wider spectrum of the structural variance can be captured by generating a large number of classes. Considering the trade-off between number of classes generated and SNR improvement by averaging subvolumes, we usually chose 50 classes as a reasonable compromise. One cycle of this processing scheme can be summarized as follows ( Fig. 1b ): Multivariate data analysis: The subvolumes were re-interpolated based on the stored alignment parameters, then masked so that only voxels representing the molecular volume were active in the analysis. An eigenvector/eigenvalue decomposition was performed and factorial coordinates of each subvolume were computed. Classification: Factorial coordinates of the subvolumes, corresponding to the most significant eigenvectors were clustered with a hierarchical ascendant algorithm. Subvolumes of the desired size were extracted from the tomograms and were averaged separately for each cluster. Alignment: The class averages were band-pass filtered and aligned with respect to each other. Rotational alignment was achieved by an orientation grid search, translational alignment by locating the position of the correlation peak. Spatial transformation: The resulting incremental changes in alignment obtained for each class average were applied to the stored alignment parameters of the constituent members of each class, and the resulting new set of parameters was stored for the next cycle. No re-interpolation of subvolumes was carried out at this point. Model calculations Density maps were computed with the program “pdb2mrc” from the EMAN package ( Ludtke et al., 1999 ). Five maps were generated from atomic coordinates which included copies of a gp120 glycoprotein simulating artificial envelope spikes: a monomer, a dimer, and three trimers. The trimers differed in the way the density was distributed in the head part and the leg part of the spikes. From these density maps, phantoms were assembled by inserting the spike maps into the surface of a spherical vesicle-like structure and filling the remaining space with a Poisson-distributed density as a model for amorphous ice. The density profile of the vesicle-like membrane structure was designed to mimic a lipid bilayer and was also generated based on atomic coordinates of a patch of lipid ( Feller et al., 1997 ). The sampling distance for the phantoms was 0.5 nm. Fifty phantoms were used for the model calculations. Each contained about 75 spikes that were randomly distributed on the spherical surface and also randomly rotated about the spike axis or surface normal. Spike positions were generated with a statistically uniform distribution on the unit sphere ( Marsaglia, 1972 ), so that each surface point was equally likely to be selected. At each computed spike position, one of the five maps was randomly selected with equal probability. With this arrangement, preferential orientations of the spikes are unlikely to occur. The phantoms were band-pass filtered in Fourier space with an apodized filter that approximated roughly the contrast transfer function of the experimental data (defocus of 4 μm at 300 kV). The location of the first zero at this defocus value corresponds to a low-pass cutoff at 1/(2.85 nm). Additional wedge filters were applied to simulate tomographic data with tilt angle ranges of ±48°, ±60°, or ±75°. Thus, the missing data amounts to a wedge with angles 84°, 60°, and 30°, respectively. For each of these three cases, subvolumes were extracted and classified. Subsequently the class memberships of the subvolumes were compared with the original input to assess the accuracy of the classification.
📊 Figures
Figure 1
(a) Flowchart of the multi-reference alignment and classification protocol. (b) Flowchart of the classification by alignment protocol.
Figure 2
(a) Section through a tomogram of myosin V in the inhibited state, adsorbed on a lipid monolayer, showing the para-crystalline, hexagonal arrangement of the u201cflower motifsu201d with one of the flo...
Figure 3
Classification of myosin V petal motifs
The classification was carried out with the aligned subvolumes. The top row shows five classes (au2013e), the number of subvolumes that contributed to the class averages is indicated at the bottom lef...
Figure 4
Classification of myosin V petal motifs
The classification was carried out with projections of the subvolumes instead of the subvolumes themselves. Refer to Fig. 3 for further details.
Figure 5
Spatial distribution of the tilt axis direction for the SIV specimen
The tilt axis direction is calculated with respect to the spike coordinate frame and mapped onto the surface of a unit sphere, which is then converted to a two-dimensional representation with a sinuso...
Figure 6
Classification of SIV spikes
Eight classes were generated (columns 1 u2013 8). Row (a): The mask used for the multivariate data analysis included the head and leg region of the spike, but not the slightly curved membrane. For row...
Figure 7
Distribution of the tilt axis directions per class
The classes were produced by the multi-reference alignment scheme. Directions of the tilt axis for the same spikes as in Fig. 5 are plotted here separately for each of eight classes. The relatively un...
Figure 8
Distribution of the tilt axis direction per class
The classes were produced by the alignment by classification scheme. For other details refer to Fig. 7 .
Figure 9
Model calculations
(au2013e) Surface renderings of density maps of artificially created spike structures. It should be noted that these spike structures are not intended to mimic actual spikes faithfully. (a) monomer, (...
Figure images are served from the NIH/NLM PubMed Central Open Access Subset or Europe PMC; copyright remains with the publishers and authors.
💬 Discussion
0 commentsNo comments yet. Be the first to start a discussion!
Leave a Comment