Abstract
Red blood cells (RBCs) (erythrocytes) are the simplest primary human cells, lacking nuclei and major organelles and instead employing about a thousand proteins to dynamically control cellular function and morphology in response to physiological cues. In this study, we define a canonical RBC proteome and interactome using quantitative mass spectrometry and machine learning. Our data reveal an RBC interactome dominated by protein homeostasis, redox biology, cytoskeletal dynamics, and carbon metabolism. We validate protein complexes through electron microscopy and chemical crosslinking and, with these data, build 3D structural models of the ankyrin/Band 3/Band 4.2 complex that bridges the spectrin cytoskeleton to the RBC membrane. The model suggests spring-like compression of ankyrin may contribute to the characteristic RBC cell shape and flexibility. Taken together, our study provides an in-depth view of the global protein organization of human RBCs and serves as a comprehensive resource for future research.
🔬 Techniques
🔭 Microscopes
🧬 Organisms
💻 Software
✨ Fluorophores
🧪 Sample Preparation
🔬 Cell Lines
🏭 Microscope Brands
🧪 Reagent Suppliers
📷 Detectors
💻 Software Details
💻 Code & Software
MSblender is a statistical tool for merging database search results from multiple database search engines for peptide identification based on...
Wrapper for running MSBlender for an experiment with multiple mxZMLs
💾 Data Repositories
🏛️ Research Organizations (ROR)
Affiliated research institutions:
📋 Methods
Data and code availability statement
Raw and processed protein mass spectrometry data have been deposited in the MASSIVE/ProteomeXchange database in entry PXD030050, available at doi: 10.25345/C5R00B . The 20S proteasome cryo-EM structure was deposited in the Electron Microscopy Data Bank as entry EMD-24822. Electron micrograph datasets have been deposited in EMPIAR as entry EMPIAR-10848. Supporting data files, including descriptions of biochemical fractionations, feature matrices, integrated 3D models, and training and test sets, have been deposited at Zenodo as accession doi: 10.5281/zenodo.6465381 .
Native protein extraction
All blood cell types and plasma were purchased from Gulf Coast Regional Blood Center (Houston, TX). The most recently drawn blood was requested for each order of cells/plasma, and the orders were shipped on ice overnight to ensure cells’ freshness. Further steps (described below) were performed to improve cell type homogeneity. Typically, 1–4 mg of extract was 0.45μm filtered (Ultrafree-MC HV Durapore PVDF, Millipore) and fractionated by HPLC chromatography. Detergent and salt were minimized where possible to avoid perturbing protein assemblies. The protein concentration was determined by Bio-Rad Protein Assay or Bio-Rad DC Protein Assay (for samples containing detergent). For additional details on types of fractionations and sample-related information, see the Zenodo data repository.
Red blood cells
According to the supplier, leukocytes were reduced via centrifugation of whole blood. However, we took additional steps to further eliminate contaminating plasma proteins and other cell types, namely leukocytes, reticulocytes, and platelets. Each blood unit was stored at 4 °C for 5 days after receipt in order to allow reticulocytes to mature into RBCs before extracting proteins ( Urbina and Palomino, 2018 ). To further reduce leukocytes and platelets, RBCs were washed with PBS pH 7.4 (Gibco, Thermo Scientific, CA, USA) at 200 ×g for 15 minutes for 3 times. At the final wash, the 1 mm top layer of the cells (residual platelets, WBCs, etc) was removed. Cell samples and buffers were kept at 0–4 °C during all steps. Purified RBCs were lysed with 5 volumes of hypotonic lysis solution (5 mM Tris-HCl pH 7.4) containing EDTA-free protease inhibitor cocktail tablets (Sigma) and phosphatase inhibitor tablets (PhosSTOP EASY, Roche). White ghosts and hemolysate were separated from each other by centrifugation at 21 000 ×g for 40 min. The protein concentration of hemolysate was measured, and the samples were snap-frozen with liquid nitrogen until further use. For ghosts preparation, excess hemoglobin (Hgb) was reduced by washing white ghosts ~10 times with a hypotonic solution until the white ghosts appeared pale pink to almost white. Ghosts were then dissolved with the appropriate detergent (see Zenodo data repository for details), and the protein concentration was measured with a detergent-compatible Bio-Rad DC Protein Assay. For ghosts dissolved in Diisobutylene/Maleic Acid (DIBMA, Cube Biotech) ( Oluwole et al., 2017 ), DIBMA was added to white ghosts to the final concentration 2.5% (50 mM HEPES pH 7.5), and the sample was rotated at 4 °C overnight ( Hagn et al., 2018 ). The ghosts dissolved in 2.5% DIBMA were then treated as for other detergent-dissolved white ghosts. For hemolysate preparations, remnant white ghosts were removed by centrifugation at 21,000 ×g for 40 min at 4 °C, and the supernatant treated with Hemoglobind (Biotech Support Group) in order to bind and remove free Hgb. Protein concentrations of the Hgb-reduced samples were measured by Bradford colorimetric assay (Bio-Rad) prior to biochemical fractionation. iPSC-derived erythroblasts Hematopoietic differentiation from iPSCs to hematopoietic stem and progenitor cells (HSPCs) was induced according to ( Leung et al., 2018 ). To induce erythroid differentiation from HSPCs, day-15 HSPCs were specified using a 2-step suspension culture system consisting of Serum-Free Expansion Medium II, 2 mM of L-glutamine, and 100 mg/mL of primocin at 37°C in normoxic, 5% carbon dioxide conditions. Between days 15 and 20, this base was supplemented with 100 ng/mL of human stem cell factor, 40 ng/mL of IGF1, 5×10 −7 M of dexamethasone, and 0.5 U/mL of hEPO and between days 20 and 25 with 4 U/mL of hEPO. Protein was extracted from day 25 differentiated cells, at which stage the differentiation cultures consisted of polychromatic and orthochromatic erythroblasts ( Vanuytsel et al., 2018 ). Platelets Prostaglandin E1 (PGE1), the aggregation inhibitor, was added to platelet rich plasma (PRP) to the final concentration of 1 mM in order to stop the activation of platelets. To remove contaminating red and white blood cells, contaminating RBCs and WBCs were pelleted at 100 × g for 10 mins at 4 °C with no brake applied. The supernatant (PRP) was then washed two times with 1 volume platelet-wash buffer: 1 volume PRP by centrifugation at 400 × g for 10 min at 4 C with no brake. The wash buffer contains 1 μM PGE1, 10 mM sodium citrate, 150 mM NaCl, 1 mM EDTA, 1% (w/v) glucose pH 7.4. Native protein complexes from platelets were obtained by incubating the washed pallet with Pierce IP lysis buffer (25 mM Tris•HCl pH 7.4, 150 mM NaCl, 1% NP-40, 1 mM EDTA, 5% glycerol) (Thermo Scientific) containing protease and phosphatase inhibitors at 4 °C for 5 minutes. The lysate was clarified at 14,000 × g for 10 mins. Plasma Plasma separated from whole blood by centrifugation was purchased and shipped in anticoagulant CPD. The received plasma still contained other blood cell types, so these cells were removed through centrifugation. Plasma was centrifuged at 5,000 × g for 15 mins at 4°C, and only the top part of supernatant was collected. The supernatant was treated with Albuviod (Biotech Support Group) to reduce the presence of serum albumin (HSA). The HSA reduced supernatant was then concentrated in a 3 kD MWCO Amicon Ultra filter unit (MilliporeSigma, Burlington, MA).
Show full methods section
Data and code availability statement
Raw and processed protein mass spectrometry data have been deposited in the MASSIVE/ProteomeXchange database in entry PXD030050, available at doi: 10.25345/C5R00B . The 20S proteasome cryo-EM structure was deposited in the Electron Microscopy Data Bank as entry EMD-24822. Electron micrograph datasets have been deposited in EMPIAR as entry EMPIAR-10848. Supporting data files, including descriptions of biochemical fractionations, feature matrices, integrated 3D models, and training and test sets, have been deposited at Zenodo as accession doi: 10.5281/zenodo.6465381 .
Native protein extraction
All blood cell types and plasma were purchased from Gulf Coast Regional Blood Center (Houston, TX). The most recently drawn blood was requested for each order of cells/plasma, and the orders were shipped on ice overnight to ensure cells’ freshness. Further steps (described below) were performed to improve cell type homogeneity. Typically, 1–4 mg of extract was 0.45μm filtered (Ultrafree-MC HV Durapore PVDF, Millipore) and fractionated by HPLC chromatography. Detergent and salt were minimized where possible to avoid perturbing protein assemblies. The protein concentration was determined by Bio-Rad Protein Assay or Bio-Rad DC Protein Assay (for samples containing detergent). For additional details on types of fractionations and sample-related information, see the Zenodo data repository.
Red blood cells
According to the supplier, leukocytes were reduced via centrifugation of whole blood. However, we took additional steps to further eliminate contaminating plasma proteins and other cell types, namely leukocytes, reticulocytes, and platelets. Each blood unit was stored at 4 °C for 5 days after receipt in order to allow reticulocytes to mature into RBCs before extracting proteins ( Urbina and Palomino, 2018 ). To further reduce leukocytes and platelets, RBCs were washed with PBS pH 7.4 (Gibco, Thermo Scientific, CA, USA) at 200 ×g for 15 minutes for 3 times. At the final wash, the 1 mm top layer of the cells (residual platelets, WBCs, etc) was removed. Cell samples and buffers were kept at 0–4 °C during all steps. Purified RBCs were lysed with 5 volumes of hypotonic lysis solution (5 mM Tris-HCl pH 7.4) containing EDTA-free protease inhibitor cocktail tablets (Sigma) and phosphatase inhibitor tablets (PhosSTOP EASY, Roche). White ghosts and hemolysate were separated from each other by centrifugation at 21 000 ×g for 40 min. The protein concentration of hemolysate was measured, and the samples were snap-frozen with liquid nitrogen until further use. For ghosts preparation, excess hemoglobin (Hgb) was reduced by washing white ghosts ~10 times with a hypotonic solution until the white ghosts appeared pale pink to almost white. Ghosts were then dissolved with the appropriate detergent (see Zenodo data repository for details), and the protein concentration was measured with a detergent-compatible Bio-Rad DC Protein Assay. For ghosts dissolved in Diisobutylene/Maleic Acid (DIBMA, Cube Biotech) ( Oluwole et al., 2017 ), DIBMA was added to white ghosts to the final concentration 2.5% (50 mM HEPES pH 7.5), and the sample was rotated at 4 °C overnight ( Hagn et al., 2018 ). The ghosts dissolved in 2.5% DIBMA were then treated as for other detergent-dissolved white ghosts. For hemolysate preparations, remnant white ghosts were removed by centrifugation at 21,000 ×g for 40 min at 4 °C, and the supernatant treated with Hemoglobind (Biotech Support Group) in order to bind and remove free Hgb. Protein concentrations of the Hgb-reduced samples were measured by Bradford colorimetric assay (Bio-Rad) prior to biochemical fractionation. iPSC-derived erythroblasts Hematopoietic differentiation from iPSCs to hematopoietic stem and progenitor cells (HSPCs) was induced according to ( Leung et al., 2018 ). To induce erythroid differentiation from HSPCs, day-15 HSPCs were specified using a 2-step suspension culture system consisting of Serum-Free Expansion Medium II, 2 mM of L-glutamine, and 100 mg/mL of primocin at 37°C in normoxic, 5% carbon dioxide conditions. Between days 15 and 20, this base was supplemented with 100 ng/mL of human stem cell factor, 40 ng/mL of IGF1, 5×10 −7 M of dexamethasone, and 0.5 U/mL of hEPO and between days 20 and 25 with 4 U/mL of hEPO. Protein was extracted from day 25 differentiated cells, at which stage the differentiation cultures consisted of polychromatic and orthochromatic erythroblasts ( Vanuytsel et al., 2018 ). Platelets Prostaglandin E1 (PGE1), the aggregation inhibitor, was added to platelet rich plasma (PRP) to the final concentration of 1 mM in order to stop the activation of platelets. To remove contaminating red and white blood cells, contaminating RBCs and WBCs were pelleted at 100 × g for 10 mins at 4 °C with no brake applied. The supernatant (PRP) was then washed two times with 1 volume platelet-wash buffer: 1 volume PRP by centrifugation at 400 × g for 10 min at 4 C with no brake. The wash buffer contains 1 μM PGE1, 10 mM sodium citrate, 150 mM NaCl, 1 mM EDTA, 1% (w/v) glucose pH 7.4. Native protein complexes from platelets were obtained by incubating the washed pallet with Pierce IP lysis buffer (25 mM Tris•HCl pH 7.4, 150 mM NaCl, 1% NP-40, 1 mM EDTA, 5% glycerol) (Thermo Scientific) containing protease and phosphatase inhibitors at 4 °C for 5 minutes. The lysate was clarified at 14,000 × g for 10 mins. Plasma Plasma separated from whole blood by centrifugation was purchased and shipped in anticoagulant CPD. The received plasma still contained other blood cell types, so these cells were removed through centrifugation. Plasma was centrifuged at 5,000 × g for 15 mins at 4°C, and only the top part of supernatant was collected. The supernatant was treated with Albuviod (Biotech Support Group) to reduce the presence of serum albumin (HSA). The HSA reduced supernatant was then concentrated in a 3 kD MWCO Amicon Ultra filter unit (MilliporeSigma, Burlington, MA).
White blood cells
In order to broadly assess WBC proteins, we analyzed a buffy coat, consisting of all WBCs (lymphocytes and granulocytes), as well as some RBCs and platelets. WBCs were separated from the other cell types using Histopaque 1077 (Sigma) according to the manufacturer’s protocol. Briefly, the buffy coat was diluted with PBS Gibco at a 1:1 ratio. Buffy coat and Histopaque were held at room temperature to ensure proper cell separation. PGE1 was added to the diluted buffy coat to a final concentration of 1 μM to prevent platelet activation and cell aggregation. WBCs were first purified using Histopaque 1077 where the “ring” of lymphocytes, the layer between plasma and histopaque after centrifugation, and the white layer containing granulocytes on top of the RBC layer at the bottom were kept for further purification steps. Contaminating RBCs were lysed by addition of hypotonic solution, and the mixture of lymphocytes and granulocytes were pelleted at 1,000 g for 10 minutes. Native protein extract of WBCs was prepared by adding Pierce IP lysis buffer, and the extract was clarified by centrifuging at 14,000 × g for 10 mins, prior to biochemical fractionation.
HPLC chromatography
Hgb-reduced hemolysate and detergent-dissolved ghosts were fractionated on a Dionex UltiMate3000 HPLC system consisting of an RS pump, Diode Array Detector, PCM-3000 pH and Conductivity Monitor, Automated Fraction Collector (Thermo Scientific, CA, USA) and a Rheodyne MXII Valve (IDEX Health & Science LLC, Rohnert Park, CA) using biocompatible PEEK tubing and various columns. The columns we used were size exclusion chromatography, ion exchange separations (mixed bed or triple-phase WAXWAXCAT), or hydrophobic interaction chromatography. The sample loaded was 1–4 mg protein as measured by the BioRad Protein Assay (hemolysate) or DC Protein Assay (detergent-dissolved ghosts) as appropriate to the sample buffer. Fractions were collected into 96-deep well plates. Size exclusion: BioSep-SEC-s4000 600 × 7.8 mm ID, particle diameter 5 μm, pore diameter 500 Å (Phenomenex, Torrance, CA) or higher Mw BioBasic AX HPLC Columns, 600 × 7.8 mm ID 5μm particle size, 300Å pore size were used. Unless otherwise specified the sample was 200 μl, low rate 0.5 ml/min, with fraction collection every 45 seconds, and mobile phase was PBS pH 7.4 (Gibco). For column calibration, molecular weight standards (Sigma -Aldrich, MWGF1000, 2–5 ug each of carbonic anhydrase (C7025), β-amylase (A8781), and bovine serum albumin (A8531)) were fractionated to ensure that the column can separate proteins based on their sizes correctly before the hemolysate/ghosts samples were fractionated. Mixed bed ion exchange: Poly CATWAX A (PolyLC Mixed-Bed WAX-WCX) 200 × 4.6 mm ID, Particle diameter 5 μm, pore diameter 100 Å (PolyLC Inc., Columbia, MD). The bed contains the cation-exchange (PolyCAT A) and anion-exchange (PolyWAX LP) materials in equal amounts. A 200–250 μl sample was loaded at ≤ 40 mM NaCl, and eluted with a 1-hour salt gradient at 0.5 ml/min with collections of 0.5 ml fractions. Gradient elution was performed with Buffer A (10 mM Tris-HCl pH 7.5, 5% glycerol, 0.01% NaN3), and 0–70% Buffer B (1.5 M NaCl in Buffer A). Triple phase ion exchange (WWC): Three columns, each 200 × 4.6 mm ID, particle diameter 5 μm, pore diameter 100 Å, were connected in series in the following order: two PolyWAX LP columns followed by a single PolyCAT A (PolyLC, Inc, Columbia, MD). Loading, buffers, and fraction collection were as for mixed bed ion exchange above with slight modifications in low rate and elution from the methods of ( Havugimana et al., 2012 ). The low rate was either 0.25 ml/min with a 120 min gradient from 5–100 %B. For the separation of nuclear extracts, the gradient was modified to a 115-minute multiphasic elution from 5–100% Buffer B. Hydrophobic interaction: The ProPac HIC-10 hydrophobic interaction chromatography column (Thermo Scientific) had the following characteristics: 4.6 × 250 mm, Amide/Ethyl (phase), 5 μm particle size, and 300 Å pore size. A 200–250 μl sample was loaded at ~250 mM (NH 4 ) 2 SO 4 , and eluted with a 100-min low salt gradient at 0.5 ml/min with collections of 0.5 ml fractions. Gradient elution was performed with Buffer A (2 M (NH 4 ) 2 SO 4 in 0.1M NaH 2 PO 4 pH 7.0), and 0–70% Buffer B (0.1 M NaH 2 PO 4 pH 7.0).
Mass spectrometry sample preparation
Samples were prepared for mass spectrometry in 96-well plate format using ultrafiltration and in-solution digestion protocols ( Wan et al., 2015 ). Plates were sealed with transparent film during incubation steps. Ultrafiltration was performed with an AcroPrep Advance 96-filter plate, 3kD MWCO (Pall) using a vacuum manifold (QIAvac 96 or Multiwell, Qiagen) at −0.75 Bar. Before filtering samples, preservatives and remnant polymers that could interfere with MS experiment were removed from the plate by sequential filtration of 100 μl LC/MS quality water and 100μl trypsin digestion buffer (50 mM Tris-HCl, pH 8.0, 2 mM CaCl 2 ). Samples were concentrated to 100 μl, diluted 2-fold with trypsin digestion buffer, and concentrated again to a final volume of 50–100 μl before being transferred back to a 96-deep well plate. 50 μl 2,2,2,-trifluoroethanol (TFE) was added and samples were reduced with TCEP (Bond-Breaker, Thermo) at a final concentration of 5 mM for 30 min. 37°C. Iodoacetamide was added to 15 mM and plates were incubated in the dark 30 min at room temperature. Alkylation was quenched by the addition of DTT to 7.5 mM. TFE was diluted to + 1 selected for collision-induced dissociation (CID). Dynamic exclusion was active with 60 s exclusion for ions selected once within a 60 s window. For some experiments, a similar top speed method was used with dynamic exclusion of 30 s for ions selected once within a 30 s window and high energy-induced dissociation (HCD) collision energy 31% stepped +/−4%. All MS2 scans were centroid and done in rapid mode. Orbitrap Lumos: top speed HCD with full precursor ion scans (MS1) collected at 120,000 m/z resolution. Monoisotopic precursor selection and charge-state screening were enabled using Advanced Peak Determination (APD), with ions of charge > + 1 selected for high energy-induced dissociation (HCD) with collision energy 30% stepped +/− 3%. Dynamic exclusion was active with 20 s exclusion for ions selected twice within a 20s window. All MS2 scans were centroid and done in rapid mode. For identification of DSSO cross-linked peptides, peptides were resolved using a reverse phase nano low chromatography system with a 115 min 3–42% acetonitrile gradient in 0.1% formic acid. The top speed method collected full precursor ion scans (MS1) in the Orbitrap at 120,000 m/z resolution for peptides of charge 4–8 and with dynamic exclusion of 60 sec after selecting once, and a cycle time of 5 sec. CID dissociation (25% energy 10 msec) of the cross-linker was followed by MS2 scans collected in the orbitrap at 30,000 m/z resolution for charge states 2–6 using an isolation window of 1.6. Peptide pairs with a targeted mass difference of 31.9721 were selected for HCD (30% energy) and collection of rapid scan rate centroid MS3 spectra in the Ion Trap.
Computational analyses of peptide mass spectra
The human reference proteome was downloaded from Uniprot.org (UniProt Consortium, 2019) in August 2018 (20,858 entries). This reference proteome can be found in the MASSIVE/ProteomeXchange database in entry PXD030050, available at doi: 10.25345/C5R00B . Mass spectral peptide matching was performed with MSGF+, X!Tandem, and Comet-2013020, each run with 10 ppm precursor tolerance, and allowing for fixed cysteine carbamidomethylation (+57.021464) and optional methionine oxidation (+15.9949). Peptide search results were integrated with MSBlender ( Kwon et al., 2011 ), https://github.com/marcottelab/msblender , https://github.com/marcottelab/run_msblender ). For DSSO cross-linked experiments, inter-protein cross-links were identified using the XlinkX ( Klykov et al., 2018 ) node in ProteomeDiscover 2.2 (ThermoScientific). FDR for all analyses was kept at 1%. RBC proteome Identification of RBC proteome In order to identify a consensus proteome of RBCs, we constructed a supervised classifier, trained on proteomics and RNA-seq data from different blood cell types (WBCs, reticulocytes, platelets, plasma proteins, and RBC precursor cells including different stages of erythroblasts). In addition to our own data, MS data from different blood cell types were retrieved from the PRIDE database and several of the datasets were reanalyzed using ProteomeDiscover 2.2 as noted in Table S6 in the Zenodo data repository. A classifier feature matrix was assembled wherein rows corresponded to proteins observed in each of 45 external MS experiments, 12 in-house CF-MS experiments (each summarizing overall protein abundances as the sum of Peptide Spectral Matches (PSMs) for a given protein across all fractions in a fractionation experiment), and 26 RNA-seq experiments (see RBC_proteome_featmat.csv in Zenodo repository), for a total of 83 features (columns). As gold standard RBC proteins, we used the set of 859 proteins that all three previous MS studies on RBCs identified ( Bryk et al., 2017 , Goodman et al., 2007 and Lange et al., 2014 ) as a positive set ( i.e ., most likely to be true RBC proteins). As a negative gold standard ( i.e ., most likely not RBC proteins), we selected all proteins found in the set of all MS and RNA-seq experiments used to assemble the feature matrix curated in Uniprot (20,504 proteins) but not found in any of the three prior RBC studies, which identified a total of 3,520 proteins across the three studies. Therefore, the total number of negative gold standard proteins is 16,894. Proteins found in at least one but not all 3 prior RBC studies were considered as unknowns and to be classified. We divided the gold standard proteins into train and test sets (80:20, train:test). We then used TPOT, an autoML wrapper of scikit-learn machine learning functions ( Olson et al., 2016 ), to perform all training steps, including identifying the best classifier and hyperparameters based on 5-fold cross-validation on the training dataset. All training steps excluded the test set. The best classifier was the RandomForest classifier with TPOT discovered hyperparameters. Application of the classifier to the entire set of proteins resulted in an RBC likelihood score/confidence score for each protein where 1 indicates the highest likelihood that the protein derives from mature RBCs and 0 indicates a non-RBC protein. Precision and recall were calculated from training (687 positive, 13,665 negative) and test (172 positive, 3,229 negative) set proteins for Figure 1A and 1B , respectively. The false discovery rate was calculated from the test set. In terms of feature importance (measured based on impurity using the scikit-learn function Feature Importance and provided in the Zenodo repository), we observed the strongest 3 features to be our CF-MS experiment on hemolysate and two published MS experiments on CD71 − and CD71 + cells (RBCs and reticulocytes, respectively) from ( Moura et al., 2018 ) reporting data from in vitro culture-derived reticulocytes (derived from CD34+ cells) and cultured reticulocytes circulated overnight in an ex vivo circulation system ( Moura et al., 2018 ). Overall, RNA-based features provided some additional power, but made a relatively minor contribution.
Defining RBC complexes based on the CF-MS datasets
Assembly of features for scoring putative protein interactions We performed 30 fractionation experiments comprising 1,944 total biochemical fractions across all fractionations. For each experiment, we assembled an elution matrix of all identified proteins (rows) by fractions (columns) containing the PSMs for each protein in each fraction, normalizing PSMs to 1 on a per protein basis for each experiment. In addition, we concatenated the elution matrices across all 1,944 fractions as one additional matrix. (Note that cross-linked data from hemolysate and ghosts were not used for training.) Next, we calculated a series of all-by-all pairwise scores between proteins for individual matrices and the concatenated matrix. We focused our analysis on well-observed proteins, so we additionally filtered for proteins with at least 60 total PSMs observed across the 30 combined fractionations. The scores/features were as follows: (1) Pearson’s r, (2) Spearman’s rho, (3) Euclidean distance, (4) Bray-Curtis similarity, (5) stationary cross-correlation, (6) covariance, and (7) hypergeometric score for the co-occurrence of proteins in fractions with repeated sampling of fractions ( Drew et al., 2017 ). All features/scores were calculated with added Poisson noise as in ( Drew et al., 2017 ). Euclidean distance and Bray-Curtis similarity scores were inverted and normalized to a max score of 1. For each individual matrix from each fractionation, features 1–6 were calculated. Additionally, the average of each type of score/feature was calculated from all individual matrices excluding the concatenated matrix. For the concatenated matrix of all fractionations, features 1–7 were calculated. Features were joined to create a final feature matrix composed of 4,131,128 rows (protein pairs) and 193 features capturing the similarities between the proteins’ elution profiles. Missing values were filled with zeros. Construction of the gold standard protein complex training and test sets We used known human protein complexes from the CORUM database ( Giurgiu et al., 2019 ) as a gold standard positive set of stable protein-protein interactions. As most of these CORUM complexes do not exist in RBCs, we further selected those CORUM complexes in which >50% of the protein subunits/members were RBC proteins at the 1% FDR level. The resulting 123 known CORUM complexes with >50% of the protein subunits belonging to the RBC proteome were divided into positive training and test complexes according to the scheme from Drew et al., 2017 . In short, we defined each positive protein interaction as a pair of proteins that are part of the same CORUM complex. A negative protein interaction was defined as a pair of proteins that are both found in the set of CORUM complexes but not within the same CORUM complex. Complexes with over 30 members were removed so as not to skew performance measurements as per ( Drew et al., 2017 ). Any overlapping interactions between training and test sets were removed such that the sets were fully disjoint. The final pairwise protein interaction training/test sets consisted of 461/299 and 14,389/11,408 positive and negative interactions, respectively. The final protein complex training/test sets consisted of 62/61 complexes. The complete lists of training/test interactions and complexes are available in the Zenodo repository.
Identification of interacting proteins by supervised machine learning
We again utilized TPOT to train our machine learning model to find pairwise interactions among members of protein complexes. We discovered optimal hyperparameters for an ExtraTree classifier with 5-fold cross-validation of the training interactions, with an area under the precision-recall curve (AUPR) of 0.45. We then trained an ExtraTree with TPOT discovered hyperparameters, and the resulting model was applied to the entire feature matrix to give a CF-MS score to each pair of proteins, with higher scores corresponding to higher confidence in the proteins interacting (specifically, being subunits in the same multi-protein assembly). Precision, recall, and false discovery rates were calculated from the 299 positive and 11,408 negative withheld test set interactions. Clustering of interacting protein pairs to define multiprotein assemblies Interaction Interaction scores above a 15% false discovery rate threshold (CF-MS score ≥ 0.17) were input into R igraph cluster_walktrap to define coherent protein complexes. The walktrap algorithm’s reweighted edges between proteins were reformatted to a dendrogram and cut at intervals to obtain a nested hierarchy of complexes as in ( McWhite et al., 2020 ). Cuts closer to the root of the dendrogram result in larger complexes and cuts closer to the tips defined smaller subcomplexes. More details on how to read the hierarchy of complexes can be found in Table S4 - RBC_interactome in the Zenodo repository.
Negative stain electron microscopy
We reduced Hgb from hemolysate which was then passed through 100 kDa filter ultracentrifugation filter (Amicron). 4 uL of hgb-reduced and filtered hemolysate was applied to a glow-discharged 400-mesh continuous carbon grid. After allowing the sample to adsorb for 1 min, the sample was negatively stained with five consecutive droplets of 2% (w/v) uranyl acetate solution, blotted to remove residual stain, and air-dried in a fume hood. Grids were imaged using an FEI Talos TEM (Thermo Scientific) equipped with a Ceta 16M detector. Micrographs were collected manually using TIA v4.14 software at a nominal magnification of x73,000, corresponding to a pixel size of 2.05 Å/pixel. CTF estimation, particle picking, and 2D class averaging were performed using both RELION v3 ( Zivanov et al., 2018 ) and cryoSPARC v2.12.4 ( Punjani et al., 2017 ). Three negative stain datasets were collected. The first dataset collected contained ~220 micrographs of pooled HPLC size exclusion fractions 1–9. ~2,500 particles were manually picked and processed in cryoSPARC to produce the TPP2 structure in Figure 3B – C . Two datasets were collected of hemolysate after being passed through a 100 kDa filter, one as prepared and the other at 1:100 dilution. For the diluted sample, ~400 micrographs were collected and ~42,500 particles were picked using Topaz ( Bepler et al., 2019 ). The resulting particles were used to generate the PRDX2 structure in Figure 3B – C and the 2D class averages in Figure S3B . For the non-dilute sample, ~230 micrographs were collected and ~1,500 proteasome particles were manually picked followed by classification in RELION Figure S3C .
Cryo-EM grid preparation and data collection
C-flat holey carbon grids (CF-1.2/1.3, Protochips Inc.) were glow-discharged for 1 min using a Solarus 950 plasma cleaner (Gatan). 2 μL of 0.2 mg/mL graphene oxide (Sigma-Aldrich) was placed onto the grids for 1 min followed by one wash with water. 3 μL of pooled and concentrated hemolysate from HPLC size exclusion fractions 10–20 was placed onto the grid, blotted for 3.5 sec with a blotting force of 0, and rapidly plunged into liquid ethane using an FEI Vitrobot MarkIV operated at 4 °C and 100% humidity. Data was acquired using an FEI Titan Krios TEM (Sauer Structural Biology Laboratory, University of Texas at Austin) operated at 300 keV with a nominal magnification of ×22,500 (1.045 Å/pixel) and defocus ranging from −1.09 to −2.5 μm. Dose-fractionated movies were collected using 20 frames (0.15 sec/frame) over a total of 3 sec with a dose rate of ~2.13 e − /Å 2 /sec and a total exposure of 42.58 e − /Å 2 . A total of 6,606 micrographs were automatically recorded on a K3 detector (Gatan) operated in counting mode using Leginon ( Suloway et al., 2005 ). A full description of the cryo-EM data collection parameters can be found in Table S1 .
Cryo-EM data processing
Motion correction, CTF-estimation and particle picking were performed in Warpv1.0.7 ( Tegunov and Cramer, 2019 ). ~1,000,000 extracted particles were imported into cryoSPARC v2.12.4 for 2D classification, 3D classification and non-uniform 3D refinement ( Figure S2 ). ~60,000 particles were used in the final 20S proteasome reconstruction. The nominal resolution of the map using the gold-standard Fourier Shell Correlation (FSC) at 0.143 is 3.35 Å ( Figure S2 ). A previously solved X-ray crystal structure of the human 20S proteasome, PDB 6RGQ ( Rêgo and Fonseca, 2019 ), was aligned by cross-correlation in UCSF Chimera ( Pettersen et al., 2004 ) and used as an initial model for refinement. The model was then refined through two rounds of molecular dynamics flexible fitting and real space refinement using Namdinator ( Kidmose et al., 2019 ), followed by further refinement in Phenix ( Adams et al., 2002 ). The high correlation to a previous structure shows our RBC derived model is similar to the canonical 20S proteasome. However, due to the decreasing quality of the map away from the center, we were unable to build a full and accurate model that might distinguish subtle differences between the structures.
Integrative 3D modeling
Integrative modeling of the red blood cell membrane complexes consisted of four main stages: 1) gathering data, 2) domain representation and configuring of spatial restraints, 3) system sampling and scoring of restraints, and 4) model validation, as previously described in integrative modeling work ( Gutierrez et al., 2020 ; Kim et al., 2018 ; Russel et al., 2012 ). The python interface of the Integrative Modeling Platform (IMP) was used to model the complexes ( Saltzberg et al., 2019 ) and all associated data, scripts, and outputs can be found at Zenodo data deposit. The following sections describe the method for modeling the ankyrin complex in the open-spring conformation, the ankyrin complex in the closed-spring conformation, the ankyrin complex in the open-spring conformation with metabolic enzymes bound, the ankyrin complex in the closed-spring conformation with metabolic enzymes bound, and the band 4.1 with spectrin subcomplex.
Data used for modeling
Protein structure representations were constructed from known X-ray crystal structures, modeled using I-TASSER ( Yang et al., 2015 ), or modeled domains from the Swiss-Prot database.
Table
S3 indicates the source of each of the protein’s structural models. A total of 156 intra- and intermolecular DSSO crosslinks were used to model the ankyrin complex. An additional 21 DSSO crosslinks were used to incorporate the GAPDH and PGK1 enzymes into the model. The subcomplex of spectrin and band 4.1 was modeled using 41 intra- and intermolecular DSSO crosslinks. Transmembrane regions of proteins were determined from Uniprot ( The UniProt Consortium, 2019 ) annotations. Domain representation and configuring spatial restraints Protein subunits were represented as rigid bodies, chains of rigid bodies, or beads ( Table S3 ). Band 3 was represented using rigid bodies for the transmembrane region and cytoplasmic domain of the protein. These domains were connected with a chain of flexible beads. The transmembrane region of the band 3 dimer ( Arakawa et al., 2015 ) was represented as a single rigid body. Glut1 and band 4.2 were represented as a single rigid body, due to their high C-score ( Table S3 ) from I-TASSER and agreement with intramolecular crosslinks. The transmembrane regions of the GYPA dimer have a known NMR structure and are represented as a single rigid body. The cytoplasmic and extracellular domains of the GYPA dimer were represented using a flexible chain of beads. RhAG and the two RhCE/D proteins were superimposed on the x-ray crystal structure (PDB 3HD6) and treated as a single rigid body ( Gruswitz et al., 2010 ); these proteins had high C-scores and satisfied intermolecular crosslinks between each other. ANK1 was represented as a chain of rigid bodies (both modeled and X-ray crystal structure) with the end represented as flexible beads. Residues 265 to 790 of SPTA1 were used in the model and were represented as a chain of rigid bodies, with each rigid body beginning/ending at the Uniprot annotated spectrin domains. Residues 1,585 to 2,006 of SPTB were represented similarly to SPTA1 with each annotated spectrin domain being a rigid body connected in a chain. Both GAPDH and PGK1 had available X-ray crystal structures that satisfied all intramolecular crosslinks and were represented as rigid bodies. For the band 4.1-spectrin complex, band 4.1 was represented as a single rigid body, while residues 1,929 to 2,386 of SPTA1 was represented as a chain of rigid bodies by spectrin domain and residues 46 to 741 of SPTB was represented as a chain of rigid bodies by spectrin domain. The flexible chains of beads were coarse-grained to 10 residues per bead. The DSSO crosslinkers were modeled using a length of 21 Å. The excluded volume restraint was applied to the 10 residue beads, preventing volumes from occupying the same space. The sequence connectivity restraint was applied between beads. Based on Uniprot annotations, segments of proteins were labeled as either inside the membrane (transmembrane), above the membrane (extracellular), or below the membrane (cytoplasmic) scored using a sigmoid potential. All of these restraints were incorporated into the scoring framework for the model. System sampling and scoring of restraints The 38 rigid bodies were first randomized in an initial configuration, followed by a steepest descent minimization based on connectivity to ensure that neighboring residues are close together and Monte Carlo sampling. Using integrative modeling ( Figure S5 ) ( Russel et al., 2012 ), we built many (600,000) models in parallel (unique starting positions) and tested for their agreement; the combination of annotated transmembrane regions, chemical crosslinks, and known partial structures provided sufficient spatial restraints to converge on a solution. (Note that detailed explanations of assessing sampling exhaustiveness can be found in previous literature ( Viswanath et al., 2017 )). These models were constructed using fixed protein compositions based on crosslinking evidence that implicated a limited number of proteins ( Figure 5A ). The ensembles were clustered based on their scoring parameters, including the clustering precision, which describes the variability between models in the cluster ( Table S2 ). The best scoring cluster was selected to proceed ( Table S2 ).
Model validation
The model cluster was first assessed against the input data. A crosslink is considered to be satisfied if any model in the cluster has the crosslink distance less than 40 Å. The crosslinks that are unsatisfied are believed to represent another conformational state of the complex. The convergence is determined through assessing the exhaustiveness of the sampling ( Viswanath et al., 2017 ). This protocol tests convergence of the model score, whether the model scores were drawn from the same parent distribution, whether the structural clusters include models from each sample proportional to their size (chi-squared and sampling precision), and structural similarity between the model samples ( Table S2 CCC between two sample densities).
Supplementary Material 1
📊 Figures
Figure 1.
Defining the canonical red blood cell (RBC) proteome with high accuracy from a synthesis of protein mass spectrometry and mRNA expression data.
(A) Measured protein and RNA abundances from diverse blood cell types and plasma were used as features for a machine learning classifier to assign confidence scores for proteins whether they belong to...
Figure 2.
Overview of the integrative Co-Fractionation / Mass Spectrometry (CF-MS) workflow used to determine stable RBC protein complexes.
(A) Hemolysate and white ghosts are chromatographically separated and the proteins in each fraction are identified by mass spectrometry. Elution profiles for each protein are graphically represented a...
Figure 3.
Validation of the CF-MS workflow using electron microscopy confirms intact multi-protein complexes.
( A ) Hemolysate from size exclusion chromatography was partitioned into five groups and visualized with negative stain EM. Elution profiles from corresponding mass spectrometry data were used to assi...
Figure 4.
A map of the primary RBC multiprotein complexes.
Thin circles show the clustering hierarchy of protein-protein interactions into complexes for each of 3 clustering thresholds (see the Zenodo data repository for complex memberships and annotations). ...
Figure 5.
Chemical crosslinks confirm mapped interactions and constrain 3D modeling of membrane/cytoskeletal complexes.
(A) Network plot represents crosslinked interactions among proteins (shown as gene names) in RBC cytoskeletal complexes. Line between each node indicates detected crosslink(s). Shaded lines indicate d...
Figure 6.
Reconstructions of band 3-Ank1-accessory protein complexes by integrative 3D modeling suggest Ank1 compression links the membrane to the cytoskeleton.
(A) An overview of the cytoskeletal network supporting the membrane of RBC. A pseudohexagonal network of spectrin heterotetramer underlies the membrane and is anchored to the membrane by the band 3-An...
Figure images are served from the NIH/NLM PubMed Central Open Access Subset or Europe PMC; copyright remains with the publishers and authors.
💬 Discussion
0 commentsNo comments yet. Be the first to start a discussion!
Leave a Comment