Mass Spectrometry
Online ISSN : 2186-5116
Print ISSN : 2187-137X
ISSN-L : 2186-5116
Special Issue: Proceedings of 19th International Mass Spectrometry Conference
Image and Spectral Processing for ToF-SIMS Analysis of Biological Materials
Daniel J. Graham, David G. Castner
著者情報
ジャーナル フリー HTML

2013 年 2 巻 Special_Issue 号 p. S0014

詳細
Abstract

Time-of-flight secondary ion mass spectrometry (ToF-SIMS) instruments can rapidly produce large complex data sets. Within each spectrum, there can be hundreds of peaks. A typical 256×256 pixel image contains 65,536 spectra. If this is extended to a 3D image, the number of spectra in a given data set can reach the millions. The challenge becomes how to process these large data sets while taking into account the changes and differences between all the peaks in the spectra. This is particularly challenging for biological materials that all contain the same types of proteins and lipids, just in varying concentrations and spatial distributions. This data analysis challenge is further complicated by the limitations in the ion yield of higher mass, more chemically specific species, and potentially by the processing power of typical computers. Herein we briefly discuss analysis methodologies including univariate analysis, multivariate analysis (MVA) methods, and some of the limitations of ToF-SIMS analysis of biological materials.

INTRODUCTION

Time-of-flight secondary ion mass spectrometry (ToF-SIMS) generates an information rich yet complex set of spectra. A typical spectrum can contain hundreds to thousands of peaks. With the advent of 2D and 3D imaging in modern ToF-SIMS instruments, one can quickly acquire an enormous amount of data. For example a 2D 256×256 pixel image contains 65,536 spectra. If this is extended to a 3D image, the number of spectra can increase to the millions depending on the number of 2D images acquired during the analysis. The challenge then becomes how to efficiently process and analyze this data in order to extract the pertinent information and gain insight about the sample.

When analyzing biological materials (cells, tissues, etc.), the analyst is presented with the additional challenges of low ion yields and the lack of molecular specificity of many of the resulting peaks.1–3) Low ion yields are the result of the low ionization efficiencies of ToF-SIMS. Typically less than 1% of the material removed from the surface during the primary ion bombardment is ionized. This can be particularly problematic if the analyte of interest is already present at low concentrations at the surface.

Major components present in cells and tissues include proteins and lipids. ToF-SIMS analysis of proteins has been shown to generate mainly fragments of the 20 component amino acids, and not large fragments or even small peptides. It has been shown that the relative intensity of these amino acid fragments encodes information about the protein’s composition,4–8) conformation9–13) and orientation.14–17) However, it would be challenging to determine the identity of individual proteins in a complex mixture such as is found in cells or tissues using ToF-SIMS.8) Many nice examples of lipid analysis with ToF-SIMS have been shown.3,18–25) Control spectra for many common lipids have been published and used to study cells and tissues.3,19,26) However, many of the higher mass, more chemically specific, lipid peaks are not detected, especially when one collects high spatial resolution images using ToF-SIMS. This is partially due to low primary ion currents and adjustments required to the analyzer when using the high spatial resolution settings on current ToF-SIMS systems. The need for these different settings is due to the fact that most modern ToF-SIMS instruments utilize a pulsed primary ion source that can either be optimized for high mass resolution or high spatial resolution, but not both at the same time. The J105 instrument is an exception to this rule as it uses a DC primary ion beam and therefore is capable of collecting high spatial resolution images with good mass resolution.27,28) Even with this advantage, the J105 is still limited by low ion yields from biological materials.

In spite of these limitations, significant work has been done in the analysis of biological materials with ToF-SIMS.1,3–5,8,18–21,23–26,28–50) For this work, ToF-SIMS analysts have used a variety of methods to process and understand their data. These include a combination of traditional manual analysis and the application of multivariate analysis (MVA) methods. In this manuscript, we briefly overview the current data processing methods available for the ToF-SIMS analyst, outline some of the challenges faced for ToF-SIMS analysis of biological materials, and summarize some of the needs to continue to move ToF-SIMS analysis of biological materials to the next level.

CURRENT PRACTICES

ToF-SIMS is a mass spectrometry technique, and as such involves basic mass spectrometry analysis methods such as spectral calibration, peak identification, integration of peak areas, and plotting of peak intensity maps for imaging. The most basic analysis involves searching through the spectra to locate expected peaks from a known component, or identifying dominant peaks to determine the structure of unknown components. Often peak area ratios are created to track changes in chemistry across a set of samples. Other basic analysis methods for imaging include creating region of interest (ROI) spectra from a given region within an image, or creating peak area images from a given peak or set of peaks. These methods can be useful to highlight some differences between the samples, however, focusing on individual peaks or ROIs risks adding user bias to the analysis and ignores the potential information present across the rest of the spectra. To avoid this problem and take advantage of all of the data, many users apply MVA methods to ToF-SIMS spectra and images.22,51–56) Though there is still a fair amount of exploration of new MVA methods, in general the main methods used for processing ToF-SIMS data include principal components analysis (PCA), partial least squares (PLS), multivariate curve resolution (MCR), and maximum autocorrelation factors (MAF). Herein we provide an example of ToF-SIMS analysis of mouse muscle tissue using manual analysis, PCA and MCR to illustrate the strengths and weaknesses of the data processing methods and of the ToF-SIMS instrumentation.

ToF-SIMS image analysis of mouse muscle tissue

ToF-SIMS imaging of mouse muscle tissue was carried out using an Iontof 5-100 instrument equipped with a 25 kV bismuth primary ion source and a 10 kV C60 sputter source. A positive ion image was acquired by summing the signal from 150 sequential analysis cycles using a 200 micron×200 micron raster with Bi3+ at 0.05 pA (∼1.5×1010 ions/cm2) for imaging, followed by a 10 second sputter of a 500 micron×500 micron area using C60++ at 0.45 nA (∼1.1×1013 ions/cm2). This summation protocol was used to increase the signal to noise within the image and to acquire the image data over a thickness of the tissue that would be more comparable with future optical microscopy analysis. The approximate depth of the summed image was 1.5 microns.

The image data was processed using standard manual analysis of peak area images and by MVA using PCA and MCR. For MVA, all peaks above background were selected and peak area images were generated using the Iontof software (Iontof, Münster, Germany). The image data was exported using the iontof .bif6 file format and then imported into the NESAC/BIO Imagegui (NBtoolbox, NESAC/BIO, University of Washington, http://mvsa.nb.uw.edu/) within Matlab (Mathworks, Natick, MA). MCR was carried out using the MCR-ALS gui (Roma Tauler and Anna de Juan, University of Barcelona, http://www.mcrals.info/) accessed through the Imagegui. PCA and MCR were carried out using the Poisson scaled57,58) and mean centered data. PCA scores were used as initial estimates for the MCR analysis. An explanation of PCA and MCR are beyond the scope of this manuscript. Good overviews of each method can be found in the literature.22,51,53,55,59–64) Briefly, PCA looks for the major directions of variance within the data. For PCA processing of ToF-SIMS image data, this results in finding which areas of the image are chemically distinct, and which peaks are responsible for the differences among those areas. The output from PCA is a set of score images and peak loadings. Within these plots, positive scores are shown as brighter regions and correspond with positive loadings. Negative scores are shown as darker regions and correspond with negative loadings. MCR uses an alternating least squares routine to find a set of “pure” components that describe the differences within the data set. MCR was carried out using a non-negativity constraint meaning that the resulting components are restricted to positive values. This means for the MCR component images, the brighter regions show higher intensity of the peaks shown in the component spectra, whereas darker regions show areas of lower or no intensity for those peaks.

Figure 1 shows a series of peaks selected manually by browsing through the full data set. Peaks were randomly selected that showed representative images of the patterns seen within the data. As seen in the figure, peaks could be found that are representative of the muscle cell nuclei (top row Fig. 1), the cell bodies (middle row Fig. 1), and the intercellular regions (bottom row Fig. 1). This manual analysis is fairly quick and allows one to pick out peaks of interest, however it brings definite user bias to the analysis since the analyst chooses which peaks to focus on and present.

Fig. 1. ToF-SIMS peak area images for manually selected peaks.

Top row (left to right) m/z=63, 79, 181. Middle row (left to right) m/z=33, 59, 71. Bottom row (left to right) m/z=65, 91, 107. All images are 200 micron×200 micron. The scale bars shown are 20 microns.

Figures 2–4 show the first 3 principal components (PC) from PCA. As seen in the figures, similar features as were seen in the manual analysis can be seen in the PC score images. However with PCA all of the peaks that show the same trends are found without the need of user input. As would be expected, the peaks selected in the manual analysis are found to correspond to their respective areas within the scores and loadings plots. For example, the nuclei are highlighted in the PC1 scores plot (bright areas) and are seen to correspond with peaks such as m/z 63, 79, and 181 all of which were chosen in the manual analysis. In addition several other peaks are shown to have positive loadings on PC1. These peaks show similar spatial distribution as the m/z 63, 79, and 181 peaks, with a higher relative intensity from the nuclei of the cells (data not shown). Similar trends can be seen for PC2 and PC3 where PCA highlights peaks that show the same relative intensity patterns.

Fig. 2. PCA PC 1 scores from the mouse muscle data.

PCA scores (left) and loadings (right) from the mouse muscle data. PC1 is shown to separate the cell nuclei (positive scores–bright regions) from the intercellular spaces and cell bodies (negative scores–darker regions).

Fig. 3. PCA PC2 scores from the mouse muscle data.

PCA scores (left) and loadings (right) from the mouse muscle data. PC2 is shown to separate the cell bodies and some cell edges (positive scores–bright regions) from the cell nuclei and intercellular spaces (negative scores–darker regions).

Fig. 4. PCA PC3 scores from the mouse muscle data.

PCA scores (left) and loadings (right) from the mouse muscle data. PC3 is shown to separate areas of the cell bodies and some cell edges (positive scores–bright regions) from other areas of the cell bodies (negative scores–darker regions).

One can go through the scores and loadings plots and quickly summarize the trends in the data. However, it is noted that many of the same peaks have relatively high loadings across multiple PCs. Looking at the loadings, one can see that the trends are mostly self consistent. For example the m/z 63, 79, and 181 peaks are always seen to correspond with the nuclei of the cells and the m/z 42 peak is seen to correspond with the cell body and intercellular space. However, the trends in some of the peaks are not as clear. For example, the peak at m/z 91 appears to correspond with the intercellular areas in the PC1 and PC2 scores images, but for PC3 it shows a positive loading that would correspond with the bright lines in the image. Based on their location, these lines could be intercellular regions or could be on the edges of the muscle cells.

Some analysts find PCA to be somewhat confusing due to the fact that peaks can be present in multiple PCs, and that PCA produces both positive and negative scores and loadings. This and the desire to try to resolve the data into “pure” components has led some users to apply MCR processing to their ToF-SIMS data. It is typical to use the PCA scores as an initial guess for the number of “pure” components for MCR, which means that MCR and PCA provide the same information, but present it differently. This is important to note because if there is no clear indication of chemically different regions in PCA, it is unlikely that MCR will find a logical answer. It may come up with a set of components, but they may not be chemically relevant or real. In fact, it is typical that the trends seen in MCR will also be apparent from PCA, but that MCR will present them in a visually simpler format. Figures 5–7 show the 3 MCR components calculated when using the 3 PC model from PCA as the initial guess. As seen in the figure, MCR separates the nuclei, the cell bodies and the intercellular space in the 3 component images. It is noted that component 3 also shows some contrast with the cell bodies in the lower half of the image. This appears to be highlighting a chemical difference between the different muscle cells. Figure 8 shows an RGB overlay of the three MCR component images. Very little overlap is seen within the image as seen by the lack of colors other than pure red, green, and blue. It is noted that the differences in the signal from the cell bodies are not due to differences in ion yields due to instrumentation or charging, as normalizing the data to the total intensity of each pixel does not change the general trends seen in the data (data not shown).

Fig. 5. MCR component 1 data from the mouse muscle sample.

MCR component image (left) and component spectra (right) from the mouse muscle data. MCR component 1 is shown to highlight the cell nuclei. The peaks in the spectra are mainly attributable to phosphate fragments and phosphate salts.

Fig. 6. MCR component 2 data from the mouse muscle sample.

MCR component image (left) and component spectra (right) from the mouse muscle data. MCR component 2 is shown to highlight the cell bodies. Some variation in the relative intensity of the cell bodies is seen across the image.

Fig. 7. MCR component 3 data from the mouse muscle sample.

MCR component image (left) and component spectra (right) from the mouse muscle data. MCR component 3 is shown to highlight intercellular space and some of the cell bodies.

Fig. 8. RGB overlay of MCR component images.

Red=Component 1, Green=Component 2, Blue=Component 3.

Comparing the PCA and MCR results, one can see that MCR component 1 and the positive scores from PC1 clearly highlight the cell nuclei. The PC1 positive loadings and MCR component 1 spectra look almost identical. Similar trends can be seen where MCR component 2 is seen to correspond well with the PC3 negative scores and loadings and MCR component 3 is seen to correspond well with the PC2 negative scores and loadings. This is not unexpected since the MCR results are based off the initial guess of 3 components from the PCA scores, and because both methods are looking at the variance within the data set that is caused by the differences in the relative intensities of the same peaks.

It should be noted that it is not possible to say which analysis method is better than the other. Which method is used will depend on the goals of the analysis. For some systems, a simple manual analysis of certain components with known secondary ion peaks may be sufficient. The various MVA methods are simply tools that enable the user to analyze the full ToF-SIMS data matrix and minimize user bias. PCA is a good choice for general data exploration and sometimes may be all that is required to summarize important features and trends in a data set. It is also important to remember that each MVA method carries a set of assumptions and many are dependent on the data scaling or initial guesses used.64) For example, the results from MCR will depend significantly on the number of components assumed in the initial estimate. For a more in depth discussion of the various MVA methods and their application to ToF-SIMS data the reader is referred to the current literature.22,51–56)

LIMITATIONS AND NEEDS

One thing that can be seen from the data presented in this study is that all of the peaks highlighted in the PCA and MCR analysis were relatively low mass peaks (<300 m/z). This is due to the low ion yields for high mass peaks (>300 m/z). Although one can extract information about a given sample set from analyzing the relative intensity differences of the low mass peaks, to get more chemically specific information from biological samples higher intensities of higher mass peaks are required. This is particularly true for high spatial resolution imaging where count rates of the high mass peaks are significantly lower. Thus any improvements in secondary ion yields that results from advances in new instrumentation (ions sources, detectors, etc.) and sample preparation (e.g., matrix enhanced SIMS) would provide a direct benefit for ToF-SIMS analysis of biological materials.

Even when high mass peaks can be detected, the challenge in many cases is to determine the peak identity. This can be particularly challenging for high mass peaks where the number of possible matches to a given mass can reach the thousands. Some work has been done to identify peaks related to amino acids,65,66) DNA bases,67) and lipids,3) however there is still a lack of easily accessible databases of ToF-SIMS data that one can use to help identify peaks of interest. It is not clear how this type of database should be organized or started, but the presence of such a database would be useful to the community.

Although this paper concentrated on 2D ToF-SIMS image data, the increased use of 3D data sets presents the challenge of needing increased computing power. This is a more easily addressable challenge, since it is relatively easy to purchase a computer with sufficient computing power for a reasonable price. It should be noted that this does not mean one can use a standard off the shelf system to easily process large data sets. Modern computers will do fine for most 2D data sets, however when moving to 3D analysis the user must significantly increase the computing power of their system or develop more efficient algorithms for processing the 3D data sets. Though we do not provide specific computer system recommendations since those will depend on the specific needs of the user and the data sets being processed, we do recommend that the user install the maximum amount of RAM possible for their system and choose a system with the fastest processor available. Multi-core processors are highly recommended. It is also recommended that the user install a dedicated graphics card equipped with the maximum amount of on board RAM. Furthermore, to more efficiently process the increasingly large data sets generated by ToF-SIMS, it is recommended that programs be developed which take advantage of the parallel processing capabilities provided by modern graphics cards equipped with multi-core GPUs.

CONCLUSION

ToF-SIMS analysis of biological materials has been shown to be an effective method to extract useful chemical information from a wide variety of samples. ToF-SIMS can be used for a wide range of sample types, and the current instrumentation allows analysis of samples in both the frozen hydrated and dehydrated states. MVA methods have enabled the analysis of full data sets that eliminates user bias and allows access to information from all generated peaks. This has enabled the analysis of many types of samples and produced useful information. However, there are still limitations that need to be overcome. These include the need for increased ion yields of high mass, chemically specific peaks, the generation of and access to databases of spectra from biologically relevant materials that can be used for peak identification, and the need for more efficient programs that take advantage of multi-core, multi-threaded GPU graphics cards.

Acknowledgment

The authors gratefully acknowledge the funding and facilities provided by the National ESCA and Surface Analysis Center for Biomedical Problems (NESAC/BIO) through grant EB-002027 from the National Institutes of Health. The authors also thank Dr. Nick Whitehead for providing the mouse muscle samples used for the ToF-SIMS image processing examples in this study.

REFERENCES
 
© 2013 The Mass Spectrometry Society of Japan
feedback
Top