Chem-Bio Informatics Journal
Online ISSN : 1347-0442
Print ISSN : 1347-6297
ISSN-L : 1347-0442
original
Progressing adaptation of SARS-CoV-2 to humans
Tomokazu Konishi
Author information
JOURNAL FREE ACCESS FULL-TEXT HTML
Supplementary material

2022 Volume 22 Pages 1-12

Details
Abstract

The second and subsequent waves of the coronavirus disease (COVID-19) have caused problems worldwide. Here, using objective analysis, we present the changes that occurred in severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). The virus has mutated in three major directions, resulting in three groups to date. The basic viral genome was identified in April and shared across all continents. However, the virus continued to mutate independently in each country after the closure of borders. Some variants with greater infectivity replaced the earlier ones and caused second waves of the disease. Some of them slowly entered other countries and caused epidemics. Going forward, these viruses could also serve as sources of further mutations.

1. Introduction

The evolutionary trajectory of an evolved virus is difficult to analyse from nucleotide sequences because the data have complex multivariate structures with numerous dimensions [1]. Phylogenetic trees have long been used to present relationships in among sequences [2] and many studies on severe acute respiratory syndrome coronavirus (SARS-CoV-2) [3] have employed them [4–7], but they have two drawbacks. One is the decisive lack of falsifiability that ruins objectivity, and the other is a lack of generality that makes it difficult to integrate with other data sources [1, 8–9].

Here, an analysis was performed using a more objective method, i.e., principal component analysis (PCA) [10]. PCA runs a singular value decomposition on a sequence described as a Boolean vector and obtains the principal component (PC) for the sample and base in parallel [11]. PCA finds common directions among a dataset and represents them as independent axes. These axes have an order; PC1 is common to most data, and the lower the contribution to differences in the data, the lower the order. The sample or base is presented on the axis according to its characteristics. As shown in the current study, SARS-CoV-2 data were represented by two axes as a whole. The lower axes represented mutations found in smaller areas. The components were scaled and compared with those of the influenza virus [12] as needed. In addition, the results were integrated with the date and location of collection.

2. Materials and methods

2.1 PCA

Nucleotide sequences were obtained from GISAID [13] on October 1, 2020. Only the complete sequences that did not contain N were selected. Sequences were aligned using DECIPHER [14]. The sequences were converted to a Boolean vector and subjected to PCA [1]. Sample PCs and sequence PCs are scaled based on the length of the sequence and the number of samples, respectively; those are designated as sPC here [11]. All calculations were performed using R [15]. The ID, acknowledgements, and scaled PCs of samples and bases can be downloaded from Figshare [16]. The newest version of R code and the PCA axes are publicly available at GitHub [17]. Using these methods, any sequence data can be represented on the same axes presented here, and the axes can also be updated with newer sequences. The relationship between PC values and the number of mutated bases is briefly explained in S2 Fig. Number of confirmed cases were obtained from WHO [3].

The axes were identified and used as follows. Because PCA axes are sensitive to sample bias, they were determined using 6,092 samples randomly selected from each continent in proportion to their population; 989 from Africa, 3636 from Asia, 610 from Europe, 476 from North America, 347 from South America, and 34 from Oceania; the number of sequences used from Africa was the maximum available; hence, the sampling was not random. For clarity, Fig. 1 and S1 presents only 100 samples from each continent, whereas S2, S4, and S5 Fig. present approximately 1000 samples. The contributions of the axes are presented in S3 Fig.

2.2 Distances between samples from different peaks

In influenza, only one strain has emerged over the years, with repeated mutations [12]. Differences between peaks within the same strain were found as differences between the mean sequences from both peaks. In contrast, in coronavirus, the peak comprises several different variants, especially the first peak. Averaging will cancel the characteristics of these variants. To avoid this problem, the differences were estimated among those that have similar PC1 and PC2 and are in the same group as shown in Table 1, and the differences only appeared in the lower PCs.

2.3 The rate of missense mutations

In influenza, differences in both nucleotide and amino acid sequences were presumed as differences between the average peak sequences. For coronavirus, the difference between the mean sequence and each sample was estimated, and the rate was estimated for each sample. The distribution was slightly distorted (S6 Fig.), and hence, the representative value was found as the median.

3. Results and discussion

3.1 Overview of PCA.

PCs 1 and 2 formed a triangle comprising four groups that were temporarily termed as 0–3 (Fig. 1A and S1). All groups were found on all continents. In contrast to PC1 and PC2 showing the overall situation, the PC3 axis and the lower-order axes showed differences that relate to only a small number of the samples (Fig. 1B and S2), with smaller contributions to differences in the data (S3 Fig.); many of them were detected in limited countries after April (comparing S3 and S5 Fig.). Such mutations occurred since April when countries began to restrict the movement of people across national borders. The overview of Fig. 1A was completed at the end of March (Fig. 1C). Although group 3 once appeared worldwide, it has been contained to date (Fig. 1D).

These features were similar for both RNA and protein (S2 and S5 Fig.). The rate of missense mutations, which affects amino acid sequence, was 63% (S6 Fig.). This is a fairly high rate; rather, null mutations are expected to be 61% [18] even though some differences caused by ignoring damaging mutations have been identified here. In fact, in the case of H1N1 influenza virus from 2001 to 2003 in the United States, the rate was 22%. This suggested that SARS-CoV-2 was under selective pressure to alter its protein structure.

The viral groups causing the epidemics changed over time (Fig. 2A). The second and third waves in each country were caused by new variants of the virus, which showed different patterns at lower PCs (Fig. 2 and S7). In England, the second wave was caused by a variant of the Group 0, which has since spread to many countries in Europe. The third wave occurred due to another variant of Group 1, H69-V70 (Fig. 2A and 2C: see the arrow in 2C) [19–20]. This group 1 variant shown by the arrow would later be called as Alpha [3]. Another variant that included in the third peak was in Group 0, which was unique in PC198 (Fig. 2B). This variant would later be called as Delta [3].

The H69-V70 variant in the UK [19–20]. actually involves multiple mutations, such as an additional deletion at I69 (S1 Table). The deletion itself has occurred spontaneously in multiple groups (Fig 2D); the epidemic in the UK is caused by a group 1 variant, while those in Germany and Denmark are caused by group 0. There were other group 1 variants without the deletion in the UK, which would be the ancestral variants. The group 1 H69-V70 variant in the UK is growing at the same or faster rate than the pan-European variant by replacing it (Fig 2A). These results show that this variant could be more infectious than the pan-European variant.

A group 2 variant is quickly spreading in South Africa; this is Beta variant, B.1.351 [3] (S7 Fig., panels Q and R). As it had not appeared until October, the axes of PCA had not trained for the variant; hence, the figure was made by renewing the axes. This is a novel variant that contains many mutations concentrated in the spike protein, and the magnitude of mutation is deemed comparable to that of influenza H1N1 hemagglutinin between sequential peaks (Table 1). Group 2 variants have been negligible in South Africa (S7 Fig., panel Q), and no variant was considered ancestral. Perhaps the variant was from areas where samples are rarely sequenced. Any region in the world can produce a new highly infectious variant, with potential to cross national borders. Therefore, measures such as vaccines are urgently needed on a global scale.

In many countries, mutations were detected several weeks prior to the peaks (Fig. 2 and S7). This suggested that the mutations were the cause of the peak. In any case, the variants prevalent at those peaks should certainly be more infectious than they were before the mutation, so the variants were replaced. The higher infectivity may have weakened the previously effective protections from the virus. Altered residues between the peaks are shown in S1 Table; because PCA showed differences among samples and bases in parallel, PCs for bases may help identify those differences. It may be possible to predict the next wave by monitoring the mutations. However, only a limited number of countries continue sequencing the virus; many have reported only the first peak. However, the second and subsequent peaks are caused by different variants (Fig. 2 and S7).

3.2 Magnitude of mutations.

Influenza H1N1 is highly infectious in humans [21]. It mutates continuously but does not appear in the same area during two sequential years. It is prevalent every few years because it takes time to mutate enough to survive against herd immunity in the area. However, the magnitude of changes in the meantime varies (Table 1). Perhaps, the part of the sequence that needs to be modified to escape herd immunity varies from case to case. However, the changes in coronavirus before and after the peaks were much smaller (Table 1). The difference seem to be smaller than that needed to obtain another peak of seasonal influenza, although this does not guarantee safety from exceptional reinfections [22–23]. A possible exception is the one that presently causes an epidemic in South Africa.

A case of pdm09 influenza virus that likely corresponded to SARS-CoV-2 mutations was observed in Thailand from 2009 to 2010 (S8 Fig.) [24]. The 2009 pandemic was subdued in this tropical country, and three peaks were subsequently identified toward 2010. At the second peak, only the nucleotide sequence changed, and at the third peak, the hemagglutinin protein changed (Table 1) with a high rate of missense mutations (54%). However, these changes were not highly associated with herd immunity because most of the population was not yet infected. These changes might have been caused by the adaptation of pdm09—originally a swine strain—to humans.

The numbers show mutations per 1000 amino acid residues. Results of the H1N1 influenza virus hemagglutinin protein observed in the peak seasons, and the whole genome of coronavirus reported in the peaks are shown. (2008–2009) indicates the difference between strains before 2009 and the pdm09 influenza virus. Max (PC1, PC2) is the difference between the most diverged SARS-CoV-2 groups in Fig. 1A. The England (group 1: Alpha) and South Africa (group 2: Beta) are the new variants presently spreading rapidly. The England variant Alpha and the South African one Beta are the new viral variants spreading at a rapid rate currently.

The influenza virus genome continues to mutate and not only in the part that is directly associated with infection [12]. This property may be the reason it has repeatedly deceived herd immunity for decades, allowing influenza to remain as a pandemic.

In contrast, mutations in SARS-CoV-2 were not uniform; rather, some smaller open reading frames (ORFs) such as E, M, and ORF6 were preserved (Fig. 3). These could have adapted to humans, but they are also conserved in various coronaviruses [25]. If they remain conserved, they could be subjected to herd immunity, extending the period between future epidemics. Evidently, this does not mean that these ORFs do not change forever; the rate of change is just slower.

The strain responsible for the pandemic has mutated rapidly. As of March, the bases of PCs 1 and 2 were formed, and each continent harboured the same set of mutated variants (Fig. 1). Since then, human movement has been restricted, and each country accumulated its own variants. In this process, the force that caused the mutations was probably adaptation to humans (Table 1). It caused changes in both codon usage and protein structure (S3 and S6 Fig.), but the rate of missense mutations was high (S1 Table).

Some particularly infectious variants of the virus crossed borders and became the cause of the second wave in many countries (S7 Fig.). A Group 0 variant, which has caused the epidemic in Spain in July, was detected in other EU countries about a month later, leaving the same fingerprints in the PCs (panels B, D, F, I, and L). In countries other than France, this strain formed second waves (Fig. 2 and S7). In France, a Group 2 variant was prevalent at that time (S7 Fig., panel H); currently, this pan-European variant is spreading at a rapid rate (panel I). The pan-European variant has also reached the USA, Japan, and Australia; however, the sequences of the epidemic variant are unknown in Japan and in the USA, though these countries are amidst waves of infection. Australia seems to have succeeded in controlling this variant (panel O).

The unique variants in each country would serve as the source of newer variants, which are more infectious than the previous variants; therefore, they replaced the mode (Fig. 2 and S7). If countries had not closed their national borders, many of their variants could have been replaced by other more powerful variants, as in the case of group 3. However, the unique variants survived in a narrower area; the variant that is causing the epidemic in South Africa would be an example. International events, such as the Olympic Games, which attract people in large numbers, will serve as an exchange market for these viruses. To make matters worse, coronavirus could acquire shift-type mutations, resulting in a mixture of genomes [25]. When an exported strain is mixed with a domestic strain, a new strain is created. If each of the unique mutations works independently (otherwise, replacement of the older variants with these variants would not be beneficial), the accumulation of the mutations can result in the production of more powerful viruses with cherry-picked the surviving unique mutations.

3.3 On the application of PCA to sequence data

PCA showed several notable characteristics in the evaluation of viral changes. First, the dimensionality of the original data was very large. The actual number of dimensions is smaller than the length of the sequence or number of samples. Coronaviruses contain approximately 300,000 base pairs; therefore, if the data consist of 6,000 samples, then the dimensionality is 6,000, producing 6,000 axes for the PC. Second, these data reflect a record of changes from the original data. For example, influenza virus H1N1 was a single variant that changed continuously and rarely formed branches [12]. This indicates that a highly infectious virus can easily cross national borders and, from the most potent variant at the time, lead to the emergence of a new variant that can break through the immune system. In contrast, coronaviruses showed several branches (Fig. 1), possibly because of the emergence of more infectious variants in each region while borders were closed. According to the data, changes are accumulating in coronaviruses.

The upper level of the axis records acclimation to humans, whereas the lower level records more recent mutations. PCA involves setting up new axes in the data that mathematically indicate the rotation of the original matrix [10]. The moment of rotation is determined by the product of the number of bases and number of samples that have changed. Altered sequences may revert to their original state; however, mutations that lead to stronger infectivity will be conserved. The positive orientation of PC1 and PC2 in Fig. 1 indicates the viral acclimation to humans, which was not reversed. When movements occur in a common direction, PCA collects these data and exhibits them on a higher axis. Based on this property, PCA is often used for dimensional compression. However, the lower axes may also be important and cannot be ignored.

Actually, PCA is NOT a specialized method for dimensional compression, as it does not summarize data. Rather, it is a method for efficiently observing data with multiple dimensions. PCA only rotates the data, without losing any information, showing the structure of data faithfully. In both compress and summarize, there is subjectivity involved in selecting what to choose and what to discard. Being free from those gives PCA the objectivity required in science.

The data presented here recorded the process of COVID-19, originally a virus of other animals, acclimatized to humans. There were steps of acclimation that were common to many variants recorded in PC1 and PC2, and there were also other steps that progressed independently in each country recorded on the lower PCs. These were the hidden structures of the data that has been disclosed by the PCA. As is evident from this clear result we have been able to achieve, the fact that the contribution of PC1 and PC2 was not huge is not a limitation or defect of the methodology. It is the structure of the data. In this article, I introduced the early acclimation, so I presented it in Fig.1 by scatter plot of PC1 and PC2. To see more detailed information, it would be necessary to read the table which is available through Figshare [16].

Newer mutations are recorded on lower levels of PC axes, as a smaller number of samples is available, and thus, the moment is small. For example, the delta variant belonged to group 0 in PC1 and 2 (Fig. 1A), and its unique mutations in PC198 were recorded (Fig. 2). In September 2020, the number of patients with the delta variant was still small; however, as there are more than 6,000 axes, PC198 is not in a very low position. Additionally, these contributions were very small (Fig. S3); however, because the average of the contributions is approximately 1/6,000, values as large as sub-percentages are not small. As the delta variant spread worldwide, it became prominent in PC1 according to the most recent data from November 2021 [26]. Therefore, a smaller contribution does not suggest that mutations are insignificant; rather, it is important to determine when the mutation was recorded and whether sufficient time has passed for its spread.

Problems in sequencing appear characteristically in PCA. Unread sections and undetermined bases, indicated by N, do not affect PCA because they are replaced by the average sequence [1]. In addition, if some sequencing errors occur, they are recorded in low axes, as their moment is quite small. However, an error in alignment can strongly affect the results. For example, PCA of the aligned sequences provided by GISAID 13 as of November 2021 resulted in a large fragmentation in PC1, leading to the generation of two distinct groups of data (Fig. S9). This large gap disappeared when the alignment was repeated, indicating an alignment error.

4. Conclusion

The epidemics showed much variation among the countries (Fig. 1, 2, S1, and S7). A Group 2 variant that caused the second wave in the USA (S7 Fig., panels M and N) migrated to Australia a few months later (panel P). However, in contrast to the USA where the epidemic did not settle down, Australia quickly suppressed the second wave. Furthermore, New Zealand did not even let an epidemic occur. It should be noted that Sweden, which did not enforce a lockdown, could not control the pan-European variant (panels K and L). Before this variant, the Swedish measures had worked well and the infection was avoided simply by social distancing. However, the pan-European variant was so infectious that it negated those defenses. The responsive measures of various countries should be compared in detail for future pandemic.

Figure 1. Principal components of SARS-CoV-2 amino acid sequences reported until December 1, 2020

One hundred samples are randomly selected from a continent and presented. A. PC1 and PC2. Numbers are given for a temporal explanation of the groups. B. PC3 and PC4. The earliest variants are indicated by the letter W. The numbers of samples that are separated by the axes are reduced in the lower-order axes (S3 and S4 Fig.). C and D. PC1 and PC2 at the end of March and from August 2020. A maximum of 100 samples is selected from a continent and presented. The components are scaled by the length of sequences; squares of the values are correlated with the expectation of differences per base (S2 Fig.).

Figure 2. Mutations found on PCA and the number of patients in England. A

Changes in rates in each group. The dotted grey line shows the confirmed cases. The second and third winds were caused mainly by group 0 and 1 variant, respectively. B. PC fingerprints of group 0. Prior to the second wind, the pan-European variant appeared. PC198 (black dots) higher than the black arrow indicates the H69-V70 group 0 variant. C. PC fingerprints of group 1. Prior to the third wind, H69-V70 group 1 variant (black arrow) appeared. D. H69-V70 variants found in England (blue).

Figure 3. Positions of mutations. A. Mean distances to the sample mean (protein)

The square of the distance reflects the expected mutation occurrence at the residue. The brown line indicates the moving average of 100 residues (×10 scale). B. Magnified view of the conserved region.

Table 1. The rate of mutated amino acid residues during the indicated years and location

The numbers show mutations per 1000 amino acid residues. Results of the H1N1 influenza virus hemagglutinin protein observed in the peak seasons, and the whole genome of coronavirus reported in the peaks are shown. (2008–2009) indicates the difference between strains before 2009 and the pdm09 influenza virus. Max (PC1, PC2) is the difference between the most diverged SARS-CoV-2 groups in Fig. 1A. The England (group 1) and South Africa (group 2) are the new variants presently spreading rapidly. The England variant (Group 1) and the South African one (Group 2) are the new viral variants spreading at a rapid rate currently.

Acknowledgements

We would like to thank Editage (www.editage.com) for English language editing.

References
 
International (CC BY 4.0) : The images, videos or other third party material in this article are also included in the article’s Creative Commons license.To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/

この記事はクリエイティブ・コモンズ [表示 4.0 国際]ライセンスの下に提供されています。
https://creativecommons.org/licenses/by/4.0/deed.ja
feedback
Top