2022 年 22 巻 p. 26-37
Structural biology comprises “Structured biology” based on folded structure and “Unstructured biology” based on unfolded structure. Principal of “Unstructured biology” is intrinsically disordered protein (IDP), in which the polypeptide chain is highly disordered under physiological conditions. The length of the disordered polypeptide may extend to several hundred residues or more. Further, in protein comprising multiple domains, long disordered polypeptides exist between domains, even if each domain has a regularly folded structure. Such polypeptide region is called intrinsically disordered region (IDR). Several IDPs and IDRs exist in living cells and play important biological roles in intracellular networks. Here, the biological significance of IDP and IDR from various viewpoints are described. Overall, this review aims to encourage further studies on “Unstructured biology” based on unfolded/disordered protein structure.
Intrinsically disordered protein (IDP) and intrinsically disordered region (IDR) are characterized by a lack of stable folded structure along their entire lengths and in localized regions between domains, respectively, when they exist as isolated polypeptides under physiological conditions. Until the late 1990s, IDP and IDR have been compared to "noodles" and have been regarded as "troublesome things" that hinder protein research methods such as X-ray crystallographic analysis. However, Wright et al. showed that the highly disordered polypeptides in their physiological conditions play important roles in protein-protein interactions within intracellular network [1]. Therefore, structural and functional studies on IDP and IDR as novel research targets for protein science have been initiated. This paper reviews the structural and functional properties of IDP and IDR, together with their fundamental databases, to encourage further studies on “Unstructured biology” based on unfolded/disordered protein structure.
In 1997, Wright et al. demonstrated that a disordered polypeptide chain shows a nuclear magnetic resonance (NMR) spectrum characteristic of a disordered/unfolded state, but when a protein with a regularly folded structure is added to it, the polypeptide chain presents a spectrum characteristic of a regularly folded structure [1]. This is called coupled folding and binding (Fig. 1) and attract attention as a novel molecular recognition mechanism that differs from conventional models such as lock-and-key, induced-fit, population-shift (pre-existing). This epoch-making discovery in protein science has introduced a paradigm shift in the structural and functional study of proteins. The following sections outlines the properties and functions of IDP and IDR that have been clarified so far.

Figure 1. Coupled folding and binding shown by Wright et al. as a novel molecular recognition mechanism [1]
Highly disordered phosphorylated kinase-inducible domain (pKID) of CREB forms two α-helices in the area indicated by red string upon interaction with the regularly folded KIX domain of CREB.
Coupled folding and binding mechanisms indicate that when IDP and IDR recognize a target molecule, the disordered polypeptides fluctuating between small local minimums of conformational energy tend to have a deep global minimum to form a different folded structure depending on the surface structure of the target molecule in different time and space (Fig. 2). This means that IDP and IDR have promiscuous characters that allow them to interact with multiple targets through their structural flexibility. This is why they are called a hub protein and a hub region for intracellular networks, respectively.
For example, p53, a well-known cancer-suppressing protein, is a transcription factor with a total length of 393 residues, 50% of which is IDR and the remaining is a DNA-binding domain with a regularly folded structure. Fig. 3 shows that p53 interacts with 14 target molecules including DNA, as clarified at the atomic resolution. Interestingly, among these 14 molecules, only two molecules, other than DNA, interact with the DNA-binding domain of p53, whereas the rest all interact with the IDRs of p53. Although p53 is a hub protein involved in the transcription of more than150 genes, the hub nature in its transcriptional signaling network is attributing to the high flexibility of the IDR within the p53 molecule. Notably, in Fig. 3, the C-terminal small region from residue numbers 370 to 385 interact with five target molecules, of which four molecules interact with p53 in the same region from residue numbers 375 to 385. This is a typical experimental evidence that highly flexible IDP and IDR play important roles as hub proteins in intracellular networks; many other proteins with long IDRs are also likely to be included in the molecules in intracellular networks.

Figure 2. Change in the conformational energy of IDP/IDR by coupled folding and binding

Figure 3. Disorder prediction of p53 (central region) and p53 structures in complex with 14 different partners (peripheral region)
Disorder prediction is designated by the PONDR score (up: disorder, down: order) along with the amino acid sequence of p53. In partner-free state, p53 has a regularly folded structure only in the DNA-binding domain from residue numbers 102 to 292. This figure was taken from ref [2] with minor modification.
IDP and IDR are characteristically abundant in eukaryotes. In fact, more than 60% of molecules in human are IDPs or proteins with long IDRs (Fig. 4). Further, they are more abundant in the nucleus than in the cytoplasm. Of the nuclear proteins involved in DNA replication, transcription, repair, and recombination, more than 50% are IDPs or proteins with long IDRs within the molecule. Furthermore, proteomic studies have shown that approximately 80% of human cancer-associated proteins have long IDRs or can be classified as IDPs [3].
These data indicate that the intrinsic properties of IDP and IDR may have been acquired during long molecular evolution through the highly complex protein-protein interaction required for the hub nature in intracellular networks, because of which IDP and IDR are peculiar to eukaryotes.

Figure 4. Populations of IDP and IDR (Fukuchi et al. unpublished data)
Many proteins change their functions significantly through post-translational modification; these also cause normal proteins to become abnormal, thus causing disease. In p53, the IDR has 27 sites that undergo post-translational modifications, including phosphorylation, methylation, and so on, whereas only four sites in the DNA-binding domains have an ordered structure (Fig. 5). Furthermore, as shown in Fig. 5, the four modifications in the DNA-binding domain are all phosphorylation modifications, whereas seven of the 27 sites in the IDR undergo two to four types of modifications at the same site. This is only experimental evidence indicating that post-translational modifications occur exclusively in the IDR. Although p53 was taken as an example here, this is a commonly observed phenomenon. As many diseases are thought to be caused by abnormal post-translational modifications that change protein function. The IDR is an important region for controlling protein function, and is likely to be a target for novel drug development (described later in this paper).

Figure 5. Domain structure of p53, a tumor suppressor protein
TAD: transcription activation domain, PRD: proline rich domain, DBD: DNA-binding domain,TD: tetramerization domain,BD: basic domain. P with yellow background: phosphorylation at Ser, P with light-blue background: phosphorylation at Thr. Ac: acetylation, S: SUMO: sumoylation, NEDD: neddylation, M: methylation, methylation, Ub: ubiquitination. This figure was taken from ref [4] with miner modification.
In eukaryotes, alternative splicing produces many kinds of proteins from one gene. This is a phenomenon in which introns are cut off in the transcription process and exons then bind to each other. This phenomenon occurs in different positions resulting in the formation of multiple mRNAs (proteins) from one gene. This is why the process is called alternative splicing; IDR is known be deeply involved in this phenomenon.
This phenomenon is not problematic at the mRNA level, but can cause major problems upon their translation into proteins, because the stability of protein structure changes significantly depending on the splice site. However, if the site exists in the IDR, the stability of the protein structure is likely to be unaffected significantly because of the flexibility of the IDR. Therefore, alternative splicing of IDR-encoding mRNA is not a major problem and can be acceptable even if mRNA is translated into protein. This indicates that IDR is the driving force underlying the production of many proteins from a single gene and the resultant eukaryotic biodiversity.
2.5 IDP and IDR are important targets as novel drug developmentIn 2.1, IDP and protein with a long IDR have been described to function as hub proteins in intracellular networks. In 2.3, the functional change of proteins by post-translational modifications in IDRs have been shown to often cause serious diseases. Furthermore, as mentioned in 2.2, bioinformatics analyses have shown that about 80% of human cancer-associated proteins can be classified as IDPs or have long IDRs in the molecule [3]. Considering these results, future drug discovery needs to include development of protein-protein interaction (PPI) inhibitors that inhibit IDR-mediated PPIs, together with conventional inhibitors based on lock-and-key, induced-fit, population-shift, and other mechanisms. In fact, when developing a PPI inhibitor, if both the interacting regions (interfaces) form folded structures together, the interaction at the interface is discontinuous and widespread, which often hinders inhibitor development. However, if either of the two regions is an IDR, PPI inhibitor development will be facilitated because the interaction at the interface is continuous and its range is limited. In this consideration, IDP and IDR can be novel drug targets.
IDP and IDR are also known to play key roles in inducing liquid-liquid phase separation to form protein aggregates called as membrane-less organelles in living cells. The membrane-less organelles formed by the self-aggregation of α-synuclein and β-amyloid are closely associated with neurodegenerative diseases such as Parkinson’s disease and Alzheimer’s disease, respectively. Interestingly, in Parkinson’s disease, pathogenic amyloid fibroses are thought to be formed through membrane-less organelle [5]. Further, in amyotrophic lateral sclerosis (ALS), many mutations are found in the IDRs of T-cell intracellular antigen 1 (TIA-1) [6], a RNA-binding protein, and a slight difference in TIA-1 self-aggregation is likely to be the trigger for ALS. These data indicate that high resolution structural data of self-aggregates formed through IDP and IDR are of great importance, and can facilitate the development of new drugs to treat these diseases.
2.6 Structural basis for molecular recognition by coupled folding and bindingSo far, molecular recognition by coupled folding and binding has been investigated by molecular dynamics (MD) simulation at an atomic resolution. These investigations have been led by research groups centered in the USA, and have proposed the two different models involving induced fit and population-shift (pre-existing). However, the simulations were biased to form the correct interactions for complex formation; thus there are concerns that the effects of bias will not be removed from the simulation results. Further, amino acids were replaced with spheres to express protein molecules (coarse graining) in the MD simulation, and the calculations were performed using a method that does not explicitly take hydrogen water into consideration
On the contrary, Higo et al. performed all-atom multicanonical MD (McMD) simulation to elucidate the structural basis for the coupled folding and binding mechanism [7]. They simulated the α-helical formation of IDR in nerve-specific transcriptional repressor NRSF (NRSF-IDR) upon interaction with the corepressor mSin3 as its target protein using the structure of mSin3, in complex with the NRSF-IDR as determined by NMR (Fig. 6(a)).
In the simulation, NRSF-IDR was initially placed at a sufficient distance from mSin3 to avoid the effects of bias and to form the correct interaction for complex formation. Further, a novel algorithm was developed to efficiently search the structures, and was applied to the simulation. This indicated that NRSF-IDR highly fluctuates among diverse structures (population-shift like fluctuation) and almost all these structures bind to mSin3 (Fig. 6 (b)). However, the resultant structures of the complexes between NRSF-IDR and mSin3 are different from the structure determined using NMR (Fig. 6(a)). As shown in Fig. 6(b), even if a complex structure that differs from that determined by NMR is formed, NRSF-IDR within the complex causes induced-fit like structural changes and eventually forms a complex structure between NRSF-IDR and mSin3 as determined by NMR.
Therefore, molecular recognition by coupled folding and binding between NSRF-IDR and mSin3 has been elucidated to involve both induced-fit and population-shift. This indicates that the simulation method developed by Higo et al. is useful for elucidating the molecular recognition mechanism of IDPs or proteins with a long IDR.

Figure 6. Interaction of the IDR of NRSF (NRSF-IDR) with mSin3 as determined by NMR (a) and the coupled folding and binding mechanism elucidated using all-atom McMD simulation (b)
Considering that PDB contributes significantly to the “structured biology” based on folded structures globally, construction of fundamental databases for “unstructured biology” including IDPs and IDRs is of great importance. Thus, the database for intrinsically disordered proteins (IDEAL: Intrinsically Disordered proteins with Extensive Annotations and Literatures), developed by Ota and Fukuchi in Japan [8], is particularly noteworthy. IDEAL provides a collection of knowledge on experimentally verified IDPs and IDRs. IDEAL also contains manually curated annotations on IDRs in locations, structures, and functional sites such as protein binding regions and posttranslational modification sites along with references and structural domain assignments.
In coupled folding and binding as a novel molecular recognition mechanism, IDEAL explicitly annotates IDRs as protean segments (ProS) when both unfolded and folded information is available experimentally in these regions. As mentioned above, IDRs as ProS regions are sensitive to post-translational modifications, thus altering the mode of protein-protein interactions, which affects cellular functions. Therefore, to analyze the functional characteristics of IDRs, it is necessary to register ProS in the database and to explicitly indicate such sites as post-translational modification and so on. In IDEAL, ProS and post-translational modification sites are annotated based on literature searches (Fig. 7a). Currently, version 05/Oct/2021 has been released, in which the number of entry is 1110 and these can be converted to XML for processing on a personal computer.
As proteins with long IDRs play an important role as hub proteins for various intracellular networks, it is important to comprehensively collect the available information on various intracellular networks. IDEAL organizes this information, and can visualize them as a network (Fig. 7b). For post-translational modifications, IDEAL incorporates a classification index that shows the relationship among the intracellular networks, and annotates the structural and functional changes caused by post-translational modifications in terms of interaction switching based on the structural and functional diversity of IDPs/IDRs. Currently, IDEAL collects and integrates more data on intermolecular interaction for intracellular networks to clarify the statistical properties of IDPs/IDRs as hub proteins/regions with reference to the network data obtained from public large-scale interaction databases.

Figure 7. IDEAL, a fundamental database for IDP and IDR [8]
(a) Ordered (folded), disordered (unfolded), and ProS regions of p53, a cancer-suppressing protein and (b) the intracellular network of p53.
IDP and IDR have been attracting much attention to open a new era in the field of structural biology. According to Uversky and Dunker, the protein world in which polypeptide chains are regularly folded to function is called the Foldome, whereas the world in which IDP and IDR are involved is called the Unfoldome, and a paradigm shift from the Foldome to Unfoldome is underway [2]. In the Foldome, X-ray crystallographic analysis, NMR, and more recently, cryogenic electron microscopy (CryoEM) play leading roles in protein structure analysis, whereas in the Unfoldome, NMR leads the way. However, there remains an overwhelming lack of technical methods in the Unfoldsome compared with those in the Foldsome when elucidating the highly flexible structures of IDPs and proteins with long IDRs at an atomic resolution.
Here, the author introduced an example for analysis of coupled folding and binding at the atomic level using a synergy effect between NMR and molecular dynamics (MD) simulation. However, further developments in novel methods for structural analyses other than NMR and MD simulation are necessary in the Unfoldsome. From this viewpoint, small-angle X-ray and neutron scattering (SAXS & SANS) in synergy with MD simulation are likely to be effective in elucidating the dynamic behavior of IDPs and proteins with long IDRs.
In particular, SANS analysis used in combination with MD simulation is notably effective. It is currently difficult to analyze the dynamic nature of IDPs or proteins with long IDRs mainly because of the limitations of computer resources in MD calculation. Although the details are omitted here, SANS has an isotopic effect similar to that of NMR. If all the hydrogen atoms (H) in the IDRs are replaced with deuterium atoms (D), SANS from folded domains is no longer observed in buffer solution containing 40% D2O/H2O. Eventually, only SANS from deuterated IDR, which is much smaller than that in the folded domains, is observed, thus resolving the computer resource limitations in MD calculations. The J-PARC accelerator in Japan, which produces neutrons with the highest intensity globally, has reached full power operation this year. Therefore, the environment for using SANS method is in place. In the near future, the author’s group is planning a dynamical structure analysis of multi-domain full-length proteins with IDRs using J-PARC neutrons. Overall, this review encourages further studies on “Unstructured biology” based on unfolded/disordered protein structures.