2022 年 22 巻 p. 38-45
Medium-sized molecules have attracted significant attention as new chemical modalities. In this study, we compared the performances of three methods of 3D structure generation for medium-sized molecules using free and commercial software. The benchmark dataset consisted of 2131 protein-binding ligands with molecular weights greater than 600, which were selected from the Protein Data Bank (PDB). When selecting the smallest root mean square deviation between the generated 3D conformers and the PDB ligand structures, 43% of the conformations determined with the software CORINA were within 1 Å, followed by 10% from OMEGA and 5% from RDKit. According to our results, comparing the polar solvent-accessible surface area (PSA) and normalized principal moment of inertia ratio (NPR) among the three methods, 83% of the conformers generated with CORINA were within 20 in PSA, and 53% of the conformers from CORINA were within 0.05 in the NPR1 and NPR2 spaces. Thus, we concluded that CORINA has the highest performance in terms of efficient conformer generation. We also examined 3D descriptor calculation using Mordred, which is a free descriptor computation tool. The results showed that OMEGA-generated conformers exhibited the highest success rate, indicating that OMEGA is a suitable conformer generation tool for various 3D descriptors. Our results could contribute to the selection of conformation generators for the rapid construction of various predictive models for medium-sized molecules and can be shared with the research community for further validation.
Accurate and rapid prediction of protein-binding ligand conformations is one of the key challenges in in silico drug discovery. Although there is a vast number of chemical structures that can be synthesized, experimental data on the three-dimensional (3D) structure of ligands are limited; therefore, computational approaches for generating 3D ligand conformations are essential to fill this gap [1]. Artificial intelligence (AI) technology has been rapidly developing and recently applied in various drug discovery studies, such as pharmacology and pharmacokinetics/toxicity research. In general, AI models require large amounts of training data. Therefore, it is necessary to generate the 3D conformations of compounds and perform descriptor calculations to rapidly prepare a large amount of input data.
Various methods have been proposed to generate ligand conformations [2]. Furthermore, previous studies have compared and validated 3D conformation generation tools for small-sized molecules [3,4]. Medium-sized molecules, which are attracting attention as new modalities, are promising candidates for protein-protein interaction (PPI) inhibitors [5]. Therefore, as with small-sized molecules, it is worthwhile to validate the practicality of conformer generation using medium-sized molecules [6]. It has been reported that sphere-like conformations are preferred for PPI inhibition, which are evaluated based on the normalized principal moment of inertia ratio (NPR), a typical 3D descriptor [7]. These 3D descriptors are strongly influenced by the prediction accuracy of the input 3D structure. Therefore, validation of the accuracy of conformer generation is crucial for the design of PPI inhibitors.
In this study, three software packages, namely, RDKit [8], OMEGA [9], and CORINA [10], were selected based on their practical availability in current drug discovery research for the validation of conformer generation. OMEGA and CORINA are commercial tools that are widely used by pharmaceutical companies. RDKit is a freely available software that is widely used in industry and academia and has been reported to be reasonably accurate in predicting small-sized molecules. This study reports the results of the 3D structure predictions and 3D descriptor calculations. For machine learning, one conformer must be selected for each compound. Assuming this situation, we adopted a conformer selected based on the minimum energy for the calculation of 3D descriptors. This study also reports valuable calculation data, which can be further verified by other researchers.
Data collection and analysis were performed on a personal computer running on a Microsoft Windows 10 Professional 64-bit Operating System using a 3.60-GHz Intel® Xeon® W-2133 processor with 64.00-GB random access memory.
The 3D ligand structure data of the Protein Data Bank (PDB) [11] were downloaded from LigandExpo (downloaded on September 6, 2019). The software packages used to generate the conformer were CORINA (ver. 4.4.0), OMEGA (ver. 4.1.1.1), and the EmbedmultipleConfs function of RDKit (ver. 2020.09.1.0, Python ver. 3.6.13). All software packages were set to generate up to 250 conformers. To examine the molecular property distribution and descriptor error of the generated 3D structure, Pipeline Pilot 2018 [12] was used to calculate descriptors. The polar solvent-accessible surface area (PSA) was calculated using Molecular_3D_PolarSASA in Pipeline Pilot 2018. Pipeline Pilot 2018 was used to calculate the energy of the generated conformers, 3D alignment between the PDB structure and generated conformers, and root mean square deviation (RMSD) calculation. The “Align Molecules using Substructure” component of Pipeline Pilot 2018 was used for 3D alignment and RMSD calculation. This component performs the same atom-mapping procedure as the "2D substructure search" and then aligns (superposes) on the matched atoms between the two structures without any conformational change. Since the whole structure of PDB ligand was used as a query, the whole structure of generated conformer was aligned. When using this component, the search options (i.e., parameters) were used with default settings. In addition, Mordred [13] was used to calculate 213 3D descriptors.
The PDB ligand structures of 2131 compounds with molecular weights (MW) ranging from 600 to 1200 were downloaded from LigandExpo, excluding the macrocycles. Of these compounds, only 11 and 10 compounds were approved drugs and natural products (as defined in the ChEMBL database [14]), respectively, and 18 compounds were recorded in the TIMABL database. The minimum MW is 600.37 and the maximum is 1198.2 (see Fig. 1). The histograms of other chemical properties are shown in Fig. S1.

Figure 1.Histogram of the MWs for the PDB ligand data used in this study
The vertical axis shows frequency and the horizontal axis shows MWs (range is between 600 and 1200).
The conformers of 2131 protein-binding ligands from the PDB were generated using RDKit, OMEGA, and CORINA. The number of compounds completed by RDKit, OMEGA, and CORINA is 2041 (95%), 1996 (94%), and 2090 (98%), respectively (see Table 1). The success rates of the calculations are almost identical; however, CORINA shows a slightly higher success rate than the other two methods. The average number of generated conformers for each compound is 241.2, 237.1, and 42.3 for RDKit, OMEGA, and CORINA, respectively. Note that the maximum conformer option was set to 250 in all methods; therefore, RDKit and OMEGA are faithful to the computational conditions set (see Table 1). Because CORINA is a rule-based system, conformers may be limited by rules.
Table 1. Computational result from each method
| Method | Compounds computationally succeeded | Mean number of generated conformers |
| RDKit | 2041 | 241.2 |
| CORINA | 2090 | 42.3 |
| OMEGA | 1996 | 237.1 |
When selecting the smallest RMSD between the generated conformer and the ligand structure in the PDB for each compound, 42% (875/2090) of the conformers from CORINA are within 1 Å, followed by the 11% (215/1996) from OMEGA, and 4% (84/2041) from RDKit. Further note that 70% (1461/2090) of the conformers from CORINA, 51% (1025/1996) from OMEGA, and 59% (1196/2041) from RDKit are within 2 Å. Thus, CORINA can generate the highest number of conformers that are close to the ligand structures in the PDB (Fig. 2).
When selecting the RMSD value between the generated conformer with the lowest energy and the ligand structure in the PDB, 28% (580/2090) of the conformers from CORINA are within 1 Å, followed by the 1% (19/1996) from OMEGA, and 0.3% (6/2041) from RDKit. Furthermore, 58% (1219/2090) of the conformers from CORINA, 8% (156/1996) from OMEGA, and 2% (48/2041) from RDKit are within 2 Å (Fig. 2). Therefore, the conformers generated by CORINA are close to the ligand structures in the PDB even in the case of the minimum energy selection. It has been shown that OMEGA and RDKit generate more conformers than CORINA, which may increase the probability of conformer generation with a low energy and high RMSD.

Figure 2. Histograms of the RMSD for each method
The RMSD distribution of the minimum RMSD conformer for each compound (top) (A, B, and C), and minimum energy conformer (bottom) (D, E, and F) are shown.
Computational descriptors for machine learning inputs use conformations generated with the lowest energy; therefore, the difference in the 3D descriptors of the generated conformer and actual ligand structure is important. Thus, we investigated the difference in the 3D descriptors of the generated conformers with the lowest energy and the ligand structure from the PDB, as well as calculated a descriptor important for PPI. In the PSA calculation results, 83% (1735/2090) of the conformers from CORINA are within 20 of the PSA, followed by the 47% (960/2041) from RDKit, and 40% (792/1996) from OMEGA (Fig. 3A). When dividing data into two groups by a PSA value of 500 in the PDB ligand structure, the RMS values between the PSA values of the generated conformers and PDB ligand structures are 38.9 (PSA < 500) and 39.3 (PSA ≥ 500) in CORINA, 48.3 (PSA < 500) and 103.7 (PSA ≥ 500) in OMEGA, and 48.1 (PSA < 500) and 83.0 (PSA ≥ 500) in RDKit (Figs. 3C-E). In CORINA, the difference in PSA is small regardless of the range (Fig. 3C) when compared with the other methods (Figs. 3D and 3E).

Figure 3. Descriptor difference between the PDB ligand structures and generated conformers
The histogram for the different PSA values between the PDB ligand structures and generated conformers with minimum energy is shown (A). The histogram for the different NPR space values between the PDB ligand structures and generated conformers with minimum energy is shown (B). The data from CORINA, OMEGA, and RDKit are colored in orange, green, and light blue, respectively. The PSA values calculated with the PDB ligand structures and generated conformers by CORINA (C), RDKit (D), and OMEGA (E) are plotted.
The Euclidean distance in 2D space of NPR1 and NPR2 was used to examine the calculation accuracy of the NPR descriptors. We calculated the difference in the Euclidean distance in the NPR space between the generated conformers and the PDB ligand structures. The percentage of compounds within a distance of 0.05 in the NPR space is 53% (1116/2090) for CORINA, 11% (226/1996) for OMEGA, and 15% (303/2041) for RDKit. Therefore, CORINA shows the best performance in reproducing NPR descriptors.
For example, Fig. 4 shows snapshots of the 1,3-di(N-propyloxy-a-mannopyranosyl)-carbomyl 5-methyazido-benzene structure from the PDB (HET code: EJT) (Fig. 4A) and conformers generated by CORINA (Fig. 4B), RDKit (Fig. 4C), and OMEGA (Fig. 4D). The X-ray structure and conformation generated by CORINA are similar based on visual inspection, with a difference in the NPR space of only 0.009. On the other hand, the conformations generated by OMEGA and RDKit are clearly different, with NPR space differences of 0.38 and 0.60, respectively.

Figure 4. 3D structure of the PDB ligand structure and three generated conformers
The conformation of the PDB ligand structure (A) as well as conformers generated by CORINA (B), RDKit (C), and OMEGA (D).
As for the features required for the machine learning input, it is desirable to input many features that can potentially improve the accuracy of the model. 3D descriptors can explain interactions with other molecules and are efficient in improving prediction accuracy. Therefore, it is desirable to have many descriptors to complete the calculation of all 3D structures. Mordred is widely used as a free descriptor calculation tool for features entered into machine learning.
In the following analysis, the calculation of 213 3D descriptors using Mordred was performed using the conformers generated by each method. A descriptor that fails to calculate even for one compound cannot be used as an input into machine learning; therefore, we used a descriptor that was calculated for all compounds as a calculation success descriptor.
Of the 213 3D-descriptors, 122 are successfully calculated using CORINA, 11 using RDKit, and 185 using OMEGA. Thus, OMEGA exhibits the highest performance in generating a 3D structure suitable for 3D descriptor calculations in Mordred.
CORINA was found to be the best method for generating 3D conformers close to the PDB ligand structures and showed the smallest error in the 3D descriptor calculation for medium-sized molecules. Because of the small number of conformers generated, CORINA has a high probability of selecting a structure close to the ligand structure in the PDB when selecting one of all the conformers generated, which is also suitable for machine learning using 3D descriptors as input. On the other hand, OMEGA has an advantage in the number of successful 3D descriptor calculations on Mordred and the small restrictions in terms of the 3D descriptors that can be used for machine learning.
In conclusion, it is important to select a conformer generation method depending on the type of descriptors as input into machine learning, as well as the calculation method of the descriptors. Our study will be useful as input data for drug discovery, for example, in ADME prediction [15].
We would like to thank Dr. Toshiyuki Tashiro (Lifematics Inc.) for helping with the generation of conformers using RDKit. We would like to thank Editage (www.editage.com) for English language editing.
Details of the molecular property distribution of the PDB dataset used in this study are shown in Figure S1. Scatter plots of RMSD versus molecular weight and rotatable bonds are shown in Figure S2. Table S1 shows list of options used for each method. We provide RMSDs between each generated conformer and the PDB structure. The calculation results of 3D descriptors determined using Pipeline Pilot 2018 and Mordred for each conformer with the lowest energy are also presented. Het codes and the SMILES of PDB ligands used in this study were given in csv format.