2026 Volume 76 Issue 4 Pages 387-392
Accurate extraction of organ contours from images remains a major bottleneck in quantitative analyses of crop morphology. Here, I present MEGAcontour, a Python-based graphical user interface that streamlines contour extraction and normalized elliptic Fourier descriptor (nEFD) analysis, and demonstrate its utility through genomic prediction of eggplant (Solanum melongena L.) fruit contours. To enable robust extraction under challenging imaging conditions, MEGAcontour integrates the salient object detection model U2-Net. I further fine-tuned U2-Net for eggplant fruit segmentation to improve specificity, particularly by excluding attached fruit branches from extracted contours. PCA of nEFDs indicated that the dominant axis of variation corresponded to fruit elongation. Using DNA markers, I performed leave-one-accession-out genomic prediction for selected nEFD coefficients, reconstructed contours from the predicted coefficients, and evaluated agreement with observed contours using intersection over union (IoU). IoU of reconstructed contours ranged from 0.34 to 0.99 depending on accession. These results provide a practical workflow for contour extraction and support the feasibility of genomic prediction for eggplant fruit shape in breeding programs.

The size and shape of plant organs are key traits affecting crop quality, market class, and end-use value. Consequently, plant breeders have long selected not only for organ size but also for organ shape, and recent reviews summarize genetic and breeding progress across diverse vegetable crops (Goldman et al. 2023).
Quantifying complex biological shapes requires descriptors that capture entire outlines while remaining comparable across samples. Elliptic Fourier descriptors (EFDs) provide a simple yet powerful representation of closed contours (Kuhl and Giardina 1982). Combining EFDs with principal component analysis (PCA) has become a standard approach for dissecting major axes of shape variation, with broad applications in plant organs (Furuta et al. 1995, Iwata et al. 2002, 2015, Sakamoto et al. 2019). In particular, the SHAPE software has contributed to the widespread use of EFD-based morphometrics (Iwata and Ukai 2002).
Eggplant (Solanum melongena L.) originated in South Asia and has been cultivated in Japan for over 1,000 years. Eggplant fruit shape shows substantial diversity, and previous studies have evaluated this diversity using basic measurements (Hurtado et al. 2013), QTL mapping in bi-parental populations (Nunome et al. 2001), and genome-wide association approaches based on categorical shape classes (Liu et al. 2019). However, because fruit shape varies continuously and often involves subtle differences beyond simple length and width, outline-based quantitative approaches may provide higher resolution for genetic dissection and breeding applications.
A practical bottleneck in outline-based analyses is robust contour extraction from images. Contour extraction often requires careful tuning of image-processing thresholds and can fail under variable color, texture, and shadow conditions—issues that are common in fruit phenotyping. Recent advances in salient object detection and segmentation models (Wang et al. 2024) offer a promising solution, yet these methods are not always accessible to breeders and geneticists without machine-learning expertise.
Here, I develop MEGAcontour, a Python-based graphical user interface (GUI) software that streamlines contour extraction and nEFD analysis by incorporating U2-Net-based foreground extraction and optional fine-tuning for specific targets such as eggplant fruits. Using an eggplant mini-core collection, I further evaluate the feasibility of genomic prediction of fruit contours from DNA markers by predicting selected nEFD coefficients, reconstructing contours, and assessing prediction accuracy using distance-based metrics as well as intersection over union (IoU). This study provides an end-to-end workflow for contour-based phenotyping and supports future eggplant breeding that incorporates genomic approaches.
In this study, the World Eggplant Mini Core Collection (WEC100) (Miyatake et al. 2019) was used as the material. After raising seedlings in a growth chamber, plants were grown in pots throughout the experiment under natural daylength conditions in a glass greenhouse at The University of Tokyo (Bunkyo, Tokyo, Japan) in 2022. Three plants per accession were cultivated and arranged randomly within the greenhouse. Owing to poor fruit set during an extremely hot summer, mature fruits suitable for analysis were obtained from 71 accessions.
Fruits were photographed indoors using a camera (Canon PowerShot SX710 HS) fixed in a downward-facing position above a box-shaped imaging setup with a blue background. Fruits were placed 35 cm below the camera. Each fruit was positioned so that the calyx faced the same direction and the fruit surface was as horizontal as possible relative to the camera. External light was blocked and images were captured using the internal lighting of the box without flash. A total of 387 images were selected. The original image size was 5,184 × 3,888 pixels (width × height). More detailed information is available in Supplemental Text 1.
Development of contour extraction softwareMEGAcontour was developed in Python v3.9.13. The GUI includes core functions of SHAPE (Iwata and Ukai 2002): (1) grayscale conversion, binarization, contour detection of target objects, and chain-code encoding; (2) normalized elliptic Fourier transformation of chain codes; and (3) PCA. To semi-automate contour extraction, MEGAcontour integrates U2-Net (Qin et al. 2020) for semantic segmentation, separating foreground from background to enable contour extraction from the segmented region. This was implemented using the remove function of the rembg package (https://github.com/danielgatis/rembg). The MEGAcontour software is available from a GitHub repository (https://github.com/sobaniki/MEGAcontour).
Fine-tuning of U2-Net for eggplant fruit segmentationTo assess whether the default pretrained U2-Net was sufficient for eggplant images and to improve exclusion of attached fruit branches, U2-Net was fine-tuned for eggplant fruit segmentation. Using EISeg v1.1.1 (https://github.com/PaddleCV-SIG/EISeg), 387 mask images were generated including eggplant fruits with attached fruit branches and calyx. Without augmentation, the dataset was split into Train:Validation:Test = 301:43:43 for mini-batch training. The model was implemented in PyTorch. Input images were resized to 320 × 320 pixels. Image colors were normalized with RGB mean = 0.485:0.456:0.406 and std = 0.229:0.224:0.225. The loss function was BCEWithLogitsLoss and the optimizer was Adam. The model was trained with a batch size of 4 and a learning rate of 1e-4 on a computer with NVIDIA GeForce RTX 4070 Ti SUPER 16 GB. A fixed random seed, was used for dataset splitting and training where applicable. The best model based on the validation loss was selected within 30 epochs. The model was integrated into MEGAcontour as a custom U2-Net backend. The segmentation performance was evaluated using pixel-wise F1 score and Intersection-over-Union (IoU). For evaluation, the predicted foreground probability map was threshold at 0.5 to obtain a binary mask, and the corresponding ground-truth mask was binarized at 0.5 after resizing to 320 × 320 pixels.
nEFD calculation, coefficient selection, and clustering visualizationNormalized EFDs (nEFDs) were calculated for each accession as the mean across replicate fruits. The number of coefficients equals four times the number of harmonics (a, b, c, d). Under normalization (invariant to size, rotation, and starting point), the leading coefficients (a1–c1) take fixed constants. Because higher-order coefficients tend to be more affected by noise and because b–c components represent asymmetric variation (Iwata et al. 1998), five harmonics (20 coefficients total) were used and PCA/genomic prediction targeted nine coefficients (a2–a5 and d1–d5). For visualization, clustering of 71 accessions was performed using only PC1–2 with the mclust package. Using the R function prcomp, PCA was performed after the input variables were zero centered and not scaled.
Genomic prediction models, imputation, reconstruction, and evaluationDNA marker genotype data (811 markers) were obtained from Miyatake et al. (2019). Genomic prediction was implemented in R v4.4.2. Five univariate models were evaluated: GBLUP, Gaussian kernel (GK), Bayesian ridge regression (BRR), BayesB, and random forest (RF). Missing genotypes were imputed and the additive relationship matrix was calculated using rrBLUP (Endelman 2011). GBLUP, GK, BRR and BayesB were fitted using BGLR (Pérez and de los Campos 2014) with burn-in = 15,000, nIter = 30,000 and thin = 5. RF was fitted using ranger with default settings (e.g., num.trees = 500) (Wright and Ziegler 2017).
Three multivariate models (mGBLUP, mGK, mBRR) were also implemented using BGLR with the function Multitrait. The covariance matrix for model residuals was defined with type = “UN (unstructured)” and df0 = 5 (i.e., the default setting). A multivariate RF (mRF) was implemented using randomForestSRC (https://github.com/kogalur/randomForestSRC). The function rfsrc was used with the default setting.
Predicted values for the nine coefficients were combined with the fixed 11 coefficients to obtain 20 coefficients, and contour coordinates (x, y) with 500 equally spaced points were reconstructed by inverse elliptic Fourier transformation as closed polygons. Reconstructed coordinates were centered, rotation-aligned, and scaled to a common size according to nEFDs. Prediction error was evaluated using mean symmetric distance (MSD), Hausdorff distance, Procrustes RMSE, and IoU. IoU was calculated as the area of intersection divided by the area of union between the observed and reconstructed closed polygons. Because IoU was computed using vector polygon geometry, no rasterization step or rasterization resolution was used. Prediction performance was assessed using leave-one-accession-out cross-validation (LOOCV).
To improve the efficiency of contour extraction, I developed the GUI software MEGAcontour (Fig. 1). Extracting the contour of the target object from an image is the most critical step for shape analysis. For high-accuracy contour extraction, it is desirable that the target object and the background are clearly distinguishable in color and each is visually uniform. However, eggplant fruits exhibit genetic variation in color ranging from white and green to purple, and some accessions show non-uniform coloration across the fruit surface (Fig. 2A). In addition, eggplant fruits have substantial thickness, and eliminating shadows entirely within the imaging box is not straightforward. Under these conditions, simple grayscale conversion and binarization were insufficient for robust contour extraction (Fig. 2B, 2C).

The MEGAcontour software for organ shape analysis.

Contour extraction of eggplant fruits. A: an original image. B: a failure of binarization based on color threshold. C: a wrong contour extraction. D: a fruit detection by a pretrained U2-Net model. E: a removal of the fruit branch by a fine-tuned U2-Net model. F: a good contour extraction by a fine-tuned U2-Net model.
In contrast, U2-Net successfully extracted eggplant fruit contours even under these challenging conditions (Fig. 2D). For typical objects such as fruits, the pretrained U2-Net model is often sufficient. However, when images included attached fruit branches, contours of the branch were extracted together with the fruit. Therefore, U2-Net was fine-tuned for extracting eggplant fruit contour (Supplemental Fig. 1). The fine-tuned model showed better performance than the pretrained model for both validation and test subsets (e.g., IoU for validation was 0.895 using pretrained and 0.973 using fine-tuned models) (Supplemental Table 1). In fact, the fine-tuned U2-Net model was able to extract only the fruit contour while excluding the fruit branch (Fig. 2E, 2F).
Characteristics of fruit contours in eggplant mini-core collectionThe extracted contours were transformed to nEFDs composed of 20 harmonics (80 coefficients). Of four kinds of coefficients (a, b, c, and d), the a and d coefficients mainly explained the variance between accessions while the b and c coefficients were involved in the variance within each accession (Supplemental Table 2). In addition, the first five harmonics explained 99.8% of the variance up to the 20th harmonics (Supplemental Fig. 2). Iwata et al. (1998) indicated that the asymmetrical variations in roots explained by the b and c coefficients may be a mere artifact. Considering these results, only the a and d coefficients of the first five harmonics were used for the following analyses. Because the a1 coefficient is fixed, nine coefficients (a2–a5 and d1–d5) were selected.
The nine selective coefficients were analyzed by PCA. In this study, 98.7% of contour variation of eggplant fruits was explained by PC1 (Fig. 3A). PC1 primarily corresponded to the aspect ratio (length-to-width ratio). PC2 was associated with the degree of swelling in the lower part of the fruit (left side in the figure), capturing differences between elliptical and ovoid shapes. PC3 reflected the extent of “shoulder” development and could represent features such as pouch-like shapes. Although contour variation was continuous, PC1–2 suggested a possible separation among round, ovoid and elongated fruit types (Fig. 3B).

Principal component analysis of nEFDs from eggplant fruit contours. A: eggplant shape variations explained by PCA. Red, pink, black, light-blue and blue lines indicate mean + 2SD, mean + 1SD, mean, mean – 1SD and mean – 2SD, respectively. B: reconstructed shapes along with PC1–2. Red, blue and green color represented three clusters based on PC1–2.
Genomic prediction of eggplant fruit contours was conducted with univariate and multivariate models. IoU showed an inverse relationship with distance-based error metrics (e.g., MSD, Hausdorff distance, Procrustes RMSE) (Supplemental Figs. 3, 4). All univariate prediction models showed similar performance based on IoU (Fig. 4A) and other indicators (Supplemental Figs. 5–7). The accuracy of multivariate models was not improved compared with univariate models although some of the nine nEFDs showed correlations (Fig. 4B, Supplemental Fig. 8). Through all predictions, the maximum IoU was 0.99 and the minimum was 0.34 (Fig. 5). Some accessions with ovoid or elliptical fruits achieved relatively high IoU. All prediction models exhibited similar trends, although some accessions showed distinct differences among models (e.g., WEC046 with IoU = 0.79–0.92).

Violin plots for genomic prediction accuracy based on IoU. A: univariate models. B: multivariate models. IoU is shown as percentage ranging from 0 to 100. Values in plots indicate the median.

Fruit contours reconstructed from predicted nEFDs. Solid lines with three colors based on the cluster represent original contours. Gray dashed lines represent the reconstruction from the predicted value by the best univariate model based on IoU. The color of accession names represents the best model. Deep red: GBLUP, sky blue: GK, deep green: BRR, purple: BayesB, and orange: RF.
Organ shapes are important traits in vegetable and other crops. Simple parameters for targeting specific organs can be measured using some specialized software (Rodríguez et al. 2010). On the other hand, SHAPE is the standard software using EFD (Iwata and Ukai 2002), which can be applied to analyzing various shape contours. The bottleneck of organism shape analysis may be the object contour extraction from images, because the optimization of extraction processes can be difficult due to the quality of images. Recently, the advancement of salient object detection models, which can be used for contour extraction, has rapidly progressed (Qin et al. 2020, Wang et al. 2024). However, the utilization of these deep learning models may not be easy for many breeders and geneticists. The software development incorporating the new models is important for various users on shape analysis.
Eggplant fruits are often categorized into round, ovoid, pouch-shaped, and elongated. However, the validity of these categories remains unclear. The variation of whole eggplant fruit shapes can mostly be explained by the length-to-width ratio, which can correspond to the difference among round, elongated and intermediate types (Fig. 3A). Although the proportion of PC2–3 was limited, these components may correspond to ovoid and pouch-shaped, respectively. The results indicate the reasonability of the classical category on eggplant fruit shapes.
Genomic predictions on plant organ shapes were performed in rice grains (Iwata et al. 2015) and sorghum seeds (Sakamoto et al. 2019). The results showed that the shape of some eggplant fruits can accurately be predicted using DNA markers (Fig. 5). On the other hand, the prediction accuracy was different in some cases though fruit shapes were similar between accessions. Several reasons for the accuracy difference can be considered. First, even if fruit shapes are similar, the underlying genetic causes may differ (i.e., the genes involved may be diverse). For example, while it has been established that many genes are involved in rice grain shape, the underlying molecular mechanism of most genes remains unclear (Huang et al. 2013). Second, the target population includes groups with significantly different levels of linkage disequilibrium (LD). In fact, Miyatake et al. (2019) suggested that the WEC100 accessions can be classified into four clusters corresponding to their geographical distribution. The WEC100 may be a complex population for demonstrating the potential of GP for complex traits such as fruit shape. Third, the sample size in this study may be too small to capture whole genetic variations involved in eggplant fruit shape. The third question is also discussed below.
The effectiveness of multivariate models was not observed (Fig. 4B), suggesting a limitation due to sample size in this study. For multivariate genomic predictions, breeding schemes considering heritability, genetic correlation and sample size are necessary (Calus and Veerkamp 2011). A recent study suggests that univariate models can be more robust than multivariate models except the case that the correlation among variates is effectively leveraged (Mbebi et al. 2025). Another drawback of multivariate models is their computational complexity. When the number of variables exceeds several hundred, methods such as MegaLMM become necessary (Runcie et al. 2021). Therefore, even when multivariate models can be used for GP, their utility should be carefully considered. Unfortunately, studies on genomic prediction remain unexplored in eggplant. For the utilization of the diverse eggplant genetic resources to breeding, the evaluation of gene bank accessions based on genomic prediction should be further developed such as rice (Tanaka et al. 2021).
Although this study showed the practical analysis workflow of plant organ shape, there are also some challenges. First, extracting the contours of complex objects will likely require extensive fine-tuning or more advanced object recognition models. Although MEGAcontour provides a solution for researchers who are not familiar with programming (Fig. 1), the explosive progress in multimodal vision models may change the situation. Next, the importance of data quality and quantity (e.g., the quality of images and the number of accessions in this study) will likely become relatively greater than that of analytical methods. Efficiently acquiring high-quality data is a challenge in phenotyping, which can directly influence the development of plant breeding. Finally, improvements in GP will likely be driven by more advanced understanding of underlying genetic mechanisms. Current GP models treat marker genotypes as simple inputs and do not explicitly consider variant function. Recently, prediction models that incorporate variant functions have been explored (Zheng et al. 2024), and their application to predicting complex traits such as plant organ shape is anticipated.
M.I. planned the experiment design and performed all experiments including plant cultivation, image acquisition, software development and data analysis, and wrote the manuscript.
This study was partly supported by JSPS KAKENHI (22K14884) and Kieikai Research Foundation.