2026 年 76 巻 4 号 p. 349-360
This study proposes a genomic-assisted breeding framework for sweet maize using a low-density SNP chip. After genotyping 95 inbred lines (68 retained with 8,484 SNPs), we assessed genetic diversity and population structure, revealing four distinct clusters. Using this information, 25 representative lines were crossed in incomplete diallel crossing designs to generate 102 F1 hybrids for phenotyping. An exploratory genome-wide association study (GWAS) on this F1 population identified 15 candidate loci for agronomic traits. Trait-specific positive correlations were observed between parental genetic distance and hybrid performance. This integrated strategy supported the release of four elite cultivars. The GWAS findings are preliminary and require independent validation due to the modest population size. Nevertheless, this work demonstrates that integrating low-density SNP data for diversity analysis and heterotic grouping effectively guides parental selection and accelerates cultivar development.

Sweet maize (Zea mays L. var. saccharata), commonly known as sweet corn, is distinguished by its high sugar content and characteristically sweet taste (Revilla et al. 2021, Szymanek 2009). Unlike grain maize, which is harvested at physiological maturity when kernels are dry and hard, sweet maize is harvested at the milk stage when kernels remain tender and moist, making it a preferred crop for direct human consumption (Ye et al. 2023). Globally, sweet maize has emerged as a valuable crop not only for fresh consumption and processing but also for its high consumer demand and adaptability to sustainable food systems (Sidahmed et al. 2025).
Beyond direct consumption, sweet maize is also used as a key raw material in various processed products, including canned corn and cornmeal, and contributes to biofuel production. From a scientific perspective, sweet maize serves as a valuable model for genetic research (Revilla et al. 2021). Its relatively small genome and the availability of advanced genetic resources (USDA Ag Data Commons 2020) make it suitable for studying complex agronomic traits. Research on sweet maize has enhanced understanding of biological processes including photosynthesis, development, and stress responses (Dang et al. 2023).
Despite its nutritional and economic importance, sweet maize production faces challenges, including susceptibility to pests and diseases, sensitivity to environmental stressors, and the need for improved yield and quality traits (Revilla et al. 2021, USDA-ARS 2022). Addressing these issues requires continuous cultivar improvement.
The need for studying genetic diversity and heterosis patternsThe development of superior cultivars fundamentally relies on effective exploitation of genetic variation. Genetic diversity forms the foundation of biodiversity and enables populations to adapt to changing environments. In plant breeding, wider genetic diversity provides more opportunities to improve key traits, including yield and abiotic/biotic stress tolerance (Dang et al. 2023, Yu et al. 2022). Therefore, a comprehensive assessment of genetic diversity and population structure is essential in modern breeding programs (Baldauf and Hochholdinger 2023, Xiao et al. 2022).
For hybrid crop breeding, understanding and exploiting heterosis is paramount. Heterosis refers to the tendency for hybrid offspring to outperform the mid-parent or better parent (Xiao et al. 2021) and contributes to yield improvement and food security (Sadaiah et al. 2013). It is influenced by parental genetic distance and combining ability (Wan et al. 2022, Xiao et al. 2021). Thus, elucidating heterosis patterns and identifying optimal parental combinations are essential for efficient hybrid breeding. Heterotic groups represent genetically distinct germplasm pools that, when crossed, tend to produce superior hybrids.
To fully harness genetic diversity and heterosis, an integrated approach combining genomic analysis with phenotypic evaluation is required (Crossa et al. 2017), allowing informed rather than empirical parental selection.
Low-cost SNP chips and the objectives of this studyRecent advances in genotyping technologies have made such an integrated approach feasible. Single nucleotide polymorphisms (SNPs) are the most abundant form of genetic variation and provide a rich resource for genetic studies. SNP chips enable high-throughput genotyping and have revolutionized genetic studies. In particular, low-density SNP arrays have improved accessibility of genomic tools for breeding programs (Dou et al. 2023, Guo et al. 2022, Xu et al. 2020, Yu et al. 2022), allowing analysis of larger sample sizes and increasing statistical power for genetic diversity assessment, population structure inference, and genome-wide association studies (GWAS).
However, a streamlined pipeline that integrates low-density SNP data to simultaneously (i) delineate heterotic groups, (ii) predict combining ability, and (iii) perform GWAS directly on the derived F1 hybrid population—thereby closing the loop from genomic analysis to hybrid prediction—remains insufficiently documented in sweet maize. GWAS using F1 hybrids, while not yet mainstream, offers the advantage of directly dissecting the genetic basis of heterosis and hybrid performance in the target breeding population. Therefore, this study employed a custom low-cost SNP chip to genotype a diverse panel of sweet maize inbred parental lines, and further utilized the derived F1 hybrid population for genetic analysis and breeding applications.
Our specific objectives were to: (1) assess the genetic diversity and population structure; (2) evaluate hybrid performance and combining ability; (3) investigate the correlation between parental genetic distance and heterosis; and (4) perform a preliminary genome-wide scan to identify candidate genomic regions associated with agronomic traits via GWAS in the F1 hybrids. Ultimately, the aim is to construct and validate a practical breeding process from genomic analysis to variety selection. The overarching aim was to establish a practical, genomics-informed breeding pipeline that translates genomic insights into efficient parental selection and, ultimately, cultivar development.
This study utilized a panel of 95 sweet maize inbred lines originating from diverse geographical regions representing wide genetic variability. The inbred lines were selected based on phenotypic performance, pedigree information, and breeding value to maximize genetic diversity. These materials represent multiple genetic backgrounds, encompassing endosperm types (su, se, sh2), maturity groups (early, mid, late) and kernel colors (yellow, white, mixed).
The inbred lines used in this study are listed in Supplemental Table 1 along with their group assignment. To facilitate further validation, the genetic materials are available to the research community. Interested researchers are encouraged to contact the corresponding author for access or collaboration inquiries.
SNP chip genotyping and quality controlGenomic DNA was extracted from young leaf tissues and genotyped using a custom SNP chip containing 9,433 markers distributed across the maize genome. The chip was designed based on the B73 reference genome (version 4) assembly, with markers uniformly distributed across the maize genome. Genotyping quality control was conducted using PLINK (version 1.9) (Chang et al. 2015, Purcell et al. 2007), excluding SNPs with call rates below 95% or minor allele frequencies (MAF) below 5%. Inbred lines with genotyping call rates below 97% were also excluded. After quality checks, 68 inbred lines and 8,484 SNP loci were retained. The raw SNP genotype data for the inbred lines have been deposited in the ScienceDB database repository (DOI: 10.57760/sciencedb.37070) as Supplemental Table 2.
Genetic diversity and population structure analysisGenetic diversity within the germplasm was assessed by calculating heterozygosity, MAF, and polymorphism information content (PIC) using PowerMarker (version 3.25) (Liu and Muse 2005). Genetic distances among inbred lines were computed using Roger’s algorithm (Rogers 1972), and a neighbor-joining (NJ) dendrogram was constructed.
Population structure was inferred using STRUCTURE software (version 2.3.4) (Falush et al. 2003, Pritchard et al. 2000). The optimal K was determined using the Evanno ΔK method (Evanno et al. 2005) implemented in STRUCTURE HARVESTER (Earl and vonHoldt 2012). Principal component analysis (PCA) in TASSEL (5.0) (Bradbury et al. 2007) complemented clustering results. Visualization of population structure and PCA outputs was performed in R (version 4.3.2) (R Core Team 2023). The inferred population structure and genetic relationships were used to inform the selection of parental lines for crossing, as described in the next section.
Incomplete diallel crossbreeding and phenotypic data collectionIncomplete diallel crossing designs (Wu and Matheson 2004) were implemented to systematically evaluate the combining ability and heterosis patterns among selected parental lines. The selection of the 25 parental inbred lines was based on the following comprehensive criteria: (1) to maximize the representation of the four distinct genetic clusters identified in the population structure analysis described in the Results (Under Population structure and genetic relationships), ensuring wide genetic diversity; (2) to encompass a wide range of phenotypic variation for key agronomic traits (e.g., plant height, ear characteristics); and (3) to incorporate important breeding values, such as disease resistance and adaptation, based on prior phenotypic investigations. These selected lines are listed in Supplemental Table 3.
The 102 F1 hybrids were generated through three independent but partially overlapping incomplete diallel crossing designs (Populations I, II, and III), which shared a subset of common parental lines to connect the datasets while each also included unique parents to expand the genetic and phenotypic space sampled. A randomized complete block design was adopted across three years at two locations: Xi’an and Hainan. Each hybrid was planted in single rows, with three replications. Rows were 4 m in length with 0.6 m between rows and 0.25 m between plants.
Phenotypic data were collected for agronomic and yield traits. Agronomic traits measured included plant height (PH), ear height (EH), and ear height coefficient (EHC), while yield traits included ear length (EL), ear diameter (ED), ear row number (ER), kernel number per row (KNPR), and kernel number per ear (KNPE). All traits were measured following standard protocols.
To obtain robust phenotypic values for genetic analyses, the raw data were processed as follows. First, the multi-environment mean was calculated for each hybrid-trait combination. Subsequently, a linear mixed model was fitted using the lme4 package (Bates et al. 2015) in R (version 4.3.2) (R Core Team 2023), with F1 hybrid genotype as a random effect and environment (location-year) as a fixed effect. Variance components were estimated using Restricted Maximum Likelihood (REML) (Patterson and Thompson 1971). The Best Linear Unbiased Prediction (BLUP) for each genotype was then extracted (Bernardo 1996). The BLUP value for each F1 hybrid, estimated from the linear mixed model described above, was used as the phenotypic input for all subsequent genetic analyses, including combining ability analysis (general combining ability, GCA; specific combining ability, SCA) and GWAS. These BLUP values integrate information across environments and provide robust estimates of genotypic performance.
Genome-wide association analysis of F1 hybridsGWAS was performed as an initial scan to identify candidate genomic regions associated with key agronomic and yield traits. The analysis utilized the BLUP values for each F1 hybrid, as described in the preceding subsection. Hybrid genotypes were inferred from their parental genotypes: for SNPs where both parents were homozygous and identical, the hybrid was assigned the same genotype; where parents were homozygous but different, the hybrid was coded as heterozygous; and for SNPs with parental heterozygosity, the hybrid genotype was set as missing. Markers with <20% missing data were retained, resulting in 7,280 SNPs for association testing.
The GWAS was conducted using the GAPIT3 package in R (Wang and Zhang 2021). To account for population structure and kinship, we employed three robust models: the Mixed Linear Model (MLM), Compressed Mixed Linear Model (CMLM), and Bayesian-information and Linkage-disequilibrium Iteratively Nested Keyway (BLINK) (Huang et al. 2019). The Q matrix (population structure covariates) was derived from parental lines and used as an approximation of hybrid structure; therefore associations were treated as exploratory. Models with high false-positive rates (GLM, general linear model) or excessive correction (FarmCPU, Fixed and random model Circulating Probability Unification (Liu et al. 2016)) were excluded from final locus selection, but their results, along with those from MLM, CMLM, and BLINK, are presented in Supplemental Fig. 1 for transparency.
A significance threshold of p < 1/N (where N is the number of valid markers for each trait) was applied. Candidate genes were identified by examining annotations within a 100-kb window of each significant SNP, based on the B73 RefGen_v4 coordinates from MaizeGDB 2024 (Shamimuzzaman et al. 2020).
Breeding practice and cultivar development pipelineTo translate the genomic insights from the preceding analyses into breeding outcomes, a practical pipeline was implemented, integrating steps from germplasm screening to hybrid selection. Parental lines for constructing hybrid combinations were selected based on the criteria detailed in the incomplete diallel crossing subsection (representing divergent genetic clusters, phenotypic variation, and breeding value), which were informed by the initial genetic diversity and population structure assessment described in the SNP genotyping and genetic diversity analysis subsections above.
Candidate hybrids generated from these crosses were evaluated in the multi-environment trials as described in the incomplete diallel crossing subsection. Elite hybrids that consistently outperformed commercial check varieties were advanced for development. The agronomic performance of these superior hybrids, which validates the effectiveness of this pipeline, is presented in the Results section (Performance of elite hybrids).
Genotyping of 95 sweet maize inbred lines using the custom SNP chip resulted in retention of 68 inbred lines and 8,484 high-quality SNP loci after quality control. The mean MAF across all loci was 0.238, with 61.3% of SNP loci having MAF values greater than 0.20 (Fig. 1A). The average heterozygosity across loci was 0.099 (Fig. 1B), and the mean PIC was 0.307 (Fig. 1C). Detailed information on MAF, heterozygosity, and PIC values is provided in Supplemental Table 4, and the genetic distance matrix for all pairwise comparisons is available in Supplemental Table 5 (both accessible via the ScienceDB repository, DOI: 10.57760/sciencedb.37070).

Genotypic data distribution of 68 sweet maize inbred lines based on 8,484 SNPs. (A) MAF, minor allele frequencies. (B) Heterozygosity. (C) PIC, polymorphism information content.
Cluster analysis based on genetic distance grouped the inbred lines into four distinct clusters (Fig. 2A). Genetic diversity statistics varied among these clusters (Fig. 2B, 2C). Population structure analysis determined the optimal number of subpopulations to be K = 4, based on the lowest cross-validation error (Fig. 2D), and displayed the estimated membership coefficients for each line (Fig. 2E). Principal component analysis (PCA) showed a three-dimensional distribution of individuals that was consistent with the four-cluster pattern (Fig. 2F).

Comprehensive analysis of genetic diversity and population structure among 68 sweet maize inbred lines. (A) Neighbor-joining phylogenetic tree constructed based on Roger’s genetic distance, showing four major genetic clusters (A–D) indicated by color branches. (B) Heatmap of pairwise Roger’s genetic distances between inbred lines; blue to red gradient reflects increasing genetic distance. (C) Box-plot comparing the distribution of genetic distances within each of the four clusters (A–D). (D) Cross-validation (CV) error for different assumed numbers of sub-populations (K = 1–5). (E) Population structure bar plot at K = 4; each vertical bar represents one inbred line, partitioned into colored segments proportional to its inferred ancestry from the four sub-populations. (F) Principal component analysis (PCA) scatter plot based on 8,484 SNP markers.
Table 1 summarizes phenotypic variation and heritability estimates for eight agronomic and yield-related traits across the 102 F1 hybrids. Coefficients of variation (CV) ranged from 8.05% (ED) to 23.14% (KNPE). Broad-sense heritability (H2) estimates ranged from 0.28 (KNPE) to 0.58 (PH), while narrow-sense heritability (h2) estimates ranged from 0.03 (KNPR) to 0.18 (PH). Most traits exhibited approximately normal distributions, with skewness and kurtosis values below 1.0 for all traits except ED and EHC (Table 1, Fig. 3A). Hybrids grown in Xi’an (long-day conditions) exhibited higher mean values for most traits than those grown in Hainan (short-day conditions) (Fig. 3B). Phenotypic values for inter-group crosses were significantly higher than those for intra-group crosses (Fig. 3C). BLUP values demonstrated high consistency across trial years (Fig. 3D). The raw phenotypic data for each F1 hybrid across all environments (locations and years) are provided in Supplemental Table 6, and the BLUP values used for phenotypic correlation and combining ability analyses are provided in Supplemental Table 7.
| Trait | Range | Mean | Skewness | Kurtosis | CVa (%) | H2b | h2c |
|---|---|---|---|---|---|---|---|
| PH (cm) | 120.00–275.00 | 204.34 ± 26.96 | –0.34 | –0.21 | 13.19 | 0.58 | 0.18 |
| EH (cm) | 7.00–120.00 | 78.75 ± 16.41 | –0.27 | 0.20 | 20.84 | 0.50 | 0.14 |
| EHC (%) | 3.54–58.47 | 38.54 ± 6.38 | –0.27 | 1.22 | 16.55 | 0.45 | 0.10 |
| EL (cm) | 8.00–20.70 | 13.89 ± 2.11 | –0.20 | –0.13 | 15.19 | 0.46 | 0.13 |
| ED (cm) | 2.29–5.62 | 4.10 ± 0.33 | –0.13 | 3.12 | 8.05 | 0.52 | 0.09 |
| ER | 10.00–26.00 | 17.81 ± 2.37 | 0.11 | 0.04 | 13.31 | 0.35 | 0.16 |
| KNPR | 18.00–52.00 | 37.00 ± 6.36 | –0.30 | –0.21 | 17.19 | 0.34 | 0.03 |
| KNPE | 216.00–1170.00 | 660.97 ± 152.94 | 0.06 | –0.29 | 23.14 | 0.28 | 0.06 |
a CV, coefficient of variation.
b H2, broad-sense heritability.
c h2, narrow-sense heritability.
PH, plant height; EH, ear height; EHC, ear height coefficient; EL, ear length; ED, ear diameter; ER, ear row number; KNPR, kernel number per row; KNPE, kernel number per ear.

Phenotypic distribution, environmental effect, heterosis performance, and data reliability assessment for eight agronomic and yield-related traits in sweet maize hybrids. (A) Histograms showing the frequency distribution of the traits. The black curve represents the fitted normal distribution. (B) Box-plot comparison of trait values between two trial locations (Xi’an vs. Hainan). (C) Box-plot comparison of trait values between intra-cluster and inter-cluster crosses. (D) Quantile-Quantile (Q-Q) plot comparing the distribution of multi-year mean phenotypic values against the BLUP values for all hybrids; BLUP, best linear unbiased prediction. The diagonal line (y = x) represents perfect agreement.
Analysis of variance (ANOVA) was performed to evaluate the effects of environmental factors and genetic crosses. Table 2 presents the F-values from the combined ANOVA across three hybrid populations (I, II, and III), with significance levels indicated for year, location, cross type (intra-/inter-cluster), hybrid genotype, and genotype-by-environment interaction. Location effects were significant (p < 0.05 to p < 0.001) for all traits across all three populations, except for EHC in Populations II and III. Cross type (intra-/inter-cluster) effects were significant for the majority of traits, particularly for PH, EH, EL, ED, KNPR and KNPE, with significance levels ranging from p < 0.05 to p < 0.001. Hybrid genotype effects were highly significant (p < 0.001) for all traits in Population I and for most traits in Populations II and III. Genotype-by-environment interaction effects were significant for selected traits in Populations II and III, but were generally non-significant in Population I. Year effects were non-significant for all traits across all three populations.
| Source of variation | Hybrid population | D.F.a | PH | EH | EHC | EL | ED | ER | KNPR | KNPE |
|---|---|---|---|---|---|---|---|---|---|---|
| Year | I | 2 | 0.29 | 1.17 | 1.05 | 0.02 | 1.06 | 1.99 | 0.13 | 0.42 |
| II | 2 | 0.04 | 010 | 0.28 | 1.63 | 0.12 | 0.29 | 2.01 | 2.10 | |
| III | 2 | 0.03 | 0.20 | 0.15 | 0.31 | 0.67 | 0.02 | 0.95 | 0.68 | |
| Locb | I | 1 | 21.78*** | 20.06*** | 5.10* | 36.76*** | 25.59*** | 14.13*** | 10.63** | 20.00*** |
| II | 1 | 89.69*** | 30.14*** | 0.05 | 67.99*** | 26.03*** | 12.21*** | 93.46*** | 102.63*** | |
| III | 1 | 39.25*** | 14.26*** | 0.26 | 33.86*** | 6.81** | 7.54** | 65.44*** | 63.58*** | |
| Intra/inter-cluster crosses | I | 1 | 13.86*** | 16.28*** | 5.77* | 14.86*** | 17.05*** | 0.30 | 59.99*** | 28.85*** |
| II | 1 | 11.77*** | 8.01** | 0.36 | 6.77** | 10.15** | 1.13 | 2.04*** | 9.68** | |
| III | 1 | 10.73*** | 5.33* | 0.02 | 1.98 | 10.57*** | 0.12 | 4.26* | 3.38 | |
| Hybrids | I | 41 | 15.73*** | 9.70*** | 5.52*** | 6.56*** | 9.04*** | 5.29*** | 8.47*** | 8.43*** |
| II | 35 | 5.84*** | 4.35*** | 3.66** | 5.56*** | 4.62*** | 3.46*** | 1.66* | 1.70* | |
| III | 23 | 6.43*** | 5.89*** | 6.36*** | 4.60*** | 4.91*** | 2.68*** | 2.73*** | 1.65 | |
| Hybrids X Locc | I | 41 | 0.99 | 0.678 | 0.75 | 1.09 | 1.09 | 0.44 | 0.539 | 0.483 |
| II | 35 | 3.58*** | 3.04*** | 1.79** | 5.55*** | 1.03 | 1.18 | 3.38*** | 1.25 | |
| III | 23 | 7.35*** | 4.14*** | 3.59*** | 3.41*** | 0.94 | 2.88*** | 3.31*** | 3.67*** | |
| Error | I | 246 | 23.31 | 15.46 | 5.68 | 2.01 | 0.23 | 2.60 | 6.97 | 176.48 |
| II | 214 | 28.21 | 16.71 | 6.54 | 2.01 | 0.39 | 2.20 | 5.51 | 124.97 | |
| III | 121 | 26.31 | 16.59 | 7.36 | 2.10 | 0.39 | 2.11 | 6.53 | 150.76 |
a D.F., degrees of freedom.
b Loc, location.
c Hybrids X Loc, interaction between hybrids and location.
*, **, *** significant at p < 0.05, 0.01, and 0.001, respectively.
Table 3 presents the F-values from the analysis of variance for SCA, including the effects of intra-/inter-cluster crosses, hybrids, female parent, male parent, and SCA. Intra-/inter-cluster cross effects were highly significant (p < 0.001) for most traits across all three populations. Hybrid effects were extremely significant (p < 0.001) for all traits, with F-values presented in scientific notation. Female parent effects were significant for the majority of traits in Populations I and III, but were less frequently significant in Population II. Male parent effects were significant for selected traits, particularly KNPE across all populations. SCA effects were significant for most traits in Population I, but were less pronounced in Populations II and III.
| Source of variation | Hybrid population | D.F.a | PH | EH | EHC | EL | ED | ER | KNPR | KNPE |
|---|---|---|---|---|---|---|---|---|---|---|
| Intra/inter-cluster crosses | I | 1 | 95.54*** | 103.45*** | 22.63*** | 101.69*** | 14.64*** | 23.52*** | 172.38*** | 124.61*** |
| II | 1 | 91.11*** | 34.80*** | 1.36 | 52.78*** | 23.35*** | 2.16 | 49.10*** | 34.58*** | |
| III | 1 | 79.22*** | 103.33*** | 6.88** | 13.64*** | 34.23*** | 4.03* | 21.03*** | 14.75*** | |
| Hybrids | I | 41 | 8.07E33*** | 7.93E33*** | 1.35E34*** | 7.29E33*** | 6.86E33*** | 6.39E33*** | 8.32E33*** | 9.22E33*** |
| II | 35 | 5.03E33*** | 6.14E33*** | 9.26E33*** | 6.50E33*** | 1.04E34*** | 8.38E33*** | 7.65E33*** | 5.52E33*** | |
| III | 23 | 9.03E33*** | 2.67E34*** | 8.24E33*** | 9.60E33*** | 7.15E33*** | 7.28E33*** | 6.743E33*** | 6.65E33*** | |
| Female | I | 6 | 1.51 | 2.50* | 2.17* | 2.29* | 2.73* | 0.88 | 5.83*** | 3.30** |
| II | 5 | 2.19 | 1.54 | 1.13 | 2.50* | 2.67* | 1.20 | 2.95* | 3.23** | |
| III | 5 | 2.68* | 2.48* | 2.94* | 2.64* | 6.93*** | 1.39 | 3.53** | 5.87*** | |
| Male | I | 5 | 3.24** | 1.12 | 0.46 | 2.69* | 2.43* | 2.30* | 4.75*** | 5.52*** |
| II | 5 | 0.51 | 1.73 | 1.07 | 1.43 | 3.51** | 2.22 | 2.15 | 7.61*** | |
| III | 3 | 2.78* | 1.15 | 0.60 | 1.99 | 4.67** | 3.74* | 4.94** | 3.48* | |
| SCA | I | 30 | 9.05*** | 4.77*** | 1.64* | 4.90* | 9.10*** | 2.21** | 8.80*** | 7.38*** |
| II | 25 | 1.66* | 2.01** | 1.29 | 2.81*** | 3.19*** | 1.56 | 1.70* | 1.13 | |
| III | 13 | 3.44*** | 2.24* | 1.80 | 2.36** | 4.70*** | 1.04 | 1.85* | 1.81 | |
| Error | I | 251 | 12.34 | 6.17 | 0.79 | 0.78 | 0.20 | 0.59 | 4.94 | 108.88 |
| II | 215 | 3.87 | 3.43 | 0.47 | 0.59 | 0.39 | 0.30 | 0.87 | 18.12 | |
| III | 143 | 8.73 | 3.16 | 0.98 | 0.47 | 0.21 | 0.01 | 1.39 | 28.64 |
F-values for the “Hybrids” source are presented in scientific notation due to exceptionally small error variances.
a D.F., degrees of freedom.
*, **, *** significant at p < 0.05, 0.01, and 0.001, respectively.
Significant positive correlations were observed between parental genetic distance and both phenotypic performance and SCA for multiple yield-related traits (Fig. 4). Genetic distance showed significant positive correlations with PH, EH, EHC, ED, KNPR and KNPE, while no significant correlation was detected for EL and ER (Fig. 4A). Genetic distance also exhibited significant positive correlations with SCA effects for all eight target traits, with the strongest associations observed for KNPR, KNPE and EH (Fig. 4B). PH was strongly correlated with EH (r = 0.644, p < 0.001), and significant positive correlations were also observed among most yield-related traits (Fig. 4A).

Triangular heatmaps of correlation analysis between parental genetic distance (GD) and hybrid performance in sweet maize. (A) Correlations with phenotypic performance of eight traits. (B) Correlations with specific combining ability (SCA) effects of eight traits. *, **, *** significant at p < 0.05, 0.01, and 0.001, respectively.
An exploratory GWAS was performed on 102 F1 hybrids to identify candidate genomic regions associated with key agronomic and yield traits. To minimize false positives, MLM (Yu et al. 2006), CMLM (Zhang et al. 2010) and BLINK (Huang et al. 2019) models, which incorporate population structure (Q) and kinship (K), were employed as the primary analytical approaches. The GWAS was conducted using the GAPIT3 package in R (Wang and Zhang 2021). A significance threshold of p < 1/N (where N is the number of valid markers for each trait) was applied.
A total of 15 significant SNP-trait associations were identified, representing 12 distinct candidate genomic regions across chromosomes 1, 2, 5, 6, 7, and 10. The number of valid markers varied among traits (ranging from 960 to 1,216) due to trait-specific missing data filtering criteria. Table 4 summarizes these candidate regions, including their physical positions, MAF, phenotypic variance explained (PVE), detection models, and putative candidate genes within 100 kb.
| Trait | Chra | Physical position (bp) | Nb | p Value | MAFc | PVEd | GWAS model | Candidate gene (within 100 kb) | Gene annotation / Putative function |
|---|---|---|---|---|---|---|---|---|---|
| PH | 2 | 936,801 | 960 | 3.58E-04 | 0.20 | 0.12 | MLM, CMLM, BLINK | – | – |
| 2 | 1,083,914 | 960 | 3.58E-04 | 0.20 | 0.12 | MLM, CMLM | – | – | |
| 6 | 169,733,977 | 960 | 4.32E-04 | 0.22 | 0.64 | MLM, CMLM, BLINK | Zm00001d039086 | CHUP1-like protein, chloroplast positioning | |
| 10 | 4,248,287 | 960 | 9.59E-04 | 0.25 | 0.09 | MLM, CMLM | Zm00001d023378 | ZmTBL10, plant height regulation | |
| 10 | 13,148,937 | 960 | 1.87E-04 | 0.12 | 0.35 | MLM, CMLM | – | – | |
| EH | 1 | 38,969,624 | 1,216 | 3.16E-06 | 0.30 | 0.11 | BLINK | – | – |
| 5 | 3,612,515 | 1,216 | 1.69E-05 | 0.43 | 0.15 | BLINK | Zm00001d013035 | smt1/1, sterol methyltransferase, brassinosteroid signaling | |
| 6 | 169,733,977 | 1,216 | 2.04E-05 | 0.22 | 0.64 | MLM, CMLM, BLINK | Zm00001d039086 | CHUP1-like protein, chloroplast positioning | |
| 5 | 132,637,362 | 1,216 | 3.26E-07 | 0.18 | 0.08 | BLINK | – | – | |
| EHC | 5 | 132,637,362 | 960 | 1.33E-05 | 0.18 | 0.08 | BLINK | – | – |
| ER | 2 | 936,801 | 960 | 2.28E-04 | 0.20 | 0.05 | BLINK | – | – |
| 10 | 4,248,287 | 1,216 | 6.86E-05 | 0.25 | 0.18 | MLM, CMLM | Zm00001d023378 | ZmTBL10, plant height regulation | |
| KNPR | 1 | 47,788,881 | 960 | 5.75E-04 | 0.11 | 0.16 | MLM, BLINK | Zm00001d028824 | ZmGATL5, pectin biosynthesis |
| 1 | 47,790,506 | 960 | 5.75E-04 | 0.11 | 0.16 | MLM, BLINK | Zm00001d028824 | ZmGATL5, pectin biosynthesis | |
| KNPE | 7 | 176,067,682 | 960 | 2.80E-07 | 0.33 | 0.68 | BLINK | – | – |
a Chr, chromosome.
b N, number of valid markers.
c MAF, minor allele frequency.
d PVE, phenotypic variance explained.
The significance threshold applied to the p value listed in the “p Value” column is p < 1/N, where N is the number of valid markers for each trait (shown in the “N” column). Because N differs among traits due to trait-specific missing data filtering, the threshold value is trait-specific. Because N differs among traits due to trait-specific missing data filtering, the threshold value is trait-specific.
“–” indicates that no known or annotated gene was identified within the 100 kb window based on the B73 RefGen_v4 reference genome (MaizeGDB). Only genes with functional annotations or literature support are listed.
For plant architecture traits, a pleiotropic region on chromosome 2 (936,801–1,083,914 bp) was associated with both PH and ER. Another locus on chromosome 6 (169,733,977 bp) was associated with both PH and EH, located adjacent to Zm00001d039086, which encodes a Chloroplast Unusual Positioning 1 (CHUP1)-like protein involved in chloroplast positioning (Oikawa et al. 2003). A region on chromosome 10 (4,248,287 bp) was associated with both PH and ER, located near Zm00001d023378 (ZmTBL10), a gene previously reported to influence plant height in maize (Wang et al. 2023). An EH-associated SNP on chromosome 5 (3,612,515 bp) was identified near Zm00001d013035 (smt1/1), a sterol methyltransferase homolog involved in brassinosteroid biosynthesis. Additionally, a locus on chromosome 5 (132,637,362 bp) was associated with both EH and EHC.
For yield-related traits, two significant SNPs on chromosome 1 (47,788,881 and 47,790,506 bp) were associated with KNPR, both located near Zm00001d028824 (ZmGATL5), a gene implicated in pectin biosynthesis. A region on chromosome 7 (176,067,682 bp) was associated with KNPE, suggesting a potential Quantitative Trait Locus (QTL) for kernel development.
The Manhattan and quantile-quantile (Q-Q) plots for all traits across the five GWAS models (GLM, FarmCPU (Liu et al. 2016), MLM, CMLM, and BLINK) are provided in Supplemental Fig. 1. Candidate genes were identified by examining annotations within a 100-kb window of each significant SNP, based on the B73 RefGen_v4 coordinates from MaizeGDB 2024 (Portwood et al. 2019, Shamimuzzaman et al. 2020). These exploratory findings provide candidate genomic regions for future validation and marker-assisted selection, but require confirmation in independent populations. The candidate genes listed in Table 4 should be considered putative, as their functional relevance to the associated traits remains to be experimentally validated.
Performance of elite hybridsFour elite hybrids—Shaan K512, Shaan K8143, Shaan K7128 (Guo et al. 2024), and Shaan K9148—were selected from the evaluated pool based on heterotic grouping, and were subsequently advanced through the formal cultivar release pipeline. Their agronomic performance was evaluated in multi-environment trials and compared with that of a currently cultivated commercial check variety.
Table 5 summarizes the agronomic traits, yield performance, and parental genetic distances of these four registered hybrids. The parental genetic distances, derived from SNP marker data (Supplemental Table 5), ranged from 0.39 to 0.44, indicating moderate to high genetic divergence between the parental lines from different heterotic groups. All four elite hybrids exhibited superior agronomic performance compared to the check variety across multiple yield-related traits. Shaan K9148 showed the highest fresh ear yield (1216.80 ± 65.74 kg/667 m2), representing a 22.05% yield increase over the check (Table 5). These performance data provide direct empirical validation that the integrated genomic pipeline effectively guides parental selection and accelerates commercial cultivar development in sweet maize.
| Variety name | Female | Male | GDa | PH (cm) |
EH (cm) |
ER | KNPR | Fresh ear yield (kg/667 m2) |
Yield increase |
|---|---|---|---|---|---|---|---|---|---|
| Shaan K512 | 85 | 87 | 0.39 | 205.80 ± 14.59 | 92.40 ± 6.61 | 16.87 ± 1.03 | 40.03 ± 2.02 | 1115.45 ± 50.16 | 11.88% |
| Shaan K8143 | 92 | 81 | 0.44 | 181.44 ± 5.00 | 61.17 ± 8.31 | 18.33 ± 1.53 | 33.22 ± 1.35 | 1076.10 ± 50.06 | 7.94% |
| Shaan K7128 | 83 | 18 | 0.39 | 216.50 ± 9.19 | 84.00 ± 2.41 | 15.60 ± 0.57 | 39.90 ± 4.10 | 1180.37 ± 80.05 | 18.40% |
| Shaan K9148 | 94 | 69 | 0.42 | 215.15 ± 7.28 | 98.20 ± 2.13 | 17.80 ± 0.48 | 39.00 ± 2.41 | 1216.80 ± 65.74 | 22.05% |
| Check Variety | – | – | – | 176.28 ± 14.22 | 60.19 ± 8.62 | 14.93 ± 1.04 | 36.83 ± 3.36 | 996.96 ± 94.83 | – |
a GD, a genetic distance between parental lines was calculated based on SNP markers; data are sourced from Supplemental Table 5.
Values are presented as mean ± standard deviation from multi-environment trials.
The check variety is a currently cultivated commercial sweet maize hybrid.
The SNP genotyping revealed substantial genetic variation within the sweet maize inbred panel, as evidenced by a mean PIC of 0.307 and a balanced MAF distribution (Fig. 1C, 1A). This level of diversity is consistent with previous findings that sweet maize retains considerable polymorphism useful for improvement (Liu et al. 2003, Revilla et al. 2021). More importantly, this variation was not merely descriptive but was actively translated into breeding outcomes: the four registered commercial hybrids developed in this study were all derived from parental lines selected based on this genomic information (Table 5).
Population structure analysis consistently resolved the 68 inbred lines into four distinct subpopulations (K = 4, Fig. 2A, 2D, 2E; PCA, Fig. 2F). This clear stratification underscores the necessity of accounting for population structure in association mapping to minimize spurious associations (Yu et al. 2022). Crucially, hybrids derived from inter-group crosses significantly outperformed intra-group crosses for multiple yield-related traits (Fig. 3C), suggesting that these SNP-based clusters may represent useful heterotic patterns. This finding aligns with the core principle of hybrid breeding—that genetic divergence between parental lines enhances heterosis (Wan et al. 2022, Xiao et al. 2021).
Combining ability analysis further supported this strategy. A significant, trait-specific positive correlation was observed between parental genetic distance and SCA for KNPR, KNPE, and EH (Fig. 4B). However, genetic distance alone is insufficient for predicting hybrid performance. The four registered cultivars were all derived from crosses between parents belonging to divergent heterotic groups, with parental genetic distances ranging from 0.39 to 0.44 (Table 5). Several specific cross combinations also exhibited superior SCA for yield-related traits in the combining ability analysis (Table 3), reinforcing the role of genetic divergence and favorable non-additive effects in driving heterosis. Nevertheless, maintaining strong GCA remains crucial for long-term breeding progress and yield stability (Laude and Carena 2015).
A cost-efficient, translatable pipeline from SNP chip to cultivar releaseA major contribution of this study is the demonstration of a complete, cost-efficient breeding pipeline that integrates low-density SNP genotyping, genetic diversity analysis, heterotic grouping, combining ability assessment, and exploratory GWAS into a unified workflow—culminating in the development and official registration of four elite commercial hybrids (Table 5). While many studies have reported genomic analyses in sweet maize, few have closed the loop from marker data to marketable product.
The ultimate validation of this pipeline lies not in statistical metrics but in agronomic performance. As summarized in Table 5, all four registered hybrids—Shaan K512, Shaan K8143, Shaan K7128, and Shaan K9148—consistently outperformed the commercial check variety across multiple environments. For instance, Shaan K9148 exhibited a 22.05% yield increase over the check (Table 5). These results provide tangible proof that genomic insights derived from low-cost SNP chips can directly inform parental selection and accelerate genetic gain.
This pipeline is particularly relevant for public-sector and resource-limited breeding programs. The custom SNP chip used in this study provides genome-wide marker coverage at a fraction of the cost of high-density arrays or sequencing-based genotyping, making genomic-assisted breeding accessible to programs where cost has historically been a barrier (Xu et al. 2020, Yu et al. 2022). The approach demonstrated here—using low-density markers for diversity analysis, heterotic grouping, and targeted trait dissection—offers a replicable model for sweet maize and potentially other under-resourced crop breeding programs. This work thus provides a replicable roadmap for sweet maize breeding programs worldwide, particularly for public-sector and resource-limited initiatives aiming to enhance efficiency and genetic gain, demonstrating that meaningful genomic-assisted breeding does not necessarily require high-density arrays or large sequencing investments.
Exploratory GWAS in F1 hybrids: Candidate loci, biological plausibility, and acknowledged limitationsThis exploratory GWAS scan, performed directly on 102 F1 hybrids, identified 15 significant SNP-trait associations representing 12 distinct candidate genomic regions (Table 4). Several of these loci co-localized with genes of known function, providing biological plausibility for their putative roles in trait regulation. However, we explicitly acknowledge several important limitations of this GWAS approach and emphasize that these findings are strictly preliminary, requiring independent validation.
First, the sample size of 102 F1 hybrids derived from 25 parental lines is insufficient for GWAS. This limitation reduces statistical power and increases the risk of false positives and inflated effect sizes (the Beavis effect). The observed distortion in Q-Q plots for some trait–model combinations (e.g., GLM and FarmCPU in Supplemental Fig. 1) likely reflects this constraint. Such overestimation of effect sizes is well-documented in underpowered GWAS and underscores the need for independent validation (Korte and Farlow 2013).
Second, GWAS using F1 hybrids—while enabling direct dissection of heterotic loci in the target breeding population—is not yet a mainstream approach and introduces complexities in statistical correction. The Q matrix used in our models was derived from parental inbred lines based on STRUCTURE analysis at K = 4. However, this matrix may not fully capture the genetic stratification of the hybrid population, and we cannot rule out the possibility of overcorrection, as suggested by the downward skew of Q-Q plots for some traits under certain models (Supplemental Fig. 1).
To address these limitations, we adopted a conservative analytical strategy: (i) we retained only associations detected by robust models (MLM, CMLM, BLINK) that incorporate both population structure (Q) and kinship (K); (ii) we excluded models with high false-positive rates (GLM) or evidence of excessive correction (FarmCPU) from final locus selection; (iii) we removed all associations with PVE > 1 to minimize the inclusion of likely spurious signals; and (iv) we explicitly report MAF, p-values, and detection models for each candidate locus to ensure transparency (Table 4).
We therefore emphasize that the associations reported here should be considered strictly preliminary. They provide hypotheses and prioritize genomic regions for future investigation, but are not ready for direct implementation in marker-assisted selection without independent validation. The candidate genes listed in Table 4, while biologically plausible, require functional validation through complementary approaches such as fine-mapping, haplotype analysis, and validation in biparental populations or larger, independent diversity panels.
Notwithstanding these limitations, the consistency of several identified loci with previous GWAS reports in maize lends credence to their potential biological relevance. ZmTBL10 has been repeatedly associated with plant height in diverse maize populations (Wang et al. 2023), and ZmGATL5 has been implicated in cell wall biosynthesis and kernel traits (Li et al. 2024). This concordance with independent studies suggests that, despite the modest sample size, our exploratory scan may have captured biologically relevant regions that merit further investigation.
HLS: Conceptualization, Data curation, Formal analysis, Writing-original draft. KW: Methodology, Investigation. LHH: Methodology, Validation. SNY: Writing-review & editing. ZLL: Conceptualization, Supervision. NZ: Investigation, Validation. HZ: Visualization, Image editing. JKD: Funding acquisition, Conceptualization, Supervision.
This research was supported by the Natural Science Basic Research Program of Shaanxi (No. 2023-JC-QN-0275), the Key Industrial Chain Project of Shaanxi (2024NC-ZDCYL-01-11), the Match Program of Shaanxi Academy of Sciences (No. 2024p-09), the Science and Technology Program of Xi’an (No. 25NJSYB00048), and the Science and Technology Program of Xianyang (L2024-ZDKJ-ZDGG-NY-0005).