2026 年 14 巻 1 号 p. 92-109
Raman spectroscopy, when combined with chemometric techniques, has become a powerful tool for quality assessment, process monitoring, and optimization in the vegetable processing industry. This review provides a comprehensive overview of recent advances in the use of Raman chemometric methods for processed vegetables, emphasizing their distinct advantages over other spectroscopic techniques, such as near-infrared (NIR), Fourier-transform infrared (FT-IR), and ultraviolet-visible (UV-Vis) spectroscopy. Due to its high molecular specificity, minimal water interference, and capacity to resolve detailed information on chemical structure, Raman spectroscopy is particularly well-suited for analyzing complex vegetable matrices. Integrated chemometric approaches, including principal component analysis (PCA), partial least squares regression (PLSR), support vector machines (SVM), and convolutional neural networks (CNN), have significantly improved both data interpretation and predictive performance across tasks such as quality classification, constituent quantification, and real-time process monitoring. Despite these advancements, challenges remain, including overlapping signals, limited sample diversity, and the lack of robust systems for in-line implementation. Future research should prioritize the development of transfer learning strategies, multimodal spectroscopic integration, and establish standardized automated workflows to enhance the generalizability and scalability of the models. These advances will facilitate the commercial adoption of Raman-based technologies in vegetable processing industry.
A wide range of vegetable processing technologies have been developed and broadly implemented to ensure food safety, extend shelf life, and improve sensory and nutritional quality. Common vegetable processing methods include freezing, drying, thermal processing (such as canning, blanching, and pasteurization), fermentation, and pickling [1]. Specifically, freezing is utilized for extending shelf-life and helps preserve nutrients and sensory qualities, although it may alter cellular structure and texture, potentially leading to softening and drip loss [2, 3, 4]. Drying serves to reduce water activity and concentrate nutrients for better storage, but this process often results in the degradation of color, heat-sensitive compounds, and flavor loss [5]. Thermal processing, primarily aimed at eliminating microbes and ensuring food safety, may lead to significant vitamin loss, flavor degradation, texture softening, and the formation of Maillard byproducts [6]. Finally, fermentation is employed to enhance functionality and flavor, developing beneficial microbes and increasing nutrition, however, it presents challenges in microbial control and requires standardization [7]. These processing methods alter the physicochemical properties and structure of vegetable tissues, necessitating analytical techniques sensitive to quality-relevant molecular transformations during processing. Moreover, these processing methods support modern product formats and directly enable ready-to-eat (RTE) and ready-to-heat (RTH). Applications include fresh-cut vegetables in modified-atmosphere packaging for RTE salads [8], fermented vegetables as RTE items [9], and the pre-made foods, including self-heating hot pot and frozen dumplings, as RTH options [1]. These applications align processing with contemporary demands for convenience and nutrition, while maintaining safety and quality.
Due to the diverse impacts of various processing techniques on the physicochemical and sensory properties of vegetables, accurate and timely quality assessment is essential to ensure product consistency, optimize processing parameters, and meet consumer expectations. Conventional analytical devices, including texture analyzers, pH meters, conductivity sensors, and colorimeters, often fall short of meeting the growing need for rapid, non-destructive, and automated evaluation in industrial settings. These devices are inherently limited to single-parameter measurements, necessitating multiple sequential analyses and often requiring sample destruction, resulting in high economic costs and reduced environmental sustainability [10, 11]. To overcome these constraints and enable simultaneous assessment of multiple quality attributes, growing interest has emerged in advanced analytical strategies capable of handling high-dimensional, multivariate data, supporting more efficient and reliable quality control in vegetable processing.
Chemometrics, the application of mathematical and statistical methods to chemical data, has emerged as a powerful tool for addressing these challenges. These techniques facilitate the efficient extraction of meaningful patterns from complex datasets, enabling non-destructive, rapid, and high-throughput evaluation of food quality. In vegetable processing, chemometric techniques have been successfully integrated with spectroscopic modalities, including near-infrared (NIR), Fourier-transform infrared (FT-IR) spectroscopy, and Raman spectroscopy, to assess quality attributes and guide processing optimization. Compared to other vibrational techniques, Raman spectroscopy has unique advantages for food processing because of its low water sensitivity and high spatial resolution [12, 13, 14]. These characteristics make Raman spectroscopy particularly well-suited for assessing the quality of structurally diverse food products.
This review highlights the integration of Raman spectroscopy with chemometric techniques for the evaluation and monitoring of vegetable processing, encompassing quality classification, quantitative analysis of key constituents, and process optimization. Despite the increasing interest in this field, research on the application of Raman spectroscopy with chemometric techniques in vegetable processing remains limited. Most studies have focused on a few vegetable types and specific components, such as moisture or carotenoids, with little comparison across different chemometric methods. This review systematically summarizes recent advances in Raman-chemometric integration, highlights representative applications across different stages of vegetable processing, and critically assesses the strengths and limitations of various modeling strategies. Through this assessment, we aim to clarify the current research gaps and provide methodological guidance for advancing real-time, non-destructive quality control systems in vegetable processing industry.
Spectroscopic techniques have become indispensable tools in food quality evaluation due to their non-destructive, molecular specificity, and adaptability for real-time or near-real-time monitoring. Although various studies have demonstrated the applicability of individual spectroscopic methods, systematic comparisons that address the specific analytical challenges posed by processed vegetable matrices remain limited. In vegetable processing, these techniques provide significant advantages for detecting molecular alterations induced by thermal, enzymatic, and physicochemical treatments [15, 16, 17, 18]. This section provides a comparative overview of NIR, FT-IR, UV-Vis, and Raman spectroscopy, outlining their underlying principles, representative applications, advantages, and limitations in the context of processed vegetables. Raman spectroscopy is examined in greater depth to explain its theoretical basis and clarify the context in which it offers practical advantages. A summary of the key characteristics of these spectroscopic techniques is provided in Table 1.
| Techniques | Spectral range | Unit | Primary components | Advantages | Limitations | Applications |
|---|---|---|---|---|---|---|
| Raman | 200–3500 | cm⁻¹ | Polysaccharides, proteins, carotenoids | High molecular specificity; minimal water interference | Weak signal intensity; costly instrumentation | Cell wall structure tracking; thermal softening analysis |
| FT-IR | 400–4000 | cm⁻¹ | Pectin, proteins, lipids | Functional group identification; relatively fast | Moisture interference; strict sample preparation | Degree of methylation; protein denaturation |
| NIR | 780–2500 | nm | Moisture, sugars, total solids | High speed; good for bulk property estimation | Requires chemometric calibration; low specificity | Moisture analysis; sugar content estimation |
| UV-Vis | 200–800 | nm | Polyphenols, pigments, vitamins | Rapid and low-cost; suitable for antioxidant profiling | Low structural information; matrix scattering; underexplored in vegetable systems | Antioxidant capacity; browning assessment |
NIR spectroscopy, operating in the 780–2500 nm range, is a powerful analytical technique that detects overtone and combination vibrations of O–H, C–H, and N–H bonds. This inherent sensitivity to these bonds makes NIR spectroscopy an excellent tool for analyzing key food components such as water, sugars, and proteins. In vegetables, NIR is widely employed for rapid evaluation of moisture content, soluble solids, and texture-related properties such as firmness and water mobility. Despite its advantages, NIR spectroscopy faces significant limitations. As noted by Nicolaï et al. [19], the strong absorption of water in the NIR region produces highly convoluted spectra, complicating the identification of specific chemical constituents. Furthermore, prediction errors can become substantial when moisture content exceeds approximately 50%, limiting its accuracy in very wet samples [20]. Another key challenge lies in the reliance on chemometric models to interpret the complex spectral data. The accuracy and generalizability of these models depend heavily on the quality and representativeness of the calibration datasets [21, 22]. For example, NIR combined with partial least squares regression (PLSR) has also been used to predict moisture and protein in corn kernels across diverse international datasets [23]. PLSR model has also excellent predictive power for adenosine content in porcini mushrooms using NIR spectra [24]. Beyond vegetables, in two plum cultivars, soluble solids content and weight loss were successfully predicted from NIR spectra preprocessed with support vector machine regression (SVM) [25].
2.1.2 Fourier-Transform Infrared spectroscopy (FT-IR)FT-IR spectroscopy is a powerful technique that operates in the mid-infrared region (typically 400–4000 cm⁻¹) to measure fundamental vibrational transitions of molecular bonds. In vegetable processing, FT-IR has been applied to characterize structural and compositional changes, particularly those involving pectin integrity and protein denaturation. Using this method, researchers characterized broccoli tissues and purées by assessing the pectin degree of methylesterification through the 1740 cm⁻¹ (ester C=O) and 1600 cm⁻¹ (COO⁻) bands. Peak deconvolution enhanced spectral resolution and improved analytical accuracy in this protein-rich vegetable [26]. For several other vegetables, including tomato and potato, FT-IR combined with sequential cell-wall extraction and principal component analysis (PCA) has been used to differentiate pectin-rich from hemicellulose-enriched residues and to support species-level discrimination. However, this method may not achieve complete removal of all components, as demonstrated in pumpkin where branched pectin and hemicelluloses remained, complicating accurate mid-IR interpretation [27]. Similar to NIR, FT-IR measurements on fresh or moist vegetable matrices are strongly influenced by water absorption, which can dominate the spectra and obscure weaker analyte bands [28]. Additionally, solid and semi-solid samples often require preprocessing steps like grinding or drying, which limit the real-time applicability of FT-IR [29]. Instrumentation is also relatively expensive and less portable, further reducing its accessibility for on-site monitoring [30, 31].
2.1.3 Ultraviolet–Visible spectroscopy (UV-Vis)UV-Vis spectroscopy is a versatile analytical technique that operates in the 200-800 nm range. It measures the electronic transitions of food chromophores, such as chlorophylls, carotenoids, anthocyanins, polyphenols, and ascorbic acid [32]. The absorbance changes can be directly correlated to compositional shifts and reaction indices. In vegetables, UV-Vis is a robust tool for both nutrient and contaminant assays. For example, it can quantify β-carotene in raw carrots and sweet potatoes with excellent linearity (R2 ≈ 0.99) after acetone-based extraction [33]. It also enables rapid (≈10-15 min) determination of the pesticide flonicamid in cucumber, tomato, and bottle gourd using a bromination–leucocrystal violet reaction. However, these reagent-based, extractive assays are susceptible to matrix interferences and require validation [34]. Nonetheless, such measurements are typically rapid, cost-effective, and amenable to high-throughput workflows, especially when the sample matrix is inherently liquid or can be readily homogenized with minimal preparation. When analyzing intact or heterogeneous tissues, UV-Vis spectroscopy faces more pronounced limitations. It has low structural specificity and is highly sensitive to turbidity, scattering, and pH-dependent shifts. As a result, it is not well-suited for structural assessment and is often coupled with separation techniques, such as HPLC-UV, to improve specificity [35, 36].
2.2 Raman spectroscopy: Fundamentals and key technologies 2.2.1 Basic principles and advantagesRaman spectroscopy is an analytical technique that provides a chemically specific “fingerprint” of a sample by probing inelastic light scattering arising from molecular vibrations [37]. This process is centered around the Stokes and anti-Stokes scattering phenomena, which occur when incident laser photons interact with the sample molecules. While most photons undergo elastic Rayleigh scattering and retain their original energy, a small fraction experiences inelastic scattering, resulting in an energy shift associated with changes in molecular vibrational states. The frequency difference between the Raman-scattered light and the Rayleigh-scattered light is defined as the Raman shift, which is determined by changes in molecular vibrational energy levels [38, 39]. Particularly, a vibration is Raman active only if it results in a change in the molecule's polarizability during the transition. This principle of polarizability-based scattering grants Raman a significant advantage in analyzing aqueous systems, as the weak Raman signal of water allows for the clear detection of components dissolved within the water matrix [40].
Operationally, the usable Raman spectrum covers a fingerprint region (~200–1800 cm⁻¹) and a high-wavenumber region (~2800–3600 cm⁻¹), with a largely silent window near 1800–2400 cm⁻¹ in native tissues [41]. In vegetables, the fingerprint region resolves spectra markers of cell wall and pigments: β-1,4-glucosidic backbone (~1095 cm⁻¹), pectin ester/carboxyl (~1740/~1600 cm⁻¹), and the carotenoid triplet (~1520/1156/1000 cm⁻¹). Protein signatures are typically indexed by Amide I/III (~1650/~1230–1300 cm⁻¹) [42, 43, 44, 45]. The high-wavenumber domain (~2800–3600 cm⁻¹) complements fingerprint-region assignments: protein-related CH bands track thermal denaturation [46], and the OH envelope >3000 cm⁻¹ reflects hydrogen bonding and hydration dynamics, supporting moisture mapping in biological tissues [47, 48].
Raman spectroscopy is well-suited for analyzing aqueous samples because water exhibits only weak Raman scattering, allowing other signals to be seen more clearly [49, 50]. Relative to NIR, Raman spectroscopy achieved comparable accuracy in predicting seasonal strawberry sugar content using approximately half the number of PLSR components, reflecting greater chemometric robustness [51]. Compared with FT-IR, Raman spectroscopy enables easier sampling and true confocal, higher-resolution mapping of in-situ plant tissues including starch granules and cotyledon components [14]. Unlike UV-Vis spectroscopy, which relies on chromophores and primarily reflects signals from phenolics and antioxidants, Raman spectroscopy showed excellent correlation with alcohol content, indicating stronger specificity for non-chromophoric constituents [52]. Raman spectroscopy’s non-destructive nature and minimal sample preparation requirements, coupled with its compatibility with fiber-optic probes, facilitate real-time in-line monitoring in industrial settings [37]. These attributes support label-free, multi-component analysis enabling continuous quality assessment in vegetable processing.
2.2.2 Surface-Enhanced Raman Scattering (SERS)SERS employs plasmonic nanostructures such as Au or Ag colloids, thin films, flexible swabs, or fiber-integrated substrates, creating near-field “hot spots” that amplify scattering cross-sections by factors of 104–109. The enhancement in SERS arises mainly from electromagnetic amplification at the localized surface plasmon resonance (LSPR), with a smaller chemical charge-transfer contribution (Fig. 1). Accordingly, excitation wavelengths are typically selected to match or closely approach the substrate’s LSPR [53, 54]. This level of sensitivity enables trace determinations in aqueous plant matrices, facilitating rapid screening of pesticide residues and oxidation intermediates in vegetables [55, 56]. Besides, electrochemically assisted SERS has quantified eight triazole residues simultaneously in fresh produce, offering enhanced signal intensity and analytical reliability [57]. This technology enables systematic monitoring and studies the pesticide residues in fruits and vegetables. It not only provides an effective and innovative solution for the detection of harmful substances in agricultural products but also brings broad prospects for food safety assurance.

Raman spectral imaging acquires a full Raman spectrum at each spatial pixel and leverages vibrational fingerprints, including endogenous signals and bioorthogonal tags, to reconstruct label-free maps across the sample. Pękala et al. [58] used Raman spectral imaging combined with true component analysis to localize cellulose, pectin, and hemicellulose in apple cell walls and to distinguish acetylated from deacetylated hemicelluloses during storage. In contrast, Imaizumi et al. [59] mapped pectin (852 cm⁻¹) in carrots subjected to various blanching treatments, correlating the spatial patterns with texture measurements and atomic force microscopy observations. A representative schematic Raman chemical map illustrating the spatial distribution of pectin is shown in Fig. 2 to facilitate intuitive understanding of Raman imaging concepts. Once constrained by weak signals, fluorescence interference, low efficiency, and slow processing, Raman spectral imaging has become increasingly practical due to advances in diode lasers, optical fibers, laser-rejection filters, FT and dispersive spectrometers, and modern computing tools that have collectively overcome many of these limitations [60].

Raman spectroscopy offers significant advantages for analyzing processed vegetables, including non-destructive measurement, minimal sample preparation, and high molecular specificity that enables simultaneous monitoring of key plant constituents such as pectin, starch, proteins and carotenoids. However, practical implementation faces several challenges. First, thermal or enzymatic processing can induce baseline drift and peak broadening primarily due to protein denaturation and cell-wall softening. Second, pigments such as chlorophyll in leafy greens and carotenoids in carrots and tomatoes, along with phenolic compounds, produce strong sample-dependent autofluorescence that can overwhelm the intrinsically weak Raman signals [61]. Third, biological variability across cultivars, harvest seasons, and tissue types introduce spectral heterogeneity that complicates quantitative predictions. These challenges necessitate reliable chemometric strategies to transform raw Raman data into fit-for-purpose process metrics [62]. Effective workflows must address fluorescence suppression, baseline correction, and biological variability through targeted preprocessing and multivariate modeling. This chapter systematically reviews chemometric techniques for Raman-based vegetable analysis, organized around preprocessing methods, classification and regression strategies, and validation approaches tailored to processing scenarios. Figure 3 illustrates a series of steps involved in the Raman-chemometric analysis procedure.

Preprocessing strategies for vegetable processing applications must address matrix-specific effects including fluorescence backgrounds from pigments and plant metabolites, noise arising from biological heterogeneity, intensity scatter from moisture variability, peak shifts induced by processing treatments, and band overlap resulting from compositional complexity.
Fluorescence interference is often pronounced in vegetable matrices due to chlorophyll in leafy greens such as spinach and kale, carotenoids in carrots and tomatoes, and phenolic metabolites in onions and cruciferous vegetables [45]. Thermal processing complicates Raman baselines by introducing background interference and altering spectral profiles. This has been evidenced by heating studies on frozen carrots subjected to boiling, steaming and microwave processing, which showed elevated baseline, and peak broadening of carotenoid bands [63]. Similar effects were observed in soybean boiling experiments where ethanol internal-standard normalization improved the repeatability of time-course spectra [64]. Commonly used baseline-correction algorithms include polynomial fitting for smooth backgrounds, asymmetric least squares and its adaptive variant airPLS for broad fluorescence interference, and rolling-circle filtering for curved baselines. For example, Matteini et al. [65] employed rolling-circle filtering to suppress chlorophyll-induced fluorescence in fresh vegetable leaves, improving nitrate prediction accuracy for quality assessment of leafy salads.
Vegetable tissues exhibit biological variability in cell size, water distribution, and compositional gradients, all of which contribute to spectral noise. Surface roughness and rapid acquisition during process monitoring further reduce signal-to-noise ratios. Savitzky-Golay smoothing (SG smoothing), which fits a polynomial over a moving window to average noise while preserving peak shape, is the most widely used noise-reduction technique [62]. Applying SG smoothing with 5–15-point windows and 2nd–3rd order polynomials generally denoises spectra yet maintains peak fidelity, facilitating accurate metrics and enabling real-time feature extraction.
In addition to manually performed preprocessing steps in some studies, such as baseline correction, smoothing, and normalization, certain Raman instruments are equipped with software that automatically executes these tasks. For example, in the study by Horns et al. [66], the software Matlab 2021b (Mathworks) was used to automatically process the imported spectra, with the final data point removed as part of the preprocessing routine. The remaining data points were aggregated into bins, each comprising 10 consecutive values. All signals were automatically normalized by the spectrometer to the highest signal with an intensity of 1. The integration of instrument-level and software-based preprocessing ensures consistency and reduces human-induced variability in the Raman-based workflows.
3.2 Multivariate analysis techniques 3.2.1 Multivariate modeling approachesSpectroscopic analysis of processed vegetables is shaped by recurring analytical problems, and an effective modeling strategy begins with those problems rather than in a predefined inventory of algorithms. The first and most common difficulty is the presence of overlapping bands and weak signals often arising from fluorescence, baseline drift, and thermal or matrix-induced effects. Following baseline correction and intensity normalization, PCA is used to reveal latent structures and to assess the internal consistency of replicate spectra. When the goal is to separate mixed chemical contributions, multivariate curve resolution alternating least squares (MCR-ALS) can decompose spectra into pure component profiles and their associated concentration trends. Once the signal is clarified, PLSR supports quantitative prediction when the relationship between spectra and targets is close to linear. In contrast, support vector machines (SVM) or support vector regression (SVR) provide better discrimination when curvature and class overlap are present.
A second difficulty concerns variability across cultivars, maturity stages, and processing lots, which inflates redundancy and can obscure the analytical signal. PCA provides a first diagnosis of batch structure, and variable selection methods such as competitive adaptive reweighted sampling (CARS) or genetic algorithms reduce redundancy by directing the model toward the most informative spectral regions. With a stable set of variables in hand, PLSR is often preferred when interpretability is a priority because loadings can be linked to specific spectral bands. If the class boundaries remain nonlinear after variable selection, SVM or random forest can improve classification performance at the cost of transparency. These choices should be verified with batch-wise or external lot validation so that apparent gains are not confined to a single data partition.
A third difficulty is the imbalance between high spectral dimensionality and limited sample sizes, which increases the risk of overfitting in predictive models. Dimensionality reduction with PCA and sparse variable selection are practical safeguards against overfitting, while model tuning should rely on nested cross validation or other rigorous resampling strategies. Simple baselines such as MLR or PLSR provide an anchor for honest error estimation. Escalation to SVM or random forest is justified only when diagnostics show systematic nonlinearity that the linear models cannot capture.
Raman imaging introduces a fourth difficulty related to spatially resolved chemistry. In this setting, MCR-ALS yields pure component spectrum and spatial maps corresponding to chemically interpretable features, while a PCA followed by SVM enables rapid pixel wise segmentation for quality control. Performance reporting should integrate accurate measures with evidence of chemical interpretability so that maps are more than statistical partitions. The significance of this challenge is first illustrated by published applications. For example, drying treatments in thermally processed sweet potatoes were first visualized by PCA, after which PLSR predicted carotenoid content with high accuracy as reported by Sebben et al. [67]. In forensic studies of food concealment, PCA separated cocaine related spectral clusters across complex matrices while a correlation-based method in wavenumber space provided similarity scoring as reported by Assi et al. [68]. In plant health monitoring, the combined use of PCA with SVM and MCR-ALS enabled tracking of pesticide penetration and bacterial infection fronts in fruit tissues, achieving high classification accuracy as reported by Luo et al. [25] and by Tian et al. [69].
Overall, these examples illustrate that PCA helps exploratory analysis, data quality checks, and batch-level diagnostics. MCR-ALS resolves spectral mixtures into chemically interpretable components. PLSR provides interpretable regression when linearity assumptions are valid while SVM and random forest strengthen nonlinear discrimination where required. Variable selection reduces redundancy and controls variance. In the remainder of this section, and in Table 2, each method is indexed to the specific problem it addresses allowing readers to navigate from problem to tool rather than from tool to problem.
| Method | Type | Application | Advantages | Limitations | References |
|---|---|---|---|---|---|
| PCA | Unsupervised learning | Exploratory analysis, data visualization | Simple, highly interpretable | Not suitable for prediction | [67] |
| PLSR | Supervised regression | Quantitative determination of constituents | Good model interpretability, tolerance to small datasets | Limited performance with nonlinear relationships | [71] |
| SVM | Supervised classification | Adulterant detection, residue classification | Effective for nonlinear, high-dimensional feature spaces | Model complexity, requires parameter optimization | [68] |
| MCR-ALS | Decomposition modeling | Component separation and localization | Handles overlapping peaks, multi-component systems | Sensitive to initialization and convergence criteria | [69] |
| CNN | Deep learning-based classification | Spectral imaging classification | End-to-end learning eliminates manual preprocessing | High demand for annotated data and computational resources | [75] |
Performance evaluation should mirror the analytical goal and the structure of the data, and model selection should follow from that evaluation. When the target is an interpretable regression based on a limited number of informative bands, linear models such as PLSR and MLR remain strong choices because their parameters can be directly linked to band assignments, as summarized by Gowen et al. [70]. When preliminary exploration indicates curvature or persistent class overlap, nonlinear approaches such as SVM or random forest often deliver higher predictive accuracy. In such cases, these models can capture complex decision boundaries that linear methods may miss, especially when class separation depends on interactions or higher-order effects.
In pesticide residue prediction under nonlinear spectral conditions, SVM achieved higher coefficients of determination and lower prediction errors than PLSR, as reported by Huang et al. [71]. Random forest can capture interactions among variables, although this comes at the cost of reduced model transparency, as discussed by Merusi et al. [72] and by Purcaro et al. [73].
Robust evaluation is essential in high dimensional Raman spectroscopy, particularly when working with limited sample sizes. Variable selection through competitive adaptive reweighted sampling or genetic algorithms, and dimensionality reduction through PCA, help mitigate overfitting [74]. Validation should be nested, at minimum, at the batch level, so that tuning decisions are separated from final error estimation. Reporting should include standard indices for regression such as the coefficient of determination and the root mean square error of prediction, and for classification such as accuracy and the F1 score, together with results on external lots when available.
Deep learning has recently been applied to Raman spectra and imaging and can outperform classical chemometric pipelines classical pipelines when large and well curated datasets are available. An optimized convolutional neural network achieved perfect classification in culture media discrimination, outperforming PCA combined with SVM. The performance gains were linked to the use of batch normalization and carefully designed convolutional blocks, as reported by Wan et al. [75].
In conclusion, a practical rule is to begin with linear models when the relationship between spectra and targets is approximately linear and when band-level interpretability matters. Nonlinear models are warranted when exploration analysis and diagnostic checks reveal systematic structure in the residuals that linear approaches fail to capture. Deep learning is best reserved for data-rich scenarios and should be accompanied by strong validation against independent batches. Throughout, evaluation should be aligned with the intended deployment setting, and Table 2 serves as a guide linking each method to the specific problem it addresses and to the performance metrics most informative for that application.
Raman spectroscopy combined with chemometric modeling has demonstrated significant potential in various aspects of vegetable processing. It offers several advantages, particularly in vegetable processing environments, including non-destructive analysis that preserves sample integrity, real-time monitoring capabilities to support continuous production workflows, and minimal sample preparation requirements that reduce processing time and cost. The ability of this technique to provide molecular information about chemical composition, structural changes, and quality parameters makes it especially suitable for the complex and dynamic nature of vegetable processing operations, where rapid decision-making is essential for maintaining product quality and safety standards.
4.1 Quality assessment and classificationFor classification tasks, PCA-LDA, PLS-DA, and soft independent modeling of class analogy (SIMCA) have been successfully applied to discriminate between red pepper ripening stages and paprika varieties. In a comparative study, both PCA-LDA and PLS-DA achieved high precision in training (95–100%) and test sets (90–100%), while SIMCA yielded the highest prediction accuracy overall (95–100%) for spectral patterns related to carotenoid content across maturity stages [76]. Although PCA-LDA and PLS-DA showed comparable overall performance, PCA-LDA demonstrated greater stability when classifying genetically similar samples, whereas PLS-DA was more susceptible to misclassification due to spectral overlap in such cases [77]. In carotenoid quantification, PLSR models have demonstrated superior performance compared to alternative approaches. Wang et al. [78] developed PLSR models based on characteristic band spectra of carotene in carrots, achieving optimal accuracy (R2 = 0.95, RMSE = 4.66 mg/kg) compared to PCR and LS-SVM methods. Sebben et al. [79] applied PCA and PLSR to evaluate carotenoid degradation in sweet potatoes during thermal processing, achieving high predictive performance (R2 = 0.90 for hot-air and 0.88 for microwave drying) while demonstrating the effectiveness of PCA in visualizing sample clustering. Wen et al. [80] extended this approach to fresh-cut Chinese yam, developing optimized PLSR models for four quality indicators: moisture content, water activity, polysaccharide content, and microbial load.
4.2 Real-time processing monitoringProcess monitoring with Raman spectroscopy has evolved from isolated demonstrations to problem driven workflows that target specific transformations in vegetable matrices. Studies on plant-based materials and food systems have shown how Raman signals capture both chemical transformations and structural reorganizations occurring during processing. Corn husks are not conventional vegetables, yet they offer a useful plant matrix for method validation. Raman signals successfully captured the transition from cellulose I to cellulose II, confirming the feasibility of process monitoring in plant-based matrices [81]. In liquid food systems, soymilk production confirmed that Raman spectroscopy combined with chemometric models can monitor protein content and secondary structure dynamics during boiling. PLSR achieved an average R2 of 0.96 and a residual predictive deviation of 3.56, which supported accurate estimates of soluble protein and the detection of thermal denaturation through alpha helix unfolding [82]. Earlier work by Lu et al. [83] demonstrated that Raman spectroscopy combined with PLSR and SIMCA, can track intracellular uptake and tissue disruption in garlic exposed to antibacterial agents, highlighting method’s sensitivity to subtle physiological changes during processing.
During fruit ripening, Raman spectroscopy enables non-destructive tracking of starch degradation and sugar accumulation and supports calibration models that predict ripeness markers to inform harvest decisions. In Norwegian Aroma and Elstar apples, Raman spectra captured starch degradation and sugar accumulation, while calibration models predicted ripeness markers with practical accuracy. These findings support the use of non-destructive starch indexing and harvest grading [84]. Additional studies in pears demonstrated the detection of pathogen-induced cell wall degradation without the need for external labeling. These cellular-level observations complement bulk compositional tracking and together they illustrate how Raman methods provide real time process insight across structural scales [45].
Process Raman instrumentation is commercially available and is already used for real time monitoring and control in manufacturing environments. Current platforms integrate stable lasers, robust spectrometers, and embedded chemometric models. Some systems also feature probes that tolerate high temperature and high pressure enabling reliable inline operation. The same capabilities can support food and agricultural processing, where variable temperature, flow, and turbidity require stable optics and robust calibration protocols. With calibration transfer, external batch validation, and routine maintenance in place, Raman spectroscopy provides a practical basis for process analysis and supports automated interventions such as endpoint detection and hold time adjustment in processed vegetables.
4.3 Quantitative analysis of key constituentsThe integration of Raman spectroscopy with chemometric models enables the effective quantification of a wide array of chemical constituents in vegetables, including water, sugars, carotenoids, nitrates, and other bioactive compounds. In leafy vegetables such as spinach, Raman spectroscopy combined with PLSR and MLR models have successfully quantified nitrate levels, yielding robust predictive performance (R2 > 0.80) [65]. Carrots have been used as model vegetables in numerous studies because of their diverse phytochemical composition. FT-Raman spectroscopy has been employed to simultaneously quantify α- and β-carotene, glucose, fructose, sucrose, and polyacetylenes, achieving R2 values exceeding 0.91. In addition, PCA of the spectral data effectively differentiated carrot root metabolic profiles by color and botanical classification, confirming the method's capability for high-resolution, multi-constituent profiling [85]. Raman-based calibration models have also enabled the non-destructive ranking of carotenoid content across 332 orange carrot cultivars with high accuracy (R2 = 0.86, RMSECV = 20.5 ppm) [86]. Moreover, the potential of Raman spectroscopy for simultaneous classification and quantification was demonstrated by Moe Htet et al. [87], who applied PLS-DA and SIMCA to accurately discriminate multiple types of vegetable oils. They also developed a PLSR model to quantify α-tocopherol with high correlation to HPLC measurements (R2 > 0.95). This dual-function approach highlights the utility of Raman spectroscopy for simultaneous classification and quantification in the multi-target analysis of complex food matrices.
Collectively these studies highlight the efficacy of Raman spectroscopy in facilitating high-throughput, non-destructive quantification of various chemical constituents in vegetable processing. This capability supports the application of Raman spectroscopy in both quality assurance and product characterization processes.
Raman spectroscopy, compared with other vibrational techniques such as NIR, FT-IR, and UV-Vis spectroscopy, demonstrates distinct advantages in the context of vegetable processing. Its high molecular specificity, resistance to water interference, and minimal sample preparation make it particularly suitable for complex, moisture-rich vegetable matrices. Unlike NIR and FT-IR, Raman spectroscopy can resolve subtle spectral variations in key constituents, such as carotenoids, nitrates, and sulfur compounds, making it a powerful tool for compositional analysis and quality control. In this review, the incorporation of chemometric techniques, such as PCA, PLSR, SIMCA, SVM, and CNN, which are based on deep learning, has been demonstrated to markedly enhance the interpretability and predictive capabilities of Raman spectral analysis. These models facilitate various applications: PCA facilitates data exploration and sample clustering, PLSR enables precise regression for compositional quantification, SIMCA and SVM are particularly effective for classification tasks, and CNNs perform automated feature extraction for high-dimensional data, especially in spectral image analysis.
The practical applications of Raman–chemometric workflows in vegetable processing have been demonstrated across a range of tasks, including quality classification (determining pepper ripening stages), quantification of key components (carotenoids and nitrates), and real-time monitoring of chemical and structural changes, thermal degradation, and protein denaturation. These studies not only affirm the value of technique in laboratory settings but also suggest its increasing potential for deployment in field and industrial systems.
Despite these advancements, several challenges persist. First, spectral overlap and complex matrix interference continue to impede signal clarity, particularly in multicomponent vegetable samples. Second, the generalizability of existing models is constrained by small and narrowly defined training datasets, which are often limited to single cultivars, specific growth conditions, or processing methods. Third, the implementation of in-line, real-time industrial applications necessitates robust and standardized pipelines, which are frequently absent in current studies. To address these gaps, future research should prioritize the development of transfer learning strategies to enhance model adaptability across diverse vegetable types and conditions. The integration of multimodal spectroscopic techniques could further enhance prediction accuracy by capturing complementary molecular information. Examples include the combination of Raman spectroscopy with NIR or hyperspectral imaging. Additionally, the design of standardized, automated Raman–chemometric platforms is crucial for enabling real-time, scalable quality control in contemporary vegetable-processing environments. In conclusion, this review systematically highlights the advantages and practical applications of Raman spectroscopy integrated with chemometric methods in vegetable processing. By leveraging the strengths of diverse modeling approaches and addressing current limitations through advanced algorithmic and instrumental innovations, Raman-based systems are well positioned to play a central role in achieving precise, sustainable, and intelligent food production.