The Journal of Toxicological Sciences
Online ISSN : 1880-3989
Print ISSN : 0388-1350
ISSN-L : 0388-1350
Original Article
Development of a decision flow-based read-across approach for assessing skin sensitization using publicly available tools
Yuri HatakeyamaKosuke ImaiHayato NishidaShiho OedaTomomi AtobeMorihiko Hirota
Author information
JOURNAL FREE ACCESS FULL-TEXT HTML
Supplementary material

2026 Volume 51 Issue 9 Pages 501-515

Details
Abstract

Animal tests like the guinea pig maximization test (GPMT) and murine local lymph node assay (LLNA) were historically used to assess skin sensitization in cosmetics. The EU banned animal testing for cosmetics in 2013, prompting the development of alternative in vitro methods and new approach methodologies (NAMs) based on adverse outcome pathways (AOPs). These NAMs are integrated into Defined Approaches (DAs) to evaluate skin sensitization hazards and potency, which challenge is incorporating them into next-generation risk assessment (NGRA) frameworks. To address this challenge, we aim to achieve quantitative risk assessment in humans. The NGRA framework proposed by Cosmetics Europe emphasizes read-across methods, which predict toxicity using data from similar substances. This approach, which utilizes data from analogues, is described as an effective way to reduce uncertainty in risk assessment. This paper proposes a transparent read-across method using the OECD QSAR Toolbox, ranked by reliability based on structural similarity, protein binding, sensitization alerts, and percutaneous absorption. The sensitization data of the extracted analogues prioritize the LLNA EC3 values, followed by predicted EC3 values from Derek nexus or GPMT data. A case study using p-isobutyl-α-methyl hydrocinnamaldehyde confirmed the target's classification as a moderate sensitizer. The proposed method ensures transparency by relying on publicly available tools, reducing uncertainty in NGRA.

INTRODUCTION

Animal tests for assessing sensitization potential, such as the guinea pig maximization test (GPMT; Magnusson and Kligman, 1969; OECD Test Guideline 406, OECD, 1992) and the murine local lymph node assay (LLNA; OECD Test Guideline 429, OECD, 2010), were historically used to evaluate the safety of cosmetics. The LLNA is based on measuring cell proliferation in draining lymph nodes following repeated topical administration of the test substance to the ear during the induction phase of skin sensitization (Basketter et al., 2002). This assay can distinguish between non-sensitizers and weak, moderate, strong and extremely strong (extreme) sensitizers. Predictions of safe levels of human exposure using the quantitative risk assessment approach (QRA), based on the positive threshold in the LLNA (EC3 value) have been widely adopted by the International Fragrance Association (IFRA) for cosmetics and fragrances (Api et al., 2008, 2020a). As a result, many chemicals (>300) have been evaluated using the LLNA (Gerberick et al., 2005; Kern et al., 2010). These extensive databases serve as a valuable resource for developing and validating alternative in vitro and in silico methods to predict sensitization potential.

However, these animal tests have been prohibited in the European Union (EU) since the EU directive for the replacement of animal tests in the safety testing of cosmetic ingredients came into effect in 2013. Consequently, various alternative in vitro tests have been developed, and the safety assessment of cosmetics has shifted towards the application of new approach methodologies (NAMs) that address the mechanistic key events of adverse outcome pathways (AOPs) related to skin sensitization (OECD, 2024a, 2024b, 2025a). Based on the consensus that multiple NAMs are needed to replace data from animal models, NAMs have been integrated into Defined Approaches (DAs), which derive information on the hazard and potency of skin sensitization (OECD, 2025b). The current challenge lies in integrating these NAMs and DAs into a next-generation risk assessment (NGRA) framework to derive an evidence-based point of departure (PoD) for human risk assessment. We have previously developed Artificial Neural Network (ANN) models to predict LLNA EC3 values using characteristic data from multiple in vitro studies (Hirota et al., 2015, 2018; Hatakeyama et al., 2025; Imai et al., 2025). Case studies using ANN prediction models for skin sensitization have been reported in OECD guidance documents and reports from the U.S. Environmental Protection Agency (U.S. EPA) (U.S. EPA, 2020; Strickland et al., 2022; OECD, 2023). Additionally, SARA-ICE, which predicts PoD values, was recently included in GL497 (Reinke et al., 2024; OECD, 2025b). The hierarchical NGRA framework proposed by Cosmetics Europe provides structured guidance for conducting NGRA for skin sensitization (Gilmour et al., 2020; Gilmour et al., 2023, Fig. 1). The ultimate goal of NGRA is to quantitatively evaluate the safety of chemical substances (such as cosmetic ingredients) in humans based on scientific evidence, without conducting any animal testing. In this framework, read-across based on appropriate analogues is considered a key element of the first stage. Read-across predictions can be combined with limited or inconsistent component-specific NAM data to derive and integrate PoD values. Additionally, confirming the results of skin sensitization NAMs and DAs can enhance the reliability of the NGRA approach (Gautier et al., 2020, 2023; Assaf Vandecasteele et al., 2021). An advantage of the read-across approach is its ability to overcome the technical limitations of NAMs (e.g., cytotoxicity, solubility, applicability) and derive more appropriate quantitative PoD values for risk assessment. The read-across approach aims to predict the hazard and potency of chemical toxicity endpoints by utilizing NAMs, animals, or human data from substances with similar properties. Several articles provide guidance on identifying analogues and performing dependable read-across (Patlewicz et al., 2019; Rovida et al., 2021; Alexander-White et al., 2022), but objectively identifying analogues continues to be difficult. The definition of similarity is intricate because it relies on various factors like chemical structural characteristics, physicochemical properties, protein reactivity, mode of action, or similar metabolic profiles (Schultz et al., 2009; Patlewicz et al., 2013). Moreover, the importance of the parameters to be taken into account might require adjustment based on the specific endpoint being assessed (Cronin et al., 2022; Moustakas et al., 2022). Therefore, the selection of analogues has so far primarily relied on expert judgment. In the context of NGRA, read-across requires a clear explanation of the method used to identify analogues and an open rationale for their ultimate choice. There have already been several reports on read-across in skin sensitization (Gautier et al., 2020; Gautier et al., 2023; Suzuki et al., 2025). This paper proposes a read-across approach for skin sensitization using entirely publicly available tools. This ensures high transparency, allowing anyone to trace the process and enabling objective evaluation. Since this read-across method can derive the PoD, it could potentially be integrated not only into Tier 0 (Identify use scenario) but also into Tier 2 (Risk assessment) of the NGRA framework. Furthermore, consistent documentation regarding the selection of analogues and read-across predictions will facilitate their use in NGRA and ultimately support their acceptance for regulatory purposes.

Fig. 1

Next generation risk assessment framework for skin sensitization risk assessment. This figure was created based on Fig. 1 presented in Gilmour et al. (2020).

MATERIALS AND METHODS

Substance selection

The substances with both in vivo and in vitro data were selected based on the studies by Sebastian Hoffmann et al. (2018) and Hirota et al. (2018). Given that our objective is to evaluate cosmetic ingredients, we specifically selected 12 compounds considered relevant to cosmetics from these studies for further assessment.

Procedure for direct peptide reactivity assay (DPRA)

DPRA data were kindly provided by Procter & Gamble Company (Strombeek-Bever, Belgium), the developer of DPRA as described previously (Hirota et al., 2018). The assays had been performed as described elsewhere (Nukada et al., 2013; Gerberick et al., 2007). The test chemicals were briefly incubated with peptides containing lysine (Ac-RFAAKAA-COOH) or cysteine (Ac-RFAACAACOOH) in the dark for 24 hr at 25°C. Then, the peptides were quantified by reverse-phase HPLC with UV detection. The percentages of peptides depletion and remaining peptides were calculated for each sample by comparing the amount of non-reacted peptides in the reaction sample to the amount of peptides in the peptide-alone control samples. The mean percentage depletion value was calculated as the average of peptide depletion for cysteine and lysine (percentage of cysteine depletion and percentage of lysine depletion, respectively), and the classification of chemicals as sensitizer or non-sensitizer was based on the OECD guidelines (OECD, 2025a). In the following analysis, when the value for the percentage of remaining cysteine or lysine peptides was negative, zero (0) or unity (1), a value of 0.01 was used as a substitute.

Procedure for KeratinoSens™ assay

The KeratinoSens™ data were provided by Givaudan Schweiz AG, the developer of KeratinoSens™ as described previously (Hirota et al., 2018). The assays had been performed as described elsewhere (Emter et al., 2010; OECD, 2024a). Briefly, cells were grown for 24 hr in 96‐well plates, and the medium was then replaced with medium containing the test chemical. Each compound was tested at 12 binary dilutions in the range from 0.98 to 2000 μM. Cells were incubated for 48 hr with the test chemical, and then the luciferase activity and cytotoxicity (MTT assay; Mossmann, 1983) were determined. The gene induction was compared with that of the dimethyl sulfoxide controls, and the wells with statistically significant induction over the threshold of 1.5 (i.e., 50% enhanced gene activity) were determined. The EC1.5 and EC3 values (concentration in μM for induction above these thresholds) were determined by linear extrapolation as described in the OECD guidelines (OECD, 2024a). In the prediction model, chemicals with significant gene induction (>1.5‐fold) at a concentration below 1000 μM and at a concentration at which the cells retain 70% viability in at least two of three repetitions were rated positive.

Procedure for human cell line activation test (h-CLAT)

h‐CLAT data were taken from previous reports (Nukada et al., 2012; Takenouchi et al., 2015; Jaworska et al., 2015). The assays had been performed as described elsewhere (OECD, 2024b). Briefly, THP‐1 cells were plated at 1 × 106 cells/mL in a 24‐well plate and treated for 24 hr with the test chemicals. The test dose was determined based on the test concentration providing a cell viability of 75% (CV75) in the cytotoxicity test. Cells were washed and FcR was blocked. Staining was done with fluorescein isothiocyanate (FITC)‐conjugated antihuman CD86 antibody (clone Fun‐1; BD Bioscience, San Diego, CA, USA) or FITC‐conjugated anti‐human CD54 antibody (clone 6.5B5; DAKO, Glostrup, Denmark) at 4°C for 30 min. FITC‐labeled mouse IgG1 (clone DAK‐G01; DAKO) was used as an isotype control. The cells were washed, and the expression of cell surface antigens was analyzed by flow cytometry. The relative fluorescence intensity, calculated according to Equation 1, was used as an indicator of CD86 and CD54 expression.

Relative fluorescence intensity (%) = (MFI of chemical-treated cells – MFI of chemical-treated isotype control cells) / (MFI of vehicle control cells – MFI of vehicle isotype control cells) × 100 (1)

The thresholds for CD86 and CD54 expression (EC150 for CD86 and EC200 for CD54) were determined as described previously (Nukada et al., 2012). EC150 and EC200 were converted from μg/mL to molar unit (μM). In the following analysis, when EC150 or EC200 could not be calculated because of negative results, a value of 100,000 (μM) was used as a substitute value. The minimum induction threshold (MIT) was determined as the smaller of EC150 and EC200.

Toxtree

Toxtree version 2.5.0, a quantitative structure–activity relationship (QSAR) tool for mechanism‐based prediction of skin sensitization potential, was downloaded from the website of the Joint Research Center of EURL‐ ECVAM (http://toxtree.sourceforge.net). Toxtree determines whether defined skin sensitization alerts are present in the test chemical structure (Enoch et al., 2008). In the following analysis, the results of Toxtree were converted into descriptors for incorporation into the ANN analysis as follows: No alert: 1, SN2: 2, Acyl transfer agent: 3, Michael acceptor: 4, Schiff base formation: 5, SNAr: 6, More than two alerts: 7.

OECD QSAR Toolbox

To obtain structurally similar compounds to the target compound, the OECD QSAR Toolbox will be used. The OECD QSAR Toolbox is a system designed to support assessments based on a category approach and is freely available for download from the OECD website. Additionally, it contains various toxicological test data provided by different countries, making it useful for referencing sensitization data for both the target compound and similar compounds. In this study, version 4.5 of the OECD QSAR Toolbox was utilized.

In vitro prediction using the Artificial Neural Network (ANN) model

In this study, we used an ANN model to predict EC3 values (Hatakeyama et al., 2025). This ANN model incorporates the remaining rates of cysteine and lysine (%) obtained from DPRA, the EC1.5 (μM) from KeratinoSensTM, the MIT (μM) and CV75 (μM) test results from h-CLAT, as well as alerts obtained from Toxtree as descriptors to derive LLNA EC3 (μM/cm2).

In silico prediction by Derek Nexus

For in silico prediction, we used Derek nexus, an in silico qualitative toxicity prediction system developed by Lhasa Limited (2022). Unlike quantitative toxicity prediction methods that use statistical analysis, Derek Nexus is a knowledge-based system that defines empirical rules for substructural toxicity correlations derived from known information, such as public databases (Macmillan et al., 2023). Derek nexus is utilized in the ITS defined approach in OECD GL497 (OECD, 2025b). In this study, version 6.2.0 of Derek Nexus was used, with Derek KB 2022 1.0 serving as the knowledge base.

Quantitative risk assessment (QRA) process

The QRA process has been widely used by IFRA for testing cosmetic and fragrance ingredients (Api et al., 2008; Basketter et al., 2008; Api et al., 2020a). The QRA method consists of four key stages. The first stage is derivation of the no expected sensitization induction level (NESIL). The NESIL is generally calculated from the LLNA EC3 value but must be converted to μg/cm2 because the unit for EC3 values is expressed as a percentage. Given that the average ear area of a mouse is 1 cm2 and the applied volume in the LLNA is 25 μL, if the EC3 value is 1%, the NESIL is calculated to be 250 μg/cm2. The second stage involves the application of the sensitization assessment factor (SAF) to account for uncertainties in determining the NESIL. SAFs are derived on the product type and include considerations such as inter-individual variability, product composition, frequency and duration of use, and the condition of the skin (related to the skin site(s) where the product will be applied). For example, the SAF for category 5 (products applied to the face and body using the hands (palms), primarily leave-on) is set to 100. The third stage is determination of the acceptable exposure level (AEL). The AEL is calculated from the NESIL and the SAF using the formula: AEL = NESIL/SAF. The final stage is estimation of the upper concentration level by comparing the AEL with the consumer exposure level (CEL). For Category 5, the CEL is set at 3.02 mg/cm2/day. In other words, the upper concentration level can be calculated using the following formula:

RESULTS

Method for extracting analogues

To extract analogues, the OECD QSAR Toolbox v.4.5 was used, with a criterion set at 80% structural similarity using the Dice coefficient. There are several methods for calculating structural similarity, but there are two main reasons for using the Dice method. The first reason is that it is the default setting in the OECD QSAR Toolbox v.4.5 and is considered a widely used method. The second reason is that the Dice method emphasizes the importance of common sub-structures within the structure, making it suitable for evaluating skin sensitization, where substructures are key. Regarding the cutoff value for structural similarity, a value of 70% is often used in general scientific literature and industry practices, and it is also mentioned in the SCCS notes of guidance (SCCS, 2023). However, a 70% cutoff results in too many compounds being identified, increasing low-similarity noise even if the purpose is to investigate concerns. The difference in the number of analogues extracted at 70% and 80% structural similarity cutoff values is shown in Table 1. Based on the discussions in subsequent case studies and sorting that takes into account parameters other than structural similarity, it was considered that a structural similarity cutoff of 80% is sufficient. However, if the number of compounds with 80% structural similarity is extremely low it may be necessary to lower the cutoff value to 70% for extraction.

Table 1. The difference in the number of analogues extracted at 70% and 80% structural similarity cutoff values.


Sorting of analogues based on reliability rank

In the section on the method for extracting analogues, analogues were extracted based on an 80% structural similarity rate. However, the similarity between the target and the analogue in the context of read-across for skin sensitization is not determined solely by structure. Structural similarity is thought to affect the structure of antigen peptides derived from hapten-protein complexes, which are presented by dendritic cells to T cells during the sensitization induction phase. Furthermore, even if a similar substructure is present, the type of reaction with proteins and the amount of hapten affected by skin absorption are considered to be related to antigen presentation in the induction of sensitization. Therefore, we incorporated not only structural similarity but also protein binding, sensitization alerts for substructures, and percutaneous absorption as factors in determining similarity. We then defined the degree of similarity between the extracted analogues and the target as a “Reliability Rank” and ranked them accordingly. Specifically, protein binding classification is checked using the OECD QSAR Toolbox v.4.5 to see if the target and similar compounds have the same classification. Sensitization alerts are checked using Derek nexus to see if the target and similar compounds have the same alert number. Percutaneous absorption is assessed to determine if it meets the low percutaneous absorption criteria described in the SCCS notes of guidance and to see if the target and similar compounds have the same percutaneous absorption classification. The reliability rank is determined by how many of the four indicators (structural similarity, protein binding classification, percutaneous absorption, and sensitization alert) match the target. If all four indicators are met, the reliability rank is 1; if three of the four are met, the reliability rank is 2; if two of the four are met, the reliability rank is 3; if one or none of the four are met, the reliability rank is set to 4. The distribution of the number of analogues relative to the target based on this ranking is shown in Fig. 2.

Fig. 2

The distribution of the number of analogues relative to the target based on this ranking.

For each target, analogues were ranked from rank 1 to rank 4, with a tendency for more analogues to be in rank 2 and rank 3. Additionally, using Linalool as an example, four analogues from rank 2 and four analogues from rank 4 were compared (Table 2). Analogues 1 to 4 in rank 2 and analogues 5 to 8 in rank 4 are approximately equivalent in terms of structural similarity, with a similarity rate of around 80%. However, regarding the extracted sensitization data, the analogues in rank 2 were closer to the target, whereas some analogues in rank 4 were classified as non-sensitizing compounds, which differed from the sensitization data of the target. Based on these findings, it is suggested that when extracting analogues, ranking based not only on structural similarity but also on substructure considerations is important.

Table 2. Representative examples of Linalool analogues (4 compounds with reliability rank 2 and 4 compounds with reliability rank 4).


Sensitization data used for the extracted analogues

For the analogues extracted using the OECD QSAR Toolbox, the available sensitization data include GPMT and LLNA. The type of sensitization data available for the analogue with the highest reliability rank for each target is shown in Table 3.

Table 3. The type of sensitization data for the analogue with the highest reliability rank for each target.


In many cases, both LLNA and GPMT data were available. When LLNA data is present, it is considered possible to adopt the LLNA EC3 value of the analogue as the result of the read-across. On the other hand, when GPMT data is available, it is important to note that GPMT can evaluate sensitization potential as Non-Sensitizer, Weak, Moderate, Strong, or Extreme. If only GPMT data is available, we decided to apply a general margin of 10 to the prediction results from the ANN model we previously reported, but only in cases where the GPMT data indicated strong sensitization potential (Hatakeyama et al., 2025; Imai et al., 2025).

However, for some compounds, there were cases where the extracted analogues did not have any sensitization data. Since it is quite possible that extracted analogues may lack sensitization data, we decided to also incorporate predictions from Derek Nexus to ensure a sufficiently predictive system, even when actual in vivo sensitization data is not available. Derek Nexus is a knowledge-based toxicity prediction system that can extract up to 10 structurally similar analogue compounds with the same sensitization alerts as the target compound and predict the LLNA EC3 value based on their weighted average (Macmillan et al., 2023). It is also a software tool used in OECD GL497.

Based on the above considerations, we decided to adopt the measured LLNA data, measured GPMT data, and predicted LLNA data from Derek Nexus as the sensitization data for the extracted analogues.

Criteria for the reliability rank of adopted analogues

As shown in Table 2, even if the structural similarity rate is close, the sensitization risk cannot be considered equivalent if the reliability rank is low. Therefore, we established criteria for adopting analogues. The results of the read-across for 12 different targets are shown in Table 4. When comparing the most similar analogue with an analogue ranked one level lower, there was little difference in the LLNA and GPMT data. However, differences were observed in the predicted EC3 values from Derek Nexus. The Derek-predicted EC3 value for the most similar analogue generally aligned with the target’s published EC3, whereas the Derek-predicted EC3 value for the analogue ranked one level lower showed discrepancies with the target’s published EC3. Specifically, for resorcinol, the published EC3 was 5.5% (moderate), while the Derek-predicted EC3 value for the most similar analogue was 5.8% (moderate). In contrast, the Derek-predicted EC3 value for the analogue ranked one level lower was 0.02% (extreme). The reliability rank for the most similar analogue was 2, while the analogue ranked one level lower had a reliability rank of 3.

Table 4. The results of the read-across for 12 targets.


Similarly, for hydroxycitronellal, the published EC3 was 33% (weak), while the Derek-predicted EC3 value for the most similar analogue was 17% (weak). In contrast, the Derek-predicted EC3 value for the analogue ranked one level lower was 0.086% (extreme). The reliability rank for the most similar analogue was 1, and the analogue ranked one level lower had a reliability rank of 2, which is not particularly low. However, the key difference was that the sensitization alerts did not match.

Skin sensitization is considered to affect the structure of antigen peptides derived from hapten-protein complexes presented by dendritic cells to T cells during the sensitization induction phase, making the structural similarity of substructures that trigger sensitization an important factor. Therefore, matching sensitization alerts is considered to improve the accuracy of predictions. The compounds identified as the most similar analogues had a reliability rank of 1 or 2, and most of their sensitization alerts matched those of the target. On the other hand, when examining compounds ranked one level lower, their reliability rank was 3, and many of their sensitization alerts did not match those of the target.

Based on the above, the analogues to be adopted were defined as those with a reliability rank of 2 or higher and with matching sensitization alerts.

Prioritization of sensitization data to be adopted

Since multiple sensitization data may exist for the extracted analogues, we considered the prioritization of the three types of sensitization data: LLNA, GPMT, and Derek predictions. For each target, Table 5 shows the information for the analogue that exhibits the strongest sensitization data among those with matching sensitization alerts and the highest reliability rank.

Table 5. Sensitization information for the target and analogue.


First, when comparing the measured EC3 values from LLNA for the analogue with the predicted EC3 values from Derek for the analogue, there was little difference in their closeness to the EC3 value of the target substance. Considering that Derek’s EC3 is a predicted value, we decided to use LLNA data for the analogue as the primary reference (Macmillan et al., 2023). Furthermore, due to the nature of GPMT as a hazard assessment method, there were cases where the predicted categories for the target and the analogue differed. For example, resorcinol is originally a moderate sensitizer, but the analogue’s GPMT data classified it as a non-sensitizer. Similarly, hydroquinone is originally classified as a strong sensitizer, but the analogue’s GPMT data classified it as a moderate sensitizer. Additionally, while LLNA evaluates the sensitization induction phase, GPMT evaluates the skin reaction during the sensitization elicitation phase, highlighting a methodological difference. Considering the above, we prioritized the data to be adopted as follows: first, LLNA; second, Derek predictions; and third, GPMT. The read-across results for all 12 compounds used in the validation are provided in the Supplemental Material 1-12.

Case study

The commonly used fragrance compound p-isobutyl-α-methyl hydrocinnamaldehyde (CAS No. 6658-48-6) was selected for the case study. This compound is listed in the annex of OECD GL497 with LLNA and in vitro skin sensitization test data. Additionally, it was evaluated by RIFM in 2020 (Api et al., 2020b), making it suitable choice for a case study compound. Using the OECD QSAR Toolbox, 291 analogues with a structural similarity rate of 80% or higher were identified. Among them, 52 analogues had a reliability rank of 1 or 2, 34 analogues had matching sensitization alerts, and 8 analogues had measured LLNA data (Fig. 3). These 8 analogues are listed in Table 6. While the measured EC3 value of the target compound is 9.5% (moderate), the LLNA EC3 values of the analogues ranged from a minimum of 1.0% (moderate) to a maximum of 37% (weak), with no analogues showing significant deviation. Adopting a more conservative value did not result in underestimation, and the classification remained within the same category of moderate. Furthermore, converting the predicted EC3 value of 1% into an acceptable concentration in face cream using the QRA process yielded 0.08%. Table 7 shows the prediction results combining in vitro and in silico predictions that we have previously proposed. In addition, the detailed read-across results for p-isobutyl-α-methyl hydrocinnamaldehyde are provided in the Supplemental Material 13.

Fig. 3

Databases and criteria used to identify suitable analogues for p-isobutyl-α-methyl hydrocinnamaldehyde.

Table 6. Read-across results for p-isobutyl-α-methyl hydrocinnamaldehyde.


Table 7. Skin sensitization evaluation results of p-isobutyl-α-methyl hydrocinnamaldehyde.


DISCUSSION

In this study, we established a read-across method using publicly available tools. The overall framework is shown in Fig. 4. The read-across method proposed in this report is innovative in that it ranks similarity not only based on structural similarity but also by considering protein binding, the presence of substructures indicating sensitization, and dermal absorption. Since skin sensitization is considered to affect the structure of antigen peptides derived from hapten-protein complexes presented by dendritic cells to T cells, the structural similarity of substructures that trigger sensitization is an important factor. It is well known that individuals sensitized to p-phenylenediamine often experience cross-reactions with similar oxidative hair dyes. Thus, due to the nature of toxicity, substructural similarity is important for skin sensitization. A scientific limitation of the LLNA is that it focuses on the induction phase in the skin sensitization AOP and cannot verify the elicitation phase. However, by utilizing the read-across method proposed in this paper, we can reference the sensitization data of analogues, which provides an advantage in covering the elicitation phase, including cross-reactions. The read-across method proposed here uses protein binding classifications from the QSAR Toolbox and sensitization alerts from Derek Nexus to sort extracted similar compounds, applying weighting that considers substructural similarity, which suggests a more accurate evaluation. Among the four parameters used to determine the reliability rank in this read-across method, matching sensitization alerts was prioritized as the most critical. As described in the section on the criteria for the reliability rank of adopted analogues, there are two reasons for this. First, since the structure of antigen peptides derived from hapten-protein complexes presented by dendritic cells to T cells is considered to be affected, the structural similarity of the substructure that triggers sensitization was deemed a crucial factor. Second, as shown in Table 4, the sensitization data of analogous compounds with non-matching sensitization alerts showed significant discrepancies compared to those of the target.

Fig. 4

Overview of the read-across approach using publicly available tools.

In the case study section, we conducted a read-across using p-isobutyl-α-methyl hydrocinnamaldehyde as an example. The measured EC3 value of p-isobutyl-α-methyl hydrocinnamaldehyde is 9.5%, while the lowest measured EC3 value among the extracted similar compounds was 1.0%, resulting in the same classification of moderate. p-isobutyl-α-methyl hydrocinnamaldehyde was evaluated by RIFM in 2020 (Api et al., 2020b), and in this literature, the maximum acceptable concentration in face moisturizer products was calculated to be 0.25%. Meanwhile, the lowest predicted EC3 value from our previously proposed in vitro and in silico tests, as well as the read-across method proposed here, was 1.0%. Based on RIFM's QRA conversion, this translates to a permissible concentration of up to 0.08% in face creams, which does not significantly deviate from the case study results by RIFM and does not result in underestimation.

In addition to its accuracy, the read-across method proposed in this report is commendable for not using any internal data and relying entirely on publicly available tools. This ensures traceability for anyone, thereby providing transparency and enabling an appropriate evaluation free from personal bias. Furthermore, it would contribute to reducing uncertainty in skin sensitization NGRA by combining it with other PoD prediction models (e.g. SARA-ICE, ANN model).

In summary, this report confirms that an appropriate read-across process leads to the identification of suitable analogues with high-quality skin sensitization data, enabling the derivation of the Point of Departure (PoD) for skin sensitization with high reliability. Although this study focused exclusively on cosmetic ingredients because our objective was the toxicity evaluation of cosmetic materials, careful consideration, including the selection of compounds for verification, is required when exploring applications to different fields, such as pharmaceuticals, in the future. Moving forward, we aim to conduct more case studies to contribute to skin sensitization NGRA.

Funding

This research was conducted using research cost of SHISEIDO CO., LTD.

Conflict of interest

All authors are employees of SHISEIDO CO., LTD.

Data availability

The data in this study are included in the article/supplementary materials. Contact the corresponding author(s) directly to request the underlying data.

Author contributions

Yuri Hatakeyama: Conceptualization, Methodology, Data curation, Investigation, Writing – original draft.

Kosuke Imai: Investigation.

Hayato Nishida: Data curation, Conceptualization.

Shiho Oeda: Investigation.

Tomomi Atobe: Data curation, Conceptualization.

Morihiko Hirota: Conceptualization, Supervision, Writing-review and editing.

Ethical approval and consent to participate

Not applicable.

Patient consent for publication

Not applicable.

REFERENCES
 
2026 Author(s)

This article is licensed under a Creative Commons [Attribution 4.0 International] license.
https://creativecommons.org/licenses/by/4.0/
feedback
Top