産業保健法学会誌 英文誌
Online ISSN : 2758-4755
Print ISSN : 2758-4771
Position Paper
The Governance Gap in Psychosocial Risk Evaluation
Yannick A. Metzler Svenja MüllerLuis TorresAditya Jain
著者情報
ジャーナル オープンアクセス HTML

2026 年 5 巻 2 号 論文ID: pp.26-004

詳細
Abstract

The assessment and management of work-related psychosocial risks is mandated across many jurisdictions, yet a critical governance gap persists: there is little consensus on the point at which measured psychosocial hazard exposure constitutes an actionable risk requiring intervention. While physical and chemical risk assessment benefits from occupational exposure limits established through transparent, deliberative processes, psychosocial risk evaluation currently lacks equivalent mechanisms. The question of when exposure becomes an unacceptable risk is answered implicitly through methodological choices that are seldom articulated or justified. This position paper addresses this governance gap by distinguishing three logics of risk evaluation: predictive (risk as a statistical association with outcomes), normative (risk as a deviation from reference standards), and weighted (risk as frequency × severity prioritization). We analyze the conceptual foundations, governance implications, and practical trade-offs of each approach. Because each logic embeds different assumptions about when intervention is warranted, method choice is not a neutral technical decision but a governance choice that encodes assumptions about acceptable working conditions and employer obligations. We call for explicit governance processes that make methodological choices visible, debatable, and accountable, including documentation requirements, sector-level guidance on method selection, and inspector training on methodological adequacy. The challenge is not a lack of knowledge but the need for coordinated action to develop governance infrastructure for psychosocial risk evaluation comparable to that achieved in chemical and physical risk assessment.

Abbreviations
COPSOQ

Copenhagen Psychosocial Questionnaire

HSE

Health and Safety Executive

JEM

job-exposure matrix

LOAEL

lowest observed adverse effect level

NOAEL

no observed adverse effect level

OEL

occupational exposure limit

OSH

occupational safety and health

ROC

receiver operating characteristic

SLIC

Senior Labour Inspectors Committee

1. Introduction

The World Health Organization (WHO) and the International Labour Organization have documented the substantial global burden of disease and harm attributable to psychosocial occupational hazards, which impose enormous costs on workers, organizations, and societies.1,2) In response, many jurisdictions have enacted or strengthened legal requirements by defining psychosocial hazards and mandating employers to assess and manage psychosocial risks as part of their occupational health and safety obligations.3) The European Union’s Framework Directive 89/391/EEC, for example, explicitly requires risk assessment of all workplace hazards, including those of a psychosocial nature.4) Similar provisions exist in national legislation across multiple countries.5)

These legislative frameworks embody a clear promise: systematic, evidence-based identification of psychosocial hazards, followed by rigorous evaluation of risks and implementation of preventive measures, will maintain and foster employee health, safety, and well-being. The underlying logic mirrors established practice in physical and chemical risk assessment.6) However, despite two decades of instrument development and validation research enabling the reliable identification of psychosocial hazards,79) organizations implementing psychosocial risk assessments face a fundamental ambiguity: while there is some consensus that certain exposures, such as excessive working hours or extreme workload, constitute clear risks,10,11) there is limited agreement on when a measured psychosocial hazard more generally constitutes an actionable risk that requires intervention. We use the term “actionable risk” to denote the thought of a limit at which measured exposure to a psychosocial hazard is deemed sufficiently severe, prevalent, or predictive of harm that organizational intervention becomes warranted or required.

The absence of clear, standardized criteria for triggering risk mitigation actions can create a regulatory gap that may weaken the consistency, comparability, and enforceability of psychosocial risk management. This is not merely a technical issue about cutoff scores or statistical significance. It is a governance decision point, a juncture where methodological choices encode assumptions about acceptable working conditions, employer obligations, and the protection owed to workers.

In toxicology and industrial hygiene, analogous decision points are managed through explicit governance processes. Expert committees synthesize dose–response evidence, apply assessment factors to account for uncertainty and vulnerability, consider technological feasibility and socioeconomic impacts, and publish transparent rationales for proposed exposure limits.12,13) Uncertainty refers to limitations in the available data, such as extrapolation from animal studies to humans, from high to low doses, or from short-term to long-term exposure, which introduce imprecision into risk estimates. Vulnerability refers to the recognition that certain workforce subgroups (e.g., due to age, genetics, pre-existing health conditions, or concurrent exposures) may be more susceptible to harm at given exposure levels. Assessment factors are applied to account for both, ensuring that derived limits protect not only the average worker but also more sensitive individuals under conditions of incomplete knowledge.

This process is neither purely scientific nor purely political; it is explicitly recognized as a deliberate synthesis of scientific evidence, expert judgments, and societal priorities. Yet such structured governance mechanisms for psychosocial risk assessment have been limited to initiatives such as the development of the WHO guidelines for mental health at work.10) The question of when exposure becomes an unacceptable risk is instead answered implicitly through methodological choices that are not always articulated, debated, or justified.14) Furthermore, differences in the use of terminology in relation to the psychosocial work environment also exacerbate the challenge, and the need for clearer terminology and further development of the EU legal framework on psychosocial risks at work has repeatedly been highlighted in various scientific publications.3,1517) This is a prior and equally important challenge, one that remains contested and that shapes what is measured across different methodological traditions. While this definitional challenge inevitably influences the content of risk evaluations, the present paper focuses on the governance of the evaluation process itself: how measured exposures, irrespective of their definitional basis, are translated into risk determinations and intervention decisions.

We discuss how different methods of translating measured exposure into risk probability—specifically, what we term predictive, normative, and weighted logics—reflect distinct assumptions about acceptable risk and encode different criteria for when intervention becomes necessary. Drawing on the existing literature and established frameworks, we analyze the conceptual foundations, governance implications, and practical trade-offs of each approach. Our goal is not to determine which method is optimal, since each reflects defensible assumptions and serves legitimate purposes. Rather, we aim to clarify what each approach assumes, what it prioritizes, and what trade-offs it entails, thereby supporting more informed and transparent methodological choices in research, practice, and especially in regulation. These are not trivial technical details; they entail consequential decisions that shape resource allocation, intervention timing, and, ultimately, the protection afforded to employees.

2. Psychosocial Risk Assessment and the Problem of Risk Evaluation

2-1. The risk assessment and risk management process

In occupational health and safety, risk assessment involves hazard identification, risk analysis and evaluation, and the review of controls.18) These activities form part of a broader risk management process that includes the implementation of control measures and ongoing monitoring to support continual improvement and enhance health, safety, and well-being.19) Ultimately, aiming at improving health, safety, and well-being within a cycle of continuous improvement, this risk management framework has proven highly effective.20)

The first explicit attempt to apply the risk management framework using the hazard-risk-harm typology was made by Cox and Cox21) in a review for the UK Health & Safety Executive and guidance for the World Health Organization. Subsequent guidance published by the WHO—the Psychosocial Risk Management European Framework15)—and corresponding ISO standards22) linked psychosocial risk management to organizational management and operational processes. These frameworks enable organizations and policymakers to systematically manage psychosocial risks. Despite their valuable contributions, however, they do not provide specific guidance on risk evaluation, the point at which a psychosocial hazard becomes a risk requiring action.

At the core of the risk assessment process lies the fundamental understanding of risk as the product of the likelihood (or probability) of a hazard causing harm and the severity (or consequence) of the resulting harm.23) This understanding assumes that risk can be quantified by combining the likelihood of harm occurring with the magnitude of that harm. In physical and chemical risk assessment, this approach rests on decades of research establishing OELs, the maximum level of exposure that employees may experience without harm. Such thresholds emerge from research on dose–response relationships, epidemiological evidence, and consensus processes involving toxicologists, industrial hygienists, and regulatory bodies.12,13) The result is a relatively transparent governance structure: exposures below the threshold (e.g., noise <85 db(A) over 8 hours24)) are considered acceptable, while exposures above the limit trigger mandatory intervention.

2-2. The absence of equivalent standards for psychosocial hazards

Psychosocial risk assessment currently lacks such standardization. Numerous assessment methodologies and tools have been developed over the past two decades. Widely used instruments like the COPSOQ25) or the Health and Safety Executive’s Management Standards Indicator Tool26) allow the reliable identification of hazards such as quantitative or emotional demands, job control, role conflicts, or social support. However, no equivalent to OELs exists for psychosocial risk factors.6) A score of 65 on a “quantitative demands” scale (0–100) does not carry the same regulatory weight as 85 db(A) of noise exposure.

Currently, while many organizations measure psychosocial hazards consistently, the critical translation step from measured exposure to actionable risk mitigation remains methodologically ambiguous and organizationally discretionary, as organizations seldom estimate exposure–outcome relationships or conduct other forms of systematic risk evaluation.14) This gap has practical and regulatory enforcement consequences. Some guidance does exist; for example, quantitative reference databases provide benchmarks in the sense of JEMs,27,28) and instruments like the HSE Management Standards offer “states to be achieved.” However, these fall short of providing clear, enforceable thresholds but rather correspond to justified recommendations. Without unambiguous criteria for determining “how much is too much,” organizations may struggle to determine which hazards require immediate attention, which can be monitored, and which fall within acceptable bounds.

In 2012, the European SLIC conducted an inspection campaign on psychosocial risks, which highlighted that the number of workplaces including psychosocial risk assessments had increased and that knowledge of psychosocial risks had also increased among labor inspectors in all countries.29) In 2018, SLIC adopted a non-binding guide on the quality of risk assessment and risk management measures with regard to the prevention of psychosocial risks.30) While this guidance includes a section on evaluating and prioritizing risks, it does not specify the approach(es) for doing so. Again, the absence of standardized risk evaluation thus becomes a governance problem, not merely a technical one. This is particularly salient given that Jain et al.31) have shown that the introduction of specific national-level legislation on psychosocial risks and work-related stress is associated with more organizations implementing actions.

3. Three Logics of Psychosocial Risk Evaluation

3-1. Overview and rationale

In the absence of consensus and standardization, organizations and researchers have developed different approaches to risk evaluation. Three broad categories of practical procedures can be distinguished: (1) uniform cut-off procedures based on exposure frequency, (2) reference values derived from JEMs, and (3) threshold-based approaches analogous to OELs for chemical agents.32) While this categorization usefully distinguishes technical procedures, we propose a complementary framework that emphasizes the underlying logic of risk evaluation, that is, the fundamental assumptions about when exposure constitutes an actionable risk.

We distinguish three such logics: (1) predictive (risk as a continuous function of exposure), (2) normative (risk as a deviation from acceptable standards), and (3) weighted (risk as prioritization across multiple factors). These logics cut across the above-mentioned procedural categories but highlight the governance implications embedded in methodological choices. Each logic answers the fundamental question differently: when does intervention become necessary? The term “logics” rather than “methods” or “approaches” is deliberate. These represent different conceptual foundations for understanding risk, not merely technical variations in calculation. The same statistical procedure might be employed under different logics, with quite different interpretive frameworks and implications for action. For example, a regression model could be used purely for prediction (predictive logic) or to identify hazards exceeding threshold effects (normative logic). It is also important to note that these logics are not mutually exclusive in practice. Most validated measures of psychosocial hazards have been developed through some form of predictive logic assessment, demonstrating statistical associations with health outcomes, even when they are subsequently used within normative frameworks that compare scores to reference values. The logics thus represent different emphases and decision rules rather than entirely separate measurement traditions.

3-2. Predictive logic: risk as statistical association

3-2-1. Core assumptions and operationalizations

The predictive approach treats risk probability as a continuous function of exposure, estimated through statistical modeling of the relationship between psychosocial hazards and health outcomes. The notion of what constitutes a psychosocial hazard is rooted in the original conception of work-related stressors: job demands that carry an elevated likelihood of impairing employee health, safety, and well-being on a population-based level.33) Resting on the fundamental expectation that risk increases as exposure intensifies, any statistical attempt to quantify associations between psychosocial hazards and employee health can be subsumed under this category—whether through correlation analyses, structural equation modeling, or machine learning algorithms.

What unites these diverse techniques is the premise that all measured exposures contribute information about risk, and that greater exposure (in frequency, intensity, or duration) predicts a greater likelihood of harm. This logic is exemplified by the extensive body of longitudinal cohort research and meta-analyses examining psychosocial work exposures and health outcomes. A meta-review of 72 meta-analytic reviews synthesized evidence on associations between psychosocial work factors and outcomes such as cardiovascular diseases, mental disorders, and musculoskeletal conditions.34) Large-scale pooled analyses across multiple European cohorts have likewise been conducted, establishing quantitative associations between job strain and outcomes such as coronary heart disease and stroke.35,36) These studies exemplify the predictive logic: they estimate hazard–outcome relationships statistically and use the magnitude of the associations (e.g., relative risks, odds ratios, and hazard ratios) to characterize risk.

This approach maximizes predictive power by using all available information simultaneously and accounting for the fact that psychosocial risk rarely stems from a single factor. From a prevention perspective, such models can identify which hazards most strongly predict outcomes within a specific organizational context. By examining the magnitude of the respective statistical coefficients, information about the relative importance of each hazard is obtained.14,37)

3-2-2. Governance implications

The predictive logic can serve multiple governance functions. It provides a basis for ranking or prioritizing hazards at the team, unit, department, or organizational levels, offering guidance on where intervention resources might be most effectively deployed. When outcome data on, for example, job burnout are available, predictive models can identify which specific hazards most strongly contribute to health impairment in a particular context, enabling targeted rather than generic interventions.

However, the predictive logic does not inherently provide a clear action threshold. Because all hazards contribute to the outcomes to some degree, the question of when intervention is required remains open. It is worth noting that not all areas of occupational health and safety risk require a binary intervene/don’t intervene decision. Sometimes, graduated responses or continuous improvement approaches are appropriate. Nevertheless, for regulatory purposes where clear duties must be established, the absence of explicit thresholds creates challenges. A work unit with slightly elevated demands, marginally reduced control, and somewhat lower social support may show an elevated predicted risk score, but does this constitute an unacceptable risk requiring immediate intervention or an acceptable deviation within normal workplace variation? The predictive logic provides excellent discrimination for ranking work units or hazards by predicted risk but offers limited guidance for establishing bright-line regulatory requirements.

3-2-3. Limitations

The predictive approach faces important conceptual and empirical limitations that currently distinguish psychosocial risk assessment from chemical or physical risk assessment. First, even minimal exposure to psychosocial hazards may not be without consequences. While comprehensive dose–response curves for psychosocial hazards remain largely absent,38) concepts from toxicology such as NOAEL and LOAEL indicate that establishing clear thresholds below which exposure is entirely safe is often difficult.39,40)

Second, the predictive approach confronts challenges related to vulnerability and uncertainty that are systematically addressed in OEL setting but remain underdeveloped for psychosocial hazards. In toxicology, vulnerability refers to workforce subgroups who may be more sensitive to exposure due to genetics, age, pre-existing health conditions, or concurrent exposures; OELs are deliberately set to protect these vulnerable populations, not just the average worker.12,41) For psychosocial hazards, variables like age, gender, or prior mental health status are traditionally treated as statistical covariates35,36) rather than as indicators of differential vulnerability requiring protective threshold adjustments.

Similarly, uncertainty arising from inter-individual variability, dose–response ambiguity, and duration extrapolation42) is formally quantified and addressed through assessment factors in chemical risk assessment. In psychosocial risk assessment, analogous uncertainties manifest, for example, as methodological inconsistencies, measurement invariance problems, and cross-sectional design limitations,43,44) yet these are rarely translated into explicit uncertainty adjustments in risk probability estimates.

Third, predictive calculations require data on outcomes. In organizational settings, this might be available when validated questionnaires like the COPSOQ25) are used, since these instruments typically include items on self-reported health and other outcomes. Some organizations also use data from health screening or occupational health surveillance programs to establish hazard–harm relationships. However, when the relevance of specific psychosocial hazards is inferred indirectly from population-level coefficients derived from epidemiological research, relying solely on these coefficients may overlook important variability in job-specific susceptibility and contextual factors that moderate hazard-outcome relationships in particular organizational settings. Moreover, the strength of statistical associations can vary considerably depending on which outcomes are selected. It remains a legitimate point of debate whether mental health outcomes are the only appropriate choice in work-related contexts.6)

3-3. Normative logic: risk as deviation from acceptable standards

3-3-1. Core assumptions and operationalizations

The threshold-based approach attempts to resolve the ambiguity of continuous risk estimates by introducing an explicit reference standard. Within this logic, risk exists only when exposure exceeds an acceptable level, typically defined by population norms, industry benchmarks, or regulatory guidance based on OELs. In practice, this often involves comparing organization-specific psychosocial hazard levels against reference values derived from representative samples. Hazards scoring worse than the reference threshold are classified as risks requiring attention, whereas those within the acceptable range are not.

JEMs represent one important operationalization of this logic. JEMs provide averaged exposure profiles across occupational classifications, enabling systematic comparison between local conditions and population norms.27) The large amount of available reference values from the COPSOQ exemplifies this approach at the instrument level.45) Similarly, the HSE Management Standards define “states to be achieved” representing desirable conditions, against which organizational scores can be compared.46) Existing studies have examined the validity of JEMs for psychosocial job stressors,47,48) demonstrating both their utility and their limitations for risk assessment purposes.49) For instance, good agreement has been found between JEM-assigned and individually reported exposures for job demands and job control, but only moderate agreement for job insecurity and fairness of pay.47) Another study48) similarly demonstrated acceptable predictive validity for depression but identified notable misclassification risks, particularly for exposures that vary substantially within occupational categories. A more recent study49) showed that consistency in psychosocial hazard ratings across similar jobs was moderate to high for task-embedded and organizationally structured hazards such as leadership quality, environmental conditions, and role conflicts, but substantially lower for discretion-dependent dimensions such as social relations and degrees of freedom, reinforcing that JEMs function as useful screening tools rather than definitive risk determinants.

This method has strong intuitive and regulatory appeal. It mirrors the logic of OELs in physical risk assessment: define an acceptable level, measure actual exposure, intervene when the threshold is exceeded. The decision rule is transparent, defensible, and facilitates communication with employees, managers, and inspectors. It also aligns with legal frameworks that frame employer obligations as duties to maintain conditions within acceptable bounds.

3-3-2. Governance implications

From a regulatory perspective, the normative logic offers a critical advantage: it establishes an explicit decision rule that creates accountability and enables enforcement. When exposures exceed defined benchmarks, employers can be held responsible, and workers can point to objective standards when advocating for improvements. The approach enables systematic comparison across organizations, facilitating benchmarking, regulatory oversight, and identification of sectors or occupations requiring targeted policy attention. Such threshold-based approaches make visible the decision rule underlying risk classification—a transparency that predictive models, with their continuous probability estimates, do not inherently provide.

3-3-3. Limitations

However, it should be noted that comparing local conditions with reference values from JEMs is not equivalent to the understanding, quality, or specificity provided by OELs. JEMs provide exposure profiles based on averaged self-reports of individual employees across occupational classifications, typically following the International Standard Classification of Occupations (ISCO)50) or comparable standards. These usually represent job-level snapshots, an “as-is” state without indicating whether observed exposures are health-protective or merely typical.

Moreover, some psychosocial factors vary systematically with task characteristics (e.g., job autonomy and development opportunities), while others depend on local social contexts (e.g., leadership quality and workplace relationships) that may fluctuate considerably even within the same occupational classification.3,51) Population or sectoral norms therefore provide a benchmark for comparison but not necessarily a harm-based threshold. Most available references are specific to the tool or questionnaire applied for hazard identification, or represent very general cohorts.48,52) Since available references for psychosocial hazards are mostly established through self-reports based on questionnaires, it is crucial that data are collected using validated instruments. While averaging individual-level responses is often assumed to reduce random measurement error and individual response variability, it does not fully eliminate systematic self-report bias, such as social desirability, mood-dependent reporting, or shared method variance linked to organizational climate or survey context.7,43)

The logic itself introduces additional tensions. Threshold placement becomes a critical—and potentially arbitrary—decision if not guided by regulatory debate. Why should, for example, the 50th percentile of a reference population define an acceptable level? Should thresholds vary by industry, occupation, or country? The approach also produces sharp discontinuities: a work unit scoring just below the threshold requires no action, while one scoring marginally above demands intervention, despite negligible practical differences. Furthermore, reference-based selection may exclude hazards that are predictive of harm but not statistically worse than reference norms. A hazard at the 60th percentile may still contribute to poor health outcomes yet would not be flagged for intervention under this logic.

An alternative approach uses ROC curves to determine optimal cut-points by evaluating the trade-off between sensitivity and specificity in predicting adverse outcomes,32) though this addresses the question of where to place thresholds rather than whether thresholds should be used at all. Such approaches are generally not preventive in nature but mark a point at which harm has already occurred.6)

3-4. Weighted logic: risk as prioritization

3-4-1. Core assumptions and operationalizations

The classical risk approach adapts the frequency–severity formulation to psychosocial risk assessment. Here, risk is estimated by combining two components: the frequency (or prevalence) of exposure to a hazard and the severity of its potential consequences. The product of frequency and severity yields a risk score that allows prioritization across multiple hazards. Scholars have proposed a specific operationalization that calculates risk as the squared correlation between a hazard and an outcome, multiplied by the mean exposure level on the hazard scale.53) This logic aligns more closely with engineering safety practice by explicitly quantifying both the prevalence of exposure and the empirical harm potential as separate, transparent components of risk.32)

In a comparative analysis14) of established psychosocial risk-evaluation methods (including regression-based approaches, threshold deviation selection, and a frequency-severity framework) applied to a cross-sectional organizational sample assessing hazards such as quantitative demands, emotional demands, role conflicts, and social support (COPSOQ), the weighted logic emerged as the most promising. Nevertheless, the choice of method led to sizable differences in which hazards were prioritized for intervention, underscoring that risk evaluation is not a neutral technical step but a consequential methodological choice with direct implications for organizational action. This study represents one of the few systematic comparisons of risk evaluation methods in psychosocial assessment. This has been developed further by designing a risk matrix for psychosocial hazards, using odds ratios from logistic regression to scale severity and assigning hazards into cells based on empirical associations with health outcomes.32)

Such approaches support resource allocation decisions by identifying high-frequency–high-severity hazards for priority intervention. The method is widely used in practice and enjoys intuitive accessibility for non-specialists. It also has practical advantages in smaller samples, where more advanced statistical models may fail to meet their assumptions, as it relies only on bivariate correlations rather than the simultaneous estimation of multiple predictors. Statistically, the multiplicative formula resembles an interaction effect as found in prominent psychological theories of work stress and design.54)

3-4-2. Governance implications

The weighted logic introduces distinct challenges that extend beyond technical operationalization. Most fundamentally, prioritization does not necessarily establish acceptability. A risk-based approach tells organizations which hazards warrant highest priority but provides no absolute threshold indicating when working conditions are acceptable. This creates an accountability gap: whereas threshold-based approaches establish clear duties to act when exposures exceed defined limits, weighted approaches currently offer no comparable benchmark.

The context-dependency of risk rankings further complicates regulatory oversight. Because rankings depend on the specific distribution of exposures and outcome associations within each organizational unit, what constitutes “high priority” naturally varies across organizations. This variability makes cross-organizational comparison and sector-level regulation difficult, at least if such uniformity is a regulatory goal. Regulators cannot establish uniform intervention requirements because the risk-based logic explicitly rejects universal mandates in favor of context-specific prioritization.

3-4-3. Limitations

The weighting scheme introduces operationalization challenges. What metric or instrument best captures frequency? How should severity be quantified? Different choices have been shown to produce different risk rankings,14,37) yet these methodological decisions are often made implicitly or without justification.

Additionally, the frequency–severity weighting embeds implicit value judgments that are rarely made explicit. By multiplying frequency and severity with equal weight, the approach treats a common hazard with moderate impact as equivalent to a rare hazard with severe consequences, yet these scenarios may warrant different organizational responses. Moreover, when risk scores are calculated at the aggregate level (e.g., across entire operating areas), demographic subgroups become buried under group averages. Unlike technical safety hazards that are tied to specific equipment, machines, or locations, psychosocial hazards interact with individual characteristics such as age, gender, or health status.3,55) A hazard that severely affects a minority within a unit may produce moderate aggregate exposure means and weak overall correlations, resulting in low risk scores despite constituting serious risk for those affected. The aggregation logic inherent in calculating group-level means and correlations can thus render invisible the concentrated harm experienced by vulnerable subpopulations.

The classical risk approach has been widely adopted in organizational practice, often operationalized through semi-quantitative risk matrices that categorize hazards into discrete risk levels. Such matrices have drawn criticism for arbitrary categorization, poor rank ordering, and false precision.5658) Despite these limitations, weighted approaches align well with organizational realities of constrained resources and competing priorities. They acknowledge explicitly what threshold-based approaches often obscure: organizations cannot address all identified hazards simultaneously and must make allocation decisions. The transparency of the ranking process—showing which hazards were considered and how they were prioritized—supports accountability even as it complicates enforcement.

4. Governance Implications and the Case for Explicit Deliberation

4-1. Method choice as governance choice

The preceding outline reveals that the transition from hazard identification to risk classification is not a neutral technical step. Each logic embeds assumptions about acceptable working conditions, thresholds for action, and the prioritization of hazards. Without transparency about these choices, risk assessments across organizations remain difficult to compare, undermining regulatory oversight and limiting the ability of social partners to evaluate progress.

What does it mean, concretely, to say that method choice is a governance choice? Consider a manufacturing facility where psychosocial hazards have been measured using a validated questionnaire. Under the predictive logic, a regression model might identify “quantitative demands” as the strongest predictor of emotional exhaustion in this context, flagging it for intervention even though exposure levels are typical for the sector. Under the normative logic, the same hazard might not exceed reference values and therefore would not be classified as an actionable risk, while “influence at work,” which falls significantly below reference norms, would be flagged instead. Under the weighted logic, “role conflicts” might emerge as highest priority because it combines moderate-high exposure prevalence with meaningful outcome associations. Each approach directs attention and resources differently. Employees will experience different interventions depending on which logic the organization applies.

This has direct implications for accountability. If an employee develops a stress-related health condition and the organization can demonstrate that it conducted a risk assessment, the question becomes: was the assessment adequate? Under current conditions, there is no agreed standard for evaluating methodological adequacy. An organization using one logic may have systematically overlooked hazards that another logic would have flagged. Regulators, inspectors, and courts have limited basis for evaluating whether a particular methodological choice was defensible. This governance ambiguity is compounded by the broader macro context, where the boundaries of employer responsibilities are not clearly defined, leaving room for interpretation of legal provisions.3)

The challenge becomes even more consequential in a labor market undergoing rapid transformation, with new psychosocial risks emerging from digitalization, automation, and artificial intelligence whose impacts are not yet fully understood.59) Without explicit governance processes for methodological choices, organizations navigating novel hazards have no shared framework for determining when these new exposures cross the threshold into actionable risk.

4-2. The current situation: implicit choices without deliberation

Currently, organizations conducting psychosocial risk assessments make methodological choices—explicitly or implicitly—about how to translate exposure into risk classifications. These choices are rarely documented, justified, or subjected to scrutiny. A practitioner comparing results across two organizations cannot know whether apparent differences in risk levels reflect genuine differences in working conditions or merely differences in the logic applied to evaluate those conditions.

This situation contrasts sharply with physical and chemical risk assessment, where OELs provide a common reference point. While debates about specific OEL values continue, the pure existence of a shared framework enables meaningful comparison, regulatory enforcement, and worker protection claims. The governance processes underlying OEL development—though not perfect—offer a model of what explicit deliberation might look like. The OECD13) has documented how OEL-setting organizations operate through expert committees (such as the EU’s former Scientific Committee on Occupational Exposure Limits [SCOEL], now European Chemicals Agency’s [ECHA] Risk Assessment Committee [RAC], or the German Senate Commission for the Investigation of Health Hazards of Chemical Compounds in the Work Area [MAK Commission]), formal literature review processes, transparent criteria for evidence evaluation, public comment periods, and published rationales for proposed values. The Nordic Expert Group exemplifies international collaboration, with five Scandinavian countries using common scientific criteria documents as the basis for their respective national OELs.60) Professional organizations like American Conference of Governmental Industrial Hygienists (ACGIH) and American Industrial Hygiene Association (AIHA) have developed systematic procedures built on “rigorous science and a transparent process”.61) These processes are neither purely scientific nor purely political—they explicitly integrate value judgments about acceptable risk levels with empirical evidence about dose–response relationships.

For psychosocial hazards, no such framework currently exists. The question of when exposure becomes unacceptable risk is answered organization by organization, assessment by assessment, without explicit deliberation about the underlying logic. The result is a fragmented landscape in which the same working conditions might be classified as “acceptable” in one organization and “requiring immediate intervention” in another, with no visibility into whether this reflects genuine contextual differences or merely methodological inconsistency.

4-3. The argument for structured governance

Given the consequential nature of methodological choices in psychosocial risk evaluation, we argue for explicit governance processes that make these choices visible, debatable, and accountable. This does not necessarily mean mandating a single method; the three logics serve different purposes and may be appropriate in different contexts. Rather, it means requiring transparency about which logic is applied, why it was chosen, and what assumptions it entails. This point is more consequential than it may first appear: how risk evaluation is conducted forms the critical bridge between hazard identification and risk mitigation.62) In practice, employees experience the interventions deemed actionable, a determination shaped by the chosen evaluation method.

What might such governance look like in practice? Drawing on the OEL-setting experience, several mechanisms warrant consideration. First, expert working groups at the sector or national level could be tasked with developing guidance on acceptable probability estimation methods for psychosocial hazards, analogous to the role played by SCOEL/RAC for chemical hazards in the EU. These groups would include occupational health researchers, practitioners, social partners, and regulatory representatives. Second, documentation requirements could mandate that organizational risk assessments explicitly state which logic was applied and provide a justification for the choice. This would enable regulatory inspectors to evaluate methodological adequacy and would support meaningful comparison across organizations. Third, reference value development through tripartite deliberation could establish sector-specific benchmarks that carry greater legitimacy than instrument-specific norms derived from convenience samples. Fourth, inspector training should equip labor inspectorates to evaluate not just whether a risk assessment was conducted, but whether the methodological approach was defensible given the context.

Regulatory guidance should acknowledge that method choice is a substantive decision, not merely a technical one. Guided by the policy context, organizations should be required to document their approach to risk evaluation and to justify their choice in relation to the hazards assessed and the outcomes considered. Such documentation would enable meaningful comparison across organizations, support regulatory oversight, and strengthen the legitimacy of organizational risk assessments.

5. Recommendations for Research, Policy, and Practice

This paper underscores a dual reality. On the one hand, psychosocial risk management research has advanced rapidly, with growing conceptual clarity, improved measurement tools, and increasing recognition among policymakers and social partners of the importance of the psychosocial work environment. Organizations consistently report that psychosocial risks are more difficult to manage than traditional occupational hazards,63) and labor inspectorates and occupational health services similarly call for clearer guidance and more actionable frameworks.6466) On the other hand, significant knowledge gaps remain, particularly regarding how to translate measured psychosocial hazard exposures into defensible, transparent, and enforceable risk evaluations. There remains a lack of appropriate actions at organizational level to address them, a situation driven by persistent deficits in awareness, resources, and technical support.3) This dual reality of rapidly expanding knowledge but insufficient implementation creates a governance challenge that must be addressed through coordinated action across research, policy, and practice.

5-1. Research directions

Future research must address both foundational questions and emerging challenges in psychosocial risk management. A first priority is the need for meta-reviews that synthesize the rapidly expanding evidence base, also with respect to both the health and the organizational impact.3) The field has grown substantially, but the absence of systematic synthesis makes it difficult for policymakers and practitioners to distinguish robust findings from preliminary or context-specific results.3) Meta-reviews would help consolidate knowledge on exposure–outcome relationships, intervention effectiveness, and methodological strengths and weaknesses. They would also provide a clearer foundation for developing consensus on how to operationalize key constructs such as risk probability, which remains conceptually and methodologically fragmented.14)

A second priority is research on new and emerging psychosocial risks. There is growing evidence that psychosocial risks increasingly arise from digitalization, automation, robotization, and artificial intelligence59)—developments that reflect profound technological transformations in the world of work.63) These changes introduce new forms of work organization, surveillance, and algorithmic management that are not yet fully understood. Research should examine how these emerging risks interact with traditional hazards and how they shape exposure patterns across sectors, occupations, and demographic groups. Crucially, for these emerging hazards, none of the three evaluation logics discussed in this paper rests on an established evidence base: predictive models lack the longitudinal data needed to quantify exposure–outcome relationships, normative benchmarks have not yet been developed, and weighted approaches cannot draw on validated severity estimates. Research on emerging psychosocial risks must therefore proceed in parallel with the development of appropriate evaluation frameworks, making the governance questions raised in this paper particularly pressing.

Third, there is a pressing need for rigorous intervention evaluation studies. National legislation increases organizational action, but current interventions often strengthen job resources without reducing job demands.31) High-quality evaluations—using longitudinal designs, mixed methods, and clear theories of change—are essential to identify what works, for whom, and under what conditions.67) Crucially, future research should also investigate how the choice of risk evaluation method affects intervention effectiveness, not just predictive accuracy. If different methods lead organizations to prioritize different hazards, then method choice may shape the very interventions implemented and their success or failure.

Finally, policy evaluation studies are essential. Given the diversity of national approaches, systematic evaluations of legislative and regulatory frameworks are needed to assess clarity, enforceability, and impact on organizational behavior. Such studies can provide evidence for future harmonization and support the development of more effective regulatory instruments. They can also help clarify how regulatory frameworks influence organizational choices about risk evaluation, an area where greater transparency and methodological consensus are urgently needed.

5-2. Policy directions

Translating the governance implications identified above into concrete policy action requires targeted interventions at multiple levels. First, the SLIC guidance on psychosocial risk assessment should be revised to address method selection explicitly, providing inspectors and organizations with a framework for understanding the governance implications of different approaches. Second, regulators should consider requiring organizations to document which probability estimation logic they apply and to justify that choice, enhancing transparency and enabling meaningful oversight. Third, the development of sector-specific reference values through tripartite deliberation—involving researchers, social partners, and regulatory authorities—could provide benchmarks that carry greater legitimacy than instrument-specific norms. Fourth, training programs for labor inspectors should include explicit attention to methodological adequacy in psychosocial risk assessment, equipping inspectors to evaluate not just procedural compliance but substantive quality. Fifth, EU-level guidance could clarify expectations regarding acceptable probability estimation methods, recognizing that method choice is a substantive decision requiring justification rather than a purely technical matter left to organizational discretion. By embedding methodological clarity into regulatory frameworks, policymakers can help ensure that psychosocial risk management becomes more consistent, accountable, and aligned with the preventive principles of EU OSH legislation.

5-3. Practice implications

Translating conceptual clarity into practical action is essential for organizations, inspectorates, and occupational health services. The development of accessible tools that support practitioners in selecting and justifying probability estimation methods is crucial. Existing labor inspection tools and guidance should be revised, promoted, and integrated into national inspection systems. These tools can support consistent enforcement and help organizations meet their obligations under the OSH Framework Directive. They can also help organizations understand the governance implications of methodological choices and link risk assessment results to preventive action.

Furthermore, stakeholder competencies must be strengthened. Employers, workers, OSH professionals, and labor inspectors require training to understand psychosocial risk management, interpret risk assessments, and implement preventive measures. Such training should include explicit attention to the different logics of risk evaluation and their implications for intervention priorities.

For regulatory practice, the analysis presented here suggests that risk-based (weighted) methods may be most appropriate as complements to threshold-based (normative) requirements rather than substitutes: normative thresholds establish minimum acceptable conditions, while risk-based prioritization guides intervention sequencing among hazards requiring attention. Organizations and practitioners should consider which logic best fits their regulatory context, data availability, and organizational capacity, while remaining transparent about the choice made and its implications.

6. Conclusion

In this paper, we argued that psychosocial risk evaluation, that is, determining at which point a hazard becomes a risk requiring intervention, is not a neutral technical step but a governance decision with consequential implications for worker protection, resource allocation, and regulatory enforcement. We have distinguished three logics of risk evaluation—predictive, normative, and weighted—and shown that each embeds different assumptions about when intervention is warranted and how accountability should be structured.

The absence of explicit governance processes for these methodological choices creates a regulatory gap that undermines the promise of systematic, comparable, and enforceable psychosocial risk management. Unlike physical and chemical risk assessment, which for many hazards benefits from structured deliberation and transparent threshold-setting processes, psychosocial risk assessment lacks equivalent mechanisms. The question of when exposure becomes unacceptable risk is answered implicitly, organization by organization, without visibility into the underlying logic applied.

The challenge ahead is not a lack of knowledge, but the need for coordinated, transparent, and accountable action. In the near term, this means organizations documenting their methodological choices and regulatory guidance acknowledging method selection as a substantive issue. In the longer term, it means developing governance infrastructure for psychosocial risk evaluation that approaches the transparency achieved in chemical and physical risk assessment. These are not merely technical questions to be resolved by methodologists but also governance questions that deserve explicit and inclusive deliberation.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Availability of data and materials

Not applicable.

Funding

The authors did not receive support from any organization for the submitted work. No funding was received to assist with the preparation of this manuscript. No funding was received for conducting this study.

Authors’ contributions

Yannick Metzler: conceptualization, investigation, project administration, writing—original draft; Svenja Müller: conceptualization, writing—original draft; Luis Torres: validation, writing—review and editing; Aditya Jain: conceptualization, supervision, writing—review and editing.

All authors have read and approved the final manuscript and agree to be accountable for all aspects of the work.

Competing interests

The authors have no conflicts of interest to declare that are relevant to the content of this article.

Use of AI in the paper

AI-based tools (DeepL/Claude) were used solely for language editing. The authors take full responsibility for the scientific content and the accuracy of all data and analyses.

References
 
© 2026 The Japan Association of Occupational Health Law
feedback
Top