Article ID: 2026PCP0002
Micro-expressions are subtle, involuntary facial movements that reveal a person's underlying emotional state. Their automatic recognition has attracted growing attention in the deep learning community, yet remains challenging due to their subtle appearance and the scarcity of large-scale annotated datasets. This paper proposes FREED-UNet, a class-conditional diffusion model for synthesizing high-fidelity apex-frame images to support micro-expression recognition. The model integrates FreeU-based skip-connection reweighting and conditional self-attention to enhance both structural integrity and fine-grained facial details. To ensure validity, we introduce a multi-stage structural filtering pipeline that combines MediaPipe detection and BiSeNet facial part segmentation. Experiments on the CASME II and newly collected TUMME datasets demonstrate that incorporating FREED-UNet-generated images into training improves the recognition accuracy of SE-DenseNet-cc from 86.74% to 90.76%. These findings highlight the effectiveness of diffusion-based data augmentation in advancing micro-expression recognition. We also report FID/KID to quantify visual fidelity.