2026 Volume 44 Issue 4 Pages 161-168
This study investigates the feature representations of medical vision-language models (VLMs) to determine if medical domain-specific training truly enhances their discriminative power for downstream tasks. While numerous medical VLMs have been proposed, their feature spaces remain underexplored. We analyzed the feature distributions of representative medical and non-medical VLMs across eight diverse medical imaging datasets using dimensionality reduction and linear discriminators. Our findings reveal that medical specialization does not consistently guarantee superior image representations. Instead, non-medical VLMs employing large language models (LLMs) as text encoders demonstrate competitive or superior classification performance. Furthermore, both medical and non-medical VLMs are strongly affected by background biases, such as overlaid text strings on images, which they misinterpret as discriminative features. The results suggest that enhancing the text encoder with an LLM contributes more significantly to model performance than large-scale pre-training on noisy medical images.