JMA Journal
Online ISSN : 2433-3298
Print ISSN : 2433-328X
Opinion: Artificial Intelligence in Medicine
Artificial Intelligence in Medical Writing: Subtle Errors and Their Complex Consequences
Shigeki Matsubara, Daisuke Matsubara, Kazuhiko Kotani
Author information
JOURNAL OPEN ACCESS
Supplementary material

2026 Volume 9 Issue 4 Pages 992-996

Details
Abstract

Although artificial intelligence (AI; such as ChatGPT) can improve manuscript readability, AI-related false descriptions (so-called hallucinations) and incorrect reference retrieval have been repeatedly reported. We tested (1) whether ChatGPT, when provided with medically correct inputs, generates a manuscript containing medically incorrect statements―particularly whether it misunderstands or overlooks “subtle but important” medical issues―and (2) whether ChatGPT cites inappropriate references. We input bullet points on placenta percreta, tasked ChatGPT-5 with generating a mini-review, and asked it to confirm whether the output was medically correct. We introduced a small, deliberate trap. In percreta, a recent conceptual change has gained increased attention: this condition is considered to result from uterine abnormality rather than abnormal placental invasion. This etiopathological shift could be misinterpreted as implying “weaker adherence” and therefore “less difficult surgery,” leading to the notion that percreta could be managed at secondary-level institutions. The ChatGPT-generated manuscript cited appropriate references and was almost medically correct, except for one critical issue: it stated that percreta could be managed in secondary-level institutions, which is incorrect and potentially dangerous. During the “confirmation” stage, ChatGPT raised a caution regarding this issue, but not in a definitive manner. An additional experiment was conducted on hypercholesterolemia. ChatGPT again failed to address an important issue, familial hypercholesterolemia, which requires a management strategy different from that for non-familial hypercholesterolemia. Overall, ChatGPT generated a linguistically appealing manuscript with largely correct context, but it produced incorrect and potentially dangerous statements in clinically critical areas. When incorrect statements are subtle rather than obvious, authors, journals, and readers may fail to recognize them. This is paradoxical: advances in AI may reduce “apparent” errors while generating less recognizable ones. When evaluating AI-assisted manuscripts, careful review by individuals with deep domain knowledge is mandatory.

Content from these authors
© 2026 Japan Medical Association

JMA Journal applies the Creative Commons Attribution License to all works published by the journal. Anyone may download, reuse, copy, reprint, distribute, or modify articles published in the journal, if they cite the original authors and source. No permission is required from the publisher.
https://creativecommons.org/licenses/by/4.0/
Previous article Next article
feedback
Top