Ilmu Komputer & AI editorial
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
The core problem
Innovation
ChatGPT-3.5 achieved complete agreement in 29 cases (48.4%), partial agreement in 14 cases (23.3%), and no agreement in 17 cases (28.3%). Gemini showed complete agreement in 20 cases (33%), partial agreement in 9 cases (15%), and no agreement in 31 cases (52%). The external pathologist had complete agreement in 36 cases (60%), partial agreement in 17 cases (28%), and no agreement in 7 cases (12%). Significant differences in diagnostic agreement were found between the LLMs and the pathologist (P < 0.001). The distribution of agreement levels is summarized in the following table:
| Evaluator | Complete Agreement | Partial Agreement | No Agreement |
|-----------------|--------------------|-------------------|--------------|
| ChatGPT-3.5 | 29 (48.4%) | 14 (23.3%) | 17 (28.3%) |
| Gemini | 20 (33%) | 9 (15%) | 31 (52%) |
| External Pathologist | 36 (60%) | 17 (28%) | 7 (12%) |
A chi-square test confirmed the significant difference in agreement levels between the LLMs and the human pathologist, with a p-value less than 0.001.
Why it matters
The results indicate that while ChatGPT-3.5 and Gemini can provide accurate diagnoses in certain instances, their overall performance is insufficient for reliable use in real-life clinical settings. ChatGPT-3.5 outperformed Gemini, with a higher rate of complete agreement (48.4% vs. 33%) and a lower rate of no agreement (28.3% vs. 52%). However, both LLMs lagged behind the external pathologist, who achieved 60% complete agreement. The significant difference (P < 0.001) underscores the current limitations of LLMs in diagnostic accuracy. Potential benefits of LLMs include rapid analysis of complex data and assistance in differential diagnosis, but biases and constraints such as reliance on textual descriptions without images, potential for hallucination, and lack of clinical context must be addressed. The study's findings align with the need for cautious integration of LLMs as adjunct tools rather than replacements for human expertise. Future research should explore multimodal LLMs that incorporate histopathological images and larger, more diverse datasets to improve generalizability. The equation for diagnostic agreement rate can be expressed as:
This study highlights the importance of rigorous validation before clinical deployment of LLMs in pathology.
Who should read this
Opening member content…