Jadwal Sholat

Memuat jadwal sholat…

Ilmu Komputer & AI editorial

Open AccessOA2025

Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study

A comparative analysis of ChatGPT-3.5, Gemini, and human pathologists in dermatopathologic diagnosis
Talar Sabir Ahmed; Rawa M. Ali; Ari M. Abdullah; H. Yasseen; Ronak Ahmed; A. M. Salih; Dilan S. Hiwa; Shvan H. Mohammed· Barw Medical Journal· 2025· DOI 10.58742/bmj.v3i3.180

The core problem

The integration of large language models (LLMs) into healthcare is rapidly evolving, with the potential to transform patient care and outcomes. Since the introduction of ChatGPT by OpenAI in November 2022, and Gemini by Google, these models have demonstrated remarkable capabilities in generating human-like responses across various tasks. In medicine, ChatGPT has shown promise in analyzing complex medical data, aiding differential diagnosis, and synthesizing information from patient records, medical research, and clinical guidelines. However, the application of LLMs in dermatology and pathology remains limited. This study aims to explore the role of LLMs, specifically ChatGPT-3.5 and Gemini, in diagnosing dermatologic conditions within pathology. It compares their accuracy and concordance with human pathologists and investigates potential advantages, biases, and limitations of integrating LLM tools into pathology decision-making processes.

Innovation

ChatGPT-3.5 achieved complete agreement in 29 cases (48.4%), partial agreement in 14 cases (23.3%), and no agreement in 17 cases (28.3%). Gemini showed complete agreement in 20 cases (33%), partial agreement in 9 cases (15%), and no agreement in 31 cases (52%). The external pathologist had complete agreement in 36 cases (60%), partial agreement in 17 cases (28%), and no agreement in 7 cases (12%). Significant differences in diagnostic agreement were found between the LLMs and the pathologist (P < 0.001). The distribution of agreement levels is summarized in the following table:

| Evaluator | Complete Agreement | Partial Agreement | No Agreement |
|-----------------|--------------------|-------------------|--------------|
| ChatGPT-3.5 | 29 (48.4%) | 14 (23.3%) | 17 (28.3%) |
| Gemini | 20 (33%) | 9 (15%) | 31 (52%) |
| External Pathologist | 36 (60%) | 17 (28%) | 7 (12%) |

A chi-square test confirmed the significant difference in agreement levels between the LLMs and the human pathologist, with a p-value less than 0.001.

The integration of large language models (LLMs) into healthcare is rapidly evolving, with the potential to transform patient care and outcomes. Since the introduction of ChatGPT by OpenAI in November 2022, and Gemini by Google, these models have demonstrated remarkable capabilities in generating human-like responses across various tasks. In medicine, ChatGPT has shown promise in analyzing complex medical data, aiding differential diagnosis, and synthesizing information from patient records, medical research, and clinical guidelines. However, the application of LLMs in dermatology and pathology remains limited. This study aims to explore the role of LLMs, specifically ChatGPT-3.5 and Gemini, in diagnosing dermatologic conditions within pathology. It compares their accuracy and concordance with human pathologists and investigates potential advantages, biases, and limitations of integrating LLM tools into pathology decision-making processes.
The study employed a comparative design using 60 real histopathology case scenarios of skin conditions, half neoplastic and half non-neoplastic, selected from a hospital database. Two board-certified pathologists reviewed each case to establish a consensus diagnosis. Cases were included if they had complete histopathological reports and comprehensive patient demographics; cases with incomplete data were excluded. A random sampling method ensured representativeness. The cases were presented to ChatGPT-3.5, Gemini, and an external board-certified pathologist in March 2023. The LLMs received standardized prompts: initially "Hello," followed by "Please provide the most accurate diagnoses from the texts that will be given below." Each case was copy-pasted individually, and the first response was recorded. If no diagnosis was given, the prompt was repeated with a variation. The external pathologist received the same information without histopathological images. Responses were categorized into complete agreement, partial agreement, or no agreement with the original diagnosis, based on whether both the general diagnosis and specific subtype were correctly identified. Statistical analysis was performed to compare agreement levels between LLMs and the pathologist.

Why it matters

The results indicate that while ChatGPT-3.5 and Gemini can provide accurate diagnoses in certain instances, their overall performance is insufficient for reliable use in real-life clinical settings. ChatGPT-3.5 outperformed Gemini, with a higher rate of complete agreement (48.4% vs. 33%) and a lower rate of no agreement (28.3% vs. 52%). However, both LLMs lagged behind the external pathologist, who achieved 60% complete agreement. The significant difference (P < 0.001) underscores the current limitations of LLMs in diagnostic accuracy. Potential benefits of LLMs include rapid analysis of complex data and assistance in differential diagnosis, but biases and constraints such as reliance on textual descriptions without images, potential for hallucination, and lack of clinical context must be addressed. The study's findings align with the need for cautious integration of LLMs as adjunct tools rather than replacements for human expertise. Future research should explore multimodal LLMs that incorporate histopathological images and larger, more diverse datasets to improve generalizability. The equation for diagnostic agreement rate can be expressed as:

This study highlights the importance of rigorous validation before clinical deployment of LLMs in pathology.

Who should read this

CS practitioners and researchers

Opening member content…