Untitled
Open-weight large language models (LLMs) can be deployed locally at low cost, making them attractive for patient education in resource-constrained and non-English settings that remain underrepresented in current LLM benchmarking. Whether medical domain specialization improves patient-facing answers in such settings is unclear. We compared one general-purpose open-weight model (GPT-OSS-20B) with on
Open-weight large language models (LLMs) can be deployed locally at low cost, making them attractive for patient education in resource-constrained and non-English settings that remain underrepresented in current LLM benchmarking. Whether medical domain specialization improves patient-facing answers in such settings is unclear. We compared one general-purpose open-weight model (GPT-OSS-20B) with one medically fine-tuned open-weight model (MedGemma-27B-Instruct) for thyroid cancer patient education in Turkish. Sixty Turkish patient questions about thyroid cancer were answered by both models. Five endocrinologists, blinded to model identity and study hypotheses, rated each response on 5-point Likert scales for Accuracy, Completeness, Clarity, Clinical Utility, and Satisfaction. Primary inference used per-question median ratings (N = 60 paired observations per criterion) with Wilcoxon signed-rank tests and Holm adjustment; effect size was rank-biserial correlation (RBC), and location shift was estimated with Hodges–Lehmann differences. Inter-rater reliability was assessed using ICC (2, k), and ceiling-aware summaries included perfect-score and top-box analyses. GPT-OSS-20B achieved higher question-level median ratings than MedGemma-27B-Instruct across all five criteria after Holm correction. The largest differences were observed for Satisfaction (median 5.0 vs. 4.0; RBC = 0.788; Holm-adjusted p < 0.001) and Completeness (median 5.0 vs. 4.0; RBC = 0.599; Holm-adjusted p < 0.001). For Accuracy, Clarity and Clinical Utility, Hodges–Lehmann location shifts were 0 points because of frequent ties at the top of the scale, although rank-based effect sizes remained moderate to large (RBC 0.45–0.64). Inter-rater reliability was good and comparable across models (ICC (2, k) ≈ 0.74–0.80). Ceiling aware reporting showed consistently higher perfect-score proportions for GPT-OSS-20B across criteria, with the most pronounced gaps in Satisfaction and Completeness. In this head-to-head comparison, the general-purpose GPT-OSS-20B was rated more highly than the medically fine-tuned MedGemma-27B-Instruct on all five criteria, with the largest and most consistent differences in Satisfaction and Completeness. Medical domain specialization therefore did not confer an advantage for this patient-facing task in Turkish. Because only two models were compared, these results should not be read as a general claim about general-purpose versus medically fine-tuned models as classes.