Abstract
Objective
Large language models (LLMs) are increasingly used by patients and the public to access medical information. In pulmonary medicine, where information needs span a wide range of chronic and potentially serious diseases, the quality of artificial intelligence-generated communication requires systematic evaluation.
Methods
This cross-sectional comparative study evaluated responses to 37 standardized pulmonary medicine questions covering major respiratory domains. Each question was answered by four LLMs-Gemini 2.5, GPT-4, Claude Sonnet 4, and DeepSeek v3.2-and a board-certified professor of pulmonology with more than 25 years of clinical experience. Six blinded raters, including three pulmonologists and three laypersons, assessed anonymized responses using a five-point Likert scale across conceptually aligned domains: accuracy/perceived reliability, clarity, and scientific adequacy/practical usefulness. Inter-rater reliability was assessed using intraclass correlation coefficients (ICC). Comparative analyses were performed using repeated-measures one way analysis of variance with Bonferroni-adjusted pairwise comparisons and Friedman tests.
Results
Inter-rater reliability indicated slight single-measure agreement (ICC 0.09-0.12) and fair-to-moderate average-measure agreement (ICC 0.36-0.45), which is consistent with the subjective nature of communication-focused assessment. Overall, LLM-generated responses received higher communication-quality ratings than physician-generated responses across evaluation domains. Gemini 2.5 and GPT-4 achieved the most favorable scores, particularly on accuracy/perceived reliability and clarity. Differences between LLMs and the physician were more pronounced among lay raters than among expert raters. Among expert raters, physician-generated responses were rated among the highest for accuracy and scientific adequacy, although this pattern was not observed for clarity.
Conclusion
Contemporary LLMs may provide clear, reliable, and practically useful information on pulmonary medicine, particularly as perceived by lay users, when evaluated from both expert and lay perspectives. These findings support their adjunctive use in patient education and health communication and emphasize that LLMs cannot replace clinical judgment, professional accountability, or physician oversight.
Introduction
Effective medical communication is a core component of high-quality healthcare delivery, since the accuracy, reliability, and clarity of the information shared with patients and caregivers directly shape clinical understanding and decision-making. Prior work has consistently demonstrated that clear and accurate responses to patient inquiries are associated with greater patient satisfaction, improved treatment adherence, and a stronger physician-patient relationship, making comprehensible and trustworthy communication central to patient-centered care(1, 2). This is particularly consequential in populations with limited health literacy, where access to accurate information from trustworthy sources is essential for both treatment success and patient safety(3).
In parallel, the integration of the internet and digital media into everyday life has reshaped how patients and their relatives seek health information, with online searches for symptoms and diagnoses giving rise to the so-called “Dr. Google” phenomenon(4). Much of the health-related material available online, however, is inaccurate, incomplete, or contextually misleading(5), with the potential to provoke unnecessary anxiety, misdirected care-seeking, or distrust toward medical advice. From a public health perspective, the dissemination of scientifically accurate, comprehensible, and trustworthy health information through digital channels has therefore become both an individual and a collective responsibility(6).
Within this evolving information ecosystem, artificial intelligence (AI) has progressively moved from a peripheral analytic tool to a core component of contemporary medical informatics. Earlier generations of AI demonstrated value in image-based diagnostic support, bioinformatics pipelines, and the interpretation of large-scale data from genetic biobanks, where machine learning approaches enabled scalable analyses of complex, heterogeneous datasets that would be impractical to perform manually(7-9). Building on this foundation, AI systems have since been embedded into clinical decision-support workflows, risk stratification models, and individualized treatment planning informed by genomic, proteomic, and radiomic data(10, 11). What distinguishes the current phase, however, is the migration of AI from structured data and pixel-level analysis toward the communication layer of medicine-the space in which patients, caregivers, and clinicians exchange information, weigh options, and arrive at decisions. This shift reframes AI not only as an analytic engine but also as an active participant in health information delivery and brings the quality, transparency, and safety of AI-generated communication into the foreground of medical informatics research.
Beyond diagnostics and data analytics, large language models (LLMs) are transforming the communication layer of medicine. Leveraging advanced natural language processing, these systems have shown promise in medical question answering, summarization of clinical dialogues, and analysis of unstructured text from electronic health records(11, 12). Their use in healthcare has expanded rapidly, with domain-specific implementations developed to improve the accuracy, clarity, and empathy of responses to medical queries(11). At the same time, their deployment raises non-trivial concerns regarding factual reliability, contextual nuance, data security, algorithmic bias, and the risk of confidently incorrect outputs-issues that underscore the need for safe, transparent, and human-centered integration of AI technologies into clinical practice(7). Systematic evaluation of LLM-generated responses, especially within clinically meaningful, domain-specific contexts such as pulmonology, is therefore essential before such tools can be responsibly recommended for patient-facing use(13).
These considerations are particularly salient in pulmonary medicine. Respiratory conditions-including asthma, chronic obstructive pulmonary disease (COPD), pneumonia, pulmonary thromboembolism, interstitial lung diseases, obstructive sleep apnea, tuberculosis, and lung cancer-account for a substantial share of the global disease burden and result in ongoing public information needs. Many of these conditions are chronic, require sustained self-management, and involve decisions regarding inhaler use, oxygen therapy, smoking cessation, vaccination, surgical options, and end-of-life care. The accessibility, accuracy, and clarity of the information delivered to patients and their relatives therefore have direct implications for adherence, safety, and informed decision-making. Despite the rapid uptake of LLMs in this space, the available evidence remains fragmented: many existing evaluations examine a single model, rely exclusively on expert reviewers, or assess technical accuracy without considering how lay audiences perceive the information. Studies comparing multiple state-of-the-art LLMs head-to-head with an experienced physician on the same set of publicly relevant questions-using blinded evaluations by both clinicians and laypersons across complementary criteria of accuracy, clarity, and scientific or practical adequacy-remain limited. This gap is consequential because expert and lay perspectives often diverge: clinicians tend to weigh factual completeness and technical precision, whereas patients and the broader public appraise responses through the lens of perceived reliability, comprehensibility, and real-world usefulness. Capturing both perspectives is therefore essential to characterize the realistic role of LLMs in patient-facing health communication and decision-support tasks.
The present study was designed to address this gap. We conducted a comparative, blinded evaluation in which four contemporary LLMs (Gemini 2.5, GPT-4, Claude Sonnet 4, and DeepSeek v3.2) and a board-certified pulmonology professor with more than 25 years of clinical experience responded to the same set of 37 standardized pulmonary medicine questions derived from publicly observable information-seeking behavior. Using conceptually aligned criteria (accuracy/perceived reliability, clarity, and scientific adequacy/practical usefulness), anonymized responses were independently rated by three expert pulmonologists and three laypersons. By examining inter-rater reliability and systematically comparing model and physician performance across audience types, the study provides an informatics-oriented assessment of how well current LLMs support communication about pulmonary medicine to non-specialist users, and aims to clarify their potential-and limitations-as complementary tools alongside experienced clinicians in health information delivery and clinical decision-support workflows.
Materials and Methods
This study was a cross-sectional, comparative evaluation designed to assess the quality and reliability of responses generated by four LLMs and a human pulmonology expert.
Design and Procedure
This study did not involve patients, patient data, identifiable personal information, biological samples, or any clinical intervention. The raters evaluated anonymized written responses; therefore, ethics committee approval was not required according to institutional and national regulations. Informed consent is not applicable, as the study did not involve patients or identifiable personal data.
Question Selection
A total of 37 standardized questions related to chest diseases were developed by a pulmonology specialist who had no conflicts of interest with either the expert or layperson raters. The questions were designed to represent a balanced distribution across major subspecialties within pulmonary medicine. Particular attention was paid to ensuring that the set encompassed diverse, clinically relevant domains: asthma (3), COPD (4), pneumonia (4), pulmonary thromboembolism (4), interstitial lung diseases (3), obstructive sleep apnea (4), lung cancer (7), tuberculosis (2), and other respiratory topics (6). In addition to its clinical relevance, the study was designed to reflect real-world patterns of health information consumption, focusing on publicly accessible questions and communication formats that influence health-related understanding beyond clinical settings. The complete 37-item question set is provided as supplementary material; formal psychometric content validation (e.g., a content-validity index) was not performed.
The questions’ content was informed by Google Trends data, social media discussions, and frequently asked questions directed to a pulmonologist’s professional social media account, to ensure that the questions reflected publicly relevant, real-world, and digitally mediated health concerns. This approach provided a representative sampling of topics commonly encountered in both clinical practice and public health communication contexts.
Participants and Evaluation Procedure
Six independent raters participated in the evaluation: three expert pulmonologists (E1, E2, and E3, with 7, 8, and 10 years of clinical experience, respectively) and three laypersons (L1, L2, and L3, a 32-year-old university-educated teacher, a 53-year-old employee and former amateur footballer, and a 45-year-old homemaker with a high-school education). All raters were blinded to the identity of the content source (LLM-generated vs. professor-written).
Four LLMs (Gemini 2.5, GPT-4, Claude Sonnet 4, and DeepSeek v3.2) were selected because they represent the most advanced publicly available generative models at the time of the study (prompts submitted in August 2025), encompassing diverse architectures, training paradigms, and user interfaces. Their inclusion enabled a balanced assessment across distinct LLM frameworks. A pulmonology professor with more than 25 years of clinical experience provided physician-generated responses. The physician is actively engaged in clinical practice, holds an administrative role, and is involved in faculty development and educator training activities. Because a single expert served as the human comparator, this benchmark reflects the performance of one highly experienced pulmonologist rather than that of physicians in general.
Each of the six raters evaluated the anonymized answers of the five respondents (four LLMs and one human professor). Ratings were assigned on a five-point Likert scale (1=lowest, 5=highest) based on the designated evaluation criteria for each group. None of the raters had any conflict of interest with either the professor or the LLM selection process. All evaluations were performed independently, and raters were blinded to the identity of each response source to ensure unbiased judgment based solely on content quality.
Conceptual Alignment of Evaluation Criteria
The three criteria used by layperson raters-perceived reliability, clarity, and practical usefulness-were designed to conceptually parallel the expert criteria of accuracy, clarity, and scientific adequacy, respectively. Perceived reliability corresponds to accuracy, as both reflect the extent to which information is regarded as correct and trustworthy, albeit from different perspectives: experts assess factual correctness, whereas laypersons judge credibility and the confidence it inspires. Clarity was retained consistently across groups, reflecting the mutual importance of linguistic simplicity and structural transparency. Finally, practical usefulness parallels scientific adequacy: experts evaluate the scientific rigor and completeness of the answer, whereas laypersons perceive the same quality in terms of real-world applicability and utility.
The criteria used for layperson evaluations were adapted from the methodological framework applied in the study by Charide et al.(14), which focused on public and user-centered assessment approaches in health communication research. In parallel, the expert evaluation criteria were determined following the methodological approach used by Akçay et al.(15) in their study evaluating the accuracy and reliability of ChatGPT 4’s responses in thoracic surgery. This mapping ensures conceptual equivalence across evaluator groups while maintaining relevance to their respective levels of expertise. Because expert accuracy (factual correctness) and lay perceived reliability are related but non-identical constructs, inter-rater reliability is reported separately for experts and laypersons, and the pooled coefficient is interpreted with caution.
Procedural Neutrality and Data Collection
The questions were submitted to the LLMs by an independent engineer specializing in software development, who is one of the authors of this article. This engineer had no background in medicine and no prior acquaintance or conflict of interest with the participating pulmonology professor. Each LLM was reset to a new session before being questioned to prevent memory retention or contextual bias, ensuring that each interaction represented a clean, independent exchange. All prompts were entered using the same computer, under identical network conditions and internet speed, to standardize response latency and interface performance. The engineer took no part in the writing, scoring, or evaluation stages and did not participate as either an expert or layperson rater, thereby ensuring full procedural neutrality and minimizing potential bias throughout data collection.
Statistical Analysis
Inter-rater reliability was assessed using intraclass correlation coefficients (ICC) to quantify agreement among raters. ICC values estimate the proportion of variance in ratings that is attributable to true score differences rather than to random error or rater subjectivity. Both single-measure and average-measure ICCs (two-way random effects model, absolute agreement type) were computed to capture individual and group-level reliability(16).
Interpretation of ICC values followed established benchmarks, with values below 0.20 indicating slight agreement, 0.21-0.40 fair agreement, 0.41-0.60 moderate agreement, 0.61-0.80 substantial agreement, and 0.81-1.00 almost perfect agreement(17).
Comparative analyses across the five response sources were performed using repeated-measures one way analysis of variance (ANOVA) because each rater provided scores for all sources, resulting in a within-subjects (dependent) data structure. This approach allowed comparison of mean ratings across conditions while accounting for the interdependence of repeated observations from the same raters. Bonferroni-adjusted pairwise comparisons were applied to control for type I error in multiple testing. To corroborate these findings and to address potential deviations from normality, Friedman tests were additionally conducted as non-parametric equivalents. The unit of analysis was the per-question mean across raters (n=37), with response source as the within-subject factor (degrees of freedom=4, 144). Sphericity was assessed (Greenhouse-Geisser), and a linear mixed-effects model with crossed random effects for questions and raters was acknowledged as a valid alternative.
To ensure robustness and to explore subgroup consistency, three complementary analytic models were applied: (1) all raters (E1-E3, L1-L3), (2) experts only (E1-E3), and (3) laypersons only (L1-L3). For each model, Bonferroni-adjusted pairwise comparisons were used to compare the five response sources across the three evaluation criteria. In these analyses, accuracy for expert raters was conceptually aligned with perceived reliability for layperson raters; clarity was shared by both groups; and scientific adequacy for expert raters was aligned with practical usefulness for layperson raters.
The primary analyses included all six raters, with subgroup analyses conducted separately for expert and layperson raters. Statistical analyses were conducted using IBM SPSS Statistics version 26 (IBM Corp., Armonk, NY, USA), with recomputation and verification performed using Python and the pingouin package.
Results
Inter-rater Reliability
Inter-rater reliability was assessed using a two-way random-effects model (absolute agreement). As summarized in Table 1, single-measure ICCs ranged from 0.09 to 0.12 and average-measure ICCs ranged from 0.36 to 0.45 across accuracy/perceived reliability, clarity, and scientific adequacy/practical usefulness (all p<0.001).
According to commonly applied interpretative frameworks, these values correspond to slight agreement for single measures and to fair-to-moderate agreement for average measures, indicating variability in individual ratings but improved consistency when scores were aggregated. Differences in scoring patterns across evaluators and domains were statistically significant (p<0.001), reflecting heterogeneity in perception rather than systematic measurement error.
Overall, these findings demonstrate that while individual-level agreement was limited, aggregated ratings showed greater consistency, underscoring the subjective nature of communication-based evaluations and the influence of evaluator background on inter-rater agreement metrics.
Summary of Evaluator Scores
Across evaluation domains, the relative ranking of mean scores remained stable, with LLM-generated responses generally receiving higher ratings than physician-generated responses. Differences in mean ratings across models and evaluator groups reflected variability in perceived quality rather than isolated or extreme scoring behavior. The observed score distributions indicate coherent, interpretable evaluation patterns across all six raters and support the internal consistency of the comparative assessment framework.
When all six raters were included, all four LLMs received higher mean ratings across all evaluation criteria than did physician-generated responses. Among the LLMs, Gemini 2.5 and GPT-4 achieved the highest overall mean scores, followed by DeepSeek v3.2 and Claude Sonnet 4. Physician-generated responses consistently received lower average ratings, with the largest differences observed in the scientific adequacy domain. Measures of dispersion were modest across all models, indicating a relatively consistent pattern of scoring among evaluators (Table 2a).
Among expert raters, the single pulmonologist was rated among the highest, rather than below the models; the physician received an accuracy/perceived reliability mean (4.66), nearly identical to Gemini 2.5, which had the highest mean score (4.67), and a scientific adequacy mean (4.13) comparable to or higher than most LLMs. Clarity scores were similar across sources. Standard deviations were small, indicating consistent expert judgments (Table 2b).
Among layperson raters, differences in mean ratings across response sources were more pronounced. Across all evaluation criteria, LLM-generated responses received higher mean scores than physician-generated responses, with the largest differences observed in accuracy and clarity. Gemini 2.5 achieved the highest and most consistent mean ratings among the LLMs. In contrast, ratings assigned to physician-generated responses showed greater dispersion, reflected by higher standard deviations, indicating increased variability in evaluations by non-expert participants (Table 2c).
When mean ratings across all 37 questions were examined within each evaluation domain (accuracy and perceived reliability, clarity, and scientific adequacy or practical usefulness), LLM-generated responses generally showed higher mean scores, whereas physician-generated responses tended to cluster toward lower values. Overlapping data points indicate identical or highly similar scores assigned by different models to specific questions, reflecting comparable scoring patterns across models rather than isolated deviations. These distributions provide a visual summary of score alignment across questions and evaluation domains (Figure 1).
Comparative Analyses
A repeated-measures ANOVA demonstrated a significant main effect of response source across all rater configurations. When all six raters were included, statistically significant differences were observed among the five response sources for all three evaluation criteria (p<0.001). Differences were most pronounced for accuracy and perceived reliability, both of which showed the largest effect size (partial η2=0.26). For all raters: accuracy/perceived reliability, F (4.144)=12.84, partial η2=0.263; clarity, F (4.144)=10.47, partial η2=0.225; scientific adequacy/practical usefulness, F (4.144)=11.89, partial η2=0.248 (all p<0.001). Subgroup analyses for expert and layperson raters are presented in Supplementary Tables 1-3.
When analyses were restricted to expert pulmonologists, significant effects of response source remained for accuracy (perceived reliability) and for scientific adequacy or practical usefulness, whereas differences in clarity were of smaller magnitude. In contrast, among layperson raters, significant differences among response sources were consistently observed for all evaluation criteria (p<0.001), with larger effect sizes indicating greater differentiation in mean ratings.
Across subgroup analyses, the direction and magnitude of differences varied by rater expertise, with LLM advantages most evident among lay raters. Detailed ANOVA results for each configuration are provided in Supplementary Tables 1-3.
Given the significant main effects observed, Bonferroni-adjusted post-hoc tests were performed to determine which specific model comparisons accounted for these differences. Detailed pairwise statistics, including mean differences, 95% confidence intervals, t-values, degrees of freedom, Cohen’s dz, and Bonferroni-adjusted p-values, are provided in Supplementary Tables 4-6.
When all raters were included, Bonferroni-adjusted pairwise comparisons showed that Gemini 2.5 and GPT-4 were rated significantly higher than the physician on all three criteria (p≤0.015), whereas DeepSeek v3.2 did not differ significantly from the physician on any criterion, and Claude Sonnet 4 differed only for clarity (p=0.010). Among the models, Gemini 2.5 scored significantly higher than DeepSeek v3.2 and Claude Sonnet 4 on accuracy/perceived reliability and scientific adequacy/practical usefulness, and higher than GPT-4 on those two criteria, with no significant difference in clarity (Table 3a).
Among expert raters, no LLM was rated significantly higher than the physician on any criterion after Bonferroni adjustment; however, the physician was rated significantly higher than Claude Sonnet 4 on accuracy/perceived reliability (p=0.046) and scientific adequacy (p=0.004). Experts rated all sources similarly for clarity. These findings indicate that, among professional evaluators, the experienced pulmonologist was judged to be at least as accurate and scientifically adequate as the strongest LLMs (Table 3b).
Layperson raters’ evaluations revealed statistically significant differences between response sources across all evaluation domains (p<0.001). Across measures of accuracy and perceived reliability, clarity, and scientific adequacy or practical usefulness, LLM-generated responses received higher mean ratings than physician-generated responses, with Gemini 2.5 and GPT-4 showing slightly higher mean ratings than Claude Sonnet 4 and DeepSeek v3.2. These differences were most pronounced for accuracy and for perceived reliability and clarity. Overall, lay evaluators consistently assigned higher ratings to LLM-generated content across domains, reflecting differences in perceived quality and usability rather than isolated rating effects (Table 3c).
Discussion
The present study demonstrated slight single-measure inter-rater reliability (ICC 0.09-0.12) and fair-to-moderate average-measure (ICC 0.36-0.45) inter-rater reliability, which is acceptable within the context of subjective, communication-focused assessments. Across evaluation domains, LLM responses received higher mean ratings than the physician did overall; however, this pattern was driven by lay raters. Expert raters rated the physician among the highest for accuracy and scientific adequacy, and no LLM significantly exceeded the physician’s performance. Answer length was also positively associated with ratings across sources (Pearson r≈0.73), identifying response length/format as a potential confounder. Mean response lengths differed across sources, with Gemini 2.5 producing the longest responses (approximately 33 words), followed by DeepSeek v3.2 (approximately 21 words), GPT-4 (approximately 20 words), the physician-generated responses (approximately 18 words), and Claude Sonnet 4 (approximately 14 words). Repeated-measures ANOVA confirmed a significant main effect of response source across all criteria, which was most pronounced for accuracy and perceived reliability. The divergence between lay and expert raters underscores the strong influence of evaluator background on communication-focused assessments.
These findings are broadly consistent with prior studies examining LLM performance in patient-oriented medical communication, while also highlighting important nuances. In a comparable study evaluating responses to rheumatology-related patient questions, both rheumatologists and patients assessed LLM-generated and clinician-generated answers. Although patient ratings did not differ significantly between the two approaches in terms of comprehensiveness or readability, rheumatologists rated LLM responses lower with respect to accuracy and completeness, and preference rates differed markedly between patients and experts(13). In contrast, our study found higher ratings for LLM-generated responses, particularly among layperson raters, whereas differences within the expert subgroup were less pronounced. Notably, no statistically significant differences were found in clarity scores between LLM and physician responses.
Similar patterns have been reported in other clinical contexts. An emergency medicine study comparing LLM-generated patient education responses with those produced by orthopedic clinicians found more favorable ratings for AI-generated content across multiple domains, aligning with the communication-related advantages observed in our study(18). Likewise, investigations focusing on patient questions related to lung cancer surgery reported high ratings for both accuracy and clarity of LLM responses, while emphasizing that such systems should be used as educational or supportive tools rather than as independent clinical decision-makers(15).
In contrast, studies assessing diagnostic and triage performance present a more cautious perspective. Evaluations comparing LLMs with physicians and laypersons have shown that while LLMs may outperform non-experts and provide information perceived as more useful than conventional search engine outputs, they continue to underperform physicians in diagnostic accuracy and complex clinical reasoning tasks(19). Similarly, a recent systematic review reported that although LLMs can achieve accuracy comparable to specialists in selected imaging-related tasks, they remain limited in patient-centered communication and clinical reasoning, reinforcing the ongoing necessity of human clinical expertise(11). Consistent with these observations, Liu et al.(20) demonstrated that LLMs performed below senior physicians in pulmonary diagnostic tasks, highlighting the distinction between theoretical knowledge synthesis and integrative, experience-based diagnostic reasoning.
Beyond comparisons between physicians and AI systems, our findings also point to differences among LLMs. Although Gemini 2.5 achieved higher mean ratings in several domains, these differences did not consistently reach statistical significance in assessments by expert evaluators. Among laypersons, however, Gemini 2.5 was rated more favorably than other models, suggesting that communicative fluency and interpretability may play a stronger role in public perception than domain-specific accuracy alone. This observation aligns with findings from intensive care medicine, where GPT-4-based models demonstrated high accuracy but were also associated with increased computational demands and concerns regarding confidently incorrect outputs, underscoring the importance of ongoing validation prior to broader clinical adoption(21).
Multiple evaluations of medical AI systems have further identified persistent challenges, including logical inconsistencies, contextual fragmentation, and limitations in deep diagnostic reasoning(22, 23). As Turner(24) has noted, trust remains an ethical concept that is difficult to ascribe to algorithmic systems lacking moral agency. In this context, the present findings support the view that LLMs should be regarded as adjunctive tools designed to augment-rather than replace-physicians. Their greatest potential lies in supporting medical communication and information delivery, while human clinicians remain essential for clinical judgment, accountability, empathy, and holistic decision-making at the core of medical practice.
Study Limitations
This study has several limitations. First, physician responses were obtained from a single pulmonologist, which may limit the generalizability of the findings to broader clinical practice. Second, the number of evaluators was limited, restricting the extent to which the results can be considered representative of larger expert and lay populations. Third, all primary outcome measures were based on subjective ratings of communication quality rather than on objective assessments of factual correctness or clinical validity. In addition, although prompts were predefined and AI-generated responses were archived at the time of evaluation, LLMs are continuously updated, which may affect reproducibility across different model versions. The study focused on perceived communication quality and the usefulness of information; therefore, the findings should not be interpreted as reflecting diagnostic accuracy, clinical reasoning performance, or decision-making capability. In addition, perceived accuracy was rated by humans without gold standard (e.g., guideline-based) adjudication of factual correctness; the human comparator was a single pulmonologist, limiting generalizability to physicians as a group; the question set did not undergo formal content validation; response length differed across sources and was positively associated with ratings, representing a potential format confounder; repeated-measures ANOVA, rather than mixed-effects modeling, was used for the primary comparisons. The linguistic and contextual characteristics of the questions and evaluators may also limit generalizability to other languages or health care settings.
Conclusion
This study suggests that current LLMs can generate pulmonology-related medical information that is perceived as clear and scientifically adequate by both expert and non-expert audiences. Although some models received higher communication-related ratings, these results should be interpreted in light of the subjective nature of the evaluations and do not imply equivalence to clinical expertise. The advantage of LLMs was driven mainly by lay raters; among expert raters, the single pulmonologist was rated at least as highly as the strongest models for accuracy and scientific adequacy.
Importantly, the results underscore that LLMs are best viewed as supportive tools rather than as substitutes for physicians. Human clinicians remain essential for clinical reasoning, contextual judgment, empathy, and professional accountability. Accordingly, the integration of AI-assisted information delivery with physician oversight may enhance patient communication while preserving the central role of human expertise in medical practice.
Supplementary Tables: https://d2v96fxpocvxx.cloudfront.net/cf9d60d6-523c-458a-a2e6-78728d3ffbb0/content-images/e99fe721-803e-4f4f-9249-125ccbd841b5.pdf


