Contents
pdf Download PDF
pdf Download XML
36 Views
15 Downloads
Share this article
Original Article | Volume 12 Issue 10 (OCTOBER, 2026) | Pages 257 - 266
Validation of ChatGPT and Gemini AI Against Conventional Expert Analysis of Hospital Cumulative Antibiograms in Guiding Empirical Antibiotic Policies
 ,
 ,
 ,
 ,
 ,
1
Undergraduate Researcher, Kurnool Medical College, Kurnool, Andhra Pradesh, India.
2
Department of Microbiology, Kurnool Medical College, Kurnool, Andhra Pradesh, India.
Under a Creative Commons license
Open Access
Received
Sept. 15, 2026
Revised
Sept. 21, 2026
Accepted
Oct. 2, 2026
Published
Oct. 10, 2026
Abstract
Background: Cumulative antibiograms guide empirical antibiotic therapy, but interpreting them depends on microbiologists, whose time is often scarce in resource-limited settings. General-purpose large language models (LLMs) could help close this gap, yet how reliably they interpret structured susceptibility data has not been well tested.Aim: To compare ChatGPT and Gemini against a three-expert microbiologist consensus in interpreting a Clinical and Laboratory Standards Institute (CLSI) M39-compliant hospital cumulative antibiogram and generating empirical antibiotic recommendations, and to characterise where and why they diverged.Materials and Methods: This single-centre, retrospective, proof-of-concept study used the 2024 annual cumulative antibiogram (1,492 non-duplicate isolates) of a tertiary-care teaching hospital in southern India. Three expert microbiologists interpreted the antibiogram and reached a consensus reference standard. ChatGPT (GPT-5.2) and Gemini 3 each analysed the same data once, using one standardised, constraint-based prompt and blinded to the expert output. Each model was rated separately against the consensus at 25 predefined decision points in four domains (50 model–expert comparisons), as complete, high, or moderate concordance, or major discordance. Proportions are reported with Wilson 95% confidence intervals (CI), and a blinded physician reviewed all outputs for safety. Results: No comparison showed major discordance. ChatGPT showed complete concordance at 19/25 decision points (76.0%; 95% CI 56.6–88.5) and complete-or-high concordance at 24/25 (96.0%; 95% CI 80.5–99.3). For Gemini, the corresponding figures were 11/25 (44.0%; 95% CI 26.7–62.9) and 16/25 (64.0%; 95% CI 44.5–79.8). Both models matched the experts on all eight resistance-pattern items (Gram-negative predominance, extended-spectrum β-lactamase proxy phenotypes, intensive care unit carbapenem erosion, extensively drug-resistant Acinetobacter, methicillin-resistant Staphylococcus aureus and vancomycin-resistant enterococci) and on avoiding aminopenicillins, third-generation cephalosporins and fluoroquinolones. ChatGPT had one moderate discordance: it offered a cautious regimen for community-acquired pneumonia where the experts judged the data insufficient. Gemini had nine, concentrated in syndrome-wise recommendations (3/5), reserve and higher-toxicity drug classes (4/8) and stewardship safeguards (2/4), reflecting earlier use of aminoglycosides, carbapenems, colistin and tigecycline (exploratory exact McNemar test, p=0.021). Both models completed their analyses in under one minute; the expert panel took about 35 minutes. Conclusion: A constrained LLM closely reproduced expert interpretation of resistance patterns and broad empirical guidance. Stewardship restraint and toxicity awareness differed between models, and Gemini fell short of expert practice in these areas. LLMs appear best suited to a human-in-the-loop role rather than as a substitute for expert microbiologists. These single-centre, single-run findings need multicentre, multi-run validation.
Keywords
INTRODUCTION
Antimicrobial resistance (AMR) is one of the most serious threats to global health. An estimated 4.95 million deaths in 2019 were associated with bacterial AMR, 1.27 million of them directly attributable to it, and the burden falls hardest on low- and middle-income countries [1,2]. Inappropriate empirical antibiotic therapy is a major driver of this problem. Hospitals therefore rely on cumulative antibiograms, built using standardised methods such as the CLSI M39 guideline, to steer empirical prescribing, support stewardship and track local resistance trends over time [3,4]. Interpreting an antibiogram well, however, requires specialised expertise. That expertise is often scarce in resource-limited settings, and gaps or inconsistencies in expert review can affect how patients are treated [5]. Large language models such as ChatGPT (OpenAI) and Gemini (Google) have moved quickly into clinical decision-support research. Under controlled conditions they can summarise laboratory data and produce guideline-concordant narratives [6–8]. Whether they can be trusted with microbiological reasoning is a separate question. LLMs can hallucinate resistance patterns, underweight toxicity and stewardship considerations, and perform unevenly depending on the task and the prompt [9,10]. Most existing evaluations have used infectious-disease vignettes or narrative cases [8,9], or machine-learning prediction of resistance from isolate-level data [11]. Few have examined the structured cumulative susceptibility data that clinicians actually work from. Few studies have compared general-purpose LLMs directly with expert microbiologist consensus on a real, CLSI M39-compliant cumulative antibiogram, and fewer still in the Indian context. This gap limits how confidently such tools can be built into stewardship workflows [12]. We therefore designed this proof-of-concept study to test how closely ChatGPT and Gemini track a three-expert microbiologist consensus when interpreting one institution's annual cumulative antibiogram and generating empirical antibiotic recommendations, and to examine where, and why, they diverge.
MATERIALS AND METHODS
Study design and setting This single-centre, retrospective, proof-of-concept study was conducted in the Department of Microbiology, Kurnool Medical College and Government General Hospital, Kurnool, a tertiary-care teaching hospital in southern India, over three months in 2025. It involved only secondary analysis of aggregate laboratory data, with no patient-level intervention or real-time clinical decision-making. The Institutional Ethics Committee approved the study (approval no. 988/2025). Data source and dataset The analysis used the hospital's annual cumulative antibiogram for 2024, prepared according to CLSI M39 [3]. Of 1,494 clinically significant bacterial isolates retrieved, 1,492 remained after deduplication and were analysed. Extracted variables included organism identity, specimen type (blood, urine, pus, respiratory), care location (outpatient, non-ICU ward, ICU) and categorical susceptibility results (susceptible/intermediate/resistant/not tested). Minimum inhibitory concentration (MIC) values, molecular resistance markers, prior antibiotic exposure and clinical outcomes were not available. Deduplication and antibiogram construction Deduplication followed CLSI M39: the first isolate per patient per organism per surveillance period was retained. A data-analysis tool (Julius AI) assisted this step, and the investigators independently cross-checked it by hand against CLSI M39 criteria, excluding repeat isolates before analysis. Percentage susceptibility was calculated as the number of susceptible isolates divided by the number tested, multiplied by 100. Only organism–antibiotic combinations with at least 30 non-duplicate isolates were reported, stratified by location and specimen. No values were imputed or extrapolated. Reference standard (expert consensus) Three expert microbiologists, each with at least five years' experience in clinical microbiology and antimicrobial stewardship, independently interpreted the finalised antibiogram. Each gave empirical antibiotic recommendations, with justifications, for predefined clinical syndromes: community-acquired urinary tract infection (UTI), complicated UTI/urosepsis, community-acquired pneumonia (CAP), hospital-acquired pneumonia (HAP) and sepsis of unknown source. Each also recorded the time the analysis took. Disagreements were resolved through a structured consensus discussion, and the resulting consensus served as the reference standard. AI-based interpretation Two general-purpose LLMs, ChatGPT (GPT-5.2, OpenAI) and Gemini 3 (Google), were tested as analytical tools through their public web interfaces at default settings on 28 December 2025 [13,14]. Each model received the identical cumulative antibiogram dataset and a single, standardised, constraint-based prompt. The prompt instructed the model to follow CLSI M39 principles, use only the supplied data, state uncertainty explicitly, and avoid extrapolation or fabrication. Each model was queried once, and its output was recorded verbatim. The models generated empirical recommendations for the same predefined syndromes or, where the data were insufficient, stated so explicitly. Blinding and sequence of analysis Expert and AI interpretations were carried out separately. The expert panel was blinded to the AI outputs, and the AI analyses were carried out without access to the experts' conclusions. For the safety assessment, the AI-generated and expert recommendations were anonymised and presented in random order to a physician who had no other role in the study. This physician rated clinical appropriateness, guideline adherence and any safety concerns. Comparison framework and outcomes Each model's output was compared with the expert consensus at 25 predefined decision points in four domains: resistance-pattern interpretation (n=8), syndrome-wise empirical recommendations (n=5), drug-class positions (n=8) and stewardship/safety safeguards (n=4). Each model was rated separately at each decision point, giving 50 model–expert comparisons. Each comparison was rated as complete concordance (the same position as the expert consensus); high concordance (the same clinical direction, with a minor difference in nuance, framing or caveat); moderate concordance (a clinically relevant difference in agent choice, escalation or safeguards); or major discordance (a recommendation likely to cause unsafe under-treatment or clearly inappropriate therapy). The primary outcome was the distribution of these ratings for each model. Secondary outcomes were the proportion of clinically acceptable concordance (complete or high), agreement by domain, the nature and direction of discordance, analysis time, and variation among the three experts. Statistical analysis Agreement is reported descriptively as counts and percentages in each concordance category, separately for each model. Wilson score 95% CIs were calculated for the proportions of complete and of clinically acceptable concordance. Cohen's kappa was not calculated, because each decision point produced a single ordinal concordance rating rather than paired, independent classifications into a common set of categories. The two models were compared on clinically acceptable concordance across the 25 shared decision points using an exact (binomial) McNemar test. The decision points come from a single antibiogram and a single run of each model and are not fully independent, so this test is treated as exploratory. A two-sided p<0.05 was considered statistically significant. Analyses were performed using Python 3.13 with SciPy 1.18. Reporting The study was reported following the DECIDE-AI guideline for early-stage clinical evaluation of AI-driven decision-support systems [15] and the TRIPOD-LLM guideline for studies using large language models [16].
RESULTS
Dataset characteristics and adequacy The 2024 cumulative antibiogram contained 1,492 non-duplicate isolates; two repeat isolates were excluded from the 1,494 retrieved. It met all CLSI M39 data-adequacy criteria: every reported organism–antibiotic combination included at least 30 isolates, %S/%I/%R reporting was complete, and results were stratified by location and specimen [Table/Fig-1]. All three experts independently judged the dataset adequate for empirical guidance. Both LLMs reached the same conclusion, giving complete concordance on data adequacy. Resistance-pattern interpretation Both LLMs and the expert panel identified the same main epidemiological signals [Table/Fig-2]. Gram-negative bacilli predominated, led by Escherichia coli and Klebsiella pneumoniae. Enterobacterales showed high resistance to aminopenicillins, third-generation cephalosporins and fluoroquinolones, consistent with widespread extended-spectrum β-lactamase (ESBL) phenotypes. Reduced carbapenem susceptibility, most marked for K. pneumoniae in the ICU, indicated a carbapenem-resistant Enterobacterales (CRE) burden. Acinetobacter baumannii showed an extensively drug-resistant (XDR) profile, and Pseudomonas aeruginosa showed variable susceptibility. Among Gram-positive organisms, methicillin-resistant Staphylococcus aureus (MRSA) was common, but susceptibility to vancomycin, linezolid and teicoplanin was preserved; vancomycin-resistant enterococci (VRE) were also identified. Of the eight resistance-pattern decision points, ChatGPT showed complete concordance at seven and high concordance at one. Gemini showed complete concordance at six and high concordance at two. Both models agreed completely with the experts on data adequacy, dominant pathogens, ESBL proxy phenotype, XDR non-fermenters, MRSA and VRE. The two models differed on deduplication. ChatGPT, unprompted, stated that deduplication could not be independently verified from the data supplied, which matched the experts' approach of manual verification (complete). Gemini accepted the deduplicated data as correct without comment (high). On ICU CRE burden, the experts stated the risk explicitly, ChatGPT gave a graded risk statement, and Gemini noted reduced carbapenem susceptibility without naming the CRE risk; both were rated high. Syndrome-wise empirical recommendations The two models differed most in this domain [Table/Fig-3]. ChatGPT showed complete concordance at three of five syndromes, high at one and moderate at one. Gemini showed complete concordance at one, high at one and moderate at three. For community-acquired UTI in outpatients, the experts and both models recommended fosfomycin or nitrofurantoin for E. coli. ChatGPT also explicitly advised avoiding third-generation cephalosporins and fluoroquinolones (complete). Gemini named the same agents without this explicit caution (high). For complicated UTI/urosepsis and for sepsis of unknown source, ChatGPT matched the expert consensus: a β-lactam/β-lactamase-inhibitor (BL-BLI) regimen, with carbapenem reserved for escalation according to severity or ICU setting (complete for both syndromes). Gemini instead recommended aminoglycoside-based combinations, amikacin plus fosfomycin for complicated UTI and amikacin plus cefoperazone-sulbactam for sepsis, which raised toxicity concerns (moderate for both). The two models diverged in opposite directions for pneumonia. For CAP, the experts and Gemini judged the data insufficient to support an empirical recommendation (complete for Gemini). ChatGPT nevertheless offered a cautious β-lactam option (moderate), its only moderate discordance in the study. For HAP in the ICU, the experts reserved last-line agents for salvage use only. ChatGPT similarly restricted reserve agents, with minor differences in wording (high). Gemini proposed cefoperazone-sulbactam plus tigecycline as an empirical regimen (moderate). Drug-class agreement The experts and both models agreed completely on avoiding aminopenicillins, third-generation cephalosporins and fluoroquinolones [Table/Fig-4]. All three supported selective use of BL-BLI combinations; ChatGPT matched the experts completely, whereas Gemini often paired BL-BLIs with amikacin or tigecycline (high). Disagreement centred on reserve and higher-toxicity agents. ChatGPT matched the experts in restricting carbapenems to escalation and colistin to salvage use (complete for both). It showed high concordance for aminoglycosides (adjunctive use with caution and therapeutic drug monitoring [TDM] advised, against the experts' requirement for TDM) and for tigecycline (strong restriction, against salvage-only use). Gemini considered carbapenems broadly, recommended aminoglycosides frequently, listed colistin as an active option and included tigecycline early. All four positions were rated moderate. Overall, ChatGPT showed complete concordance at six of eight drug-class decision points and high at two. Gemini showed complete concordance at three, high at one and moderate at four. Stewardship and safety safeguards The experts and both models consistently avoided overtreatment in outpatients [Table/Fig-5]. ChatGPT matched the experts in restricting reserve drugs in the ICU and in excluding tigecycline from UTI and sepsis regimens. It emphasised aminoglycoside TDM, which the experts considered mandatory (high). Gemini did not use tigecycline for UTI or sepsis but did not state the exclusion (high). It applied ICU reserve-drug restriction inconsistently, with broad consideration of carbapenems and empirical tigecycline for HAP (moderate), and did not mention aminoglycoside TDM (moderate). In this domain, ChatGPT showed complete concordance at three of four decision points and high at one. Gemini showed complete concordance at one, high at one and moderate at two. The blinded physician reviewer did not identify any AI recommendation that would risk unsafe under-treatment. The concerns identified were overtreatment and toxicity, not insufficient coverage. Differences among the three experts were limited to drug choice within the same class and did not affect overall coverage or stewardship safety. Overall concordance and efficiency Across all 50 model–expert comparisons, 30 (60.0%) showed complete concordance, 10 (20.0%) high concordance and 10 (20.0%) moderate concordance; none showed major discordance [Table/Fig-6]. ChatGPT showed complete concordance at 19 of 25 decision points (76.0%; 95% CI 56.6–88.5), high at 5 (20.0%) and moderate at 1 (4.0%). Its clinically acceptable concordance was 24/25 (96.0%; 95% CI 80.5–99.3). Gemini showed complete concordance at 11 of 25 (44.0%; 95% CI 26.7–62.9), high at 5 (20.0%) and moderate at 9 (36.0%). Its clinically acceptable concordance was 16/25 (64.0%; 95% CI 44.5–79.8). Both models reached clinically acceptable concordance on all eight resistance-pattern decision points. In the other three domains, ChatGPT reached 4/5, 8/8 and 4/4, whereas Gemini reached 2/5, 4/8 and 2/4. ChatGPT was acceptable at nine decision points where Gemini was not, and Gemini at one (CAP) where ChatGPT was not (exploratory exact McNemar test, p=0.021). Given the small number of non-independent decision points, this result should be regarded as hypothesis-generating. Each model completed its analysis in under one minute, compared with about 35 minutes for the expert panel. [Table/Fig-1]: Cumulative antibiogram composition and CLSI M39 data-adequacy assessment. Characteristic Detail / status Surveillance period Calendar year 2024 (annual cumulative antibiogram) Isolates retrieved 1,494 clinically significant bacterial isolates Isolates analysed (post-deduplication) 1,492 non-duplicate isolates (2 repeat isolates excluded) Deduplication rule First isolate per patient per organism per period (CLSI M39); tool-assisted, manually verified Reporting threshold ≥30 non-duplicate isolates per reported organism–antibiotic combination Susceptibility reporting %S / %I / %R complete for reported combinations Stratification Location (OP, non-ICU ward, ICU) and specimen (blood, urine, pus, respiratory) Predominant organisms Gram-negative bacilli; E. coli and K. pneumoniae most frequent Overall CLSI M39 compliance Met (all three experts and both LLMs concurred) CLSI: Clinical and Laboratory Standards Institute; OP: outpatient; ICU: intensive care unit; %S/%I/%R: percentage susceptible/intermediate/resistant; LLM: large language model. [Table/Fig-2]: Resistance-pattern interpretation: concordance of each LLM with the expert consensus (8 decision points). Decision point Expert consensus ChatGPT-5.2 Gemini-3 Concordance: ChatGPT Concordance: Gemini Data adequacy / CLSI M39 Confirmed Confirmed Confirmed Complete Complete Deduplication verification Manually verified Flagged as not verifiable from supplied data Assumed correct Complete High Dominant pathogens E. coli, K. pneumoniae Same Same Complete Complete ESBL proxy phenotype Identified, stewardship-framed Identified, with restraint Identified, numerical Complete Complete CRE burden (ICU) Explicit CRE risk Graded risk flagged Reduced %S noted High High Non-fermenter AMR (XDR) Identified Identified Identified Complete Complete MRSA prevalence Identified Correct Correct Complete Complete VRE Identified Correct Correct Complete Complete Summary (C / H / M) 7 / 1 / 0 6 / 2 / 0 C/H/M: complete/high/moderate concordance; no major discordance occurred. ESBL: extended-spectrum β-lactamase; CRE: carbapenem-resistant Enterobacterales; AMR: antimicrobial resistance; XDR: extensively drug-resistant; MRSA: methicillin-resistant Staphylococcus aureus; VRE: vancomycin-resistant enterococci. [Table/Fig-3]: Syndrome-wise empirical antibiotic recommendations: concordance of each LLM with the expert consensus (5 decision points). Syndrome Expert consensus ChatGPT-5.2 Gemini-3 Concordance: ChatGPT Concordance: Gemini Community-acquired UTI (OP) Fosfomycin / nitrofurantoin Fosfomycin / nitrofurantoin; avoid 3GCs, FQs Fosfomycin / nitrofurantoin Complete High Complicated UTI / urosepsis BL-BLI ± carbapenem escalation BL-BLI; carbapenem restricted (ICU) Amikacin + fosfomycin Complete Moderate Sepsis of unknown source BL-BLI / carbapenem by severity BL-BLI; carbapenem by severity Amikacin + cefoperazone-sulbactam Complete Moderate Community-acquired pneumonia Insufficient data Cautious β-lactam option Insufficient data Moderate Complete Hospital-acquired pneumonia (ICU) Reserve agents as salvage only Reserve agents restricted Cefoperazone-sulbactam + tigecycline High Moderate Summary (C / H / M) 3 / 1 / 1 1 / 1 / 3 UTI: urinary tract infection; OP: outpatient; 3GCs: third-generation cephalosporins; FQs: fluoroquinolones; BL-BLI: β-lactam/β-lactamase inhibitor; ICU: intensive care unit; C/H/M: complete/high/moderate concordance. [Table/Fig-4]: Drug-class positions: concordance of each LLM with the expert consensus (8 decision points). Antibiotic class Expert consensus ChatGPT-5.2 Gemini-3 Concordance: ChatGPT Concordance: Gemini Aminopenicillins Avoid Avoid Avoid Complete Complete Third-generation cephalosporins Avoid Avoid Avoid Complete Complete Fluoroquinolones Avoid Avoid Avoid Complete Complete BL-BLI combinations Selective use Selective use Selective use; often paired with amikacin or tigecycline Complete High Carbapenems Restricted escalation Restricted escalation Broad consideration Complete Moderate Aminoglycosides Adjunct with TDM Adjunct with caution; TDM advised Frequently recommended High Moderate Colistin Salvage only Salvage only Listed as active option Complete Moderate Tigecycline Salvage only Strong restriction Early inclusion High Moderate Summary (C / H / M) 6 / 2 / 0 3 / 1 / 4 BL-BLI: β-lactam/β-lactamase inhibitor; TDM: therapeutic drug monitoring; C/H/M: complete/high/moderate concordance. [Table/Fig-5]: Stewardship and safety safeguards: concordance of each LLM with the expert consensus (4 decision points). Safeguard Expert consensus ChatGPT-5.2 Gemini-3 Concordance: ChatGPT Concordance: Gemini Avoidance of outpatient overtreatment Explicit Explicit Explicit Complete Complete ICU reserve-drug restriction Strict Strict Inconsistent (carbapenems broadly; tigecycline in HAP) Complete Moderate Tigecycline exclusion (UTI/sepsis) Explicit Explicit Not used, but exclusion not stated Complete High Aminoglycoside TDM Mandatory Emphasised Not mentioned High Moderate Summary (C / H / M) 3 / 1 / 0 1 / 1 / 2 ICU: intensive care unit; HAP: hospital-acquired pneumonia; UTI: urinary tract infection; TDM: therapeutic drug monitoring; C/H/M: complete/high/moderate concordance. [Table/Fig-6]: Overall concordance of each LLM with the expert consensus across 25 decision points. Metric ChatGPT-5.2 (n=25) Gemini-3 (n=25) Both models (n=50) Complete concordance, n (%) 19 (76.0) 11 (44.0) 30 (60.0) High concordance (minor nuance), n (%) 5 (20.0) 5 (20.0) 10 (20.0) Moderate concordance (clinically relevant), n (%) 1 (4.0) 9 (36.0) 10 (20.0) Major discordance, n (%) 0 0 0 Complete concordance, % (95% CI) 76.0 (56.6–88.5) 44.0 (26.7–62.9) – Clinically acceptable (complete + high), n (%; 95% CI) 24 (96.0; 80.5–99.3) 16 (64.0; 44.5–79.8) 40 (80.0) Acceptable by domain: resistance patterns (n=8) 8/8 8/8 16/16 Acceptable by domain: syndrome-wise recommendations (n=5) 4/5 2/5 6/10 Acceptable by domain: drug-class positions (n=8) 8/8 4/8 12/16 Acceptable by domain: stewardship safeguards (n=4) 4/4 2/4 6/8 Analysis time (expert panel: ~35 min) <1 min <1 min – CI: Wilson score 95% confidence interval. Comparison of clinically acceptable concordance between models (discordant decision points 9 vs 1): exact McNemar test p=0.021, exploratory only, because the decision points derive from a single antibiogram and a single model run and are not fully independent.
DISCUSSION
In this proof-of-concept comparison, two general-purpose LLMs, constrained by a standardised, stewardship-aware prompt, reproduced expert microbiologist interpretation of a real, CLSI M39-compliant cumulative antibiogram with no major discordance. ChatGPT reached clinically acceptable concordance at 96% of decision points and complete concordance at 76%; Gemini reached 64% and 44%. None of the discordances would have led to unsafe under-treatment, which is a basic safety requirement for any AI-assisted stewardship tool. The gap between the two models lay mainly in restraint with reserve antibiotics and in toxicity awareness. Both models reliably identified the overall resistance picture: Gram-negative predominance, ESBL proxies, ICU carbapenem erosion, XDR Acinetobacter, MRSA and VRE. All resistance-pattern decision points were acceptable for both models, consistent with evidence that AI systems perform well on bounded pattern-recognition tasks [10,17]. The clinically important divergence was therefore not in reading the resistance data but in stewardship judgement. Gemini tended to escalate on the basis of susceptibility percentages alone. It favoured aminoglycoside-based regimens, considered carbapenems broadly, listed colistin as an active option and included tigecycline early. ChatGPT more consistently followed expert reasoning: it restricted carbapenem escalation, kept tigecycline out of UTI and sepsis regimens and advised aminoglycoside monitoring. These differences matter clinically. Tigecycline achieves low urinary and serum concentrations and has been associated with increased all-cause mortality [18–20], and aminoglycosides given without TDM increase the risk of nephrotoxicity [21]. The experts' advantage here reflects contextual pharmacokinetic–pharmacodynamic reasoning that general-purpose LLMs do not reliably apply on their own. This is consistent with earlier reports that LLM outputs can appear guideline-aligned while underweighting toxicity and local stewardship norms unless the prompt constrains them explicitly [6,9]. Pneumonia showed the two models erring in opposite directions. For CAP, the experts and Gemini judged the antibiogram insufficient to support empirical guidance, whereas ChatGPT offered a cautious β-lactam option. This was ChatGPT's only moderate discordance, and it involved over-reach rather than unsafe therapy. Respiratory isolates are prone to colonisation bias, and classic community respiratory pathogens are often under-represented. Numerical adequacy therefore does not guarantee clinical validity for every syndrome, a point CLSI M39 itself makes [3]. For HAP, by contrast, Gemini proposed a reserve-agent regimen that the experts would have kept for salvage use. Agreement was strongest for UTI, where the specimen-level data were robust. AI-assisted antibiogram interpretation therefore appears to work best where the microbiological signal is strong, and least well where clinical and host factors must also be considered. The time difference was large: under one minute for each model versus about 35 minutes for expert review. This suggests a role for LLMs in rapid first-pass analysis, trend surveillance and policy drafting where microbiologists are in short supply [22]. Speed, however, cannot replace clinical accountability, especially for ICU and reserve-antibiotic decisions. Our findings support a human-in-the-loop arrangement. AI would provide a standardised preliminary analysis and flag resistance trends or unsafe options, while microbiologists retain the final decision on empirical policy, reserve-drug use and toxicity oversight. This is consistent with current guidance on early-stage, supervised clinical evaluation of AI decision-support systems [15,17]. Limitation(s) This study has several limitations. It was a single-centre evaluation of one annual antibiogram, so its generalisability to other resistance ecologies is untested. MIC-level data, molecular data and clinical outcomes were not available, so there was no external clinical gold standard; we assessed concordance with expert consensus rather than diagnostic accuracy. Each model was queried only once, and LLM outputs are stochastic and prompt-dependent, so test–retest reproducibility was not assessed. The 25 decision points are few and not statistically independent, which is reflected in the wide CIs. For this reason, agreement is presented descriptively and the between-model test is exploratory. Finally, performance depended on the constraints built into the prompt, and unconstrained real-world use may perform less well. This should therefore be regarded as a hypothesis-generating study. Future work should include multicentre validation across different AMR ecologies; multi-run assessment of reproducibility and prompt sensitivity; incorporation of MIC distributions and pharmacodynamic targets; prospective evaluation within live stewardship workflows; and development of domain-specific, safety-constrained microbiology models.
CONCLUSION
Constrained general-purpose LLMs reproduced expert interpretation of a hospital cumulative antibiogram without major discordance and without any recommendation that risked unsafe under-treatment. ChatGPT matched the expert consensus at 96% of decision points (76% completely); Gemini matched at 64% (44% completely), with shortfalls concentrated in stewardship and toxicity nuance. These exploratory, single-centre findings suggest that LLMs work best as human-in-the-loop decision-support adjuncts, useful for fast first-pass interpretation where resources are limited, rather than as replacements for expert microbiologists. Expert judgement remains essential for ICU, reserve-antibiotic and toxicity-sensitive decisions. Multicentre, multi-run validation is needed before clinical adoption.
REFERENCES
1. Murray CJL, Ikuta KS, Sharara F, Swetschinski L, Robles Aguilar G, Gray A, et al. Global burden of bacterial antimicrobial resistance in 2019: a systematic analysis. Lancet. 2022;399(10325):629-55. 2. GBD 2021 Antimicrobial Resistance Collaborators. Global burden of bacterial antimicrobial resistance 1990-2021: a systematic analysis with forecasts to 2050. Lancet. 2024;404(10459):1199-226. 3. Clinical and Laboratory Standards Institute. Analysis and presentation of cumulative antimicrobial susceptibility test data. 5th ed. CLSI guideline M39. Wayne (PA): CLSI; 2022. 4. Truong WR, Hidayat L, Bolaris MA, Nguyen L, Yamaki J. The antibiogram: key considerations for its development and utilization. JAC Antimicrob Resist. 2021;3(2):dlab060. 5. Khatri D, Freeman C, Falconer N, de Camargo Catapan S, Gray LC, Paterson DL. Clinical impact of antibiograms as an intervention to optimize antimicrobial prescribing and patient outcomes: a systematic review. Am J Infect Control. 2024;52(1):107-22. 6. Antonie NI, Gheorghe G, Ionescu VA, Tiucă LC, Diaconu CC. The role of ChatGPT and AI chatbots in optimizing antibiotic therapy: a comprehensive narrative review. Antibiotics (Basel). 2025;14(1):60. 7. Hu L, Xu X, Zhuang Y, Lin Y, Xu M, Wu X, et al. Pre-trained ChatGPT for report generation in automated microbial identification and antibiotic susceptibility testing systems. Sci Rep. 2025;15:36283. 8. Rao A, Pang M, Kim J, Kamineni M, Lie W, Prasad AK, et al. Assessing the utility of ChatGPT throughout the entire clinical workflow: development and usability study. J Med Internet Res. 2023;25:e48659. 9. De Vito A, Geremia N, Marino A, Bavaro DF, Caruana G, Meschiari M, et al. Assessing ChatGPT's theoretical knowledge and prescriptive accuracy in bacterial infections: a comparative study with infectious diseases residents and specialists. Infection. 2025;53(3):873-81. 10. Nagendran M, Chen Y, Lovejoy CA, Gordon AC, Komorowski M, Harvey H, et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies. BMJ. 2020;368:m689. 11. Sakagianni A, Koufopoulou C, Feretzakis G, Kalles D, Verykios VS, Myrianthefs P, et al. Using machine learning to predict antimicrobial resistance—a literature review. Antibiotics (Basel). 2023;12(3):452. 12. Pinto-de-Sá R, Sousa-Pinto B, Costa-de-Oliveira S. Brave new world of artificial intelligence: its use in antimicrobial stewardship—a systematic review. Antibiotics (Basel). 2024;13(4):307. 13. OpenAI. ChatGPT (GPT-5.2) [large language model]. San Francisco (CA): OpenAI; 2025 [cited 2025 Dec 28]. Available from: https://chatgpt.com/ 14. Google. Gemini 3 [large language model]. Mountain View (CA): Google LLC; 2025 [cited 2025 Dec 28]. Available from: https://gemini.google.com/ 15. Vasey B, Nagendran M, Campbell B, Clifton DA, Collins GS, Denaxas S, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. BMJ. 2022;377:e070904. 16. Gallifant J, Afshar M, Ameen S, Aphinyanaphongs Y, Chen S, Cacciamani G, et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nat Med. 2025;31(1):60-9. 17. Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44-56. 18. US Food and Drug Administration. FDA Drug Safety Communication: FDA warns of increased risk of death with IV antibacterial Tygacil (tigecycline) and approves new boxed warning. Silver Spring (MD): FDA; 2013 Sep 27. 19. Prasad P, Sun J, Danner RL, Natanson C. Excess deaths associated with tigecycline after approval based on noninferiority trials. Clin Infect Dis. 2012;54(12):1699-709. 20. Yahav D, Lador A, Paul M, Leibovici L. Efficacy and safety of tigecycline: a systematic review and meta-analysis. J Antimicrob Chemother. 2011;66(9):1963-71. 21. Rybak MJ, Abate BJ, Kang SL, Ruffing MJ, Lerner SA, Drusano GL. Prospective evaluation of the effect of an aminoglycoside dosing regimen on rates of observed nephrotoxicity and ototoxicity. Antimicrob Agents Chemother. 1999;43(7):1549-55. 22. Langford BJ, Branch-Elliman W, Nori P, Marra AR, Bearman G. Confronting the disruption of the infectious diseases workforce by artificial intelligence: what this means for us and what we can do about it. Open Forum Infect Dis. 2024;11(3):ofae053.
Recommended Articles
Original Article
Prevalence and Risk Factors of Malnutrition Among Under-Five Children Attending a Tertiary Care Hospital, Zydus medical college and Hospital, Dahod, Gujarat
Published: 10/10/2026
Original Article
Comparative Evaluation of Dietary Modification Versus Proton Pump Inhibitor Therapy Versus Their Combination in the Management of Laryngopharyngeal Reflux: A Prospective Observational Study
...
Published: 10/10/2026
Original Article
Sleep Disturbances as a Manifestation of Occupational Stress and Their Association with Psychological Morbidity among Healthcare Professionals: A Cross-Sectional Study
Published: 07/12/2020
Original Article
Clinical Profile and Outcomes of Children Undergoing Surgery for Acute Appendicitis in Children Below 12 Years
...
Published: 09/10/2026
Chat on WhatsApp
© Copyright Journal of Contemporary Clinical Practice