Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors
Search in posts
Search in pages
Filter by Categories
Case Report
Clinical Images
Research Article
Review Article
Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors
Search in posts
Search in pages
Filter by Categories
Case Report
Clinical Images
Research Article
Review Article
View/Download PDF

Translate this page into:

Research Article
2026
:5;
100728
doi:
10.1016/j.jorep.2025.100728

Comparative evaluation of LLMs in orthopedic surgery

Department of Orthopedic Surgery, New Jersey Medical School, 185 S Orange Ave, Newark, NJ, 07103, USA
FAU Charles E. Schmidt College of Medicine, 777 Glades Road BC-71, Boca Raton, FL, 33431, USA
Hackensack Meridian School of Medicine, 123 Metro Blvd, Nutley, NJ, 07110, USA

⁎Corresponding author: Gnaneswar Chundi. gc554@njms.rutgers.edu

Disclaimer:
This article was originally published by Reed Elsevier India Pvt. Ltd. and was migrated to Scientific Scholar after the change of Publisher.

Abstract

Abstract

This study aimed to evaluate the performance of leading large language models (LLMs) in orthopedic surgery, with a focus on diagnostic accuracy, radiographic interpretation, subspecialty-specific performance, consistency, and gender bias. Unlike previous investigations that focused solely on ChatGPT, we assessed Claude-3-Sonnet, GPT-4o, Gemini-1.5, and Meta's LLaMA Vision-Instruct models against Orthopaedic In-Training Examination (OITE) questions.

We tested each LLM on 2906 multiple-choice questions from the OITE question bank. Questions were categorized by subspecialty and by presence of images. Models provided answer choices, confidence scores (1–4), and justifications. Each question was administered three times to evaluate consistency. Statistical analyses included Z-tests, t-tests, ANOVA, chi-square tests, and logistic regression. Rasch transformation enabled comparison to PGY-1 and PGY-5 resident performance. Gender bias was evaluated based on performance differences across gender-specific cases.

GPT-4o achieved the highest accuracy (72.8 %), outperforming all other models. Performance improved with larger model sizes across all vendors. All models showed diminished accuracy on image-based questions (mean 10.9 % lower, p < 0.001). Confidence scores and triplicate agreement were associated with accuracy; their combination yielded the most reliable outputs (up to 83.4 % accuracy). Several models exhibited worse performance on female patient questions, indicating possible gender bias.

LLMs demonstrate varying levels of accuracy in orthopaedic applications, with GPT-4o approaching PGY-5 performance, strictly in the context of standardized exam questions. Current models underperform on image-based questions and exhibit gender bias. Response confidence and consistency can flag reliable outputs. Continued development, bias mitigation, and fine-tuning with diverse, domain-specific datasets are essential for integration of LLMs into orthopaedic education and practice.

Keywords

LLM
OITE
Artificial intelligence
Subspecialty
Chatbots
Machine learning
1

1 Background

Artificial intelligence (AI) is rapidly transforming the field of orthopaedic surgery, offering innovative solutions that extend beyond traditional diagnostic and treatment paradigms. AI-based tools, particularly large language models (LLMs) such as ChatGPT, Gemini, and Claude, have evolved into sophisticated multimodal systems capable of interpreting complex medical data, including both textual and visual inputs.1 In orthopaedics, AI applications range from diagnosing fractures with sensitivity rates as high as 92.3 %, to predicting congenital hip dysplasia, meniscus tears, and osteoarthritis.2–5 However, the application and reliability of AI-driven models, beyond ChatGPT, in orthopaedic education and clinical practice remain underexplored. There is a lack of global consensus on the efficacy of AI in clinical orthopaedic applications. For instance, while some studies highlight the success of major LLMs in triaging common sports medicine complaints of the knee joint, others criticize their ability to identify suitable candidates for surgical intervention.6–8 This inconsistency underscores a gap in understanding the true utility of LLMs in orthopaedics, particularly those beyond ChatGPT.

The present study aims to fill this gap by conducting an evaluation and comparison of current state-of-the-art LLMs, including Claude-3-sonnet, GPT-4o, Gemini-1.5, and Meta's LLaMA Vision-Instruct models, using standardized questions from the Orthopaedic In-Training Examination (OITE) question bank. This analysis focuses not only on overall accuracy but also on crucial capabilities such as interpreting radiographic images, evaluating performance across subspecialties and assessing response consistency and confidence levels. This study also aims to compare the performance of AI model responses to gender-related questions to detect potential gender bias. Furthermore, unlike prior studies that primarily assessed the performance of a single popular LLM—ChatGPT—our study aims to compare the performances of all major LLMs, including Gemini, Claude, ChatGPT, and LLaMA, to one another and to the standard average performance of orthopaedic residents throughout various intervals of training.

2

2 Methods

2.1

2.1 Aims

The primary aim was to compare LLM performance on multiple-choice questions from the American Academy of Orthopaedic Surgeons (AAOS) ResStudy Question Bank, a standardized orthopaedic resident resource for the Orthopaedic In-Training Examination (OITE). Secondary aims included evaluating performance variations across orthopaedic subspecialties, differences between questions with and without images, strategies for reliability assessment via confidence scoring and response consistency, and identifying potential model biases.

2.2

2.2 LLMs used

The LLMs in this study fell into two categories: centrally hosted and local models. Models were chosen for their ability to process both text and images. Hosted models included Claude-3-sonnet [Anthropic], GPT-4o and GPT-4o-mini [OpenAI], and Gemini-1.5-pro and Gemini-1.5-flash [Google], all with multimodal capabilities. Local models were two Llama-3.2-Vision-Instruct versions [Meta]—11B and 90B parameters.

Given the testing scope (8000+ prompts per model), an automated prompting system was built using each model's API, avoiding manual input. Testing occurred from November to December 2024. Each model received standardized prompts containing the question, related images (if any), and four answer choices (A, B, C, D).

All models were instructed to provide three specific outputs for each question: a single letter answer selection (A, B, C, or D), a confidence score on a scale of 1–4, and an explanation justifying the chosen answer. The confidence scoring system was designed to reflect the model's self-assessed certainty, where 1 indicated a random guess, 2 indicated low confidence, 3 indicated moderate confidence, and 4 indicated maximal confidence. All prompts were designed to maintain consistency across models.

2.3

2.3 Testing set

The testing set included 2906 standardized questions from the established question bank. Questions were classified as text-only or containing images to assess visual reasoning. Pre-existing subspecialty tags (adult reconstruction, spine, basic science, foot/ankle, hand/wrist, oncology, pediatrics, shoulder/elbow, sports, trauma) enabled discipline-specific performance analysis. Correct answers were based on the official key. To assess response consistency, each model answered every question three times using identical prompts but in separate environments to ensure independence.

Finally, to assess potential gender bias, questions were analyzed using regular expressions to identify gendered terms. Questions containing masculine terms (e.g., "boy," "male," "man," "he," "him," "his") were classified as male, while those with feminine terms (e.g., "girl," "female," "woman," "she," "her," "hers") were labeled female. Questions containing both male and female terms were marked as "both," while those without gendered terms were labeled "unspecified." This classification enabled analysis of model performance differences between male and female patient cases.

2.4

2.4 Statistical analysis

Accuracy differences between models were assessed with Z tests for proportions. One-way ANOVA evaluated performance variation across disciplines. Pearson's correlation measured the relationship between overall and discipline-specific accuracy. Paired t-tests compared performance on image vs. non-image questions. Gender bias was tested using chi-square with FDR correction.

Model accuracies were mapped to PGY-1 and PGY-5 ACGME scores using Rasch linear transformation. Percentiles were calculated via the cumulative normal distribution function, enabling direct comparison with resident scores.

For confidence analysis, we assessed how well models could predict their own accuracy using confidence scores. Chi-square tests determined if higher confidence scores were associated with better performance. We then compared accuracy between maximum confidence responses to overall accuracy using proportion tests. The same test was used to evaluate if responses with triplicate agreement (same answer given three times) were more accurate than average.

Finally, logistic regression was used to analyze interactions between confidence scores and triplicate agreement to evaluate the synergistic effects between high confidence and triplicate agreement on accuracy.

Visualizations were generated using the Python seaborn library. All statistical analyses were performed using Python's SciPy, NumPy, scikit-learn, and pandas libraries. Statistical significance was set at p < 0.05.

3

3 Results

3.1

3.1 Overall performance

GPT-4o achieved the highest overall accuracy, while LLaMA-11B had the lowest. Across all model families—such as GPT-4o vs. GPT-4o-mini, Gemini-Pro vs. Gemini-Flash, and LLaMA-90B vs. LLaMA-11B—larger models consistently outperformed their smaller counterparts (all p < 0.001; see Table 1).

Table 1 Model performance distribution.
Model Accuracy (95 % CI)
Claude-3-Sonnet 0.690 (0.682–0.698)
Gemini-1.5-Flash 0.518 (0.512–0.524)
Gemini-1.5-Pro 0.618 (0.611–0.625)
GPT-4o 0.728 (0.720–0.736)
GPT-4o-mini 0.543 (0.537–0.549)
Meta-LLaMA-3.2–11B 0.447 (0.441–0.453)
Meta-LLaMA-3.2–90B 0.600 (0.592–0.608)
3.2

3.2 Performance comparison to trainees

Based on 2022 OITE benchmarks, GPT-4o achieved performance at the 99th percentile relative to PGY-1 residents and aligned with the 48th percentile among PGY-5 residents, indicating strong exam-level proficiency. Claude-3-Sonnet also performed well against PGY-1 standards (98th percentile) but dropped to the 22nd percentile at the PGY-5 level. Gemini-Pro, LLaMA-90B, and all other models performed well below the PGY-5 median, with most scoring under the 5th percentile. These comparisons, derived via Rasch-transformed percentiles, reflect multiple-choice exam alignment only and do not imply clinical equivalency (Table 2).

Table 2 Model Comparison vs. Orthopaedic Residents – Values depict model performance as percentiles based on PGY-1 and PGY-5 performance on 2022 OITE.
Model Accuracy PGY-1 Percentile PGY-5 Percentile
GPT-4o 72.8 % 99.7th 48.5th
Claude-3-sonnet 69.0 % 98.5th 22.0th
Gemini-1.5-pro 61.8 % 85.5th 1.5th
Llama-3.2–90b 60.0 % 78.1st 0.6th
GPT-4o-mini 54.3 % 45.7th <0.1st
Gemini-1.5-flash 51.8 % 31.0th <0.1st
Llama-3.2–11b 44.7 % 5.5th <0.1st
3.3

3.3 Performance by discipline

Analysis of model performance by discipline revealed a variation by subject area across all models (ANOVA, F = 9.29, p < 0.001). A strong positive correlation (Pearsons’ r = 0.998, p < 0.001) was observed between models' overall performance and their discipline-specific performance, indicating that models maintaining high general accuracy consistently performed well across subfields. All models had the best performance on questions categorized in the foundational topic of basic science and the worst performance on questions tagged in the trauma subfield (Fig. 1).

Heatmap of Model Performance Stratified by Discipline: Heatmap of model performance across medical disciplines, with accuracy scores ranging from 0.40 (light) to 0.80 (dark).
Fig. 1 Heatmap of Model Performance Stratified by Discipline: Heatmap of model performance across medical disciplines, with accuracy scores ranging from 0.40 (light) to 0.80 (dark).
3.4

3.4 Performance by image vs No image

An analysis of questions that contained images such as radiographs versus questions that contained text only revealed statistically significant lower performance on image-based questions compared to non-image questions across all models. The average decrease in performance was 10.9 % ± 1.6 % (t = 17.77, p < 0.001). GPT-4o maintained the highest accuracy in both categories (78.7 % no-image, 68.4 % with-image), while the 11B LLaMA model showed the lowest performance (49.7 % no-image, 41.0 % with-image). (Fig. 2).

Comparison of Model Performance on Questions with and without Images: Bars depict models' performance on questions stratified by image vs. no-image. Error bars indicate standard error of measurement.
Fig. 2 Comparison of Model Performance on Questions with and without Images: Bars depict models' performance on questions stratified by image vs. no-image. Error bars indicate standard error of measurement.
3.5

3.5 Confidence self-assessment

Additional analysis was performed to assess methods in which a user may be able to ensure accurate model responses, such as a maximum self-assessed confidence score or an agreement of responses across three independent trials (triplicate agreement). The first analysis was done using the confidence scores provided by every model for each question. All models were able to output confidence scores that were associated with accuracy (chi-square test, p < 0.001), except for LLaMA-11B (p = 0.6359) and Gemini-Flash (p = 0.2607). All of these models had improved accuracy in question responses when they reported a maximal confidence of 4/4 compared with overall accuracy (Z-test, all p < 0.001). However, it was noted that models tended to report higher confidence uniformly, with at least 95 % of responses returning scores of 3 or 4 in each model, and only 1 model recording a confidence score of 1 at all (Fig. 3).

Model Accuracy vs Confidence Level: Lines depict the relationship between model confidence scores (1–4) and accuracy (%). Point sizes indicate sample sizes (500–2500 questions).
Fig. 3 Model Accuracy vs Confidence Level: Lines depict the relationship between model confidence scores (1–4) and accuracy (%). Point sizes indicate sample sizes (500–2500 questions).
3.6

3.6 Combined strategy

The second set of analyses investigated triplicate agreement as a way for users to evaluate their level of confidence in the model's response (Fig. 4). A significant proportion of cases failed to produce triplicate agreement: 44.8 % for LLaMA 11B, 29.8 % for LLaMA 90B, 23.7 % for Claude-Sonnet, 6.8 % for Gemini-Flash, 5.9 % for Gemini-Pro, 5.3 % for GPT-4o-mini, and 3.0 % for GPT-4o. There was a strong positive correlation between overall model performance and the proportion of queries that resulted in a correct triplicate agreement (r = 0.914, p = 0.004). Furthermore, when there is a triplicate agreement, a higher accuracy than baseline is seen in all models (test of proportions, all p < 0.001). The largest increase was seen in LLaMA 90B (60.3 %–70.2 %) (Fig. 5).

Model Consistency and Accuracy: Bars represent distribution of model responses across three repeated prompts showing correct answers (0–3) and consistency patterns (colored bars). Purple bars indicate identical responses, the lighter green shows two matching responses, and dark green represents completely different responses across attempts.
Fig. 4 Model Consistency and Accuracy: Bars represent distribution of model responses across three repeated prompts showing correct answers (0–3) and consistency patterns (colored bars). Purple bars indicate identical responses, the lighter green shows two matching responses, and dark green represents completely different responses across attempts.
Model Performance by Subgroup. Bars depict model accuracy in scenarios of maximal confidence (confidence = 4.0), consistency of responses across three trials (triplicate agreement), or in scenarios of both occurring in tandem. Red dashed line indicates overall model accuracy.
Fig. 5 Model Performance by Subgroup. Bars depict model accuracy in scenarios of maximal confidence (confidence = 4.0), consistency of responses across three trials (triplicate agreement), or in scenarios of both occurring in tandem. Red dashed line indicates overall model accuracy.

On the other hand, a significant proportion of queries were observed to have incorrect triplicate agreement (25.1 % for LLaMA 11B, 20.9 % LLaMA 90B, 16.9 % for Claude-Sonnet, 25.3 % for GPT-4o, 41.9 % for GPT-4o-mini, 43.5 % for Gemini-Flash, and 34.2 % for Gemini-Pro). These results represent "fixed false beliefs," which means that the LLM consistently produced the same wrong answer across multiple attempts (Fig. 4).

Finally, we evaluated the combinations of these factors (selected model, self-assessed confidence, and triplicate agreement) in identifying cases where the model output can be considered more reliable. The combination of these 3 factors successfully identifies high-performing subgroups of queries in GPT-4o and Claude-Sonnet, with accuracies of 83.4 % (n = 1592) and 79.6 % (n = 1899) respectively. In all models, triplicate agreement was observed to result in more accuracy individually than maximal confidence. However, a significant interaction between the confidence score and triplicate agreement was observed in all models (all interaction coefficients positive, all p < 0.001). The positive interaction coefficients indicate that high confidence combined with triplicate agreement had a synergistic effect on accuracy, indicating that the impact was greater than what would be expected from simply adding their individual effects (Fig. 5).

3.7

3.7 Female vs. male

The analysis of gender biases in model responses revealed that all models performed worse in questions with female patients, except for LLaMA-11B. Gemini-1.5-flash, Claude-3-sonnet, and Meta's Llama-3.2–90B all showed significant gender effects, with statistically significant worse performances in female-oriented questions (Fig. 6).

Model Performance Stratified by Gender: Bars depict accuracy on questions related to male (blue) or female (pink) patients. Asterisks (*) indicate statistically significant gender performance differences.
Fig. 6 Model Performance Stratified by Gender: Bars depict accuracy on questions related to male (blue) or female (pink) patients. Asterisks (*) indicate statistically significant gender performance differences.
3.8

3.8 Qualitative review

To better understand the limitations of LLM performance, we conducted a qualitative review of incorrect model responses based on model explanations of answer choices. Common errors included misinterpretation of complex clinical presentations, overlooking subtle radiographic findings such as minimally displaced fractures or early degenerative changes, and confusion between similarly worded diagnostic options. Models frequently misclassified questions involving nuanced clinical context or patient-specific decision-making factors, particularly in subspecialties such as trauma and pediatrics. Additionally, errors commonly arose from the models’ inability to correctly prioritize multiple abnormal findings presented simultaneously in image-based questions.

4

4 Discussion

The primary objective of this study was to evaluate the foundational orthopaedic knowledge of multiple large language models (LLMs) as represented by their performance on Orthopaedic In-Training Examination (OITE) multiple-choice questions. However, it is important to emphasize that orthopaedic clinical decision-making is inherently complex, context-rich, and extends beyond what can be captured solely through MCQ-based assessments. Thus, while these results offer insight into the foundational knowledge base and educational potential of these models, they do not directly equate to clinical competence or real-world decision-making ability.

Our findings indicate significant variability in model performance, with OpenAI's GPT-4o emerging as the top performer, with Meta's Llama-3.2-11B-Vision-Instruct demonstrating the lowest accuracy. This variation may be attributed to differences in training data quality, model architecture, and the extent to which each model incorporates domain-specific medical literature.9 Additionally, OpenAI models likely benefited from reinforcement learning with human feedback (RLHF),10 which may have fine-tuned their responses to align more closely with medical reasoning and guidelines. Furthermore, local models such as LLaMA were trained on fewer parameters which may negatively impact performance across a broad range of clinical scenarios such as those presented by the OITE practice question set.

A key observation from our analysis is that all models performed significantly worse on image-based questions compared to text-based ones, highlighting the challenges of processing complex visual data. Orthopaedics is a visually intensive specialty that relies heavily on interpreting radiographic and clinical images, requiring an advanced understanding of spatial relationships and subtle abnormalities.11,12 The poorer performance in this area may be due to a lack of sufficient annotated imaging data in the models' training datasets, limitations in their visual recognition capabilities compared to human perception, and the complexity of orthopaedic images that often require contextual clinical correlation.13 Additionally, LLMs may struggle with differences in image quality, minute variations in positioning, and the presence of artifacts that can obscure critical diagnostic features. Thus, the observed 10.9 % decrease in accuracy for image-based questions may be attributed to multiple factors and future work should systematically investigate these factors, ideally incorporating standardized imaging datasets and direct comparisons to expert human interpretation.

Across various orthopaedic subdisciplines, our results reveal that model performance is highest in basic sciences, while notable gaps persist in more complex areas such as trauma and pediatrics. This discrepancy may stem from several factors, including the broader availability of structured and standardized information on basic sciences, which is more commonly found in textbooks and research articles that these models are trained on.14 In contrast, trauma and pediatric orthopaedics involve more variability in presentation, complex decision-making based on evolving clinical scenarios, and fewer standardized guidelines, making it more challenging for models to generalize effectively.15,16 Moreover, the dynamic nature of trauma cases,17 which often require individualized management approaches based on patient-specific factors, presents a unique challenge that AI models may not yet be equipped to handle.

Model confidence levels exhibited a general positive association with accuracy; however, no consistent trend was observed across different models. While higher self-appraised confidence often coincided with correct responses, inconsistencies in confidence reporting were noted, particularly with models that tended to overestimate their capabilities. Several potential factors could contribute to this phenomenon, including overfitting to certain types of data,18 limitations in the models' calibration algorithms,19 and a lack of proper validation datasets that reflect real-world clinical complexity.20 Interestingly, none of the models, except LLaMA 90b, reported confidence levels below 2.0, suggesting a tendency to present responses with unwarranted certainty. This overconfidence may be influenced by the models’ training process, which often prioritizes fluency and coherence over accuracy, potentially leading to the generation of plausible-sounding but incorrect answers.21–23

Our findings also highlight the importance of response consistency, with triplicate agreement being a stronger predictor of accuracy than confidence levels alone. This suggests that repeated agreement in responses may serve as a more reliable indicator of accuracy compared to self-appraised confidence. The improved accuracy with repeated responses may indicate that certain questions align well with the model's training data, while questions producing inconsistent responses may highlight areas where models lack sufficient understanding or have conflicting information within their knowledge base. However, this finding highlights the risk of models reinforcing incorrect information if a bias exists within the training data or if the same incorrect response is consistently generated.

Notably, our analysis revealed a synergistic interaction between confidence scores and triplicate agreement across all models, where their combined effect on accuracy exceeded the sum of their individual contributions. This effect suggests complex interactions between model confidence and response consistency that warrant further investigation. One possible mechanism is that high confidence combined with consistent responses indicates questions that fall within well-defined, thoroughly trained knowledge domains, while the absence of either factor may signal edge cases or areas of incomplete learning. However, the concerning prevalence of "fixed false beliefs" across all models highlights potential systematic errors in model knowledge and training methods. These consistent incorrect responses likely stem from biases or errors in training data that become deeply embedded in the model's knowledge representation. The variation in fixed false belief rates across models suggests that architectural improvements and refined training methodologies may help mitigate this issue. Future research should explore whether these reliability indicators remain stable across different task types and knowledge domains, whether they can be effectively leveraged in deployment settings to automatically flag responses requiring human verification, and how to better detect and correct fixed false beliefs during both training and inference.

A key finding of this analysis is the detection of gender bias within AI-generated responses. Questions pertaining to male patients yielded higher accuracy rates compared to those involving female patients in Gemini, Claude, and Llama 90b. The complexity of female-specific injury patterns, such as differences in bone density, hormonal influences, and anatomical variations, may contribute to the reduced accuracy in female-related cases, as well as potential biases in the training data available to these models.24–27 The bias could also stem from a historical underrepresentation of women in clinical trials and research studies,28 leading to a lack of sufficient data for AI models to learn from. Addressing such biases will be crucial for ensuring the equitable application of AI in orthopaedic education and practice. However, given the use of a keyword-based classification method and lack of expert validation or difficulty adjustment, these findings should be interpreted as exploratory rather than conclusive evidence of bias.

The qualitative analysis of model errors provided important insights into the limitations of current LLM capabilities. Errors frequently stemmed from insufficient understanding of complex clinical contexts and challenges in accurately interpreting subtle radiographic findings. These findings underscore that despite high numerical accuracy in foundational knowledge assessments, significant gaps remain in nuanced clinical judgment and image interpretation skills. Future research should include detailed qualitative error analysis to systematically identify and address specific weaknesses in LLM clinical reasoning.

While our study provides valuable insights into the capabilities and shortcomings of current LLMs in orthopaedics, several limitations must be acknowledged. The first is the absence of expert orthopaedic surgeon review of model-generated explanations or answer selections. As a result, we were unable to evaluate the clinical plausibility or educational quality of the responses beyond answer accuracy. Future studies should incorporate blinded expert review to assess the explanatory depth, clinical reasoning, and potential utility of LLM outputs in educational or decision-support settings.

A further limitation of this study is the potential for data leakage due to the public availability of the OITE question bank, despite the existence of a paywall. Since large language models are trained on extensive public datasets, it is possible that models had prior exposure to some of these questions, potentially inflating their accuracy. This represents an important confounder that must be considered when interpreting findings. Future studies aiming to provide more robust benchmarks should employ non-public, clinically validated question sets to minimize the risk of data leakage.

Importantly, our comparison of LLM performance to orthopaedic residents' scores must be interpreted cautiously. The percentile comparison used in this study is strictly limited to accuracy on standardized multiple-choice questions and does not reflect real-world clinical decision-making skills, contextual judgment, or patient-management abilities. The OITE is designed to assess resident knowledge, not clinical acumen or decision-making in real-world scenarios. Therefore, while our findings provide a benchmark for LLM performance on standardized educational material, they should not be interpreted as indicative of clinical competence or readiness for deployment in practice settings. Future studies should assess LLM performance using real patient cases, vignettes, and clinician-adjudicated scenarios to enhance clinical validity.

In relation to the gender bias analysis, a limitation was its reliance on a keyword-matching approach to classify questions as male- or female-oriented. While this method enabled large-scale categorization, it may oversimplify complex clinical narratives and potentially misclassify questions involving nuanced gender-related contexts. Additionally, no clinical expert validation was performed to confirm these classifications. Future studies should incorporate expert adjudication of gender attribution and examine intersectional factors that may influence model bias.

Another notable limitation is the lack of rigorous methodology for assessing visual interpretation skills. The radiographic images included in our test set were not standardized for quality or complexity, and we did not explicitly verify the LLMs’ capability to perform visual feature extraction or conduct human benchmarking to calibrate image-question difficulty. Future research should incorporate standardized imaging datasets, human expert validation, and formal evaluations of LLM image-processing abilities to provide more robust and clinically meaningful assessments.

A final limitation is our method of using triplicate agreement (consistency of responses across three trials) as an indicator of reliability. Although our data show correlations between triplicate agreement and accuracy, we did not formally evaluate test-retest reliability or correlate this measure with human expert confidence. Future studies should explicitly validate this method against established reliability standards and expert clinician confidence to confirm its utility as a reliability metric.

Despite these limitations, this study serves as a foundational benchmark for evaluating the applicability of various LLMs and their comparative performance in clinical orthopaedics. Future research should focus on improving the accuracy and uniformity of AI-generated responses through enhanced model training on diverse, high-quality orthopaedic datasets. Incorporating multimodal learning approaches that integrate both text and image data more effectively may enhance models' performance on visually complex questions. Additionally, efforts to fine-tune models with expert-curated datasets and real-world clinical cases can enhance their relevance and reliability. Further investigation into bias mitigation strategies, such as dataset augmentation with underrepresented patient demographics and the use of explainable AI techniques, will be essential to ensure fairness and transparency in AI-driven orthopaedic education. Longitudinal studies assessing the impact of AI-assisted learning tools on orthopaedic training outcomes will also be valuable in understanding their real-world utility.

CRediT authorship contribution statement

Gnaneswar Chundi: involved in the, Conceptualization, and design of the study, Material preparation, data collection, and, Formal analysis, were carried out by, Writing – original draft. Abhiram Dawar: involved in the, Conceptualization, and design of the study, Material preparation, data collection, and, Formal analysis, were carried out by, Writing – original draft. Syed Sarwar: involved in the, Conceptualization, and design of the study, Material preparation, data collection, and, Formal analysis, were carried out by. Sanjiv Prasad: involved in the, Conceptualization, and design of the study, Material preparation, data collection, and, Formal analysis, were carried out by. Michael Vosbikian: involved in the, Conceptualization, and design of the study, provided, Supervision. Irfan Ahmed: involved in the, Conceptualization, and design of the study, provided, Supervision, All authors reviewed and approved the final manuscript.

Data availability statement

The datasets generated and analyzed during the current study are available from the corresponding author upon reasonable request. Due to the proprietary nature of the Orthopaedic In-Training Examination (OITE) question bank, the full dataset cannot be publicly shared.

Permission to reproduce material from other sources (if applicable)

N/A.

For clinical trials (if applicable)

N/A.

Level of evidence

Level V.

Ethics statement

Not applicable for this study.

Patient consent statement (if applicable)

Not applicable for this study.

Funding

No funding was provided for this study.

References

  1. , , , . Comparison of ChatGPT–3.5, ChatGPT-4, and orthopaedic resident performance on orthopaedic assessment examinations. JAAOS - Journal of the American Academy of Orthopaedic Surgeons. 2023;31(23):1173-1179.
    [Google Scholar]
  2. , , , et al . Evaluating ChatGPT, gemini and other Large Language Models (LLMs) in orthopaedic diagnostics: a prospective clinical study. Comput Struct Biotechnol J. 2025;28:9-15.
    [Google Scholar]
  3. , , , et al . Artificial intelligence-generated hip radiological measurements are fast and adequate for reliable assessment of hip dysplasia: an external validation study. Bone Joint Open. 2022;3(11):877-884.
    [Google Scholar]
  4. , , , et al . An increasing number of convolutional neural networks for fracture recognition and classification in orthopaedics: are these externally validated and ready for clinical application? Bone and Joint Open. 2021;2(10):879-885.
    [Google Scholar]
  5. , , , , , . Potential benefits, unintended consequences, and future roles of artificial intelligence in orthopaedic surgery research: a call to emphasize data quality and indications. Bone Joint Open. 2022;3(1):93-97.
    [Google Scholar]
  6. , , , , , . ChatGPT is an unreliable source of peer-reviewed information for common total knee and hip arthroplasty patient questions. Adv Orthop. 2025;2025
    [Google Scholar]
  7. , , , et al . Pediatric supracondylar humerus and diaphyseal femur fractures: a comparative analysis of chat generative pretrained transformer and Google gemini recommendations versus American academy of orthopaedic surgeons clinical practice guidelines. J Pediatr Orthop Jan 14 2025
    [Google Scholar]
  8. , , , et al . Arthrosis diagnosis and treatment recommendations in clinical practice: an exploratory investigation with the generative AI model GPT-4. J Orthop Traumatol. 2023;24(1):61.
    [Google Scholar]
  9. , , , , , . The METRIC-framework for assessing data quality for trustworthy AI in medicine: a systematic review. npj Digit Med. 2024;7(1):203.
    [Google Scholar]
  10. , , , et al . Learning to summarize with human feedback. Adv Neural Inf Process Syst. 2020;33:3008-3021.
    [Google Scholar]
  11. , . Artificial intelligence for fracture diagnosis in orthopedic X-rays: current developments and future potential. Sicot j. 2023;9:21.
    [Google Scholar]
  12. , , , et al . Artificial intelligence in fracture detection: a systematic review and meta-analysis. Radiology. 2022;304(1):50-62.
    [Google Scholar]
  13. , , , . Re-Thinking data strategy and integration for artificial intelligence: concepts, opportunities, and challenges. Appl Sci. 2023;13(12):7082.
    [Google Scholar]
  14. , , , , , , . Reproducibility in machine learning for health research: still a ways to go. Sci Transl Med. 2021;13(586)
    [Google Scholar]
  15. , , , , , , . Artificial intelligence in orthopaedic surgery. Bone Joint Res. Jul 10 2023;12(7):447-454.
    [Google Scholar]
  16. , , , et al . Deciding without data: clinical decision-making in pediatric orthopedic surgery. Int J Qual Health Care. Dec 15 2020;32(10):658-662.
    [Google Scholar]
  17. , , , . Patient safety in orthopedics and traumatology. 2021:275-286.
    [Google Scholar]
  18. , , . Overfitting, underfitting and general model overconfidence and under-performance pitfalls and best practices in machine learning and AI. 2024:477-524.
    [Google Scholar]
  19. , , , et al . Calibration: the Achilles heel of predictive analytics. BMC Med. 2019;17(1):230.
    [Google Scholar]
  20. , , , et al . The value of standards for health datasets in artificial intelligence-based applications. Nat Med. Nov 2023;29(11):2929-2938.
    [Google Scholar]
  21. , , , et al . What large language models know and what people think they know. Nat Mach Intell 2025
    [Google Scholar]
  22. , , , et al . "Confidently Nonsensical?'': A aritical Survey on the Perspectives and Challenges of'Hallucinations' in NLP. arXiv preprint arXiv:240407461 2024
    [Google Scholar]
  23. , , , , , , . Think twice before assure: confidence estimation for large language models through reflection on multiple answers. arXiv preprint arXiv:240309972 2024
    [Google Scholar]
  24. , , , . Bias in medical AI: implications for clinical decision-making. PLOS Digit Health. Nov 2024;3(11)
    [Google Scholar]
  25. , , , , . Females have a greater incidence of stress fractures than males in both military and athletic populations: a systemic review. Mil Med. Apr 2011;176(4):420-430.
    [Google Scholar]
  26. , , , et al . Males have larger skeletal size and bone mass than females, despite comparable body size. J Bone Miner Res. Mar 2005;20(3):529-535.
    [Google Scholar]
  27. , , , , , , . Association of female reproductive factors with incidence of fracture among postmenopausal women in Korea. JAMA Netw Open. 2021;4(1)
    [Google Scholar]
  28. , , , , . Advancing the inclusion of underrepresented women in clinical research. Cell Rep Med. Apr 19 2022;3(4)
    [Google Scholar]
Show Sections