
IThe o1-preview model was tested on diagnostic reasoning tasks. On New England Journal of Medicine (NEJM) clinicopathological conferences (143 cases), the model included the correct diagnosis in its differential 78% of the time. Its top-1 accuracy was 52%. For NEJM Healer cases (80 responses), the model achieved a perfect revised-IDEA score in 78 cases. This compares to 47 for GPT-4, 28 for attending physicians, and 16 for residents. This highlights significant Large language Model performance and clinical reasoning Tasks in diagnostics.
The o1-preview model was tested on diagnostic reasoning tasks. On New England Journal of Medicine (NEJM) clinicopathological conferences (143 cases), the model included the correct diagnosis in its differential 78% of the time. Its top-1 accuracy was 52%. For NEJM Healer cases (80 responses), the model achieved a perfect revised-IDEA score in 78 cases. This compares to 47 for GPT-4, 28 for attending physicians, and 16 for residents. This highlights significant Large language Model performance and clinical reasoning Tasks in diagnostics.
The evaluation also included management reasoning scenarios. The o1-preview model achieved a median score of 86%. In contrast, GPT-4 alone scored 42%. Physicians using GPT-4 scored 41%, and those with conventional resources scored 34%. This demonstrates superior LLM Clinical Reasoning Performance in managing patient care scenarios compared to other groups.
The model's capabilities were further assessed using 76 real emergency department (ED) cases. O1 produced diagnoses rated “exact/very close” in 67%–83% of cases. This performance spanned across three diagnostic stages. It surpassed two attending physicians at each stage. This shows strong Large language Model performance and clinical reasoning Tasks in practical, complex settings.
These results show that current LLMs have surpassed most existing clinical reasoning benchmarks. However, these evaluations reflect isolated cognitive assessments. They do not represent real-world clinical integration. The question remains whether AI-assisted reasoning improves patient outcomes. This requires future prospective trials. Thus, the full impact of LLMs on patient care needs further study
| Other Related Links | |
| AI Technical Performance vs Human Performance | MMLU: Massive Multitask Language Understanding |
| Closed vs Open-Weight AI Models | AI for Science in 2025 – Shaping Future Discoveries |
.

