Readable research linked to original sources
Articles
Browse normalized, publication-ready research with direct links to its evidence and source.
Loading articles data…
Readable research linked to original sources
Browse normalized, publication-ready research with direct links to its evidence and source.
Loading articles data…
Readable research linked to original sources
Browse normalized, publication-ready research with direct links to its evidence and source.
5 articles
Newest firstIn this study, we introduce a novel benchmark comprising over 9,000 real-world, point-of-care, multilingual, and multimodal clinical question-answer pairs sourced from frontline health workers in Nigeria. Using the dataset, we compare local general practitioners to multiple leading open and closed LLMs. Our results reveal several critical insights into the suitability of LLMs as clinical decision support systems i…
Open article record in new tab ↗To characterise the potential learning effects from a GenAI-based clinical decision support tool (CDST), we examined clinician behaviour within a cluster-randomised trial. The tool, AI Consult, parsed clinician notes written (in real-time) to document patient encounters and would raise green, yellow, or red flags to indicate no, potential, or critical risks of harm (respectively) in decisions the clinician made. O…
Open article record in new tab ↗Evaluating the outputs of generative AI (GenAI) models in healthcare remains a significant bottleneck for the safe and scalable deployment of these tools. Human expert raters remain the gold standard for assessing the accuracy, contextual appropriateness, and empathy of AI-generated responses, but their assessments are costly, inconsistent, and difficult to scale. The concept of "LLM-as-a-judge" systems, i.e., AI…
Open article record in new tab ↗BackgroundLarge language models (LLMs) show promise on healthcare tasks, yet most evaluations emphasize multiple-choice accuracy rather than open-ended reasoning. Evidence from low-resource settings remains limited. MethodsWe benchmarked five LLMs (GPT-4.1, Gemini-2.5-Flash, DeepSeek-R1, MedGemma, and o3) against Kenyan clinicians, using a randomly subsampled dataset of 507 vignettes (from a larger pool of 5,107 c…
Large language models (LLMs) have demonstrated strong performance in medical contexts; however, existing benchmarks often fail to reflect the real-world complexity of low-resource health systems accurately. This study developed a dataset of 5,609 clinical questions contributed by 101 community health workers (CHWs) across four Rwandan districts and compared responses generated by five large language models (LLMs)…