Readable research linked to original sources
Articles
Browse normalized, publication-ready research with direct links to its evidence and source.
Loading articles data…
Readable research linked to original sources
Browse normalized, publication-ready research with direct links to its evidence and source.
Loading articles data…
Readable research linked to original sources
Browse normalized, publication-ready research with direct links to its evidence and source.
15 articles
Newest firstTo characterise the potential learning effects from a GenAI-based clinical decision support tool (CDST), we examined clinician behaviour within a cluster-randomised trial. The tool, AI Consult, parsed clinician notes written (in real-time) to document patient encounters and would raise green, yellow, or red flags to indicate no, potential, or critical risks of harm (respectively) in decisions the clinician made. O…
Open article record in new tab ↗IntroductionPolypharmacy in older adults is associated with increased risks of adverse drug events and functional decline. Discharge summaries often contain deprescribing recommendations, but these are frequently overlooked due to documentation complexity. ObjectiveTo develop and validate a two-stage hybrid system combining rule-based natural language processing (NLP) and large language model (LLM) for automated e…
Open article record in new tab ↗IntroductionTimely, protocol-adherent clinical decisions are crucial for reducing neonatal mortality in low-resource settings. Translating extensive national guidelines into bedside practice remains challenging. ObjectiveWe developed and evaluated AIFYA, a human-supervised, large language model (LLM)-based clinical decision support system (CDSS) aligned with Kenyas national newborn care protocols. MethodsThis pros…
Generative artificial intelligence (GenAI) applications have been at the forefront of clinical documentation assistants, aiming to reduce physician notetaking burden. However, GenAI systems are resource-intensive, and deployment in low-resource healthcare settings can be challenging and cost prohibitive. We present a symbolic reasoning model (SRM) for detecting chief complaints from clinical conversations and eval…
Open article record in new tab ↗Evaluating the outputs of generative AI (GenAI) models in healthcare remains a significant bottleneck for the safe and scalable deployment of these tools. Human expert raters remain the gold standard for assessing the accuracy, contextual appropriateness, and empathy of AI-generated responses, but their assessments are costly, inconsistent, and difficult to scale. The concept of "LLM-as-a-judge" systems, i.e., AI…
Open article record in new tab ↗BackgroundLarge language models (LLMs) show promise on healthcare tasks, yet most evaluations emphasize multiple-choice accuracy rather than open-ended reasoning. Evidence from low-resource settings remains limited. MethodsWe benchmarked five LLMs (GPT-4.1, Gemini-2.5-Flash, DeepSeek-R1, MedGemma, and o3) against Kenyan clinicians, using a randomly subsampled dataset of 507 vignettes (from a larger pool of 5,107 c…
Large language models (LLMs) have demonstrated strong performance in medical contexts; however, existing benchmarks often fail to reflect the real-world complexity of low-resource health systems accurately. This study developed a dataset of 5,609 clinical questions contributed by 101 community health workers (CHWs) across four Rwandan districts and compared responses generated by five large language models (LLMs)…
This study explores the potential of using large language models to assist content analysis by conducting a case study to identify adverse events (AEs) in social media posts. The case study compares ChatGPT's performance with human annotators' in detecting AEs associated with delta-8-tetrahydrocannabinol, a cannabis-derived product. Using the identical instructions given to human annotators, ChatGPT closely approx…
Open article record in new tab ↗BackgroundThe potential of generative artificial intelligence (GenAI) to augment clinical consultation services in clinical microbiology and infectious diseases (ID) is being evaluated. MethodsThis cross-sectional study evaluated the performance of four GenAI chatbots (GPT-4.0, a Custom Chatbot based on GPT-4.0, Gemini Pro, and Claude 2) by analysing 40 unique clinical scenarios synthesised from real-life clinical…
BackgroundThe launch of the Chat Generative Pre-trained Transformer (ChatGPT) in November 2022 has attracted public attention and academic interest to large language models (LLMs), facilitating the emergence of many other innovative LLMs. These LLMs have been applied in various fields, including healthcare. Numerous studies have since been conducted regarding how to employ state-of-the-art LLMs in health-related s…
BackgroundGenerative artificial intelligence (AI) technology has the revolutionary potentials to augment clinical practice and telemedicine. The nuances of real-life patient scenarios and complex clinical environments demand a rigorous, evidence-based approach to ensure safe and effective application. MethodsWe present a protocol for the systematic evaluation of generative AI large language models (LLMs) as chatbo…
Open article record in new tab ↗BackgroundUsing artificial intelligence (AI) to help clinical diagnoses has been an active research topic for more than six decades. Past research, however, has not had the scale and accuracy for use in clinical decision making. The power of AI in large language model (LLM)-related technologies may be changing this. In this study, we evaluated the performance and interpretability of Generative Pre-trained Transfor…
The safety of large language models (LLMs) as mental health chatbots is not fully established. This study evaluated the risk escalation responses of publicly available ChatGPT conversational agents when presented with prompts of increasing depression severity and suicidality. The average referral point to a human was at the midpoint of escalating prompts. However, most agents only definitively recommended professi…
Open article record in new tab ↗ObjectiveTo evaluate the clinical potential of large language models (LLMs) in ophthalmology using a more robust benchmark than raw examination scores. Materials and methodsGPT-3.5 and GPT-4 were trialled on 347 questions before GPT-3.5, GPT-4, PaLM 2, LLaMA, expert ophthalmologists, and doctors in training were trialled on a mock examination of 87 questions. Performance was analysed with respect to question subje…
The remarkable performance of ChatGPT, launched in November 2022, has significantly impacted the field of natural language processing, inspiring the application of large language models as supportive tools in clinical practice and research worldwide. Although ChatGPT recently scored high on the United States Medical Licensing Examination, its performance on medical licensing examinations of other nations, especial…
Open article record in new tab ↗