Readable research linked to original sources
Articles
Browse normalized, publication-ready research with direct links to its evidence and source.
Loading articles data…
Readable research linked to original sources
Browse normalized, publication-ready research with direct links to its evidence and source.
Loading articles data…
Readable research linked to original sources
Browse normalized, publication-ready research with direct links to its evidence and source.
2 articles
Newest firstEvaluating the outputs of generative AI (GenAI) models in healthcare remains a significant bottleneck for the safe and scalable deployment of these tools. Human expert raters remain the gold standard for assessing the accuracy, contextual appropriateness, and empathy of AI-generated responses, but their assessments are costly, inconsistent, and difficult to scale. The concept of "LLM-as-a-judge" systems, i.e., AI…
Open article record in new tab ↗Large language models (LLMs) have demonstrated strong performance in medical contexts; however, existing benchmarks often fail to reflect the real-world complexity of low-resource health systems accurately. This study developed a dataset of 5,609 clinical questions contributed by 101 community health workers (CHWs) across four Rwandan districts and compared responses generated by five large language models (LLMs)…