Readable research linked to original sources
Articles
Browse normalized, publication-ready research with direct links to its evidence and source.
Loading articles data…
Readable research linked to original sources
Browse normalized, publication-ready research with direct links to its evidence and source.
Loading articles data…
Readable research linked to original sources
Browse normalized, publication-ready research with direct links to its evidence and source.
3 articles
Newest firstBackgroundStandardized evaluation of agentic artificial intelligence (AI) for medication management is lacking. Given the potential lethality of medication errors endorsed or missed by AI, performance evaluation constructs are essential. The purpose of this evaluation was to develop a standardized grading framework for performance evaluation of medication management tasks. MethodsA mixed-methods approach was under…
Open article record in new tab ↗BackgroundElectrolyte replacement is ubiquitous in the acute care setting, but its familiarity cannot belie that even small dosing errors with potassium can cause lethal cardiac arrhythmias. Recently, MedAgentBench offered a benchmark for agentic artificial intelligence (AI) including the ability to correctly dose potassium based on a single rule; however, this does not adequately reflect the clinical complexity o…
Open article record in new tab ↗BackgroundLarge language models (LLMs) are increasingly used in healthcare, but standardized benchmarks fail to capture their validity and safety in real-world scenarios. Evaluating their quality and reliability is critical for safe integration into practice. MethodsFour fictitious clinical vignettes (orthopedics, pediatrics, gynecology, psychiatry) were developed by independent specialists and tested in four conv…