Readable research linked to original sources
Articles
Browse normalized, publication-ready research with direct links to its evidence and source.
Loading articles data…
Readable research linked to original sources
Browse normalized, publication-ready research with direct links to its evidence and source.
Loading articles data…
Readable research linked to original sources
Browse normalized, publication-ready research with direct links to its evidence and source.
33 articles
Newest firstBackground The increasing prevalence of cannabis use has motivated researchers to develop computational behavioral models that predict usage patterns and related health impacts in naturalistic environments. However, the opaque nature of many artificial intelligence (AI) systems limits users' ability to interpret outputs and undermines trust. Existing explainable artificial intelligence techniques often remain over…
Open article record in new tab ↗Introduction People who inject drugs (PWID) face a high risk for serious infections, yet International Classification of Diseases (ICD) codes fail to identify this population. Large language models (LLM) offer a promising alternative by extracting information from unstructured clinical text. This study evaluated the diagnostic performance of off-the-shelf LLMs in identifying PWID and related attributes from hospit…
ObjectivesNatural language processing (NLP) can enable scalable extraction of clinically relevant information from unstructured radiology reports retrieved from electronic healthcare data warehouses, but reliance on externally hosted models may pose cost, privacy, and deployment challenges. We compared self-hosted discriminative and generative NLP pipelines for automated extraction of Prostate Imaging and Reportin…
BackgroundLarge language models (LLMs) are increasingly deployed in healthcare, where they may adopt different stakeholder perspectives, yet the effect of role-prompting on clinical ethical reasoning remains poorly characterized. MethodsWe evaluated three frontier LLMs: Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro across 25 ethically complex medical cases. Each model responded from three stakeholder perspectives (…
BackgroundStandardized evaluation of agentic artificial intelligence (AI) for medication management is lacking. Given the potential lethality of medication errors endorsed or missed by AI, performance evaluation constructs are essential. The purpose of this evaluation was to develop a standardized grading framework for performance evaluation of medication management tasks. MethodsA mixed-methods approach was under…
Open article record in new tab ↗IntroductionGenerative artificial intelligence (AI) can produce realistic clinical scenarios on demand and deliver immediate, individualized feedback, yet its use to teach ethical reasoning, rather than to address the ethics of AI itself, remains underexplored in interprofessional healthcare education. AimThis pilot study examined how interprofessional healthcare students perceived an AI-enhanced, case-based platf…
Open article record in new tab ↗Large language models (LLMs) such as ChatGPT are rapidly reshaping healthcare education and simulation-based training in non-technical skills (NTS), yet no bibliometric analysis has mapped this landscape. We searched seven open-access databases (OpenAlex, PubMed, Europe PMC, Crossref, Semantic Scholar, CORE, DOAJ) for English-language publications from January 2020 to March 2026. From 100,277 initial records, a se…
BackgroundElectrolyte replacement is ubiquitous in the acute care setting, but its familiarity cannot belie that even small dosing errors with potassium can cause lethal cardiac arrhythmias. Recently, MedAgentBench offered a benchmark for agentic artificial intelligence (AI) including the ability to correctly dose potassium based on a single rule; however, this does not adequately reflect the clinical complexity o…
Open article record in new tab ↗ImportanceLarge language models are increasingly explored as clinical decision-support tools in orthodontics, yet existing evaluations have been confined to knowledge-based question answering where reported accuracy ranges from 18% to 100%. No study has evaluated performance on the computational and classificatory tasks that define daily diagnostic work. Furthermore, 84.3% of published healthcare large language mo…
BackgroundLarge Language Model (LLM) chatbots are increasingly used for exercise and fitness topics, yet users experience with these tools remains understudied. MethodsThis study is a national survey of U.S. adults who have used an LLM chatbot for exercise-related topics in the past month. Participants answered questions about the exercise-related topics for which they used LLM chatbots, their perceptions of these…
Open article record in new tab ↗BackgroundElectronic health records (EHRs) with clinical decision support tools are now ubiquitous in healthcare organizations. Clinical foundation models (CFMs) pretrained on large-scale, heterogeneous structured EHR data have emerged as a powerful approach to improve predictive performance and generalizability. Meanwhile, large language models (LLMs) pretrained on broad data sources are being applied to an expan…
Open article record in new tab ↗The rapid development of large language models (LLMs) has stimulated growing interest in their use for medical question answering and clinical decision support. However, compared with frontier proprietary systems, the empirical understanding of lightweight open-source LLMs in medical settings remains limited, particularly under resource-constrained experimental conditions. To address this gap, we introduce MedScop…
Open article record in new tab ↗PurposeTo evaluate whether large language models (LLMs) can enhance clinician-patient communication by simplifying radiology reports to improve patient readability and comprehension. MethodsA randomised controlled trial was conducted at a single healthcare service for patients undergoing X-ray, ultrasound or computed tomography between May 2025 and June 2025. Participants were randomised in a 1:1 ratio to receive…
Large language models (LLMs) are increasingly used for qualitative thematic analysis, yet evidence on their performance in analysing focus-group data, where polyvocality and context complicate coding, remains limited. Given the increasing role of such models in thematic analysis, there is a need for methodological frameworks that enable systematic, metric-based comparisons between human and model-based analyses. W…
Open article record in new tab ↗Large language models (LLMs) are increasingly explored as tools for healthcare research and data analysis. However, their applicability to structured public health datasets, especially in non-English contexts, remains underexamined. We systematically evaluated 11 state-of-the-art LLMs on their ability to generate executable Python code for analytical queries over Czech public health datasets, focusing on incidence…
Open article record in new tab ↗ObjectivesThe aim of the present study is to systematically investigate the phenomenon of Conformity Bias in contemporary LLMs, specifically evaluating how repeated probing with incorrect information influences model outputs in a clinical context. Methods4 LLMs including GPT-4o, Gemini-1.5 Flash, Claude-3 Haiku, and GPT-o1 were systematically evaluated through 20 clinical questions focused on ocular disease treatm…
Open article record in new tab ↗BackgroundLarge language models (LLMs) are increasingly used in healthcare, but standardized benchmarks fail to capture their validity and safety in real-world scenarios. Evaluating their quality and reliability is critical for safe integration into practice. MethodsFour fictitious clinical vignettes (orthopedics, pediatrics, gynecology, psychiatry) were developed by independent specialists and tested in four conv…
BackgroundMissing data is a persistent challenge in digital health research, and traditional approaches like Multiple Imputation by Chained Equations (MICE) may not capture complex patterns. While large language models (LLMs) could offer a viable alternative, their use in this context remains understudied. Moreover, a critical gap remains in embedding human-centred artificial intelligence (AI) approaches that inte…
BackgroundGeneral-purpose large language models (LLMs) have rapidly evolved from experimental tools into widely adopted components of healthcare. Their proliferation - accelerated by the "ChatGPT effect" - has sparked intense interest across patient-facing specialties. Among these, dermatology provides a high-visibility use case through which to assess LLM capabilities, evaluation practices, and adoption trends. O…
Open article record in new tab ↗Substance use disorders (SUD) are a leading cause of psychiatric hospitalization among adolescents, yet the underlying diagnostic profiles and comorbidities remain poorly characterized. Here, we applied a transformer-based language model to 4,849 hospital discharge records from adolescents (aged 11-18) admitted with mental health and SUD in Spain between 2016 and 2020. We generated dense clinical embeddings and id…