Accuracy of diagnostic codes and algorithms used to identify rheumatoid arthritis and juvenile idiopathic arthritis in electronic health records: systematic review and meta-analysis
1School of Pharmacy and Biomolecular Sciences, Royal College of Surgeons in Ireland (RCSI), Dublin, Ireland
2School of Medicine, Royal College of Surgeons in Ireland (RCSI), Dublin, Ireland
3Dept Public Health & Epidemiology, School of Population Health, Royal College of Surgeons in Ireland (RCSI), Dublin, Ireland
4Department of General Practice, Royal College of Surgeons in Ireland (RCSI), Dublin, Ireland
5Department of Rheumatology, Beaumont Hospital/Royal College of Surgeons in Ireland (RCSI), Dublin, Ireland
*Corresponding Author Frank Moriarty Telephone/Fax number: +3531-402-8575 e-mail: frankmoriarty@rcsi.comAbstract
Objective
This systematic review aimed to assess the diagnostic accuracy of algorithms used to identify rheumatoid arthritis (RA) and juvenile idiopathic arthritis (JIA) in electronic health records (EHRs).
Methods
We searched MEDLINE, Embase, and CENTRAL databases and included studies that validated case definitions against a reference standard such as rheumatologist-confirmed diagnosis or ACR/EULAR classification criteria. Title/abstract screening, full-text review, data extraction and quality assessment were all completed in duplicate. Results were synthesised narratively and using a bivariate random-effects meta-analysis of sensitivity and specificity.
Results
A total of 35 studies were included. Algorithms varied widely in complexity, ranging from single ICD codes to combinations including disease-modifying antirheumatic drugs (DMARDs), hospitalisation records, and specialist diagnosis. Algorithms combining ICD codes with DMARD prescriptions (pooled sensitivity= 0.79 95% CI 0.61-0.90, specificity= 0.96 95% CI 0.72-1.00, PPV= 0.78 95% CI 0.63-0.88) or requiring an ICD code assigned by a rheumatologist (pooled sensitivity= 0.91 95% CI 0.70-0.98, specificity= 0.94 95% CI 0.49-1.00, PPV= 0.70 95% CI 0.64-0.75) showed the highest accuracy, with balanced sensitivity, specificity, and positive predictive value (PPV). Less restrictive algorithms demonstrated high sensitivity but lower PPV. Substantial heterogeneity was observed across studies, likely due to differences in algorithm structure, data sources, and validation methods. Despite this variability, we used conceptually coherent categories to allow for meaningful synthesis, prioritising clinical interpretability.
Conclusions
These findings support the use of more specific algorithms when diagnostic certainty is essential and highlight the need for further validation of high-performing algorithms across diverse healthcare systems.
Significance and Innovations
- ▪ This is the first comprehensive systematic review to evaluate and synthesize the accuracy of algorithms used to identify rheumatoid arthritis and juvenile idiopathic arthritis in electronic health records (EHRs), addressing a growing need as real-world data become increasingly central in rheumatology research.
- ▪ The findings provide critical guidance for researchers and clinicians on the strengths and limitations of commonly used case definitions, helping improve validity of studies using administrative or EHR data.
- ▪ By categorizing algorithms based on their components and reference standards, this review offers a practical framework for selecting the most appropriate algorithm depending on the study purpose and data source.
- ▪ The review highlight gaps in validation efforts and emphasizes the need to validate high-performing algorithms across diverse healthcare settings and evolving coding systems, ensuring accurate disease identification in current and future research.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Clinical Protocols
Funding Statement
This research was funded by the Health Research Board Applied Programme Grant Awards (APRO-2023-028) scheme.
1.Introduction
Rheumatoid arthritis (RA) and juvenile idiopathic arthritis (JIA) are autoimmune diseases characterised by joint inflammation, and systemic complications, leading to disability, and reduced quality of life1. Despite differences in onset and classification criteria, they share key pathogenic mechanisms, clinical manifestations and therapeutic strategies2,3. A challenge in RA/JIA research is obtaining robust epidemiological data to understand the natural history of disease and impact of therapy on disease outcomes4. Electronic health records (EHRs), have greatly facilitated studying these conditions in large populations, providing comprehensive data on patient demographics, symptoms, diagnoses, and treatment history5,6, in a real-world setting.
However, the reliability and validity of research based on EHR data depend on the accuracy of the diagnostic codes and algorithms used to identify the condition7,8. Combining multiple diagnostic codes with clinical data, such as prescriptions and laboratory results, enhances RA identification accuracy. For instance, algorithms incorporating at least three physician diagnostic codes, achieve specificity and positive predictive values (PPVs) of over 90%9. The inclusion of disease-modifying antirheumatic drug (DMARD) prescriptions with diagnostic codes, significantly improves case identification10. Standardised algorithms that integrate International Classification of Diseases (ICD) codes with other clinical parameters have also been assessed for their utility in diverse healthcare systems11.
Despite advances in identifying RA/JIA cases, limitations persist due to coding inaccuracy, population heterogeneity, and differences across healthcare systems and methodological approaches. A 2013 systematic review, part of the Mini-Sentinel pilot project sponsored by the Food and Drug Administration (FDA), reported substantial variability in the accuracy of algorithms (PPV 34%-97%) applied to EHRs12. The data sources of interest were limited to the US or Canada reflecting Mini-Sentinel’s focus on US-based data for surveillance system13, thus limiting the global applicability of the findings.
The aim of this systematic review is to synthesise and critically assess evidence on the accuracy of diagnostic codes and algorithms used to identify RA and JIA in EHRs and other administrative databases. By expanding the focus to include other regions, this review provides a more comprehensive evaluation of the tools used to identify the disease across various key data sources in the context of the significant changes in RA/JIA treatment and outcomes over the past decade.
2.Methods
The study protocol was registered on PROSPERO (Registration ID: CRD42024619130). The systematic review is reported in line with the PRISMA-DTA statement14.
2.1Eligibility criteria
We included studies assessing the accuracy of codes and algorithms identifying RA and JIA in EHRs or other administrative databases against a “reference standard” based on clinician-documented diagnosis, clinician-conducted chart review, another form of clinician confirmed diagnosis, or application of the American College of Rheumatology/ European League Against Rheumatism (ACR/EULAR) criteria (e.g., to GP records, medical charts, office records) regardless of study design. Studies relying on self-reported diagnoses as the reference standard were excluded, as previously recommended15. Eligible studies had to clearly report algorithm components and describe methods for accuracy assessment. We excluded: case reports, editorials, commentaries, and narrative reviews. While systematic reviews were not included, relevant ones were used for reference checking. No restrictions were placed on language or publication date.
2.2Information sources
We searched MEDLINE (PubMed interface), EMBASE and the Cochrane Central Register for Controlled Trials. Search strategies used medical subject headings (MeSH) terms and text-words aligned with our eligibility criteria (Supplementary Appendix 1) and were adapted for each database. To ensure literature saturation, we manually reviewed references of eligible studies and included articles from prior systematic reviews.
2.3Search strategy
Our strategy was based on the Mini-Sentinel project search strategies12,16, and updated with additional terms. We reviewed the PubMed MeSH terms, as well as abstracts and full-texts of primary studies from a prior systematic review to identify relevant terminology. The search was developed collaboratively and validated by confirming it retrieved all articles from the Chung et al review12. After finalizing the MEDLINE search, it was adapted to the other databases. Full search strategies and validation process details are in Supplementary Appendix 2. Searches were conducted on November 11, 2024.
2.4Study records
2.4.1Data management
We used Covidence software for record management and deduplication. Screening questions for title/abstract review were developed and tested. Reviewers completed training and pilot screening (50 abstracts by five reviewers) to calibrate eligibility criteria.
2.4.2Study selection
Two reviewers (of CSH, JB, YA, MuF, FM) independently screened titles/abstracts against eligibility criteria. Full-texts of potentially eligible studies were reviewed in duplicate. Disagreements were resolved by discussion or a third reviewer. Reasons for full-text exclusions were documented.
2.4.3Data collection process and items
Two reviewers independently extracted data using a standardized template based on the STARD criteria17, adapted for administrative databases. Discrepancies were resolved through discussion or a third reviewer. Extracted data included study characteristics, algorithms definitions, validation methods, and accuracy measures. The template is available as a data supplement (see availability of data statement). Classifications for target condition are provided in Supplementary Appendix 3.
2.5Diagnostic accuracy measures
Primary accuracy measures were sensitivity, specificity, PPV, and negative predictive value (NPV). When not reported, these were calculated from two-by-two tables. For practical interpretation, accuracy was categorized is high (≥80%), moderate (60-80%), or low (<60%)18. The unit of assessment was the individual. Further details are provided in Supplementary Appendix 3.
2.6Risk of bias and applicability
Quality was assessed using the QUADAS-2 tool, which examines risk of bias and applicability across four domains: patient selection, index test, reference standard, and flow/timing19. Domains were rated as low, unclear or high for each domain. Overall study bias followed QUADAS-2 guidance (i.e., any high-risk domain rendered the study high risk overall)14. Two reviewers independently assessed study quality, resolving disagreements through discussion.
2.7Synthesis of results
We categorized algorithms based on the number and type of codes used to identify RA/JIA (e.g., ICD, prescriptions, diagnostic tests, procedures). Classification followed Shrestha et al20 with modifications: Algorithms were categorised as either restrictive or less restrictive. For further analysis, we developed sub-categories among restrictive algorithms based on expected accuracy and usage frequency in included studies. These were: RA diagnosis codes and DMARD prescription codes, ≥2 ICD codes for RA, ≥1 ICD code for RA by rheumatologist, and ≥1 ICD hospitalisation code for RA (full description are in Supplementary Appendix 3).
- ▪ Less restrictive algorithms: required a single diagnostic code from an outpatient visit or unspecified source (e.g., single ICD code).
- ▪ Restrictive algorithms: required multiple codes or inclusion of specific types (e.g., procedures, prescriptions, tests, or hospitalisation codes).
3.Results
3.1Study selection
The search yielded 4,109 records; after removing duplicates, 3,452 were screened based on eligibility criteria, and 108 underwent full-text review. Seventy-seven articles were excluded, reasons provided in Figure 1. Reference checks of the 31 eligible articles identified 4 additional studies, totalling 35 included in the review (Figure 1).
3.2Study characteristics
Most studies were from the US and Canada (n=21), followed by Europe, (mainly Italy, UK, Denmark and Sweden; n=10), Asia (Japan and South Korea; n=3), and Australia (n=1). The majority were retrospective cohorts (n=25), followed by cross-sectional (n=6) and case-control studies (n=4). Two-thirds (n=23, 66%) were published since the 2013 review (Supplementary Appendix 4-Table 1).
The included studies (n=35) used diverse electronic databases, including national health insurance claims-i.e. Korean National Health Insurance database24, Medicare10,25–27, and Ontario Health Insurance Plan9,28; hospital discharge records-i.e. Stockholm County Medical Information System29, Danish National Patient Registry30,31, Partners Healthcare database used by the Brigham and Women’s Hospital and the Massachusetts General Hospital32–34, and Canadian Institute for Health Information Discharge Abstract Database9,28; and institutional EHRs-i.e. Vanderbilt University Medical Center’s Synthetic Derivate32,35,36, University of California Los Angeles Health System37, Mayo Clinic Biobank38, and from specific rheumatological institutions or disease registries11,39–41. Some also incorporated pharmacy data 30,42,43, and physician billing records41,44,45.
Most studies focused on RA patients, though five examined JIA35,39,45–47. Samples sizes ranged from small cohorts, i.e., 3148 and 9446 participants, to large datasets exceeding 55,000 individuals24. Women comprised 50-80% of participants. Mean age varied, typically between 50-70 years for RA, and <19 years for JIA. Full study characteristics are detailed in Supplementary Appendix 4 - Table 1.
RA case identification relied mainly on ICD codes (Table 1), with ICD-9 (62.8%) and ICD-10 (48.6%) being the most used (Supplementary Appendix 4-Table 1). Some older datasets also used ICD-8 codes29,30. Additional criteria were applied in several studies, such as medications (54.3%), and laboratory markers (22.8%) like rheumatoid factor (RF) and anticyclic citrullinated peptide (anti-CCP) antibodies34,49 (Table 1). Some studies implemented complex algorithms, including the eMERGE algorithm, which combined ICD codes with RF values and other autoimmune disease exclusions37,38. The algorithm developed by Liao et al, incorporating ICD codes, prescriptions and laboratory values34 was validated in other settings by Carroll et al32, and Huang et al33 (Supplementary Appendix 4-Table 1).
Most studies (71.4%) used a rheumatologist’s clinical diagnosis as the reference standard (Table 1), while others applied established classification criteria including the 1987 ACR criteria, 2010 ACR/EULAR criteria, and for JIA, the 2001 International League of Associations for Rheumatology (ILAR) criteria (Supplementary Appendix 4-Table 1). Four studies relied on medical records review by trained nurses or researchers, without clear specification of the reference standard26,27,36,49 (Table 1).
3.3Risk of bias and applicability
Risk of bias and applicability concerns varied across studies, with patient selection and flow/timing being the most affected domains (Table 2). Patient selection showed the highest risk, mainly due to non-random inclusion, use of ICD codes without additional verification, and concerns regarding representativeness. The index test domain showed a low risk, generally due to clearly defined and consistently applied algorithms, though some studies lacked clarity or did not report blinding to the reference standard. The reference standard domain showed the lowest risk, using established standards like rheumatologist diagnosis or ACR/EULAR criteria, though some had unclear validation methods. Flow and timing had mixed results, with high risk in studies that did not include all participants in analysis or did not assess all diagnostic accuracy measures due to study design.
Overall, six studies were rated high risk of bias25–27,29,40,48 due to limitations in multiple domains, potentially leading to misclassification and reduced generalisability. In contrast, studies by Carroll et al32, Cho et al24, Liao et al34, Singh et al50, and Widdifield et al 20149 had low risk, with strong design across all domains (Table 2.)
Applicability concerns were generally low, suggesting good relevance to real-world settings. However, studies using specialised databases or narrow inclusion criteria (i.e., paediatric populations, single-centre databases, incident RA cases) had moderate concerns as they may not generalise to broader RA populations (Table 2).
3.4Results of individual studies
3.4.1Rheumatoid Arthritis codes and algorithms
Individual study results are detailed in Supplementary Appendix 5-Table 2. Across studies, over 180 definitions/algorithms were evaluated. Percentage of studies that used less restrictive, restrictive algorithms and their categories are presented in Figure 2.
Less restrictive algorithms typically had high sensitivity but low specificity, increasing false positives. For example, Almutairi et al51 found 90% sensitivity but only 28.5% specificity using ≥1 RA primary code in hospital discharge records, highlighting the risk of over-identification. Other studies28,44,50 showed similar trends, with sensitivities >90% and specificities below 65%.
Adding DMARD prescriptions to ICD codes improved specificity and PPV while maintaining good sensitivity. Cho et al24 achieved 85% sensitivity and 70% specificity using ICD-10 codes with biologic, DMARD, or NSAID prescriptions, maintaining >85% PPV values in all developed algorithms. Similarly, Singh et al50 reported 85% sensitivity, 83% specificity, and 81% PPV when using a similar approach. Linauskas et al30 showed PPV increased from 61.9% to 87.7% when requiring ≥1 DMARD prescription.
Requiring ≥2 ICD RA codes, separated over time, improved specificity while maintaining good sensitivity, though decreasing PPV. Hanly et al44, reported 83% sensitivity, 81% specificity, and 52% PPV using two RA physician codes ≥2 months apart. Similar trends were seen in Kubota et al52 (84% sensitivity, 99% specificity, and 59% PPV) and Widdifield et al9 (84% sensitivity, 99% specificity, and 46% PPV).
Requiring ≥1 hospitalisation record improved specificity and PPV but greatly reduced sensitivity. Convertino et al53 found 95% specificity, 81% PPV but only 37% sensitivity when requiring 1 RA hospitalisation or emergency admission code. Similarly, Hanly et al44 achieved 98.5% specificity, 77% PPV and 20.7% sensitivity when using ≥1 hospitalisation code. Similar patterns were noted by Wididifield et al9,28.
Algorithms requiring at least one rheumatologist diagnosis code showed overall moderate to high sensitivity and specificity, with moderate PPV. Widdiefield et al 20149 and 201328, achieved 81% and 99% sensitivity, 99% and 77% specificity and 51% and 68% PPV, respectively. Similarly, Hanly et al44 reported 88% sensitivity, 75.5% specificity, and 47% PPV when using ≥1 RA code by a rheumatologist. In contrast, Carrara et al11 reported 99.8% specificity but only 13% sensitivity when combining one RA rheumatologist certification with other ICD-9 codes.
The eMERGE algorithm-combining ICD codes, RF values, and other autoimmune disease exclusions-demonstrated high specificity (95-99%) and PPV (94-97%) in both studies evaluating it (Zheng et al 202237 and Kronzer et al 202038, respectively). However, sensitivity varied likely due to differences in sample size and study population. Zheng et al37 (n=4,766, UCLA Health System) reported a higher sensitivity (72%), while Kronzer et al38 (n= 497, Mayo Clinic Biobank included in the Rochester Epidemiology Project) achieved a sensitivity of 53% in a smaller and more geographically limited sample. However, Kronzer et al38 found sensitivity increased with RA duration and EHR history, reaching 71% for patients with ≥10 years of disease. Overall, the eMERGE algorithm performed well, with sensitivity maybe influenced by dataset scope.
The RA algorithm initially developed by Liao et al34, validated by Carroll et al32 and Huang et al33, consistently showed moderate sensitivity (51-79%) and high PPV (>87%) across different settings showing a good balance between measures for identifying RA cases. Huang et al33 enhanced the model by incorporating ICD-10 codes and newer biologic treatments, achieving 77% sensitivity, 95% specificity, and 91% PPV. Overall, the algorithm maintains high accuracy, with sensitivity improving slightly as case definitions and coding systems evolved.
Both the eMERGE and Liao algorithms used laboratory data (RF, anti-CCP), which improved specificity and PPV, and had limited effect on sensitivity. Paltta et al42 reported 98% PPV when incorporating anti-citrullinated protein antibodies (ACPA) positivity in their algorithm, compared with ICD codes alone (82%) or ICD codes and DMARDs (89%). Singh et al also showed improved sensitivity (88%), specificity (91%), and PPV (93%) when adding positive RF to their algorithm, compared to ICD codes alone (sensitivity: 100%; specificity: 55%; PPV: 66%) or ICD codes and DMARDs (sensitivity: 85%; specificity: 83%; PPV: 81%)50, suggesting that laboratory data helps reduce false positives while maintaining a balanced improvement in the accuracy measures.
3.4.2Juvenile Idiopathic Arthritis codes and algorithms
JIA studies evaluated similar EHR-based algorithms (Supplementary Appendix 5-Table 2). Less restrictive algorithms, like ≥1 JIA diagnosis code, showed high sensitivity but moderate to low PPV. Harrold et al39 and Peterson et al35 both reported 100% sensitivity with PPVs of 45% and 48%, respectively.
Only one study assessed the accuracy of DMARD prescription alone as a definition. Thomas et al47 found 40% sensitivity, but 100% specificity and PPV when requiring ≥1 DMARD prescription without prior alternative indication.
Requiring ≥2 ICD codes for JIA improved PPV while maintaining good sensitivity. Reported sensitivity ranged from 72-93% and PPV from 60-100% across studies35,39,45–47. Peterson et al35 noted that increasing the number of JIA codes (≥4) raised PPV (83%) without substantial decrease in sensitivity (87%), compared to using ≥1 code (PPV=48%). Shiff et al45 found specificity improved (85-92%) with more ICD codes over longer intervals.
Two studies assessed algorithms requiring ≥1 JIA code from a rheumatologist. Harrold et al39 showed that requiring a rheumatologist visit improved PPV from 69% to 90% while maintaining good sensitivity (81%). Stringer et al46 showed similar sensitivity (86%) but decreased PPV (58%), which increased to 65% when using two JIA codes ≥8 weeks apart within 2 years, maintaining good sensitivity (81%).
4.Discussion
4.1Summary of evidence
This systematic review analysed the accuracy of algorithms for identifying RA and JIA in EHRs. Overall, algorithms combining ICD codes with DMARD prescriptions, or those that include ≥1 ICD code assigned by rheumatologist, demonstrated the highest accuracy, balancing sensitivity, specificity, and PPV. These algorithms are more reliable for identifying confirmed RA cases, though rheumatologist-assigned codes may miss cases diagnosed in primary care, potentially leading to underestimation of disease prevalence. Less restrictive algorithms are useful for initial case identification (highly sensitive) but may require additional criteria to enhance specificity and PPV. Requiring ≥2 ICD codes is a well-balanced approach, reducing false positives (higher specificity) without significantly lowering sensitivity, though achieving moderate PPV. Hospitalisation-based codes are highly specific for identifying severe RA cases but may exclude many non-hospitalised, mild/moderate cases, making it less suitable for general RA case identification.
For JIA, a single ICD code is highly sensitive but may overestimate prevalence due to false positives. Requiring ≥2 spaced ICD codes improves accuracy and seems to be the best approach, providing a good balance of sensitivity, specificity and PPV in EHRs or administrative data. However, validation evidence for JIA remains limited, highlighting the need for further research.
Our findings align with and extend previous systematic reviews on RA and connective tissue diseases algorithms in EHRs. Compared to the 35 studies we included, the 2013 review by Chung et al12 included 9 studies and found that algorithms combining diagnostic and prescription data-particularly those including specialist codes, medication or laboratory data-achieved higher PPVs (>80%), while ICD-only algorithms had more variable performance across settings, generally with low PPVs. Similar trends were seen in studies for other conditions, but generally lower PPVs (<65%) in general populations20,54, probably reflecting the challenges of identifying these conditions due to non-specific coding and overlap with other autoimmune conditions. Similarly, a systematic review by Crossfield et al55 on gout also highlighted variability in methods and performance, reflecting a lack of standardisation in defining rheumatologic conditions using real-world data. Compared to these conditions, RA appears to be more accurately identified in EHRs, especially when algorithms include medication data or are supported by specialist confirmation.
The accuracy of hospitalisation-based algorithms must be interpreted within the context of evolving healthcare delivery. RA hospitalisations have declined significantly over the past two decades due to earlier diagnosis, improved outpatient care, and widespread use of biologic therapies56–58. As a result, performance of algorithms using hospitalisation codes in historic data may be less representative of current patient populations. Similarly, treatment patterns have evolved, with earlier DMARD initiation, combined therapies (e.g., biologics and DMARDs)59–61, and the development of targeted-synthetic DMARDs (e.g., JAK inhibitors), which may affect the predictive value of medication-based algorithms applied to contemporary prescribing data.
The transition from ICD-9 to ICD-10 introduced more specific diagnosis codes, potentially influencing algorithm performance. In this review, 62.8% of studies examined ICD-9 and 48.6% examined ICD-10 codes, reflecting variability in study periods, country-level implementation, differences in capacity to adopt new coding systems, and the use of historical data sources. Algorithms developed with ICD-9 codes may not translate directly to ICD-10 without re-validation62,63, particularly for specific or rare disease subtypes. These changes underscore the need for regular update and validation of algorithms to align with current coding systems and clinical practice.
From a practical perspective, the selection of an algorithm for identifying RA or JIA should be guided by the purpose of the study and the nature of the data source. For studies aiming to evaluate treatment outcomes, assess adverse events, or conduct comparative effectiveness studies, algorithms with high specificity may be preferable to ensure the correct identification of true cases. In contrast, for surveillance or studies assessing prevalence or burden of disease, algorithms prioritising higher sensitivity may be preferable, even at lower PPV.
Future research should focus on validating existing algorithms, such as the eMERGE37,38 and Liao et al34 algorithms, across diverse healthcare systems and populations. While both show consistent diagnostic accuracy in diverse settings, the generalisability of the Liao algorithm beyond academic health centres, remains uncertain. Validation using community-based data is needed to ensure broader applicability. Clear and transparent reporting of algorithm components is essential for replication, external validation, and broader implementation in different contexts. Additionally, the development and validation of algorithms using machine learning or natural language processing techniques hold promise for further improving case ascertainment, particularly in settings where structured coding is limited.
4.2Limitations
The studies reviewed exhibited substantial methodological variability, making direct comparisons difficult. Differences in sample size, sampling methods, settings, population, and data sources likely contributed to the high statistical heterogeneity in several analyses. Algorithm definitions varied widely, even within similar categories, highlighting a lack of standardisation in algorithm development and validation methods. Despite this, meta-analyses were feasible by grouping algorithms into broad, but conceptually coherent categories, prioritising simplicity and clinical interpretability. While many incorporated medication use, most studies grouped DMARDs as a single category, precluding an evaluation of the contribution of specific drugs or therapeutic classes to the accuracy. Several studies did not report full accuracy measures due to sampling designs, which involved evaluating only algorithm-positive cases, and thus only the PPV. The PPV depends on disease prevalence within the population being studied. Several studies selected their populations from hospital settings or rheumatologic clinics, where the prevalence of RA is likely to be higher than in the general population, potentially leading to an overestimation of the PPVs. Another source of variability may be the data used in each study and purpose for which it was originally generated, i.e., for claims/reimbursement versus clinical care. Therefore, findings should be interpreted cautiously, and ideally, validated in general populations to better understand their performance across varying levels of disease prevalence. Finally, the search focused on studies explicitly reporting diagnostic validation, potentially excluding relevant studies that addressed this in broader investigations (i.e., assessed an algorithm accuracy as an intermediate objective as part of an EHR-based study of RA outcomes). A protocol deviation was the omission of grey literature search due the nature of the review and to the expectation that relevant studies would be published primarily in peer-reviewed journals.
This systematic review incorporates comprehensive and contemporary evidence from diverse geographic areas. To our knowledge, is the first systematic review to comprehensively synthesise and quantitatively analyse RA and JIA algorithm accuracy in EHRs without geographic restriction across diverse healthcare settings. It adheres to PRISMA guidelines, systematically categorised algorithms by restrictiveness to facilitate clinically meaningful comparisons, and used the QUADAS-2 tool for transparent quality assessment.
4.3Conclusions
This review highlights the variability in accuracy of algorithms used to identify RA and JIA in EHRs. Algorithms combining diagnostic codes and prescription data, particularly DMARDs, or a rheumatologist diagnosis demonstrated the highest overall performance, with a good balance of sensitivity, specificity, and PPV. Less restrictive algorithms may aid initial case identification but require refinement for specific research purposes. The choice of algorithm should be tailored to the study objective and data source. Future efforts should prioritise external validation of high-performing algorithms across diverse settings to improve standardisation and reliability in real-world data research.
Supporting information
Data Availability
This research is based on published literature. All data used in this research are already included in the article or the supplementary material. The data extraction template is available at the Open Science Framework: OSF | Data Extraction Template.xlsx.
Acknowledgment
The authors acknowledge the use of large language model to assist with English grammar and spelling during the preparation of the manuscript.
Availability of data
This research is based on published literature. All data used in this research are already included in the article or the supplementary material. The data extraction template is available at the Open Science Framework: OSF | Data Extraction Template.xlsx.