RCT-Reviewer: Free Machine Learning Risk of Bias Assessment Tool

RCT-Reviewer is a modernized, standalone version of RobotReviewer, designed as a third-party reference tool for Risk of Bias assessment. It builds upon RobotReviewer's original machine learning models trained on 12,808 randomized controlled trials (RCTs).

Online RCT-Reviewer busy? Use mirror alternative below

Use mirror

Why Use RCT-Reviewer?

RCT-Reviewer is designed as an automated ML third-party tiebreaker for Cochrane RoB and systematic reviews. Unlike generative AI, this structured machine learning model provides an instant, objective, and data-driven third opinion to resolve ties without the risk of hallucination.

  1. Near-Human Accuracy

    The system achieves 71.0% accuracy for Risk of Bias judgments, performing within <8% of human expert consensus (which stands at 78.3%) [1].

  2. Highly Precise Extraction

    In a randomized Cochrane user trial, the models demonstrated 87% Precision and 90% Recall for identifying the exact text snippets supporting the bias judgment [2].

  3. Validated Acceptance

    Real-world feasibility studies show that human reviewers accept the tool's judgments at a rate equal to that of their human peers (Risk Ratio 1.02) [3].

  4. Rigorous Methodology

    Developed by Marshall, Kuiper, and Wallace, the models were trained on 12,808 clinical trial PDFs using "distant supervision" to ensure high-quality classification without prohibitive manual labeling costs [1,4].

How to cite RCT-Reviewer in your methods section:
"Risk of bias was assessed independently by two reviewers. Disagreements were resolved by consensus, or where consensus could not be reached, by using the machine learning tool, RCT-Reviewer (Sahu, 2026) as a third-party tiebreaker. RCT-Reviewer utilizes structured RobotReviewer ML models (Marshall et al., 2017) trained on 12,808 RCTs to provide data-driven bias judgments and highlight supporting text snippets."

Validated Against the Original RobotReviewer

RCT-Reviewer has been independently validated against the original 2017 RobotReviewer implementation using a five-tier validation harness. Key results:

  1. Strong Predictive Validity

    On 751 human-labelled Clinical Hedges records, RCT-Reviewer achieved 91.5% accuracy (94.1% sensitivity, 88.1% specificity, 0.925 F1, 0.966 ROC AUC).

  2. 100% Fidelity to the Original

    Risk of Bias judgments agree with the original 2017 code on all 6,018 document × domain comparisons (κ = 1.0), with identical sentence scores and vectorizer outputs.

  3. Robust PDF Parsing

    1,000/1,000 PDFs parsed successfully, with a median processing time of 1.57 seconds per PDF.

  4. External Human Validity

    Evaluated against human consensus on a 313-trial open-access subset (Tian 2024), demonstrating comparable agreement (κ 0.12–0.48) to the original tool.

Want more info? The complete methodology, statistical results and reproducibility data are documented in the Validation repository README.

References

  1. Marshall IJ, Kuiper J, Wallace BC. RobotReviewer: evaluation of a system for automatically assessing bias in clinical trials. Journal of the American Medical Informatics Association. 2016;23(1):193-201. doi

  2. Soboczenski F, et al. Machine learning to help researchers evaluate biases in clinical trials: a prospective, randomized user study. BMC Medical Informatics and Decision Making. 2019;19(1):96. doi

  3. Nussbaumer-Streit B, et al. Automating risk of bias assessment in systematic reviews: a real-time mixed methods comparison of human researchers to a machine learning system. BMC Medical Research Methodology. 2022;22:160. doi

  4. Marshall I, Kuiper J, Wallace B. Automating Risk of Bias Assessment for Clinical Trials. IEEE Journal of Biomedical and Health Informatics. 2015;19(4):1406-1412. doi

Citation

If you use this software in your research, please cite both RCT-Reviewer and the original RobotReviewer paper. Select a reference and format below to copy or download your citation.