Md. Asif Uddin

    Projects04 of 15

    PRISMA-Local

    Screening a systematic review on your own hardware

    Screening a systematic review means ranking thousands of records without missing one, often on data that cannot leave the institution. Small open models fine-tuned to score each record in a single token, an honest result against two baselines, and a precision bug that hid inside a good AUC.

    Open the demoSource code

    The problem

    A systematic review begins with thousands of database records that two people screen by hand. Recall has to be near-total, because one missed study invalidates the review.

    So the useful output is a ranking with a stopping rule, not a yes-or-no classifier. And because records often cannot leave the institution, it has to run locally.

    How it works

    • Small open models are fine-tuned to score a title and abstract against the review’s protocol in a single generated token.
    • The score is the raw margin logit(INCLUDE) − logit(EXCLUDE), which turns a classifier into a ranker.
    • QLoRA fits a 7B model onto 16 GB; LoRA trains two independent 3B screeners.
    • vLLM serves both adapters from one copy of the base model, with prefix caching making it 1.40× faster.

    How it was tested

    The models were trained on 67 of the 73 reviews in SYNERGY and tested on the other 6, held out entirely. Testing used each review’s natural prevalence: 13,665 records, of which 4.6% are includes. The transfer is cross-topic, with zero labels from the review being screened.

    Results

    It works, and it is not good enough.

    • Work saved over sampling at 95% recall (WSS@95%) is 0.331, against 0.129 zero-shot.
    • That is level with a TF-IDF logistic regression on the same reviews.
    • It is half of what within-review active learning reaches.

    Both baselines are computed here rather than cited, so the gap is stated by me and not found by a reviewer.

    The finding worth keeping: a precision bug

    Taking the decision through a bf16 softmax collapsed 13,665 records onto 45 distinct scores. Recomputing the same margin in fp32 gave 13,619.

    AUC moved by 0.002. Recall in the top 10% moved by 26%. The metric that survived the bug is the one screening papers lead with.

    Score resolution is now logged for every run, and a collapsed run is treated as broken rather than as a weak model.

    Built with

    PyTorch, PEFT, vLLM and Gradio. QLoRA 4-bit fine-tuning of Qwen2.5-7B and LoRA fine-tuning of Qwen2.5-3B. The in-browser demo takes RIS in and gives RIS out.