Artificial intelligence can reduce the manual burden of literature screening, but evidence reviews in oncology pose a difficult test. Eligibility criteria often combine disease sub-types, line of therapy, biomarker status, demographics, geography, and research design. Performance measured in a simple review may therefore give little indication of how the same model behaves when the population, intervention – comparator, outcomes, and research-design criteria become more granular.
The analysis evaluated a supervised machine-learning model across 28 oncology-focused literature reviews, with individual searches yielding between 587 and 3,653 records. Performance varied substantially. Decision match rates ranged from 52% to 87%, recall from 17% to 98%, precision from 8% to 59%, and F-scores from 11% to 73%. Reviews with simpler eligibility structures generally achieved higher recall, whereas demographic granularity, specific mutation criteria, line-of-therapy restrictions, and heterogeneous research designs reduced recall and precision. The model also did not provide reasons for exclusion, which made disagreement resolution more labor-intensive for human reviewers.
The key contribution is the identification of boundary conditions for supervised screening rather than a single average accuracy estimate. Across multiple real-world oncology reviews, the results show that screening performance depends strongly on the complexity of the review question. This supports a more cautious, human-supervised use of artificial intelligence, with structured filters and staged application of eligibility criteria where appropriate, rather than assuming that automation will deliver uniform performance across evidence-review tasks.




