Reframing the SWAR around risk-of-bias assessment
Retrospective entry — reconstructed from project records.Retrospective entry — reconstructed from project records.
The project began with a practical question: could generative AI contribute usefully to one of the most demanding and judgement-heavy parts of a systematic review?
The SWAR was embedded within the Regional Anaesthesia Living Systematic Review, which examines randomised controlled trials of regional anaesthesia interventions and chronic post-surgical pain.
The first major decision was to stop treating the work as a generic data-extraction exercise. Risk-of-bias assessment is not simply the recovery of facts from a paper. It requires the reviewer to identify the exact result being assessed, follow the structure of the appraisal tool, distinguish missing reporting from evidence of poor methods, and make transparent judgements under uncertainty. The SWAR registration therefore needed to be rewritten around the evaluation of AI-supported risk-of-bias assessment itself.
The intended comparison was between an independently generated AI assessment and the human review process. The evaluation would examine agreement, consistency, types of error and the time required. From the outset, the human reviewers were not positioned as an infallible gold standard. Their final consensus or adjudicated decision would be the comparator, while disagreements would remain data to be examined rather than inconvenient noise to be erased.
This reframing established the central question that still governs the project: not whether AI can produce a plausible-looking appraisal, but whether it can produce an appraisal that is accurate enough to be useful, transparent enough to audit and structured enough to reproduce.