Research

Human reviewers verify, compare and adjudicate AI outputs, but human consensus remains a comparator rather than unquestioned truth.

As the project developed, the human role required clearer definition. Human involvement could not be reduced to a final glance at an AI-generated table, but neither should human consensus be treated as automatically correct. The working sequence is: human appraisal first, AI appraisal independently, structured comparison, and human adjudication of discrepancies. Keeping the AI blind to reviewer judgements protects independence. Keeping the human assessments visible in the later comparison allows the project to examine where disagreement occurs and whether it reflects an AI error, a human error, ambiguous reporting or a genuine judgement call within RoB 2.The working sequence is: human appraisal first, AI appraisal independently, structured comparison, and human adjudication of discrepancies. Keeping the AI blind to reviewer judgements protects independence. Keeping the human assessments visible in the later comparison allows the project to examine where disagreement occurs and whether it reflects an AI error, a human error, ambiguous reporting or a genuine judgement call within RoB 2.

Every AI output remains subject to human verification before it could contribute to a real review decision. This is not because the human reviewer is a ceremonial safety officer. It is because the task combines deterministic elements with structured judgement. Applicability and response legality can be parameterised; interpretation of evidence, materiality and justified departures from default algorithms still require accountable judgement.

The evaluation therefore separates agreement from accuracy. High agreement can indicate shared error. Low agreement can identify an AI failure, a human inconsistency or a poorly specified assessment target. Discrepancy is not simply a defect to be resolved; it is one of the main sources of methodological information produced by the SWAR.