Research

Building, testing and verifying the AI workflow A clean and auditable AI environment was established, four studies were used to test the workflow, and a struct

The first calibration work required the AI environment to be treated as an experimental condition rather than as an informal conversation. A clean, blank ChatGPT account was therefore used, with no access to earlier project discussions, human risk-of-bias assessments, reviewer comments or previous attempts at the task. This separation was important because the AI intervention needed to be reproducible. An output influenced by prior conversations, stored context or knowledge of the human reviewers’ decisions could not be treated as an independent assessment. The model had to encounter each study through a controlled and consistent evidence package. Each study was therefore supplied using a standardised input package containing the citation, the relevant methods and results text, the chronic post-surgical pain outcome data, and any available protocol, trial registry or statistical analysis plan material. The exact result being assessed had to be defined in advance so that the AI did not select a different comparison, outcome measure, time point or analysis from the one being evaluated by the human reviewers.

Where relevant documents or information were unavailable, their absence had to be stated explicitly. The AI was instructed to use only the supplied material, avoid web searches and outside knowledge, and write “No information in supplied text” where the evidence required to answer a signalling question was absent. This was intended to prevent two common failures. First, the model could not fill gaps using plausible but unverified knowledge about the study or trial. Second, missing information could not quietly disappear beneath a confident narrative. Absence of evidence had to remain visible in the output. The prompt and accompanying instructions were to be locked before the formal evaluation. Human operators could prepare the standardised evidence package, upload the material and run the prompt, but they could not edit, correct or improve the AI response before it was compared with the human assessment.

The complete output had to be retained verbatim. This included uncertainty, formatting errors, incomplete answers, unsupported claims, routing failures and any correction made during the AI’s own verification stage. A polished version created after human review would no longer represent the intervention that had actually been delivered. An oversight log was introduced to distinguish implementation failures from methodological judgements. This allowed the project to record problems such as incomplete uploads, missing source documents, truncated responses, incorrect prompt versions, use of the wrong RoB 2 pathway, failed runs and the need for reruns without confusing them with disagreement over the substantive risk-of-bias judgement.

Timing was also separated into distinct components: preparation of the study package, AI generation, checking, transfer of the output into the comparison materials, independent human assessment, consensus or adjudication, and any required reruns. This prevented the rapid production of model-generated text from being presented as the total time cost of the workflow. The scientifically important work included defining the task, preparing the evidence, controlling the assessment target and checking whether the process had been followed correctly. The governing principle was simple: if the AI output is the intervention, polishing it afterwards changes the intervention.

Running the four-study calibration pilot

The first calibration run used four studies: Amr 2024, Hu 2025, Kono 2025 and Zheng 2025. These studies formed a pilot set through which the prompt, evidence-package workflow and transfer of the AI output into the RoB 2 comparison materials could be tested before the formal evaluation began.

The work took approximately four hours, including around two hours devoted to preparing and refining the rules and instructions. That distribution of time was revealing. The model could generate an answer quickly, but the central methodological work lay elsewhere: defining the target result, deciding what evidence the model would receive, translating RoB 2 into an executable sequence, checking whether the correct pathway had been followed and determining how failures would be preserved and analysed. The four pilot studies were designated as calibration studies and excluded from the primary agreement analysis. Their purpose was not to increase the apparent size of the evaluation dataset. They were included specifically because they were permitted to change the intervention.

Any modification arising from the calibration process had to be recorded with a version number, date and rationale. Once calibration was complete, the parameterisation and accompanying source pack would be frozen before being applied to the formal study set. The pilot demonstrated that a clean AI run was feasible and that the resulting assessments could be transferred into the existing RoB 2 comparison structure. It also exposed a more important problem: plausibility was not sufficient evidence of methodological validity.

The AI could produce clear, confident and professionally formatted reasoning while still following an invalid route through the tool, answering a signalling question that should not have been active, selecting an unsupported response option or making an inference that was not justified by the supplied material. This shifted the focus of development. The quality of the prose was no longer the central concern. The more important question was whether the process used to reach the judgement was legal within RoB 2: whether the correct questions had been activated, whether conditional routes had been followed, whether the evidence supported each response and whether the final judgement was consistent with the official algorithm. A polished answer produced through the wrong pathway remained a wrong answer. The next stage therefore concentrated on process legality rather than rhetorical quality.

Making verification part of the intervention

A structured verification loop was added after the initial AI assessment. It required the model to examine five areas: the evidential support for each response, missing or unclear information, consistency with RoB 2 logic, hallucinated or unsupported claims, and the final verified domain and overall judgements. Verification was embedded within the intervention rather than applied afterwards as an invisible human correction process. This distinction was critical. The project was not merely interested in the AI’s final answer; it also sought to determine whether a structured self-check could identify and correct failures arising during the initial assessment.

Verification could not become a licence to tidy the record. Where the AI identified an error, it had to state what the original problem had been and provide the corrected response or judgement explicitly. Both the initial assessment and the verified output were retained. Silent correction would have made the final answer appear more reliable while removing the evidence required to understand how the failure arose. It would also have prevented later analysis from distinguishing between an AI system that reached the correct answer immediately and one that reached it only after detecting and repairing its own mistake.

The same principle governed human oversight. Human reviewers could flag an incomplete assessment, unsupported inference, incorrect Domain 2 variant, failure to follow the prompt or other implementation problem. They could not edit the saved AI output. Where a new run was necessary, it had to be recorded as a rerun, with the reason and relevant version information preserved. This created two analytically distinct questions. First, how well did the AI perform on its initial attempt? Second, how effectively did the structured verification stage identify and correct its own errors?

Collapsing the initial and verified assessments into a single final judgement would conceal an important part of the mechanism being tested. A model that frequently made errors but reliably detected them during verification would have a different performance profile from one that rarely made errors at all. Both might produce the same final judgement, but they would not represent the same process or carry the same implementation risks.

The verification loop improved transparency, but it did not solve every problem. A verifier could still approve a coherent answer that had followed the wrong branch of the RoB 2 tool. Internal consistency was not enough if the assessment had begun from an invalid route. Verification therefore had to do more than reread the conclusion. It needed to reconstruct the pathway: identify the active questions, confirm that inactive questions had remained unanswered, check the conditional routing, verify the response codes and ensure that the final judgement followed from the official RoB 2 logic.

The calibration work on 19 May therefore established three foundations for the formal evaluation. The AI environment had to be clean and controlled; calibration studies had to be allowed to expose and change the intervention without entering the primary analysis; and verification had to be treated as an observable component of the AI workflow rather than as a hidden human repair stage.

Together, these decisions shifted the project away from asking whether an AI could produce an impressive-looking risk-of-bias assessment. The more demanding question was whether it could produce an assessment through a controlled, transparent and reproducible process—and whether its failures could be detected without erasing the evidence that they had occurred.

  • Journal phase: Planning, first review and analysis
  • Entry status: Retrospective
  • Project stage: AI workflow design and calibration