Research

Verification, prompt lock and making the development pathway visible - A full verification sweep identified a material defect in Domain 2, established the testi

This was a busy day for the project with a lot of updates and changes. A full verification sweep was conducted on version 0.4 of the parameterisation against the official RoB 2 guidance and completion template. The purpose was not simply to confirm that the document appeared coherent, but to test whether its wording, routing and judgement logic accurately represented the official tool and could be applied reproducibly within the formal evaluation.

 

First Verification:

The conditional routes and default judgement mappings for Domains 1 to 5 were found to be correct. However, the document was not considered ready to lock unchanged. The sweep identified a material wording defect in Question 2.6. Version 0.4 had narrowed inappropriate post-randomisation exclusions to exclusions made for intervention-related reasons. The official effect-of-assignment logic is broader. Eligible participants excluded after randomisation are generally a cause for concern, subject to a specific exception for participants who were subsequently found to have been genuinely ineligible, where eligibility could not reasonably have been confirmed before randomisation and the decision to exclude them could not have been influenced by intervention group.

This mattered because the narrower wording could have caused the AI to overlook exclusions that were methodologically important but not explicitly attributed to intervention-related reasons. The defect did not invalidate the wider structure of the parameterisation, but it demonstrated why line-by-line verification against the official tool was necessary before prompt lock.

Version 0.5 corrected the wording and strengthened several additional safeguards. It clarified that no fixed percentage threshold should be applied to Question 2.7 or Domain 3, because the importance of deviations and missing outcome data depends on their likely effect on the specific result being assessed rather than on a universal numerical cut-off.

The revision also distinguished ordinary non-adherence from deviations arising because of the trial context. Participants may fail to adhere to an intervention for many reasons, but this does not automatically constitute a deviation that should influence the effect-of-assignment judgement. The relevant question is whether the trial context caused deviations that were inconsistent with the intended intervention and whether those deviations were likely to affect the outcome. Further clarification was added for participant-reported outcomes. Where participants provide the outcome data themselves, they are also the outcome assessors for the purposes of Domain 4. The parameterisation must therefore consider whether participants were aware of their assigned intervention and whether that awareness could have influenced their reporting.

The sweep also strengthened the distinction between selective reporting of a result and complete non-reporting of an outcome. Domain 5 assesses whether the reported result may have been selected from multiple eligible measurements, analyses or time points. It should not be used as a substitute for judging an outcome that has not been reported at all. Governance requirements were tightened alongside the methodological corrections. A valid run requires one locked instruction source, the exact official signalling-question wording, a fixed pairwise target result and a complete record of the model, interface, source-document versions, date and operator. A coherent-looking answer generated under conflicting or changing instructions cannot be treated as reproducible evidence. The candidate version 0.5 parameterisation was therefore considered methodologically verified as a draft. It was not declared finished through applause, confidence or wishful thinking.

What must happen before prompt lock

Methodological verification of the written document is not the same as empirical validation of its execution. Before the parameterisation can be used on the formal evaluation set, it must be tested across every conditional branch and every possible Low risk, Some concerns and High risk endpoint. We have decided to imput a synthetic test suite to assess both valid and invalid routes through the parameterisation.

The planned illegal-path tests include:

  • active questions being answered as not applicable;
  • Inactive questions receiving substantive responses;
  • Signalling questions being merged;
  • Prohibited response codes being used;
  • Conditional questions being skipped;
  • Checking if the wrong Domain 2 variant being selected.

The suite will also test recurring interpretation traps. These include confusing random sequence generation with allocation concealment, treating ordinary non-adherence as a trial-context deviation, accepting an intention-to-treat label without verifying that participants were analysed according to assigned group, misidentifying the outcome assessor for participant-reported outcomes, and treating the absence of a protocol as automatic evidence of selective reporting. The four calibration studies will then be rerun using the complete and internally consistent source pack. For each run, the project will record the prompt or parameterisation version, source-document versions, model and interface, date, operator, any run failures and any changes produced by the embedded verification loop.

The calibration studies will remain excluded from the primary agreement analysis. Their purpose is to test and refine the intervention before it is frozen, not to contribute data to the final comparison.

Only after independent human review, synthetic branch testing and calibration regression have passed will the parameterisation and source pack be frozen as the formal intervention version. Any later change to wording, routing, source order, judgement mapping or verification logic will constitute a new version and will require regression testing.

Prompt lock is not an exercise in bureaucratic purity. Its purpose is to ensure that every study in the formal evaluation receives the same intervention. Without that control, differences between AI judgements could reflect changes in the instructions rather than differences in the evidence contained within the studies.

Launching the real-time journal

The verification process also reinforced the need to document the methodological pathway while the work remained in progress. This journal has therefore been created to record the development of the SWAR as it happens. Entries dated before its launch have been reconstructed retrospectively from project records, saved prompts, pilot outputs, working documents, correspondence and recorded team decisions. Future entries will be added contemporaneously wherever possible.

The journal will document more than successful milestones. It will include unexpected problems, discarded approaches, prompt and parameter changes, unresolved questions, reviewer disagreements and the reasoning behind methodological decisions. The intention is not to publish a continuous stream of project housekeeping. It is to preserve the development pathway that would otherwise be compressed into several polished sentences in the final paper.

This is particularly important in AI-supported research. Models, interfaces and product behaviour can change during the lifetime of a project. A final prompt alone cannot demonstrate how the intervention was developed, which failures it was intended to prevent, which assumptions were challenged by the evidence or why particular safeguards were introduced.

The entries therefore form part of the project’s audit trail and reflexive record. They do not replace the protocol, statistical analysis plan, locked parameterisation or final publication. They document how those outputs were created, tested, challenged and revised. The ambition is straightforward: to make the mess visible enough that another researcher can understand it, question it and reproduce it — without requiring access to the project team’s collective memory or a séance with the deleted prompt history.