Adnan Agha, Muhammad Jalil, Eram Anwar, Omran Bakoush
REACT-AI and its judge pipeline are built to be reusable. They close the gap between accuracy benchmarks and genuine reasoning assessment and provide a template for studying reasoning mode effects.
BACKGROUND: Most clinical reasoning evaluations in large language models (LLMs) score only the final answer, usually multiple-choice accuracy, which explains little about model reasoning. Two developments stress this gap: reasoning-optimized models are now common, and several expose an explicit extended thinking control. Whether that deliberation improves the reasoning process remains untested with a validated instrument.
OBJECTIVE: This protocol introduces and aims to validate the Rapid Evaluation Assessment of Clinical Reasoning Tool (REACT)-AI, a Behaviorally Anchored Rating Scale (BARS) with 13 subdomains for process-oriented assessment of AI clinical reasoning, that is, the quality of the externalized reasoning rather than its faithfulness to the model's internal computation. It also builds a conflict-of-interest-controlled LLM-as-judge pipeline and tests whether extended thinking improves reasoning quality.
METHODS: This 2-phase, prospective, comparative study has a within-model thinking-mode factor. Six flagship models are run in standard and extended thinking or reasoning modes, giving 12 conditions: 4 providers are toggled within the same model, and 2 pair a standard model with a dedicated reasoning model. In phase 1, 5 standardized urgent care vignettes are answered under all 12 conditions, 3 runs per condition (180 AI outputs), and by an independent expert clinician panel (n=5) and senior and junior medical students. Every output is scored by dual-blind human raters and, in parallel, by LLM judges, with no model judging its own family. Acceptance criteria for deploying the judge at scale are a weighted κ of at least 0.60 and an intraclass correlation coefficient above 0.75. In phase 2, the validated pipeline scores 3600 AI outputs. A self-correction turn and a paired Gulf English condition run alongside and are reported separately, with the AI metacognition assessment rubric and disinformation generation rate as secondary instruments.
RESULTS: As of June 2026, ethics approval was obtained (ERSC_2025_6124), and the study is registered on the Open Science Framework under embargo. Responses to the 5 phase 1 vignettes were collected from senior and junior medical students between January and June 2026; 173 scripts were received and remain sealed, unopened, and unscored. No AI outputs have been generated, and the expert panel remains unrecruited. Following the June 1, 2026, model-version lock, AI generation and scoring will begin in September 2026, with phase 1 calibration through December 2026 and phase 2 from January to April 2027; results are expected in winter 2027. The study will report the validated instrument, judge calibration against human experts, self-preference bias by model family, and the effect of extended thinking on reflection and metacognition, including null findings.
CONCLUSIONS: REACT-AI and its judge pipeline are built to be reusable. They close the gap between accuracy benchmarks and genuine reasoning assessment and provide a template for studying reasoning mode effects.
TRIAL REGISTRATION: OSF Registries osf.io/jep5t; https://osf.io/jep5t.
INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/103220.