Zhehan Jiang
Qian et al. showed single LLMs can generate acceptable knowledge-based questions but struggle with higher-order reasoning. We argue this is architectural: decomposing item development into specialized agents for drafting, critique, and iterative adversarial refinement improves quality. In blinded evaluation for China's National Medical Licensing Examination, multi-agent outputs received 57.7% of expert preferences, compared with 42.3% for the single-model baseline, indicating a viable path to surpass current limits in AI-assisted item generation.