Wenwen Fu, Yuanchang Ye, Ganggang Yang, Yanwen Wei, Lina Zhang
Multiple-choice question answering (MCQA) systems are often required to act on uncertain evidence while still providing a concise decision to a user. We introduce a risk-aware MCQA interface that converts normalized option probabilities into conformal prediction sets, allowing the system to retain every answer option supported at a user-selected risk level. The option-wise conformal p-value provides an interpretable rank-based view of this decision. Across five instruction-tuned language models and common-question subsets of MMLU and MMLU-Pro, the procedure is evaluated with 100 random 1:1 calibration-test splits at target miscoverage levels [Formula: see text]. At [Formula: see text], mean empirical miscoverage ranges from 0.0390 to 0.0399 on MMLU and from 0.0469 to 0.0487 on MMLU-Pro; the same target-tracking pattern persists at the larger operating points and across diverse subjects. Prediction sets contract smoothly as the allowable risk increases, while the more demanding MMLU-Pro questions retain broader sets. These results position conformal p-values as a practical reliability layer for MCQA systems that must communicate calibrated ambiguity rather than conceal it behind a single top-ranked answer.