科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Segmentation, classification, and synthesis for brain tumors and traumatic brain injuries : MICCAI 2025 Challenges: BraTS-Lighthouse 2025 and AIMS-TBI 2025, held in conjunction with MICCAI 2025, Daejeon, South Korea, September 23, 2025,...2026-01-01

My Model Is Better Than Yours! Statistically-aware Ranking for Fair Benchmarking of AI Models.

Spyridon Bakas, Siddhesh Thakur, Ujjwal Baid, Akis Linardos, Sarthak Pati, Jimit Doshi, Russell T Shinohara

原始摘要(英文原文)· Original abstract
The proliferation of artificial intelligence (AI) in healthcare has triggered the need of fair benchmarking. This, in turn, has inspired the rise of computational challenges, where participants worldwide submit models under standardized evaluation protocols. However, ranking AI models -particularly when declaring winners- raises questions about their difference from the next best (but lower-ranked) model. Here, we present a statistically-aware ranking framework, PermRanker , designed to be agnostic to computational workloads (e.g., segmentation, classification, registration) and underlying data types (e.g., 2D pathology images, 3D MRI scans, or even non-imaging data). PermRanker provides fair and informative benchmarking of AI models, based on two stages: (i) a ranking score, based on case-wise cumulative rankings aggregated across multiple metrics and testing cases, and (ii) a rigorous statistical significance analysis via pairwise permutation testing across the ranked order of the AI models. This framework has served as the official ranking mechanism for over 33 international challenges between 2017 and 2025, including the BraTS, FeTS, and ISLES challenges. While mainly applied in biomedical AI challenges, PermRanker aims to address the unmet need of fair and informative benchmarking of AI models beyond this scope, tackle actual real-world conditions, and hence contribute in streamlining clinical translation of AI models.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

My Model Is Better Than Yours! Statistically-aware Ranking for Fair Benchmarking of AI Models. — 科研速览 Science Skim