科研速览 · Science Skim继续刷下去 · Keep skimming →
2026-07-31· Computer science

Running KI-Benchmark Deutsch: A Name-Blind, Cross-Vendor Evaluation of Large Language Models on German Business Tasks v1

Ideal Syka

原始摘要(英文原文)· Original abstract
KI-Benchmark Deutsch is a recurring computational benchmark for evaluating large language models on realistic German-language business and administrative tasks. Each run evaluates a fixed model roster on 24 held-out tasks across six categories using fixed rubrics, reference answers where applicable, and a name-blind cross-vendor judge panel with leave-one-provider-family-out assignment. This protocol defines answer generation, judge eligibility, score parsing, aggregation, coverage safeguards, quality assurance, and versioned publication. Applying it yields model- and category-level aggregate scores on a 0–100 scale plus a machine-readable snapshot; held-out tasks and raw answers remain private to limit contamination and gaming.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Running KI-Benchmark Deutsch: A Name-Blind, Cross-Vendor Evaluation of Large Language Models on German Business Tasks v1 — 科研速览 Science Skim