F. Carli, P. Rusina, L. Ai, L. Kuechenhoff, P. K. P. To, E. M. McDonagh, S. Lobentanzer, F. Petroni, A. Dugourd, D. Ochoa, J. Saez-Rodriguez
Language models and agents are increasingly used in biomedicine, but current benchmarks reward correct answers even when the underlying reasoning is flawed. Here we introduce Karenina, an open-source framework that turns expert knowledge into multi-dimensional evaluations of questions, conversations and autonomous agents. Illustrated in Question-Answer pairs, multi-turn conversations and autonomous data-analysis, these dimensions together moves evaluation beyond scoring, enabling trustworthy decision-making with AI in biomedicine.