科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Empirical Software Engineering2025-12-05· Computer science

Evaluating the quality of GenAI applications in software engineering: a multi-case study

Liang Yu, Emil Alégroth, Panagiota Chatzipetrou, Tony Gorschek

原始摘要(英文原文)· Original abstract
Abstract Context Generative AI (GenAI) is increasingly adopted in software development for tasks such as document generation, data analysis, and code generation. However, evaluating the quality of GenAI applications becomes challenging, as traditional quality measurements may not be fully applicable. Objective In this study, we explore how practitioners evaluate the quality of GenAI applications and investigate quality evaluation techniques. Method We conducted a multi-case study in three industrial projects from software development companies. We examined four GenAI application domains: document generation, data analysis and insight generation, customer service, and code generation. Data were collected through three workshops and 23 semi-structured interviews with industrial practitioners. Results We identified fourteen GenAI use cases and 28 metrics currently used to evaluate the quality of GenAI applications’ outputs. We synthesized the identified metrics’ usage patterns and challenges based on the collected data. Conclusions This study presents practical insights into using metrics to measure GenAI-based system qualities in real industrial settings. Our findings indicate that practitioners use custom-built and context-specific metrics; combining these with academic metrics can strengthen GenAI system quality evaluation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating the quality of GenAI applications in software engineering: a multi-case study — 科研速览 Science Skim