科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ ZDM2026-04-07· Quality (philosophy)

No one-size-fits-all: a study of prompt techniques and large language models to enhance AI’s mathematics educational quality

Sebastian Schorcht, Fabian Anton Müller, Nils Buchholtz

原始摘要(英文原文)· Original abstract
Abstract Large Language Models (LLMs) are increasingly discussed as tools for supporting mathematical problem solving. However, existing research predominantly evaluates LLM performance in terms of correctness, while the mathematics educational quality of AI-generated worked-out solutions remains largely unexplored. This study investigates how the mathematics educational quality (a set of quality criteria relating to teaching and learning mathematics) of AI-generated problem solutions can be influenced by model choice and prompt engineering. Four LLMs (Gemini 1.5 Pro (Advanced), Claude 3.5 Sonnet, ChatGPT-o3 mini, and DeepSeek-R1) were tested on six problem-solving tasks from the domains of number and algebra using four prompt techniques (Zero Shot, Chain of Thought, Persona, and Retrieval-Augmented Generation). In total, 2880 solutions were analyzed, combining human expert coding of content-related, process-related, and pedagogical-contextual quality with binary logistic regression. Results show that content-related quality is mainly driven by model type and task characteristics, whereas process-related and pedagogical-contextual quality depend strongly on prompt design, particularly Persona prompting. Overall, no single model or prompt technique performs optimally across all dimensions, indicating that effective educational use of LLMs requires context-sensitive combinations of models, prompts, and tasks.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

No one-size-fits-all: a study of prompt techniques and large language models to enhance AI’s mathematics educational quality — 科研速览 Science Skim