科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Next research.2026-04-10· Computer science

Evaluation of LLMs for mathematical problem solving

Ruonan Wang, Runxi Wang, Yunwen Shen, Chengfeng Wu, Qinglin Zhou, Rohitash Chandra

原始摘要(英文原文)· Original abstract
Large Language Models (LLMs) have shown strong performance on a range of educational tasks, but their potential to solve difficult mathematical problems requires further evaluation. In this study, we compare three prominent LLMs, GPT-4o, DeepSeek-V3, and Gemini-2.0, on three mathematics datasets of varying complexity: GSM8K, MATH500, and the MIT OpenCourseWare dataset. We take a five-dimensional approach based on the Structured Chain-of-Thought (SCoT) framework to assess final answer correctness, step completeness, step validity, intermediate calculation accuracy, and problem comprehension. The results indicate that Gemini-2.0 achieved the strongest overall performance across the three datasets and performed particularly well on advanced questions from the MIT OpenCourseWare dataset. DeepSeek-V3 was competitively strong in well-structured domains such as optimisation, but showed less consistent performance in statistical inference tasks. Overall, the findings suggest that current LLMs can perform competitively on structured mathematical tasks, but face clear limitations on problems requiring sustained multistep reasoning and advanced symbolic manipulation. Supplementary qualitative analysis also revealed model-specific weaknesses: GPT-4o lacked sufficient precision or explanatory detail in some cases, DeepSeek-V3 often provided condensed solutions with limited intermediate detail, and Gemini-2.0 showed reduced flexibility on more complex mathematical problems.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluation of LLMs for mathematical problem solving — 科研速览 Science Skim