Jonas Lesigang, Jakob Pietschnig
Multimodal large language models (MLLMs) have become increasingly proficient at solving complex problems. Here, we systematically compare performance between humans and two of the currently most widely used MLLMs, ChatGPT-5 and Gemini 3, on mental rotation and figural reasoning tests using identical instructions for humans and online interface-instructed MLLMs. Furthermore, we aimed to assess the effect of prompting configurations (self-consistency, segmented uploads, different contexts, explicit Chain-of-Thought prompting, German vs. English prompts) on MLLM performance. To this end, we assessed a human online sample (total N = 430) as well as ChatGPT-5 and Gemini 3 performance on five mental rotation and figural reasoning tests in a cross-sectional comparative design. MLLM performance was evaluated via percentile ranks. The human online sample outperformed MLLMs in both mental rotation (MLLM percentile rank ranges 0 to 7 on most tests and conditions) as well as figural reasoning (MLLM percentile ranks ranges 0 to 79) tests. Prompting configuration changes increased ChatGPT-5 but not unequivocally Gemini 3 performance (mean score change ranges: 0.67 to 2.33 and -1.67 to 1.50 points, respectively). Combining best-performing prompting configurations in single comparisons yielded largest increases in both models (3.83 and 2.33 points, respectively). Our results indicate that humans may currently outperform widely used MLLMs in mental rotation and figural reasoning tasks, especially if they require complex visuospatial abilities.