Masaru Shirasuna, Yuto Yoshida
Computational capacity and knowledge of humans are more limited than those of large language models (LLMs). However, in buzzer quizzes, human trivia experts can often identify the correct answer even from insufficient information such as only a few words. Investigating how they can make fast and accurate judgments is expected to highlight new characteristics of human intelligence, but little is known about that. In this exploratory and case-based analysis, we predicted that trivia experts and LLMs would differ in which words/phrases in a question are important for identifying the answer, and compared experts' performance with LLMs' performance in Japanese trivia questions through an LLM-as-a-judge approach: We regarded LLMs as evaluators and then used their outputs as comparative tools for experts' evaluations. First, we constructed a quiz question processing system that tokenized question texts based on morphological analysis and then numerically evaluated the importance of each token using GPT-4o/GPT-4.1. Then, we conducted a behavioral experiment wherein actual trivia experts were asked to numerically evaluate the importance of each token, just as LLMs had performed. As a result, trivia experts treated a variety of words/phrases as important, even if each word/phrase did not appear to be strongly associated with the answer. Their evaluation scores tended to accumulate faster than those of the LLMs. This may indicate that trivia experts can make inductive inferences faster (e.g., finding a common concept even from few items). More advanced question-answering systems may be designed by applying trivia experts' cognitive processes to LLMs' information processing, and our findings may provide a scaffolding toward such goals.