Kexin Huang, Chunlei Liu, Jiaqin Yang
Creativity research faces three persistent bottlenecks: divergent-thinking scoring is labour-intensive, cognitive models of creativity remain underspecified, and laboratory tasks fall short of real-world creative achievement. Large language models (LLMs) offer potential solutions, but the field lacks a structured framework for evaluating them. This selective narrative review (January 2018-June 2026) applies the generation-capability-assessment (GCA) framework, whose three axes are operationalised through descriptive criteria with provisional heuristic thresholds. On the generation axis, LLMs exceed average human performance on divergent-thinking tasks in most independent comparisons (Hedges' g ≈ 0.5-2.6), an advantage qualified by fluency dependency, the superiority of top-performing humans at scale, and a novelty-typicality trade-off. On the capability axis, LLMs simulate some task-level associative behaviour and can generate hypotheses for human research, but there is no evidence that they instantiate human-like creative mechanisms. On the assessment axis, automated scoring shows promising reliability and convergent validity for specific languages and tasks (ICC ≥ 0.80 and r ≥ 0.70 in selected studies), but cross-language generalisation is largely untested and individual-level use is unsupported. Human-AI co-creativity may benefit from a division of labour, although social-affective dimensions may matter more than cognitive support. We provide a GCA reporting protocol and identify research priorities.