科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Data in brief2026-08-01

A dataset for studying the vocabulary of picturebooks: A supplement to the children's picturebook lexicon (CPB-Lex).

Clarence Green, Kathleen Keogh

原始摘要(英文原文)· Original abstract
This data-in-brief article shares a large-scale lexical resource derived from children's picturebook read‑aloud videos publicly available online, supplementing the related Children's Picturebook Lexicon (CPB‑Lex) by substantially increasing corpus size, and providing richer metadata such as genre information. Using automatically generated English closed captions as a proxy for picturebook text, we collected, filtered, and processed 5874 unique read‑aloud transcripts, yielding a dataset of approximately 3.49 million word tokens and 44,617 word types. Each transcript was tokenized and aggregated into document-term matrices (DTMs), one representing the full corpus, and additional DTMs for narrative and informational subsets defined via genre labels assigned by a large language model. From these lexical resources, word, lemma and text frequencies are derived and shared. Additional metadata include video titles, source channels, and view counts (where available), supporting analyses of the different types of picturebook language input patterns for educational, computation and cognitive science research. The dataset enables researchers to estimate children's exposure to specific words and lexical distributions in early shared reading, to select stimuli for experimental work (e.g., lexical decision or naming tasks), and to model aspects of vocabulary growth and reading development. The data is openly available on Mendeley Data, the Open Science Framework and a webpage platform for data sharing: https://cgg-projects.github.io/CPB_LEX/.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A dataset for studying the vocabulary of picturebooks: A supplement to the children's picturebook lexicon (CPB-Lex). — 科研速览 Science Skim