Clarence Green, Kathleen Keogh
This data-in-brief article shares a large-scale lexical resource derived from children's picturebook read‑aloud videos publicly available online, supplementing the related Children's Picturebook Lexicon (CPB‑Lex) by substantially increasing corpus size, and providing richer metadata such as genre information. Using automatically generated English closed captions as a proxy for picturebook text, we collected, filtered, and processed 5874 unique read‑aloud transcripts, yielding a dataset of approximately 3.49 million word tokens and 44,617 word types. Each transcript was tokenized and aggregated into document-term matrices (DTMs), one representing the full corpus, and additional DTMs for narrative and informational subsets defined via genre labels assigned by a large language model. From these lexical resources, word, lemma and text frequencies are derived and shared. Additional metadata include video titles, source channels, and view counts (where available), supporting analyses of the different types of picturebook language input patterns for educational, computation and cognitive science research. The dataset enables researchers to estimate children's exposure to specific words and lexical distributions in early shared reading, to select stimuli for experimental work (e.g., lexical decision or naming tasks), and to model aspects of vocabulary growth and reading development. The data is openly available on Mendeley Data, the Open Science Framework and a webpage platform for data sharing: https://cgg-projects.github.io/CPB_LEX/.