Toshizumi Muta
Supplementary materials for the paper "Embedding qualitative data in LLM semantic space: A conceptual history, R tutorial, and empirical calibration" (Muta, 2026). This project holds the code, the archived embedding matrices, and the results files from which every statistic reported in the paper can be recomputed. The manuscript itself is on PsyArXiv rather than here, so that it has one version history rather than two. The paper asks whether commercial large-language-model embedding APIs can serve as a measuring instrument for survey data that quantitative analysis usually discards: nominal category labels and open-ended text. No text is generated and no respondents are simulated. Structure is fixed from theory and tested confirmatorily, never estimated from the embeddings. Eight pre-specified demonstrations span four psychological instruments, occupational labels, semantic-differential terms, and respondent-written narratives, across three providers (Gemini, Voyage, OpenAI) and, where a validated translation exists, two languages. Six further analyses, added after the confirmatory grid, test specific claims from the generative-psychometrics literature against response data. Contents: 30 R scripts; 23 archived embedding matrices; 42 results files; the four published figures; and qualembed, the R package (development continues at github.com/PsycholoStudio/qualembed; the copy here is a frozen snapshot of the version that produced these analyses). Reproducibility note. Commercial embedding models are versioned products and may be retired, so reproducibility attaches to the archived matrices rather than to the APIs. A verification script checks every number in the manuscript against the results files without any API access or API key. Two deliberate omissions are documented in the README. The working embedding cache is not deposited, because it is keyed by the text embedded and would therefore contain participants' own sentences; an embedding vector does not protect the text it came from. And the Japanese renderings of the Big Five and Schwartz materials that appeared in an earlier version of this work were withdrawn in full, having been produced during drafting rather than taken from a validated published translation; no result here rests on them. All analyses run in R. Code and package: GPL-3. Text and figures: CC-BY 4.0.