Katja Gorlinski
Does the quality of thinking behind data selection affect a model's "character"? We train the same model on four different sets of texts, all drawn from one shared pool, and give every trained copy the same exam. The only thing that changes between the sets is how the texts were chosen: at random, by an attentive human with no procedure, and by a second person following a fixed written protocol. A fourth, deliberately flattering set checks that the measurement works at all. Forty training runs. Every criterion is fixed before the first run, and a null result is published with the same precision as a success. The full frozen document, the figures and the budget are in the attached files.