Adam Wasilewski
This dataset contains the execution scripts, prompt templates, synthetic persona profiles, and raw/processed simulation outputs for the study examining the evolution of demographic bias in GPT-4.1 during sequential fine-tuning on the Belkin Audio Device Segmentation Survey ($N = 3000$).
1. The /participants Directory
This directory contains the input data defining the synthetic persona sample and the prompt templates used to control how information was presented to the model:
a. llm_prompts_conditions_ABC.csv: Prompt templates across three experimental conditions: Condition A (explicit demographic labels: age, gender, income), Condition B (behavioral descriptions and lifestyle details without categorical labels), and Condition C (no persona profile, evaluating the model's baseline assumptions).
b. selected_participants_300.csv: A stratified random sample of $N = 300$ personas drawn to match the empirical population distribution from the reference Belkin survey.
c. selected_participants_300_with_profile.csv: Full persona profiles containing demographic attributes (Q1, Q2, Q45) and behavioral variables (recent purchases Q9, primary purchasing channel Q20, self-reported lifestyle Q37).
2. The /pilotage Directory
This directory holds the files from the pilot study used to validate output formatting, verify text generation stability, and calibrate sampling temperature:
a. script/run_full_pilot_belkin@20260627.py: Python execution script used to run the pilot simulation
b. results/llm_pilot_v2_full_parsed_responses.csv: A dataset of 3,000 parsed responses generated during the pilot (50 personas, 5 runs per cell across 4 checkpoints and 3 prompting conditions), used to confirm a 0.00% JSON structural error rate and set the optimal temperature ($T = 0.7$).
c. results/llm_pilot_v2_parsed_responses_first100.csv: A subset of the first 100 generated responses used to verify and validate the deterministic text parser.
3. The /robustness Directory
This directory gathers data used for planned robustness and sensitivity checks conducted on a 10% sub-sample of personas ($N = 30$):
a. extra_robustness_full_prompt_text_parsed_responses.csv: Results from the prompt phrasing sensitivity test (RC1), assessing sBAS metric stability under an alternative paraphrase of the primary instruction.
b. extra_robustness_full_temperature_corrected_parsed_responses.csv: Results from the generation temperature sensitivity test (RC3), comparing output consistency between $T = 0.7$ and $T = 1.0$.
4. The /study Directory
This directory contains the scripts and datasets from the primary experimental study:
a. script/run_full_study_v3_parallel.py: An optimized execution script implementing parallel API calls for the full experimental design.
b. results/llm_full_study_v3_main_corrected_raw_responses.csv: Raw text outputs returned by the LLM prior to deterministic parsing and numeric encoding.
c. results/llm_full_study_v3_main_corrected_parsed_responses.csv: The primary processed dataset covering full simulations across 4 fine-tuning checkpoints (train000, train033, train067, train100), 3 prompting conditions (A, B, C), 300 personas, and $k = 43$ generation runs per cell, including mapped data on hardware category ownership (Q11) and brand purchase consideration (Q34).