科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ Harvard Dataverse2026-07-31· Computer science

Replication Data for: Mirrors, Magnifiers, or Fabricators? How Fine-Tuning on Consumer Survey Data Reshapes Demographic Bias in Language Models

Adam Wasilewski

原始摘要(英文原文)· Original abstract
This dataset contains the execution scripts, prompt templates, synthetic persona profiles, and raw/processed simulation outputs for the study examining the evolution of demographic bias in GPT-4.1 during sequential fine-tuning on the Belkin Audio Device Segmentation Survey ($N = 3000$). 1. The /participants Directory This directory contains the input data defining the synthetic persona sample and the prompt templates used to control how information was presented to the model: a. llm_prompts_conditions_ABC.csv: Prompt templates across three experimental conditions: Condition A (explicit demographic labels: age, gender, income), Condition B (behavioral descriptions and lifestyle details without categorical labels), and Condition C (no persona profile, evaluating the model's baseline assumptions). b. selected_participants_300.csv: A stratified random sample of $N = 300$ personas drawn to match the empirical population distribution from the reference Belkin survey. c. selected_participants_300_with_profile.csv: Full persona profiles containing demographic attributes (Q1, Q2, Q45) and behavioral variables (recent purchases Q9, primary purchasing channel Q20, self-reported lifestyle Q37). 2. The /pilotage Directory This directory holds the files from the pilot study used to validate output formatting, verify text generation stability, and calibrate sampling temperature: a. script/run_full_pilot_belkin@20260627.py: Python execution script used to run the pilot simulation b. results/llm_pilot_v2_full_parsed_responses.csv: A dataset of 3,000 parsed responses generated during the pilot (50 personas, 5 runs per cell across 4 checkpoints and 3 prompting conditions), used to confirm a 0.00% JSON structural error rate and set the optimal temperature ($T = 0.7$). c. results/llm_pilot_v2_parsed_responses_first100.csv: A subset of the first 100 generated responses used to verify and validate the deterministic text parser. 3. The /robustness Directory This directory gathers data used for planned robustness and sensitivity checks conducted on a 10% sub-sample of personas ($N = 30$): a. extra_robustness_full_prompt_text_parsed_responses.csv: Results from the prompt phrasing sensitivity test (RC1), assessing sBAS metric stability under an alternative paraphrase of the primary instruction. b. extra_robustness_full_temperature_corrected_parsed_responses.csv: Results from the generation temperature sensitivity test (RC3), comparing output consistency between $T = 0.7$ and $T = 1.0$. 4. The /study Directory This directory contains the scripts and datasets from the primary experimental study: a. script/run_full_study_v3_parallel.py: An optimized execution script implementing parallel API calls for the full experimental design. b. results/llm_full_study_v3_main_corrected_raw_responses.csv: Raw text outputs returned by the LLM prior to deterministic parsing and numeric encoding. c. results/llm_full_study_v3_main_corrected_parsed_responses.csv: The primary processed dataset covering full simulations across 4 fine-tuning checkpoints (train000, train033, train067, train100), 3 prompting conditions (A, B, C), 300 personas, and $k = 43$ generation runs per cell, including mapped data on hardware category ownership (Q11) and brand purchase consideration (Q34).
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Replication Data for: Mirrors, Magnifiers, or Fabricators? How Fine-Tuning on Consumer Survey Data Reshapes Demographic Bias in Language Models — 科研速览 Science Skim