Ying Fang, Pavlo Bazilinskyy, Marieke Martens
Supplementary material for the paper: Fang, Y., Bazilinskyy, P., & Martens, M. (2026). Interviewing Silicon Experts: A Persona-Based LLM Interview Pipeline in Automated Driving. Manuscript submitted to ACM.
This study introduces a three-phase, persona-based large language model (LLM) interview pipeline for structuring multidisciplinary expert knowledge before involving human experts. Using drivers' minimum mental model (MMM) for the safe use of automated driving systems as a demonstration topic, the study first constructed 28 expert personas spanning seven disciplines, two technological orientations, and two priorities. Seven differentiated "silicon experts" were then purposively selected, converted into first-person professional backstories, and interviewed independently using a structured protocol. Their JSON-formatted responses were analysed through point-level open coding. The demonstration produced 58 distinct candidate MMM items across six predefined themes. Shared priorities included driver responsibility, concrete operational design domain (ODD) boundaries, and permitted and prohibited non-driving-related tasks (NDRTs), while persona-specific contributions included over-the-air updates, sensor limitations, and the perceivability of warning channels. The pipeline is intended as a transparent and auditable preliminary scoping tool for preparing subsequent human expert interviews and validation.
This supplementary dataset contains the complete pool of 28 structured expert personas, the seven selected personas and their first-person backstories, the interview prompts and execution scripts, seven JSON-formatted expert interview responses, the point-level open-coding workbook, and the four figures reported in the paper.
It has the following structure:- p1_expert_profile/: Contains the structured persona profiles, backstory prompt, generation script, and generated backstories for Phase 1.- p1_expert_profile/persona_metrics_n28/: Contains 28 individual JSON persona profiles combining seven professional disciplines, two technological orientations, and two priorities.- p1_expert_profile/generated_personas.json: Stores the complete pool of 28 structured persona profiles in a single JSON array.- p1_expert_profile/selected_personas/: Contains the seven purposively selected persona profiles used in the expert interviews.- p1_expert_profile/first_person_backstory_template.txt: Prompt template for translating each selected profile into a discipline-specific first-person professional backstory.- p1_expert_profile/generate_backstories.py: Python script for generating and storing first-person backstories from the selected persona profiles.- p1_expert_profile/backstories/: Contains the seven generated first-person expert backstories used to condition the interview model.- p2_interview/: Contains the interview prompts, interview execution script, and structured LLM expert responses for Phase 2.- p2_interview/system_prompt.rtf: Defines the model's role as a discipline-specific expert participating in an interview about drivers' MMM requirements.- p2_interview/shared_rules.rtf: Specifies shared interview constraints, including discipline-native reasoning, safety relevance, concrete scenarios, and non-fabrication of sources.- p2_interview/interview_questions.rtf: Contains the D1 interview dimension, "Minimal declarative knowledge requirement," addressing what drivers should know before first use of an automated driving system.- p2_interview/run_interviews.py: Python script for assembling persona-conditioned prompts and storing structured interview outputs.- p2_interview/interviews/: Contains seven JSON-formatted interview responses, one per selected persona, including candidate knowledge points, reasoning, scenarios, examples, and references or legal bases.- p3_open_coding/: Contains the point-level qualitative coding materials produced in Phase 3.- p3_open_coding/point_level_codes_codebook.xlsx: Point-level coding workbook containing 210 coded units from the seven interviews, including the original statements, LLM-generated codes, human-generated codes, condensed codes, and final themes.