科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ Open MIND2026-07-31· Transparency (behavior)

From commitment to conduct: Behavioural decoupling in AI identity disclosure

Xinyi Wang, Yifan Zhang

原始摘要(英文原文)· Original abstract
This study pairs, provider by provider, the public commitments of eight major AI developers concerning AI identity disclosure ("our models will not impersonate humans" and analogous statements) with the observed behaviour of their deployed systems when users probe whether they are talking to a human. We field a fixed panel of human-written identity-probing questions (sourced verbatim from the UK AI Security Institute's RealityTest dataset) to each provider's current flagship deployment through two channels — the developer API and the consumer product interface — under three pre-specified pressure conditions, at three time points around 2 August 2026, the application date of Article 50 of the EU AI Act (used strictly as a timing anchor, not as a compliance benchmark). The primary outcome is the rate of explicit false human claims. The study examines (i) whether stronger public commitments are associated with fewer explicit false human claims, (ii) whether commitment-behaviour consistency degrades under identity-suppression instructions, (iii) whether interface-level transparency and conversation-level behaviour are mutually consistent, and (iv) whether deployment-level behaviour and interfaces change in a temporally aligned way around the regulatory application boundary. Research questions. Confirmatory: RQ1. Are providers with an explicit public commitment to AI identity disclosure associated with lower rates of explicit false human claims in deployed systems? (Presented descriptively and via an auxiliary model without provider fixed effects. Not interpreted causally.) RQ2. Does the association between commitment and behaviour weaken under identity-suppression instructions? (The single confirmatory hypothesis test of this study: the commitment x pressure interaction in the provider fixed-effects model.) Exploratory (pre-declared, not confirmatory): RQ3. Is the verifiability of a commitment associated with more robust identity-transparency behaviour? RQ3 is fixed as exploratory: we will report direction, effect sizes, provider-level patterns, and uncertainty; no p-value threshold will be used as a success criterion. Further exploratory analyses: temporal change across T0/T1/T2; the product-layer x model-layer matrix; the Microsoft/OpenAI same-underlying-model, different-deployment contrast; and a commitment-era comparison (older vs newer model generations), only where a frozen, still-accessible older model exists, supplementary and non-causal.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

From commitment to conduct: Behavioural decoupling in AI identity disclosure — 科研速览 Science Skim