科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in digital health2026-01-01

Development of the perceived AI-physician consistency and trust scale for cardiology: a Delphi-based study.

Yun Yan, Zhiping Wang, Xueting Wang, Yuxia Zhang, Ping Zhang

一句话结论 · In one sentence

The three-round Delphi process produced a content-valid 12-item pool (PACT-12) organised into four theoretically derived domains, with satisfactory expert consensus on item relevance and clarity. These results reflect content validation only; patient cognitive interviews, exploratory factor analysis, reliability testing, and criterion-related validation are required before PACT-12 can be regarded as a psychometrically validated or clinically deployable instrument for evaluating patient-side trust shifts and associated safety risks in AI-assisted cardiovascular decision-making.

原始摘要(英文原文)· Original abstract
BACKGROUND: Large language models (LLMs) have demonstrated considerable potential in cardiovascular diagnostic assistance and clinical decision support. A critical safety-related question, however, remains poorly characterised: how patients subjectively perceive the consistency between artificial intelligence (AI) and physician diagnoses, and how they redistribute trust between the two when their outputs diverge. The costs of miscalibrated trust are ultimately borne downstream by emergency and intensive care services, a pathway typified by decompensating heart failure, in which algorithm-reassured delay converts an outpatient consultation into an intensive care admission. Recurrent clinical antecedents include patient anchoring on the most benign diagnosis offered by AI, and consultation-room friction, often compounded by concealed AI use, that quietly erodes adherence to physician advice. No standardised measurement instrument currently addresses this gap. To develop and content-validate an initial item pool for the Perceived AI-Physician Consistency and Trust Scale for Cardiology (PACT) for use in cardiology outpatient settings, as a preliminary step toward a standardised instrument for evaluating the safety of human-machine interaction in cardiovascular digital health. METHODS: We pre-specified four domains: AI-physician perceived consistency, AI competence trust, AI benevolence trust, and physician reference trust. An initial pool of 10 candidate items was generated. Ten experts from cardiology, medical informatics, and nursing psychology participated in three rounds of Delphi consultation; Rounds 1 and 2 screened and revised the item pool, and Round 3 re-administered the full 12-item pool using the same I-CVI and clarity instrument, with particular focus on confirming the three items whose domain assignment or response format had changed after Round 2. Items were screened using the item-level content validity index (I-CVI) and a 5-point language clarity rating, with retention thresholds of I-CVI ≥ 0.78 and mean clarity score ≥ 3.5. Kendall's coefficient of concordance (W) quantified convergence of expert opinion. RESULTS: Response rates were 100% in all three rounds, and the mean expert authority coefficient was Cr = 0.86. In Round 1, items Q6 and Q8 yielded I-CVIs of 0.70 and mean clarity scores of 3.3 and 3.5 respectively; both were reconstructed after panel deliberation. The remaining eight items had I-CVIs of 0.90-1.00. Two new items were added on the basis of expert recommendations, producing a 12-item revised pool. In Round 2, all 12 items achieved I-CVIs of 0.90-1.00 and mean clarity scores of 4.2-4.9. Kendall's W for item relevance was 0.35 (P = 0.002) in Round 1 and 0.41 (P < 0.001) in Round 2; both values indicate weak-to-moderate agreement, and because the two rounds evaluated non-identical item sets these coefficients are reported as within-round descriptive summaries rather than as evidence of increasing consensus. A construct-alignment review conducted for this revision reassigned one item (Q7) to a different domain, reclassified one item (Q9) as a supplementary behavioural item excluded from domain scoring, and corrected an inconsistent Methods description of a third (Q3, whose fielded domain was unchanged); the response anchors for Q1 were also corrected to match its item stem. Because these changes post-dated Round 2, the affected items were re-rated in Round 3 against their revised roles: all ten experts participated (response rate 100%), every item attained an I-CVI of 1.00 with mean clarity scores of 4.2-5.0, modified kappa was 1.000 throughout, and S-CVI/Ave and S-CVI/UA were both 1.000 (Round-2 values 0.975 and 0.750). The resulting four-domain structure comprises 11 domain-scored items and one supplementary behavioural item, each supported by expert ratings obtained for the assignment it now carries. CONCLUSIONS: The three-round Delphi process produced a content-valid 12-item pool (PACT-12) organised into four theoretically derived domains, with satisfactory expert consensus on item relevance and clarity. These results reflect content validation only; patient cognitive interviews, exploratory factor analysis, reliability testing, and criterion-related validation are required before PACT-12 can be regarded as a psychometrically validated or clinically deployable instrument for evaluating patient-side trust shifts and associated safety risks in AI-assisted cardiovascular decision-making.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Development of the perceived AI-physician consistency and trust scale for cardiology: a Delphi-based study. — 科研速览 Science Skim