Muhammed Emin Boylu, Beyza Zeynep Seçkin, Elifnaz Uyar, Fatma Betül Boylu, Faruk Kırbıyık, İsmet Kırpınar
Concordance was low, and the most influential factor was whether a specific diagnostic label had been recorded at all. Requiring a specific category on the consultation form and structured delirium screening are plausible service-level targets; the temporal and machine-learning findings require prospective confirmation.
OBJECTIVE: Recognition of psychiatric morbidity by non-psychiatric physicians is a quality indicator in consultation-liaison (C-L) psychiatry. We examined concordance between the provisional entries of referring physicians and the routine clinical diagnoses of C-L psychiatrists, and identified correlates of concordance, suicide-related presentations and delirium.
METHOD: We retrospectively reviewed 4556 adult psychiatric consultations (3070 patients) between January 2018 and October 2022. The C-L diagnosis was treated as an operational clinical reference, not a gold standard. Concordance was quantified using unweighted Cohen's kappa with a patient-clustered bootstrap. Four logistic-regression models with cluster-robust standard errors were fitted, with sensitivity analyses restricted to first consultations and specific provisional diagnoses.
RESULTS: Exact agreement occurred in 27.5 % of consultations (κ = 0.228, 95 % CI 0.209-0.248), rising to 37.2 % (κ = 0.296) among referrals with a specific provisional diagnosis and falling to 23.2 % (κ = 0.183) among index consultations. Agreement was lowest for adjustment disorder (0.6 %) and bipolar disorder with a manic episode (2.4 %). Non-diagnostic entries formed 42.2 % of referrals and dominated discordance. Delirium-assessment recording was strongly non-random (100 % in 2018-2019 versus 12.8 % in 2021); full-cohort analysis removed the gastrointestinal-delirium association (odds ratio 21.77 to 0.79) and reduced the COVID-19-period association to 1.64. Removing provisional-entry variables reduced random-forest AUROC from 0.875 to 0.738.
CONCLUSIONS: Concordance was low, and the most influential factor was whether a specific diagnostic label had been recorded at all. Requiring a specific category on the consultation form and structured delirium screening are plausible service-level targets; the temporal and machine-learning findings require prospective confirmation.