Weijing Dang, Liang Lyu, Yong Shen, Ji Wang, Yi Yu, Qirui Wang, Mingjin Zhang, Jiajia Zheng, Dawei Liu
Repeatable acquisition of smiling and rest-related facial states is essential for longitudinal three-dimensional (3D) facial imaging, yet the effect of task instructions may depend on the state and measurement. This within-participant study compared conventional (CONV) and state-specific phoneme-guided (PHON) acquisition in 162 adults. For SMILE, CONV elicited a posed smile, whereas PHON used /tʃ/ articulation followed by maintenance of the resulting perioral configuration; for the REST-related condition, CONV elicited relaxed rest, whereas PHON used sustained /m/ to define a closed-lip reference configuration rather than unconstrained physiological rest. Each protocol-state condition was recorded twice. Within-session repeatability was characterized by absolute error between repeated acquisitions, technical error of measurement, intraclass correlation coefficient [ICC(2,1)], and Bland-Altman agreement. State-specific protocol differences were assessed using paired Wilcoxon signed-rank tests with Holm correction. During SMILE, PHON reduced median error for the 3D Sn-Gn distance (3.79 vs. 2.36 mm), interlabial distance (3.48 vs. 1.84 mm), and sagittal lip step (1.33 vs. 0.66 mm; adjusted p < 0.001), but increased commissure-distance error (1.50 vs. 1.89 mm; adjusted p = 0.007). During the REST-related condition, no prespecified linear outcome differed after correction. Phoneme guidance changed, rather than uniformly improved, repeatability, and its effect depended on facial state and outcome.