Ian T. Adams, Kyle McLean, Paige E. Vaughn, Alexis Fabila, Geoffrey Alpert
Artificial intelligence (AI) tools that analyze body-worn camera footage are increasingly used to monitor police behavior, yet the behavioral measures they generate have not been validated against human perception. We conduct the first human-perception validation of an AI-based police professionalism measure using a national sample of 3,020 U.S. adults. Respondents watched body-worn camera videos pre-classified by Truleo’s natural language processing system as Below Standard, Standard, or High professionalism and rated officer professionalism and respectful communication. Cross-classified multilevel models (n = 15,100) show that respondents rate Below Standard encounters 0.87 points lower than Standard on a three-point scale (p < .001), but cannot distinguish Standard from High professionalism. This negativity bias means unprofessional conduct is immediately recognizable, while positive markers of exemplary professionalism are not apparent to untrained observers. These results provide partial validation: agencies can treat Below Standard flags as reliable, while distinguishing adequate from exemplary performance requires human judgment.