David van der Vloed
Automatic speaker recognition (ASR) is a method used in forensic speaker comparison (FSC) casework. In FSC cases two voice recordings are compared to make an inference about speaker identity. When employing ASR for this task, the case audio is compared by software, resulting in a case score. This case score needs to be converted (calibrated) to a likelihood ratio, the formal metric of strength of evidence. An LR calculation set (calibration set), i.e. audio data representative of the case audio files, is needed to train a mapping between scores and likelihood ratios specific to the case. What audio can be deemed as representative of the case audio and what audio conditions to use as selection criteria is a subjective, case-specific decision by the forensic practitioner. This decision can be informed by research into the effect of audio conditions on the score-to-likelihood-ratio mappings and the resulting likelihood ratios. Using forensically relevant data, this type of research is done in this work. Multiple conditions are investigated in this manner, along with an analysis of the influence of number of speakers in the LR calculation set to serve as a baseline for variability. The results suggest duration differences below 40 s should be accounted for. Ethnolect differences, whether recordings are made in a car and speech durations above about 40 s appear not impactful. The same result is found for gender and the difference between operational data / lab data, except when cross-comparing those conditions.