André Miguel Viegas Oliveira, Rafael Geraldo Dos Santos, Helder Daniel, José Valente de Oliveira
Portuguese is spoken by approximately 300 million people, making it one of the most widely spoken languages in the world, with speakers distributed across every continent. Despite its global reach, Portuguese remains significantly underrepresented in current synthetic speech detection technologies. This dataset aims to address this critical gap in the existing detection solutions. This dataset may be used in tasks related to biometric security, speaker verification, anti-spoofing systems, and deepfake protection, supporting the development and evaluation of models for Portuguese language speech data. The Synthetic Speech Detection - Portuguese (SSD_PT) dataset consists of 242,156 audio samples in the Portuguese language, containing short speech segments ranging from 1 to 7 s each. The dataset is organized into two balanced classes: genuine speech samples (bonafide) and synthetic speech samples (spoof). For the Brazilian Portuguese portion of the dataset, the genuine speech samples were collected from publicly available repositories, namely VoxForge, Mozilla Common Voice, and OpenSLR. As no publicly available repositories with comparable coverage were found for the remaining Portuguese varieties, the genuine speech samples for the remaining portions of the dataset were collected from television interviews, radio programs, and podcasts representing different regions of the Lusophone world, including mainland Portugal, the Azores, Madeira, Angola, and Mozambique. The synthetic speech samples were generated from genuine speech material. For Brazilian Portuguese, a text-to-speech model based on Coqui TTS (XTTS-v2) was used. For other Portuguese variants, voice conversion techniques were applied using the RVC (Retrieval-based Voice Conversion) framework, requiring the creation of speaker-specific models trained on multiple audio segments to capture individual vocal characteristics. All audio files were standardized to a 16 kHz sampling rate and converted to mono format. The dataset is split into training, development, and evaluation subsets, accompanied by protocol files that associate each sample with its corresponding label.