Nidhal Tarhouni, Ahmed Bayoudh, Amira Mahfoudhi, Bilel Hadrich, Karim Kriaa, Imen Kallel
Peptidylarginine deiminase 4 (PAD4) is an increasingly prominent therapeutic target in oncology, inflammatory disease, and neutrophil extracellular trap (NET)-associated pathologies, yet public bioactivity data for PAD4 inhibitors remain fragmented across multiple repositories with substantial redundancy and inconsistent annotation. Here, we present PAD4-DB, a curated structure-activity relationship resource integrating 3093 unique inhibitors (consensus pIC50 range 2.00-8.52; median 6.84) from PubChem, ChEMBL, and BindingDB through a reproducible pipeline encompassing structure standardization, activity normalization, source-independence assessment, and deduplication. Quantitative analysis of 358,416 compound pairs with Tanimoto similarity ≥0.6 indicates that the sampled PAD4 SAR landscape is predominantly smooth: only 94 pairs meeting the stringent activity-cliff criterion (Tanimoto ≥ 0.8; |ΔpIC50| ≥ 2.0) were identified, representing 0.026% of all related pairs and 0.78% of the 12,071 cliff-candidate pairs. Of these, 80 (85.1%) received structural support from matched molecular pair (MMP) analysis. Severe cliffs were non-randomly distributed, with four hub compounds collectively accounting for 53.2% of all severe cliff pairs, while 96.8% of multi-member scaffold series remained completely smooth. Provenance analysis further showed that 82.9% of the dataset was classified as pipeline-dependent under the provenance-scoring framework across repositories rather than independent source measurements, underscoring the importance of provenance-aware confidence weighting in downstream modeling. PAD4-DB therefore provides a reproducible foundation for PAD4 inhibitor discovery and a curated benchmark for evaluating similarity-based and machine-learning approaches in a chemically structured SAR landscape containing rare but highly concentrated activity cliffs.