Gunikhan Sonowal, Parag Rughani, Gaurav Gogia
The exponential expansion of digital evidence across traditional computing systems, mobile devices, cloud infrastructures, and IoT ecosystems presents new issues for forensic investigators. Modern forensic acquisitions frequently produce multi-terabyte datasets with substantial inter-case redundancy, which significantly affects storage capacity, processing time, and investigation costs. Data deduplication is a potential approach; however, its use in forensic contexts is limited by special operational requirements, such as total data integrity, court-admissible audit trails, and resistance to anti-forensic manipulation. This paper provides a comprehensive survey and evaluation of forensic data deduplication methods. We propose a five-dimensional forensic taxonomy that classifies techniques based on granularity, forensic integrity, temporal placement, verification methodology, and scope & scale. Using multi-criteria decision analysis, we create a Deduplication Reliability Index (DRI) to compare existing systems to forensic requirements statistically. We also systematically identify six types of adversarial threats particular to forensic deduplication, such as hash collision exploitation, deduplication oracle attacks, and reference mapping manipulation. Our analysis shows that file-level and fixed-block deduplication have the highest DRI scores (4.35-4.30), indicating that they are optimally suited for forensic adoption, whereas semantic, probabilistic, and lossy techniques score less than 3.0 due to poor verification and integrity assurances. The report concludes by proposing crucial research goals, such as standardized validation frameworks, adversarial-resilient deduplication structures, and efficient zero-knowledge verification algorithms for forensics.