Khaled Kechida, Hamoudi Kalla, Mourad Brik
Considerable advances in software and hardware technologies now enable information systems to integrate heterogeneous data from multiple sources. Such integration helps address data scarcity by combining related datasets. However, the absence of standard formats and the unstructured nature of many datasets limit automated processing, interoperability, and practical analysis. Heterogeneity also arises because each facility manages its own data using different information systems, creating a major challenge for integration. The objective of this study is to integrate heterogeneous data sources using relevant attributes (e.g., name, date of birth, e-mail), which distinguishes our approach, rather than relying on identification keys that make integration manual and hard to generalise. The purpose of our method is to integrate all heterogeneous data for each individual into a unified database. For patients, it centralises medical data, supports care continuity, and cuts redundant examination costs. This approach can also be applied to education (academic records), commerce (customer data), and epidemiological studies. A case study illustrates the merging of data into a standardised XML (eXtensible Markup Language) format via a data processing pipeline, demonstrating the applicability of the proposed approach, and discusses the main challenges encountered and possible solutions. This methodology enhances efficiency, optimises resources, and informs decision-making.