Anas Fauzi Masykuri, Wiwit Suryanto, Theodosius Marwan Irnaka, Bayu Pranata
Earthquake catalogues are fundamental data resources for seismic hazard assessment, tectonic studies, and the development of seismological methods. In tectonically active regions such as Indonesia, longterm earthquake records are particularly important for documenting seismicity across diverse geological settings and time periods. However, earthquake information for Indonesia is distributed across historical compilations, national catalogues, and international databases that differ in temporal coverage, reporting standards, magnitude scales, and data structures (Wichmann, 1918;Martin et al., 2022;ISC, 2025;USGS, 2025;GFZ, 2025). These differences complicate data integration and limit the direct reuse of available datasets.Historical earthquake information provides a valuable understanding of pre-instrumental seismicity but is often characterised by sparse spatial coverage and heterogeneous reporting. Concerning Indonesia, early documentary compilations (Wichmann, 1918) and macroseismic databases such as the Gempa Nusantara dataset (Martin et al., 2022) constitute the primary sources of historical seismic records. Instrumental earthquake data are systematically compiled at the national level by the Indonesian Agency for Meteorology, Climatology, and Geophysics (BMKG), with an openly accessible national catalogue covering the period 1998-2024(BMKG, 2025)). This catalogue has supported a range of studies on Indonesian seismicity, including investigations of tectonic processes, magnitude completeness, and seismic hazard parameters (Diantari, 2017;Hutchings and Mooney, 2021;Lewerissa et al., 2021).International earthquake information for the Indonesian region is provided by global catalogues maintained through the International Seismological Centre (ISC), USGS, and the German Research Centre for Geosciences (GFZ). These catalogues extend temporal coverage and improve documentation of regional and offshore events, but differ from national and historical datasets in event identification practices, magnitude reporting, and metadata structure.Comparable efforts to compile integrated or unified earthquake catalogues have been commenced in other tectonically active regions, including Greece, Northern Algeria, Guatemala, and the Caucasus region (Makropoulos et al., 2012;Boudebouda et al., 2024;González-Negreros and Gáspar-Escribano, 2024;Vorobieva et al., 2023Vorobieva et al., , 2024)). Although many of these studies apply magnitude homogenisation, differences among magnitude types reflect distinct physical principles used to quantify earthquake size (Kanamori, 1983). Integrated catalogues are widely used in downstream applications such as seismicity analysis, b-value mapping, magnitude completeness estimation, and probabilistic seismic hazard assessment (Wiemer and Wyss, 2000;Woessner and Wiemer, 2005;Sobiesiak et al., 2007;Taroni et al., 2021).The data report presents an integrated and standardised earthquake catalogue of Indonesia covering the period from 88 AD to 2024. The dataset compiles historical earthquake records, a national instrumental catalogue from BMKG, and international earthquake catalogues from ISC, USGS, and GFZ into a unified data structure. Data standardisation is applied separately to each source catalogue before catalogue integration to harmonise formats, coordinate reference systems, parameter units, and metadata fields, while preserving original magnitude values, magnitude types, and source identifiers.During the process of the analysis, no magnitude-scale homogenisation is applied. The resulting multimagnitude catalogue is intended to support transparent reuse in seismicity studies, regional-scale analyses, and methodological applications.In the context of an integrated catalogue, an earthquake event signifies the best-integrated representation of reported observations compiled from heterogeneous sources, rather than a fully resolved physical rupture. Each catalogue entry is constructed through the systematic integration of available reports describing the same occurrence in time and space, while preserving the original parameters and metadata provided by the source catalogues.Event identity in the catalogue is defined operationally through spatio-temporal matching criteria, not via physical source reconstruction or seismic waveform analysis. This method reflects the primary purpose of the dataset as a transparent archival compilation, designed to maximise traceability and reproducibility across historical and instrumental periods, respectively.Individual entries may represent aggregated descriptions derived from documentary or macroseismic sources, rather than a single, instrumentally constrained rupture, particularly for historical earthquakes. As a result, catalogue events should be interpreted as curated representations of reported seismic occurrences, rather than as physically unified seismic sources.The catalogue philosophy is intentionally conservative and non-interpretative in the context of the discussion. No attempt is made to resolve source complexity, reconcile conflicting physical models, or infer rupture characteristics. Rather, priority is placed on consistency, documentation, and explicit acknowledgment of uncertainty, supporting informed reuse and further methodological development by end users.The integrated and standardised earthquake catalogue covers the Indonesian region and adjacent tectonic domains, spanning approximately 94°E-142°E in longitude and 15°S-7°N in latitude. The temporal coverage extends from 88 AD to 2024, comprising both historical (pre-instrumental) and instrumental earthquake records.Instrumental events are generally associated with well-defined epicentral coordinates and origin times derived from national as well as international seismic networks. Historical earthquake records, primarily derived from documentary sources, are included to preserve the long-term seismic history of Indonesia. However, these records often show lower spatial and temporal precision. These historical events are treated as descriptive records rather than uniformly resolved seismic observations.The heterogeneous spatial and temporal data availability of the catalogue is shown in Figure 1a, 1b, and 1c. The apparent scarcity of events before 1900 reflects data sparsity and reporting limitations rather than data exclusion. All geographic coordinates are reported in the WGS84 reference system, while depth values are expressed in kilometres where available.Earthquake data are compiled from multiple authoritative sources representing historical, national, and international earthquake catalogues. Historical earthquake information is derived from documentary compilations and curated historical databases, including early works (Wichmann, 1918) and the Gempa Nusantara database. This provides macroseismic observations for historical earthquake in Indonesia between 1546 and 1950 (Martin et al., 2022).National instrumental earthquake data are obtained from the Indonesian BMKG through the curated dataset Indonesian Earthquake Catalogue (BMKG), 1998-2024. This is publicly available via Mendeley Data with a persistent DOI (BMKG, 2025).International earthquake information is sourced from catalogues maintained by the USGS (2025), ISC (2025), and GFZ (2025). These datasets complement national records by extending temporal coverage and improving documentation of regional and offshore events. All data sources are documented in the dataset metadata and shown in Table 1. This reports the number of events from each catalogue before and after duplicate event identification.The dataset is provided as a single UTF-8 encoded comma-separated values (CSV) file, where each row represents one earthquake event following multi-source integration and standardisation. Columns correspond to standardised event parameters, including event identification, origin time, geographic location, focal depth, reported magnitude values and types, source catalogue information, and supplementary remarks, particularly for historical events.Event identifiers are designed to ensure uniqueness across the integrated catalogue. Each collection entry includes a unique event identifier derived from the source catalogue, allowing direct traceability and unambiguous referencing after integration. For historical records lacking original identifiers, internally generated identifiers are assigned to maintain traceability while preserving source information. Missing or unavailable values are retained as empty fields to clearly distinguish data absence from numerical zero values. A complete description of all fields and data definitions is provided in the accompanying metadata file.The temporal distribution of events reflects pronounced variation in data density through time.Instrumental records dominate the catalogue after the early twentieth century, corresponding to the expansion of seismic monitoring networks. Earlier historical periods contain fewer events and show irregular reporting and variable temporal resolution.Magnitude information varies across sources and time periods, with multiple magnitude types reported depending on data availability and institutional practice. Original magnitude values and types are preserved to ensure transparency and reproducibility. The catalogue is suited to descriptive and exploratory analyses across long temporal scales, while acknowledging differences in data completeness as well as uncertainty.The Mc is estimated using the MAXC method to provide a descriptive diagnosis of temporal changes in data reporting. Importantly, Mc is not computed for the entire catalogue as a single entity. The analysis is restricted to the three magnitude types with the largest number of events, namely MLv_BMKG, M_BMKG, and mb_ISC, as shown in Figure 1d, 1e, and 1f. This is to avoid mixing heterogeneous magnitude scales and reporting practices in a single frequency-magnitude distribution. Therefore, the resulting Mc versus time curves are intended as catalogue diagnostics, reflecting changes in data availability, reporting practices, and network development in each dominant data stream, rather than as physically unified detection thresholds. No magnitude homogenisation is implied, and the Mc values are not interpreted as a completeness level applicable across all magnitude types or historical periods.The dataset is released as a structured CSV file accompanied by detailed metadata documentation. The standardised tabular format facilitates direct use with common statistical, geospatial, and seismological analysis tools without the need for additional preprocessing. Dataset versioning and long-term accessibility are ensured through deposition in a public data repository.Earthquake data were collected from historical documentary sources, a national earthquake catalogue, and multiple international earthquake databases. All datasets were acquired in the original formats and preserved as primary inputs before any processing or integration.Historical earthquake events were compiled from documentary sources and curated historical databases based on macroseismic observations, including early documentary compilations (Wichmann, 1918) and the Gempa Nusantara dataset. This provided systematically curated macroseismic information for historical earthquake in Indonesia between 1546 and 1950 (Martin et al., 2022).National instrumental earthquake records were obtained from the Indonesian Earthquake Catalogue released by the Indonesian BMKG, covering the period 1998-2024 (BMKG, 2025). This dataset served as the primary national reference for instrumental seismicity and provided standardised parameters suitable for integration with other sources.International instrumental earthquake data were obtained from global catalogues maintained by the USGS, ISC, and GFZ. ISC data were sourced specifically from the ISC Bulletin, which provided reviewed global earthquake data compiled from multiple contributing agencies. These global datasets were used to complement national records by extending temporal coverage and providing additional information for regional as well as offshore earthquake in and around the Indonesian region.All collected datasets were consolidated into a unified data structure through systematic integration procedures. Data standardisation was applied before catalogue combining, with each source catalogue processed separately to harmonise formats, coordinate reference systems, parameter units, and metadata fields. Source identifiers and catalogue references were retained to ensure that each event remained traceable to its originating dataset. Figure 2 showed the workflow used for multi-source data acquisition, standardisation, and duplicate event identification in the compilation of the earthquake catalogue.Data standardisation was performed to harmonise heterogeneous datasets into a consistent and interoperable catalogue structure while preserving the original characteristics of each data source. The standardisation process focused on matching data formats, measurement units, coordinate reference systems, and metadata fields.Event time information was harmonised to Coordinated Universal Time (UTC) for instrumental earthquake records where original time references were available. Historical earthquake records were exempted from time-zone conversion, and the original reported temporal information was preserved to maintain data authenticity and reflect the limitations of pre-instrumental sources.Geographic coordinates were standardised to a common geographic coordinate system, and depth values were converted to consistent units where necessary. Magnitude values and magnitude types were retained as reported in the original catalogues. No magnitude-scale conversion or full magnitude homogenisation was applied. Rather, the catalogue preserved multiple magnitude types from different sources in a standardised metadata framework.The integration of multiple earthquake catalogues led to duplicate records for the same seismic event.Potential duplicate events were identified using spatial and temporal thresholds (Δt ≤ 60 s and Δd ≤ 50 km), consistent with values commonly applied in regional earthquake catalogue integration studies (Weatherill et al., 2016;Holmgren et al., 2023). When duplicate candidates were identified, event selection followed predefined source-priority rules while preserving original metadata for traceability.For historical records, these thresholds were applied pragmatically and with explicit acknowledgment of the limitations.Data validation procedures were implemented to assess internal consistency, traceability, and usability of the integrated earthquake catalogue rather than to independently reassess individual events. Validation focused on basic metadata attributes, including event date and time formats, geographic coordinates, depth values, magnitude information, and source identifiers.Cross-comparisons among datasets were conducted to identify obvious inconsistencies, implausible parameter ranges, and potential spatio-temporal duplicates arising from multi-source integration. Event identifiers and associated metadata were examined to confirm consistency between reported origin time, location, and source catalogue information.Historical earthquake records were treated separately during validation in this study. Given the inherent limitations of documentary and macroseismic sources, no modern instrumental validation criteria were imposed on historical events, and original reported information was preserved. The validation process aimed to improve catalogue coherence and transparency while maintaining the original characteristics and documented uncertainties of the source data.Technical validation was conducted to assess the internal coherence, completeness, and usability of the integrated and standardised earthquake catalogue. The validation strategy focused on evaluating the consistency of compiled records across multiple data sources, rather than on independent verification or reinterpretation of individual events. No re-estimation of seismological parameters was performed during the process of this study.Validation procedures included systematic checks of key event attributes, such as origin time, geographic location, depth, and reported magnitude information. These checks were designed to identify implausible parameter ranges, formatting inconsistencies, and residual spatio-temporal duplications that persisted following catalogue integration. Review statistics and distributional patterns were examined to ensure that standardised parameters showed physically reasonable ranges and continuity through time.The heterogeneous nature of the dataset across historical and instrumental periods was explicitly considered. For historical earthquake records, validation prioritised documentation completeness and traceability to sources rather than numerical precision. Events were retained even when spatial or temporal uncertainties were high, provided that source attribution and descriptive information were visibly documented.Concerning instrumental records, cross-source consistency was evaluated by comparing overlapping parameters reported by different catalogues. Identified discrepancies were assessed to ensure that retained parameters provided a coherent and traceable event description. The validation process supported reproducibility at the level of data compilation and standardisation, while acknowledging that uncertainties inherent to seismic observations and historical documentation could not be fully eliminated. Maps were generated using QGIS, and descriptive plots as well as visualisations were produced using Python.This dataset was intended for descriptive, exploratory, and methodological applications that require a standardised compilation of earthquake records over long temporal periods at regional to national scales. It was explicitly designed as a transparent, non-homogenised archival catalogue rather than as a physically unified seismic reference catalogue.The completeness, precision, and reporting characteristics of the catalogue varied substantially through time. Instrumental earthquake records dominated after the early twentieth century, reflecting the development and expansion of seismic monitoring networks in Indonesia. Consequently, historical earthquake records derived from documentary and macroseismic sources were fewer and commonly characterised by lower spatial and temporal resolution. These historical records were provided to preserve long-term seismic context and should be interpreted as descriptive observations rather than instrumentally constrained seismic events.The catalogue was not used directly for quantitative seismological analyses such as probabilistic seismic hazard assessment, b-value estimation, seismicity rate modelling, or recurrence analysis without substantial additional processing. These applications required expert-driven procedures including magnitude homogenisation, uncertainty modelling, completeness assessment modified to specific magnitude types, and rigorous event validation (Wiemer and Wyss, 2000;Woessner and Wiemer, 2005;Sobiesiak et al., 2007;Taroni et al., 2021).Magnitude information was reported exactly as provided by the source catalogues and included multiple magnitude types that were not physically equivalent. No magnitude-scale homogenisation was applied, and magnitude values were not treated as directly comparable across different sources or time periods. Estimates of Mc were intended solely as descriptive catalogue diagnostics reflecting changes in data availability and reporting practices, rather than as physically meaningful detection thresholds.Users were responsible for applying appropriate filtering, preprocessing, and methodological choices consistent with individual-specific study objectives. Earthquake data were collected from historical documentary sources, national catalogues released by BMKG, and international earthquake databases. Each source catalogue was standardised independently before combining. Duplicate events were identified using spatial and temporal criteria, leading to a unified catalogue that preserved original magnitude information without magnitude-scale homogenisation.Table 1. Number of events contributed by each catalogue source before and after the identification of duplicate event process used to compile the unified earthquake catalogue.