科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of exposure science & environmental epidemiology2026-08-17

A proposed organization of multifaceted knowledge in data repositories: structuring record collections and evidence sets for data lakehouses and moonshots.

Adam W Kiefer, Snigdha Sompalli, Jason P Mihalik, Robert Hubal

一句话结论 · In one sentence

Scientists supported through large research programs, such as those funded by government initiatives or consortia, are increasingly encouraged (or required) to share their diverse data in public repositories in FAIR format. The proposed framework provides a flexible, cloud-based platform that meets these requirements by enabling the harmonization and secure sharing of multifaceted data across studies. Designed for scalability and interoperability with established data models, the platform supports both current research needs and a flexible structure responsive to those needs for use across future large-scale precision health initiatives.

原始摘要(英文原文)· Original abstract
BACKGROUND: We are developing a cloud-based neurohealth data platform to aggregate and store data from large, complementary historical and ongoing studies focused on traumatic brain injury (TBI), a condition shaped by complex interactions among biological, environmental and contextual factors, and to enable secure, flexible querying across studies. These multifaceted datasets span clinical, physiological, behavioral, and environmental domains, mirroring the complexity of comparable multisite efforts such as the exposome moonshot. OBJECTIVE: Define a scalable and interoperable framework for organizing and harmonizing multifaceted data derived from different sources, purposes, and time scales, supporting precision neurohealth research and translational applications. METHODS: We designed and have begun implementing the Precision Health for Enhanced Neuro-Operational Modeling (PHENOM) platform, using a hybrid lakehouse architecture deployed on the public cloud. PHENOM employs automated extract-transform-load pipelines, metadata templates, and standardized biomedical terminologies to support FAIR-compliant data integration and cross-domain interoperability. We have mapped our findings to our datasets of interest in traumatic brain injury (TBI) and to the knowledge being generated for moonshot exposome efforts. RESULTS: The PHENOM framework organizes data into four general "bins": (1) One-time capture, (2) regular or continual lab or site visits, (3) personalized electronic health records, and (4) collective measures. The architecture enables ingesting multimodal data and supports AI-driven analytic tools and digital twin modeling through public cloud services. Initial deployment has demonstrated automated harmonization across multiple TBI studies and scalable query performance across datasets. SIGNIFICANCE: Scientists supported through large research programs, such as those funded by government initiatives or consortia, are increasingly encouraged (or required) to share their diverse data in public repositories in FAIR format. The proposed framework provides a flexible, cloud-based platform that meets these requirements by enabling the harmonization and secure sharing of multifaceted data across studies. Designed for scalability and interoperability with established data models, the platform supports both current research needs and a flexible structure responsive to those needs for use across future large-scale precision health initiatives. IMPACT STATEMENT: This Perspective proposes a cloud framework for structuring multifaceted data derived from complementary or consortia-related studies using four general "bins": (1) One-time capture, (2) regular or continual lab or site visits, (3) personalized health records, and (4) collective measures. The proposed framework offers a flexible structure for use across many existing research programs that are increasingly encouraged (or required) to share their diverse data in public repositories in FAIR format.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A proposed organization of multifaceted knowledge in data repositories: structuring record collections and evidence sets for data lakehouses and moonshots. — 科研速览 Science Skim