Project Leads: This email address is being protected from spambots. You need JavaScript enabled to view it.This email address is being protected from spambots. You need JavaScript enabled to view it.

The exposome, defined as the totality of environmental exposures across the human lifespan, is increasingly recognised as a major determinant of disease risk, contributing substantially to premature mortality from chronic conditions such as cardiovascular disease, respiratory illness, and cancer. Large-scale cohort initiatives, including UK Biobank, NAKO, and All of Us, have advanced the characterisation of the external exposome (also referred to as place-based exposures) by integrating satellite data, geospatial modelling, and linked administrative records to participants, creating new opportunities to study the interaction between gene and environment. 

However, despite growing interest in GxE, the lack of a large-scale synthetic training dataset coupling population-wide genomics with rich place-based exposure data currently hampers the ability to benchmark computational models, develop generalised methods, and bridge genomics, exposomics, and epidemiology communities. Synthetic data overcomes the need for special access permissions required for participant data, providing a readily accessible and interoperable resource for global methods development. 

This project proposes to develop an open, synthetic, multi-modal dataset integrating genomic, phenotypic, and exposomic data. The data will be generated using statistical and computational models (for creating synthetic genomic data e.g., msprime and stdpopsim) for phenotypic data, e.g., GEPSi that mimic the structure and correlations of real-world data without compromising privacy, thereby enabling unrestricted use. 

The project is co-led by developers of CLUES, an open-source framework for accessing environmental data and integrating it with genetic and phenotypic data. CLUES provides exposomic data spanning domains such as climate, atmospheric pollution, green and blue space, urbanicity, and socioeconomic exposures. CLUES will serve as the basis for the exposome data integration in the project. Outputs will be developed in alignment with the GA4GH Human Exposome Data Standards Study Group, ensuring interoperability with international standards for representing place-based exposure data. 

During the hackathon, we aim to prototype the synthetic data generation pipeline, define a machine-readable schema for the exposomic domains, and benchmark initial integration with genomic simulation approaches, producing a FAIR-compliant resource that supports method development, training, and benchmarking without the barriers associated with controlled-access data.

This project directly supports ELIXIR's growing efforts at the intersection of environmental science and human health, in particular the Environmental Impact Focus Group, which brings together climate science, biodiversity, toxicology, and health data with geolocation. It further aligns with ELIXIR through the development of a machine-readable schema and software prototype consistent with ELIXIR data standards. 

By extending established synthetic-data paradigms from genomics to include environmental exposures, the project addresses a recognised gap in FAIR, openly accessible exposomic resources, and contributes to de.NBI and ELIXIR Germany's mission of providing standardised, sustainable bioinformatics infrastructure for the life sciences community.