Project Leads: This email address is being protected from spambots. You need JavaScript enabled to view it. , This email address is being protected from spambots. You need JavaScript enabled to view it.
DNA and RNA foundation models (FMs) have boomed in recent years, but their practical usefulness and the understanding of their inner workings remain limited. Models are frequently evaluated on outdated or confounded benchmark data, and comparisons often lack clear baseline performance. Furthermore, pilot studies reveal significant inconsistencies, showing large performance swings across seemingly similar tasks. To address these issues, this hackathon project will provide tools for collecting and quality-controlling (QCing) current datasets, alongside developing infrastructure to compare and interpret the performance of different FMs.
Building on the organizing labs’ established expertise in RNA bioinformatics and task-specific machine learning, the project focuses on RNA foundation models across a broad range of use cases. It directly aligns with the ELIXIR RNA Data Focus Group (emphasizing long-read epi-transcriptomics) and NFDI initiatives like GHGA, which seek infrastructure to select, deploy, and interpret genomic foundation and frontier models.
The project is structured around three core objectives:
- Curated Benchmark Datasets: Building on the Genomic Benchmarks toolkit, the team will develop open, FAIR-compliant infrastructure for collecting and QCing datasets. Utilizing data from established task-specific models (e.g., Ribonanza, SpliceAI, RiboNN), the benchmarks will cover RNA-binding protein target prediction, translation level prediction, RNA modification (epitranscriptomics), and splice site/variant prediction. All datasets will be released with FAIR metadata.
- Modular Evaluation Workflow: Pretrained whole-genome models (Nucleotide Transformer v2/v3, Evo v2) and RNA-specific models (RiNALMo, RNAfm, Ortrus, NucleicBERT, RNAErnie) will be integrated into a modular workflow based on the biotrainer architecture. The open workflow will support both zero-shot inference and lightweight fine-tuning while remaining extensible to new models and tasks.
- Interpretability Workflows: Standardized workflows will evaluate what foundation models actually learn and when they are applicable. By examining mRNA feature embeddings across local tasks (protein–RNA interactions) and long-range dependency tasks (translation, modifications), the project will relate embeddings to documented confounders to test whether models generalize to debiased variants or merely latch onto artifacts.
Expected Biohackathon Outcomes & Target Audience:
Deliverables include state-of-the-art "AI-ready" datasets with QC metadata, infrastructure to fine-tune and evaluate FMs, standardized interpretability workflows, and benchmark profiles characterizing task-specific capabilities. We welcome participants with expertise in genomics pipelines, running or fine-tuning foundation models, or domain-specific RNA bioinformatics.
