
ReferenceSeeker
ReferenceSeeker – Rapid Identification of Suitable Reference Genomes
ReferenceSeeker is a tool for the rapid identification of closely related reference genomes for bacterial, archaeal, fungal, protozoan, and viral genome sequences. It applies a scalable hierarchical approach that combines fast k-mer profile-based database searches with the calculation of Average Nucleotide Identity (ANI) values to identify the most suitable reference genomes for downstream analyses. By reducing the number of candidate genomes requiring detailed comparison, ReferenceSeeker provides a fast and efficient solution for reference genome selection.
Key benefits
Rapid identification of closely related reference genomes
Combines fast k-mer-based screening with accurate ANI calculations
Scalable approach suitable for large reference genome collections
Supports multiple taxonomic groups, including bacteria, archaea, fungi, protozoa, and viruses
Reduces computational effort by focusing on the most promising candidates
Facilitates reproducible and standardized reference genome selection
Applications
Selection of suitable reference genomes for comparative genomics
Identification of closely related genomes for microbial characterization
Support for genome assembly validation and quality assessment
Reference genome selection for variant calling and phylogenetic analyses
Taxonomic and evolutionary studies based on genome similarity
Preparation of downstream bacterial and microbial genomics workflows
Intended use
ReferenceSeeker is intended for microbiologists, bioinformaticians, microbial genomicists, and comparative genomics researchers who require rapid and reliable identification of suitable reference genomes. It is particularly suited for users working with newly assembled genomes and seeking high-quality references for comparative analyses, annotation, phylogenetics, and genome characterization.
Contact:
Website https://github.com/oschwengers/referenceseeker/issues
