DivBrowse—interactive visualization and exploratory data analysis of variant call matrices

2023 | journal article. A publication with affiliation to the University of Göttingen.

Jump to: Cite & Linked | Documents & Media | Details | Version history

Cite this publication

​DivBrowse—interactive visualization and exploratory data analysis of variant call matrices​
König, P.; Beier, S.; Mascher, M.; Stein, N.; Lange, M. & Scholz, U.​ (2023) 
GigaScience12 art. giad025​.​ DOI: https://doi.org/10.1093/gigascience/giad025 

Documents & Media

License

GRO License GRO License

Details

Authors
König, Patrick; Beier, Sebastian; Mascher, Martin; Stein, Nils; Lange, Matthias; Scholz, Uwe
Abstract
Abstract Background The sequencing of whole genomes is becoming increasingly affordable. In this context, large-scale sequencing projects are generating ever larger datasets of species-specific genomic diversity. As a consequence, more and more genomic data need to be made easily accessible and analyzable to the scientific community. Findings We present DivBrowse, a web application for interactive visualization and exploratory analysis of genomic diversity data stored in Variant Call Format (VCF) files of any size. By seamlessly combining BLAST as an entry point together with interactive data analysis features such as principal component analysis in one graphical user interface, DivBrowse provides a novel and unique set of exploratory data analysis capabilities for genomic biodiversity datasets. The capability to integrate DivBrowse into existing web applications supports interoperability between different web applications. Built-in interactive computation of principal component analysis allows users to perform ad hoc analysis of the population structure based on specific genetic elements such as genes and exons. Data interoperability is supported by the ability to export genomic diversity data in VCF and General Feature Format 3 files. Conclusion DivBrowse offers a novel approach for interactive visualization and analysis of genomic diversity data and optionally also gene annotation data by including features like interactive calculation of variant frequencies and principal component analysis. The use of established standard file formats for data input supports interoperability and seamless deployment of application instances based on the data output of established bioinformatics pipelines.
Issue Date
2023
Journal
GigaScience 
eISSN
2047-217X
Language
English

Reference

Citations


Social Media