VLDB 2026 Research / reviewers in the wild / expert
Maximillian G. Marin
dblp:271/0895
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-9108-3328ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
comparative genomics |
1.6 | 2 | 2025 | Pitfalls of bacterial pan-genome analysis approaches: a case study of Mycobacterium tuberculosis and two less clonal bacterial species · Bioinform. 2025 Exploring gene content with pangene graphs · Bioinform. 2024 |
Bioinformatics and computational biology › comparative genomics › pangenomics
pan-genome analysis |
0.9 | 1 | 2025 | Pitfalls of bacterial pan-genome analysis approaches: a case study of Mycobacterium tuberculosis and two less clonal bacterial species · Bioinform. 2025 |
Bioinformatics and computational biology › comparative genomics › pangenomics
pangenome graph |
0.8 | 1 | 2024 | Exploring gene content with pangene graphs · Bioinform. 2024 |
Bioinformatics and computational biology › comparative genomics
pangenomics |
0.8 | 1 | 2024 | Exploring gene content with pangene graphs · Bioinform. 2024 |
Bioinformatics and computational biology › genomics
variant calling |
0.6 | 1 | 2022 | Benchmarking the empirical accuracy of short-read sequencing across the M. tuberculosis genome · Bioinform. 2022 |
Bioinformatics and computational biology
genome annotation |
0.3 | 1 | 2025 | Pitfalls of bacterial pan-genome analysis approaches: a case study of Mycobacterium tuberculosis and two less clonal bacterial species · Bioinform. 2025 |
Bioinformatics and computational biology › genomics › genome sequencing
whole genome sequencing |
0.2 | 1 | 2022 | Benchmarking the empirical accuracy of short-read sequencing across the M. tuberculosis genome · Bioinform. 2022 |
Methods — techniques the papers use, named apart from their topics
pan-genome pipeline comparison · 0.9protein sequence alignment · 0.8graph construction · 0.8repetitive sequence masking · 0.6mapping quality filtering · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pitfalls of bacterial pan-genome analysis approaches: a case study of Mycobacterium tuberculosis and two less clonal bacterial speciesabstractSUMMARY: Pan-genome analysis is a fundamental tool for studying bacterial genome evolution; however, the variety in methods used to define and measure the pan-genome poses challenges to the interpretation and reliability of results. Using Mycobacterium tuberculosis, a clonally evolving bacterium with a small accessory genome, as a model system, we systematically evaluated sources of variability in pan-genome estimates. Our analysis revealed that differences in assembly type (short-read versus hybrid), annotation pipeline, and pan-genome software, significantly impact predictions of core and accessory genome size. Extending our analysis to two additional bacterial species, Escherichia coli and Staphylococcus aureus, we observed consistent tool-dependent biases but species-specific patterns in pan-genome variability. Our findings highlight the importance of integrating nucleotide- and protein-level analyses to improve the reliability and reproducibility of pan-genome studies across diverse bacterial populations. AVAILABILITY AND IMPLEMENTATION: Panqc is freely available under an MIT license at https://github.com/maxgmarin/panqc. Maximillian G. Marin, Natalia Quinones-Olvera, Christoph Wippel, Mahboobeh Behruznia, Brendan M. Jeffrey, Michael Harris, Brendon C. Mann, Alex Rosenthal, Karen R. Jacobson, Robin M. Warren, Conor J. Meehan, Maha R. Farhat |
Bioinform. | 1 |
| 2024 | Exploring gene content with pangene graphsabstractMOTIVATION: The gene content regulates the biology of an organism. It varies between species and between individuals of the same species. Although tools have been developed to identify gene content changes in bacterial genomes, none is applicable to collections of large eukaryotic genomes such as the human pangenome. RESULTS: We developed pangene, a computational tool to identify gene orientation, gene order, and gene copy-number changes in a collection of genomes. Pangene aligns a set of input protein sequences to the genomes, resolves redundancies between protein sequences and constructs a gene graph with each genome represented as a walk in the graph. It additionally finds subgraphs, which we call bibubbles, that capture gene content changes. Applied to the human pangenome, pangene identifies known gene-level variations and reveals complex haplotypes that are not well studied before. Pangene also works with high-quality bacterial pangenome and reports similar numbers of core and accessory genes in comparison to existing tools. AVAILABILITY AND IMPLEMENTATION: Source code at https://github.com/lh3/pangene; prebuilt pangene graphs can be downloaded from https://zenodo.org/records/8118576 and visualized at https://pangene.bioinweb.org. Maximillian G. Marin, Maha R. Farhat |
Bioinform. | 2 |
| 2022 | Benchmarking the empirical accuracy of short-read sequencing across the M. tuberculosis genomeabstractMOTIVATION: Short-read whole-genome sequencing (WGS) is a vital tool for clinical applications and basic research. Genetic divergence from the reference genome, repetitive sequences and sequencing bias reduces the performance of variant calling using short-read alignment, but the loss in recall and specificity has not been adequately characterized. To benchmark short-read variant calling, we used 36 diverse clinical Mycobacterium tuberculosis (Mtb) isolates dually sequenced with Illumina short-reads and PacBio long-reads. We systematically studied the short-read variant calling accuracy and the influence of sequence uniqueness, reference bias and GC content. RESULTS: Reference-based Illumina variant calling demonstrated a maximum recall of 89.0% and minimum precision of 98.5% across parameters evaluated. The approach that maximized variant recall while still maintaining high precision (<99%) was tuning the mapping quality filtering threshold, i.e. confidence of the read mapping (recall = 85.8%, precision = 99.1%, MQ ≥ 40). Additional masking of repetitive sequence content is an alternative conservative approach to variant calling that increases precision at cost to recall (recall = 70.2%, precision = 99.6%, MQ ≥ 40). Of the genomic positions typically excluded for Mtb, 68% are accurately called using Illumina WGS including 52/168 PE/PPE genes (34.5%). From these results, we present a refined list of low confidence regions across the Mtb genome, which we found to frequently overlap with regions with structural variation, low sequence uniqueness and low sequencing coverage. Our benchmarking results have broad implications for the use of WGS in the study of Mtb biology, inference of transmission in public health surveillance systems and more generally for WGS applications in other organisms. AVAILABILITY AND IMPLEMENTATION: All relevant code is available at https://github.com/farhat-lab/mtb-illumina-wgs-evaluation. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Maximillian G. Marin, Roger Vargas, Michael Harris, Brendan M. Jeffrey, L. Elaine Epperson, David Durbin, Michael Strong, Max Salfinger, Zamin Iqbal, Irada Akhundova, Sergo Vashakidze, Valeriu Crudu, Alex Rosenthal, Maha R. Farhat |
Bioinform. | 1 |