EDBT 2026 Demo / reviewers in the wild / expert
Xiaofang Jiang
dblp:140/3988
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-0955-8284ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › sequence analysis
DNA language model |
0.9 | 1 | 2025 | The impact of tokenizer selection in genomic language models · Bioinform. 2025 |
Bioinformatics and computational biology › sequence analysis › DNA language model
genome tokenization |
0.9 | 1 | 2025 | The impact of tokenizer selection in genomic language models · Bioinform. 2025 |
Bioinformatics and computational biology › statistical genetics
genotype-phenotype association |
0.7 | 1 | 2023 | Evolink: a phylogenetic approach for rapid identification of genotype-phenotype associations in large-scale microbial multispecies data · Bioinform. 2023 |
Bioinformatics and computational biology › genomics
microbial genomics |
0.7 | 1 | 2023 | Evolink: a phylogenetic approach for rapid identification of genotype-phenotype associations in large-scale microbial multispecies data · Bioinform. 2023 |
Bioinformatics and computational biology › genomics
computational genomics |
0.3 | 1 | 2025 | The impact of tokenizer selection in genomic language models · Bioinform. 2025 |
Bioinformatics and computational biology › sequence analysis
sequence classification |
0.3 | 1 | 2025 | The impact of tokenizer selection in genomic language models · Bioinform. 2025 |
Methods — techniques the papers use, named apart from their topics
state space model · 0.9k-mer tokenization · 0.9character tokenization · 0.9byte pair encoding · 0.9phylogenetics · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The impact of tokenizer selection in genomic language modelsabstractMOTIVATION: Genomic language models have recently emerged as a new method to decode, interpret, and generate genetic sequences. Existing genomic language models have utilized various tokenization methods, including character tokenization, overlapping and nonoverlapping k-mer tokenization, and byte-pair encoding, a method widely used in natural language models. Genomic sequences differ from natural language because of their low character variability, complex and overlapping features, and inconsistent directionality. These features make subword tokenization in genomic language models significantly different from both traditional language models and protein language models. RESULTS: This study explores the impact of tokenization in genomic language models by evaluating their downstream performance on 44 classification fine-tuning tasks. We also perform a direct comparison of byte pair encoding and character tokenization in Mamba, a state-space model. Our results indicate that character tokenization outperforms subword tokenization methods on tasks that rely on nucleotide-level resolution, such as splice site prediction and promoter detection. While byte-pair tokenization had stronger performance on the SARS-CoV-2 variant classification task, we observed limited statistically significant differences between tokenization methods on the remaining downstream tasks. AVAILABILITY AND IMPLEMENTATION: Detailed results of all benchmarking experiments are available in https://github.com/leannmlindsey/DNAtokenization. Training datasets and pretrained models are available at https://huggingface.co/datasets/leannmlindsey. Datasets and processing scripts are available at doi: 10.5281/zenodo.16287401 and doi: 10.5281/zenodo.16287130. LeAnn Lindsey, Nicole L. Pershing, Anisa Habib, Keith Dufault-Thompson, W. Zac Stephens, Anne J. Blaschke, Xiaofang Jiang, Hari Sundar |
Bioinform. | 7 |
| 2023 | Evolink: a phylogenetic approach for rapid identification of genotype-phenotype associations in large-scale microbial multispecies dataabstractMOTIVATION: The discovery of the genetic features that underly a phenotype is a fundamental task in microbial genomics. With the growing number of microbial genomes that are paired with phenotypic data, new challenges, and opportunities are arising for genotype-phenotype inference. Phylogenetic approaches are frequently used to adjust for the population structure of microbes but scaling them to trees with thousands of leaves representing heterogeneous populations is highly challenging. This greatly hinders the identification of prevalent genetic features that contribute to phenotypes that are observed in a wide diversity of species. RESULTS: In this study, Evolink was developed as an approach to rapidly identify genotypes associated with phenotypes in large-scale multispecies microbial datasets. Compared with other similar tools, Evolink was consistently among the top-performing methods in terms of precision and sensitivity when applied to simulated and real-world flagella datasets. In addition, Evolink significantly outperformed all other approaches in terms of computation time. Application of Evolink on flagella and gram-staining datasets revealed findings that are consistent with known markers and supported by the literature. In conclusion, Evolink can rapidly detect phenotype-associated genotypes across multiple species, demonstrating its potential to be broadly utilized to identify gene families associated with traits of interest. AVAILABILITY AND IMPLEMENTATION: The source code, docker container, and web server for Evolink are freely available at https://github.com/nlm-irp-jianglab/Evolink. Yiyan Yang, Xiaofang Jiang |
Bioinform. | 2 |
| 2021 | High-throughput sequencing of SARS-CoV-2 in wastewater provides insights into circulating variants
Rafaela S. Fontenele, Simona Kraberger, James Hadfield, Erin M. Driver, Devin Bowes, LaRinda A. Holland, Temitope O. Faleye, Sangeet Adhikari, Rosa Inchausti, Wydale K. Holmes, Stephanie Deitrick, Darrell Duty, Aruni Bhatnagar, Ray A. Yeager, Rochelle H. Holm, Kevin Dixon, Tim Constantine, Melissa A. Wilson, Efrem S. Lim, Xiaofang Jiang, Rolf U. Halden, Matthew Scotch, Arvind Varsani |
AMIA | 22 |