Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Albina Khasanova

dblp:426/2034 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0000-0001-6652-012XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
metagenomics
0.912025
SetBERT: the deep learning platform for contextualized embeddings and explainable predictions from high-throughput sequencing · Bioinform. 2025
Bioinformatics and computational biology › metagenomics
taxonomic classification
0.912025
SetBERT: the deep learning platform for contextualized embeddings and explainable predictions from high-throughput sequencing · Bioinform. 2025
Bioinformatics and computational biology › sequence analysis
high-throughput sequencing data analysis
0.312025
SetBERT: the deep learning platform for contextualized embeddings and explainable predictions from high-throughput sequencing · Bioinform. 2025

Methods — techniques the papers use, named apart from their topics

transformer · 0.9pre-training · 0.9deep learning · 0.9
YearPublicationVenuePosition
2025 SetBERT: the deep learning platform for contextualized embeddings and explainable predictions from high-throughput sequencing
abstract
MOTIVATION: High-throughput sequencing (HTS) is a modern sequencing technology used to profile microbiomes by sequencing thousands of short genomic fragments from the microorganisms within a given sample. This technology presents a unique opportunity for artificial intelligence to comprehend the underlying functional relationships of microbial communities. However, due to the unstructured nature of HTS data, nearly all computational models are limited to processing DNA sequences individually. This limitation causes them to miss out on key interactions between microorganisms, significantly hindering our understanding of how these interactions influence the microbial communities as a whole. Furthermore, most computational methods rely on post-processing of samples which could inadvertently introduce unintentional protocol-specific bias. RESULTS: Addressing these concerns, we present SetBERT, a robust pre-training methodology for creating generalized deep learning models for processing HTS data to produce contextualized embeddings and be fine-tuned for downstream tasks with explainable predictions. By leveraging sequence interactions, we show that SetBERT significantly outperforms other models in taxonomic classification with genus-level classification accuracy of 95%. Furthermore, we demonstrate that SetBERT is able to accurately explain its predictions autonomously by confirming the biological-relevance of taxa identified by the model. AVAILABILITY AND IMPLEMENTATION: All source code is available at https://github.com/DLii-Research/setbert. SetBERT may be used through the q2-deepdna QIIME 2 plugin whose source code is available at https://github.com/DLii-Research/q2-deepdna.
David W. Ludwig II, Christopher Guptil, Nicholas R. Alexander, Kateryna Zhalnina, Edi M.-L. Wipf, Albina Khasanova, Nicholas A. Barber, Wesley Swingley, Donald M. Walker, Joshua L. Phillips
Bioinform.6