Ying Wang 0005

dblp:94/3104-5 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0001-8766-5950ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 8 since 2021
YearPublicationVenuePosition
2026 IPRLLM: Iterative Prompt Refinement with Large Language Models for Enhancing Phage-Host Association Prediction
Feng Zhou 0021, Jiaxing Bai, Ying Wang 0005
ISBRA (2)5
2025 PhyMicroNet: Inferring Directed 'Species Interactions in Microbiomes from Longitudinal Abundance Data Using Physics-Informed Neural Networks
abstract
Microbial communities are shaped by complex interspecies interactions that sustain a dynamic balance. Understanding the roles and relationships among these species is essential for uncovering the functional mechanisms underlying microbial ecosystems. In this study, we present PhyMicroNet, a self-supervised learning framework designed to infer directed species interactions from longitudinal microbial abundance data. By integrating both data-driven constraints and physical principles, PhyMicroNet embeds the Generalized Lotka-Volterra (GLV) equations into the learning process to ensure ecological plausibility of the estimated interactions. To enhance consistency across multiple biological copies and avoid artificial discontinuities caused by directly concatenating fragmented trajectories, we incorporate a Siamese network structure. This design promotes convergence of inferred parameters under varying initial conditions while preserving the continuity required by ecological models. We evaluated PhyMicroNet on simulated datasets across different levels of species richness and temporal resolution, and further validated its performance using real human gut microbiome data. These results demonstrate that PhyMicroNet consistently outperforms four established baselines not only in overall inference accuracy, but also maintains higher robustness when data are sparse, noisy, or collected over limited temporal sampling.
Bingyao Lin, Tingzhi Deng, Xiaobing Huang, Ying Wang 0005
BIBM4
2025 BIOTIC: a Bayesian framework to integrate single-cell multi-omics for transcription factor activity inference and improve identity characterization of cells
abstract
Understanding cell destiny requires unraveling the intricate mechanism of gene regulation, where transcription factors (TFs) play a pivotal role. However, the actual contribution of TFs, that is TF activity, is not only determined by TF expression, but also accessibility of corresponding chromatin regions. Therefore, we introduce BIOTIC, an advanced Bayesian model with a well-established gene regulation structure that harnesses the power of single-cell multi-omics data to model the gene expression process under the control of regulatory elements, thereby defining the regulatory activity of TFs with variational inference. We demonstrated that the TF activity inferred by BIOTIC can serve as a characterization of cell identity, and outperforms baseline methods for the tasks of cell typing, cell development tracking, and batch effect correction. Additionally, BIOTIC trained on multi-omics data can flexibly be applied to the scenario where merely single-cell transcriptome sequencing is available, to infer TF activity and annotate the cell type by mapping the query cell into the reference TF activity space, as an emerging application of cell atlases. The structure of BIOTIC has been determined to be adaptable for the inclusion of additional biological factors, allowing for flexible and more comprehensive gene regulation analysis. BIOTIC introduces a pioneering biological-mechanism-driven framework to infer TF activity and elucidate cell identity states at gene regulatory level, paving the way for a deeper understanding of the complex interplay between TFs and gene expression in living systems.
Shengquan Chen, Xiaobing Huang, Ying Wang 0005
Briefings Bioinform.7
2025 ViTax: adaptive hierarchical viral taxonomy classification with a taxonomy belief tree on a foundation model
abstract
Viruses exert a profound influence on both human health and the global ecosystem, yet they remain largely unexplored. Precise taxonomic classification of viral sequences is essential for discovering novel viruses, elucidating their functions, and assessing their implications for public health and environmental monitoring. Traditional taxonomy methods based on genome references are limited by the vast number of unexplored viruses, rapid mutation rates, and high genetic diversity. Additionally, highly imbalanced species distribution and significant variances in inter-species genomic distances across taxonomic units pose challenges to classifier training. Conceptualizing genomic sequences as sentences in a natural language, large language models provide novel approaches for extracting intrinsic viral genome characteristics. In this study, we introduce ViTax, a virus taxonomy classification tool powered by HyenaDNA, a large language foundation model for long-range genomic sequences at single nucleotide resolution. ViTax integrates supervised prototypical contrastive learning to address the highly imbalanced distributions across various taxonomic clades and demonstrates superior performance to current leading methods in virus taxonomy, particularly significant for long sequences. Moreover, ViTax designs a belief mapping tree using the Lowest Common Ancestor algorithm to adaptively assign a sequence to the lowest taxonomy clade with confidence. For the open-set problem, where sequences belong to novel and unexplored genera, ViTax can adaptively assign them to a higher level of known taxonomy with outstanding performance. These capabilities make ViTax a robust tool for advancing the accuracy and reliability of viral taxonomy classification. The code is available at https://github.com/Ying-Lab/ViTax.
Yushuang He, Feng Zhou 0021, Jiaxing Bai, Yichun Gao, Xiaobing Huang, Ying Wang 0005
Briefings Bioinform.6
2022 Metric learning for comparing genomic data with triplet network
abstract
Many biological applications are essentially pairwise comparison problems, such as evolutionary relationships on genomic sequences, contigs binning on metagenomic data, cell type identification on gene expression profiles of single-cells, etc. To make pair-wise comparison, it is necessary to adopt suitable dissimilarity metric. However, not all the metrics can be fully adapted to all possible biological applications. It is necessary to employ metric learning based on data adaptive to the application of interest. Therefore, in this study, we proposed MEtric Learning with Triplet network (MELT), which learns a nonlinear mapping from original space to the embedding space in order to keep similar data closer and dissimilar data far apart. MELT is a weakly supervised and data-driven comparison framework that offers more adaptive and accurate dissimilarity learned in the absence of the label information when the supervised methods are not applicable. We applied MELT in three typical applications of genomic data comparison, including hierarchical genomic sequences, longitudinal microbiome samples and longitudinal single-cell gene expression profiles, which have no distinctive grouping information. In the experiments, MELT demonstrated its empirical utility in comparison to many widely used dissimilarity metrics. And MELT is expected to accommodate a more extensive set of applications in large-scale genomic comparisons. MELT is available at https://github.com/Ying-Lab/MELT.
Yang Young Lu, Renhao Lin, Zizi Yang, Ying Wang 0005
Briefings Bioinform.7
2021 Corrigendum to: Assessment of metagenomic assemblers based on hybrid reads of real and simulated metagenomic sequences
abstract
The first version of this article did not include Shanfeng Zhu's funding from Project 111. This has now been corrected. The authors regret the error.
Ying Wang 0005, Jed A. Fuhrman, Fengzhu Sun, Shanfeng Zhu
Briefings Bioinform.2
2021 CRAFT: Compact genome Representation toward large-scale Alignment-Free daTabase
abstract
MOTIVATION: Rapid developments in sequencing technologies have boosted generating high volumes of sequence data. To archive and analyze those data, one primary step is sequence comparison. Alignment-free sequence comparison based on k-mer frequencies offers a computationally efficient solution, yet in practice, the k-mer frequency vectors for large k of practical interest lead to excessive memory and storage consumption. RESULTS: We report CRAFT, a general genomic/metagenomic search engine to learn compact representations of sequences and perform fast comparison between DNA sequences. Specifically, given genome or high throughput sequencing data as input, CRAFT maps the data into a much smaller embedding space and locates the best matching genome in the archived massive sequence repositories. With 102-104-fold reduction of storage space, CRAFT performs fast query for gigabytes of data within seconds or minutes, achieving comparable performance as six state-of-the-art alignment-free measures. AVAILABILITY AND IMPLEMENTATION: CRAFT offers a user-friendly graphical user interface with one-click installation on Windows and Linux operating systems, freely available at https://github.com/jiaxingbai/CRAFT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yang Young Lu, Jiaxing Bai, Ying Wang 0005, Fengzhu Sun
Bioinform.4
2021 MAT2: manifold alignment of single-cell transcriptomes with cell triplets
abstract
MOTIVATION: Aligning single-cell transcriptomes is important for the joint analysis of multiple single-cell RNA sequencing datasets, which in turn is vital to establishing a holistic cellular landscape of certain biological processes. Although numbers of approaches have been proposed for this problem, most of which only consider mutual neighbors when aligning the cells without taking into account known cell type annotations. RESULTS: In this work, we present MAT2 that aligns cells in the manifold space with a deep neural network employing contrastive learning strategy. Compared with other manifold-based approaches, MAT2 has two-fold advantages. Firstly, with cell triplets defined based on known cell type annotations, the consensus manifold yielded by the alignment procedure is more robust especially for datasets with limited common cell types. Secondly, the batch-effect-free gene expression reconstructed by MAT2 can better help annotate cell types. Benchmarking results on real scRNA-seq datasets demonstrate that MAT2 outperforms existing popular methods. Moreover, with MAT2, the hematopoietic stem cells are found to differentiate at different paces between human and mouse. AVAILABILITY AND IMPLEMENTATION: MAT2 is publicly available at https://github.com/Zhang-Jinglong/MAT2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ying Wang 0005, Xing-Ming Zhao
Bioinform.3
2020 Assessment of metagenomic assemblers based on hybrid reads of real and simulated metagenomic sequences
abstract
In metagenomic studies of microbial communities, the short reads come from mixtures of genomes. Read assembly is usually an essential first step for the follow-up studies in metagenomic research. Understanding the power and limitations of various read assembly programs in practice is important for researchers to choose which programs to use in their investigations. Many studies evaluating different assembly programs used either simulated metagenomes or real metagenomes with unknown genome compositions. However, the simulated datasets may not reflect the real complexities of metagenomic samples and the estimated assembly accuracy could be misleading due to the unknown genomes in real metagenomes. Therefore, hybrid strategies are required to evaluate the various read assemblers for metagenomic studies. In this paper, we benchmark the metagenomic read assemblers by mixing reads from real metagenomic datasets with reads from known genomes and evaluating the integrity, contiguity and accuracy of the assembly using the reads from the known genomes. We selected four advanced metagenome assemblers, MEGAHIT, MetaSPAdes, IDBA-UD and Faucet, for evaluation. We showed the strengths and weaknesses of these assemblers in terms of integrity, contiguity and accuracy for different variables, including the genetic difference of the real genomes with the genome sequences in the real metagenomic datasets and the sequencing depth of the simulated datasets. Overall, MetaSPAdes performs best in terms of integrity and continuity at the species-level, followed by MEGAHIT. Faucet performs best in terms of accuracy at the cost of worst integrity and continuity, especially at low sequencing depth. MEGAHIT has the highest genome fractions at the strain-level and MetaSPAdes has the overall best performance at the strain-level. MEGAHIT is the most efficient in our experiments. Availability: The source code is available at https://github.com/ziyewang/MetaAssemblyEval.
Ying Wang 0005, Jed A. Fuhrman, Fengzhu Sun, Shanfeng Zhu
Briefings Bioinform.2
2017 Improving contig binning of metagenomic data using (d2S) oligonucleotide frequency dissimilarity
abstract
BACKGROUND: Metagenomics sequencing provides deep insights into microbial communities. To investigate their taxonomic structure, binning assembled contigs into discrete clusters is critical. Many binning algorithms have been developed, but their performance is not always satisfactory, especially for complex microbial communities, calling for further development. RESULTS: According to previous studies, relative sequence compositions are similar across different regions of the same genome, but they differ between distinct genomes. Generally, current tools have used the normalized frequency of k-tuples directly, but this represents an absolute, not relative, sequence composition. Therefore, we attempted to model contigs using relative k-tuple composition, followed by measuring dissimilarity between contigs using [Formula: see text]. The [Formula: see text] was designed to measure the dissimilarity between two long sequences or Next-Generation Sequencing data with the Markov models of the background genomes. This method was effective in revealing group and gradient relationships between genomes, metagenomes and metatranscriptomes. With many binning tools available, we do not try to bin contigs from scratch. Instead, we developed [Formula: see text] to adjust contigs among bins based on the output of existing binning tools for a single metagenomic sample. The tool is taxonomy-free and depends only on k-tuples. To evaluate the performance of [Formula: see text], five widely used binning tools with different strategies of sequence composition or the hybrid of sequence composition and abundance were selected to bin six synthetic and real datasets, after which [Formula: see text] was applied to adjust the binning results. Our experiments showed that [Formula: see text] consistently achieves the best performance with tuple length k = 6 under the independent identically distributed (i.i.d.) background model. Using the metrics of recall, precision and ARI (Adjusted Rand Index), [Formula: see text] improves the binning performance in 28 out of 30 testing experiments (6 datasets with 5 binning tools). The [Formula: see text] is available at https://github.com/kunWangkun/d2SBin . CONCLUSIONS: Experiments showed that [Formula: see text] accurately measures the dissimilarity between contigs of metagenomic reads and that relative sequence composition is more reasonable to bin the contigs. The [Formula: see text] can be applied to any existing contig-binning tools for single metagenomic samples to obtain better binning results.
Ying Wang 0005, Yang Young Lu, Fengzhu Sun
BMC Bioinform.1