EDBT 2026 Demo / reviewers in the wild / expert
Xiaobing Huang
dblp:27/1470
· DBLP profile ↗
10ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PhyMicroNet: Inferring Directed 'Species Interactions in Microbiomes from Longitudinal Abundance Data Using Physics-Informed Neural NetworksabstractMicrobial communities are shaped by complex interspecies interactions that sustain a dynamic balance. Understanding the roles and relationships among these species is essential for uncovering the functional mechanisms underlying microbial ecosystems. In this study, we present PhyMicroNet, a self-supervised learning framework designed to infer directed species interactions from longitudinal microbial abundance data. By integrating both data-driven constraints and physical principles, PhyMicroNet embeds the Generalized Lotka-Volterra (GLV) equations into the learning process to ensure ecological plausibility of the estimated interactions. To enhance consistency across multiple biological copies and avoid artificial discontinuities caused by directly concatenating fragmented trajectories, we incorporate a Siamese network structure. This design promotes convergence of inferred parameters under varying initial conditions while preserving the continuity required by ecological models. We evaluated PhyMicroNet on simulated datasets across different levels of species richness and temporal resolution, and further validated its performance using real human gut microbiome data. These results demonstrate that PhyMicroNet consistently outperforms four established baselines not only in overall inference accuracy, but also maintains higher robustness when data are sparse, noisy, or collected over limited temporal sampling. Bingyao Lin, Tingzhi Deng, Xiaobing Huang, Ying Wang 0005 |
BIBM | 3 |
| 2025 | BGC Prediction on Minimizing Generalization Error Upper Bound by Collaborative Belief from Fused Biological Language RepresentationsabstractBiosynthetic gene clusters (BGCs) encode enzymes responsible for the synthesis of microbial secondary metabolites. These metabolites constitute a major source of clinically important drugs. Accurate BGC identification and classification are essential for natural product discovery. Existing tools often rely on single-level features and show limited sensitivity to novel cluster architectures. We present CoBe-BGC, a deep learning framework designed specifically for precise BGC detection and product classification across diverse genomes. It integrates gene-level and Pfam domain-level biological language representations through a Collaborative Belief Fusion strategy. Guided by the Generalization Error Upper Bound, CoBe-BGC adaptively adjusts the contribution of each modality according to its estimated confidence. Across benchmark datasets, CoBe-BGC achieved AUROC scores of 0.960 on the nine genome dataset and 0.958 on the six-genome dataset demonstrating reliability across different genomic contexts. Applied to 94 fungal genomes from the phylum Chytridiomycota, CoBe-BGC identified a diverse set of candidate BGCs and classifies them into known biosynthetic categories, including NRPs, polyketides and RiPPs. CoBe-BGC's capabilities in BGC detection reveal the potential to improve prediction accuracy and functional resolution in complex or poorly annotated genomes. By integrating diverse genomic features, it also provides a scalable solution for large-scale genome mining and natural product discovery. Zixuan Xiao, Tingzhi Deng, Xiaobing Huang |
BIBM | 3 |
| 2025 | BIOTIC: a Bayesian framework to integrate single-cell multi-omics for transcription factor activity inference and improve identity characterization of cellsabstractUnderstanding cell destiny requires unraveling the intricate mechanism of gene regulation, where transcription factors (TFs) play a pivotal role. However, the actual contribution of TFs, that is TF activity, is not only determined by TF expression, but also accessibility of corresponding chromatin regions. Therefore, we introduce BIOTIC, an advanced Bayesian model with a well-established gene regulation structure that harnesses the power of single-cell multi-omics data to model the gene expression process under the control of regulatory elements, thereby defining the regulatory activity of TFs with variational inference. We demonstrated that the TF activity inferred by BIOTIC can serve as a characterization of cell identity, and outperforms baseline methods for the tasks of cell typing, cell development tracking, and batch effect correction. Additionally, BIOTIC trained on multi-omics data can flexibly be applied to the scenario where merely single-cell transcriptome sequencing is available, to infer TF activity and annotate the cell type by mapping the query cell into the reference TF activity space, as an emerging application of cell atlases. The structure of BIOTIC has been determined to be adaptable for the inclusion of additional biological factors, allowing for flexible and more comprehensive gene regulation analysis. BIOTIC introduces a pioneering biological-mechanism-driven framework to infer TF activity and elucidate cell identity states at gene regulatory level, paving the way for a deeper understanding of the complex interplay between TFs and gene expression in living systems. Shengquan Chen, Xiaobing Huang, Ying Wang 0005 |
Briefings Bioinform. | 5 |
| 2025 | ViTax: adaptive hierarchical viral taxonomy classification with a taxonomy belief tree on a foundation modelabstractViruses exert a profound influence on both human health and the global ecosystem, yet they remain largely unexplored. Precise taxonomic classification of viral sequences is essential for discovering novel viruses, elucidating their functions, and assessing their implications for public health and environmental monitoring. Traditional taxonomy methods based on genome references are limited by the vast number of unexplored viruses, rapid mutation rates, and high genetic diversity. Additionally, highly imbalanced species distribution and significant variances in inter-species genomic distances across taxonomic units pose challenges to classifier training. Conceptualizing genomic sequences as sentences in a natural language, large language models provide novel approaches for extracting intrinsic viral genome characteristics. In this study, we introduce ViTax, a virus taxonomy classification tool powered by HyenaDNA, a large language foundation model for long-range genomic sequences at single nucleotide resolution. ViTax integrates supervised prototypical contrastive learning to address the highly imbalanced distributions across various taxonomic clades and demonstrates superior performance to current leading methods in virus taxonomy, particularly significant for long sequences. Moreover, ViTax designs a belief mapping tree using the Lowest Common Ancestor algorithm to adaptively assign a sequence to the lowest taxonomy clade with confidence. For the open-set problem, where sequences belong to novel and unexplored genera, ViTax can adaptively assign them to a higher level of known taxonomy with outstanding performance. These capabilities make ViTax a robust tool for advancing the accuracy and reliability of viral taxonomy classification. The code is available at https://github.com/Ying-Lab/ViTax. Yushuang He, Feng Zhou 0021, Jiaxing Bai, Yichun Gao, Xiaobing Huang, Ying Wang 0005 |
Briefings Bioinform. | 5 |
| 2025 | CellTypeHGCN: a heterogeneous graph convolutional network for cell typing on single-cell RNA-seq data with reference TRNabstractAbstract Accurate cell annotation in single-cell RNA sequencing (scRNA-seq) data is a fundamental task for characterizing cellular states and identities. Existing approaches largely rely on gene expression similarity, but fail to capture the complex molecular dependencies that drive transcriptional programs. To address this, we propose CellTypeHGCN, a heterogeneous graph convolutional network (HGCN) framework thats integrates gene expression with transcriptional regulatory network (TRN) to explicitly model the regulatory architecture. TRN is naturally represented as heterogeneous graphs consisting of two regulatory element types, transcription factors and target genes, which are linked by directed regulatory relationships. Through HGCN, CellTypeHGCN can capture cellular representations while incorporating regulatory dependencies, thereby encoding the regulatory information that fundamentally determines cellular identity. Evaluations on six benchmark datasets demonstrate that CellTypeHGCN consistently outperforms seven state-of-the-art methods in cell type annotation. By integrating the prior TRN as the additional constraint, CellTypeHGCN can also robustly perform cross-platform cell type annotation and novel cell type identification. Moreover, the framework of CellTypeHGCN can be flexibly extended to single-cell ATAC-seq data, where it achieves superior performance to existing annotation tools. Overall, CellTypeHGCN provides a scalable, transferable, and mechanistically informed framework for accurate cell type annotation across different single-cell data modalities. References 1. Aibar S, et al. ‘SCENIC: single-cell regulatory network inference and clustering.’ Nature Methods 2017; 14(11):1083–1086. 2. Yin Q, et al. ‘scGraph: a graph neural network-based approach to automatically identify cell types.’ Bioinformatics 2022; 38(11):2996–3003. 3. Cha J and Lee I. ‘Single-cell network biology for resolving cellular heterogeneity in human diseases.’ Experimental & molecular medicine 2020; 52(11):1798–1808. Yongyu Long, Xiaobing Huang |
Briefings Bioinform. | 4 |
| 2024 | Research on Minimum Receptive Field of Convolutional Neural Networks for Energy Analysis Attack
Jifang Jin, Ziran Nie, Bingqi Xie, Xiaobing Huang, Xiaoyi Duan |
ICDF2C (2) | 5 |
| 2022 | PIWI-interacting RNAs in human diseases: databases and computational modelsabstractPIWI-interacting RNAs (piRNAs) are short 21-35 nucleotide molecules that comprise the largest class of non-coding RNAs and found in a large diversity of species including yeast, worms, flies, plants and mammals including humans. The most well-understood function of piRNAs is to monitor and protect the genome from transposons particularly in germline cells. Recent data suggest that piRNAs may have additional functions in somatic cells although they are expressed there in far lower abundance. Compared with microRNAs (miRNAs), piRNAs have more limited bioinformatics resources available. This review collates 39 piRNA specific and non-specific databases and bioinformatics resources, describes and compares their utility and attributes and provides an overview of their place in the field. In addition, we review 33 computational models based upon function: piRNA prediction, transposon element and mRNA-related piRNA prediction, cluster prediction, signature detection, target prediction and disease association. Based on the collection of databases and computational models, we identify trends and potential gaps in tool development. We further analyze the breadth and depth of piRNA data available in public sources, their contribution to specific human diseases, particularly in cancer and neurodegenerative conditions, and highlight a few specific piRNAs that appear to be associated with these diseases. This briefing presents the most recent and comprehensive mapping of piRNA bioinformatics resources including databases, models and tools for disease associations to date. Such a mapping should facilitate and stimulate further research on piRNAs. Liang Chen 0021, Rongzhen Li, Ning Liu 0014, Xiaobing Huang, Garry Wong |
Briefings Bioinform. | 5 |
| 2018 | Design and implementation of DeepDSL: A DSL for deep learning
Xiaobing Huang |
Comput. Lang. Syst. Struct. | 2 |
| 2017 | DeepDSL: A Compilation-based Domain-Specific Language for Deep Learning
Xiaobing Huang |
ICLR (Poster) | 2 |
| 2013 | PIR: A Domain Specific Language for Multimedia RetrievalabstractMultimedia retrieval is a problem domain involving salient features extraction, machine learning, indexing, and retrieval. There are a variety of implementations for these tasks, which are difficult to compose and reuse due to the interface and language incompatibility. Because of this low reusability, researchers often have to implement their experiments from scratch and the resulting programs are not optimized for efficiency and cannot be easily adapted for parallelization. In this paper, we present PIR (Pipeline Information Retrieval), a domain specific language (DSL) for multimedia feature manipulation. The goal is to unify the programming tasks for feature-related programming in multimedia retrieval experiments by hiding the programming details under a flexible layer of domain specific interface. This DSL enables us to optimize the feature-related tasks by compiling the DSL programs into pipeline graphs, which can be executed using a variety of strategies to eliminate redundant computation and enable parallelization and change propagation. Xiaobing Huang |
ISM | 1 |