VLDB 2026 Research / reviewers in the wild / expert
Xiuzhen Huang
dblp:52/6518
· DBLP profile ↗
23ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-0498-3131ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 4 since 2021Theory of computation · 10 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking large language models for identifying transcription factor regulatory interactionsabstractMOTIVATION: Transcription factors (TFs) and their target genes form regulatory networks that control gene expression and influence diverse biological processes and disease outcomes. Although multiple computational methods and curated databases have been developed to identify TF-target interactions, they often require specialized expertise. Large language models (LLMs) chatbots offer a more accessible alternative for querying TF-target interactions. In this study, we benchmarked four prominent LLMs, Anthropic's Claude 3.5 Sonnet, Google's Gemini 1.0 Pro, OpenAI's GPT-4o, and Meta's Llama3 8b, using 8432 literature-curated human TF-target interactions. We examined four regulatory categories: bidirectional, ambiguous, self-regulated, and unidirectional interactions. RESULTS: Under single-turn queries, Claude 3.5 Sonnet and GPT-4o outperformed the others, with balanced accuracies reaching 50.0 ± 7.6% (GPT-4o, self-regulated) and 48.2 ± 1.0% (Claude 3.5 Sonnet, unidirectional). Zero-temperature settings generally enhanced reproducibility, and multi-turn prompting improved performance for most models, increasing Claude 3.5 Sonnet's accuracy on self-regulated pairs by 32.6%. Excluding TF-target pairs with all unknown regulation types also generally improved accuracy, with unidirectional regulation reaching near 70% balanced accuracy in some cases. We also benchmarked Anthropic's Claude 3.5 Sonnet, Google's Gemini 2.0 Flash, OpenAI's GPT-4o, and Meta's Llama3 using 5148 experimentally derived TF-target interactions. Claude 3.5 Sonnet consistently outperformed the other models across conditions. Our findings highlight that prompt engineering and strategic use of model parameters consistently influence LLM chatbots' performance on TF-target identifications. This study establishes a benchmarking framework and demonstrates the potential of pre-trained general-purpose LLMs to support regulatory biology research, especially for researchers without extensive computational expertise. AVAILABILITY AND IMPLEMENTATION: The literature-based TF-target interactions ground truth were obtained from TRRUST v2 human dataset (www.grnpedia.org/trrust). The experimental derived TF-target interactions ground truth were obtained from TFLink Home Sapiens small-scale interaction table (https://tflink.net/). Processed TF-target interactions data and the analytical pipeline has been compiled as an interactive Python notebook file and is available at https://github.com/pengpclab/LLM-TF-interactions. Lake Noel, Yi-Wen Hsiao, Yimeng He, Andrew Hung, Xiaojiang Cui, Edward Ray, Jason H. Moore, Pei-Chen Peng, Xiuzhen Huang |
Bioinform. | 9 |
| 2022 | CrossCas: A Novel Cross-Platform Approach for Predicting Cascades in Online Social Networks with Hidden Markov ModelabstractInformation sharing through online social networks (OSNs) facilitates quick discovery and consumption of information online. Many OSNs such as Facebook, Twitter provide resharing or reposting features, which allows users to share others' content with their own friends or followers. As content is shared from person to person, cascades of information-sharing can occur. There are many existing works focusing on analyzing and characterizing the cascades in OSNs. However, previous works focus on the analysis and characterization of cascades without providing a solution to accurately predict cascades. Although some methods for cascade prediction have been proposed recently, their methods work in social networks such as Facebook (or Twitter), and do not work well simultaneously in multiple OSNs such as Software Social Network (SSN) GitHub, Twitter and Reddit because GitHub, Twitter and Reddit have different social activity patterns. In this paper, we first perform a thorough analysis of cascades in multiple OSNs: GitHub, Twitter and Reddit, and identify the cascades of information-sharing. We then propose CrossCas, a novel cross-platform approach for predicting cascades in multiple OSNs with Hidden Markov Model (HMM). The experimental results show that our proposed method achieves high performance. Xiaonan Zhang 0001, Richard A. Aló, Xiuzhen Huang, Long Cheng 0003, Feng Deng |
GLOBECOM | 4 |
| 2022 | Spatial Pyramid Pooling With 3D Convolution Improves Lung Cancer DetectionabstractLung cancer is the leading cause of cancer deaths. Low-dose computed tomography (CT)screening has been shown to significantly reduce lung cancer mortality but suffers from a high false positive rate that leads to unnecessary diagnostic procedures. The development of deep learning techniques has the potential to help improve lung cancer screening technology. Here we present the algorithm, DeepScreener, which can predict a patient's cancer status from a volumetric lung CT scan. DeepScreener is based on our model of Spatial Pyramid Pooling, which ranked 16th of 1972 teams (top 1 percent)in the Data Science Bowl 2017 competition (DSB2017), evaluated with the challenge datasets. Here we test the algorithm with an independent set of 1449 low-dose CT scans of the National Lung Screening Trial (NLST)cohort, and we find that DeepScreener has consistent performance of high accuracy. Furthermore, by combining Spatial Pyramid Pooling and 3D Convolution, it achieves an AUC of 0.892, surpassing the previous state-of-the-art algorithms using only 3D convolution. The advancement of deep learning algorithms can potentially help improve lung cancer detection with low-dose CT scans. Jason L. Causey, Xianghao Chen, Wei Dong 0003, Karl Walker, Jake A. Qualls, Jonathan W. Stubblefield, Jason H. Moore, Yuanfang Guan, Xiuzhen Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 10 |
| 2022 | An Ensemble of U-Net Models for Kidney Tumor Segmentation With CT ImagesabstractWe present here the Arkansas AI-Campus solution method for the 2019 Kidney Tumor Segmentation Challenge (KiTS19). Our Arkansas AI-Campus team participated the KiTS19 Challenge for four months, from March to July of 2019. This paper provides a summary of our methods, training, testing and validation results for this grand challenge in biomedical imaging analysis. Our deep learning model is an ensemble of U-Net models developed after testing many model variations. Our model has consistent performance on the local test dataset and the final competition independent test dataset. The model achieved local test Dice scores of 0.949 for kidney and tumor segmentation, and 0.601 for tumor segmentation, and the final competition test earned Dice scores 0.9470 and 0.6099 respectively. The Arkansas AI-Campus team solution with a composite DICE score of 0.7784 has achieved a final ranking of top fifty worldwide, and top five among the United States teams in the KiTS19 Competition. Jason L. Causey, Jonathan W. Stubblefield, Jake A. Qualls, Jennifer Fowler, Lingrui Cai, Karl Walker, Yuanfang Guan, Xiuzhen Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2022 | EditorialabstractThis special section gives the opportunity to know recent advances in the application of intelligent optimization algorithms in genomics and precision medicine. Precision medicine is designed to optimize the pathway for diagnosis, therapeutic intervention, and prognosis by using multidimensional biological datasets that capture individual variability in genes, function, and environment. Recent advances in -omics technologies provide substantial novel opportunities to study and/or identify biomarkers of chronic diseases by interpreting multi-omics data, including transcriptomics, epigenomics, genomics, and proteomics, that, together may improve understanding of precision medicine. Precision medicine is drugs or treatments designed for small groups, rather than large populations, based on characteristics, such as medical history, genetic makeup, and data recorded by wearable devices. The use of genomic data can support precision medicine to enable clinicians to predict the most appropriate course of action quickly, efficiently, and accurately for a patient. This offers clinicians the opportunity to tailor early interventions to each patient more carefully. Xiuzhen Huang, Yu Zhang 0150, Xuan Guo 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | runibic: a Bioconductor package for parallel row-based biclustering of gene expression dataabstractMotivation: Biclustering is an unsupervised technique of simultaneous clustering of rows and columns of input matrix. With multiple biclustering algorithms proposed, UniBic remains one of the most accurate methods developed so far. Results: In this paper we introduce a Bioconductor package called runibic with parallel implementation of UniBic. For the convenience the algorithm was reimplemented, parallelized and wrapped within an R package called runibic. The package includes: (i) a couple of times faster parallel version of the original sequential algorithm, (ii) much more efficient memory management, (iii) modularity which allows to build new methods on top of the provided one and (iv) integration with the modern Bioconductor packages such as SummarizedExperiment, ExpressionSet and biclust. Availability and implementation: The package is implemented in R and is available from Bioconductor (starting from version 3.6) at the following URL http://bioconductor.org/packages/runibic with installation instructions and tutorial. Supplementary information: Supplementary data are available at Bioinformatics online. Patryk Orzechowski, Artur Panszczyk, Xiuzhen Huang, Jason H. Moore |
Bioinform. | 3 |
| 2018 | EBIC: an evolutionary-based parallel biclustering algorithm for pattern discoveryabstractMotivation: Biclustering algorithms are commonly used for gene expression data analysis. However, accurate identification of meaningful structures is very challenging and state-of-the-art methods are incapable of discovering with high accuracy different patterns of high biological relevance. Results: In this paper, a novel biclustering algorithm based on evolutionary computation, a sub-field of artificial intelligence, is introduced. The method called EBIC aims to detect order-preserving patterns in complex data. EBIC is capable of discovering multiple complex patterns with unprecedented accuracy in real gene expression datasets. It is also one of the very few biclustering methods designed for parallel environments with multiple graphics processing units. We demonstrate that EBIC greatly outperforms state-of-the-art biclustering methods, in terms of recovery and relevance, on both synthetic and genetic datasets. EBIC also yields results over 12 times faster than the most accurate reference algorithms. Availability and implementation: EBIC source code is available on GitHub at https://github.com/EpistasisLab/ebic. Supplementary information: Supplementary data are available at Bioinformatics online. Patryk Orzechowski, Moshe Sipper, Xiuzhen Huang, Jason H. Moore |
Bioinform. | 3 |
| 2016 | BinPacker: Packing-Based De Novo Transcriptome Assembly from RNA-seq DataabstractHigh-throughput RNA-seq technology has provided an unprecedented opportunity to reveal the very complex structures of transcriptomes. However, it is an important and highly challenging task to assemble vast amounts of short RNA-seq reads into transcriptomes with alternative splicing isoforms. In this study, we present a novel de novo assembler, BinPacker, by modeling the transcriptome assembly problem as tracking a set of trajectories of items with their sizes representing coverage of their corresponding isoforms by solving a series of bin-packing problems. This approach, which subtly integrates coverage information into the procedure, has two exclusive features: 1) only splicing junctions are involved in the assembling procedure; 2) massive pell-mell reads are assembled seemingly by moving a comb along junction edges on a splicing graph. Being tested on both real and simulated RNA-seq datasets, it outperforms almost all the existing de novo assemblers on all the tested datasets, and even outperforms those ab initio assemblers on the real dog dataset. In addition, it runs substantially faster and requires less memory space than most of the assemblers. BinPacker is published under GNU GENERAL PUBLIC LICENSE and the source is available from: http://sourceforge.net/projects/transcriptomeassembly/files/BinPacker_1.0.tar.gz/download. Quick installation version is available from: http://sourceforge.net/projects/transcriptomeassembly/files/BinPacker_binary.tar.gz/download. Ting Yu 0010, Bingqiang Liu, Rick McMullen, Pengyin Chen, Xiuzhen Huang |
PLoS Comput. Biol. | 8 |
| 2014 | Stochastic k-Tree Grammar and Its Application in Biomolecular Structure Modeling
Liang Ding 0007, Xingran Xue, Xiuzhen Huang, Russell L. Malmberg, Liming Cai |
LATA | 4 |
| 2010 | Fixed-Parameter Approximation: Conceptual Framework and Approximability Results
Liming Cai, Xiuzhen Huang |
Algorithmica | 2 |
| 2008 | A practical comparison of two K-Means clustering algorithmsabstractBACKGROUND: Data clustering is a powerful technique for identifying data with similar characteristics, such as genes with similar expression patterns. However, not all implementations of clustering algorithms yield the same performance or the same clusters. RESULTS: In this paper, we study two implementations of a general method for data clustering: k-means clustering. Our experimentation compares the running times and distance efficiency of Lloyd's K-means Clustering and the Progressive Greedy K-means Clustering. CONCLUSION: Based on our implementation, not just in processing time, but also in terms of mean squared-difference (MSD), Lloyd's K-means Clustering algorithm is more efficient. This analysis was performed using both a gene expression level sample and on randomly-generated datasets in three-dimensional space. However, other circumstances may dictate a different choice in some situations. Gregory A. Wilkin, Xiuzhen Huang |
BMC Bioinform. | 2 |
| 2008 | Parameterized Complexity and Biopolymer Sequence ComparisonabstractThe paper surveys parameterized algorithms and complexities for computational tasks on biopolymer sequences, including the problems of longest common subsequence, shortest common supersequence, pairwise sequence alignment, multiple sequencing alignment, structure–sequence alignment and structure–structure alignment. Algorithm techniques, built on the structural-unit level as well as on the residue level, are discussed. Liming Cai, Xiuzhen Huang, Frances A. Rosamond, Yinglei Song |
Comput. J. | 2 |
| 2007 | Polynomial time approximation schemes and parameterized complexity
Jianer Chen, Xiuzhen Huang, Iyad Kanj, Ge Xia |
Discret. Appl. Math. | 2 |
| 2006 | Lower Bounds and Parameterized Approach for Longest Common Subsequence
Xiuzhen Huang |
COCOON | 1 |
| 2006 | Maximum common subgraph: some upper bound and lower bound resultsabstractBACKGROUND: Structure matching plays an important part in understanding the functional role of biological structures. Bioinformatics assists in this effort by reformulating this process into a problem of finding a maximum common subgraph between graphical representations of these structures. Among the many different variants of the maximum common subgraph problem, the maximum common induced subgraph of two graphs is of special interest. RESULTS: Based on current research in the area of parameterized computation, we derive a new lower bound for the exact algorithms of the maximum common induced subgraph of two graphs which is the best currently known. Then we investigate the upper bound and design techniques for approaching this problem, specifically, reducing it to one of finding a maximum clique in the product graph of the two given graphs. Considering the upper bound result, the derived lower bound result is asymptotically tight. CONCLUSION: Parameterized computation is a viable approach with great potential for investigating many applications within bioinformatics, such as the maximum common subgraph problem studied in this paper. With an improved hardness result and the proposed approaches in this paper, future research can be focused on further exploration of efficient approaches for different variants of this problem within the constraints imposed by real applications. Xiuzhen Huang, Jing Lai, Steven F. Jennings |
BMC Bioinform. | 1 |
| 2006 | Strong computational lower bounds via parameterized complexity
Jianer Chen, Xiuzhen Huang, Iyad Kanj, Ge Xia |
J. Comput. Syst. Sci. | 2 |
| 2006 | Efficient Parameterized Algorithms for Biopolymer Structure-Sequence AlignmentabstractComputational alignment of a biopolymer sequence (e.g., an RNA or a protein) to a structure is an effective approach to predict and search for the structure of new sequences. To identify the structure of remote homologs, the structure-sequence alignment has to consider not only sequence similarity, but also spatially conserved conformations caused by residue interactions and, consequently, is computationally intractable. It is difficult to cope with the inefficiency without compromising alignment accuracy, especially for structure search in genomes or large databases. This paper introduces a novel method and a parameterized algorithm for structure-sequence alignment. Both the structure and the sequence are represented as graphs, where, in general, the graph for a biopolymer structure has a naturally small tree width. The algorithm constructs an optimal alignment by finding in the sequence graph the maximum valued subgraph isomorphic to the structure graph. It has the computational time complexity O[k(t)N(2)] for the structure of N residues and its tree decomposition of width t. Parameter k, small in nature, is determined by a statistical cutoff for the correspondence between the structure and the sequence. This paper demonstrates a successful application of the algorithm to RNA structure search used for noncoding RNA identification. An application to protein threading is also discussed. Yinglei Song, Xiuzhen Huang, Russell L. Malmberg, Ying Xu 0001, Liming Cai |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2005 | W-Hardness Under Linear FPT-Reductions: Structural Properties and Further Applications
Jianer Chen, Xiuzhen Huang, Iyad Kanj, Ge Xia |
COCOON | 2 |
| 2005 | Efficient Parameterized Algorithm for Biopolymer Structure-Sequence Alignment
Yinglei Song, Xiuzhen Huang, Russell L. Malmberg, Ying Xu 0001, Liming Cai |
WABI | 3 |
| 2005 | Tight lower bounds for certain parameterized NP-hard problems
Jianer Chen, Benny Chor, Michael R. Fellows, Xiuzhen Huang, David W. Juedes, Iyad Kanj, Ge Xia |
Inf. Comput. | 4 |
| 2004 | Tight Lower Bounds for Certain Parameterized NP-Hard ProblemsabstractBased on the framework of parameterized complexity theory, we derive tight lower bounds on the computational complexity for a number of well-known NP-hard problems. We start by proving a general result, namely that the parameterized weighted satisfiability problem on depth-t circuits cannot be solved in time n/sup o(k)/poly(m), where n is the circuit input length, m is the circuit size, and k is the parameter, unless the (t - l)-st level W[t $1] of the W-hierarchy collapses to FPT. By refining this technique, we prove that a group of parameterized NP-hard problems, including weighted SAT, dominating set, hitting set, set cover, and feature set, cannot be solved in time n/sup o(k)/poly(m), where n is the size of the universal set from which the k elements are to be selected and m is the instance size, unless the first level W[l] of the W-hierarchy collapses to FPT. We also prove that another group of parameterized problems which includes weighted q-SAT (for any fixed q /spl ges/ 2), clique, and independent set, cannot be solved in time n/sup o(k)/ unless all search problems in the syntactic class SNP, introduced by Papadimitriou and Yannakakis, are solvable in subexponential time. Note that all these parameterized problems have trivial algorithms of running time either n/sup k/ poly(m) or O(n/sup k/). Jianer Chen, Benny Chor, Michael R. Fellows, Xiuzhen Huang, David W. Juedes, Iyad Kanj, Ge Xia |
CCC | 4 |
| 2004 | Polynomial Time Approximation Schemes and Parameterized Complexity
Jianer Chen, Xiuzhen Huang, Iyad Kanj, Ge Xia |
MFCS | 2 |
| 2004 | Linear FPT reductions and computational lower boundsabstractWe develop new techniques for deriving very strong computational lower bounds for a class of well-known NP-hard problems, including weighted satisfiability, dominating set, hitting set, set cover, clique, and independent set. For example, although a trivial enumeration can easily test in time O(nk) if a given graph of n vertices has a clique of size k, we prove that unless an unlikely collapse occurs in parameterized complexity theory, the problem is not solvable in time f(k) no(k) for any function f, even if we restrict the parameter value k to be bounded by an arbitrarily small function of n. Under the same assumption, we prove that even if we restrict the parameter values k to be Θ(μ(n)) for any reasonable function μ, no algorithm of running time no(k) can test if a graph of n vertices has a clique of size k. Similar strong lower bounds are also derived for other problems in the above class. Our techniques can be extended to derive computational lower bounds on approximation algorithms for NP-hard optimization problems. For example, we prove that the NP-hard distinguishing substring selection problem, for which a polynomial time approximation scheme has been recently developed, has no polynomial time approximation schemes of running time f(1/ε)no(1/ε) for any function f unless an unlikely collapse occurs in parameterized complexity theory. Jianer Chen, Xiuzhen Huang, Iyad Kanj, Ge Xia |
STOC | 2 |