VLDB 2026 Research / reviewers in the wild / expert
Bingqiang Liu
dblp:14/3443
· DBLP profile ↗
26ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 6 since 2021Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Live Demonstration: An Energy-efficient SoC for Correlative Scan Matching Based 2D-LiDAR SLAM
Yulong Tan, Zixuan Shen, Bingqiang Liu, Chao Wang 0096 |
ISCAS | 4 |
| 2025 | Live Demonstration: An Area and Energy Efficient Reconfigurable Cryptographic Accelerator Based SoC Design for Securing IoT DevicesabstractThis demonstration presents an energy and area efficient Reconfigurable Cryptographic Accelerator (RCA) SoC for secure communication in IoT devices. Built on a ZYNQ-7000 development board, the platform supports multiple block ciphers (DES, AES, SM4) and Hash functions (SHA-1, SHA-2, SM3). Users can follow prompts on the OLED screen to select the cryptographic algorithm via buttons and input data through a keyboard or choose large text files from SD card. The ARM Core and accelerator execute the cryptographic operation simultaneously, and energy efficiency is calculated based on power and computing time, showcasing the improved computing speed and energy efficiency of the proposed accelerator. Xvpeng Zhang, Bingqiang Liu, Lingyun Hu, Zixuan Shen, Zaisheng He, Dengke Xu, Bah-Hwee Gwee, Chao Wang 0096 |
ISCAS | 2 |
| 2025 | MetaMDA: explainable prediction of microbe-drug association utilizing random walks on a microbe-metabolite-drug heterogeneous networkabstractMOTIVATION: Human-associated microbes play a critical role in physiological processes and disease development, including cancer. Predicting microbe-drug associations (MDAs) can aid drug discovery and personalized medicine. However, existing methods cannot predict MDAs involving microbes or drugs absent from labeled data, and they fail to model the underlying biological mechanisms between microbes and drugs. To address these limitations, we propose a novel computational framework, named MetaMDA, for predicting MDAs by performing random walks on a microbe-metabolite-drug heterogeneous network. MetaMDA first constructs a heterogeneous graph that integrates microbes, metabolites, and drugs, enabling the modeling of complex biological interactions. A random walk algorithm with tailored transition probabilities is subsequently applied to the graph, effectively capturing features from multiple node types on a unified scale. RESULTS: Experimental results across multiple datasets demonstrate that MetaMDA consistently outperforms state-of-the-art methods, achieving an average improvement of 26%. Notably, we show MetaMDA's unique ability to predict MDAs involving microbes or drugs absent from labeled data, as illustrated by associations related to acarbose. Furthermore, mechanistic analysis of MetaMDA provides biological explanations for the associations between Escherichia coli and escitalopram, highlighting its potential to reveal a deeper mechanistic understanding of microbe-drug associations. AVAILABILITY AND IMPLEMENTATION: The code and datasets are available on Zenodo https://doi.org/10.5281/zenodo.17348446 and GitHub https://github.com/wqlyt17/MetaMDA. Xintian Miao, Bingqiang Liu |
Bioinform. | 5 |
| 2025 | An Energy- and Resource-Efficient Parallel-Pipelined Pedestrian Detector With Multiscale Image Computation Scheduling for Always-On Intelligent Edge DevicesabstractHistogram of Oriented Gradients (HOG) and linear Support Vector Machine (SVM) have been widely used for pedestrian detection in applications like video surveillance, automatic driving, and intelligent robots. However, in Internet of Things (IoT) applications relying on intelligent edge devices, it is a big challenge to design a high frame-rate multi-scale pedestrian detector without sacrificing precision under strictly resource-limited and energy-constrained conditions. This paper proposes a HOG-SVM-based pedestrian detector with a novel multi-scale image scheduling method based parallel-pipelined multi-detector architecture to maintain a high frame rate with small hardware overhead, and an optimized inter-module pipeline design to minimize pipeline cycles and on-chip buffer costs. Besides, a fine-grained block-score Multiply-Accumulate (MAC) segmentation and mapping method is proposed for the SVM-classifier MAC array to reduce resource overhead while maintaining the same throughput. FPGA validation shows that as compared to the state-of-the-art design with 12 multi-scale detectors, our proposed design achieves a frame rate of 288 fps using only 2 parallel detectors, which reduces LUT, FF, BRAM, and Digital Signal Processor (DSP) usage by 68.9%, 76.1%, 63.3%, and 94.4%, respectively, while improving energy efficiency by 48.6%. ASIC implementation further improves the energy efficiency by 97% and at the same time increases the frame rate to 400 fps at 200 MHz. Zixuan Shen, Bingqiang Liu, Yulong Tan, Yuanjin Zheng, Chao Wang 0096, Jiang Tang |
IEEE Internet Things J. | 3 |
| 2025 | An Energy-Efficient, High-Frame-Rate, and Reconfigurable EKF-SLAM Processor With Full Acceleration for Autonomous Mobile RobotsabstractIn many intelligent edge applications involving Autonomous Mobile Robots (AMRs), efficient and real-time localization and mapping is a fundamental issue. Extended Kalman Filter Simultaneous Localization and Mapping (EKFSLAM) algorithm is a classic and successful solution to realize localization and mapping, while it is computationally intensive and poses a challenge for real-time tasks in small and micro robots. To address this issue, this work proposes an energy-efficient, highframe-rate, and reconfigurable EKF-SLAM processor. Firstly, a heterogeneous dual-core architecture is proposed to enable full acceleration of both matrix operations and nonlinear calculations in EKF-SLAM at the hardware architecture level. Secondly, a Reconfigurable Matrix Accelerator (RMA) and Reconfigurable Nonlinear Accelerator (RNA) are proposed to maximize data reuse and support diverse nonlinear functions at the data flow level. Thirdly, a data property-aware strategy is proposed at the data property level, which exploits matrix symmetry, sparsity, and dependency to reduce storage significantly and eliminate redundant computations. FPGA validation results show that the proposed design can achieve a frame rate of 774 fps and an energy efficiency of 0.66 mJ/frame, when performing mapping processes involving 60 landmarks at 100 MHz. Bingqiang Liu, Yequan Zhao, Minjie Bao, Zhendong Fan, Dingcheng Jiang, Zixuan Shen, Yulong Tan, Zaisheng He, Dengke Xu, Ke Wang 0028, Chao Wang 0096, Lining Sun |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | Live Demonstration: A High-frame-rate and Energy-efficient SIFT Feature Extraction Accelerator Based SoC Design for AMR ApplicationsabstractThis demonstration presents a high-frame-rate and energy-efficient Scale-Invariant Feature Transform (SIFT) feature extraction accelerator based System on Chip (SoC) design. The platform implementing SIFT-based object recognition consists of an OV5640 camera, a SIFT hardware accelerator based on the ZYNQ-7000 SoC, and a personal computer (PC). The feature points and recognition results are displayed on the monitor in real-time at 60 frames per second (fps) with QVGA resolution for Autonomous Mobile Robot (AMR) applications. Zhenhui Duan, Bingqiang Liu, Zehua Yin, Zixuan Shen, Xupeng Zhang, Zaisheng He, Chao Wang 0096 |
ISCAS | 2 |
| 2024 | Live Demonstration: A Reconfigurable, Energy-efficient and High-frame-rate EKF-SLAM Accelerator Based SoC Design for Autonomous Mobile Robot ApplicationsabstractThis demonstration shows a Extend Kalman Filter-Simultaneous Localization And Mapping (EKF-SLAM) accelerator based System On Chip (SoC) design for Autonomous Mobile Robots (AMR). The AMR platform consists of a multi-sensor system with a wheel encoder and LiDAR, and a ZYNQ-7000 FPGA based SoC featuring an EKF-SLAM hardware accelerator. This AMR system achieves real-time SLAM with significant energy efficient improvement against the state-of-the-art designs. Dingcheng Jiang, Bingqiang Liu, Ao Hu, Yequan Zhao, Minjie Bao, Zhendong Fan, Zixuan Shen, Ke Wang 0028, Chao Wang 0096 |
ISCAS | 2 |
| 2024 | Enhancer-driven gene regulatory networks inference from single-cell RNA-seq and ATAC-seq dataabstractDeciphering the intricate relationships between transcription factors (TFs), enhancers, and genes through the inference of enhancer-driven gene regulatory networks (eGRNs) is crucial in understanding gene regulatory programs in a complex biological system. This study introduces STREAM, a novel method that leverages a Steiner forest problem model, a hybrid biclustering pipeline, and submodular optimization to infer eGRNs from jointly profiled single-cell transcriptome and chromatin accessibility data. Compared to existing methods, STREAM demonstrates enhanced performance in terms of TF recovery, TF-enhancer linkage prediction, and enhancer-gene relation discovery. Application of STREAM to an Alzheimer's disease dataset and a diffuse small lymphocytic lymphoma dataset reveals its ability to identify TF-enhancer-gene relations associated with pseudotime, as well as key TF-enhancer-gene relations and TF cooperation underlying tumor cells. Yang Li 0089, Anjun Ma, Yizhong Wang, Cankun Wang, Hongjun Fu, Bingqiang Liu, Qin Ma 0003 |
Briefings Bioinform. | 7 |
| 2024 | CEMIG: prediction of the cis-regulatory motif using the de Bruijn graph from ATAC-seqabstractSequence motif discovery algorithms enhance the identification of novel deoxyribonucleic acid sequences with pivotal biological significance, especially transcription factor (TF)-binding motifs. The advent of assay for transposase-accessible chromatin using sequencing (ATAC-seq) has broadened the toolkit for motif characterization. Nonetheless, prevailing computational approaches have focused on delineating TF-binding footprints, with motif discovery receiving less attention. Herein, we present Cis rEgulatory Motif Influence using de Bruijn Graph (CEMIG), an algorithm leveraging de Bruijn and Hamming distance graph paradigms to predict and map motif sites. Assessment on 129 ATAC-seq datasets from the Cistrome Data Browser demonstrates CEMIG's exceptional performance, surpassing three established methodologies on four evaluative metrics. CEMIG accurately identifies both cell-type-specific and common TF motifs within GM12878 and K562 cell lines, demonstrating its comparative genomic capabilities in the identification of evolutionary conservation and cell-type specificity. In-depth transcriptional and functional genomic studies have validated the functional relevance of CEMIG-identified motifs across various cell types. CEMIG is available at https://github.com/OSU-BMBL/CEMIG, developed in C++ to ensure cross-platform compatibility with Linux, macOS and Windows operating systems. Yizhong Wang, Yang Li 0089, Cankun Wang, Chan-Wang Jerry Lio, Qin Ma 0003, Bingqiang Liu |
Briefings Bioinform. | 6 |
| 2023 | An Energy-Efficient, Resource-Efficient and High Frame-Rate End-to-End Pedestrian Detector Using HOG-SVM for Intelligent Edge DevicesabstractThis paper proposes a Histogram of Oriented Gradients-Support Vector Machine (HOG-SVM) based pedestrian detector with an end-to-end fully-pipelined architecture to achieve a high frame rate by improving the throughput, and reduce the power consumption by minimizing the data movement. To further improve the energy efficiency under the high frame rate, a bit-width pruning method is used to remove the gray-scale converter's redundant data bit width, and a block-score normalization is employed to significantly reduce the normalizer's required divisions. The reduced computation amount also saves the hardware overhead while maintaining the same calculation accuracy. Besides, a modeling and analysis method of the SVM-classifier-Multiply-ACcumulate (MAC) array is proposed to further improve the energy efficiency and save the logic resources, by optimizing the array size with a hardware utilization of 98.4% while maintaining the same throughput. The FPGA implementation results of$640\times 480$video show a high frame rate of up to 439 fps @143 MHz and a high energy efficiency of 0.76 nJ/pixel with 46.7% fewer LUTs, 22.4% fewer registers, 88.3% fewer DSPs, compared to the state-of-the-art design. The ASIC implementation in 55 nm also confirms a high energy efficiency of 0.35 nJ/pixels at 613 fps and 200 MHz as well as a hardware overhead of 177 k gates and 108 Kbits SRAM. Jianhui Song, Bingqiang Liu, Zixuan Shen, Fengwei An, Chao Wang 0096, Jiang Tang |
IECON | 3 |
| 2022 | Energy-Efficient Intelligent Pulmonary Auscultation for Post COVID-19 Era Wearable Monitoring Enabled by Two-Stage Hybrid Neural NetworkabstractThis paper proposes an energy-efficient intelligent pulmonary auscultation system for post COVID-19 era wearable monitoring. This system consists of a tightly coupled two-stage hybrid neural network (TC-TSHNN) model and a corresponding multi-task training paradigm to improve prediction accuracy and generalization ability based on the fact that the number of COVID-19 patients is far less than that of normal people. At the first stage, two-category coarse classification is performed to identify normal and abnormal lung sounds. If the lung sound is abnormal, the second stage would be triggered to perform a four-category fine-grained classification. Besides, discrete wavelet transform is utilized for feature extraction, denoising and data reduction. In addition, advanced lightweight convolutional neural networks are used to reduce the model’s computation and improve the model’s performance. The hybrid network model can achieve 92% computation reduction and energy saving compared with a direct four-category classification when the input lung sound is normal, which is the majority of cases. Experiment results with inter-patient classification on the COVID-19 lung sound dataset from Tongji Hospital in Wuhan City and the ICBHI’17 dataset show that the proposed TC-TSHNN model can significantly reduce power consumption while maintaining competitive performance against the state-of-the-art work. Bingqiang Liu, Ziyuan Wen, Hongling Zhu, Jinsheng Lai, Jiajun Wu 0006, Heng Ping, Wenqing Liu, Guoyi Yu, Zuozhu Liu, Hesong Zeng, Chao Wang 0096 |
ISCAS | 1 |
| 2022 | An Energy-Efficient SIFT Based Feature Extraction Accelerator for High Frame-Rate Video ApplicationsabstractVisual feature extraction is a key technology of computer vision for intelligent video processing. Efficient feature extraction is a fundamental problem in computer vision applications. Scale-Invariant Feature Transform (SIFT) is one of the most popular feature extraction algorithms because SIFT features are invariant to image scale and rotation and robust to changes in illumination and noise. However, SIFT is a computationally-intensive and power-hungry algorithm, which needs to be accelerated by efficient hardware design to achieve both high-speed feature extraction and high energy efficiency for many high frame-rate video applications at Artificial-intelligent Internet of Things edges. In this work, an energy-efficient SIFT based feature extraction accelerator is proposed. In the Gaussian pyramid and Differences of Gaussian (DoG) pyramid construction process, three design methods are proposed to reduce power consumption and improve information fidelity: a fast and slow dual clock domain design method with a reconfigurable design strategy is proposed to reduce the computation resources; a partial sum reuse design method is proposed to further reduce the computation resources and the amount of computation; a dynamic padding design method is proposed to solve the problem of information loss at image edges and corners after convolution operation. In the keypoint descriptor generation process, an optimized algorithm using circular region and polar coordinates is proposed to parallelize the main orientation assignment and descriptor generation to achieve high-speed processing, while maintaining a comparable matching accuracy with the state-of-the-art designs. The experiment results show that the proposed SIFT hardware accelerator is able to extract features by up to 162 frames per second ($640\times 480$pixels) under 100 MHz, with the power consumption of 364.26 mW and energy efficiency of 2.25 mJ/frame based on 180 nm technology, which is suitable for many high frame-rate AIoT applications including autonomous driving cars and unmanned aerial vehicles. Bingqiang Liu, Zehua Yin, Xvpeng Zhang, Xiaofeng Hu, Guoyi Yu, Yuanjin Zheng, Chao Wang 0096, Xuecheng Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | The functional determinants in the organization of bacterial genomesabstractBacterial genomes are now recognized as interacting intimately with cellular processes. Uncovering organizational mechanisms of bacterial genomes has been a primary focus of researchers to reveal the potential cellular activities. The advances in both experimental techniques and computational models provide a tremendous opportunity for understanding these mechanisms, and various studies have been proposed to explore the organization rules of bacterial genomes associated with functions recently. This review focuses mainly on the principles that shape the organization of bacterial genomes, both locally and globally. We first illustrate local structures as operons/transcription units for facilitating co-transcription and horizontal transfer of genes. We then clarify the constraints that globally shape bacterial genomes, such as metabolism, transcription and replication. Finally, we highlight challenges and opportunities to advance bacterial genomic studies and provide application perspectives of genome organization, including pathway hole assignment and genome assembly and understanding disease mechanisms. Zhaoqian Liu, Jingtong Feng, Bin Yu 0007, Qin Ma 0003, Bingqiang Liu |
Briefings Bioinform. | 5 |
| 2021 | Network analyses in microbiome based on high-throughput multi-omics dataabstractTogether with various hosts and environments, ubiquitous microbes interact closely with each other forming an intertwined system or community. Of interest, shifts of the relationships between microbes and their hosts or environments are associated with critical diseases and ecological changes. While advances in high-throughput Omics technologies offer a great opportunity for understanding the structures and functions of microbiome, it is still challenging to analyse and interpret the omics data. Specifically, the heterogeneity and diversity of microbial communities, compounded with the large size of the datasets, impose a tremendous challenge to mechanistically elucidate the complex communities. Fortunately, network analyses provide an efficient way to tackle this problem, and several network approaches have been proposed to improve this understanding recently. Here, we systemically illustrate these network theories that have been used in biological and biomedical research. Then, we review existing network modelling methods of microbial studies at multiple layers from metagenomics to metabolomics and further to multi-omics. Lastly, we discuss the limitations of present studies and provide a perspective for further directions in support of the understanding of microbial communities. Zhaoqian Liu, Anjun Ma, Ewy A. Mathé, Marlena Merling, Qin Ma 0003, Bingqiang Liu |
Briefings Bioinform. | 6 |
| 2021 | A novel computational framework for genome-scale alternative transcription units predictionabstractAlternative transcription units (ATUs) are dynamically encoded under different conditions and display overlapping patterns (sharing one or more genes) under a specific condition in bacterial genomes. Genome-scale identification of ATUs is essential for studying the emergence of human diseases caused by bacterial organisms. However, it is unrealistic to identify all ATUs using experimental techniques because of the complexity and dynamic nature of ATUs. Here, we present the first-of-its-kind computational framework, named SeqATU, for genome-scale ATU prediction based on next-generation RNA-Seq data. The framework utilizes a convex quadratic programming model to seek an optimum expression combination of all of the to-be-identified ATUs. The predicted ATUs in Escherichia coli reached a precision of 0.77/0.74 and a recall of 0.75/0.76 in the two RNA-Sequencing datasets compared with the benchmarked ATUs from third-generation RNA-Seq data. In addition, the proportion of 5'- or 3'-end genes of the predicted ATUs, having documented transcription factor binding sites and transcription termination sites, was three times greater than that of no 5'- or 3'-end genes. We further evaluated the predicted ATUs by Gene Ontology and Kyoto Encyclopedia of Genes and Genomes functional enrichment analyses. The results suggested that gene pairs frequently encoded in the same ATUs are more functionally related than those that can belong to two distinct ATUs. Overall, these results demonstrated the high reliability of predicted ATUs. We expect that the new insights derived by SeqATU will not only improve the understanding of the transcription mechanism of bacteria but also guide the reconstruction of a genome-scale transcriptional regulatory network. Zhaoqian Liu, Wen-Chi Chou, Laurence Ettwiller, Qin Ma 0003, Bingqiang Liu |
Briefings Bioinform. | 7 |
| 2021 | Prediction of protein-protein interactions based on elastic net and deep forest
Bin Yu 0007, Cheng Chen 0051, Zhaomin Yu, Anjun Ma, Bingqiang Liu |
Expert Syst. Appl. | 6 |
| 2020 | QUBIC2: a novel and robust biclustering algorithm for analyses and interpretation of large-scale RNA-Seq dataabstractMOTIVATION: The biclustering of large-scale gene expression data holds promising potential for detecting condition-specific functional gene modules (i.e. biclusters). However, existing methods do not adequately address a comprehensive detection of all significant bicluster structures and have limited power when applied to expression data generated by RNA-Sequencing (RNA-Seq), especially single-cell RNA-Seq (scRNA-Seq) data, where massive zero and low expression values are observed. RESULTS: We present a new biclustering algorithm, QUalitative BIClustering algorithm Version 2 (QUBIC2), which is empowered by: (i) a novel left-truncated mixture of Gaussian model for an accurate assessment of multimodality in zero-enriched expression data, (ii) a fast and efficient dropouts-saving expansion strategy for functional gene modules optimization using information divergency and (iii) a rigorous statistical test for the significance of all the identified biclusters in any organism, including those without substantial functional annotations. QUBIC2 demonstrated considerably improved performance in detecting biclusters compared to other five widely used algorithms on various benchmark datasets from E.coli, Human and simulated data. QUBIC2 also showcased robust and superior performance on gene expression data generated by microarray, bulk RNA-Seq and scRNA-Seq. AVAILABILITY AND IMPLEMENTATION: The source code of QUBIC2 is freely available at https://github.com/OSU-BMBL/QUBIC2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Juan Xie, Anjun Ma, Bingqiang Liu, Sha Cao, Cankun Wang, Chi Zhang 0021, Qin Ma 0003 |
Bioinform. | 4 |
| 2019 | Interpretation of differential gene expression results of RNA-seq data: review and integrationabstractDifferential gene expression (DGE) analysis is one of the most common applications of RNA-sequencing (RNA-seq) data. This process allows for the elucidation of differentially expressed genes across two or more conditions and is widely used in many applications of RNA-seq data analysis. Interpretation of the DGE results can be nonintuitive and time consuming due to the variety of formats based on the tool of choice and the numerous pieces of information provided in these results files. Here we reviewed DGE results analysis from a functional point of view for various visualizations. We also provide an R/Bioconductor package, Visualization of Differential Gene Expression Results using R, which generates information-rich visualizations for the interpretation of DGE results from three widely used tools, Cuffdiff, DESeq2 and edgeR. The implemented functions are also tested on five real-world data sets, consisting of one human, one Malus domestica and three Vitis riparia data sets. Adam McDermaid, Brandon Monier, Bingqiang Liu, Qin Ma 0003 |
Briefings Bioinform. | 4 |
| 2019 | MetaQUBIC: a computational pipeline for gene-level functional profiling of metagenome and metatranscriptomeabstractMOTIVATION: Metagenomic and metatranscriptomic analyses can provide an abundance of information related to microbial communities. However, straightforward analysis of this data does not provide optimal results, with a required integration of data types being needed to thoroughly investigate these microbiomes and their environmental interactions. RESULTS: Here, we present MetaQUBIC, an integrated biclustering-based computational pipeline for gene module detection that integrates both metagenomic and metatranscriptomic data. Additionally, we used this pipeline to investigate 735 paired DNA and RNA human gut microbiome samples, resulting in a comprehensive hybrid gene expression matrix of 2.3 million cross-species genes in the 735 human fecal samples and 155 functional enriched gene modules. We believe both the MetaQUBIC pipeline and the generated comprehensive human gut hybrid expression matrix will facilitate further investigations into multiple levels of microbiome studies. AVAILABILITY AND IMPLEMENTATION: The package is freely available at https://github.com/OSU-BMBL/metaqubic. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anjun Ma, Minxuan Sun, Adam McDermaid, Bingqiang Liu, Qin Ma 0003 |
Bioinform. | 4 |
| 2019 | MetaQUBIC: a computational pipeline for gene-level functional profiling of metagenome and metatranscriptomeabstractBioinformatics (2019) doi: 10.1093/bioinformatics/btz414, 35, 4474–4477. An incomplete supplementary data file was published alongside the above article. This has now been replaced with the complete version. Anjun Ma, Minxuan Sun, Adam McDermaid, Bingqiang Liu, Qin Ma 0003 |
Bioinform. | 4 |
| 2019 | Protein-protein interaction sites prediction by ensemble random forests with synthetic minority oversampling techniqueabstractMOTIVATION: The prediction of protein-protein interaction (PPI) sites is a key to mutation design, catalytic reaction and the reconstruction of PPI networks. It is a challenging task considering the significant abundant sequences and the imbalance issue in samples. RESULTS: A new ensemble learning-based method, Ensemble Learning of synthetic minority oversampling technique (SMOTE) for Unbalancing samples and RF algorithm (EL-SMURF), was proposed for PPI sites prediction in this study. The sequence profile feature and the residue evolution rates were combined for feature extraction of neighboring residues using a sliding window, and the SMOTE was applied to oversample interface residues in the feature space for the imbalance problem. The Multi-dimensional Scaling feature selection method was implemented to reduce feature redundancy and subset selection. Finally, the Random Forest classifiers were applied to build the ensemble learning model, and the optimal feature vectors were inserted into EL-SMURF to predict PPI sites. The performance validation of EL-SMURF on two independent validation datasets showed 77.1% and 77.7% accuracy, which were 6.2-15.7% and 6.1-18.9% higher than the other existing tools, respectively. AVAILABILITY AND IMPLEMENTATION: The source codes and data used in this study are publicly available at http://github.com/QUST-AIBBDRC/EL-SMURF/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Yu 0007, Anjun Ma, Cheng Chen 0051, Bingqiang Liu, Qin Ma 0003 |
Bioinform. | 5 |
| 2019 | Computational Prediction of Sigma-54 Promoters in Bacterial Genomes by Integrating Motif Finding and Machine Learning StrategiesabstractSigma factor, as a unit of RNA polymerase holoenzyme, is a critical factor in the process of gene transcriptional regulation. It recognizes the specific DNA sites and brings the core enzyme of RNA polymerase to the upstream regions of target genes. Therefore, the prediction of the promoters for a particular sigma factor is essential for interpreting functional genomic data and observation. This paper develops a new method to predict sigma-54 promoters in bacterial genomes. The new method organically integrates motif finding and machine learning strategies to capture the intrinsic features of sigma-54 promoters. The experiments on E. coli benchmark test set show that our method has good capability to distinguish sigma-54 promoters from surrounding or randomly selected DNA sequences. The applications of the other three bacterial genomes indicate the potential robustness and applicative power of our method on a large number of bacterial genomes. The source code of our method can be freely downloaded at https://github.com/maqin2001/PromotePredictor. Bingqiang Liu, Xiangrong Liu, Jichang Wu, Qin Ma 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | An algorithmic perspective of de novo cis-regulatory motif finding based on ChIP-seq dataabstractTranscription factors are proteins that bind to specific DNA sequences and play important roles in controlling the expression levels of their target genes. Hence, prediction of transcription factor binding sites (TFBSs) provides a solid foundation for inferring gene regulatory mechanisms and building regulatory networks for a genome. Chromatin immunoprecipitation sequencing (ChIP-seq) technology can generate large-scale experimental data for such protein-DNA interactions, providing an unprecedented opportunity to identify TFBSs (a.k.a. cis-regulatory motifs). The bottleneck, however, is the lack of robust mathematical models, as well as efficient computational methods for TFBS prediction to make effective use of massive ChIP-seq data sets in the public domain. The purpose of this study is to review existing motif-finding methods for ChIP-seq data from an algorithmic perspective and provide new computational insight into this field. The state-of-the-art methods were shown through summarizing eight representative motif-finding algorithms along with corresponding challenges, and introducing some important relative functions according to specific biological demands, including discriminative motif finding and cofactor motifs analysis. Finally, potential directions and plans for ChIP-seq-based motif-finding tools were showcased in support of future algorithm development. Bingqiang Liu, Yang Li 0089, Adam McDermaid, Qin Ma 0003 |
Briefings Bioinform. | 1 |
| 2016 | BinPacker: Packing-Based De Novo Transcriptome Assembly from RNA-seq DataabstractHigh-throughput RNA-seq technology has provided an unprecedented opportunity to reveal the very complex structures of transcriptomes. However, it is an important and highly challenging task to assemble vast amounts of short RNA-seq reads into transcriptomes with alternative splicing isoforms. In this study, we present a novel de novo assembler, BinPacker, by modeling the transcriptome assembly problem as tracking a set of trajectories of items with their sizes representing coverage of their corresponding isoforms by solving a series of bin-packing problems. This approach, which subtly integrates coverage information into the procedure, has two exclusive features: 1) only splicing junctions are involved in the assembling procedure; 2) massive pell-mell reads are assembled seemingly by moving a comb along junction edges on a splicing graph. Being tested on both real and simulated RNA-seq datasets, it outperforms almost all the existing de novo assemblers on all the tested datasets, and even outperforms those ab initio assemblers on the real dog dataset. In addition, it runs substantially faster and requires less memory space than most of the assemblers. BinPacker is published under GNU GENERAL PUBLIC LICENSE and the source is available from: http://sourceforge.net/projects/transcriptomeassembly/files/BinPacker_1.0.tar.gz/download. Quick installation version is available from: http://sourceforge.net/projects/transcriptomeassembly/files/BinPacker_binary.tar.gz/download. Ting Yu 0010, Bingqiang Liu, Rick McMullen, Pengyin Chen, Xiuzhen Huang |
PLoS Comput. Biol. | 5 |
| 2015 | Revisiting operons: an analysis of the landscape of transcriptional units in E. coliabstractBACKGROUND: Bacterial operons are considerably more complex than what were thought. At least their components are dynamically rather than statically defined as previously assumed. Here we present a computational study of the landscape of the transcriptional units (TUs) of E. coli K12, revealed by the available genomic and transcriptomic data, providing new understanding about the complexity of TUs as a whole encoded in the genome of E. coli K12. RESULTS AND CONCLUSION: Our main findings include that (i) different TUs may overlap with each other by sharing common genes, giving rise to clusters of overlapped TUs (TUCs) along the genomic sequence; (ii) the intergenic regions in front of the first gene of each TU tend to have more conserved sequence motifs than those of the other genes inside the TU, suggesting that TUs each have their own promoters; (iii) the terminators associated with the 3' ends of TUCs tend to be Rho-independent terminators, substantially more often than terminators of TUs that end inside a TUC; and (iv) the functional relatedness of adjacent gene pairs in individual TUs is higher than those in TUCs, suggesting that individual TUs are more basic functional units than TUCs. Xizeng Mao, Qin Ma 0003, Bingqiang Liu, Ying Xu 0001 |
BMC Bioinform. | 3 |
| 2013 | An integrated toolkit for accurate prediction and analysis of cis-regulatory motifs at a genome scaleabstractMOTIVATION: We present an integrated toolkit, BoBro2.0, for prediction and analysis of cis-regulatory motifs. This toolkit can (i) reliably identify statistically significant cis-regulatory motifs at a genome scale; (ii) accurately scan for all motif instances of a query motif in specified genomic regions using a novel method for P-value estimation; (iii) provide highly reliable comparisons and clustering of identified motifs, which takes into consideration the weak signals from the flanking regions of the motifs; and (iv) analyze co-occurring motifs in the regulatory regions. RESULTS: We have carried out systematic comparisons between motif predictions using BoBro2.0 and the MEME package. The comparison results on Escherichia coli K12 genome and the human genome show that BoBro2.0 can identify the statistically significant motifs at a genome scale more efficiently, identify motif instances more accurately and get more reliable motif clusters than MEME. In addition, BoBro2.0 provides correlational analyses among the identified motifs to facilitate the inference of joint regulation relationships of transcription factors. AVAILABILITY: The source code of the program is freely available for noncommercial uses at http://code.google.com/p/bobro/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qin Ma 0003, Bingqiang Liu, Chuan Zhou 0009, Yanbin Yin, Ying Xu 0001 |
Bioinform. | 2 |