Guoqing Lu

dblp:97/1388 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 7 first-author · 4 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PR-TDMPC: Preference-Based Reinforcement Learning for Humanoid Control
Tianxiao Yao, Maonian Wu, Guoqing Lu, Shaojun Zhu, Bo Zheng 0004
ICIC (2)3
2026 Poster: Hybrid Frequency Crossover for Multi-Scale Field Reconstruction in Wireless Sensor Networks
Guoqing Lu, Yixuan Sun, Yiwen Jiang, Dongxu Xia
SECON1
2026 Dynamic-K adaptation framework for energy transfer coordination in wireless sensor networks
Guoqing Lu, Bernard Butler
Ad Hoc Networks1
2025 Subsurface Natural Fracture Identification Using an Integrated Ensemble Learning Method
abstract
Natural fractures play a crucial role in the storage and seepage of oil shale. However, identifying fractures using conventional logging techniques presents challenges due to complex response characteristics and severe data imbalance. Here, we propose a highly accurate integrated ensemble learning method, called BSI-extreme gradient boosting (XGBoost), for identifying the natural fracture development, which combines several steps including isolation forests (iForests), synthetic minority oversampling techniques (SMOTEs), and XGBoost, and incorporates rock brittleness as a controlling factor in the model construction process. The proposed model effectively addresses several challenges encountered in fracture identification, including complex logging response characteristics, low precision and recall of fractured labels, and excessive sensitivity of ensemble learning to noise. To do so, the relationship between fracture density and brittle mineral content is analyzed through core analysis and X-ray diffraction (XRD). Then, conventional logging and rock brittleness are used as features for training the model. Herein, by screening the outliers of iForest, SMOTE oversampling, and feature selection, optimal hyperparameters of the model are obtained through the grid search method. The results demonstrated that using BSI-XGBoost, the testing set achieved an accuracy of 92.45%. Comparatively, this accuracy is 4.86% higher than the original XGBoost model and 3.73% higher than the B-XGBoost model, which incorporated brittleness curves but did not include oversampling and outlier removal. Collectively, this workflow provided an effective method for intelligent identification of fractures in oil shale with high accuracy based on easily accessible conventional logging curves.
Guoqing Lu, Lianbo Zeng, Xiaoxuan Chen, Mehdi Ostadhassan, Yangkang Chen
IEEE Trans. Geosci. Remote. Sens.1
2024 Fracture Identification Based on Graph Pooling and Graph Construction in Continental Shale
abstract
Identification of natural fractures in continental shale is a problematic task by conventional logging. To address this issue, a method known as Frac-gPCC that combines graph pooling, graph construction, and node classification to train the identification model is proposed in this study. This method integrates existing geological knowledge into the graph structure and model network construction and captures the topology information of a single fracture and fractured zone through the calculation of graphs. The model effectively addresses the issue of incorrectly identifying samples as nonfractures when the number of fracture samples is significantly less than nonfractures. Concurrently, this method utilizes a large number of unlabeled samples effectively to incorporate them in model training, avoiding the problem of limited labeled samples. The identification process is divided into three steps: first, reconstruct the logging curve and use the graph pooling section to filter and fuse the node information within a certain depth interval to enhance the characteristics of the tool response to fractures. Second, it integrates the relationship between fractures and lithologies, as well as the spatial distribution characteristics of the strata, into the structure of the global graph. Finally, the nodes of the constructed global graph are classified through the node classification section, which naturally supports combination generalization. This method is applied in the Fengcheng Formation of the Mahu Sag, Western China. The results showed that the identification accuracy for testing data reaches 91.96%. Collectively, this reflects the superiority of Frac-gPCC in fracture identification, providing a successful workflow for continental shale characterization.
Guoqing Lu, Lianbo Zeng, Mehdi Ostadhassan, Shaoqun Dong
IEEE Trans. Geosci. Remote. Sens.1
2022 Identifying host-specific amino acid signatures for influenza A viruses using an adjusted entropy measure
abstract
BACKGROUND: Influenza A viruses (IAV) exhibit vast genetic mutability and have great zoonotic potential to infect avian and mammalian hosts and are known to be responsible for a number of pandemics. A key computational issue in influenza prevention and control is the identification of molecular signatures with cross-species transmission potential. We propose an adjusted entropy-based host-specific signature identification method that uses a similarity coefficient to incorporate the amino acid substitution information and improve the identification performance. Mutations in the polymerase genes (e.g., PB2) are known to play a major role in avian influenza virus adaptation to mammalian hosts. We thus focus on the analysis of PB2 protein sequences and identify host specific PB2 amino acid signatures. RESULTS: Validation with a set of H5N1 PB2 sequences from 1996 to 2006 results in adjusted entropy having a 40% false negative discovery rate compared to a 60% false negative rate using unadjusted entropy. Simulations across different levels of sequence divergence show a false negative rate of no higher than 10% while unadjusted entropy ranged from 9 to 100%. In addition, under all levels of divergence adjusted entropy never had a false positive rate higher than 9%. Adjusted entropy also identifies important mutations in H1N1pdm PB2 previously identified in the literature that explain changes in divergence between 2008 and 2009 which unadjusted entropy could not identify. CONCLUSIONS: Based on these results, adjusted entropy provides a reliable and widely applicable host signature identification approach useful for IAV monitoring and vaccine development.
Kent M. Eskridge, Shunpu Zhang, Guoqing Lu
BMC Bioinform.4
2022 Analysis and Variants of Broad Learning System
abstract
The broad learning system (BLS) is designed based on the technology of compressed sensing and pseudo-inverse theory, and consists of feature nodes and enhancement nodes, has been proposed recently. Compared with the popular deep learning structures, such as deep neural networks, BLS has the ability of rapid incremental learning and can remodel the system without the usual tedious retraining process. However, given that BLS is still in its infancy, it still needs analysis, improvements, and verification. In this article, we first analyze the principle of fast incremental learning ability of BLS in depth. Second, in order to provide an in-depth analysis of the BLS structure, according to the novel structure design concept of deep neural networks, we present four brand-new BLS variant networks and their incremental realizations. Third, based on our analysis of the effect of feature nodes and enhancement nodes, a new BLS structure with a semantic feature extraction layer has been proposed, which is called SFEBLS. The experimental results show that SFEBLS and its variants can increase the accuracy rate on the NORB dataset 6.18%, Fashion-MNIST dataset by 3.15%, ORL data by 5.00%, street view house number dataset by 12.88%, and CIFAR-10 dataset by 18.42%, respectively, and the four brand-new BLS variant networks also obviously outperform the original BLS.
Liang Zhang 0010, Guoqing Lu, Peiyi Shen, Mohammed Bennamoun, Syed Afaq Ali Shah, Qiguang Miao, Guangming Zhu 0001, Ping Li 0030, Xiaoyuan Lu
IEEE Trans. Syst. Man Cybern. Syst.3
2017 A systems biology approach for modeling microbiomes using split graphs
abstract
With the recent advances in sequencing technology, researchers now have opportunities to study microbiomes associated with various environments. Recent studies have shown that the composition of microbiomes in our bodies and our environments play a significant role in our health. For example, 90% of human DNA is composed of bacterial microbiomes. In this study, we propose a systems biology approach using split graphs to analyze the composition of microbiomes and the impact of such composition on the health and growth of organisms living in associated environments. We focus on a case study related to the composition of microbiomes in fish guts and its impact on various growth parameters for three types of fish. The proposed model explores features in the aquatic ecosystem including correlations among its microorganisms and their abundance levels. The results of the study show that single or groups of bacteria are significantly associated with multiple growth phenotypes in different gut portions of the fish. We also identify bacterial clusters that provide new insight to functional relevance of these bacteria and their contribution to the fish gut microbial ecosystem.
Su Yeon Kim, Ishwor Thapa, Guoqing Lu, Lifeng Zhu, Hesham Ali 0001
BIBM3
2016 Model-based clustering with certainty estimation: implication for clade assignment of influenza viruses
abstract
BACKGROUND: Clustering is a common technique used by molecular biologists to group homologous sequences and study evolution. There remain issues such as how to cluster molecular sequences accurately and in particular how to evaluate the certainty of clustering results. RESULTS: We presented a model-based clustering method to analyze molecular sequences, described a subset bootstrap scheme to evaluate a certainty of the clusters, and showed an intuitive way using 3D visualization to examine clusters. We applied the above approach to analyze influenza viral hemagglutinin (HA) sequences. Nine clusters were estimated for high pathogenic H5N1 avian influenza, which agree with previous findings. The certainty for a given sequence that can be correctly assigned to a cluster was all 1.0 whereas the certainty for a given cluster was also very high (0.92-1.0), with an overall clustering certainty of 0.95. For influenza A H7 viruses, ten HA clusters were estimated and the vast majority of sequences could be assigned to a cluster with a certainty of more than 0.99. The certainties for clusters, however, varied from 0.40 to 0.98; such certainty variation is likely attributed to the heterogeneity of sequence data in different clusters. In both cases, the certainty values estimated using the subset bootstrap method are all higher than those calculated based upon the standard bootstrap method, suggesting our bootstrap scheme is applicable for the estimation of clustering certainty. CONCLUSIONS: We formulated a clustering analysis approach with the estimation of certainties and 3D visualization of sequence data. We analysed 2 sets of influenza A HA sequences and the results indicate our approach was applicable for clustering analysis of influenza viral sequences.
Shunpu Zhang, Kevin Beland, Guoqing Lu
BMC Bioinform.4
2010 Applying neural networks to classify influenza virus antigenic types and hosts
abstract
Influenza viruses continue to evolve rapidly and are responsible for seasonal epidemics and occasional, but catastrophic, pandemics. We recently demonstrated the use of decision tree and support vector machine methods in classifying pandemic swine flu viral strains with high accuracy. Here, we applied the technique of artificial neural networks for the prediction of important influenza virus antigenic types (H1, H3, and H5) and hosts (Human, Avian, and Swine), which fulfills a critical need for a computational system for influenza surveillance. A comprehensive experiment on different k-mers and different binary encoding types showed classification based upon frequencies of k-mer nucleotide strings performed better than transformed binary data of nucleotides. It has been found for the first time that the accuracy of virus classification varies from host to host and from gene segment to gene segment. In particular, compared to avian and swine viruses, human influenza viruses can be classified with high accuracy, which indicates influenza virus strains might have become well adapted to their human host and hence less variation occurs in human viruses. In addition, the accuracy of host classification varies from genome segment to segment, achieving the highest values when using the HA and NA segments for human host classification. This research, along with our previous studies, shows machine learning techniques play an indispensable role in virus classification.
Pavan Kumar Attaluri, Zhengxin Chen, Guoqing Lu
CIBCB3
2008 Highlighting computations in bioscience and bioinformatics: review of the Symposium of Computations in Bioinformatics and Bioscience (SCBB07)
abstract
The Second Symposium on Computations in Bioinformatics and Bioscience (SCBB07) was held in Iowa City, Iowa, USA, on August 13-15, 2007. This annual event attracted dozens of bioinformatics professionals and students, who are interested in solving emerging computational problems in bioscience, from China, Japan, Taiwan and the United States. The Scientific Committee of the symposium selected 18 peer-reviewed papers for publication in this supplemental issue of BMC Bioinformatics. These papers cover a broad spectrum of topics in computational biology and bioinformatics, including DNA, protein and genome sequence analysis, gene expression and microarray analysis, computational proteomics and protein structure classification, systems biology and machine learning.
Guoqing Lu
BMC Bioinform.1
2008 An improved string composition method for sequence comparison
abstract
BACKGROUND: Historically, two categories of computational algorithms (alignment-based and alignment-free) have been applied to sequence comparison-one of the most fundamental issues in bioinformatics. Multiple sequence alignment, although dominantly used by biologists, possesses both fundamental as well as computational limitations. Consequently, alignment-free methods have been explored as important alternatives in estimating sequence similarity. Of the alignment-free methods, the string composition vector (CV) methods, which use the frequencies of nucleotide or amino acid strings to represent sequence information, show promising results in genome sequence comparison of prokaryotes. The existing CV-based methods, however, suffer certain statistical problems, thereby underestimating the amount of evolutionary information in genetic sequences. RESULTS: We show that the existing string composition based methods have two problems, one related to the Markov model assumption and the other associated with the denominator of the frequency normalization equation. We propose an improved complete composition vector method under the assumption of a uniform and independent model to estimate sequence information contributing to selection for sequence comparison. Phylogenetic analyses using both simulated and experimental data sets demonstrate that our new method is more robust compared with existing counterparts and comparable in robustness with alignment-based methods. CONCLUSION: We observed two problems existing in the currently used string composition methods and proposed a new robust method for the estimation of evolutionary information of genetic sequences. In addition, we discussed that it might not be necessary to use relatively long strings to build a complete composition vector (CCV), due to the overlapping nature of vector strings with a variable length. We suggested a practical approach for the choice of an optimal string length to construct the CCV.
Guoqing Lu, Shunpu Zhang
BMC Bioinform.1
2006 GenomeBlast: a web tool for small genome comparison
abstract
BACKGROUND: Comparative genomics has become an essential approach for identifying homologous gene candidates and their functions, and for studying genome evolution. There are many tools available for genome comparisons. Unfortunately, most of them are not applicable for the identification of unique genes and the inference of phylogenetic relationships in a given set of genomes. RESULTS: GenomeBlast is a Web tool developed for comparative analysis of multiple small genomes. A new parameter called "coverage" was introduced and used along with sequence identity to evaluate global similarity between genes. With GenomeBlast, the following results can be obtained: (1) unique genes in each genome; (2) homologous gene candidates among compared genomes; (3) 2D plots of homologous gene candidates along the all pairwise genome comparisons; and (4) a table of gene presence/absence information and a genome phylogeny. We demonstrated the functions in GenomeBlast with an example of multiple herpesviral genome analysis and illustrated how GenomeBlast is useful for small genome comparison. CONCLUSION: We developed a Web tool for comparative analysis of small genomes, which allows the user not only to identify unique genes and homologous gene candidates among multiple genomes, but also to view their graphical distributions on genomes, and to reconstruct genome phylogeny. GenomeBlast runs on a Linux server with 4 CPUs and 4 GB memory. The online version of GenomeBlast is available to public by using a Web browser with the URL http://bioinfo-srv1.awh.unomaha.edu/genomeblast/.
Guoqing Lu, Liying Jiang, Resa M. K. Helikar, Thaine W. Rowley, Etsuko N. Moriyama
BMC Bioinform.1
2006 AffyMiner: mining differentially expressed genes and biological knowledge in GeneChip microarray data
abstract
BACKGROUND: DNA microarrays are a powerful tool for monitoring the expression of tens of thousands of genes simultaneously. With the advance of microarray technology, the challenge issue becomes how to analyze a large amount of microarray data and make biological sense of them. Affymetrix GeneChips are widely used microarrays, where a variety of statistical algorithms have been explored and used for detecting significant genes in the experiment. These methods rely solely on the quantitative data, i.e., signal intensity; however, qualitative data are also important parameters in detecting differentially expressed genes. RESULTS: AffyMiner is a tool developed for detecting differentially expressed genes in Affymetrix GeneChip microarray data and for associating gene annotation and gene ontology information with the genes detected. AffyMiner consists of the functional modules, GeneFinder for detecting significant genes in a treatment versus control experiment and GOTree for mapping genes of interest onto the Gene Ontology (GO) space; and interfaces to run Cluster, a program for clustering analysis, and GenMAPP, a program for pathway analysis. AffyMiner has been used for analyzing the GeneChip data and the results were presented in several publications. CONCLUSION: AffyMiner fills an important gap in finding differentially expressed genes in Affymetrix GeneChip microarray data. AffyMiner effectively deals with multiple replicates in the experiment and takes into account both quantitative and qualitative data in identifying significant genes. AffyMiner reduces the time and effort needed to compare data from multiple arrays and to interpret the possible biological implications associated with significant changes in a gene's expression.
Guoqing Lu, The V. Nguyen, Yuannan Xia, Mike Fromm 0002
BMC Bioinform.1
2004 Vector NTI, a balanced all-in-one sequence analysis suite
abstract
Vector NTI is a well-balanced desktop application integrated for molecular sequence analysis and biological data management. It has a centralised database and five application modules: Vector NTI, AlignX, BioAnnotator, ContigExpress and GenomBench. In this review, the features and functions available in this software are examined. These include database management, primer design, virtual cloning, alignments, sequence assembly, 3D molecular viewer and internet tools. Some problems encountered when using this software are also discussed. It is hoped that this review will introduce this software to more molecular biologists so they can make better-informed decisions when choosing computational tools to facilitate their everyday laboratory work. This tool can save time and enhance analysis but it requires some learning on the user's part and there are some issues that need to be addressed by the developer.
Guoqing Lu, Etsuko N. Moriyama
Briefings Bioinform.1
2004 The Hera database and its use in the characterization of endoplasmic reticulum proteins
abstract
MOTIVATION: Information concerning endoplasmic reticulum (ER) proteins is widely dispersed and cannot be easily and rapidly processed by the biological community. We present a comprehensive database of human ER proteins, called Human ER Aperçu (Hera). The Hera database was constructed by exhaustively searching through public databases and the scientific literature for ER proteins. RESULTS: Hera was used for the analysis of characteristics common to all human ER proteins. Our results show that a high proportion of ER proteins (59%) have at least one transmembrane domain and display physical characteristics consistent with this observation. In addition, one-third of ER proteins contain known ER retrieval or retention signals and 70% of ER proteins contain a signal peptide or anchor. Finally, 85% of ER proteins contain at least one InterPro motif. The most abundant InterPro motifs in ER proteins represent many of the most well-characterized functions of the ER.
Michelle S. Scott, Guoqing Lu, Michael T. Hallett, David Y. Thomas
Bioinform.2