VLDB 2026 Research / reviewers in the wild / expert
Xiaoqing Peng
dblp:55/6794
· DBLP profile ↗
26ranked-venue papers
12as first author
16since 2021 · last 2025
0000-0002-4099-5183ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 9 first-author · 15 since 2021Computer networks · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Chinese medical named entity recognition method considering length diversity of entities
Long Lyu, Weifu Chang, Yuexin Zhao, Xiaoqing Peng |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Noninvasive Diagnosis of Cancer Based on the Heterogeneity and Fragmentation Features of Cell-Free Mitochondrial DNA
Xiaoqing Peng, Wenlong Jie, Junjie You, Wentong Feng |
ISBRA (2) | 1 |
| 2024 | AGImpute: imputation of scRNA-seq data based on a hybrid GAN with dropouts identificationabstractMOTIVATION: Dropout events bring challenges in analyzing single-cell RNA sequencing data as they introduce noise and distort the true distributions of gene expression profiles. Recent studies focus on estimating dropout probability and imputing dropout events by leveraging information from similar cells or genes. However, the number of dropout events differs in different cells, due to the complex factors, such as different sequencing protocols, cell types, and batch effects. The dropout event differences are not fully considered in assessing the similarities between cells and genes, which compromises the reliability of downstream analysis. RESULTS: This work proposes a hybrid Generative Adversarial Network with dropouts identification to impute single-cell RNA sequencing data, named AGImpute. First, the numbers of dropout events in different cells in scRNA-seq data are differentially estimated by using a dynamic threshold estimation strategy. Next, the identified dropout events are imputed by a hybrid deep learning model, combining Autoencoder with a Generative Adversarial Network. To validate the efficiency of the AGImpute, it is compared with seven state-of-the-art dropout imputation methods on two simulated datasets and seven real single-cell RNA sequencing datasets. The results show that AGImpute imputes the least number of dropout events than other methods. Moreover, AGImpute enhances the performance of downstream analysis, including clustering performance, identifying cell-specific marker genes, and inferring trajectory in the time-course dataset. AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/xszhu-lab/AGImpute. Xiaoshu Zhu, Shuang Meng, Gaoshi Li, Jianxin Wang 0001, Xiaoqing Peng |
Bioinform. | 5 |
| 2023 | Noninvasive diagnosis of hepatocellular carcinoma by integrating the genetic, epigenetic and fragmentation features of cell-free DNAsabstractCell-free DNAs (cfDNAs) released from the apoptosis and necrosis cells provide abundant information for screening cancers. Recent studies have revealed different types of features in cfDNAs which can be applied to auxiliary diagnosis of cancers. In this work, we proposed a staked ensemble model to effectively integrate three classes of features extracted from the bisulfite sequencing of cfDNA samples, including four types of fragmentation features (fragment size coverage, fragment size distribution, end motif, breakpoint motif), one type of genetic feature (Copy Number variation), and one type of epigenetic feature (methylation fragment level). For each type of features, a base model is constructed which ensembles six machine learning algorithms to predict cancer scores for every sample. Then, the cancer scores from six base models are ensembled to generate the final stacked model. To demonstrate the effectiveness of the staked ensemble model, it was applied to discriminate samples with hepatocellular carcinoma (HCC) and health controls based on the whole genome bisulfite sequencing data of cfDNAs. The results show that it outperforms four popular methods with single or multi types of features, achieving a sensitivity of 93.33%, a specificity of 93.40%, a accuracy of 95.67% and an AUC of 0.9815 on the validation dataset. Conclusively, the staked model can effectively integrate six types of features of cfDNAs and has great potential for screening cancers. Xiaoqing Peng, Junjie You, Wanxin Cui, Mengxi Zou |
BIBM | 1 |
| 2023 | CATAD: exploring topologically associating domains from an insight of core-attachment structureabstractIdentifying topologically associating domains (TADs), which are considered as the basic units of chromosome structure and function, can facilitate the exploration of the 3D-structure of chromosomes. Methods have been proposed to identify TADs by detecting the boundaries of TADs or identifying the closely interacted regions as TADs, while the possible inner structure of TADs is seldom investigated. In this study, we assume that a TAD is composed of a core and its surrounding attachments, and propose a method, named CATAD, to identify TADs based on the core-attachment structure model. In CATAD, the cores of TADs are identified based on the local density and cosine similarity, and the surrounding attachments are determined based on boundary insulation. CATAD was applied to the Hi-C data of two human cell lines and two mouse cell lines, and the results show that the boundaries of TADs identified by CATAD are significantly enriched by structural proteins, histone modifications, transcription start sites and enzymes. Furthermore, CATAD outperforms other methods in many cases, in terms of the average peak, boundary tagged ratio and fold change. In addition, CATAD is robust and rarely affected by the different resolutions of Hi-C matrices. Conclusively, identifying TADs based on the core-attachment structure is useful, which may inspire researchers to explore TADs from the angles of possible spatial structures and formation process. Xiaoqing Peng, Mengxi Zou, Xiangyan Kong, Yu Sheng |
Briefings Bioinform. | 1 |
| 2023 | IsoCell: An Approach to Enhance Single Cell Clustering by Integrating Isoform-Level Expression Through Orthogonal ProjectionabstractSingle cell RNA sequencing (scRNA-seq) provides a powerful approach for profiling transcriptomes at single cell resolution. An essential application of scRNA-seq is the discovery of cell types with the aid of clustering analysis. Currently, existing single cell clustering methods are exclusively based on gene-level expression data, without considering alternative splicing information. It has been shown that alternative splicing has an important influence on biological processes such as cell differentiation and cell cycle. We therefore hypothesize that adding information about alternative splicing may help enhance single cell clustering. This motivates us to develop a way to integrate isoform-level expression and gene-level expression. We report an approach to enhance single cell clustering by integrating isoform-level expression through orthogonal projection. First, we construct an orthogonal projection matrix based on gene expression data. Second, isoforms are projected to the gene space to remove the redundant information between them. Third, isoform selection is performed based on the residual of the projected expression and the selected isoforms are combined with gene expression data for subsequent clustering. We applied our method to sixteen scRNA-seq datasets. We find that alternative splicing contains differential information among cell types and can be integrated to enhance single cell clustering. Compared with using only gene-level expression data, the integration of isoform-level expression leads to better clustering performances for most of the datasets. The integration of isoform-level expression also has potential in the detection of novel cell subgroups. Our study shows that integrating isoform and gene-level expression is a promising way to improve single cell clustering. The IsoCell R package is freely available at both Github (https://github.com/genemine/IsoCell) and Zenodo (https://zenodo.org/record/4395707). Yingyi Liu, Hong-Dong Li, Yunpei Xu, Yi-Wei Liu, Xiaoqing Peng, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | ADmeth: A Manually Curated Database for the Differential Methylation in Alzheimer's DiseaseabstractAlzheimer's disease (AD) is the most common neurodegenerative disease. More and more evidence show that DNA methylation is closely related to the pathological mechanism of AD. Many AD-associated differentially methylated genes, regions and CpG sites have been identified in recent researches, which may have great potential in clinical research. However, there is no dedicated database to collect AD-related differential methylation up to now. To provide a reference to researchers, we design a database named ADmeth by manually curating relevant articles, which contains a total of 16,709 AD-related differentially methylated items identified from different brain regions and different cell types in the blood, involving 209 genes, 2,229 regions and 14,271 CpG sites. The ADmeth database provides user-friendly pages to search, submit and download data. We hope that the ADmeth database can facilitate researchers to select candidate AD-associated methylation markers in revealing the pathological mechanism of AD and promote the cell-free DNA based non-invasive diagnosis of AD. The ADmeth database is available at http://www.biobdlab.cn/ADmeth. Xiaoqing Peng, Wanxin Cui, Binrong Ding, Qingtong Lyu, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | An integration of deep learning with feature fusion for protein-protein interaction predictionabstractHigh-throughput biological and large-scale experiments for protein–protein interaction (PPI) identification provides valuable information about PPI networks but are time-consuming and limited in determining PPI between different species. We propose the integration of deep learning with feature fusion for PPI prediction, combining handcrafted features and protein sequence embedding for protein sequence representation. Our proposed method is evaluated on the Yeast full, Yeast core, Human, and eight independent datasets. The experimental results show that our method achieves 95.8% accuracy on the Yeast core, 99.2% accuracy on Human, and 100% accuracy on eight independent datasets. We also perform extensive comparisons with other existing outstanding methods and demonstrate the superior ability of the proposed method. Hoai-Nhan Tran, Phuc-Xuan-Quynh Nguyen, Xiaoqing Peng, Jianxin Wang 0001 |
BIBM | 3 |
| 2022 | STgcor: A Distribution-Based Correlation Measurement Method for Spatial Transcriptome Data
Xiaoshu Zhu, Liyuan Pang, Wei Lan 0001, Shuang Meng, Xiaoqing Peng |
ISBRA | 5 |
| 2022 | Metrics for evaluating differentially methylated region sets predicted from BS-seq dataabstractInvestigating differentially methylated regions (DMRs) presented in different tissues or cell types can help to reveal the mechanisms behind the tissue-specific gene expression. The identified tissue-/disease-specific DMRs also can be used as feature markers for spotting the tissues-of-origins of cell-free DNA (cfDNA) in noninvasive diagnosis. In recent years, many methods have been proposed to detect DMRs. However, due to the lack of benchmark DMRs, it is difficult for researchers to choose proper methods and select desirable DMR sets for downstream studies. The application of DMRs, used as feature markers, can be benefited by the longer length of DMRs containing more CpG sites when a threshold is given for the methylation differences of DMRs. According to this, two metrics ($Qn$ and $Ql$), in which the CpG numbers and lengths of DMRs with different methylation differences are weighted differently, are proposed in this paper to evaluate the DMR sets predicted by different methods on BS-seq data. DMR sets predicted by eight methods on both simulated datasets and real BS-seq datasets are evaluated by the proposed metrics, the benchmark-based metrics, and the enrichment analysis of biological data, including genomic features, transcription factors and histones. The rank correlation analysis shows that the $Qn$ and $Ql$ are highly correlated to the benchmark metrics for simulated datasets and the biological data enrichment analysis for real BS-seq data. Therefore, with no need for additional biological data, the proposed metrics can help researchers selecting a more suitable DMR set on a certain BS-seq dataset. Xiaoqing Peng, Hongze Luo, Xiangyan Kong, Jianxin Wang 0001 |
Briefings Bioinform. | 1 |
| 2022 | Drug repositioning based on multi-view learning with matrix completionabstractDetermining drug indications is a critical part of the drug development process. However, traditional drug discovery is expensive and time-consuming. Drug repositioning aims to find potential indications for existing drugs, which is considered as an important alternative to the traditional drug discovery. In this article, we propose a multi-view learning with matrix completion (MLMC) method to predict the potential associations between drugs and diseases. Specifically, MLMC first learns the comprehensive similarity matrices from five drug similarity matrices and two disease similarity matrices based on the multi-view learning (ML) with Laplacian graph regularization, and updates the drug-disease association matrix simultaneously. Then, we introduce matrix completion (MC) to add some positive entries in original association matrix based on low-rank structure, and re-execute the multi-view learning algorithm for association prediction. At last, the prediction results of the above two operations are integrated as the final output. Evaluated by 10-fold cross-validation and de novo tests, MLMC achieves higher prediction accuracy than the current state-of-the-art methods. Moreover, case studies confirm the ability of our method in novel drug-disease association discovery. The codes of MLMC are available at https://github.com/BioinformaticsCSU/MLMC. Contact: [email protected]. Yixin Yan, Mengyun Yang, Guihua Duan, Xiaoqing Peng, Jianxin Wang 0001 |
Briefings Bioinform. | 5 |
| 2022 | AlzCode: a platform for multiview analysis of genes related to Alzheimer's diseaseabstractMOTIVATION: Alzheimer's disease (AD) is a complex brain disorder with risk genes incompletely identified. The candidate genes are dominantly obtained by computational approaches. In order to obtain biological insights of candidate genes or screen genes for experimental testing, it is essential to assess their relevance to AD. A platform that integrates different types of omics data and approaches would facilitate the analysis of candidate genes and is in great need. RESULTS: We report AlzCode, a platform for multiview analysis of genes related to AD. First, this platform integrates a rich collection of functional genomic data, including expression data of AD samples (gene expression, single-cell RNA-seq data and protein expression), AD-specific biological networks (co-expression networks and functional gene networks), neuropathological and clinical traits (CERAD score, Braak staging score, Clinical Dementia Rating, cognitive function and clinical severity) and general data such as protein-protein interaction, regulatory networks, sequence similarity and miRNA-target interactions. These data provide basis for analyzing genes from different views. Second, the platform integrates multiple approaches designed for the various types of data. We implement functions to analyze both individual genes and gene sets. We also compare AlzCode with two existing platforms for AD analysis, which are Agora and AD Atlas. We pinpoint the features of each platform and highlight their differences. This platform would be valuable to the understanding of AD genetics and pathological mechanisms. AVAILABILITY AND IMPLEMENTATION: AlzCode is freely available at: http://www.alzcode.xyz. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cui-Xiang Lin, Hong-Dong Li, Shannon Erhardt, Jun Wang 0153, Xiaoqing Peng, Jianxin Wang 0001 |
Bioinform. | 6 |
| 2021 | SCOTCluster: Deep Clustering with Optimal Transport for Large-scale Single-cell RNA-seq DataabstractSingle-cell RNA sequencing (scRNA-seq) presents cell heterogeneity in a high resolution to explore cell development. The high dimension and the high noise in scRNAseq data bring some computational challenges. By introducing optimal transmission regularization, we proposed a novel deep clustering method, called SCOTCluster, which accurately learned low-dimensional representations. In SCOTCluster, a joint training strategy was designed by integrating AutoEncoder and soft k-means. Notably, to improve simultaneously the accuracy and robustness, the optimal transmission was introduced in the objective function of soft k-means, and entropy regularization and Sinkhorn iterative algorithm were performed to constraint the cluster size. To test the performance, we compared SCOTCluster with five state-of the-art methods on 16 real large-scale scRNA-seq datasets. The experimental results showed that SCOTCluster improved the training stability and clustering performance. Faning Long, Xiaoqing Peng, Jianxin Wang 0001, Xiaoshu Zhu |
BIBM | 3 |
| 2021 | ScDA: A Denoising AutoEncoder Based Dimensionality Reduction for Single-cell RNA-seq Data
Xiaoshu Zhu, Yongchang Lin, Jianxin Wang 0001, Xiaoqing Peng |
ISBRA | 5 |
| 2021 | Identifying the tissues-of-origin of circulating cell-free DNAs is a promising way in noninvasive diagnosticsabstractAdvances in sequencing technologies facilitate personalized disease-risk profiling and clinical diagnosis. In recent years, some great progress has been made in noninvasive diagnoses based on cell-free DNAs (cfDNAs). It exploits the fact that dead cells release DNA fragments into the circulation, and some DNA fragments carry information that indicates their tissues-of-origin (TOOs). Based on the signals used for identifying the TOOs of cfDNAs, the existing methods can be classified into three categories: cfDNA mutation-based methods, methylation pattern-based methods and cfDNA fragmentation pattern-based methods. In cfDNA mutation-based methods, the SNP information or the detected mutations in driven genes of certain diseases are employed to identify the TOOs of cfDNAs. Methylation pattern-based methods are developed to identify the TOOs of cfDNAs based on the tissue-specific methylation patterns. In cfDNA fragmentation pattern-based methods, cfDNA fragmentation patterns, such as nucleosome positioning or preferred end coordinates of cfDNAs, are used to predict the TOOs of cfDNAs. In this paper, the strategies and challenges in each category are reviewed. Furthermore, the representative applications based on the TOOs of cfDNAs, including noninvasive prenatal testing, noninvasive cancer screening, transplantation rejection monitoring and parasitic infection detection, are also reviewed. Moreover, the challenges and future work in identifying the TOOs of cfDNAs are discussed. Our research provides a comprehensive picture of the development and challenges in identifying the TOOs of cfDNAs, which may benefit bioinformatics researchers to develop new methods to improve the identification of the TOOs of cfDNAs. Xiaoqing Peng, Hong-Dong Li, Fang-Xiang Wu, Jianxin Wang 0001 |
Briefings Bioinform. | 1 |
| 2021 | Protein interaction networks: centrality, modularity, dynamics, and applications
Xiangmao Meng, Xiaoqing Peng, Yaohang Li, Min Li 0007 |
Frontiers Comput. Sci. | 3 |
| 2019 | Detecting protein complex based on hierarchical compressing network embeddingabstractDetecting protein complexes from protein-protein interaction (PPI) networks provides biologists an opportunity to efficiently understand the cellular organizations and functions. Existing computational methods just focus on mining high-density regions as the protein complexes by searching the local topological information of a PPI network and ignore the global topological information. To address this limitation, in this study, we present a novel protein complex detection method based on hierarchical compressing network embedding, named DPC-HCNE. The proposed method can preserve both the local topological information and global topological information of a PPI network. To evaluate the performance of our method, DPC-HCNE is compared with other eight typical clustering algorithms to detect protein complexes on two yeast datasets. The experimental results show that DPC-HCNE outperforms those state-of-the-art complex detection methods. Xiangmao Meng, Xiaoqing Peng, Fang-Xiang Wu, Min Li 0007 |
BIBM | 2 |
| 2019 | A Global Similarity Learning for Clustering of Single-Cell RNA-Seq DataabstractSingle-cell RNA-seq (scRNA-seq) data analysis is a powerful tool for biological researches. Similarity plays an important role in clustering scRNA-seq data. Existing similarity measurements are mainly based on local distance information that is calculated between directly connected node pairs, or shared nearest neighbours' information, without considering the global information. Therefore, these similarity measurements may be not very accurate based on the insufficient information. Based on multi-kernel indices in a global feature space and path-based similarity, we proposed a new similarity measurement for single-cell clustering, called multi-kernel and path-based global similarity (MPGS). In MPGS, global information was incorporated by a new feature space from Spearman correlation coefficient, and a global similarity matrix calculated by multi-kernel. A path-based similarity metric was designed to expand the relevant node range. Based on this similaritiy, a modified Louvain community detection method was applied to cluster the scRNA-seq data, named MPGS-Louvain. To validate the performance of MPGS, the clustering performances of several clustering methods combined with different similarity measurements were compared. To demonstrate the performance of MPGS-Louvain, we compared MPGS-Louvain and five scRNA-seq clustering methods on twenty scRNA-seq datasets. The experimental results showed that MPGS outperformed other similarity measurements, and MPGS-Louvain achieved better performance on these datasets. It can be observed that MPGS provided a new insight to improve the accuracy of clustering scRNA-seq data by considering the global information in similarity measurement. MPGS-Louvain automatically detected clusters accurately without prior knowledge. Xiaoshu Zhu, Lilu Guo, Yunpei Xu, Hong-Dong Li, Xingyu Liao, Fang-Xiang Wu, Xiaoqing Peng |
BIBM | 7 |
| 2017 | Protein-protein interactions: detection, reliability assessment and applicationsabstractProtein-protein interactions (PPIs) participate in all important biological processes in living organisms, such as catalyzing metabolic reactions, DNA replication, DNA transcription, responding to stimuli and transporting molecules from one location to another. To reveal the function mechanisms in cells, it is important to identify PPIs that take place in the living organism. A large number of PPIs have been discovered by high-throughput experiments and computational methods. However, false-positive PPIs have been introduced too. Therefore, to obtain reliable PPIs, many computational methods have been proposed. Generally, these methods can be classified into two categories. One category includes the methods that are designed to determine new reliable PPIs. The other one is designed to assess the reliability of existing PPIs and filter out the unreliable ones. In this article, we review the two kinds of methods for detecting reliable PPIs, and then focus on evaluating the performance of some of these typical methods. Later on, we also enumerate several PPI network-based applications with taking a reliability assessment of the PPI data into consideration. Finally, we will discuss the challenges for obtaining reliable PPIs and future directions of the construction of reliable PPI networks. Our research will provide readers some guidance for choosing appropriate methods and features for obtaining reliable PPIs. Xiaoqing Peng, Jianxin Wang 0001, Wei Peng 0004, Fang-Xiang Wu, Yi Pan 0001 |
Briefings Bioinform. | 1 |
| 2016 | Drug repositioning based on comprehensive similarity measures and Bi-Random walk algorithmabstractMOTIVATION: Drug repositioning, which aims to identify new indications for existing drugs, offers a promising alternative to reduce the total time and cost of traditional drug development. Many computational strategies for drug repositioning have been proposed, which are based on similarities among drugs and diseases. Current studies typically use either only drug-related properties (e.g. chemical structures) or only disease-related properties (e.g. phenotypes) to calculate drug or disease similarity, respectively, while not taking into account the influence of known drug-disease association information on the similarity measures. RESULTS: In this article, based on the assumption that similar drugs are normally associated with similar diseases and vice versa, we propose a novel computational method named MBiRW, which utilizes some comprehensive similarity measures and Bi-Random walk (BiRW) algorithm to identify potential novel indications for a given drug. By integrating drug or disease features information with known drug-disease associations, the comprehensive similarity measures are firstly developed to calculate similarity for drugs and diseases. Then drug similarity network and disease similarity network are constructed, and they are incorporated into a heterogeneous network with known drug-disease interactions. Based on the drug-disease heterogeneous network, BiRW algorithm is adopted to predict novel potential drug-disease associations. Computational experiment results from various datasets demonstrate that the proposed approach has reliable prediction performance and outperforms several recent computational drug repositioning approaches. Moreover, case studies of five selected drugs further confirm the superior performance of our method to discover potential indications for drugs practically. AVAILABILITY AND IMPLEMENTATION: http://github.com//bioinfomaticsCSU/MBiRW CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huimin Luo, Jianxin Wang 0001, Min Li 0007, Xiaoqing Peng, Fang-Xiang Wu, Yi Pan 0001 |
Bioinform. | 5 |
| 2015 | An efficient method to identify essential proteins for different species by integrating protein subcellular localization informationabstractEssential proteins are indispensable to maintain life activities in living organisms, and play important roles in the studies of pathology, synthetic biology, and drug design. Many computational methods are employed to identify essential proteins from Protein-protein Interaction Networks (PINs). In this paper, considering the different importance of protein-protein interactions which take place in different subcellular compartments, a Compartment Importance Centrality (CIC) method is proposed to detect essential proteins by integrating protein subcellular localization information. The experiments were carried on four species (Saccharomyces cerevisiae, Homo sapiens, Mus musculus and Drosophila melanogaster), and the performance of CIC was compared with other centrality methods, including the centrality methods solely based on topology and the ones combining both topology and other biological knowledge. The results show that CIC method has better performance to predict essential protein on four species. Furthermore, different from methods which overfits with the features of essential proteins of one species and may perform poor for other species, CIC has a wide applicable scope to identify essential proteins for different species. Xiaoqing Peng, Jianxin Wang 0001, Jiancheng Zhong, Yi Pan 0001 |
BIBM | 1 |
| 2015 | Re-alignment of the unmapped reads with base quality scoreabstractMOTIVATION: Based on the next generation genome sequencing technologies, a variety of biological applications are developed, while alignment is the first step once the sequencing reads are obtained. In recent years, many software tools have been developed to efficiently and accurately align short reads to the reference genome. However, there are still many reads that can't be mapped to the reference genome, due to the exceeding of allowable mismatches. Moreover, besides the unmapped reads, the reads with low mapping qualities are also excluded from the downstream analysis, such as variance calling. If we can take advantages of the confident segments of these reads, not only can the alignment rates be improved, but also more information will be provided for the downstream analysis. RESULTS: This paper proposes a method, called RAUR (Re-align the Unmapped Reads), to re-align the reads that can not be mapped by alignment tools. Firstly, it takes advantages of the base quality scores (reported by the sequencer) to figure out the most confident and informative segments of the unmapped reads by controlling the number of possible mismatches in the alignment. Then, combined with an alignment tool, RAUR re-align these segments of the reads. We run RAUR on both simulated data and real data with different read lengths. The results show that many reads which fail to be aligned by the most popular alignment tools (BWA and Bowtie2) can be correctly re-aligned by RAUR, with a similar Precision. Even compared with the BWA-MEM and the local mode of Bowtie2, which perform local alignment for long reads to improve the alignment rate, RAUR also shows advantages on the Alignment rate and Precision in some cases. Therefore, the trimming strategy used in RAUR is useful to improve the Alignment rate of alignment tools for the next-generation genome sequencing. AVAILABILITY: All source code are available at http://netlab.csu.edu.cn/bioinformatics/RAUR.html. Xiaoqing Peng, Jianxin Wang 0001, Zhen Zhang 0024, Qianghua Xiao, Min Li 0007, Yi Pan 0001 |
BMC Bioinform. | 1 |
| 2015 | Sparse K-best detector for generalised space shift keying in large-scale multiple-input-multiple-output systemsabstractIn this study, the authors propose a low complexity detector for the generalised space shift keying (GSSK) in large‐scale multiple‐input–multiple‐output systems. To be concrete, they propose a sparse K ‐best (SK) detector based on the breadth‐first category of sphere detector (referred to K ‐best sphere decoding). The author's detector is inspired by the fact that the GSSK signal is naturally a sparse zero‐one vector since only a few antennas are activated at the transmitter. Different with the conventional K ‐best detector searching all the transmit antennas, their proposed SK detector investigates only a few promising candidates which are activated antennas at the transmitter. Overall, their proposed SK detector exploits not only the sparsity of the GSSK signal but also the constraint on its non‐zero values. Therefore the restricted isometry property‐based performance analysis shows that is effective in detecting the GSSK signal. Moreover, the empirical results show that their detector performs much better than the sparse algorithms‐based normalised compressive sensing (NCS) detectors while exhibits only slightly higher complexity than the latter (the low‐complexity orthogonal matching pursuit‐based NCS detector). Xiaoqing Peng, Weimin Wu 0003, Jun Sun 0020, Yingzhuang Liu |
IET Commun. | 1 |
| 2015 | Sparsity-aware, channel order-blind pilot placement with channel estimation in orthogonal frequency division multiplexing systemsabstractEquispaced pilot arrangement is the most popular scheme for pilot‐aided transmission in orthogonal frequency division multiplexing systems. In this study, the authors argue that a non‐equispaced pilot pattern may outperform its equispaced counterpart, if they fully take into account the sparsity of the channel impulse response which is inherent in wireless channels. More specifically, a sparsity‐aware pilot arrangement scheme based on the coherence criterion is investigated in this study. To address the resulting non‐deterministic polynomial (NP)‐hard combinatorial optimisation problem, they propose an efficient local search algorithm. For channel estimation, they convert it to a sparse recovery problem. To enhance the applicability of the authors scheme, that is, when there is no prior knowledge about the channel order, they propose to employ the Bayesian information criterion to estimate the channel order first and then recover the sparse channel vector via existing low‐complexity methods, for example, orthogonal matching pursuit. By combining the above pilot arrangement scheme with channel estimation, their scheme exhibits substantially better performance in comparison with the conventional equispaced schemes with linear (or spline) interpolation, in terms of total number of pilot symbols and bit error rate. Xiaoqing Peng, Weimin Wu 0003, Jun Sun 0020, Yingzhuang Liu, F. Y. Li |
IET Commun. | 1 |
| 2011 | Active Protein Interaction Network and Its Application on Protein Complex DetectionabstractIn recent years, more and more attentions are focused on modelling and analyzing dynamic network. Some researchers attempted to extract dynamic network by combining the dynamic information from gene expression data or subcellular localization data with protein network. However, the dynamics of proteins' presence does not guarantee the dynamics of interactions, since the presence of a protein does not indicate the protein's activity. The activity of a protein is closely connected with its function. Thus only the dynamics of proteins activity ensure the dynamics of interaction. The gene expression of a cellular process or cycle carries more information than only the dynamics of proteins' presence. We assume that a protein is active when its expression values are near its maximum expression value, since the expression quantity will decrease after it has performed its function that leads a feedback for controlling the expression quantity. In this paper, we proposed a method to identify active time points for each protein in a cellular process or cycle by using a 3-sigma principle to compute an active threshold for each gene according to the characteristics of its expression curve. Combined the activity information and protein interaction network, we can construct an active protein interaction network (APPI). To demonstrate the efficiency of APPI network model, we applied it on complex detection. Compared with single threshold time series networks, APPI network achieves a better performance on protein complex prediction. Jianxin Wang 0001, Xiaoqing Peng, Min Li 0007, Yi Pan 0001 |
BIBM | 2 |
| 2006 | An Operational Semantics of an Event-Driven System-Level SimulatorabstractAs a system-level modelling language, SystemC possesses some new and interesting features such as delayed notifications, notification cancelling, notification overriding and delta-cycle. It is challenging to formalise SystemC. In this paper, we first select a kernel subset of SystemC and study its operational semantics. Based on the operational semantics we define a bisimulation relation, from which program equivalence is explored. Finally, we present a set of algebraic laws for the subset language, which can be proved based on the operational semantics model via bisimulation Xiaoqing Peng, Huibiao Zhu, Jifeng He 0001, Naiyong Jin |
SEW | 1 |