Peng Qiu

dblp:51/3600 · DBLP profile ↗
← Back
43ranked-venue papers
9as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 33 · 8 first-author · 10 since 2021Systems, architecture and hardware · 5 · 1 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deconvolving cell-type-specific gene expression profiles from bulk RNA-seq samples
abstract
Bulk RNA sequencing (bulk RNA-seq) and single-cell RNA sequencing (scRNA-seq) are two important high-throughput sequencing platforms that have wide applications in biomedical research. Bulk RNA-seq reflects the average gene expression of all cells in the sample at a low experimental cost, whereas scRNA-seq enables transcriptomics profiling at a single-cell level, although with higher experimental costs. To integrate the strengths of both sequencing approaches and capitalize on the wealth of existing bulk RNA-seq datasets, we developed a U-Net-based deep learning algorithm, BLUE, to deconvolve bulk RNA-seq samples into cell-type proportions and cell-type-specific gene expression profiles. Built upon a U-Net backbone, BLUE leverages its powerful feature extraction and representation learning capabilities to achieve accurate predictions for cell-type-specific gene expression profiles, which significantly outperform existing deconvolution algorithms. Given the accurate prediction from BLUE, we developed an integrative framework for subtyping cancer patients and identifying cell-type-specific gene signatures that can function as prognostic biomarkers for cancer.
Sichen Zhu, Zhengqi Wang, Kevin D. Bunting, Peng Qiu
PLoS Comput. Biol.4
2025 AmpLyze: A Deep Learning Model for Predicting the Hemolytic Concentration
abstract
In antimicrobial peptide development, red-blood-cell lysis ($\text{HC}_{50}$) is the principal safety barrier, but existing in silico tools stop at a binary toxicity classification. Here we propose a new method, AmpLyze, that closes this gap by predicting the actual$\text{HC}_{50}$value from protein sequence alone and explaining the residues that drive toxicity. The model couples residue-level ProtT5/ESM2 embeddings with sequence-level descriptors in dual local and global branches, aligned by a cross-attention module and trained with log-cosh loss for robustness to assay noise. The optimal AmpLyze model reaches a PCC of 0.756 and an MSE of 0.987, outperforming classical regressors and the state-of-the-art. Ablations confirm that both branches are essential, and cross-attention adds a further$1 \% \text{PCC}$and 3% MSE improvement. Expected-Gradients attributions reveal known toxicity hotspots and suggest safer substitutions. By turning hemolysis assessment into a quantitative, sequence-based, and interpretable prediction, AmpLyze facilitates AMP design and offers a practical tool for early-stage toxicity screening.
Peng Qiu, Hanqi Feng, Mengchun Zhang, Barnabás Póczos
BIBM1
2025 Diffusion Generative Modeling for Spatially Resolved Gene Expression Inference from Histology Images
abstract
Spatial Transcriptomics (ST) allows a high-resolution measurement of RNA sequence abundance by systematically connecting cell morphology depicted in Hematoxylin and eosin (H\&E) stained histology images to spatially resolved gene expressions. ST is a time-consuming, expensive yet powerful experimental technique that provides new opportunities to understand cancer mechanisms at a fine-grained molecular level, which is critical for uncovering new approaches for disease diagnosis and treatments. Here, we present $\textbf{Stem}$ ($\underline{\textbf{S}}$pa$\underline{\textbf{T}}$ially resolved gene $\underline{\textbf{E}}$xpression inference with diffusion $\underline{\textbf{M}}$odel), a novel computational tool that leverages a conditional diffusion generative model to enable in silico gene expression inference from H&E stained images. Through better capturing the inherent stochasticity and heterogeneity in ST data, $\textbf{Stem}$ achieves state-of-the-art performance on spatial gene expression prediction and generates biologically meaningful gene profiles for new H&E stained images at test time. We evaluate the proposed algorithm on datasets with various tissue sources and sequencing platforms, where it demonstrates clear improvement over existing approaches. $\textbf{Stem}$ generates high-fidelity gene expression predictions that share similar gene variation levels as ground truth data, suggesting that our method preserves the underlying biological heterogeneity. Our proposed pipeline opens up the possibility of analyzing existing, easily accessible H&E stained histology images from a genomics point of view without physically performing gene expression profiling and empowers potential biological discovery from H&E stained histology images. Code is available at: https://github.com/SichenZhu/Stem.
Sichen Zhu, Molei Tao, Peng Qiu
ICLR4
2025 Protocol Compliance in Popular RTC Applications
abstract
Real-time communication (RTC) has been prevalent since COVID-19, supporting billions of video calls and voice chat interactions. Protocols such as STUN, TURN, RTP, RTCP, and QUIC play a critical role in transmitting RTC media in various applications. Based on standardized protocol specifications, in this paper, we investigate the extent of protocol compliance by analyzing the network traffic in real-world one-on-one calls. We capture and filter RTC traffic, design a Deep Packet Inspection framework to identify all messages for RTC media transmission, and systematically evaluate each message's compliance against protocol specifications. Our analysis of six popular RTC applications—Zoom, FaceTime, WhatsApp, Facebook Messenger (i.e., Messenger), Discord and Google Meet—reveals that: 1) None of the studied applications strictly follow all RTC protocol specifications, and existing protocol implementations, except for QUIC, have some level of non-compliance; 2) Existing applications either implement proprietary protocols or modify existing message types to achieve the desired protocol functionality.
Peiqing Chen, Peng Qiu, Lambda, Zaoxing Liu
IMC2
2025 Motion Control of a Hybrid Self-Reconfigurable Wheel-Legged Dual-Arm Robot
abstract
Current wheeled bipedal robots face significant mobility challenges when traversing discontinuous terrain such as gaps and step-like obstacles, and suffer from substantial dynamic inefficiencies. This paper presents a hybrid self-reconfigurable wheel-legged dual-arm robot equipped with an active docking mechanism, enabling transitions between wheeled bipedal and multi-wheel-legged configurations. Based on a self-developed robotic platform, this work addresses key control challenges in articulated multi-wheel-legged mode and proposes a novel distributed operation paradigm for wheeled bipedal robots. Each module utilizes its manipulators for stable grasping of elevated objects and collaborative tasks, while the multi-unit system achieves efficient, high-load, and stable locomotion. To manage the control complexities in multimodal operation, we develop a unified modular control architecture integrating Virtual Model Control (VMC) and Linear Quadratic Regulator (LQR). For the articulated multi-wheel-legged mode, a body-posture controller regulates global body configuration, and a turning controller adjusts the wheelbase and roll angle via distributed actuation to manage the passive degrees of freedom (DoF) at the articulation points. Experimental validation using a physical prototype confirms the effectiveness and practicality of the proposed approach.
Hong Du, Peng Qiu, Yi Yang 0009, Wenjie Song 0001
IROS3
2024 Quantifying the clusterness and trajectoriness of single-cell RNA-seq data
abstract
Among existing computational algorithms for single-cell RNA-seq analysis, clustering and trajectory inference are two major types of analysis that are routinely applied. For a given dataset, clustering and trajectory inference can generate vastly different visualizations that lead to very different interpretations of the data. To address this issue, we propose multiple scores to quantify the "clusterness" and "trajectoriness" of single-cell RNA-seq data, in other words, whether the data looks like a collection of distinct clusters or a continuum of progression trajectory. The scores we introduce are based on pairwise distance distribution, persistent homology, vector magnitude, Ripley's K, and degrees of connectivity. Using simulated datasets, we demonstrate that the proposed scores are able to effectively differentiate between cluster-like data and trajectory-like data. Using real single-cell RNA-seq datasets, we demonstrate the scores can serve as indicators of whether clustering analysis or trajectory inference is a more appropriate choice for biological interpretation of the data.
Hong Seo Lim, Peng Qiu
PLoS Comput. Biol.2
2024 Hierarchical marker genes selection in scRNA-seq analysis
abstract
When analyzing scRNA-seq data containing heterogeneous cell populations, an important task is to select informative marker genes to distinguish various cell clusters and annotate the clusters with biologically meaningful cell types. In existing analysis methods and pipelines, marker genes are typically identified using a one-vs-all strategy, examining differential expression between one cell cluster versus the combination of all other cell clusters. However, this strategy applied to cell clusters belonging to closely related cell types often generates overlapping marker genes, which capture the common signature of closely related cell clusters but provide limited information for distinguishing them. To address the limitations of the one-vs-all strategy, we propose a hierarchical marker gene selection strategy that groups similar cell clusters and selects marker genes in a hierarchical manner. This strategy is able to improve the accuracy and interpretability of cell type identification in single-cell RNA-seq data.
Peng Qiu
PLoS Comput. Biol.2
2023 UPCoL: Uncertainty-Informed Prototype Consistency Learning for Semi-supervised Medical Image Segmentation
Wenjing Lu, Jiahao Lei, Peng Qiu, Rui Sheng, Jinhua Zhou, Xinwu Lu, Yang Yang 0030
MICCAI (4)3
2023 Quantitative model for genome-wide cyclic AMP receptor protein binding site identification and characteristic analysis
abstract
Cyclic AMP receptor proteins (CRPs) are important transcription regulators in many species. The prediction of CRP-binding sites was mainly based on position-weighted matrixes (PWMs). Traditional prediction methods only considered known binding motifs, and their ability to discover inflexible binding patterns was limited. Thus, a novel CRP-binding site prediction model called CRPBSFinder was developed in this research, which combined the hidden Markov model, knowledge-based PWMs and structure-based binding affinity matrixes. We trained this model using validated CRP-binding data from Escherichia coli and evaluated it with computational and experimental methods. The result shows that the model not only can provide higher prediction performance than a classic method but also quantitatively indicates the binding affinity of transcription factor binding sites by prediction scores. The prediction result included not only the most knowns regulated genes but also 1089 novel CRP-regulated genes. The major regulatory roles of CRPs were divided into four classes: carbohydrate metabolism, organic acid metabolism, nitrogen compound metabolism and cellular transport. Several novel functions were also discovered, including heterocycle metabolic and response to stimulus. Based on the functional similarity of homologous CRPs, we applied the model to 35 other species. The prediction tool and the prediction results are online and are available at: https://awi.cuhk.edu.cn/∼CRPBSFinder.
Yigang Chen 0001, Yang-Chi-Dung Lin, Yijun Luo, Xiao-Xuan Cai, Peng Qiu, Shi-Dong Cui, Hsi-Yuan Huang, Hsien-Da Huang
Briefings Bioinform.5
2022 FUSSNet: Fusing Two Sources of Uncertainty for Semi-supervised Medical Image Segmentation
Jinyi Xiang, Peng Qiu, Yang Yang 0030
MICCAI (8)2
2021 Identification of Protein Markers Predictive of Drug-Specific Survival Outcome in Cancers
Shuting Lin, Jie Zhou 0032, Yiqiong Xiao, Bridget Neary, Yong Teng, Peng Qiu
ISBRA6
2021 Immune-Microbiota Crosstalk Underlying Inflammatory Bowel Disease
Congmin Xu, Quoc D. Mac, Peng Qiu
ISBRA4
2021 JSOM: Jointly-evolving self-organizing maps for alignment of biological datasets and identification of related clusters
abstract
With the rapid advances of various single-cell technologies, an increasing number of single-cell datasets are being generated, and the computational tools for aligning the datasets which make subsequent integration or meta-analysis possible have become critical. Typically, single-cell datasets from different technologies cannot be directly combined or concatenated, due to the innate difference in the data, such as the number of measured parameters and the distributions. Even datasets generated by the same technology are often affected by the batch effect. A computational approach for aligning different datasets and hence identifying related clusters will be useful for data integration and interpretation in large scale single-cell experiments. Our proposed algorithm called JSOM, a variation of the Self-organizing map, aligns two related datasets that contain similar clusters, by constructing two maps-low-dimensional discretized representation of datasets-that jointly evolve according to both datasets. Here we applied the JSOM algorithm to flow cytometry, mass cytometry, and single-cell RNA sequencing datasets. The resulting JSOM maps not only align the related clusters in the two datasets but also preserve the topology of the datasets so that the maps could be used for further analysis, such as clustering.
Hong Seo Lim, Peng Qiu
PLoS Comput. Biol.2
2020 Feature selection algorithms for predicting preeclampsia: A comparative approach
abstract
Preeclampsia is a disease that complicates a large number of pregnancies. In this study, a comparative approach is taken to understand how dimension reduction methods and time-series summary methods can be useful for predicting preeclampsia based on proteomics data. Here, the dimension reduction methods are the Imperialist Competitive Algorithm and the gene clustering method of the Sample Progression Discovery algorithm, the time-series summary methods included a simple overall average and a 3-point summary corresponding to the three trimesters of pregnancy. These approaches achieved similar prediction accuracy around 90% in two independent datasets. Separate analysis of data from each trimester showed an interesting result that it is easier to predict preeclampsia based on proteomics data of the first two trimesters of pregnancy rather than the last trimester.
Jose F. Carreño, Peng Qiu
BIBM2
2020 Image-based early predictions of functional properties in cell manufacturing
abstract
Effective cell manufacturing is essential to realizing the full potential of cell-based therapies but faces a multitude of challenges. One of the major challenges is the identification of critical quality attributes (CQAs), especially ones that enable early predictions of functional properties of the final products. The main goal of this study is to develop machine learning models for early predictions of the functional properties of mesenchymal stromal/stem cells(MSCs) in cell manufacturing. Deep learning models are trained and tested for image-based prediction of functional property-Collagen II expression after chondrogenic differentiation-of MSCs cells. During the MSC expansion, images of culturing wells were collected daily in the first six days, and the Collagen II level was assayed at the end of differentiation, following expansion. For each day, a deep learning model was trained with images from a specific experimental condition, and each model was tested with images from the same condition and also from other conditions. The trained neural network models showed 70-90 percent accuracy. Most of the models across different days and conditions show high consistency, especially models trained with images past day 2 of cell culture. Such consistency suggests that models are picking up similar features in predicting chondrogenesis capability. Our study highlighted the potential of deep neural network models used for early predictions of the functional properties of MSCs in cell manufacturing.
Hong Seo Lim, Madeline E. Smerchansky, Jingxuan Zhou, Paramita Chatterjee, Angela C. Jimenez, Krishnendu Roy, Peng Qiu
BIBM8
2020 Leveraging TCGA gene expression data to build predictive models for cancer drug response
abstract
BACKGROUND: Machine learning has been utilized to predict cancer drug response from multi-omics data generated from sensitivities of cancer cell lines to different therapeutic compounds. Here, we build machine learning models using gene expression data from patients' primary tumor tissues to predict whether a patient will respond positively or negatively to two chemotherapeutics: 5-Fluorouracil and Gemcitabine. RESULTS: We focused on 5-Fluorouracil and Gemcitabine because based on our exclusion criteria, they provide the largest numbers of patients within TCGA. Normalized gene expression data were clustered and used as the input features for the study. We used matching clinical trial data to ascertain the response of these patients via multiple classification methods. Multiple clustering and classification methods were compared for prediction accuracy of drug response. Clara and random forest were found to be the best clustering and classification methods, respectively. The results show our models predict with up to 86% accuracy; despite the study's limitation of sample size. We also found the genes most informative for predicting drug response were enriched in well-known cancer signaling pathways and highlighted their potential significance in chemotherapy prognosis. CONCLUSIONS: Primary tumor gene expression is a good predictor of cancer drug response. Investment in larger datasets containing both patient gene expression and drug response is needed to support future work of machine learning models. Ultimately, such predictive models may aid oncologists with making critical treatment decisions.
Evan A. Clayton, Toyya A. Pujol, John F. McDonald 0002, Peng Qiu
BMC Bioinform.4
2020 Shape-to-graph mapping method for efficient characterization and classification of complex geometries in biological images
abstract
With the ever-increasing quality and quantity of imaging data in biomedical research comes the demand for computational methodologies that enable efficient and reliable automated extraction of the quantitative information contained within these images. One of the challenges in providing such methodology is the need for tailoring algorithms to the specifics of the data, limiting their areas of application. Here we present a broadly applicable approach to quantification and classification of complex shapes and patterns in biological or other multi-component formations. This approach integrates the mapping of all shape boundaries within an image onto a global information-rich graph and machine learning on the multidimensional measures of the graph. We demonstrated the power of this method by (1) extracting subtle structural differences from visually indistinguishable images in our phenotype rescue experiments using the endothelial tube formations assay, (2) training the algorithm to identify biophysical parameters underlying the formation of different multicellular networks in our simulation model of collective cell behavior, and (3) analyzing the response of U2OS cell cultures to a broad array of small molecule perturbations.
William Pilcher, Anastasia Zhurikhina, Olga Chernaya, Peng Qiu, Denis Tsygankov
PLoS Comput. Biol.6
2020 Classification of Antibacterial Peptides Using Long Short-Term Memory Recurrent Neural Networks
abstract
Antimicrobial peptides are short amino acid sequences that may be antibacterial, antifungal, and antiviral. Most machine learning methodologies applied to identifying antibacterial peptides have developed feature vectors of identical lengths for each peptide in a given dataset although the peptides themselves may differ in number of amino acids. Features are often chosen which represent certain periodic patterns in the peptide sequence without any initial guidance as to whether such patterns are relevant for the classification task at hand. This can result in the construction of a large number of irrelevant features in addition to relevant features. To help alleviate these issues, we choose to extract a feature vector from individual amino acid feature representations through the application of bidirectional Long Short-Term Memory recurrent neural networks. The Long Short-Term Memory network recursively iterates along both directions of the given amino acid sequence and ultimately extracts a finite length feature vector that is then used to classify the peptide. This work demonstrates the application of Long Short-Term Memory recurrent neural networks to classification of antibacterial peptides and compares it to a Random Forest classifier and a k-nearest neighbor classifier.
Michael Youmans, John Christian Givhan Spainhour, Peng Qiu
IEEE ACM Trans. Comput. Biol. Bioinform.3
2019 An Efficient Cell Segmentation Algorithm Based on Unsupervised Clustering and Morphology
abstract
In traditional Chinese medicine processing domain, the ‘firepower’ and ‘proper roasting’ are vital for medicinal efficacy of Chinese herbs. We take phellodendron microscopic images to determine if the roasting is appropriate according to some features of crystal fiber and stone cell, these features include color, geometrical appearance, and so on. The first step of feature analysis is to segment objects accurately. In the paper, we propose an efficient unsupervised algorithm to segment stone cell and crystal fiber in phellodendron microscopic images automatically without any human's intervene. Firstly, the superpixels is adopted, then the method extracts features such as RGB, Sobel, Harris and so on, clustering superpixels patches according to the characteristic of phellodendron, lastly morphological operations are implemented to produce segmented stone cell and crystal fiber. The experimental evaluation shows that the proposed algorithm can segment objects automatically with about 87% precision and 93% recall.
Dawei Qiu, Xuelan Zhang, Peng Qiu, Yibo Feng, Huifen Li
BIBM4
2019 Deconvolving multiplexed protease signatures with substrate reduction and activity clustering
abstract
Proteases are multifunctional, promiscuous enzymes that degrade proteins as well as peptides and drive important processes in health and disease. Current technology has enabled the construction of libraries of peptide substrates that detect protease activity, which provides valuable biological information. An ideal library would be orthogonal, such that each protease only hydrolyzes one unique substrate, however this is impractical due to off-target promiscuity (i.e., one protease targets multiple different substrates). Therefore, when a library of probes is exposed to a cocktail of proteases, each protease activates multiple probes, producing a convoluted signature. Computational methods for parsing these signatures to estimate individual protease activities primarily use an extensive collection of all possible protease-substrate combinations, which require impractical amounts of training data when expanding to search for more candidate substrates. Here we provide a computational method for estimating protease activities efficiently by reducing the number of substrates and clustering proteases with similar cleavage activities into families. We envision that this method will be used to extract meaningful diagnostic information from biological samples.
Qinwei Zhuang, Brandon Alexander Holt, Gabriel A. Kwong, Peng Qiu
PLoS Comput. Biol.4
2018 Multi-objective service composition model based on cost-effective optimization
Ying Huo, Peng Qiu, Jiyou Zhai, Dajuan Fan, Huanfeng Peng
Appl. Intell.2
2017 The relative importance of data points in systems biology and parameter estimation
abstract
Estimating model parameters is a crucial step to understand the behavior of biological systems. To perform parameter estimation, a commonly used formulation is the least square method that minimizes the mean squared error. This method finds the model parameters that minimize the sum of the squared error between experimental data and model predictions. However, such a formulation can misguide parameter estimation and the understanding of the system. This is mainly because least square formulation typically treats all data points equally, while the reality is that not all data points are of equal importance. Another common issue in systems biology is that the amount of experimental data is almost always limited compared to the model complexity, making parameter estimation challenging and ill-conditioned. Ignoring the relative importance of data points may amplify the ill-conditioned nature of the problem. Therefore, we propose to give different weight to each data point when formulating the least square cost function. The weight of each data point is defined by an uncertainty measure for the data point given the others, quantifying each data point's unique information that cannot be inferred from other data points. To test our algorithm, we used a G1/S transition model with two dynamic variables and 12 parameters, developed a sampling algorithm to obtain collections of parameter settings close to the best fit, and demonstrated the benefits of the proposed weighted cost function formulation.
Jenny Jeong, Peng Qiu
BIBM2
2017 Long short-term memory recurrent neural networks for antibacterial peptide identification
abstract
Antimicrobial peptides are short amino acid sequences with antibacterial, antifungal, and antiviral properties. Antibacterial peptides have the possibility to form a new class of antibiotics to aid in combating bacterial antibiotic resistance. Most machine learning methodologies applied to the task of identifying antimicrobial peptides have applied features representing the presence or absence of certain periodic patterns in the amino acid sequence. This requires considering different periodicities for each feature and leads to a large number of features many of which are likely irrelevant to the classification task at hand. Also as the peptides vary in length it is difficult to develop a feature vector of identical finite length representing all the sequences. An easy way to circumvent both of these problems is provided by recurrent neural networks. In this work we choose to extract a feature vector through the application of bidirectional Long Short-Term Memory (LSTM) recurrent neural networks from features representing individual amino acids within each sequence. The LSTM network recursively iterates along both directions of each amino acid sequence and extracts two finite vectors whose concatenation yields the finite length vector representation of the amino acid sequence. As this is done during the training of the network on the classification task, the representation extracted is more likely to be relevant for distinguishing between the classes. This work demonstrates the LSTM approach to classification of antibacterial peptides and compares it to a Random Forest classifier and a k-nearest neighbor classifier.
Michael Youmans, John Christian Givhan Spainhour, Peng Qiu
BIBM3
2017 Power loss and efficiency analysis of non-isolated weinberg converter
abstract
Non-isolated Weinberg Converter is suitable for battery discharge regulator due to the advantages of continuous current, no RHP zeroes and high efficiency. The working principle under overlap and non-overlap modes with stray parameters is analyzed in this paper. In order to estimate the heat distribution, the power loss is calculated and the trend of change in efficiency is achieved. In addition, the experimental results are given to verify the advantages of Non-isolated Weinberg Converter and the effectiveness of power loss analysis by a 600W prototype.
Peng Qiu, Haihong Yu, Kai Tong, Jiazhuo Xuan
IECON2
2017 Comparison of DC fault handling strategies for hybrid HVDC system
abstract
The hybrid high voltage direct current (Hybrid HVDC) system based on line commutated converter (LCC) and modular multilevel converter (MMC) is a feasible solution for long distance bulk power transmission. In order to handle the dc-side short-circuit fault, two kinds of methods can be adopted. The first one is to replace half-bridge sub-modules (HBSMs) with full-bridge SMs or clamp-double SMs which have the capability of dc fault clearance, and the second one is to arrange diodes at the dc port of MMC. In this paper, detailed processes of the dc fault solutions are described and discussed. The typical testing system is built in PSCAD/EMTDC, and the system performances under different strategies are compared. In addition, the non-block dc fault handling strategy based on full-bridge SMs is analyzed and tested. The results show that this kind of solution is not suitable for the bipolar HVDC system and the fault current cannot be blocked thoroughly.
Jihong Li, Peng Qiu, Huangqing Xiao, Gaoren Liu, Zheng Xu 0011
IECON3
2017 Research on flexible medium-voltage DC distribution technology based shore-to-ship power supply system
abstract
The development of shore-to-ship power supply technology can effectively reduce port pollution and promote the construction of green port. This paper briefly introduces the technical scheme of the common used inverter for shore power, and then puts forward the technical scheme of multi-terminal medium-voltage DC (MVDC), according to the transmission characteristics of the shore-to-ship power supply. Considering the different demand of converters connected to the distribution network and the ship grid, the paper discusses the suitable converter schemes. Based on the current operation demand of the distribution network and the new energy development trend of ports and ships in future, advantages and disadvantages of the inverter technology and the multi-terminal MVDC technology are compared, which demonstrates the feasibility and superiority of the multi-terminal MVDC technology in shore-to-ship power supply applications.
Xiaohua Xuan, Peng Qiu, Kai Tong, Jiazhuo Xuan, Daozhuo Jiang
IECON4
2017 GDISC: a web portal for integrative analysis of gene-drug interaction for survival in cancer
abstract
Summary: Survival analysis has been applied to The Cancer Genome Atlas (TCGA) data. Although drug exposure records are available in TCGA, existing survival analyses typically did not consider drug exposure, partly due to naming inconsistencies in the data. We have spent extensive effort to standardize the drug exposure data, which enabled us to perform survival analysis on drug-stratified subpopulations of cancer patients. Using this strategy, we integrated gene copy number data, drug exposure data and patient survival data to infer gene-drug interactions that impact survival. The collection of all analyzed gene-drug interactions in 32 cancer types are organized and presented in a searchable web-portal called gene-drug Interaction for survival in cancer (GDISC). GDISC allows biologists and clinicians to interactively explore the gene-drug interactions identified in the context of TCGA, and discover interactions associated to their favorite cancer, drug and/or gene of interest. In addition, GDISC provides the standardized drug exposure data, which is a valuable resource for developing new methods for drug-specific analysis. Availability and Implementation: GDISC is available at https://gdisc.bme.gatech.edu/. Contact: [email protected].
John Christian Givhan Spainhour, Juho Lim, Peng Qiu
Bioinform.3
2016 Identification of gene-drug interactions that impact patient survival in TCGA
abstract
BACKGROUND: With the advent of large scale biological data collection for various diseases, data analysis pipelines and workflows need to be established to build frameworks for integrative analysis. Here the authors present a pipeline for identifying disease specific gene-drug interactions using CNV (Copy Number Variation) and clinical data from the TCGA (The Cancer Genome Atlas) project. Two cancer types were selected for analysis, LGG (Brain lower grade glioma) and GBM (Glioblastoma multiforme), due to the possible progression from LGG to GBM in some cases. The copy number and clinical data were then used to preform survival analysis on a gene by gene basis on sub-populations of patients exposed to a given drug. RESULTS: Several gene-drug interactions are identified, where the copy number of a gene is associated to survival of a patient exposed to a certain drug. Both Irinotecan/HAS2 (Hyaluronan synthase 2) and Bevacizumab/PGAM1 (Phosphoglycerate mutase 1) are interactions found in this study with independent confirmation. Independent work in colon, breast cancer and leukemia (Györffy, Breast Cancer Res Treat 123:725-731, 2010; Mueller, Mol Cancer Ther 11:3024-3032, 2010; Hitosugi, Cancer Cell 13:585-600, 2012) showed these two interactions can lead to increased survival. CONCLUSION: While the pipeline produced several possible interactions where increased survival is linked to normal or increased copy number of a given gene for patients treated with a given drug, no instance of low copy number or full deletion was linked to increased survival. The development of this pipeline shows a promising utility to identify possible beneficial gene-drug interactions that could improve patient survival and may illustrate some of the problems inherent in this kind of analysis on these data.
John Christian Givhan Spainhour, Peng Qiu
BMC Bioinform.2
2016 Bridging Mechanistic and Phenomenological Models of Complex Biological Systems
abstract
The inherent complexity of biological systems gives rise to complicated mechanistic models with a large number of parameters. On the other hand, the collective behavior of these systems can often be characterized by a relatively small number of phenomenological parameters. We use the Manifold Boundary Approximation Method (MBAM) as a tool for deriving simple phenomenological models from complicated mechanistic models. The resulting models are not black boxes, but remain expressed in terms of the microscopic parameters. In this way, we explicitly connect the macroscopic and microscopic descriptions, characterize the equivalence class of distinct systems exhibiting the same range of collective behavior, and identify the combinations of components that function as tunable control knobs for the behavior. We demonstrate the procedure for adaptation behavior exhibited by the EGFR pathway. From a 48 parameter mechanistic model, the system can be effectively described by a single adaptation parameter τ characterizing the ratio of time scales for the initial response and recovery time of the system which can in turn be expressed as a combination of microscopic reaction rates, Michaelis-Menten constants, and biochemical concentrations. The situation is not unlike modeling in physics in which microscopically complex processes can often be renormalized into simple phenomenological models with only a few effective parameters. The proposed method additionally provides a mechanistic explanation for non-universal features of the behavior.
Mark K. Transtrum, Peng Qiu
PLoS Comput. Biol.2
2015 Flexible power distribution unit - A novel power electronic transformer development and demonstration for distribution system
abstract
Power electronic transformer (PET) is an emerging new type of power converter in recent years. It has not only the basic functions of power transformation and isolation, but also the extra functions of power quality control. A novel power electronics transformer for distribution system named flexible power distribution unit is proposed in this paper, and the energy exchange mechanism between the network and load is revealed. Finally, a 100kW 600Vac/220Vac/110Vdc medium frequency isolated prototype is developed and demonstrated in laboratory. The experimental results show that the proposed structure and control strategy are feasible.
Qing Duan, Jianhua Wang 0001, Binshi Gu, Baojian Ji, Peng Qiu, Jun You
IECON6
2015 Unsupervised Discovery of Subspace Trends
abstract
This paper presents unsupervised algorithms for discovering previously unknown subspace trends in high-dimensional data sets without the benefit of prior information. A subspace trend is a sustained pattern of gradual/progressive changes within an unknown subset of feature dimensions. A fundamental challenge to subspace trend discovery is the presence of irrelevant data dimensions, noise, outliers, and confusion from multiple subspace trends driven by independent factors that are mixed in with each other. These factors can obscure the trends in conventional dimension reduction & projection based data visualizations. To overcome these limitations, we propose a novel graph-theoretic neighborhood similarity measure for detecting concordant progressive changes across data dimensions. Using this measure, we present an unsupervised algorithm for trend-relevant feature selection, subspace trend discovery, quantification of trend strength, and validation. Our method successfully identified verifiable subspace trends in diverse synthetic and real-world biomedical datasets. Visualizations derived from the selected trend-relevant features revealed biologically meaningful hidden subspace trend(s) that were obscured by irrelevant features and noise. Although our examples are drawn from the biological domain, the proposed algorithm is broadly applicable to exploratory analysis of high-dimensional data including visualization, hypothesis generation, knowledge discovery, and prediction in diverse other applications.
Peng Qiu, Badrinath Roysam
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Improving the sensitivity of sample clustering by leveraging gene co-expression networks in variable selection
abstract
BACKGROUND: Many variable selection techniques have been proposed for the clustering of gene expression data. While these methods tend to filter out irrelevant genes and identify informative genes that contribute to a clustering solution, they are based on criteria that do not consider the potential interactive influence among individual genes. Motivated by ensemble clustering, there is a strong interest in leveraging the structure of gene networks for gene selection, so that the relationship information between genes can be effectively utilized, while the selected genes are expected to preserve all the possible clustering structures in the data. RESULTS: We present a new filter method that uses the gene connectivity in the gene co-expression network as the evaluation criteria for variable selection. The gene connectivity measures the importance of the genes in term of their expression similarity with others in the co-expression network. The hard threshold and soft threshold transformations are employed to construct the gene co-expression networks. Both simulation studies and real data analysis have shown that the network based on soft thresholding is more effective in selecting relevant variables and provides better clustering results compared to the hard thresholding transformation and two other canonical filter methods for variable selection. Furthermore, a new module analysis approach is proposed to reveal the higher order organization of the gene space, where the genes of a module share significant topological similarity and are associated with a consensus partition of the sample space. We demonstrate that the identified modules can lead to biologically meaningful sample partitions that might be missed by other methods. CONCLUSIONS: By leveraging the structure of gene co-expression network, first we propose a variable selection method that selects individual genes with top connectivity. Both simulation studies and real data application have demonstrated that our method has better performance in terms of the reliability of the selected genes and sample clustering results. In addition, we propose a module recovery method that can help discover novel sample partitions that might be hidden when performing clustering analyses using all available genes. The source code of our program is available at http://nba.uth.tmc.edu/homepage/liu/netVar/.
Zixing Wang 0003, F. Anthony San Lucas, Peng Qiu
BMC Bioinform.3
2014 Unfold High-Dimensional Clouds for Exhaustive Gating of Flow Cytometry Data
abstract
Flow cytometry is able to measure the expressions of multiple proteins simultaneously at the single-cell level. A flow cytometry experiment on one biological sample provides measurements of several protein markers on or inside a large number of individual cells in that sample. Analysis of such data often aims to identify subpopulations of cells with distinct phenotypes. Currently, the most widely used analytical approach in the flow cytometry community is manual gating on a sequence of nested biaxial plots, which is highly subjective, labor intensive, and not exhaustive. To address those issues, a number of methods have been developed to automate the gating analysis by clustering algorithms. However, completely removing the subjectivity can be quite challenging. This paper describes an alternative approach. Instead of automating the analysis, we develop novel visualizations to facilitate manual gating. The proposed method views single-cell data of one biological sample as a high-dimensional point cloud of cells, derives the skeleton of the cloud, and unfolds the skeleton to generate 2D visualizations. We demonstrate the utility of the proposed visualization using real data, and provide quantitative comparison to visualizations generated from principal component analysis and multidimensional scaling.
Peng Qiu
IEEE ACM Trans. Comput. Biol. Bioinform.1
2014 BM-SNP: A Bayesian Model for SNP Calling Using High Throughput Sequencing Data
abstract
A single-nucleotide polymorphism (SNP) is a sole base change in the DNA sequence and is the most common polymorphism. Detection and annotation of SNPs are among the central topics in biomedical research as SNPs are believed to play important roles on the manifestation of phenotypic events, such as disease susceptibility. To take full advantage of the next-generation sequencing (NGS) technology, we propose a Bayesian approach, BM-SNP, to identify SNPs based on the posterior inference using NGS data. In particular, BM-SNP computes the posterior probability of nucleotide variation at each covered genomic position using the contents and frequency of the mapped short reads. The position with a high posterior probability of nucleotide variation is flagged as a potential SNP. We apply BM-SNP to two cell-line NGS data, and the results show a high ratio of overlap ( >95 percent) with the dbSNP database. Compared with MAQ, BM-SNP identifies more SNPs that are in dbSNP, with higher quality. The SNPs that are called only by BM-SNP but not in dbSNP may serve as new discoveries. The proposed BM-SNP method integrates information from multiple aspects of NGS data, and therefore achieves high detection power. BM-SNP is fast, capable of processing whole genome data at 20-fold average coverage in a short amount of time.
Yanxun Xu, Yuan Yuan 0003, Marcos R. Estecio, Jean-Pierre Issa, Peng Qiu, Shoudan Liang
IEEE ACM Trans. Comput. Biol. Bioinform.6
2012 CytoSPADE: high-performance analysis and visualization of high-dimensional cytometry data
abstract
MOTIVATION: Recent advances in flow cytometry enable simultaneous single-cell measurement of 30+ surface and intracellular proteins. CytoSPADE is a high-performance implementation of an interface for the Spanning-tree Progression Analysis of Density-normalized Events algorithm for tree-based analysis and visualization of this high-dimensional cytometry data. AVAILABILITY: Source code and binaries are freely available at http://cytospade.org and via Bioconductor version 2.10 onwards for Linux, OSX and Windows. CytoSPADE is implemented in R, C++ and Java. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Additional documentation available at http://cytospade.org.
Michael D. Linderman, Zach Bjornson, Erin F. Simonds, Peng Qiu, Robert V. Bruggner, Ketaki Sheode, Teresa H. Meng, Sylvia K. Plevritis, Garry P. Nolan
Bioinform.4
2012 Identification of markers associated with global changes in DNA methylation regulation in cancers
abstract
DNA methylation exhibits different patterns in different cancers. DNA methylation rates at different genomic loci appear to be highly correlated in some samples but not in others. We call such phenomena conditional concordant relationships (CCRs). In this study, we explored DNA methylation patterns in 12 common cancers using data of 2434 patient samples collected by The Cancer Genome Atlas project. We developed an exploratory method to characterize CCRs in the methylation data and identified the 200 gene markers whose on-and-off statuses in DNA methylation are most significantly associated with drastic changes in CCRs throughout the genome. Clustering analysis of the methylation data of the 200 markers showed that they are tightly associated with cancer subtypes. We also generated a library of the significant CCRs that may be of interest to future studies of the regulation network of DNA methylation in cancer.
Peng Qiu
BMC Bioinform.1
2012 Optimal experiment selection for parameter estimation in biological differential equation models
abstract
BACKGROUND: Parameter estimation in biological models is a common yet challenging problem. In this work we explore the problem for gene regulatory networks modeled by differential equations with unknown parameters, such as decay rates, reaction rates, Michaelis-Menten constants, and Hill coefficients. We explore the question to what extent parameters can be efficiently estimated by appropriate experimental selection. RESULTS: A minimization formulation is used to find the parameter values that best fit the experiment data. When the data is insufficient, the minimization problem often has many local minima that fit the data reasonably well. We show that selecting a new experiment based on the local Fisher Information of one local minimum generates additional data that allows one to successfully discriminate among the many local minima. The parameters can be estimated to high accuracy by iteratively performing minimization and experiment selection. We show that the experiment choices are roughly independent of which local minima is used to calculate the local Fisher Information. CONCLUSIONS: We show that by an appropriate choice of experiments, one can, in principle, efficiently and accurately estimate all the parameters of gene regulatory network. In addition, we demonstrate that appropriate experiment selection can also allow one to restrict model predictions without constraining the parameters using many fewer experiments. We suggest that predicting model behaviors and inferring parameters represent two different approaches to model calibration with different requirements on data and experimental cost.
Mark K. Transtrum, Peng Qiu
BMC Bioinform.2
2011 Discovering Biological Progression Underlying Microarray Samples
abstract
In biological systems that undergo processes such as differentiation, a clear concept of progression exists. We present a novel computational approach, called Sample Progression Discovery (SPD), to discover patterns of biological progression underlying microarray gene expression data. SPD assumes that individual samples of a microarray dataset are related by an unknown biological process (i.e., differentiation, development, cell cycle, disease progression), and that each sample represents one unknown point along the progression of that process. SPD aims to organize the samples in a manner that reveals the underlying progression and to simultaneously identify subsets of genes that are responsible for that progression. We demonstrate the performance of SPD on a variety of microarray datasets that were generated by sampling a biological process at different points along its progression, without providing SPD any information of the underlying process. When applied to a cell cycle time series microarray dataset, SPD was not provided any prior knowledge of samples' time order or of which genes are cell-cycle regulated, yet SPD recovered the correct time order and identified many genes that have been associated with the cell cycle. When applied to B-cell differentiation data, SPD recovered the correct order of stages of normal B-cell differentiation and the linkage between preB-ALL tumor cells with their cell origin preB. When applied to mouse embryonic stem cell differentiation data, SPD uncovered a landscape of ESC differentiation into various lineages and genes that represent both generic and lineage specific processes. When applied to a prostate cancer microarray dataset, SPD identified gene modules that reflect a progression consistent with disease stages. SPD may be best viewed as a novel tool for synthesizing biological hypotheses because it provides a likely biological progression underlying a microarray dataset and, perhaps more importantly, the candidate genes that regulate that progression.
Peng Qiu, Andrew J. Gentles, Sylvia K. Plevritis
PLoS Comput. Biol.1
2009 An Activity-Subspace Approach for Estimating the Integrated Input Function and Relative Distribution Volume in PET Parametric Imaging
abstract
Dynamic positron emission tomography (PET) imaging technique enables the measurement of neuroreceptor distributions corresponding to anatomic structures, and thus, allows image-wide quantification of physiological and biochemical parameters. Accurate quantification of the concentration of neuroreceptor has been the objective of many research efforts. Compartment modeling is the most widely used approach for receptor binding studies. However, current compartment-model-based methods often either require intrusive collection of accurate arterial blood measurements as the input function, or assume the existence of a reference region. To obviate the need for the input function or a reference region, in this paper, we propose to estimate the input function. We propose a novel concept of activity subspace, and estimate the input function by the analysis of the intersection of the activity subspaces. Then, the input function and the distribution volume (DV) parameter are refined and estimated iteratively. Thus, the underlying parametric image of the total DV is obtained. The proposed method is compared with a blind estimation method, iterative quadratic maximum-likelihood (IQML) via simulation, and the proposed method outperforms IQML. The proposed method is also evaluated in a brain PET dataset.
Peng Qiu, Z. Jane Wang 0001, K. J. Ray Liu, Zsolt Szabo
IEEE Trans. Inf. Technol. Biomed.1
2008 A robust method for QRS detection based on modified p-spectrum
abstract
In this paper, we propose a robust method based on the modified p-spectrum to detect heart beats in ECG signals, which is also referred as QRS detection in the literature. QRS detection is an old problem that has been studied for several decades. In the literature, there are many methods based on various forms of transformations and thresholding, which require predetermined or fine-tuned thresholds. Another class of existing methods are based on machine learning, which require carefully labeled training data. In our study, we propose the modified p-spectrum for QRS detection, which does not require either predetermined thresholds or training data, and can operate in real-time. The proposed method is evaluated using the MIT-BIH arrhythmia database.
Peng Qiu, K. J. Ray Liu
ICASSP1
2007 Dependence network modeling for biomarker identification
abstract
MOTIVATION: Our purpose is to develop a statistical modeling approach for cancer biomarker discovery and provide new insights into early cancer detection. We propose the concept of dependence network, apply it for identifying cancer biomarkers, and study the difference between the protein or gene samples from cancer and non-cancer subjects based on mass-spectrometry (MS) and microarray data. RESULTS: Three MS and two gene microarray datasets are studied. Clear differences are observed in the dependence networks for cancer and non-cancer samples. Protein/gene features are examined three at one time through an exhaustive search. Dependence networks are constructed by binding triples identified by the eigenvalue pattern of the dependence model, and are further compared to identify cancer biomarkers. Such dependence-network-based biomarkers show much greater consistency under 10-fold cross-validation than the classification-performance-based biomarkers. Furthermore, the biological relevance of the dependence-network-based biomarkers using microarray data is discussed. The proposed scheme is shown promising for cancer diagnosis and prediction. AVAILABILITY: See supplements: http://dsplab.eng.umd.edu/~genomics/dependencenetwork/
Peng Qiu, Z. Jane Wang 0001, K. J. Ray Liu, Zhang-Zhi Hu, Cathy H. Wu
Bioinform.1
2006 Polynomial model approach for resynchronization analysis of cell-cycle gene expression data
abstract
MOTIVATION: Identification of genes expressed in a cell-cycle-specific periodical manner is of great interest to understand cyclic systems which play a critical role in many biological processes. However, identification of cell-cycle regulated genes by raw microarray gene expression data directly is complicated by the factor of synchronization loss, thus remains a challenging problem. Decomposing the expression measurements and extracting synchronized expression will allow to better represent the single-cell behavior and improve the accuracy in identifying periodically expressed genes. RESULTS: In this paper, we propose a resynchronization-based algorithm for identifying cell-cycle-related genes. We introduce a synchronization loss model by modeling the gene expression measurements as a superposition of different cell populations growing at different rates. The underlying expression profile is then reconstructed through resynchronization and is further fitted to the measurements in order to identify periodically expressed genes. Results from both simulations and real microarray data show that the proposed scheme is promising for identifying cyclic genes and revealing underlying gene expression profiles. AVAILABILITY: Contact the authors. SUPPLEMENTARY INFORMATION: Supplementary data are available at: http://dsplab.eng.umd.edu/~genomics/syn/
Peng Qiu, Z. Jane Wang 0001, K. J. Ray Liu
Bioinform.1
2005 Ensemble dependence model for classification and prediction of cancer and normal gene expression data
abstract
MOTIVATION: DNA microarray technologies make it possible to simultaneously monitor thousands of genes' expression levels. A topic of great interest is to study the different expression profiles between microarray samples from cancer patients and normal subjects, by classifying them at gene expression levels. Currently, various clustering methods have been proposed in the literature to classify cancer and normal samples based on microarray data, and they are predominantly data-driven approaches. In this paper, we propose an alternative approach, a model-driven approach, which can reveal the relationship between the global gene expression profile and the subject's health status, and thus is promising in predicting the early development of cancer. RESULTS: In this work, we propose an ensemble dependence model, aimed at exploring the group dependence relationship of gene clusters. Under the framework of hypothesis-testing, we employ genes' dependence relationship as a feature to model and classify cancer and normal samples. The proposed classification scheme is applied to several real cancer datasets, including cDNA, Affymetrix microarray and proteomic data. It is noted that the proposed method yields very promising performance. We further investigate the eigenvalue pattern of the proposed method, and we discover different patterns between cancer and normal samples. Moreover, the transition between cancer and normal patterns suggests that the eigenvalue pattern of the proposed models may have potential to predict the early stage of cancer development. In addition, we examine the effects of possible model mismatch on the proposed scheme.
Peng Qiu, Z. Jane Wang 0001, K. J. Ray Liu
Bioinform.1