EDBT 2026 Demo / reviewers in the wild / expert
Wenyuan Zhao
dblp:96/2772
· DBLP profile ↗
24ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing Leaky Private Information Retrieval Codes to Achieve O(log K) Leakage Ratio ExponentabstractWe study the problem of leaky private information retrieval (L-PIR), where the amount of privacy leakage is measured by the pure differential privacy parameter, referred to as the leakage ratio exponent. Unlike the previous L-PIR proposed by Samy et al., which is merely a re-allocation of the clean (low-cost) retrieval pattern within the generalized TSC family, we show that the active pure-DP constraints couple adjacent Hamming-weight layers of the random key, which reduces the optimization to a layered problem whose optimum is geometric across these layers. As a result, only cyclic permutations are needed without loss of optimality, and lower-Hamming weight keys should be assigned higher probabilities. This new scheme provides a significant improvement, leading to anO(logK) leakage ratio exponent with fixed download costD, in contrast to the previous art that only achieves a Θ(K) exponent, whereKis the number of messages. Wenyuan Zhao, Yu-Shin Huang, Chao Tian 0002, Alexander Sprintson |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | From Deep Additive Kernel Learning to Last-Layer Bayesian Neural Networks via Induced Prior ApproximationabstractWith the strengths of both deep learning and kernel methods like Gaussian Processes (GPs), Deep Kernel Learning (DKL) has gained considerable attention in recent years. From the computational perspective, however, DKL becomes challenging when the input dimension of the GP layer is high. To address this challenge, we propose the Deep Additive Kernel (DAK) model, which incorporates i) an additive structure for the last-layer GP; and ii) induced prior approximation for each GP unit. This naturally leads to a last-layer Bayesian neural network (BNN) architecture. The proposed method enjoys the interpretability of DKL as well as the computational advantages of BNN. Empirical results show that the proposed approach outperforms state-of-the-art DKL methods in both regression and classification tasks. Wenyuan Zhao, Tie Liu 0002, Rui Tuo 0001, Chao Tian 0002 |
AISTATS | 1 |
| 2025 | Optimizing Leaky Private Information Retrieval Codes to Achieve $O(\log K)$ Leakage Ratio ExponentabstractWe study the problem of leaky private information retrieval (L- PIR), where the amount of privacy leakage is measured by the pure differential privacy parameter, referred to as the leakage ratio exponent. Unlike the previous L-PIR scheme proposed by Samy et al., which only adjusted the probability allocation to the clean (low-cost) retrieval pattern, we optimized the probabilities assigned to all the retrieval patterns jointly. It is demonstrated that the optimal probability distribution of the retrieval pattern is quite sophisticated and has a layered structure: the retrieval associated with the random key values of lower Hamming weights should be assigned higher probabilities. This new scheme provides a significant improvement, leading to an$O(\log K)$leakage ratio exponent with fixed download cost$D$and number of servers$N$, in contrast to the previous art that only achieves a$\theta(K)$exponent, where$K$is the number of messages. Wenyuan Zhao, Yu-Shin Huang, Chao Tian 0002, Alexander Sprintson |
ISIT | 1 |
| 2025 | Partial Information Decomposition via Normalizing Flows in Latent Gaussian DistributionsabstractThe study of multimodality has garnered significant interest in fields where analyzing interactions among multiple information sources can enhance predictive modeling, data fusion, and interpretability. Partial information decomposition (PID) has emerged as a useful information-theoretic framework to quantify the degree to which individual modalities independently, redundantly, or synergistically convey information about a target variable. However, existing PID methods depend on optimizing over a joint distribution constrained by estimated pairwise probability distributions, which are costly and inaccurate for continuous and high-dimensional modalities. Our first key insight is that the problem can be solved efficiently when the pairwise distributions are multivariate Gaussians, and we refer to this problem as Gaussian PID (GPID). We propose a new gradient-based algorithm that substantially enhances computational efficiency for GPID based on an alternative formulation of the underlying optimization problem. To generalize the applicability to non-Gaussian data, we learn information-preserving encoders to transform random variables of arbitrary input distributions into pairwise Gaussian random variables. Along the way, we resolved an open problem regarding the optimality of joint Gaussian solutions for GPID. Empirical validation on diverse synthetic examples demonstrates that our proposed method provides more accurate and efficient PID estimates than existing baselines. We further evaluate on a series of large-scale multimodal benchmarks to show its utility in real-world applications of quantifying PID in multimodal datasets and selecting high-performing models. Wenyuan Zhao, Adithya Balachandran, Chao Tian 0002, Paul Pu Liang |
NeurIPS | 1 |
| 2025 | Weakly Private Information Retrieval From Heterogeneously Trusted ServersabstractWe study the problem of weakly private information retrieval (PIR) when there is heterogeneity in servers’ trustworthiness under the maximal leakage (Max-L) metric and mutual information (MI) metric. A user wishes to retrieve a desired message from N non-colluding servers efficiently, such that the identity of the desired message is not leaked in a significant manner; however, some servers can be more trustworthy than others. We propose a code construction for this setting and optimize the probability distribution for this construction. For the Max-L metric, it is shown that the optimal probability allocation for the proposed scheme essentially separates the delivery patterns into two parts: a completely private part that has the same download overhead as the capacity-achieving PIR code, and a non-private part that allows complete privacy leakage but has no download overhead by downloading only from the most trustful server. The optimal solution is established through a sophisticated analysis of the underlying convex optimization problem and a reduction between the homogeneous setting and the heterogeneous setting. For the MI metric, the homogeneous case is studied first for which the code can be optimized with an explicit probability assignment, while a closed-form solution becomes intractable for the heterogeneous case. Numerical results are provided for both cases to corroborate the theoretical analysis. Wenyuan Zhao, Yu-Shin Huang, Ruida Zhou, Chao Tian 0002 |
IEEE Trans. Inf. Theory | 1 |
| 2024 | Weakly Private Information Retrieval from Heterogeneously Trusted ServersabstractWe study the problem of weakly private information retrieval (PIR) when there is heterogeneity in servers' trustfulness under the maximal leakage (Max-L) metric. A user wishes to retrieve a desired message from$N$non-colluding servers efficiently, such that the identity of the desired message is not leaked in a significant manner; however, some servers can be more trustworthy than others. We propose a code construction for this setting and optimize the probability distribution for this construction. It is shown that the optimal probability allocation for the proposed scheme essentially separates the delivery patterns into two parts: a completely private part that has the same download overhead as the capacity-achieving PIR code, and a non-private part that allows complete privacy leakage but has no download overhead by downloading only from the most trustful server. The optimal solution is established through a sophisticated analysis of the underlying convex optimization problem, and a reduction between the homogeneous setting and the heterogeneous setting. Yu-Shin Huang, Wenyuan Zhao, Ruida Zhou, Chao Tian 0002 |
ISIT | 2 |
| 2024 | Amplitude-Time Dual-View Fused EEG Temporal Feature Learning for Automatic Sleep StagingabstractElectroencephalogram (EEG) plays an important role in studying brain function and human cognitive performance, and the recognition of EEG signals is vital to develop an automatic sleep staging system. However, due to the complex nonstationary characteristics and the individual difference between subjects, how to obtain the effective signal features of the EEG for practical application is still a challenging task. In this article, we investigate the EEG feature learning problem and propose a novel temporal feature learning method based on amplitude-time dual-view fusion for automatic sleep staging. First, we explore the feature extraction ability of convolutional neural networks for the EEG signal from the perspective of interpretability and construct two new representation signals for the raw EEG from the views of amplitude and time. Then, we extract the amplitude-time signal features that reflect the transformation between different sleep stages from the obtained representation signals by using conventional 1-D CNNs. Furthermore, a hybrid dilation convolution module is used to learn the long-term temporal dependency features of EEG signals, which can overcome the shortcoming that the small-scale convolution kernel can only learn the local signal variation information. Finally, we conduct attention-based feature fusion for the learned dual-view signal features to further improve sleep staging performance. To evaluate the performance of the proposed method, we test 30-s-epoch EEG signal samples for healthy subjects and subjects with mild sleep disorders. The experimental results from the most commonly used datasets show that the proposed method has better sleep staging performance and has the potential for the development and application of an EEG-based automatic sleep staging system. Panfeng An, Jianhui Zhao 0001, Bo Du 0001, Wenyuan Zhao, Tingbao Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | A Novel Electromagnetic Positioning Prototype System With Simplified Receiver for Interventional Surgery ApplicationabstractTo meet the need of 3-D positioning of surgical instrument in interventional surgery and the requirement of smaller size sensor for the narrow blood vessels, a new electromagnetic positioning model is proposed with a simplified receiver. Based on the electromagnetic theory and geometry principle, the electromagnetic field transmitter with groups of three orthogonal coils and the simplified receiver with a smaller size than existing sensors are designed. Then, using the Biot–Savart law, the distance between receiving end and geometric center of orthogonal coils is calculated, and the spatial coordinate of receiving end is computed with the spherical intersection formula. To further reduce the positioning error, a two-round accuracy improvement algorithm is designed for selecting the optimal topology of coil groups at the transmitter end and fitting correction of the calculated distance. We implement the prototype system and perform both simulation experiments and actual experiments. The results show that our proposed new electromagnetic positioning method with a simplified receiver has higher accuracy and stronger stability, compared with the traditional positioning approach of three-coil transmitter and three-coil receiver. Peijun Zhong, Jianhui Zhao 0001, Wenyuan Zhao, Tingbao Zhang |
IEEE Internet Things J. | 4 |
| 2022 | Brain Tumor Segmentation Framework Based on Edge Cloud Cooperation and Deep Learning
Saifeng Feng, Jianhui Zhao 0001, Wenyuan Zhao, Tingbao Zhang |
ICANN (1) | 3 |
| 2022 | A Torque-Current Prediction Model Based on GRU for Circumferential Rotation Force Feedback Device
Zekang Qiu, Jianhui Zhao 0001, Chudong Shan, Wenyuan Zhao, Tingbao Zhang |
ICIC (1) | 4 |
| 2022 | IndGOterm: a qualitative method for the identification of individually dysregulated GO terms in cancerabstractIndividual pathway analysis can dissect heterogeneities among different cancer patients and provide efficient guidelines for individualized therapy. However, the existence of the batch effect brings extensive limitations for the application of many individual methods for pathway analysis. Previously, researchers proposed that methods based on within-sample relative expression ordering (REO) of the genes are notably insensitive to 'batch effects'. In this article, we focus on the Gene Ontology (GO) database and propose an individual qualitative GO term analysis method (IndGOterm) based on the REO of genes. Compared with some current widely used single-sample enrichment analysis methods, such as ssGSEA and GSVA, IndGOterm has a predominance of ignoring the batch effects caused by diverse technologies. Through the survival and drug responses analysis, we found IndGOterm could capture more terms connected to cancer than other single-sample enrichment analysis methods. Furthermore, through the application of IndGOterm, we found some terms that present different dysregulation models that manifest heterogenetic in homologous patients. Collectively, these results attested that IndGOterm could capture useful information from patients and be a useful tool to reveal the intrinsic characteristic of cancer. An open-source R statistical analysis package 'IndGOterm' is available at https://github.com/robert19960424/IndGOterm. Jiashuai Zhang, Huiting Xiao, Keru Li, Hengrui Yuan, Rongqiang Yuan, Zhiqiang Chang, Wenyuan Zhao |
Briefings Bioinform. | 10 |
| 2021 | Noisy Mammogram Classification Method Based on New Weighted Fusion FrameworkabstractConvolutional neural network (CNN) has made outstanding performance in the classification of natural light images. However, images in many fields have the characteristics of high noise, low resolution, no color information and small data set, such as mammogram, which will affect the accuracy and robustness of the model. In order to improve the classification accuracy and the noise robustness of convolution network for mammogram images, we design a novel classification model based on the new weighted fusion convolution framework. This method has been improved from the following aspects: firstly, we take the place of traditional max-pooling layer with convolution layer with increased step, which achieves the purpose of down-sampling and extracts features more rationally through back-propagation. Secondly, we fuse multi-level feature maps to make full use of the information contained in the shallow levels and deep levels. At the same time, we design a new fusion method to effectively fuse the feature maps from different layers with different sizes. Finally, our model is tested on the mammographic image analysis society (MIAS), which is a mammographic medical image dataset. The experimental results show that the average accuracy of the model is as high as 97.6%, and the convolution layer with increased step has better robustness than the traditional max-pooling layer. Jianhui Zhao 0001, Saifeng Feng, Wenyuan Zhao, Tingbao Zhang |
IJCNN | 5 |
| 2021 | Multilevel prioritization of gene regulators associated with consensus molecular subtypes of colorectal cancerabstractConsensus molecular subtypes (CMSs) are emerging as critical factor for prognosis and treatment of colorectal cancer. Gene regulators, including chromatin regulator, RNA-binding protein and transcriptional factor, are critical modulators of cancer hallmark, yet little is known regarding the underlying functional mechanism in CMSs. Herein, we identified a core set of 235 functional gene regulators (FGRs) by integrating genome, epigenome, transcriptome and interactome of CMSs. FGRs exhibited significant multi-omics alterations and impacts on cell lines growth, as well as significantly enriched cancer driver genes and pathways. Moreover, common FGRs played different roles in the context of CMSs. In accordance with the immune characteristics of CMSs, we found that the anti-tumor immune pathways were mainly activated by FGRs (e.g. STAT1 and CREBBP) in CMS1, while inhibited by FGRs in CMS2-4. FGRs mediated aberrant expression of ligands, which bind to receptor on immune cells, and modulated tumor immune microenvironment of subtypes. Intriguingly, systematic exploration of datasets using genomic and transcriptome co-similarity reveals the coordinated manner in FGRs act in CMSs to orchestrate their pathways and patients' prognosis. Expression signatures of the FGRs revealed an optimized CMS classifier, which demonstrated 88% concordance with the gold-standard classifier, but avoiding the influence of sample composition. Overall, our integrative analysis identified FGRs to regulate core tumorigenic processes/pathways across CMSs. Hailong Zheng, Liangliang Jin, Huiting Xiao, Jiashuai Zhang, Zhangxiang Zhao, Wenyuan Zhao |
Briefings Bioinform. | 10 |
| 2021 | Reference genome and annotation updates lead to contradictory prognostic predictions in gene expression signatures: a case study of resected stage I lung adenocarcinomaabstractRNA-sequencing enables accurate and low-cost transcriptome-wide detection. However, expression estimates vary as reference genomes and gene annotations are updated, confounding existing expression-based prognostic signatures. Herein, prognostic 9-gene pair signature (GPS) was applied to 197 patients with stage I lung adenocarcinoma derived from previous and latest data from The Cancer Genome Atlas (TCGA) processed with different reference genomes and annotations. For 9-GPS, 6.6% of patients exhibited discordant risk classifications between the two TCGA versions. Similar results were observed for other prognostic signatures, including IRGPI, 15-gene and ORACLE. We found that conflicting annotations for gene length and overlap were the major cause of their discordant risk classification. Therefore, we constructed a prognostic 40-GPS based on stable genes across GENCODE v20-v30 and validated it using public data of 471 stage I samples (log-rank P < 0.0010). Risk classification was still stable in RNA-sequencing data processed with the newest GENCODE v32 versus GENCODE v20-v30. Specifically, 40-GPS could predict survival for 30 stage I samples with formalin-fixed paraffin-embedded tissues (log-rank P = 0.0177). In conclusion, this method overcomes the vulnerability of existing prognostic signatures due to reference genome and annotation updates. 40-GPS may offer individualized clinical applications due to its prognostic accuracy and classification stability. Zheyang Zhang, Zhangxiang Zhao, Changjing Chen, Juxuan Zhang, Mengyue Li, Zixin Wei, Wenbin Jiang 0008, Ying Li 0028, Yingyue Cao, Wenyuan Zhao, Yunyan Gu, Qingwei Meng, Lishuang Qi |
Briefings Bioinform. | 14 |
| 2021 | An absolute human stemness index associated with oncogenic dedifferentiationabstractThe progression of cancer is accompanied by the acquisition of stemness features. Many stemness evaluation methods based on transcriptional profiles have been presented to reveal the relationship between stemness and cancer. However, instead of absolute stemness index values-the values with certain range-these methods gave the values without range, which makes them unable to intuitively evaluate the stemness. Besides, these indices were based on the absolute expression values of genes, which were found to be seriously influenced by batch effects and the composition of samples in the dataset. Recently, we have showed that the signatures based on the relative expression orderings (REOs) of gene pairs within a sample were highly robust against these factors, which makes that the REO-based signatures have been stably applied in the evaluations of the continuous scores with certain range. Here, we provided an absolute REO-based stemness index to evaluate the stemness. We found that this stemness index had higher correlation with the culture time of the differentiated stem cells than the previous stemness index. When applied to the cancer and normal tissue samples, the stemness index showed its significant difference between cancers and normal tissues and its ability to reveal the intratumor heterogeneity at stemness level. Importantly, higher stemness index was associated with poorer prognosis and greater oncogenic dedifferentiation reflected by histological grade. All results showed the capability of the REO-based stemness index to assist the assignment of tumor grade and its potential therapeutic and diagnostic implications. Hailong Zheng, Yelin Fu, Tianyi You, Wenbing Guo, Liangliang Jin, Yunyan Gu, Lishuang Qi, Wenyuan Zhao |
Briefings Bioinform. | 11 |
| 2019 | LIRAT: Layout and Image Recognition Driving Automated Mobile Testing of Cross-PlatformabstractThe fragmentation issue spreads over multiple mobile platforms such as Android, iOS, mobile web, and WeChat, which hinders test scripts from running across platforms. To reduce the cost of adapting scripts for various platforms, some existing tools apply conventional computer vision techniques to replay the same script on multiple platforms. However, because these solutions can hardly identify dynamic or similar widgets. It becomes difficult for engineers to apply them in practice. In this paper, we present an image-driven tool, namely LIRAT, to record and replay test scripts cross platforms, solving the problem of test script cross-platform replay for the first time. LIRAT records screenshots and layouts of the widgets, and leverages image understanding techniques to locate them in the replay process. Based on accurate widget localization, LIRAT supports replaying test scripts across devices and platforms. We employed LIRAT to replay 25 scripts from 5 application across 8 Android devices and 2 iOS devices. The results show that LIRAT can replay 88% scripts on Android platforms and 60% on iOS platforms. The demo can be found at: https: //github.com/YSC9848/LIRAT. Shengcheng Yu, Chunrong Fang, Yang Feng 0003, Wenyuan Zhao, Zhenyu Chen 0001 |
ASE | 4 |
| 2019 | Link synthetic lethality to drug sensitivity of cancer cellsabstractSynthetic lethal (SL) interactions occur when alterations in two genes lead to cell death but alteration in only one of them is not lethal. SL interactions provide a new strategy for molecular-targeted cancer therapy. Currently, there are few drugs targeting SL interactions that entered into clinical trials. Therefore, it is necessary to investigate the link between SL interactions and drug sensitivity of cancer cells systematically for drug development purpose. We identified SL interactions by integrating the high-throughput data from The Cancer Genome Atlas, small hairpin RNA data and genetic interactions of yeast. By integrating SL interactions from other studies, we tested whether the SL pairs that consist of drug target genes and the genes with genomic alterations are related with drug sensitivity of cancer cells. We found that only 6.26%∼34.61% of SL interactions showed the expected significant drug sensitivity using the pooled cancer cell line data from different tissues, but the proportion increased significantly to approximately 90% using the cancer cell line data for each specific tissue. From an independent pharmacogenomics data of 41 breast cancer cell lines, we found three SL interactions (ABL1-IFI16, ABL1-SLC50A1 and ABL1-SYT11) showed significantly better prognosis for the patients with both genes being altered than the patients with only one gene being altered, which partially supports the SL effect between the gene pairs. Our study not only provides a new way for unraveling the complex mechanisms of drug sensitivity but also suggests numerous potentially important drug targets for cancer therapy. Ruiping Wang 0002, Zhangxiang Zhao, Xianlong Wang 0002, Lishuang Qi, Wenyuan Zhao, Zheng Guo 0002, Yunyan Gu |
Briefings Bioinform. | 9 |
| 2018 | A landscape of synthetic viable interactions in cancerabstractSynthetic viability, which is defined as the combination of gene alterations that can rescue the lethal effects of a single gene alteration, may represent a mechanism by which cancer cells resist targeted drugs. Approaches to detect synthetic viable (SV) interactions in cancer genome to investigate drug resistance are still scarce. Here, we present a computational method to detect synthetic viability-induced drug resistance (SVDR) by integrating the multidimensional data sets, including copy number alteration, whole-exome mutation, expression profile and clinical data. SVDR comprehensively characterized the landscape of SV interactions across 8580 tumors in 32 cancer types by integrating The Cancer Genome Atlas data, small hairpin RNA-based functional experimental data and yeast genetic interaction data. We revealed that the SV interactions are favorable to cells and can predict clinical prognosis for cancer patients, which were robustly observed in an independent data set. By integrating the cancer pharmacogenomics data sets from Cancer Cell Line Encyclopedia (CCLE) and Broad Cancer Therapeutics Response Portal, we have demonstrated that SVDR enables drug resistance prediction and exhibits high reliability between two databases. To our knowledge, SVDR is the first genome-scale data-driven approach for the identification of SV interactions related to drug resistance in cancer cells. This data-driven approach lays the foundation for identifying the genomic markers to predict drug resistance and successfully infers the potential drug combination for anti-cancer therapy. Yunyan Gu, Ruiping Wang 0002, Zhangxiang Zhao, Fuduan Peng, Haihai Liang, Lishuang Qi, Wenyuan Zhao, Da Yang 0003, Zheng Guo 0002 |
Briefings Bioinform. | 11 |
| 2016 | Critical limitations of prognostic signatures based on risk scores summarized from gene expression levels: a case study for resected stage I non-small-cell lung cancerabstractMost of current gene expression signatures for cancer prognosis are based on risk scores, usually calculated as some summaries of expression levels of the signature genes, whose applications require presetting risk score thresholds and data normalization. In this study, we demonstrate the critical limitations of such type of signatures that the risk scores of samples will change greatly when they are normalized together with different samples, which would induce spurious risk classification and difficulty in clinical settings, and the risk scores of independent samples are incomparable if data normalization is not adopted. To overcome these limitations, we propose a rank-based method to extract a prognostic gene pair signature for overall survival of stage I non-small-cell lung cancer. The prognostic gene pair signature is verified in three integrated data sets detected by different laboratories with different microarray platforms. We conclude that, different from the type of signatures based on risk scores summarized from gene expression levels, the rank-based signatures could be robustly applied at the individualized level to independent clinical samples assessed in different laboratories. Lishuang Qi, Libin Chen, Yang Li 0023, Rufei Pan, Wenyuan Zhao, Yunyan Gu, Hongwei Wang 0003, Ruiping Wang 0002, Xiangqi Chen, Zheng Guo 0002 |
Briefings Bioinform. | 6 |
| 2016 | Individualized identification of disease-associated pathways with disrupted coordination of gene expressionabstractCurrent pathway analysis approaches are primarily dedicated to capturing deregulated pathways at the population level and cannot provide patient-specific pathway deregulation information. In this article, the authors present a simple approach, called individPath, to detect pathways with significantly disrupted intra-pathway relative expression orderings for each disease sample compared with the stable, normal intra-pathway relative expression orderings pre-determined in previously accumulated normal samples. Through the analysis of multiple microarray data sets for lung and breast cancer, the authors demonstrate individPath's effectiveness for detecting cancer-associated pathways with disrupted relative expression orderings at the individual level and dissecting the heterogeneity of pathway deregulation among different patients. The portable use of this simple approach in clinical contexts is exemplified by the identification of prognostic intra-pathway gene pair signatures to predict overall survival of resected early-stage lung adenocarcinoma patients and signatures to predict relapse-free survival of estrogen receptor-positive breast cancer patients after tamoxifen treatment. Hongwei Wang 0003, Lu Ao, Haidan Yan, Wenyuan Zhao, Lishuang Qi, Yunyan Gu, Zheng Guo 0002 |
Briefings Bioinform. | 5 |
| 2015 | Individual-level analysis of differential expression of genes and pathways for personalized medicineabstractMOTIVATION: The differential expression analysis focusing on inter-group comparison can capture only differentially expressed genes (DE genes) at the population level, which may mask the heterogeneity of differential expression in individuals. Thus, to provide patient-specific information for personalized medicine, it is necessary to conduct differential expression analysis at the individual level. RESULTS: We proposed a method to detect DE genes in individual disease samples by using the disrupted ordering in individual disease samples. In both simulated data and real paired cancer-normal sample data, this method showed excellent performance. It was found to be insensitive to experimental batch effects and data normalization. The landscape of stable gene pairs in a particular type of normal tissue could be predetermined using previously accumulated data, based on which dysregulated genes and pathways for any disease sample can be readily detected. The usefulness of the RankComp method in clinical settings was exemplified by the identification and application of prognostic markers for lung cancer. AVAILABILITY AND IMPLEMENTATION: RankComp is implemented in R script that is freely available from Supplementary Materials. Hongwei Wang 0003, Wenyuan Zhao, Lishuang Qi, Yunyan Gu, Pengfei Li 0002, Yang Li 0023, Zheng Guo 0002 |
Bioinform. | 3 |
| 2012 | GO-function: deriving biologically relevant functions from statistically significant functionsabstractIn high-throughput studies of diseases, terms enriched with disease-related genes based on Gene Ontology (GO) are routinely found. However, most current algorithms used to find significant GO terms cannot handle the redundancy that results from the dependencies of GO terms. Simply based on some numerical considerations, current algorithms developed for reducing this redundancy may produce results that do not account for biologically interesting cases. In this article, we present several rules used to design a tool called GO-function for extracting biologically relevant terms from statistically significant GO terms for a disease. Using one gene expression profile for colorectal cancer, we compared GO-function with four algorithms designed to treat redundancy. Then, we validated results obtained in this data set by GO-function using another data set for colorectal cancer. Our analysis showed that GO-function can identify disease-related terms that are more statistically and biologically meaningful than those found by the other four algorithms. Jing Wang 0004, Xianxiao Zhou, Jing Zhu 0004, Yunyan Gu, Wenyuan Zhao, Jinfeng Zou, Zheng Guo 0002 |
Briefings Bioinform. | 5 |
| 2012 | Temporal and spatial relationship of gene expression in the infarcted rat heartabstractBackground Gene interaction plays an important role in regulating the molecular and cellular actions in both the physiological and pathological circumstances. Following myocardial infarction (MI), cardiac repair/remodeling represented as cardiac inflammation, angiogenesis, apoptosis, fibrosis occur in the infarcted and noninfarcted myocardium. Local angiotensin system (RAS) is involved in the cardiac repair. Wenyuan Zhao, Tieqian Zhao, Yuanjian Chen |
BMC Bioinform. | 1 |
| 2010 | Extracting consistent knowledge from highly inconsistent cancer gene data sourcesabstractBACKGROUND: Hundreds of genes that are causally implicated in oncogenesis have been found and collected in various databases. For efficient application of these abundant but diverse data sources, it is of fundamental importance to evaluate their consistency. RESULTS: First, we showed that the lists of cancer genes from some major data sources were highly inconsistent in terms of overlapping genes. In particular, most cancer genes accumulated in previous small-scale studies could not be rediscovered in current high-throughput genome screening studies. Then, based on a metric proposed in this study, we showed that most cancer gene lists from different data sources were highly functionally consistent. Finally, we extracted functionally consistent cancer genes from various data sources and collected them in our database F-Census. CONCLUSIONS: Although they have very low gene overlapping, most cancer gene data sources are highly consistent at the functional level, which indicates that they can separately capture partial genes in a few key pathways associated with cancer. Our results suggest that the sample sizes currently used for cancer studies might be inadequate for consistently capturing individual cancer genes, but could be sufficient for finding a number of cancer genes that could represent functionally most cancer genes. The F-Census database provides biologists with a useful tool for browsing and extracting functionally consistent cancer genes from various data sources. Ruihong Wu, Yuannv Zhang, Wenyuan Zhao, Lixin Cheng, Yunyan Gu, Lin Zhang 0057, Jing Wang 0004, Jing Zhu 0004, Zheng Guo 0002 |
BMC Bioinform. | 4 |