Yunpeng Cai

dblp:80/6847 · DBLP profile ↗
← Back
41ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0001-8797-4243ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021
YearPublicationVenuePosition
2026 Teacher-Student Instance-Level Adversarial Augmentation for Single Domain Generalized Medical Image Segmentation
abstract
Recently, single-source domain generalization (SDG) has gained popularity in medical image segmentation. As a prominent technique, adversarial image augmentation technique can generate synthetic training data that are challenging for the segmentation model to recognize. To avoid the over-augmentation problem, existing adversarial-based works often employ augmenters with relatively simple structures for medical images, typically operating at the image level, limiting the diversity of the augmented images. In this paper, we propose a Teacher-Student Instance-level Adversarial Augmentation (TSIAA) model for generalized medical image segmentation. The objective of TSIAA is to derive domain-generalizable representations by exploring out-of-source data distributions. First, we construct an Instance-level Image Augmenter (IIAG) using several Instance-level Augmentation Modules (IAMs), which are based on the learnable constrained Bèzier transformation function. Compared to image-level adversarial augmentation, instance-level adversarial augmentation breaks the uniformity of augmentation rules across different structures within an image, thereby providing greater diversity. Then, TSIAA conducts Teacher-Student (TS) learning through an adversarial approach, alternating novel image augmentation and generalized representation learning. The former delves into out-of-source and plausible data, while the latter continuously updates both the student and teacher to ensure the original and augmented features maintain consistent and generalized characteristics. By integrating both strategies, our proposed TSIAA model achieves significant improvements over state-of-the-art methods in four challenging SDG tasks. The code can be accessed at https://github.com/Wangzs0228/TSIAA.
Zhengshan Wang, Long Chen 0001, Xuelin Xie, Yang Zhang 0053, Yunpeng Cai, Weiping Ding 0001
IEEE Trans. Medical Imaging5
2025 Lesion Localization for Medical Imaging Using Counter-factual Generation Prompt Learning
abstract
Lesion localization using machines greatly assists doctors in diagnosing diseases and providing better treatment for patients, which is significant for intelligent healthcare. Unlike natural images, the background and target objects in medical images are often difficult to distinguish, and obtaining medical annotations for training models is challenging. This makes accurate lesion localization extremely difficult. In this paper, we propose a counter-factual generation prompt learning framework for lesion localization in medical images. First, we employ the class association embedding method for separating lesion-related information from lesion-irrelevant information in medical images. By embedding different lesion-related information, we generate counterfactual samples, and obtain lesion-related knowledge based on comparison. We further process the lesion-related knowledge and obtain a prior prompt, which is then fed into a well-known segmentation network for more accurate and detailed lesion localization. To accurately acquire lesion-related knowledge, we propose an irrelevant feature similarity transfer method to reduce the interference of irrelevant knowledge. Experimental results show that our method achieves excellent lesion localization results without requiring pixel-level annotations for training, and also outperforms other existing localization algorithms.
Yi Pan 0001, Limai Jiang, Juan He 0006, Yufu Huo, Yunpeng Cai, Ruitao Xie
ICME7
2025 Accurate and Interpretable Wound Healing Progress Detection Based on a Task-Related Knowledge Refinement Learning Method
Juan He 0006, Yi Pan 0001, Zhengshan Wang, Tzu-Ming Liu, Yunpeng Cai, Long Chen 0001, Ruitao Xie
ISBRA (2)8
2025 Multi-modal Graph Diffusion Model for Depression Detection
Yufu Huo, Ruitao Xie, Yunpeng Cai
PRCV (13)3
2025 Automated Learning of Semantic Embedding Representations for Diffusion Models
abstract
Generative models capture the true distribution of data, yielding semantically rich representations. Denoising diffusion models (DDMs) exhibit superior generative capabilities, though efficient representation learning for them are lacking. In this work, we employ a multi-level denoising autoencoder framework to expand the representation capacity of DDMs, which introduces sequentially consistent Diffusion Transformers and an additional timestep-dependent encoder to acquire embedding representations on the denoising Markov chain through self-conditional diffusion learning. Intuitively, the encoder, conditioned on the entire diffusion process, compresses high-dimensional data into directional vectors in latent under different noise levels, facilitating the learning of image embeddings across all timesteps. To verify the semantic adequacy of embeddings generated through this approach, extensive experiments are conducted on various datasets, demonstrating that optimally learned embeddings by DDMs surpass state-of-the-art self-supervised representation learning methods in most cases, achieving remarkable discriminative semantic representation quality. Our work justifies that DDMs are not only suitable for generative tasks, but also potentially advantageous for general-purpose deep learning applications.
Limai Jiang, Yunpeng Cai
SDM2
2025 Weakly supervised lesion localization and attribution for OCT images with a guided counterfactual explainer model
Limai Jiang, Ruitao Xie, Juan He 0006, Huazhen Huang, Yi Pan 0001, Yunpeng Cai
Expert Syst. Appl.7
2024 Accurate Explanation Model for Image Classifiers using Class Association Embedding
abstract
Image classification is a primary task in data analy-sis where explainable models are crucially demanded in various applications. Although amounts of methods have been proposed to obtain explainable knowledge from the black-box classifiers, these approaches lack the efficiency of extracting global knowl-edge regarding the classification task, thus is vulnerable to local traps and often leads to poor accuracy. In this study, we propose a generative explanation model that combines the advantages of global and local knowledge for explaining image classifiers. We develop a representation learning method called class association embedding (CAE), which encodes each sample into a pair of separated class-associated and individual codes. Recombining the individual code of a given sample with altered class-associated code leads to a synthetic real-looking sample with preserved individual characters but modified class-associated features and possibly flipped class assignments. A building-block coherency feature extraction algorithm is proposed that efficiently separates class-associated features from individual ones. The extracted feature space forms a low-dimensional manifold that visualizes the classification decision patterns. Explanation on each individual sample can be then achieved in a counter-factual generation manner which continuously modifies the sample in one direction, by shifting its class-associated code along a guided path, until its classification outcome is changed. We compare our method with state-of-the-art ones on explaining image classification tasks in the form of saliency maps, demonstrating that our method achieves higher accuracies. The class-associated manifold not only helps with skipping local traps and achieving accurate explanation, but also provides insights to the data distribution patterns that potentially aids knowledge discovery. The code is available at https://github.com/xrtll/xAI-CODE.
Ruitao Xie, Limai Jiang, Yi Pan 0001, Yunpeng Cai
ICDE6
2024 A Weakly Supervised and Globally Explainable Learning Framework for Brain Tumor Segmentation
abstract
Machine-based brain tumor segmentation can help doctors make better diagnoses. However, the complex structure of brain tumors and expensive pixel-level annotations present challenges for automatic tumor segmentation. In this paper, we propose a counterfactual generation framework that not only achieves exceptional brain tumor segmentation performance without the need for pixel-level annotations, but also provides explainability. Our framework effectively separates class-related features from class-unrelated features of the samples, and generate new samples that preserve identity features while altering class attributes by embedding different class-related features. We perform topological data analysis on the extracted class-related features and obtain a globally explainable manifold, and for each abnormal sample to be segmented, a meaningful normal sample could be effectively generated with the guidance of the rule-based paths designed within the manifold for comparison for identifying the tumor regions. We evaluate our proposed method on two datasets, which demonstrates superior performance of brain tumor segmentation. The code is available at https://github.com/xrt11/tumor-segmentation.
Ruitao Xie, Limai Jiang, Xiaoxi He, Yi Pan 0001, Yunpeng Cai
ICME5
2024 gaBERT: An Interpretable Pretrained Deep Learning Framework for Cancer Gene Marker Discovery
Jiale Hou, Xinzhe Pang, Yunpeng Cai
ISBRA (1)5
2024 RFIR: A Lightweight Network for Retinal Fundus Image Restoration
Limai Jiang, Yunpeng Cai
ISBRA (1)3
2024 A generalized integrated framework for urban public transport operations evaluation based on interval neutrosophic TODIM and EDAS technique
Yunpeng Cai
Soft Comput.1
2024 HybAVPnet: A Novel Hybrid Network Architecture for Antiviral Peptides Prediction
abstract
Viruses pose a great threat to human production and life, thus the research and development of antiviral drugs is urgently needed. Antiviral peptides play an important role in drug design and development. Compared with the time-consuming and laborious wet chemical experiment methods, it is critical to use computational methods to predict antiviral peptides accurately and rapidly. However, due to limited data, accurate prediction of antiviral peptides is still challenging and extracting effective feature representations from sequences is crucial for creating accurate models. This study introduces a novel two-step approach, named HybAVPnet, to predict antiviral peptides with a hybrid network architecture based on neural networks and traditional machine learning methods. We adopted a stacking-like structure to capture both the long-term dependencies and local evolution information to achieve a comprehensive and diverse prediction using the predicted labels and probabilities. Using an ensemble technique with the different kinds of features can reduce the variance without increasing the bias. The experimental result shows HybAVPnet can achieve better and more robust performance compared with the state-of-the-art methods, which makes it useful for the research and development of antiviral drugs. Meanwhile, it can also be extended to other peptide recognition problems because of its generalization ability.
Ruiquan Ge, Yixiao Xia, Minchao Jiang, Gangyong Jia, Xiaoyang Jing, Ye Li 0002, Yunpeng Cai
IEEE ACM Trans. Comput. Biol. Bioinform.7
2023 VirusBERTHP: Improved Virus Host Prediction Via Attention-based Pre-trained Model Using Viral Genomic Sequences
abstract
Virus has become the most prominent cause of infectious diseases which greately threaten human health. Determining whether a viral genome can possess human host infectivity would be of great value to epidemic prevention. However, due to the highly diversified and unstructured nature of virus genomes, current bioinformatic and machine learning methods for prediction virus host infectivities are rather limited in performance. In this paper we propose an accurate virus human host infectivity prediction tool, VirusBERTHP, using an attention-based pretraining mechanism following the well-known BERT architecture, which is capable of predicting the human infectivity of a novel virus species whose genome is not in the training database. We develop a BERT-based representation learning scheme, VirusBERT, to efficiently extract the complex feature among versatile virus sequences, which show greate seperability in the feature space. We created a large curated database containing 2,948,656 unlabelled virus sequences to efficiently pre-train the VirusBERT model. Then, the VirusBERTHP model is trained with a relatively smaller set of labelled sequences corresponding to specific tasks, using a full-connected deep neural network. We adopted the model on four published virus-host classification datasets and showed that our model outperforms previous state-of-the-art methods in prediction performance. On three datasets with open-view setting where no restriction is imposed on the taxonomy of the input virus sequences, our model achieved more than 99% accuracy in predicting human host infectivity, justifying the efficiency of our method. In addition to accuracy boost, our model is adaptive to various virus sequence prediction task by seperating the pretraining and supervised learning phases. In addition, the model is adaptable to a wide range of sequence lengths from 250bps to 10k bps, expanding the application field of the model. Source code and data of our paper is available at https://github.com/wyzwyzwyz/virusBert/.
Yunzhan Wang, Yunpeng Cai
BIBM3
2023 AdaPPI: identification of novel protein functional modules via adaptive graph convolution networks in a protein-protein interaction network
abstract
Identifying unknown protein functional modules, such as protein complexes and biological pathways, from protein-protein interaction (PPI) networks, provides biologists with an opportunity to efficiently understand cellular function and organization. Finding complex nonlinear relationships in underlying functional modules may involve a long-chain of PPI and pose great challenges in a PPI network with an unevenly sparse and dense node distribution. To overcome these challenges, we propose AdaPPI, an adaptive convolution graph network in PPI networks to predict protein functional modules. We first suggest an attributed graph node presentation algorithm. It can effectively integrate protein gene ontology attributes and network topology, and adaptively aggregates low- or high-order graph structural information according to the node distribution by considering graph node smoothness. Based on the obtained node representations, core cliques and expansion algorithms are applied to find functional modules in PPI networks. Comprehensive performance evaluations and case studies indicate that the framework significantly outperforms state-of-the-art methods. We also presented potential functional modules based on their confidence.
Yunpeng Cai, Chaojie Ji, Gurudeeban Selvaraj
Briefings Bioinform.2
2023 Graph Polish: A Novel Graph Generation Paradigm for Molecular Optimization
abstract
Molecular optimization, which transforms a given input molecule X into another Y with desired properties, is essential in molecular drug discovery. The traditional approaches either suffer from sample-inefficient learning or ignore information that can be captured with the supervised learning of optimized molecule pairs. In this study, we present a novel molecular optimization paradigm, Graph Polish. In this paradigm, with the guidance of the source and target molecule pairs of the desired properties, a heuristic optimization solution can be derived: given an input molecule, we first predict which atom can be viewed as the optimization center, and then the nearby regions are optimized around this center. We then propose an effective and efficient learning framework, Teacher and Student polish, to capture the dependencies in the optimization steps. A teacher component automatically identifies and annotates the optimization centers and the preservation, removal, and addition of some parts of the molecules; a student component learns these knowledges and applies them to a new molecule. The proposed paradigm can offer an intuitive interpretation for the molecular optimization result. Experiments with multiple optimization tasks are conducted on several benchmark datasets. The proposed approach achieves a significant advantage over the six state-of-the-art baseline methods. Also, extensive studies are conducted to validate the effectiveness, explainability, and time savings of the novel optimization paradigm.
Chaojie Ji, Ruxin Wang 0001, Yunpeng Cai
IEEE Trans. Neural Networks Learn. Syst.4
2022 A machine learning model for disease risk prediction by integrating genetic and non-genetic factors
abstract
Polygenic risk score (PRS) has been widely used to identify the high-risk individuals from the general population, which would be helpful for disease prevention and early treatment. Many methods have been developed to calculate PRS by weighting and aggregating the phenotype-associated risk alleles from genome-wide association studies. However, only considering genetic effects may not be sufficient for risk prediction because the disease risk is not only related to genetic factors but also non-genetic factors, e.g., diet, physical exercise et al. But it is still a challenge to integrate these genetic and non-genetic factors into a unified machine learning framework for disease risk prediction. In this paper, we proposed PRSIMD (PRS Integrating Multi-source Data), a machine learning model that applies posterior regularization to integrate genetic and non-genetic factors to improve disease risk prediction. Also, we applied Mendelian Randomization analysis to identify the causal non-genetic risk factors for the selected diseases. We applied PRSIMD to predict type 2 diabetes and coronary artery disease from UK Biobank and observed that PRSIMD was significantly better than the existing methods to calculate PRS. In addition, we observed that PRSIMD achieved the better predictive power than the composite risk score. The codes of PRSIMD are available at: https://github.con ericcombiolab/PRSIMD
Chonghao Wang, Yunpeng Cai, Ouzhou Young, Aiping Lyu, Lu Zhang 0061
BIBM4
2022 M-US-EMRs: A Multi-modal Data Fusion Method of Ultrasonic Images and Electronic Medical Records Used for Screening of Coronary Heart Disease
Ying Nan Zuo, Shunxiang Yang, Genqiang Deng, Suisong Zhu, Yunpeng Cai
ISBRA6
2022 Smoothness Sensor: Adaptive Smoothness-Transition Graph Convolutions for Attributed Graph Clustering
abstract
Clustering techniques attempt to group objects with similar properties into a cluster. Clustering the nodes of an attributed graph, in which each node is associated with a set of feature attributes, has attracted significant attention. Graph convolutional networks (GCNs) represent an effective approach for integrating the two complementary factors of node attributes and structural information for attributed graph clustering. Smoothness is an indicator for assessing the degree of similarity of feature representations among nearby nodes in a graph. Oversmoothing in GCNs, caused by unnecessarily high orders of graph convolution, produces indistinguishable representations of nodes, such that the nodes in a graph tend to be grouped into fewer clusters, and pose a challenge due to the resulting performance drop. In this study, we propose a smoothness sensor for attributed graph clustering based on adaptive smoothness-transition graph convolutions, which senses the smoothness of a graph and adaptively terminates the current convolution once the smoothness is saturated to prevent oversmoothing. Furthermore, as an alternative to graph-level smoothness, a novel fine-grained nodewise-level assessment of smoothness is proposed, in which smoothness is computed in accordance with the neighborhood conditions of a given node at a certain order of graph convolution. In addition, a self-supervision criterion is designed considering both the tightness within clusters and the separation between clusters to guide the entire neural network training process. The experiments show that the proposed methods significantly outperform 13 other state-of-the-art baselines in terms of different metrics across five benchmark datasets. In addition, an extensive study reveals the reasons for their effectiveness and efficiency.
Chaojie Ji, Ruxin Wang 0001, Yunpeng Cai
IEEE Trans. Cybern.4
2020 A novel hybrid network of fusing rhythmic and morphological features for atrial fibrillation detection on mobile ECG signals
Xiaomao Fan, Zhejing Hu, Ruxin Wang 0001, Liyan Yin, Ye Li 0002, Yunpeng Cai
Neural Comput. Appl.6
2019 Finding High-Order Homologous Microbe Community Modules via Network Embedding
abstract
Microbial network analysis help with discovering microbe groups that covariate with environmental factors. However, microbial communities are highly diversified and localized, which poses challenges to existing correlation-based network construction methods in terms of stability and functional significance. In this paper, we propose to explore the high-level relationships in the microbial network structure with the aid of network embedding methods. Microbial function modules are then extracted by spectrum clustering on the embedded networks, rather than the original ones. By investigating the correlation between the obtained modules and the environmental factors on several real-world microbial datasets, we demonstrate that the embedded modules provide feature information of the microbial community that are distinct to traditional correlation-based network modules. Furthermore, we show that the introduction of high-order modules helps with improving the performance of prediction models comparing with using OTU features or traditional correlation-based modules alone. Our study demonstrated that high-order network modules created by network embedding can be served as a potential new biomarker for feature extraction of microbial communities.
Qianyin Li, Zhiyong Tao, Yunpeng Cai
BIBE5
2019 A parallel computational framework for ultra-large-scale sequence clustering analysis
abstract
Motivation: The rapid development of sequencing technology has led to an explosive accumulation of genomic data. Clustering is often the first step to be performed in sequence analysis. However, existing methods scale poorly with respect to the unprecedented growth of input data size. As high-performance computing systems are becoming widely accessible, it is highly desired that a clustering method can easily scale to handle large-scale sequence datasets by leveraging the power of parallel computing. Results: In this paper, we introduce SLAD (Separation via Landmark-based Active Divisive clustering), a generic computational framework that can be used to parallelize various de novo operational taxonomic unit (OTU) picking methods and comes with theoretical guarantees on both accuracy and efficiency. The proposed framework was implemented on Apache Spark, which allows for easy and efficient utilization of parallel computing resources. Experiments performed on various datasets demonstrated that SLAD can significantly speed up a number of popular de novo OTU picking methods and meanwhile maintains the same level of accuracy. In particular, the experiment on the Earth Microbiome Project dataset (∼2.2B reads, 437 GB) demonstrated the excellent scalability of the proposed method. Availability and implementation: Open-source software for the proposed method is freely available at https://www.acsu.buffalo.edu/~yijunsun/lab/SLAD.html. Supplementary information: Supplementary data are available at Bioinformatics online.
Wei Zheng 0010, Qi Mao 0001, Robert J. Genco, Jean Wactawski-Wende, Michael J. Buck, Yunpeng Cai, Yijun Sun
Bioinform.6
2018 A benchmark study of sequence alignment methods for protein clustering
abstract
BACKGROUND: Protein sequence alignment analyses have become a crucial step for many bioinformatics studies during the past decades. Multiple sequence alignment (MSA) and pair-wise sequence alignment (PSA) are two major approaches in sequence alignment. Former benchmark studies revealed drawbacks of MSA methods on nucleotide sequence alignments. To test whether similar drawbacks also influence protein sequence alignment analyses, we propose a new benchmark framework for protein clustering based on cluster validity. This new framework directly reflects the biological ground truth of the application scenarios that adopt sequence alignments, and evaluates the alignment quality according to the achievement of the biological goal, rather than the comparison on sequence level only, which averts the biases introduced by alignment scores or manual alignment templates. Compared with former studies, we calculate the cluster validity score based on sequence distances instead of clustering results. This strategy could avoid the influence brought by different clustering methods thus make results more dependable. RESULTS: Results showed that PSA methods performed better than MSA methods on most of the BAliBASE benchmark datasets. Analyses on the 80 re-sampled benchmark datasets constructed by randomly choosing 90% of each dataset 10 times showed similar results. CONCLUSIONS: These results validated that the drawbacks of MSA methods revealed in nucleotide level also existed in protein sequence alignment analyses and affect the accuracy of results.
Yunpeng Cai
BMC Bioinform.3
2018 Multiscaled Fusion of Deep Convolutional Neural Networks for Screening Atrial Fibrillation From Single Lead Short ECG Recordings
abstract
Atrial fibrillation (AF) is one of the most common sustained chronic cardiac arrhythmia in elderly population, associated with a high mortality and morbidity in stroke, heart failure, coronary artery disease, systemic thromboembolism, etc. The early detection of AF is necessary for averting the possibility of disability or mortality. However, AF detection remains problematic due to its episodic pattern. In this paper, a multiscaled fusion of deep convolutional neural network (MS-CNN) is proposed to screen out AF recordings from single lead short electrocardiogram (ECG) recordings. The MS-CNN employs the architecture of two-stream convolutional networks with different filter sizes to capture features of different scales. The experimental results show that the proposed MS-CNN achieves 96.99% of classification accuracy on ECG recordings cropped/padded to 5 s. Especially, the best classification accuracy, 98.13%, is obtained on ECG recordings of 20 s. Compared with artificial neural network, shallow single-stream CNN, and VisualGeometry group network, the MS-CNN can achieve the better classification performance. Meanwhile, visualization of the learned features from the MS-CNN demonstrates its superiority in extracting linear separable ECG features without hand-craft feature engineering. The excellent AF screening performance of the MS-CNN can satisfy the most elders for daily monitoring with wearable devices.
Xiaomao Fan, Qihang Yao, Yunpeng Cai, Fen Miao, Fangmin Sun, Ye Li 0002
IEEE J. Biomed. Health Informatics3
2017 ESPRIT-Forest: Parallel clustering of massive amplicon sequence data in subquadratic time
abstract
The rapid development of sequencing technology has led to an explosive accumulation of genomic sequence data. Clustering is often the first step to perform in sequence analysis, and hierarchical clustering is one of the most commonly used approaches for this purpose. However, it is currently computationally expensive to perform hierarchical clustering of extremely large sequence datasets due to its quadratic time and space complexities. In this paper we developed a new algorithm called ESPRIT-Forest for parallel hierarchical clustering of sequences. The algorithm achieves subquadratic time and space complexity and maintains a high clustering accuracy comparable to the standard method. The basic idea is to organize sequences into a pseudo-metric based partitioning tree for sub-linear time searching of nearest neighbors, and then use a new multiple-pair merging criterion to construct clusters in parallel using multiple threads. The new algorithm was tested on the human microbiome project (HMP) dataset, currently one of the largest published microbial 16S rRNA sequence dataset. Our experiment demonstrated that with the power of parallel computing it is now compu- tationally feasible to perform hierarchical clustering analysis of tens of millions of sequences. The software is available at http://www.acsu.buffalo.edu/∼yijunsun/lab/ESPRIT-Forest.html.
Yunpeng Cai, Wei Zheng 0010, Volker Mai, Qi Mao 0001, Yijun Sun
PLoS Comput. Biol.1
2015 Parallel Hierarchical Clustering in Linearithmic Time for Large-Scale Sequence Analysis
abstract
The rapid development of sequencing technology has led to an explosive accumulation of genomics data. Clustering is often the first step to perform in sequence analysis, and hierarchical clustering is one of the most commonly used approaches for this purpose. However, the standard hierarchical clustering method scales poorly due to its quadratic time and space complexities stemming mainly from the need of computing and storing a pairwise distance matrix. It is thus necessary to minimize the number of pairwise distances computed without degrading clustering performance. On the other hand, as high-performance computing systems are becoming widely accessible, it is highly desirable that a clustering method can be easily adapted to parallel computing environments for further speedup, which is not a trivial task for hierarchical clustering. We proposed a new hierarchical clustering method that achieves good clustering performance and high scalability on large sequence datasets. It consists of two stages. In the first stage, a new landmark-based active hierarchical divisive clustering method was proposed that partitions a large-scale sequence dataset into groups, and in the second stage, a fast hierarchical agglomerative clustering method is applied to each group. By assembling hierarchies from both stages, the hierarchy of the data can be easily recovered. Theoretical results showed that our method can recover the true hierarchy with a high probability under some mild conditions and has a linearithmic time complexity with respect to the number of input sequences. The proposed method also facilitates an efficient parallel implementation. Empirical results on various datasets showed that our method achieved clustering accuracy comparable to ESPRIT-Tree and ran faster than greedy heuristic methods.
Qi Mao 0001, Wei Zheng 0010, Li Wang 0033, Yunpeng Cai, Volker Mai, Yijun Sun
ICDM4
2014 Learning Sparse Gaussian Bayesian Network Structure by Variable Grouping
abstract
Bayesian networks (BNs) are popular for modeling conditional distributions of variables and causal relationships, especially in biological settings such as protein interactions, gene regulatory networks and microbial interactions. Previous BN structure learning algorithms treat variables with similar tendency separately. In this paper, we propose a grouped sparse Gaussian BN (GSGBN) structure learning algorithm which creates BN based on three assumptions: (i) variables follow a multivariate Gaussian distribution, (ii) the network only contains a few edges (sparse), (iii) similar variables have less-divergent sets of parents, while not-so-similar ones should have divergent sets of parents (variable grouping). We use L1regularization to make the learned network sparse, and another term to incorporate shared information among variables. For similar variables, GSGBN tends to penalize the differences of similar variables' parent sets more, compared to those not-so-similar variables' parent sets. The similarity of variables is learned from the data by alternating optimization, without prior domain knowledge. Based on this new definition of the optimal BN, a coordinate descent algorithm and a projected gradient descent algorithm are developed to obtain edges of the network and also similarity of variables. Experimental results on both simulated and real datasets show that GSGBN has substantially superior prediction performance for structure learning when compared to several existing algorithms.
Henry C. M. Leung, Siu-Ming Yiu, Yunpeng Cai, Francis Y. L. Chin
ICDM4
2013 An ontology-based approach for text mining of stroke electronic medical records
abstract
In this paper, we propose a novel ontology-based approach for text mining of EMR information retrieval. The advantage of this approach is that it is capable of handling numerous variations in nature text which essentially refer to the same identity, as well as inferring implicit information from the plain text, which are both important in data mining of medical records. We applied the approach to text mining of EMR documents for stroke patients in a Chinese medical hospital. A benchmark study on an independent test set shows that the proposed pipeline can accurately extract the vast majority of useful information from the EMR documents, including the implicit ones through ontology inference. We also carry out a primary statistical analysis on a sample EMR set to illustrate the utilization of the approach on medical studies.
Yunpeng Cai, Wenshu Luo, Zhenghui Ma, Xiaolu Yu
BIBM2
2013 Intra- and inter-sparse multiple output regression with application on environmental microbial community study
abstract
Feature selection is important for many biological studies, especially when the number of available samples is limited (in order of hundreds) while the number of input features is large (in order of millions), such as eQTL (expression quantitative trait loci) mapping, GWAS (genome wide association study) and environmental microbial community study. We study the problem of multiple output regression which leverages the underlying common relationship shared by multiple output features and propose an efficient and accurate approach for feature selection. Our approach considers both intra- and inter-group sparsities. The intergroup sparsity assumes that only small set of input features are related to the output features. The intragroup sparsity assumes that each input features may relate to multiple output features which should have different kinds of sparsity. Most existing methods do not model the intragroup sparsity well by either assuming uniform regularization on each group, i.e. each input feature relates to similar number of output features, or requiring prior knowledge of the relationship of input and output features. By modelling the regression coefficients as a mixture distributions of Laplacian and Gaussian, we can shrink group regression coefficients to be small adaptively and learn the intergroup, intragroup sparsity and shrinkage estimation patterns. Empirical studies on the synthetic and real environmental microbial community datasets show that our model has better predictions on test dataset than existing methods such as Lasso, Elastic Net, dirty model and rMTFL (robust multi-task feature learning). Moreover, by using least angle regression or coordinate descent and projected gradient descent techniques for optimization, we can obtain the optimal regression efficiently.
Henry C. M. Leung, Siu-Ming Yiu, Yunpeng Cai, Francis Y. L. Chin
BIBM4
2012 HCloud: A novel application-oriented cloud platform for preventive healthcare
abstract
As an emerging state-of-the-art technology, cloud computing has been applied to an extensive range of real life situations. Healthcare is one of such important application fields. We developed a healthcare system, named HCloud, after comprehensive analysis of requirements of healthcare, which based on cloud platform with characteristics of loose coupling algorithms modules and powerful parallel computing capabilities to compute the detail of these indicators for the purpose of preventive healthcare service. The proposed system can support huge physiological data storage and process heterogeneous data for various health care applications such as automated electrocardiogram (ECG) analysis, providing an early warning mechanism for users with chronic disease. The architecture of the cloud platform for physiological data storage, computing, data mining, and feature selections are described. Performance evaluations based on testing has demonstrated the effectiveness and usability of the system.
Xiaomao Fan, Yunpeng Cai, Ye Li 0002
CloudCom3
2012 A large-scale benchmark study of existing algorithms for taxonomy-independent microbial community analysis
abstract
Recent advances in massively parallel sequencing technology have created new opportunities to probe the hidden world of microbes. Taxonomy-independent clustering of the 16S rRNA gene is usually the first step in analyzing microbial communities. Dozens of algorithms have been developed in the last decade, but a comprehensive benchmark study is lacking. Here, we survey algorithms currently used by microbiologists, and compare seven representative methods in a large-scale benchmark study that addresses several issues of concern. A new experimental protocol was developed that allows different algorithms to be compared using the same platform, and several criteria were introduced to facilitate a quantitative evaluation of the clustering performance of each algorithm. We found that existing methods vary widely in their outputs, and that inappropriate use of distance levels for taxonomic assignments likely resulted in substantial overestimates of biodiversity in many studies. The benchmark study identified our recently developed ESPRIT-Tree, a fast implementation of the average linkage-based hierarchical clustering algorithm, as one of the best algorithms available in terms of computational efficiency and clustering accuracy.
Yijun Sun, Yunpeng Cai, Susan M. Huse, Rob Knight 0001, William G. Farmerie, Volker Mai
Briefings Bioinform.2
2010 Effective structure learning for EDA via L1-regularizedbayesian networks
abstract
The Bayesian optimization algorithm (BOA) uses Bayesian networks to explore the dependencies between decision variables of an optimization problem in pursuit of both faster speed of convergence and better solution quality. In this paper, a novel method that learns the structure of Bayesian networks for BOA is proposed. The proposed method, called L1BOA, uses L1-regularized regression to find the candidate parents of each variable, which leads to a sparse but nearly optimized network structure. The proposed method improves the efficiency of the structure learning in BOA due to the reduction and automated control of network complexity introduced with L1-regularized learning. Experimental studies on different types of benchmark problems are carried out, which show that L1BOA outperforms the standard BOA when no a-priori knowledge about the problem structure is available, and nearly achieves the best performance of BOA that applies explicit complexity controls.
Jiadong Yang, Yunpeng Cai, Peifa Jia
GECCO3
2010 Fast Implementation of ℓ1Regularized Learning Algorithms Using Gradient Descent Methods
abstract
With the advent of high-throughput technologies, ℓ1 regularized learning algorithms have attracted much attention recently. Dozens of algorithms have been proposed for fast implementation, using various advanced optimization techniques. In this paper, we demonstrate that ℓ1 regularized learning problems can be easily solved by using gradient-descent techniques. The basic idea is to transform a convex optimization problem with a non-differentiable objective function into an unconstrained non-convex problem, upon which, via gradient descent, reaching a globally optimum solution is guaranteed. We present detailed implementation of the algorithm using ℓ1 regularized logistic regression as a particular application. We conduct large-scale experiments to compare the new approach with other state-of-the-art algorithms on eight medium and large-scale problems. We demonstrate that our algorithm, though simple, performs similarly or even better than other advanced algorithms in terms of computational efficiency and memory usage.
Yunpeng Cai, Yijun Sun, Yubo Cheng, Jian Li 0001, Steve Goodison
SDM1
2009 Online Feature Selection Algorithm with Bayesian l1 Regularization
Yunpeng Cai, Yijun Sun, Jian Li 0001, Steve Goodison
PAKDD1
2009 Duple-EDA and sample density balancing
Yunpeng Cai, Hua Xu 0003, Xiaomin Sun 0001, Peifa Jia, ZeHua Liu
Sci. China Ser. F Inf. Sci.1
2008 Combining nomogram and microarray data for predicting prostate cancer recurrence
abstract
The derivation of molecular signatures indicative of disease status and behavior are required to facilitate the optimal choice of treatment for prostate cancer patients. We conducted a computational analysis of gene expression profile data obtained from 79 cases, 39 of which were classified as having disease recurrence, to investigate whether an advanced computational algorithm can derive more accurate prognostic signatures for prostate cancer. At the 90% sensitivity level, a newly derived genetic signature achieved 85% specificity. This is the first reported genetic signature to outperform a clinically used postoperative nomogram. Furthermore, a hybrid signature derived by combination of the nomogram and gene expression data significantly outperformed both genetic and clinical signatures, and achieved a specificity of 95%. Our study demonstrates the possibility of utilizing both genetic and clinical information for highly accurate prostate cancer prognosis beyond the current clinical systems, and shows that more advanced computational modeling of microarray and clinical data is warranted before clinical application of predictive signatures is considered.
Yijun Sun, Yunpeng Cai, Steve Goodison
BIBE2
2008 Semi-supervised feature selection under logistic I-RELIEF framework
abstract
We consider feature selection in the semi-supervised learning setting. This problem is rarely addressed in the literature. We propose a new algorithm as a natural extension of the recently developed Logistic I-RELIEF algorithm. The basic idea of the proposed algorithm is to modify the objective function of Logistic I-RELIEF to include the margins of unlabeled samples by following the large margin principle. Experimental results on artificial and benchmark datasets are presented to demonstrate the viability of the newly proposed method.
Yubo Cheng, Yunpeng Cai, Yijun Sun, Jian Li 0001
ICPR2
2007 Cross entropy and adaptive variance scaling in continuous EDA
abstract
This paper deals with the adaptive variance scaling issue incontinuous Estimation of Distribution Algorithms. A phenomenon is discovered that current adaptive variance scaling method in EDA suffers from imprecise structure learning. A new type of adaptation method is proposed to overcome this defect. The method tries to measure the difference between the obtained population and the prediction of the probabilistic model, then calculate the scaling factor by minimizing the cross entropy between these two distributions. This approach calculates the scaling factor immediately rather than adapts it incrementally. Experiments show that this approach extended the class of problems that can be solved, and improve the search efficiency in some cases. Moreover, the proposed approach features in that each decomposed subspace can be assigned an individual scaling factor, which helps to solve problems with special dimension property.
Yunpeng Cai, Xiaomin Sun 0001, Hua Xu 0003, Peifa Jia
GECCO1
2006 Probabilistic modeling for continuous EDA with Boltzmann selection and Kullback-Leibeler divergence
abstract
This paper extends the Boltzmann Selection, a method in EDA with theoretical importance, from discrete domain to the continuous one. The difficulty of estimating the exact Boltzmann distribution in continuous state space is circumvented by adopting the multivariate Gaussian model, which is popular in continuous EDA, to approximate only the final sampling distribution. With the minimum Kullback-Leibeler divergence principle, both the mean vector and the covariance matrix of the Gaussian model can be calibrated to preserve the features of Boltzmann selection reflecting desired selection pressure. A method is proposed to adapt the selection pressure based on measuring the successfulness of the past evolution process. These works established a formal basis that helps to build probabilistic models in continuous EDA algorithms with adaptive parameters. The framework is incorporated in both the continuous UMDA and the EMNA algorithm, and tested in several benchmark problems. The experiment results are compared with some existing EDA versions and the benefit of the proposed approach is discussed.
Yunpeng Cai, Xiaomin Sun 0001, Peifa Jia
GECCO1
2003 Technical Solutions of TsinghuAeolus for Robotic Soccer
Jinyi Yao, Ni Lao, Yunpeng Cai, Zengqi Sun
RoboCup4
2001 Global Planning from Local Eyeshot: An Implementation of Observation-Based Plan Coordination in RoboCup Simulation Games
Yunpeng Cai, Jinyi Yao, Shi Li 0002
RoboCup1
2001 Architecture of TsinghuAeolus
Jinyi Yao, Yunpeng Cai, Shi Li 0002
RoboCup3