EDBT 2026 Demo / reviewers in the wild / expert
Dao-Qing Dai
dblp:31/1717
· DBLP profile ↗
73ranked-venue papers
4as first author
16since 2021 · last 2024
0000-0003-4622-8492ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Maximizing conditional independence for unsupervised domain adaptation
Yiming Zhai, Chuan-Xian Ren, You-Wei Luo, Dao-Qing Dai |
Sci. China Inf. Sci. | 4 |
| 2024 | Cross-site prognosis prediction for nasopharyngeal carcinoma from incomplete multi-modal data
Chuan-Xian Ren, Gengxin Xu, Dao-Qing Dai, Qing-Shan Liu |
Medical Image Anal. | 3 |
| 2023 | BuresNet: Conditional Bures Metric for Transferable Representation LearningabstractAs a fundamental manner for learning and cognition, transfer learning has attracted widespread attention in recent years. Typical transfer learning tasks include unsupervised domain adaptation (UDA) and few-shot learning (FSL), which both attempt to sufficiently transfer discriminative knowledge from the training environment to the test environment to improve the model's generalization performance. Previous transfer learning methods usually ignore the potential conditional distribution shift between environments. This leads to the discriminability degradation in the test environments. Therefore, how to construct a learnable and interpretable metric to measure and then reduce the gap between conditional distributions is very important in the literature. In this article, we design the Conditional Kernel Bures (CKB) metric for characterizing conditional distribution discrepancy, and derive an empirical estimation with convergence guarantee. CKB provides a statistical and interpretable approach, under the optimal transportation framework, to understand the knowledge transfer mechanism. It is essentially an extension of optimal transportation from the marginal distributions to the conditional distributions. CKB can be used as a plug-and-play module and placed onto the loss layer in deep networks, thus, it plays the bottleneck role in representation learning. From this perspective, the new method with network architecture is abbreviated as BuresNet, and it can be used extract conditional invariant features for both UDA and FSL tasks. BuresNet can be trained in an end-to-end manner. Extensive experiment results on several benchmark datasets validate the effectiveness of BuresNet. Chuan-Xian Ren, You-Wei Luo, Dao-Qing Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Conditional Independence Induced Unsupervised Domain Adaptation
Xiao-Lin Xu, Gengxin Xu, Chuan-Xian Ren, Dao-Qing Dai, Hong Yan 0001 |
Pattern Recognit. | 4 |
| 2022 | Subtype-WESLR: identifying cancer subtype with weighted ensemble sparse latent representation of multi-view dataabstractThe discovery of cancer subtypes has become much-researched topic in oncology. Dividing cancer patients into subtypes can provide personalized treatments for heterogeneous patients. High-throughput technologies provide multiple omics data for cancer subtyping. Integration of multi-view data is used to identify cancer subtypes in many computational methods, which obtain different subtypes for the same cancer, even using the same multi-omics data. To a certain extent, these subtypes from distinct methods are related, which may have certain guiding significance for cancer subtyping. It is a challenge to effectively utilize the valuable information of distinct subtypes to produce more accurate and reliable subtypes. A weighted ensemble sparse latent representation (subtype-WESLR) is proposed to detect cancer subtypes on heterogeneous omics data. Using a weighted ensemble strategy to fuse base clustering obtained by distinct methods as prior knowledge, subtype-WESLR projects each sample feature profile from each data type to a common latent subspace while maintaining the local structure of the original sample feature space and consistency with the weighted ensemble and optimizes the common subspace by an iterative method to identify cancer subtypes. We conduct experiments on various synthetic datasets and eight public multi-view datasets from The Cancer Genome Atlas. The results demonstrate that subtype-WESLR is better than competing methods by utilizing the integration of base clustering of exist methods for more precise subtypes. Wenjing Song, Weiwen Wang 0001, Dao-Qing Dai |
Briefings Bioinform. | 3 |
| 2022 | Learning representation for multiple biological networks via a robust graph regularized integration approachabstractLearning node representation is a fundamental problem in biological network analysis, as compact representation features reveal complicated network structures and carry useful information for downstream tasks such as link prediction and node classification. Recently, multiple networks that profile objects from different aspects are increasingly accumulated, providing the opportunity to learn objects from multiple perspectives. However, the complex common and specific information across different networks pose challenges to node representation methods. Moreover, ubiquitous noise in networks calls for more robust representation. To deal with these problems, we present a representation learning method for multiple biological networks. First, we accommodate the noise and spurious edges in networks using denoised diffusion, providing robust connectivity structures for the subsequent representation learning. Then, we introduce a graph regularized integration model to combine refined networks and compute common representation features. By using the regularized decomposition technique, the proposed model can effectively preserve the common structural property of different networks and simultaneously accommodate their specific information, leading to a consistent representation. A simulation study shows the superiority of the proposed method on different levels of noisy networks. Three network-based inference tasks, including drug-target interaction prediction, gene function identification and fine-grained species categorization, are conducted using representation features learned from our method. Biological networks at different scales and levels of sparsity are involved. Experimental results on real-world data show that the proposed method has robust performance compared with alternatives. Overall, by eliminating noise and integrating effectively, the proposed method is able to learn useful representations from multiple biological networks. Weiwen Wang 0001, Chuan-Xian Ren, Dao-Qing Dai |
Briefings Bioinform. | 4 |
| 2022 | springD2A: capturing uncertainty in disease-drug association prediction with model integrationabstractMOTIVATION: Drug repositioning that aims to find new indications for existing drugs has been an efficient strategy for drug discovery. In the scenario where we only have confirmed disease-drug associations as positive pairs, a negative set of disease-drug pairs is usually constructed from the unknown disease-drug pairs in previous studies, where we do not know whether drugs and diseases can be associated, to train a model for disease-drug association prediction (drug repositioning). Drugs and diseases in these negative pairs can potentially be associated, but most studies have ignored them. RESULTS: We present a method, springD2A, to capture the uncertainty in the negative pairs, and to discriminate between positive and unknown pairs because the former are more reliable. In springD2A, we introduce a spring-like penalty for the loss of negative pairs, which is strong if they are too close in a unit sphere, but mild if they are at a moderate distance. We also design a sequential sampling in which the probability of an unknown disease-drug pair sampled as negative is proportional to its score predicted as positive. Multiple models are learned during sequential sampling, and we adopt parameter- and feature-based ensemble schemes to boost performance. Experiments show springD2A is an effective tool for drug-repositioning. AVAILABILITY AND IMPLEMENTATION: A python implementation of springD2A and datasets used in this study are available at https://github.com/wangyuanhao/springD2A. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Weiwen Wang 0001, Dao-Qing Dai |
Bioinform. | 3 |
| 2022 | Unsupervised Domain Adaptation via Discriminative Manifold PropagationabstractUnsupervised domain adaptation is effective in leveraging rich information from a labeled source domain to an unlabeled target domain. Though deep learning and adversarial strategy made a significant breakthrough in the adaptability of features, there are two issues to be further studied. First, hard-assigned pseudo labels on the target domain are arbitrary and error-prone, and direct application of them may destroy the intrinsic data structure. Second, batch-wise training of deep learning limits the characterization of the global structure. In this paper, a Riemannian manifold learning framework is proposed to achieve transferability and discriminability simultaneously. For the first issue, this framework establishes a probabilistic discriminant criterion on the target domain via soft labels. Based on pre-built prototypes, this criterion is extended to a global approximation scheme for the second issue. Manifold metric alignment is adopted to be compatible with the embedding space. The theoretical error bounds of different alignment metrics are derived for constructive guidance. The proposed method can be used to tackle a series of variants of domain adaptation problems, including both vanilla and partial settings. Extensive experiments have been conducted to investigate the method and a comparative study shows the superiority of the discriminative manifold learning framework. You-Wei Luo, Chuan-Xian Ren, Dao-Qing Dai, Hong Yan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Hyperspectral Image Classification via Discriminant Gabor Ensemble FilterabstractFor a broad range of applications, hyperspectral image (HSI) classification is a hot topic in remote sensing, and convolutional neural network (CNN)-based methods are drawing increasing attention. However, to train millions of parameters in CNN requires a large number of labeled training samples, which are difficult to collect. A conventional Gabor filter can effectively extract spatial information with different scales and orientations without training, but it may be missing some important discriminative information. In this article, we propose the Gabor ensemble filter (GEF), a new convolutional filter to extract deep features for HSI with fewer trainable parameters. GEF filters each input channel by some fixed Gabor filters and learnable filters simultaneously, then reduces the dimensions by some learnable 1×1 filters to generate the output channels. The fixed Gabor filters can extract common features with different scales and orientations, while the learnable filters can learn some complementary features that Gabor filters cannot extract. Based on GEF, we design a network architecture for HSI classification, which extracts deep features and can learn from limited training samples. In order to simultaneously learn more discriminative features and an end-to-end system, we propose to introduce the local discriminant structure for cross-entropy loss by combining the triplet hard loss. Results of experiments on three HSI datasets show that the proposed method has significantly higher classification accuracy than other state-of-the-art methods. Moreover, the proposed method is speedy for both training and testing. Ke-Kun Huang, Chuan-Xian Ren, Zhao-Rong Lai, Yu-Feng Yu 0001, Dao-Qing Dai |
IEEE Trans. Cybern. | 6 |
| 2021 | MetaCon: Meta Contrastive Learning for Microsatellite Instability Detection
Weiwen Wang 0001, Chuan-Xian Ren, Dao-Qing Dai |
MICCAI (8) | 4 |
| 2021 | DeFusion: a denoised network regularization framework for multi-omics integrationabstractWith diverse types of omics data widely available, many computational methods have been recently developed to integrate these heterogeneous data, providing a comprehensive understanding of diseases and biological mechanisms. But most of them hardly take noise effects into account. Data-specific patterns unique to data types also make it challenging to uncover the consistent patterns and learn a compact representation of multi-omics data. Here we present a multi-omics integration method considering these issues. We explicitly model the error term in data reconstruction and simultaneously consider noise effects and data-specific patterns. We utilize a denoised network regularization in which we build a fused network using a denoising procedure to suppress noise effects and data-specific patterns. The error term collaborates with the denoised network regularization to capture data-specific patterns. We solve the optimization problem via an inexact alternating minimization algorithm. A comparative simulation study shows the method's superiority at discovering common patterns among data types at three noise levels. Transcriptomics-and-epigenomics integration, in seven cancer cohorts from The Cancer Genome Atlas, demonstrates that the learned integrative representation extracted in an unsupervised manner can depict survival information. Specially in liver hepatocellular carcinoma, the learned integrative representation attains average Harrell's C-index of 0.78 in 10 times 3-fold cross-validation for survival prediction, which far exceeds competing methods, and we discover an aggressive subtype in liver hepatocellular carcinoma with this latent representation, which is validated by an external dataset GSE14520. We also show that DeFusion is applicable to the integration of other omics types. Weiwen Wang 0001, Dao-Qing Dai |
Briefings Bioinform. | 3 |
| 2021 | Hyperspectral image classification via discriminative convolutional neural network with an improved triplet loss
Ke-Kun Huang, Chuan-Xian Ren, Zhao-Rong Lai, Yu-Feng Yu 0001, Dao-Qing Dai |
Pattern Recognit. | 6 |
| 2021 | WMLRR: A Weighted Multi-View Low Rank Representation to Identify Cancer Subtypes From Multiple Types of Omics DataabstractThe identification of cancer subtypes is of great importance for understanding the heterogeneity of tumors and providing patients with more accurate diagnoses and treatments. However, it is still a challenge to effectively integrate multiple omics data to establish cancer subtypes. In this paper, we propose an unsupervised integration method, named weighted multi-view low rank representation (WMLRR), to identify cancer subtypes from multiple types of omics data. Given a group of patients described by multiple omics data matrices, we first learn a unified affinity matrix which encodes the similarities among patients by exploring the sparsity-consistent low-rank representations from the joint decompositions of multiple omics data matrices. Unlike existing subtype identification methods that treat each omics data matrix equally, we assign a weight to each omics data matrix and learn these weights automatically through the optimization process. Finally, we apply spectral clustering on the learned affinity matrix to identify cancer subtypes. Experiment results show that the survival times between our identified cancer subtypes are significantly different, and our predicted survivals are more accurate than other state-of-the-art methods. In addition, some clinical analyses of the diseases also demonstrate the effectiveness of our method in identifying molecular subtypes with biological significance and clinical relevance. Yesen Sun, Le Ou-Yang, Dao-Qing Dai |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Heterogeneous Domain Adaptation via Covariance Structured Feature TranslatorsabstractDomain adaptation (DA) and transfer learning with statistical property description is very important in image analysis and data classification. This article studies the domain adaptive feature representation problem for the heterogeneous data, of which both the feature dimensions and the sample distributions across domains are so different that their features cannot be matched directly. To transfer the discriminant information efficiently from the source domain to the target domain, and then enhance the classification performance for the target data, we first introduce two projection matrices specified for different domains to transform the heterogeneous features into a shared space. We then propose a joint kernel regression model to learn the regression variable, which is called feature translator in this article. The novelty focuses on the exploration of optimal experimental design (OED) to deal with the heterogeneous and nonlinear DA by seeking the covariance structured feature translators (CSFTs). An approximate and efficient method is proposed to compute the optimal data projections. Comprehensive experiments are conducted to validate the effectiveness and efficacy of the proposed model. The results show the state-of-the-art performance of our method in heterogeneous DA. Chuan-Xian Ren, Jiashi Feng, Dao-Qing Dai, Shuicheng Yan |
IEEE Trans. Cybern. | 3 |
| 2021 | Learning Kernel for Conditional Moment-Matching Discrepancy-Based Image ClassificationabstractConditional maximum mean discrepancy (CMMD) can capture the discrepancy between conditional distributions by drawing support from nonlinear kernel functions; thus, it has been successfully used for pattern classification. However, CMMD does not work well on complex distributions, especially when the kernel function fails to correctly characterize the difference between intraclass similarity and interclass similarity. In this paper, a new kernel learning method is proposed to improve the discrimination performance of CMMD. It can be operated with deep network features iteratively and thus denoted as KLN for abbreviation. The CMMD loss and an autoencoder (AE) are used to learn an injective function. By considering the compound kernel, that is, the injective function with a characteristic kernel, the effectiveness of CMMD for data category description is enhanced. KLN can simultaneously learn a more expressive kernel and label prediction distribution; thus, it can be used to improve the classification performance in both supervised and semisupervised learning scenarios. In particular, the kernel-based similarities are iteratively learned on the deep network features, and the algorithm can be implemented in an end-to-end manner. Extensive experiments are conducted on four benchmark datasets, including MNIST, SVHN, CIFAR-10, and CIFAR-100. The results indicate that KLN achieves the state-of-the-art classification performance. Chuan-Xian Ren, Pengfei Ge, Dao-Qing Dai, Hong Yan 0001 |
IEEE Trans. Cybern. | 3 |
| 2021 | Joint Transformation Learning via the L2, 1-Norm Metric for Robust Graph MatchingabstractEstablishing correspondence between two given geometrical graph structures is an important problem in computer vision and pattern recognition. In this paper, we propose a robust graph matching (RGM) model to improve the effectiveness and robustness on the matching graphs with deformations, rotations, outliers, and noise. First, we embed the joint geometric transformation into the graph matching model, which performs unary matching over graph nodes and local structure matching over graph edges simultaneously. Then, the L2,1-norm is used as the similarity metric in the presented RGM to enhance the robustness. Finally, we derive an objective function which can be solved by an effective optimization algorithm, and theoretically prove the convergence of the proposed algorithm. Extensive experiments on various graph matching tasks, such as outliers, rotations, and deformations show that the proposed RGM model achieves competitive performance compared to the existing methods. Yu-Feng Yu 0001, Guoxia Xu, Min Jiang 0003, Hu Zhu, Dao-Qing Dai, Hong Yan 0001 |
IEEE Trans. Cybern. | 5 |
| 2020 | Discriminative Residual Analysis for Image Set Classification With Posture and Age VariationsabstractImage set recognition has been widely applied in many practical problems like real-time video retrieval and image caption tasks. Due to its superior performance, it has grown into a significant topic in recent years. However, images with complicated variations, e.g., postures and human ages, are difficult to address, as these variations are continuous and gradual with respect to image appearance. Consequently, the crucial point of image set recognition is to mine the intrinsic connection or structural information from the image batches with variations. In this work, a Discriminant Residual Analysis (DRA) method is proposed to improve the classification performance by discovering discriminant features in related and unrelated groups. Specifically, DRA attempts to obtain a powerful projection which casts the residual representations into a discriminant subspace. Such a projection subspace is expected to magnify the useful information of the input space as much as possible, then the relation between the training set and the test set described by the given metric or distance will be more precise in the discriminant subspace. We also propose a nonfeasance strategy by defining another approach to construct the unrelated groups, which help to reduce furthermore the cost of sampling errors. Two regularization approaches are used to deal with the probable small sample size problem. Extensive experiments are conducted on benchmark databases, and the results show superiority and efficiency of the new methods. Chuan-Xian Ren, You-Wei Luo, Xiao-Lin Xu, Dao-Qing Dai, Hong Yan 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Dual Adversarial Autoencoders for ClusteringabstractAs a powerful approach for exploratory data analysis, unsupervised clustering is a fundamental task in computer vision and pattern recognition. Many clustering algorithms have been developed, but most of them perform unsatisfactorily on the data with complex structures. Recently, adversarial autoencoder (AE) (AAE) shows effectiveness on tackling such data by combining AE and adversarial training, but it cannot effectively extract classification information from the unlabeled data. In this brief, we propose dual AAE (Dual-AAE) which simultaneously maximizes the likelihood function and mutual information between observed examples and a subset of latent variables. By performing variational inference on the objective function of Dual-AAE, we derive a new reconstruction loss which can be optimized by training a pair of AEs. Moreover, to avoid mode collapse, we introduce the clustering regularization term for the category variable. Experiments on four benchmarks show that Dual-AAE achieves superior performance over state-of-the-art clustering methods. In addition, by adding a reject option, the clustering accuracy of Dual-AAE can reach that of supervised CNN algorithms. Dual-AAE can also be used for disentangling style and content of images without using supervised information. Pengfei Ge, Chuan-Xian Ren, Dao-Qing Dai, Jiashi Feng, Shuicheng Yan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Sparse approximation to discriminant projection learning and application to image classification
Yu-Feng Yu 0001, Chuan-Xian Ren, Min Jiang 0003, Man-Yu Sun, Dao-Qing Dai, Guodong Guo |
Pattern Recognit. | 5 |
| 2018 | Kernel Embedding Multiorientation Local Pattern for Image RepresentationabstractLocal feature descriptor plays a key role in different image classification applications. Some of these methods such as local binary pattern and image gradient orientations have been proven effective to some extent. However, such traditional descriptors which only utilize single-type features, are deficient to capture the edges and orientations information and intrinsic structure information of images. In this paper, we propose a kernel embedding multiorientation local pattern (MOLP) to address this problem. For a given image, it is first transformed by gradient operators in local regions, which generate multiorientation gradient images containing edges and orientations information of different directions. Then the histogram feature which takes into account the sign component and magnitude component, is extracted to form the refined feature from each orientation gradient image. The refined feature captures more information of the intrinsic structure, and is effective for image representation and classification. Finally, the multiorientation refined features are automatically fused in the kernel embedding discriminant subspace learning model. The extensive experiments on various image classification tasks, such as face recognition, texture classification, object categorization, and palmprint recognition show that MOLP could achieve competitive performance with those state-of-the art methods. Yu-Feng Yu 0001, Chuan-Xian Ren, Dao-Qing Dai, Ke-Kun Huang |
IEEE Trans. Cybern. | 3 |
| 2018 | A Peak Price Tracking-Based Learning System for Portfolio SelectionabstractWe propose a novel linear learning system based on the peak price tracking (PPT) strategy for portfolio selection (PS). Recently, the topic of tracking control attracts intensive attention and some novel models are proposed based on backstepping methods, such that the system output tracks a desired trajectory. The proposed system has a similar evolution with a transform function that aggressively tracks the increasing power of different assets. As a result, the better performing assets will receive more investment. The proposed PPT objective can be formulated as a fast backpropagation algorithm, which is suitable for large-scale and time-limited applications, such as high-frequency trading. Extensive experiments on several benchmark data sets from diverse real financial markets show that PPT outperforms other state-of-the-art systems in computational time, cumulative wealth, and risk-adjusted metrics. It suggests that PPT is effective and even more robust than some defensive systems in PS. Zhao-Rong Lai, Dao-Qing Dai, Chuan-Xian Ren, Ke-Kun Huang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Radial Basis Functions With Adaptive Input and Composite Trend Representation for Portfolio SelectionabstractWe propose a set of novel radial basis functions with adaptive input and composite trend representation (AICTR) for portfolio selection (PS). Trend representation of asset price is one of the main information to be exploited in PS. However, most state-of-the-art trend representation-based systems exploit only one kind of trend information and lack effective mechanisms to construct a composite trend representation. The proposed system exploits a set of RBFs with multiple trend representations, which improves the effectiveness and robustness in price prediction. Moreover, the input of the RBFs automatically switches to the best trend representation according to the recent investing performance of different price predictions. We also propose a novel objective to combine these RBFs and select the portfolio. Extensive experiments on six benchmark data sets (including a new challenging data set that we propose) from different real-world stock markets indicate that the proposed RBFs effectively combine different trend representations and AICTR achieves state-of-the-art investing performance and risk control. Besides, AICTR withstands the reasonable transaction costs and runs fast; hence, it is applicable to real-world financial environments. Zhao-Rong Lai, Dao-Qing Dai, Chuan-Xian Ren, Ke-Kun Huang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Regularized coplanar discriminant analysis for dimensionality reduction
Ke-Kun Huang, Dao-Qing Dai, Chuan-Xian Ren |
Pattern Recognit. | 2 |
| 2017 | Fusing landmark-based features at kernel level for face recognition
Ke-Kun Huang, Dao-Qing Dai, Chuan-Xian Ren, Yu-Feng Yu 0001, Zhao-Rong Lai |
Pattern Recognit. | 2 |
| 2017 | Discriminative multi-scale sparse coding for single-sample face recognition with occlusion
Yu-Feng Yu 0001, Dao-Qing Dai, Chuan-Xian Ren, Ke-Kun Huang |
Pattern Recognit. | 2 |
| 2017 | Discriminative multi-layer illumination-robust feature extraction for face recognition
Yu-Feng Yu 0001, Dao-Qing Dai, Chuan-Xian Ren, Ke-Kun Huang |
Pattern Recognit. | 2 |
| 2017 | Learning Kernel Extended Dictionary for Face RecognitionabstractA sparse representation classifier (SRC) and a kernel discriminant analysis (KDA) are two successful methods for face recognition. An SRC is good at dealing with occlusion, while a KDA does well in suppressing intraclass variations. In this paper, we propose kernel extended dictionary (KED) for face recognition, which provides an efficient way for combining KDA and SRC. We first learn several kernel principal components of occlusion variations as an occlusion model, which can represent the possible occlusion variations efficiently. Then, the occlusion model is projected by KDA to get the KED, which can be computed via the same kernel trick as new testing samples. Finally, we use structured SRC for classification, which is fast as only a small number of atoms are appended to the basic dictionary, and the feature dimension is low. We also extend KED to multikernel space to fuse different types of features at kernel level. Experiments are done on several large-scale data sets, demonstrating that not only does KED get impressive results for nonoccluded samples, but it also handles the occlusion well without overfitting, even with a single gallery sample per subject. Ke-Kun Huang, Dao-Qing Dai, Chuan-Xian Ren, Zhao-Rong Lai |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | A two-layer integration framework for protein complex detectionabstractBACKGROUND: Protein complexes carry out nearly all signaling and functional processes within cells. The study of protein complexes is an effective strategy to analyze cellular functions and biological processes. With the increasing availability of proteomics data, various computational methods have recently been developed to predict protein complexes. However, different computational methods are based on their own assumptions and designed to work on different data sources, and various biological screening methods have their unique experiment conditions, and are often different in scale and noise level. Therefore, a single computational method on a specific data source is generally not able to generate comprehensive and reliable prediction results. RESULTS: In this paper, we develop a novel Two-layer INtegrative Complex Detection (TINCD) model to detect protein complexes, leveraging the information from both clustering results and raw data sources. In particular, we first integrate various clustering results to construct consensus matrices for proteins to measure their overall co-complex propensity. Second, we combine these consensus matrices with the co-complex score matrix derived from Tandem Affinity Purification/Mass Spectrometry (TAP) data and obtain an integrated co-complex similarity network via an unsupervised metric fusion method. Finally, a novel graph regularized doubly stochastic matrix decomposition model is proposed to detect overlapping protein complexes from the integrated similarity network. CONCLUSIONS: Extensive experimental results demonstrate that TINCD performs much better than 21 state-of-the-art complex detection techniques, including ensemble clustering and data integration techniques. Le Ou-Yang, Min Wu 0008, Xiao-Fei Zhang, Dao-Qing Dai, Xiaoli Li 0001, Hong Yan 0001 |
BMC Bioinform. | 4 |
| 2016 | Protein complex detection based on partially shared multi-view clusteringabstractBACKGROUND: Protein complexes are the key molecular entities to perform many essential biological functions. In recent years, high-throughput experimental techniques have generated a large amount of protein interaction data. As a consequence, computational analysis of such data for protein complex detection has received increased attention in the literature. However, most existing works focus on predicting protein complexes from a single type of data, either physical interaction data or co-complex interaction data. These two types of data provide compatible and complementary information, so it is necessary to integrate them to discover the underlying structures and obtain better performance in complex detection. RESULTS: In this study, we propose a novel multi-view clustering algorithm, called the Partially Shared Multi-View Clustering model (PSMVC), to carry out such an integrated analysis. Unlike traditional multi-view learning algorithms that focus on mining either consistent or complementary information embedded in the multi-view data, PSMVC can jointly explore the shared and specific information inherent in different views. In our experiments, we compare the complexes detected by PSMVC from single data source with those detected from multiple data sources. We observe that jointly analyzing multi-view data benefits the detection of protein complexes. Furthermore, extensive experiment results demonstrate that PSMVC performs much better than 16 state-of-the-art complex detection techniques, including ensemble clustering and data integration techniques. CONCLUSIONS: In this work, we demonstrate that when integrating multiple data sources, using partially shared multi-view clustering model can help to identify protein complexes which are not readily identifiable by conventional single-view-based methods and other integrative analysis methods. All the results and source codes are available on https://github.com/Oyl-CityU/PSMVC . Le Ou-Yang, Xiao-Fei Zhang, Dao-Qing Dai, Meng-Yun Wu, Yuan Zhu 0005, Hong Yan 0001 |
BMC Bioinform. | 3 |
| 2016 | Regularized logistic regression with network-based pairwise interaction for biomarker identification in breast cancerabstractBACKGROUND: To facilitate advances in personalized medicine, it is important to detect predictive, stable and interpretable biomarkers related with different clinical characteristics. These clinical characteristics may be heterogeneous with respect to underlying interactions between genes. Usually, traditional methods just focus on detection of differentially expressed genes without taking the interactions between genes into account. Moreover, due to the typical low reproducibility of the selected biomarkers, it is difficult to give a clear biological interpretation for a specific disease. Therefore, it is necessary to design a robust biomarker identification method that can predict disease-associated interactions with high reproducibility. RESULTS: In this article, we propose a regularized logistic regression model. Different from previous methods which focus on individual genes or modules, our model takes gene pairs, which are connected in a protein-protein interaction network, into account. A line graph is constructed to represent the adjacencies between pairwise interactions. Based on this line graph, we incorporate the degree information in the model via an adaptive elastic net, which makes our model less dependent on the expression data. Experimental results on six publicly available breast cancer datasets show that our method can not only achieve competitive performance in classification, but also retain great stability in variable selection. Therefore, our model is able to identify the diagnostic and prognostic biomarkers in a more robust way. Moreover, most of the biomarkers discovered by our model have been verified in biochemical or biomedical researches. CONCLUSIONS: The proposed method shows promise in the diagnosis of disease pathogenesis with different clinical characteristics. These advances lead to more accurate and stable biomarker discovery, which can monitor the functional changes that are perturbed by diseases. Based on these predictions, researchers may be able to provide suggestions for new therapeutic approaches. Meng-Yun Wu, Xiao-Fei Zhang, Dao-Qing Dai, Le Ou-Yang, Yuan Zhu 0005, Hong Yan 0001 |
BMC Bioinform. | 3 |
| 2016 | Comparative analysis of housekeeping and tissue-specific driver nodes in human protein interaction networksabstractBACKGROUND: Several recent studies have used the Minimum Dominating Set (MDS) model to identify driver nodes, which provide the control of the underlying networks, in protein interaction networks. There may exist multiple MDS configurations in a given network, thus it is difficult to determine which one represents the real set of driver nodes. Because these previous studies only focus on static networks and ignore the contextual information on particular tissues, their findings could be insufficient or even be misleading. RESULTS: In this study, we develop a Collective-Influence-corrected Minimum Dominating Set (CI-MDS) model which takes into account the collective influence of proteins. By integrating molecular expression profiles and static protein interactions, 16 tissue-specific networks are established as well. We then apply the CI-MDS model to each tissue-specific network to detect MDS proteins. It generates almost the same MDSs when it is solved using different optimization algorithms. In addition, we classify MDS proteins into Tissue-Specific MDS (TS-MDS) proteins and HouseKeeping MDS (HK-MDS) proteins based on the number of tissues in which they are expressed and identified as MDS proteins. Notably, we find that TS-MDS proteins and HK-MDS proteins have significantly different topological and functional properties. HK-MDS proteins are more central in protein interaction networks, associated with more functions, evolving more slowly and subjected to a greater number of post-translational modifications than TS-MDS proteins. Unlike TS-MDS proteins, HK-MDS proteins significantly correspond to essential genes, ageing genes, virus-targeted proteins, transcription factors and protein kinases. Moreover, we find that besides HK-MDS proteins, many TS-MDS proteins are also linked to disease related genes, suggesting the tissue specificity of human diseases. Furthermore, functional enrichment analysis reveals that HK-MDS proteins carry out universally necessary biological processes and TS-MDS proteins usually involve in tissue-dependent functions. CONCLUSIONS: Our study uncovers key features of TS-MDS proteins and HK-MDS proteins, and is a step forward towards a better understanding of the controllability of human interactomes. Xiao-Fei Zhang, Le Ou-Yang, Dao-Qing Dai, Meng-Yun Wu, Yuan Zhu 0005, Hong Yan 0001 |
BMC Bioinform. | 3 |
| 2016 | Enhanced Local Gradient Order Features and Discriminant Analysis for Face RecognitionabstractRobust descriptor-based subspace learning with complex data is an active topic in pattern analysis and machine intelligence. A few researches concentrate the optimal design on feature representation and metric learning. However, traditionally used features of single-type, e.g., image gradient orientations (IGOs), are deficient to characterize the complete variations in robust and discriminant subspace learning. Meanwhile, discontinuity in edge alignment and feature match are not been carefully treated in the literature. In this paper, local order constrained IGOs are exploited to generate robust features. As the difference-based filters explicitly consider the local contrasts within neighboring pixel points, the proposed features enhance the local textures and the order-based coding ability, thus discover intrinsic structure of facial images further. The multimodal features are automatically fused in the most discriminant subspace. The utilization of adaptive interaction function suppresses outliers in each dimension for robust similarity measurement and discriminant analysis. The sparsity-driven regression model is modified to adapt the classification issue of the compact feature representation. Extensive experiments are conducted by using some benchmark face data sets, e.g., of controlled and uncontrolled environments, to evaluate our new algorithm. Chuan-Xian Ren, Zhen Lei 0001, Dao-Qing Dai, Stan Z. Li |
IEEE Trans. Cybern. | 3 |
| 2015 | Determining minimum set of driver nodes in protein-protein interaction networksabstractBACKGROUND: Recently, several studies have drawn attention to the determination of a minimum set of driver proteins that are important for the control of the underlying protein-protein interaction (PPI) networks. In general, the minimum dominating set (MDS) model is widely adopted. However, because the MDS model does not generate a unique MDS configuration, multiple different MDSs would be generated when using different optimization algorithms. Therefore, among these MDSs, it is difficult to find out the one that represents the true driver set of proteins. RESULTS: To address this problem, we develop a centrality-corrected minimum dominating set (CC-MDS) model which includes heterogeneity in degree and betweenness centralities of proteins. Both the MDS model and the CC-MDS model are applied on three human PPI networks. Unlike the MDS model, the CC-MDS model generates almost the same sets of driver proteins when we implement it using different optimization algorithms. The CC-MDS model targets more high-degree and high-betweenness proteins than the uncorrected counterpart. The more central position allows CC-MDS proteins to be more important in maintaining the overall network connectivity than MDS proteins. To indicate the functional significance, we find that CC-MDS proteins are involved in, on average, more protein complexes and GO annotations than MDS proteins. We also find that more essential genes, aging genes, disease-associated genes and virus-targeted genes appear in CC-MDS proteins than in MDS proteins. As for the involvement in regulatory functions, the sets of CC-MDS proteins show much stronger enrichment of transcription factors and protein kinases. The results about topological and functional significance demonstrate that the CC-MDS model can capture more driver proteins than the MDS model. CONCLUSIONS: Based on the results obtained, the CC-MDS model presents to be a powerful tool for the determination of driver proteins that can control the underlying PPI networks. The software described in this paper and the datasets used are available at https://github.com/Zhangxf-ccnu/CC-MDS . Xiao-Fei Zhang, Le Ou-Yang, Yuan Zhu 0005, Meng-Yun Wu, Dao-Qing Dai |
BMC Bioinform. | 5 |
| 2015 | Detecting Protein Complexes from Signed Protein-Protein Interaction NetworksabstractIdentification of protein complexes is fundamental for understanding the cellular functional organization. With the accumulation of physical protein-protein interaction (PPI) data, computational detection of protein complexes from available PPI networks has drawn a lot of attentions. While most of the existing protein complex detection algorithms focus on analyzing the physical protein-protein interaction network, none of them take into account the "signs" (i.e., activation-inhibition relationships) of physical interactions. As the "signs" of interactions reflect the way proteins communicate, considering the "signs" of interactions can not only increase the accuracy of protein complex identification, but also deepen our understanding of the mechanisms of cell functions. In this study, we proposed a novel Signed Graph regularized Nonnegative Matrix Factorization (SGNMF) model to identify protein complexes from signed PPI networks. In our experiments, we compared the results collected by our model on signed PPI networks with those predicted by the state-of-the-art complex detection techniques on the original unsigned PPI networks. We observed that considering the "signs" of interactions significantly benefits the detection of protein complexes. Furthermore, based on the predicted complexes, we predicted a set of signed complex-complex interactions for each dataset, which provides a novel insight of the higher level organization of the cell. All the experimental results and codes can be downloaded from http://mail.sysu.edu.cn/home/[email protected]/dai/others/SGNMF.zip. Le Ou-Yang, Dao-Qing Dai, Xiao-Fei Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2015 | Discriminative and Compact Coding for Robust Face RecognitionabstractIn this paper, we propose a novel discriminative and compact coding (DCC) for robust face recognition. It introduces multiple error measurements into regression model. They collaborate to tune regression codes of different properties (sparsity, compactness, high discriminating ability, etc.), to further improve robustness and adaptivity of the regression model. We propose two types of coding models: 1) multiscale error measurements that produces sparse and highly discriminative codes and 2) inspires within-class collaborative representation that produces sparse and compact codes. The update of codes and the combination of different errors are automatically processed. DCC is also robust to the choice of parameters, producing stable regression residuals which are crucial to classification. Extensive experiments on benchmark datasets show that DCC has promising performance and outperforms other state-of-the-art regression models. Zhao-Rong Lai, Dao-Qing Dai, Chuan-Xian Ren, Ke-Kun Huang |
IEEE Trans. Cybern. | 2 |
| 2015 | Multiscale Logarithm Difference Edgemaps for Face Recognition Against Varying Lighting ConditionsabstractLambertian model is a classical illumination model consisting of a surface albedo component and a light intensity component. Some previous researches assume that the light intensity component mainly lies in the large-scale features. They adopt holistic image decompositions to separate it out, but it is difficult to decide the separating point between large-scale and small-scale features. In this paper, we propose to take a logarithm transform, which can change the multiplication of surface albedo and light intensity into an additive model. Then, a difference (substraction) between two pixels in a neighborhood can eliminate most of the light intensity component. By dividing a neighborhood into subregions, edgemaps of multiple scales can be obtained. Then, each edgemap is multiplied by a weight that can be determined by an independent training scheme. Finally, all the weighted edgemaps are combined to form a robust holistic feature map. Extensive experiments on four benchmark data sets in controlled and uncontrolled lighting conditions show that the proposed method has promising results, especially in uncontrolled lighting conditions, even mixed with other complicated variations. Zhao-Rong Lai, Dao-Qing Dai, Chuan-Xian Ren, Ke-Kun Huang |
IEEE Trans. Image Process. | 2 |
| 2015 | Sample Weighting: An Inherent Approach for Outlier Suppressing Discriminant AnalysisabstractAs the data acquirement technologies develop rapidly, both the amount and types of data become larger and larger. However, noise and outliers usually attach to the data and then affect the real performance of leaning algorithms in data mining and pattern analysis. To address this problem, the importance of the sample itself in building the optimal subspace is explored, and then an importance-sampling-inspired method is proposed for outlier suppressing feature extraction. First, we assign each sample a weight, which is estimated by graph Laplacian, and then calculate the approximated mean for each subject. By highlighting the most subject-oriented samples, the weighted average and the scatter metrics can be measured with maximum margins and superior classification performance. The supervised information integrates local data structure with respective contributions to building the optimal subspace. The linear criterion can be extended to a nonlinear case by the kernel trick. A regularization framework is proposed to deal with the rank-deficient problem, which is usually induced by the small sample size of training set. Competitive performance of our algorithm has been validated by extensive experiments performed on the synthetic and benchmark data, including facial images and gene micro-array data. Chuan-Xian Ren, Dao-Qing Dai, Xiaofei He 0001, Hong Yan 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | Detecting temporal protein complexes from dynamic protein-protein interaction networksabstractBACKGROUND: Proteins dynamically interact with each other to perform their biological functions. The dynamic operations of protein interaction networks (PPI) are also reflected in the dynamic formations of protein complexes. Existing protein complex detection algorithms usually overlook the inherent temporal nature of protein interactions within PPI networks. Systematically analyzing the temporal protein complexes can not only improve the accuracy of protein complex detection, but also strengthen our biological knowledge on the dynamic protein assembly processes for cellular organization. RESULTS: In this study, we propose a novel computational method to predict temporal protein complexes. Particularly, we first construct a series of dynamic PPI networks by joint analysis of time-course gene expression data and protein interaction data. Then a Time Smooth Overlapping Complex Detection model (TS-OCD) has been proposed to detect temporal protein complexes from these dynamic PPI networks. TS-OCD can naturally capture the smoothness of networks between consecutive time points and detect overlapping protein complexes at each time point. Finally, a nonnegative matrix factorization based algorithm is introduced to merge those very similar temporal complexes across different time points. CONCLUSIONS: Extensive experimental results demonstrate the proposed method is very effective in detecting temporal protein complexes than the state-of-the-art complex detection techniques. Le Ou-Yang, Dao-Qing Dai, Xiaoli Li 0001, Min Wu 0008, Xiao-Fei Zhang |
BMC Bioinform. | 2 |
| 2014 | Detecting overlapping protein complexes based on a generative model with functional and topological propertiesabstractBACKGROUND: Identification of protein complexes can help us get a better understanding of cellular mechanism. With the increasing availability of large-scale protein-protein interaction (PPI) data, numerous computational approaches have been proposed to detect complexes from the PPI networks. However, most of the current approaches do not consider overlaps among complexes or functional annotation information of individual proteins. Therefore, they might not be able to reflect the biological reality faithfully or make full use of the available domain-specific knowledge. RESULTS: In this paper, we develop a Generative Model with Functional and Topological Properties (GMFTP) to describe the generative processes of the PPI network and the functional profile. The model provides a working mechanism for capturing the interaction structures and the functional patterns of proteins. By combining the functional and topological properties, we formulate the problem of identifying protein complexes as that of detecting a group of proteins which frequently interact with each other in the PPI network and have similar annotation patterns in the functional profile. Using the idea of link communities, our method naturally deals with overlaps among complexes. The benefits brought by the functional properties are demonstrated by real data analysis. The results evaluated using four criteria with respect to two gold standards show that GMFTP has a competitive performance over the state-of-the-art approaches. The effectiveness of detecting overlapping complexes is also demonstrated by analyzing the topological and functional features of multi- and mono-group proteins. CONCLUSIONS: Based on the results obtained in this study, GMFTP presents to be a powerful approach for the identification of overlapping protein complexes using both the PPI network and the functional profile. The software can be downloaded from http://mail.sysu.edu.cn/home/[email protected]/dai/others/GMFTP.zip. Xiao-Fei Zhang, Dao-Qing Dai, Le Ou-Yang, Hong Yan 0001 |
BMC Bioinform. | 2 |
| 2014 | Multilayer Surface Albedo for Face Recognition With Reference Images in Bad Lighting ConditionsabstractIn this paper, we propose a multilayer surface albedo (MLSA) model to tackle face recognition in bad lighting conditions, especially with reference images in bad lighting conditions. Some previous researches conclude that illumination variations mainly lie in the large-scale features of an image and extract small-scale features in the surface albedo (or surface texture). However, this surface albedo is not robust enough, which still contains some detrimental sharp features. To improve robustness of the surface albedo, MLSA further decomposes it as a linear sum of several detailed layers, to separate and represent features of different scales in a more specific way. Then, the layers are adjusted by separate weights, which are global parameters and selected for only once. A criterion function is developed to select these layer weights with an independent training set. Despite controlled illumination variations, MLSA is also effective to uncontrolled illumination variations, even mixed with other complicated variations (expression, pose, occlusion, and so on). Extensive experiments on four benchmark data sets show that MLSA has good receiver operating characteristic curve and statistical discriminating capability. The refined albedo improves recognition performance, especially with reference images in bad lighting conditions. Zhao-Rong Lai, Dao-Qing Dai, Chuan-Xian Ren, Ke-Kun Huang |
IEEE Trans. Image Process. | 2 |
| 2014 | Transfer Learning of Structured Representation for Face RecognitionabstractFace recognition under uncontrolled conditions, e.g., complex backgrounds and variable resolutions, is still challenging in image processing and computer vision. Although many methods have been proved well-performed in the controlled settings, they are usually of weak generality across different data sets. Meanwhile, several properties of the source domain, such as background and the size of subjects, play an important role in determining the final classification results. A transferrable representation learning model is proposed in this paper to enhance the recognition performance. To deeply exploit the discriminant information from the source domain and the target domain, the bioinspired face representation is modeled as structured and approximately stable characterization for the commonality between different domains. The method outputs a grouped boost of the features, and presents a reasonable manner for highlighting and sharing discriminant orientations and scales. Notice that the method can be viewed as a framework, since other feature generation operators and classification metrics can be embedded therein, and then, it can be applied to more general problems, such as low-resolution face recognition, object detection and categorization, and so forth. Experiments on the benchmark databases, including uncontrolled Face Recognition Grand Challenge v2.0 and Labeled Faces in the Wild show the efficacy of the proposed transfer learning algorithm. Chuan-Xian Ren, Dao-Qing Dai, Ke-Kun Huang, Zhao-Rong Lai |
IEEE Trans. Image Process. | 2 |
| 2014 | Band-Reweighed Gabor Kernel Embedding for Face Image Representation and RecognitionabstractFace recognition with illumination or pose variation is a challenging problem in image processing and pattern recognition. A novel algorithm using band-reweighed Gabor kernel embedding to deal with the problem is proposed in this paper. For a given image, it is first transformed by a group of Gabor filters, which output Gabor features using different orientation and scale parameters. Fisher scoring function is used to measure the importance of features in each band, and then, the features with the largest scores are preserved for saving memory requirements. The reduced bands are combined by a vector, which is determined by a weighted kernel discriminant criterion and solved by a constrained quadratic programming method, and then, the weighted sum of these nonlinear bands is defined as the similarity between two images. Compared with existing concatenation-based Gabor feature representation and the uniformly weighted similarity calculation approaches, our method provides a new way to use Gabor features for face recognition and presents a reasonable interpretation for highlighting discriminant orientations and scales. The minimum Mahalanobis distance considering the spatial correlations within the data is exploited for feature matching, and the graphical lasso is used therein for directly estimating the sparse inverse covariance matrix. Experiments using benchmark databases show that our new algorithm improves the recognition results and obtains competitive performance. Chuan-Xian Ren, Dao-Qing Dai, Xiaoxin Li 0001, Zhao-Rong Lai |
IEEE Trans. Image Process. | 2 |
| 2013 | Identifying Spurious Interactions and Predicting Missing Interactions in the Protein-Protein Interaction Networks via a Generative Network ModelabstractWith the rapid development of high-throughput experiment techniques for protein-protein interaction (PPI) detection, a large amount of PPI network data are becoming available. However, the data produced by these techniques have high levels of spurious and missing interactions. This study assigns a new reliably indication for each protein pairs via the new generative network model (RIGNM) where the scale-free property of the PPI network is considered to reliably identify both spurious and missing interactions in the observed high-throughput PPI network. The experimental results show that the RIGNM is more effective and interpretable than the compared methods, which demonstrate that this approach has the potential to better describe the PPI networks and drive new discoveries. Xiao-Fei Zhang, Dao-Qing Dai, Meng-Yun Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2013 | Identification of DNA-Binding and Protein-Binding Proteins Using Enhanced Graph Wavelet FeaturesabstractInteractions between biomolecules play an essential role in various biological processes. For predicting DNA-binding or protein-binding proteins, many machine-learning-based techniques have used various types of features to represent the interface of the complexes, but they only deal with the properties of a single atom in the interface and do not take into account the information of neighborhood atoms directly. This paper proposes a new feature representation method for biomolecular interfaces based on the theory of graph wavelet. The enhanced graph wavelet features (EGWF) provides an effective way to characterize interface feature through adding physicochemical features and exploiting a graph wavelet formulation. Particularly, graph wavelet condenses the information around the center atom, and thus enhances the discrimination of features of biomolecule binding proteins in the feature space. Experiment results show that EGWF performs effectively for predicting DNA-binding and protein-binding proteins in terms of Matthew's correlation coefficient (MCC) score and the area value under the receiver operating characteristic curve (AUC). Yuan Zhu 0005, Weiqiang Zhou, Dao-Qing Dai, Hong Yan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2013 | Structured Sparse Error Coding for Face Recognition With OcclusionabstractFace recognition with occlusion is common in the real world. Inspired by the works of structured sparse representation, we try to explore the structure of the error incurred by occlusion from two aspects: the error morphology and the error distribution. Since human beings recognize the occlusion mainly according to its region shape or profile without knowing accurately what the occlusion is, we argue that the shape of the occlusion is also an important feature. We propose a morphological graph model to describe the morphological structure of the error. Due to the uncertainty of the occlusion, the distribution of the error incurred by occlusion is also uncertain. However, we observe that the unoccluded part and the occluded part of the error measured by the correntropy induced metric follow the exponential distribution, respectively. Incorporating the two aspects of the error structure, we propose the structured sparse error coding for face recognition with occlusion. Our extensive experiments demonstrate that the proposed method is more stable and has higher breakdown point in dealing with the occlusion problems in face recognition as compared to the related state-of-the-art methods, especially for the extreme situation, such as the high level occlusion and the low feature dimension. Xiaoxin Li 0001, Dao-Qing Dai, Xiao-Fei Zhang, Chuan-Xian Ren |
IEEE Trans. Image Process. | 2 |
| 2012 | Reflectance Estimation Using Local Regression Methods
Weifeng Zhang 0006, Peng Yang 0006, Dao-Qing Dai, Arye Nehorai |
ISNN (1) | 3 |
| 2012 | Robust classification using ℓ2, 1-norm based regression model
Chuan-Xian Ren, Dao-Qing Dai, Hong Yan 0001 |
Pattern Recognit. | 2 |
| 2012 | Biomarker Identification and Cancer Classification Based on Microarray Data Using Laplace Naive Bayes Model with Mean ShrinkageabstractBiomarker identification and cancer classification are two closely related problems. In gene expression data sets, the correlation between genes can be high when they share the same biological pathway. Moreover, the gene expression data sets may contain outliers due to either chemical or electrical reasons. A good gene selection method should take group effects into account and be robust to outliers. In this paper, we propose a Laplace naive Bayes model with mean shrinkage (LNB-MS). The Laplace distribution instead of the normal distribution is used as the conditional distribution of the samples for the reasons that it is less sensitive to outliers and has been applied in many fields. The key technique is the L1 penalty imposed on the mean of each class to achieve automatic feature selection. The objective function of the proposed model is a piecewise linear function with respect to the mean of each class, of which the optimal value can be evaluated at the breakpoints simply. An efficient algorithm is designed to estimate the parameters in the model. A new strategy that uses the number of selected features to control the regularization parameter is introduced. Experimental results on simulated data sets and 17 publicly available cancer data sets attest to the accuracy, sparsity, efficiency, and robustness of the proposed algorithm. Many biomarkers identified with our method have been verified in biochemical or biomedical research. The analysis of biological and functional correlation of the genes based on Gene Ontology (GO) terms shows that the proposed method guarantees the selection of highly correlated genes simultaneously Meng-Yun Wu, Dao-Qing Dai, Hong Yan 0001, Xiao-Fei Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | A Framework for Incorporating Functional Interrelationships into Protein Function Prediction AlgorithmsabstractThe functional annotation of proteins is one of the most important tasks in the post-genomic era. Although many computational approaches have been developed in recent years to predict protein function, most of these traditional algorithms do not take interrelationships among functional terms into account, such as different GO terms usually coannotate with some common proteins. In this study, we propose a new functional similarity measure in the form of Jaccard coefficient to quantify these interrelationships and also develop a framework for incorporating GO term similarity into protein function prediction process. The experimental results of cross-validation on S. cerevisiae and Homo sapiens data sets demonstrate that our method is able to improve the performance of protein function prediction. In addition, we find that small size terms associated with a few of proteins obtain more benefit than the large size ones when considering functional interrelationships. We also compare our similarity measure with other two widely used measures, and results indicate that when incorporated into function prediction algorithms, our proposed measure is more effective. Experiment results also illustrate that our algorithms outperform two previous competing algorithms, which also take functional interrelationships into account, in prediction accuracy. Finally, we show that our method is robust to annotations in the database which are not complete at present. These results give new insights about the importance of functional interrelationships in protein function prediction. Xiao-Fei Zhang, Dao-Qing Dai |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | Protein Complexes Discovery Based on Protein-Protein Interaction Data via a Regularized Sparse Generative Network ModelabstractDetecting protein complexes from protein interaction networks is one major task in the postgenome era. Previous developed computational algorithms identifying complexes mainly focus on graph partition or dense region finding. Most of these traditional algorithms cannot discover overlapping complexes which really exist in the protein-protein interaction (PPI) networks. Even if some density-based methods have been developed to identify overlapping complexes, they are not able to discover complexes that include peripheral proteins. In this study, motivated by recent successful application of generative network model to describe the generation process of PPI networks and to detect communities from social networks, we develop a regularized sparse generative network model (RSGNM), by adding another process that generates propensities using exponential distribution and incorporating Laplacian regularizer into an existing generative network model, for protein complexes identification. By assuming that the propensities are generated using exponential distribution, the estimators of propensities will be sparse, which not only has good biological interpretation but also helps to control the overlapping rate among detected complexes. And the Laplacian regularizer will lead to the estimators of propensities more smooth on interaction networks. Experimental results on three yeast PPI networks show that RSGNM outperforms six previous competing algorithms in terms of the quality of detected complexes. In addition, RSGNM is able to detect overlapping complexes and complexes including peripheral proteins simultaneously. These results give new insights about the importance of generative network models in protein complexes identification. Xiao-Fei Zhang, Dao-Qing Dai, Xiaoxin Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | A New On-Board Image Codec Based on Binary Tree With Adaptive Scanning Order in Scan-Based ModeabstractRemote sensing images offer a large amount of information but require on-board compression because of the storage and transmission constraints of on-board equipment. JPEG2000 is too complex to become a recommended standard for the mission, and CCSDS-IDC fixes most of the parameters and only provides quality scalability. In this paper, we present a new, low-complexity, low-memory, and efficient embedded wavelet image codec for on-board compression. First, we propose the binary tree as a novel and robust way of coding remote sensing image in wavelet domain. Second, we develop an adaptive scanning order to traverse the binary tree level by level from the bottom to the top, so that better performance and visual effect are attained. Last, the proposed method is processed with a scan-based mode, which significantly reduces the memory requirement. The proposed method is very fast because it does not use any entropy coding and rate-distortion optimization, while it provides quality, position, and resolution scalability. Being less complex, it is very easy to implement in hardware and very suitable for on-board compression. Experimental results show that the proposed method can significantly improve peak signal-to-noise ratio compared with SPIHT without arithmetic coding and scan-based CCSDS-IDC, and is similar to scan-based JPEG2000. Ke-Kun Huang, Dao-Qing Dai |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2012 | Coupled Kernel Embedding for Low-Resolution Face Image RecognitionabstractPractical video scene and face recognition systems are sometimes confronted with low-resolution (LR) images. The faces may be very small even if the video is clear, thus it is difficult to directly measure the similarity between the faces and the high-resolution (HR) training samples. Traditional super-resolution (SR) methods based face recognition usually have limited performance because the target of SR may not be consistent with that of classification, and time-consuming SR algorithms are not suitable for real-time applications. In this paper, a new feature extraction method called Coupled Kernel Embedding (CKE) is proposed for LR face recognition without any SR preprocessing. In this method, the final kernel matrix is constructed by concatenating two individual kernel matrices in the diagonal direction, and the (semi-)positively definite properties are preserved for optimization. CKE addresses the problem of comparing multi-modal data that are difficult for conventional methods in practice due to the lack of an efficient similarity measure. Particularly, different kernel types (e.g., linear, Gaussian, polynomial) can be integrated into an uniformed optimization objective, which cannot be achieved by simple linear methods. CKE solves this problem by minimizing the dissimilarities captured by their kernel Gram matrices in the low- and high-resolution spaces. In the implementation, the nonlinear objective function is minimized by a generalized eigenvalue decomposition. Experiments on benchmark and real databases show that our CKE method indeed improves the recognition performance. Chuan-Xian Ren, Dao-Qing Dai, Hong Yan 0001 |
IEEE Trans. Image Process. | 2 |
| 2011 | Framelet Algorithms for De-Blurring Images Corrupted by Impulse Plus Gaussian NoiseabstractThis paper studies a problem of image restoration that observed images are contaminated by Gaussian and impulse noise. Existing methods for this problem in the literature are based on minimizing an objective functional having the l(1) fidelity term and the Mumford-Shah regularizer. We present an algorithm on this problem by minimizing a new objective functional. The proposed functional has a content-dependent fidelity term which assimilates the strength of fidelity terms measured by the l(1) and l(2) norms. The regularizer in the functional is formed by the l(1) norm of tight framelet coefficients of the underlying image. The selected tight framelet filters are able to extract geometric features of images. We then propose an iterative framelet-based approximation/sparsity deblurring algorithm (IFASDA) for the proposed functional. Parameters in IFASDA are adaptively varying at each iteration and are determined automatically. In this sense, IFASDA is a parameter-free algorithm. This advantage makes the algorithm more attractive and practical. The effectiveness of IFASDA is experimentally illustrated on problems of image deblurring with Gaussian and impulse noise. Improvements in both PSNR and visual quality of IFASDA over a typical existing method are demonstrated. In addition, Fast_IFASDA, an accelerated algorithm of IFASDA, is also developed. Yan-Ran Li 0001, Lixin Shen, Dao-Qing Dai, Bruce W. Suter |
IEEE Trans. Image Process. | 3 |
| 2011 | Finding Correlated Biclusters from Gene Expression DataabstractExtracting biologically relevant information from DNA microarrays is a very important task for drug development and test, function annotation, and cancer diagnosis. Various clustering methods have been proposed for the analysis of gene expression data, but when analyzing the large and heterogeneous collections of gene expression data, conventional clustering algorithms often cannot produce a satisfactory solution. Biclustering algorithm has been presented as an alternative approach to standard clustering techniques to identify local structures from gene expression data set. These patterns may provide clues about the main biological processes associated with different physiological states. In this paper, different from existing bicluster patterns, we first introduce a more general pattern: correlated bicluster, which has intuitive biological interpretation. Then, we propose a novel transform technique based on singular value decomposition so that identifying correlated-bicluster problem from gene expression matrix is transformed into two global clustering problems. The Mixed-Clustering algorithm and the Lift algorithm are devised to efficiently produce δ-corBiclusters. The biclusters obtained using our method from gene expression data sets of multiple human organs and the yeast Saccharomyces cerevisiae demonstrate clear biological meanings. Wen-Hui Yang, Dao-Qing Dai, Hong Yan 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2010 | The theoretic framework of local weighted approximation for microarray missing value estimation
Chao-Chun Liu, Dao-Qing Dai, Hong Yan 0001 |
Pattern Recognit. | 2 |
| 2010 | Incremental learning of bidirectional principal components for face recognition
Chuan-Xian Ren, Dao-Qing Dai |
Pattern Recognit. | 2 |
| 2010 | Bilinear Lanczos components for fast dimensionality reduction and feature extraction
Chuan-Xian Ren, Dao-Qing Dai |
Pattern Recognit. | 2 |
| 2010 | Multiframe Super-Resolution Reconstruction Using Sparse Directional RegularizationabstractWe present a variational approach to obtain high-resolution images from multiframe low-resolution video stills. The objective functional for the variational approach consists of a data fidelity term and a regularizer. The fidelity term is formed by adaptively mimickingl1andl2norms. The regularization uses thel1norm of the framelet coefficients of a high-resolution image with a geometric tight framelet system constructed in this paper. The tight framelet system has abilities to detect multi-orientation and multi-order variations of an image. A two-phase iterative method for super-resolution reconstruction is proposed to construct a high-resolution image. The first phase is to get an approximation of the solution (i.e., the ideal image) using the steepest descent method. The second phase is to enhance the sparsity of the approximate solution by using the soft thresholding operator with variable thresholding parameters. Numerical results based on both synthetic data and real videos show that our algorithm is efficient in terms of removing visual artifacts and preserving edges in restored images. Yan-Ran Li 0001, Dao-Qing Dai, Lixin Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Framelet Kernels With Applications to Support Vector Regression and Regularization NetworksabstractSupport vector regression and regularization networks are kernel-based techniques for solving the regression problem of recovering the unknown function from sample data. The choice of the kernel function, which determines the mapping between the input space and the feature space, is of crucial importance to such learning machines. Estimating the irregular function with a multiscale structure that comprises both the steep variations and the smooth variations is a hard problem. The result achieved by the traditional Gaussian kernel is often unsatisfactory, because it cannot simultaneously avoid underfitting and overfitting. In this paper, we present a new class of kernel functions derived from the framelet system. A framelet is a tight wavelet frame constructed via multiresolution analysis and has the merit of both wavelets and frames. The construction and approximation properties of framelets have been well studied. Our goal is to combine the power of framelet representation with the merit of kernel methods on learning from sparse data. The proposed framelet kernel has the ability to approximate functions with a multiscale structure and can reduce the influence of noise in data. Experiments on both simulated and real data illustrate the usefulness of the new kernels. Weifeng Zhang 0006, Dao-Qing Dai, Hong Yan 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2009 | Face Recognition Using Dual-Tree Complex Wavelet FeaturesabstractWe propose a novel facial representation based on the dual-tree complex wavelet transform for face recognition. It is effective and efficient to represent the geometrical structures in facial image with low redundancy. Moreover, we experimentally verify that the proposed method is more powerful to extract facial features robust against the variations of shift and illumination than the discrete wavelet transform and Gabor wavelet transform. Chao-Chun Liu, Dao-Qing Dai |
IEEE Trans. Image Process. | 2 |
| 2009 | Two-Dimensional Maximum Margin Feature Extraction for Face RecognitionabstractOn face recognition, most previous works on dimensionality reduction and classification would first transform the input image into 1-D vector, which ignores the underlying data structure and often leads to the small sample size problem. More recently, 2-D discriminant analysis has become an interesting technique which can overcome the aforementioned drawbacks. However, 2-D methods extract features based on the rows or the columns of all images, so it is possible that the features using 2-D methods still contain some redundant information. In addition, most existing 2-D methods cannot provide an automatic strategy to choose discriminant vectors. In this paper, we study the combination of 2-D discriminant analysis and 1-D discriminant analysis and propose a two-stage framework: " (2D)(2)MMC + LDA." Because the extracted features based on maximal margin criterion (MMC) is robust, stable, and efficient, in the first stage, a 2-D two-directional feature extraction technique, (2D)(2)MMC , is presented. In the second stage, the linear discriminant analysis (LDA) step is performed in the (2D)(2)MMC subspace. Experiments with Feret, Olivetti and Oracle Research Laboratory, and Carnegie Mellon University Pose, Illumination, and Expression databases are conducted to evaluate our method in terms of classification accuracy and robustness. Wen-Hui Yang, Dao-Qing Dai |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2008 | Feature Extraction and Uncorrelated Discriminant Analysis for High-Dimensional DataabstractHigh-dimensional data and the small sample size problem occur in many modern pattern classification applications such as face recognition and gene expression data analysis. To deal with such data, one important step is dimensionality reduction. Principal component analysis (PCA) and between-group analysis (BGA) are two commonly used methods, and various extensions of these two methods exist. The principle of these two approaches comes from their best approximation property. From a pattern recognition perspective, we show that PCA, which is based on the total scatter matrix, preserves linear separability, and BGA, which is based on between-class scatter matrix, retains the distance between class centroids. Moreover, we propose an automatic nonparameter uncorrelated discriminant analysis (UDA) algorithm based on the maximum margin criterion (MMC). The extracted features via UDA are statistically uncorrelated. UDA combines rank-preserving dimensionality reduction and constraint discriminant analysis and also serves as an effective solution for the small-sample-size problem. Experiments with face images and gene expression data sets are conducted to evaluate UDA in terms of classification accuracy and robustness. Wen-Hui Yang, Dao-Qing Dai, Hong Yan 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2007 | A Multi-scale Dynamically Growing Hierarchical Self-organizing Map for Brain MRI Image Segmentation
Jingdan Zhang, Dao-Qing Dai |
ISNN (2) | 2 |
| 2007 | On a New Class of Framelet Kernels for Support Vector Regression and Regularization Networks
Weifeng Zhang 0006, Dao-Qing Dai, Hong Yan 0001 |
PAKDD | 2 |
| 2007 | Local Discriminant Wavelet Packet Coordinates for Face Recognition
Chao-Chun Liu, Dao-Qing Dai, Hong Yan 0001 |
J. Mach. Learn. Res. | 2 |
| 2007 | Improved discriminate analysis for high-dimensional data and its application to face recognition
Xiaosheng Zhuang, Dao-Qing Dai |
Pattern Recognit. | 2 |
| 2007 | Face Recognition by Regularized Discriminant AnalysisabstractWhen the feature dimension is larger than the number of samples the small sample-size problem occurs. There is great concern about it within the face recognition community. We point out that optimizing the Fisher index in linear discriminant analysis does not necessarily give the best performance for a face recognition system. We propose a new regularization scheme. The proposed method is evaluated using the Olivetti Research Laboratory database, the Yale database, and the Feret database. Dao-Qing Dai, Pong C. Yuen |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2005 | Inverse Fisher discriminate criteria for small sample size problem and its application to face recognition
Xiaosheng Zhuang, Dao-Qing Dai |
Pattern Recognit. | 2 |
| 2005 | Kernel machine-based one-parameter regularized Fisher discriminant method for face recognitionabstractThis paper addresses two problems in linear discriminant analysis (LDA) of face recognition. The first one is the problem of recognition of human faces under pose and illumination variations. It is well known that the distribution of face images with different pose, illumination, and face expression is complex and nonlinear. The traditional linear methods, such as LDA, will not give a satisfactory performance. The second problem is the small sample size (S3) problem. This problem occurs when the number of training samples is smaller than the dimensionality of feature vector. In turn, the within-class scatter matrix will become singular. To overcome these limitations, this paper proposes a new kernel machine-based one-parameter regularized Fisher discriminant (K1PRFD) technique. K1PRFD is developed based on our previously developed one-parameter regularized discriminant analysis method and the well-known kernel approach. Therefore, K1PRFD consists of two parameters, namely the regularization parameter and kernel parameter. This paper further proposes a new method to determine the optimal kernel parameter in RBF kernel and regularized parameter in within-class scatter matrix simultaneously based on the conjugate gradient method. Three databases, namely FERET, Yale Group B, and CMU PIE, are selected for evaluation. The results are encouraging. Comparing with the existing LDA-based methods, the proposed method gives superior results. Pong C. Yuen, Jian Huang 0009, Dao-Qing Dai |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2003 | Regularized discriminant analysis and its application to face recognition
Dao-Qing Dai, Pong C. Yuen |
Pattern Recognit. | 1 |
| 2000 | Face Recognition Based on Local Fisher Features
Dao-Qing Dai, Guo-Can Feng, Jian-Huang Lai, Pong C. Yuen |
ICMI | 1 |
| 1998 | Human face image retrieval system for large databaseabstractAddresses the speed problem in a human face image retrieval system from a large database. A novel method based on the wavelet transform and principal component analysis (PCA) is developed and presented. The computational load of the proposed method is greatly reduced compared with the original PCA based method. Moreover, the accuracy of the proposed method is improved. Pong C. Yuen, Guo-Can Feng, Dao-Qing Dai |
ICPR | 3 |
| 1998 | Polynomial preserving algorithm for digital image interpolation
Dao-Qing Dai, Tsi-Min Shih, Foo-Tim Chau |
Signal Process. | 1 |