EDBT 2026 Demo / reviewers in the wild / expert
Minghua Deng
dblp:33/5820
· DBLP profile ↗
71ranked-venue papers
3as first author
44since 2021 · last 2026
0000-0002-9143-1898ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 41 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 21 · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 17 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Out-of-distribution generalization enhances protein function annotation for low-homology sequencesabstractUnderstanding protein functions in biological processes is pivotal for disease elucidation and drug discovery. Despite notable progress, existing approaches primarily focus on function transfer under in-distribution (ID) settings, where training and test proteins exhibit high sequence similarity. As a result, their performance often degrades when applied to novel, diverse, and low-homology protein sequences, posing a major challenge for out-of-distribution (OOD) generalization encountered in practice. Towards this end, we develop ProteinScore, a graph transformer approach tailored to improve protein function prediction in OOD settings. ProteinScore integrates a label-invariant variational subgraph generator with self-supervised contrastive learning, thereby identifying meaning substructures within proteins. By highlighting informative features while filtering out redundant ones, ProteinScore improves generalization to diverse and low-homology sequences. Experiments on datasets with both experimentally resolved and AlphaFold2-predicted structures demonstrate that ProteinScore consistently outperforms strong baselines and provides biologically meaningful interpretability through accurately identifying binding sites. In addition, ProteinScore generalizes effectively to two additional downstream tasks, drug-target interaction classification and subcellular localization prediction, achieving superior predictive performance and reliable interpretability. Yiwei Fu, Jiaxiao Chen, Haoyu Lin, Zhonghui Gu, Qingqing Long, Xiao Luo 0001, Minghua Deng |
Briefings Bioinform. | 8 |
| 2025 | Gradient-Based Sample Selection for Black-Box Universal Domain AdaptationabstractUniversal domain adaptation (UniDA) transfers knowledge from a labelled source domain to an unlabelled target domain under domain-shift and category-shift for annotation. In reality, due to privacy protection or other limits, not only source data but also pre-trained models on it may be unavailable when training on target data. In this paper, we go a step further to explore the black-box universal domain adaptation (B^2-UniDA) problem. It requires tackling the labelling task under shifts by only accessing the interface of pre-trained source models. To this end, we introduce GSS which proposes a novel sample selection criterion based on gradient descent and Bayes' Theorem to identify samples of potential unknown classes. This criterion doesn't require manually-set thresholds depending on data used and is suitable for various datasets. GSS builds an open-set classifier and enables it to estimate probabilities of belonging to each class including the unknown category and adjust estimates adaptively. To overcome class-imbalance, especially imbalance between the unknown and known classes, we propose a balancing mechanism by measuring training status and estimating DA type. In addition to distilling knowledge from source model outputs, we focus on mining the categorical structure of target domain by self-training. Experiments on benchmarks show the state-of-the-art performance of GSS compared to typical methods, including source models or source data dependent methods. Qiuyan He, Minghua Deng |
AAAI | 2 |
| 2025 | Semi-supervised contrastive learning variational autoencoder Integrating single-cell multimodal mosaic datasetsabstractAs single-cell sequencing technology became widely used, scientists found that single-modality data alone could not fully meet the research needs of complex biological systems. To address this issue, researchers began simultaneously collect multi-modal single-cell omics data. But different sequencing technologies often result in datasets where one or more data modalities are missing. Therefore, mosaic datasets are more common when we analyze. However, the high dimensionality and sparsity of the data increase the difficulty, and the presence of batch effects poses an additional challenge. To address these challenges, we proposes a flexible integration framework based on Variational Autoencoder called scGCM. The main task of scGCM is to integrate single-cell multimodal mosaic data and eliminate batch effects. This method was conducted on multiple datasets, encompassing different modalities of single-cell data. The results demonstrate that, compared to state-of-the-art multimodal data integration methods, scGCM offers significant advantages in clustering accuracy and data consistency. The source code of scGCM can be accessed at https://github.com/closmouz/scCGM . Minghua Deng |
BMC Bioinform. | 3 |
| 2024 | Distribution-Independent Cell Type Identification for Single-Cell RNA-seq Data
Yuyao Zhai, Minghua Deng |
IJCAI | 3 |
| 2024 | scPLAN: a hierarchical computational framework for single transcriptomics data annotation, integration and cell-type label refinementabstractMOTIVATION: In the past decade, single-cell RNA sequencing (scRNA-seq) has emerged as a pivotal method for transcriptomic profiling in biomedical research. Precise cell-type identification is crucial for subsequent analysis of single-cell data. And the integration and refinement of annotated data are essential for building comprehensive databases. However, prevailing annotation techniques often overlook the hierarchical organization of cell types, resulting in inconsistent annotations. Meanwhile, most existing integration approaches fail to integrate datasets with different annotation depths and none of them can enhance the labels of outdated data with lower annotation resolutions using more intricately annotated datasets or novel biological findings. RESULTS: Here, we introduce scPLAN, a hierarchical computational framework designed for scRNA-seq data analysis. scPLAN excels in annotating unlabeled scRNA-seq data using a reference dataset structured along a hierarchical cell-type tree. It identifies potential novel cell types in a systematic, layer-by-layer manner. Additionally, scPLAN effectively integrates annotated scRNA-seq datasets with varying levels of annotation depth, ensuring consistent refinement of cell-type labels across datasets with lower resolutions. Through extensive annotation and novel cell detection experiments, scPLAN has demonstrated its efficacy. Two case studies have been conducted to showcase how scPLAN integrates datasets with diverse cell-type label resolutions and refine their cell-type labels. AVAILABILITY: https://github.com/michaelGuo1204/scPLAN. Qirui Guo, Musu Yuan, Minghua Deng |
Briefings Bioinform. | 4 |
| 2024 | Continually adapting pre-trained language model to universal annotation of single-cell RNA-seq dataabstractMOTIVATION: Cell-type annotation of single-cell RNA-sequencing (scRNA-seq) data is a hallmark of biomedical research and clinical application. Current annotation tools usually assume the simultaneous acquisition of well-annotated data, but without the ability to expand knowledge from new data. Yet, such tools are inconsistent with the continuous emergence of scRNA-seq data, calling for a continuous cell-type annotation model. In addition, by their powerful ability of information integration and model interpretability, transformer-based pre-trained language models have led to breakthroughs in single-cell biology research. Therefore, the systematic combining of continual learning and pre-trained language models for cell-type annotation tasks is inevitable. RESULTS: We herein propose a universal cell-type annotation tool, called CANAL, that continuously fine-tunes a pre-trained language model trained on a large amount of unlabeled scRNA-seq data, as new well-labeled data emerges. CANAL essentially alleviates the dilemma of catastrophic forgetting, both in terms of model inputs and outputs. For model inputs, we introduce an experience replay schema that repeatedly reviews previous vital examples in current training stages. This is achieved through a dynamic example bank with a fixed buffer size. The example bank is class-balanced and proficient in retaining cell-type-specific information, particularly facilitating the consolidation of patterns associated with rare cell types. For model outputs, we utilize representation knowledge distillation to regularize the divergence between previous and current models, resulting in the preservation of knowledge learned from past training stages. Moreover, our universal annotation framework considers the inclusion of new cell types throughout the fine-tuning and testing stages. We can continuously expand the cell-type annotation library by absorbing new cell types from newly arrived, well-annotated training datasets, as well as automatically identify novel cells in unlabeled datasets. Comprehensive experiments with data streams under various biological scenarios demonstrate the versatility and high model interpretability of CANAL. AVAILABILITY: An implementation of CANAL is available from https://github.com/aster-ww/CANAL-torch. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Journal Name online. Musu Yuan, Yiwei Fu, Minghua Deng |
Briefings Bioinform. | 4 |
| 2024 | SPANN: annotating single-cell resolution spatial transcriptome data with scRNA-seq dataabstractMOTIVATION: The rapid development of spatial transcriptome technologies has enabled researchers to acquire single-cell-level spatial data at an affordable price. However, computational analysis tools, such as annotation tools, tailored for these data are still lacking. Recently, many computational frameworks have emerged to integrate single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics datasets. While some frameworks can utilize well-annotated scRNA-seq data to annotate spatial expression patterns, they overlook critical aspects. First, existing tools do not explicitly consider cell type mapping when aligning the two modalities. Second, current frameworks lack the capability to detect novel cells, which remains a key interest for biologists. RESULTS: To address these problems, we propose an annotation method for spatial transcriptome data called SPANN. The main tasks of SPANN are to transfer cell-type labels from well-annotated scRNA-seq data to newly generated single-cell resolution spatial transcriptome data and discover novel cells from spatial data. The major innovations of SPANN come from two aspects: SPANN automatically detects novel cells from unseen cell types while maintaining high annotation accuracy over known cell types. SPANN finds a mapping between spatial transcriptome samples and RNA data prototypes and thus conducts cell-type-level alignment. Comprehensive experiments using datasets from various spatial platforms demonstrate SPANN's capabilities in annotating known cell types and discovering novel cell states within complex tissue contexts. AVAILABILITY: The source code of SPANN can be accessed at https://github.com/ddb-qiwang/SPANN-torch. CONTACT: [email protected]. Musu Yuan, Qirui Guo, Minghua Deng |
Briefings Bioinform. | 5 |
| 2024 | scEVOLVE: cell-type incremental annotation without forgetting for single-cell RNA-seq dataabstractThe evolution in single-cell RNA sequencing (scRNA-seq) technology has opened a new avenue for researchers to inspect cellular heterogeneity with single-cell precision. One crucial aspect of this technology is cell-type annotation, which is fundamental for any subsequent analysis in single-cell data mining. Recently, the scientific community has seen a surge in the development of automatic annotation methods aimed at this task. However, these methods generally operate at a steady-state total cell-type capacity, significantly restricting the cell annotation systems'capacity for continuous knowledge acquisition. Furthermore, creating a unified scRNA-seq annotation system remains challenged by the need to progressively expand its understanding of ever-increasing cell-type concepts derived from a continuous data stream. In response to these challenges, this paper presents a novel and challenging setting for annotation, namely cell-type incremental annotation. This concept is designed to perpetually enhance cell-type knowledge, gleaned from continuously incoming data. This task encounters difficulty with data stream samples that can only be observed once, leading to catastrophic forgetting. To address this problem, we introduce our breakthrough methodology termed scEVOLVE, an incremental annotation method. This innovative approach is built upon the methodology of contrastive sample replay combined with the fundamental principle of partition confidence maximization. Specifically, we initially retain and replay sections of the old data in each subsequent training phase, then establish a unique prototypical learning objective to mitigate the cell-type imbalance problem, as an alternative to using cross-entropy. To effectively emulate a model that trains concurrently with complete data, we introduce a cell-type decorrelation strategy that efficiently scatters feature representations of each cell type uniformly. We constructed the scEVOLVE framework with simplicity and ease of integration into most deep softmax-based single-cell annotation methods. Thorough experiments conducted on a range of meticulously constructed benchmarks consistently prove that our methodology can incrementally learn numerous cell types over an extended period, outperforming other strategies that fail quickly. As far as our knowledge extends, this is the first attempt to propose and formulate an end-to-end algorithm framework to address this new, practical task. Additionally, scEVOLVE, coded in Python using the Pytorch machine-learning library, is freely accessible at https://github.com/aimeeyaoyao/scEVOLVE. Yuyao Zhai, Minghua Deng |
Briefings Bioinform. | 3 |
| 2024 | scBOL: a universal cell type identification framework for single-cell and spatial transcriptomics dataabstractMOTIVATION: Over the past decade, single-cell transcriptomic technologies have experienced remarkable advancements, enabling the simultaneous profiling of gene expressions across thousands of individual cells. Cell type identification plays an essential role in exploring tissue heterogeneity and characterizing cell state differences. With more and more well-annotated reference data becoming available, massive automatic identification methods have sprung up to simplify the annotation process on unlabeled target data by transferring the cell type knowledge. However, in practice, the target data often include some novel cell types that are not in the reference data. Most existing works usually classify these private cells as one generic 'unassigned' group and learn the features of known and novel cell types in a coupled way. They are susceptible to the potential batch effects and fail to explore the fine-grained semantic knowledge of novel cell types, thus hurting the model's discrimination ability. Additionally, emerging spatial transcriptomic technologies, such as in situ hybridization, sequencing and multiplexed imaging, present a novel challenge to current cell type identification strategies that predominantly neglect spatial organization. Consequently, it is imperative to develop a versatile method that can proficiently annotate single-cell transcriptomics data, encompassing both spatial and non-spatial dimensions. RESULTS: To address these issues, we propose a new, challenging yet realistic task called universal cell type identification for single-cell and spatial transcriptomics data. In this task, we aim to give semantic labels to target cells from known cell types and cluster labels to those from novel ones. To tackle this problem, instead of designing a suboptimal two-stage approach, we propose an end-to-end algorithm called scBOL from the perspective of Bipartite prototype alignment. Firstly, we identify the mutual nearest clusters in reference and target data as their potential common cell types. On this basis, we mine the cycle-consistent semantic anchor cells to build the intrinsic structure association between two data. Secondly, we design a neighbor-aware prototypical learning paradigm to strengthen the inter-cluster separability and intra-cluster compactness within each data, thereby inspiring the discriminative feature representations. Thirdly, driven by the semantic-aware prototypical learning framework, we can align the known cell types and separate the private cell types from them among reference and target data. Such an algorithm can be seamlessly applied to various data types modeled by different foundation models that can generate the embedding features for cells. Specifically, for non-spatial single-cell transcriptomics data, we use the autoencoder neural network to learn latent low-dimensional cell representations, and for spatial single-cell transcriptomics data, we apply the graph convolution network to capture molecular and spatial similarities of cells jointly. Extensive results on our carefully designed evaluation benchmarks demonstrate the superiority of scBOL over various state-of-the-art cell type identification methods. To our knowledge, we are the pioneers in presenting this pragmatic annotation task, as well as in devising a comprehensive algorithmic framework aimed at resolving this challenge across varied types of single-cell data. Finally, scBOL is implemented in Python using the Pytorch machine-learning library, and it is freely available at https://github.com/aimeeyaoyao/scBOL. Yuyao Zhai, Minghua Deng |
Briefings Bioinform. | 3 |
| 2024 | Intrinsic structure exploitation with dual alignment for universal visual recognitionabstractThis research explores the complex yet applicable domain of universal visual recognition, where unlabeled datasets may encompass additional, as-of-yet unidentified classes and present a skewed feature distribution within labeled datasets. The primary challenge lies in concurrently distinguishing instances of known classes and uncovering new classes within the unlabeled data. This difficulty is compounded by two factors: category shift and distribution shift. Current methodologies typically overlook the potential of harnessing the inherent structural relationship between labeled and unlabeled datasets, and they fail to effectively isolate the semantic variability of the newly identified classes from the established ones. To overcome these hurdles, we introduce an innovative framework, which we have termed MESA (short for Manifold structure Exploitation with double Spaces dual Alignment), to advance the field of universal visual recognition. Our approach begins by identifying geometrically and semantically related pairs of nearest neighbors between labeled and unlabeled data, which serve as anchor points to foster an intrinsic match between the two datasets. By integrating an affinity score for similarity, we establish a novel cross-view, anchor-guided, self-supervised learning paradigm. This model is designed to spread the learned labels’ information and to conjoin novel semantic insights within the predictive space. To intensify the matching accuracy for the known classes and enhance the differentiation of novel classes, we further implement a confident-prototype-driven, self-supervised learning paradigm in a cross-view manner. This strategy is aimed at revealing the macroscopic category configuration within the embedding space. MESA employs a pioneering bidirectional dual-alignment mechanism that operates simultaneously in the embedding and prediction spaces, hence providing a more robust approach to confronting challenges brought about by distribution and category shifts. Our exhaustive evaluation across several extensive image recognition benchmarks substantiates MESA’s superior performance over the latest cutting-edge methods, marking a significant leap forward in the sphere of universal visual recognition. Finally, MESA is implemented in Python using the Pytorch machine-learning library, and it is freely available at https://github.com/AimeeStar/MESA. Yuyao Zhai, Minghua Deng |
Knowl. Based Syst. | 3 |
| 2024 | Criterion-based Heterogeneous Collaborative Filtering for Multi-behavior Implicit RecommendationabstractRecent years have witnessed the explosive growth of interaction behaviors in multimedia information systems, where multi-behavior recommender systems have received increasing attention by leveraging data from various auxiliary behaviors such as tip and collect. Among various multi-behavior recommendation methods, non-sampling methods have shown superiority over negative sampling methods. However, two observations are usually ignored in existing state-of-the-art non-sampling methods based on binary regression: (1) users have different preference strengths for different items, so they cannot be measured simply by binary implicit data; (2) the dependency across multiple behaviors varies for different users and items. To tackle the above issue, we propose a novel non-sampling learning framework namedCriterion-guidedHeterogeneousCollaborativeFiltering (CHCF). CHCF introduces both upper and lower thresholds to indicate selection criteria, which will guide user preference learning. Besides, CHCF integrates criterion learning and user preference learning into a unified framework, which can be trained jointly for the interaction prediction of the target behavior. We further theoretically demonstrate that the optimization of Collaborative Metric Learning can be approximately achieved by the CHCF learning framework in a non-sampling form effectively. Extensive experiments on three real-world datasets show the effectiveness of CHCF in heterogeneous scenarios. Xiao Luo 0001, Daqing Wu, Yiyang Gu, Chong Chen 0002, Luchen Liu, Jinwen Ma, Ming Zhang 0004, Minghua Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
ACM Trans. Knowl. Discov. Data | 8 |
| 2024 | CLEAR: Cluster-Enhanced Contrast for Self-Supervised Graph Representation LearningabstractThis article studies self-supervised graph representation learning, which is critical to various tasks, such as protein property prediction. Existing methods typically aggregate representations of each individual node as graph representations, but fail to comprehensively explore local substructures (i.e., motifs and subgraphs), which also play important roles in many graph mining tasks. In this article, we propose a self-supervised graph representation learning framework named cluster-enhanced Contrast (CLEAR) that models the structural semantics of a graph from graph-level and substructure-level granularities, i.e., global semantics and local semantics, respectively. Specifically, we use graph-level augmentation strategies followed by a graph neural network-based encoder to explore global semantics. As for local semantics, we first use graph clustering techniques to partition each whole graph into several subgraphs while preserving as much semantic information as possible. We further employ a self-attention interaction module to aggregate the semantics of all subgraphs into a local-view graph representation. Moreover, we integrate both global semantics and local semantics into a multiview graph contrastive learning framework, enhancing the semantic-discriminative ability of graph representations. Extensive experiments on various real-world benchmarks demonstrate the efficacy of the proposed over current graph self-supervised representation learning approaches on both graph classification and transfer learning tasks. Xiao Luo 0001, Wei Ju 0001, Meng Qu, Yiyang Gu, Chong Chen 0002, Minghua Deng, Xian-Sheng Hua 0001, Ming Zhang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Generalized Cell Type Annotation and Discovery for Single-Cell RNA-Seq DataabstractThe rapid development of single-cell RNA sequencing (scRNA-seq) technology allows us to study gene expression heterogeneity at the cellular level. Cell annotation is the basis for subsequent downstream analysis in single-cell data mining. Existing methods rarely explore the fine-grained semantic knowledge of novel cell types absent from the reference data and usually susceptible to batch effects on the classification of seen cell types. Taking into consideration these limitations, this paper proposes a new and practical task called generalized cell type annotation and discovery for scRNA-seq data. In this task, cells of seen cell types are given class labels, while cells of novel cell types are given cluster labels instead of a unified “unassigned” label. To address this problem, we carefully design a comprehensive evaluation benchmark and propose a novel end-to-end algorithm framework called scGAD. Specifically, scGAD first builds the intrinsic correspondence across the reference and target data by retrieving the geometrically and semantically mutual nearest neighbors as anchor pairs. Then we introduce an anchor-based self-supervised learning module with a connectivity-aware attention mechanism to facilitate model prediction capability on unlabeled target data. To enhance the inter-type separation and intra-type compactness, we further propose a confidential prototypical self-supervised learning module to uncover the consensus category structure of the reference and target data. Extensive results on massive real datasets demonstrate the superiority of scGAD over various state-of-the-art clustering and annotation methods. Yuyao Zhai, Minghua Deng |
AAAI | 3 |
| 2023 | Realistic Cell Type Annotation and Discovery for Single-cell RNA-seq DataabstractThe rapid development of single-cell RNA sequencing (scRNA-seq) technologies allows us to explore tissue heterogeneity at the cellular level. Cell type annotation plays an essential role in the substantial downstream analysis of scRNA-seq data. Existing methods usually classify the novel cell types in target data as an “unassigned” group and rarely discover the fine-grained cell type structure among them. Besides, these methods carry risks, such as susceptibility to batch effect between reference and target data, thus further compromising of inherent discrimination of target data. Considering these limitations, here we propose a new and practical task called realistic cell type annotation and discovery for scRNA-seq data. In this task, cells from seen cell types are given class labels, while cells from novel cell types are given cluster labels. To tackle this problem, we propose an end-to-end algorithm framework called scPOT from the perspective of optimal transport (OT). Specifically, we first design an OT-based prototypical representation learning paradigm to encourage both global discriminations of clusters and local consistency of cells to uncover the intrinsic structure of target data. Then we propose an unbalanced OT-based partial alignment strategy with statistical filling to detect the cells from the seen cell types across reference and target data. Notably, scPOT also introduces an easy yet effective solution to automatically estimate the overall cell type number in target data. Extensive results on our carefully designed evaluation benchmarks demonstrate the superiority of scPOT over various state-of-the-art clustering and annotation methods. Yuyao Zhai, Minghua Deng |
IJCAI | 3 |
| 2023 | scGAD: a new task and end-to-end framework for generalized cell type annotation and discoveryabstractThe rapid development of single-cell RNA sequencing (scRNA-seq) technology allows us to study gene expression heterogeneity at the cellular level. Cell annotation is the basis for subsequent downstream analysis in single-cell data mining. As more and more well-annotated scRNA-seq reference data become available, many automatic annotation methods have sprung up in order to simplify the cell annotation process on unlabeled target data. However, existing methods rarely explore the fine-grained semantic knowledge of novel cell types absent from the reference data, and they are usually susceptible to batch effects on the classification of seen cell types. Taking into consideration the limitations above, this paper proposes a new and practical task called generalized cell type annotation and discovery for scRNA-seq data whereby target cells are labeled with either seen cell types or cluster labels, instead of a unified 'unassigned' label. To accomplish this, we carefully design a comprehensive evaluation benchmark and propose a novel end-to-end algorithmic framework called scGAD. Specifically, scGAD first builds the intrinsic correspondences on seen and novel cell types by retrieving geometrically and semantically mutual nearest neighbors as anchor pairs. Together with the similarity affinity score, a soft anchor-based self-supervised learning module is then designed to transfer the known label information from reference data to target data and aggregate the new semantic knowledge within target data in the prediction space. To enhance the inter-type separation and intra-type compactness, we further propose a confidential prototype self-supervised learning paradigm to implicitly capture the global topological structure of cells in the embedding space. Such a bidirectional dual alignment mechanism between embedding space and prediction space can better handle batch effect and cell type shift. Extensive results on massive simulation datasets and real datasets demonstrate the superiority of scGAD over various state-of-the-art clustering and annotation methods. We also implement marker gene identification to validate the effectiveness of scGAD in clustering novel cell types and their biological significance. To the best of our knowledge, we are the first to introduce this new and practical task and propose an end-to-end algorithmic framework to solve it. Our method scGAD is implemented in Python using the Pytorch machine-learning library, and it is freely available at https://github.com/aimeeyaoyao/scGAD. Yuyao Zhai, Minghua Deng |
Briefings Bioinform. | 3 |
| 2023 | Hierarchical graph transformer with contrastive learning for protein function predictionabstractMOTIVATION: In recent years, high-throughput sequencing technologies have made large-scale protein sequences accessible. However, their functional annotations usually rely on low-throughput and pricey experimental studies. Computational prediction models offer a promising alternative to accelerate this process. Graph neural networks have shown significant progress in protein research, but capturing long-distance structural correlations and identifying key residues in protein graphs remains challenging. RESULTS: In the present study, we propose a novel deep learning model named Hierarchical graph transformEr with contrAstive Learning (HEAL) for protein function prediction. The core feature of HEAL is its ability to capture structural semantics using a hierarchical graph Transformer, which introduces a range of super-nodes mimicking functional motifs to interact with nodes in the protein graph. These semantic-aware super-node embeddings are then aggregated with varying emphasis to produce a graph representation. To optimize the network, we utilized graph contrastive learning as a regularization technique to maximize the similarity between different views of the graph representation. Evaluation of the PDBch test set shows that HEAL-PDB, trained on fewer data, achieves comparable performance to the recent state-of-the-art methods, such as DeepFRI. Moreover, HEAL, with the added benefit of unresolved protein structures predicted by AlphaFold2, outperforms DeepFRI by a significant margin on Fmax, AUPR, and Smin metrics on PDBch test set. Additionally, when there are no experimentally resolved structures available for the proteins of interest, HEAL can still achieve better performance on AFch test set than DeepFRI and DeepGOPlus by taking advantage of AlphaFold2 predicted structures. Finally, HEAL is capable of finding functional sites through class activation mapping. AVAILABILITY AND IMPLEMENTATION: Implementations of our HEAL can be found at https://github.com/ZhonghuiGu/HEAL. Zhonghui Gu, Xiao Luo 0001, Jiaxiao Chen, Minghua Deng, Luhua Lai |
Bioinform. | 4 |
| 2023 | Clustering single-cell multi-omics data with MoClustabstractMOTIVATION: Single-cell multi-omics sequencing techniques have rapidly developed in the past few years. Clustering analysis with single-cell multi-omics data may give us novel perspectives to dissect cellular heterogeneity. However, multi-omics data have the properties of inherited large dimension, high sparsity and existence of doublets. Moreover, representations of different omics from even the same cell follow diverse distributions. Without proper distribution alignment techniques, clustering methods will encounter less separable clusters easily affected by less informative omics data. RESULTS: We developed MoClust, a novel joint clustering framework that can be applied to several types of single-cell multi-omics data. A selective automatic doublet detection module that can identify and filter out doublets is introduced in the pretraining stage to improve data quality. Omics-specific autoencoders are introduced to characterize the multi-omics data. A contrastive learning way of distribution alignment is adopted to adaptively fuse omics representations into an omics-invariant representation. This novel way of alignment boosts the compactness and separableness of clusters, while accurately weighting the contribution of each omics to the clustering object. Extensive experiments, over both simulated and real multi-omics datasets, demonstrated the powerful alignment, doublet detection and clustering ability features of MoClust. AVAILABILITY AND IMPLEMENTATION: An implementation of MoClust is available from https://doi.org/10.5281/zenodo.7306504. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Musu Yuan, Minghua Deng |
Bioinform. | 3 |
| 2023 | A Survey on Deep Hashing MethodsabstractNearest neighbor search aims at obtaining the samples in the database with the smallest distances from them to the queries, which is a basic task in a range of fields, including computer vision and data mining. Hashing is one of the most widely used methods for its computational and storage efficiency. With the development of deep learning, deep hashing methods show more advantages than traditional methods. In this survey, we detailedly investigate current deep hashing algorithms including deep supervised hashing and deep unsupervised hashing. Specifically, we categorize deep supervised hashing methods into pairwise methods, ranking-based methods, pointwise methods as well as quantization according to how measuring the similarities of the learned hash codes. Moreover, deep unsupervised hashing is categorized into similarity reconstruction-based methods, pseudo-label-based methods, and prediction-free self-supervised learning-based methods based on their semantic learning manners. We also introduce three related important topics including semi-supervised deep hashing, domain adaption deep hashing, and multi-modal deep hashing. Meanwhile, we present some commonly used public datasets and the scheme to measure the performance of deep hashing algorithms. Finally, we discuss some potential research directions in conclusion. Xiao Luo 0001, Haixin Wang 0003, Daqing Wu, Chong Chen 0002, Minghua Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2022 | Mutual Nearest Neighbor Contrast and Hybrid Prototype Self-Training for Universal Domain AdaptationabstractUniversal domain adaptation (UniDA) aims to transfer knowledge learned from a labeled source domain to an unlabeled target domain under domain shift and category shift. Without prior category overlap information, it is challenging to simultaneously align the common categories between two domains and separate their respective private categories. Additionally, previous studies utilize the source classifier's prediction to obtain various known labels and one generic "unknown" label of target samples. However, over-reliance on learned classifier knowledge is inevitably biased to source data, ignoring the intrinsic structure of target domain. Therefore, in this paper, we propose a novel two-stage UniDA framework called MATHS based on the principle of mutual nearest neighbor contrast and hybrid prototype discrimination. In the first stage, we design an efficient mutual nearest neighbor contrastive learning scheme to achieve feature alignment, which exploits the instance-level affinity relationship to uncover the intrinsic structure of two domains. We introduce a bimodality hypothesis for the maximum discriminative probability distribution to detect the possible target private samples, and present a data-based statistical approach to separate the common and private categories. In the second stage, to obtain more reliable label predictions, we propose an incremental pseudo-classifier for target data only, which is driven by the hybrid representative prototypes. A confidence-guided prototype contrastive loss is designed to optimize the category allocation uncertainty via a self-training mechanism. Extensive experiments on three benchmarks demonstrate that MATHS outperforms previous state-of-the-arts on most UniDA settings. Qianjin Du, Yihang Lou, Minghua Deng |
AAAI | 6 |
| 2022 | Evidential Neighborhood Contrastive Learning for Universal Domain AdaptationabstractUniversal domain adaptation (UniDA) aims to transfer the knowledge learned from a labeled source domain to an unlabeled target domain without any constraints on the label sets. However, domain shift and category shift make UniDA extremely challenging, mainly attributed to the requirement of identifying both shared “known” samples and private “unknown” samples. Previous methods barely exploit the intrinsic manifold structure relationship between two domains for feature alignment, and they rely on the softmax-based scores with class competition nature to detect underlying “unknown” samples. Therefore, in this paper, we propose a novel evidential neighborhood contrastive learning framework called TNT to address these issues. Specifically, TNT first proposes a new domain alignment principle: semantically consistent samples should be geometrically adjacent to each other, whether within or across domains. From this criterion, a cross-domain multi-sample contrastive loss based on mutual nearest neighbors is designed to achieve common category matching and private category separation. Second, toward accurate “unknown” sample detection, TNT introduces a class competition-free uncertainty score from the perspective of evidential deep learning. Instead of setting a single threshold, TNT learns a category-aware heterogeneous threshold vector to reject diverse “unknown” samples. Extensive experiments on three benchmarks demonstrate that TNT significantly outperforms previous state-of-the-art UniDA methods. Yihang Lou, Minghua Deng |
AAAI | 5 |
| 2022 | Neighborhood Consensus Contrastive Learning for Backward-Compatible RepresentationabstractIn object re-identification (ReID), the development of deep learning techniques often involves model updates and deployment. It is unbearable to re-embedding and re-index with the system suspended when deploying new models. Therefore, backward-compatible representation is proposed to enable ``new'' features to be compared with ``old'' features directly, which means that the database is active when there are both ``new'' and ``old'' features in it. Thus we can scroll-refresh the database or even do nothing on the database to update. The existing backward-compatible methods either require a strong overlap between old and new training data or simply conduct constraints at the instance level. Thus they are difficult in handling complicated cluster structures and are limited in eliminating the impact of outliers in old embeddings, resulting in a risk of damaging the discriminative capability of new features. In this work, we propose a Neighborhood Consensus Contrastive Learning (NCCL) method. With no assumptions about the new training data, we estimate the sub-cluster structures of old embeddings. A new embedding is constrained with multiple old embeddings in both embedding space and discrimination space at the sub-class level. The effect of outliers diminished, as the multiple samples serve as ``mean teachers''. Besides, we propose a scheme to filter the old embeddings with low credibility, further improving the compatibility robustness. Our method ensures the compatibility without impairing the accuracy of the new model. It can even improve the new model's accuracy in most scenarios. Shengsen Wu, Yihang Lou, Minghua Deng, Ling-Yu Duan |
AAAI | 6 |
| 2022 | Geometric Anchor Correspondence Mining with Uncertainty Modeling for Universal Domain AdaptationabstractUniversal domain adaptation (UniDA) aims to transfer the knowledge learned from a label-rich source domain to a label-scarce target domain without any constraints on the label space. However, domain shift and category shift make UniDA extremely challenging, which mainly lies in how to recognize both shared “known” samples and private “unknown” samples. Previous works rarely explore the intrinsic geometrical relationship between the two domains, and they manually set a threshold for the overconfident closed-world classifier to reject “unknown” samples. Therefore, in this paper, we propose a Geometric anchor-guided Adversarial and conTrastive learning framework with uncErtainty modeling called GATE to alleviate these issues. Specifically, we first develop a random walk-based anchor mining strategy together with a high-order attention mechanism to build correspondence across domains. Then a global joint local domain alignment paradigm is designed, i.e., geometric adversarial learning for global distribution calibration and subgraph-level contrastive learning for local region aggregation. Toward accurate target private samples detection, GATE introduces a universal incremental classifier by modeling the energy uncertainty. We further efficiently generate novel categories by manifold mixup, and minimize the open-set entropy to learn the “unknown” threshold adaptively. Extensive experiments on three benchmarks demonstrate that GATE significantly out-performs previous state-of-the-art UniDA methods. Yihang Lou, Minghua Deng |
CVPR | 5 |
| 2022 | DHWP: Learning High-Quality Short Hash Codes Via Weight PruningabstractHashing is widely used in large-scale image retrieval because of its efficiency in storage and computation. Although longer hash codes can lead to higher search accuracy, the retrieval cost increases linearly with the increase of the number of hash bits. Most deep hashing methods suffer from the problem of trivial solutions and usually result in highly correlated redundant hash bits in practice, which limits the performance. To obtain short hash codes with high quality for fast and accurate image retrieval, we propose a novel framework named Deep Hashing via Weight Pruning (DHWP). DHWP first trains the model with relatively long hash codes. Then it obtains shorter codes gradually by weight pruning based on four different criteria. The framework of DHWP can be applied to most deep supervised hashing models, which helps remove those redundant hash bits while retaining the representation ability of long hash codes. Extensive experimental results on two widely used benchmark datasets show that DHWP outperforms the existing state-of-the-art methods, especially for short hash codes. Zeyu Ma 0001, Yuhang Guo 0002, Xiao Luo 0001, Chong Chen 0002, Minghua Deng, Wei Cheng 0002, Guangming Lu 0002 |
ICASSP | 5 |
| 2022 | DualGraph: Improving Semi-supervised Graph Classification via Dual Contrastive LearningabstractIn this paper, we study semi-supervised graph classification, a fundamental problem in data mining and machine learning. The problem is typically solved by learning graph neural networks with pseudo-labeling or knowledge distillation to incorporate both labeled and unlabeled graphs. However, these methods usually either suffer from overconfident and biased pseudo-labels or suboptimal distillation caused by the insufficient use of unlabeled data. Inspired by the recent progress of contrastive learning and dual learning, we propose DualGraph, a principled framework to leverage unlabeled graphs more effectively for semi-supervised graph classification. DualGraph consists of a prediction module and a retrieval module to model graphs$G$and their labels$y$from opposite while complementary views (i.e., p(y | G) and p(G | y) respectively). The two modules are jointly trained via posterior regularization, which encourages their inter-module consistency on unlabeled graphs. Moreover, we improve model training for each module with a contrastive learning framework to encourage the intra-module consistency on unlabeled data. Experimental results on a range of publicly accessible datasets reveal the effectiveness of our DualGraph. Xiao Luo 0001, Wei Ju 0001, Meng Qu, Chong Chen 0002, Minghua Deng, Xian-Sheng Hua 0001, Ming Zhang 0004 |
ICDE | 5 |
| 2022 | TGNN: A Joint Semi-supervised Framework for Graph-level ClassificationabstractThis paper studies semi-supervised graph classification, a crucial task with a wide range of applications in social network analysis and bioinformatics. Recent works typically adopt graph neural networks to learn graph-level representations for classification, failing to explicitly leverage features derived from graph topology (e.g., paths). Moreover, when labeled data is scarce, these methods are far from satisfactory due to their insufficient topology exploration of unlabeled data. We address the challenge by proposing a novel semi-supervised framework called Twin Graph Neural Network (TGNN). To explore graph structural information from complementary views, our TGNN has a message passing module and a graph kernel module. To fully utilize unlabeled data, for each module, we calculate the similarity of each unlabeled graph to other labeled graphs in the memory bank and our consistency loss encourages consistency between two similarity distributions in different embedding spaces. The two twin modules collaborate with each other by exchanging instance similarity knowledge to fully explore the structure information of both labeled and unlabeled data. We evaluate our TGNN on various public datasets and show that it achieves strong performance. Wei Ju 0001, Xiao Luo 0001, Meng Qu, Yifan Wang 0014, Chong Chen 0002, Minghua Deng, Xian-Sheng Hua 0001, Ming Zhang 0004 |
IJCAI | 6 |
| 2022 | Improved Deep Unsupervised Hashing with Fine-grained Semantic Similarity Mining for Multi-Label Image RetrievalabstractIn this paper, we study deep unsupervised hashing, a critical problem for approximate nearest neighbor research. Most recent methods solve this problem by semantic similarity reconstruction for guiding hashing network learning or contrastive learning of hash codes. However, in multi-label scenarios, these methods usually either generate an inaccurate similarity matrix without reflection of similarity ranking or suffer from the violation of the underlying assumption in contrastive learning, resulting in limited retrieval performance. To tackle this issue, we propose a novel method termed HAMAN, which explores semantics from a fine-grained view to enhance the ability of multi-label image retrieval. In particular, we reconstruct the pairwise similarity structure by matching fine-grained patch features generated by the pre-trained neural network, serving as reliable guidance for similarity preserving of hash codes. Moreover, a novel conditional contrastive learning on hash codes is proposed to adopt self-supervised learning in multi-label scenarios. According to extensive experiments on three multi-label datasets, the proposed method outperforms a broad range of state-of-the-art methods. Zeyu Ma 0001, Xiao Luo 0001, Yingjie Chen 0002, Mi-Xiao Hou, Jinxing Li 0003, Minghua Deng, Guangming Lu 0002 |
IJCAI | 6 |
| 2022 | Rare variant association tests for ancestry-matched case-control data based on conditional logistic regressionabstractWith the increasing volume of human sequencing data available, analysis incorporating external controls becomes a popular and cost-effective approach to boost statistical power in disease association studies. To prevent spurious association due to population stratification, it is important to match the ancestry backgrounds of cases and controls. However, rare variant association tests based on a standard logistic regression model are conservative when all ancestry-matched strata have the same case-control ratio and might become anti-conservative when case-control ratio varies across strata. Under the conditional logistic regression (CLR) model, we propose a weighted burden test (CLR-Burden), a variance component test (CLR-SKAT) and a hybrid test (CLR-MiST). We show that the CLR model coupled with ancestry matching is a general approach to control for population stratification, regardless of the spatial distribution of disease risks. Through extensive simulation studies, we demonstrate that the CLR-based tests robustly control type 1 errors under different matching schemes and are more powerful than the standard Burden, SKAT and MiST tests. Furthermore, because CLR-based tests allow for different case-control ratios across strata, a full-matching scheme can be employed to efficiently utilize all available cases and controls to accelerate the discovery of disease associated genes. Jingjing Lyu, Xian Shi, Zengmiao Wang, Minghua Deng, Baoluo Sun, Chaolong Wang |
Briefings Bioinform. | 6 |
| 2022 | scNAME: neighborhood contrastive clustering with ancillary mask estimation for scRNA-seq dataabstractMOTIVATION: The rapid development of single-cell RNA sequencing (scRNA-seq) makes it possible to study the heterogeneity of individual cell characteristics. Cell clustering is a vital procedure in scRNA-seq analysis, providing insight into complex biological phenomena. However, the noisy, high-dimensional and large-scale nature of scRNA-seq data introduces challenges in clustering analysis. Up to now, many deep learning-based methods have emerged to learn underlying feature representations while clustering. However, these methods are inefficient when it comes to rare cell type identification and barely able to fully utilize gene dependencies or cell similarity integrally. As a result, they cannot detect a clear cell type structure which is required for clustering accuracy as well as downstream analysis. RESULTS: Here, we propose a novel scRNA-seq clustering algorithm called scNAME which incorporates a mask estimation task for gene pertinence mining and a neighborhood contrastive learning framework for cell intrinsic structure exploitation. The learned pattern through mask estimation helps reveal uncorrupted data structure and denoise the original single-cell data. In addition, the randomly created augmented data introduced in contrastive learning not only helps improve robustness of clustering, but also increases sample size in each cluster for better data capacity. Beyond this, we also introduce a neighborhood contrastive paradigm with an offline memory bank, global in scope, which can inspire discriminative feature representation and achieve intra-cluster compactness, yet inter-cluster separation. The combination of mask estimation task, neighborhood contrastive learning and global memory bank designed in scNAME is conductive to rare cell type detection. The experimental results of both simulations and real data confirm that our method is accurate, robust and scalable. We also implement biological analysis, including marker gene identification, gene ontology and pathway enrichment analysis, to validate the biological significance of our method. To the best of our knowledge, we are among the first to introduce a gene relationship exploration strategy, as well as a global cellular similarity repository, in the single-cell field. AVAILABILITY AND IMPLEMENTATION: An implementation of scNAME is available from https://github.com/aster-ww/scNAME. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Minghua Deng |
Bioinform. | 3 |
| 2022 | scMRA: a robust deep learning method to annotate scRNA-seq data with multiple reference datasetsabstractMOTIVATION: Single-cell RNA-seq (scRNA-seq) has been widely used to resolve cellular heterogeneity. After collecting scRNA-seq data, the natural next step is to integrate the accumulated data to achieve a common ontology of cell types and states. Thus, an effective and efficient cell-type identification method is urgently needed. Meanwhile, high-quality reference data remain a necessity for precise annotation. However, such tailored reference data are always lacking in practice. To address this, we aggregated multiple datasets into a meta-dataset on which annotation is conducted. Existing supervised or semi-supervised annotation methods suffer from batch effects caused by different sequencing platforms, the effect of which increases in severity with multiple reference datasets. RESULTS: Herein, a robust deep learning-based single-cell Multiple Reference Annotator (scMRA) is introduced. In scMRA, a knowledge graph is constructed to represent the characteristics of cell types in different datasets, and a graphic convolutional network serves as a discriminator based on this graph. scMRA keeps intra-cell-type closeness and the relative position of cell types across datasets. scMRA is remarkably powerful at transferring knowledge from multiple reference datasets, to the unlabeled target domain, thereby gaining an advantage over other state-of-the-art annotation methods in multi-reference data experiments. Furthermore, scMRA can remove batch effects. To the best of our knowledge, this is the first attempt to use multiple insufficient reference datasets to annotate target data, and it is, comparatively, the best annotation method for multiple scRNA-seq datasets. AVAILABILITY AND IMPLEMENTATION: An implementation of scMRA is available from https://github.com/ddb-qiwang/scMRA-torch. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Musu Yuan, Minghua Deng |
Bioinform. | 3 |
| 2022 | GHNN: Graph Harmonic Neural Networks for semi-supervised graph-level classification
Wei Ju 0001, Xiao Luo 0001, Zeyu Ma 0001, Minghua Deng, Ming Zhang 0004 |
Neural Networks | 5 |
| 2022 | Profiling transcription factor activity dynamics using intronic reads in time-series transcriptome dataabstractActivities of transcription factors (TFs) are temporally modulated to regulate dynamic cellular processes, including development, homeostasis, and disease. Recent developments of bioinformatic tools have enabled the analysis of TF activities using transcriptome data. However, because these methods typically use exon-based target expression levels, the estimated TF activities have limited temporal accuracy. To address this, we proposed a TF activity measure based on intron-level information in time-series RNA-seq data, and implemented it to decode the temporal control of TF activities during dynamic processes. We showed that TF activities inferred from intronic reads can better recapitulate instantaneous TF activities compared to the exon-based measure. By analyzing public and our own time-series transcriptome data, we found that intron-based TF activities improve the characterization of temporal phasing of cycling TFs during circadian rhythm, and facilitate the discovery of two temporally opposing TF modules during T cell activation. Collectively, we anticipate that the proposed approach would be broadly applicable for decoding global transcriptional architecture during dynamic processes. Lingfeng Xue, Wen Huang 0011, Minghua Deng |
PLoS Comput. Biol. | 4 |
| 2022 | Improve Deep Unsupervised Hashing via Structural and Intrinsic Similarity LearningabstractHashing has attracted increasing attention in image retrieval recently due to its storage and computational efficiency. Although several deep unsupervised hashing methods have been proposed lately, their effectiveness is far from satisfactory in practice owing to two drawbacks. On the one hand, they mostly construct binary similarity matrices which could neglect the confidence differences among multiple similarity signals. On the other hand, they ignore the desired properties of hash codes (i.e., independence and robustness). In this paper, we propose an effective unsupervised hashing method calledHashing viaStructural andIntrinsic siMilarity learning (HashSIM) to tackle these issues in an end-to-end manner. Specifically, HashSIM utilizes both highly and normally confident image pairs to jointly build a continuous similarity matrix, which guides hash code learning via structural similarity learning. Moreover, inspired by contrastive learning, we impose an intrinsic similarity learning objective, which can maximally satisfy the independence and robustness properties of hash bits. Extensive experiments on three popular benchmark datasets demonstrate that our HashSIM outperforms a broad range of state-of-the-art baselines. Xiao Luo 0001, Zeyu Ma 0001, Wei Cheng 0002, Minghua Deng |
IEEE Signal Process. Lett. | 4 |
| 2021 | Graph Contrastive ClusteringabstractRecently, some contrastive learning methods have been proposed to simultaneously learn representations and clustering assignments, achieving significant improvements. However, these methods do not take the category information and clustering objective into consideration, thus the learned representations are not optimal for clustering and the performance might be limited. Towards this issue, we first propose a novel graph contrastive learning framework, and then apply it to the clustering task, resulting in the Graph Constrastive Clustering (GCC) method. Different from basic contrastive clustering that only assumes an image and its augmentation should share similar representation and clustering assignments, we lift the instance-level consistency to the cluster-level consistency with the assumption that samples in one cluster and their augmentations should all be similar. Specifically, on the one hand, we propose the graph Laplacian based contrastive loss to learn more discriminative and clustering-friendly features. On the other hand, we propose a novel graph-based contrastive learning strategy to learn more compact clustering assignments. Both of them incorporate the latent category information to reduce the intra-cluster variance as well as increase the inter-cluster variance. Experiments on six commonly used datasets demonstrate the superiority of our proposed approach over the state-of-the-art methods.1 Huasong Zhong, Jianlong Wu, Chong Chen 0002, Jianqiang Huang 0001, Minghua Deng, Liqiang Nie, Zhouchen Lin, Xian-Sheng Hua 0001 |
ICCV | 5 |
| 2021 | Composition-Enhanced Graph Collaborative Filtering for Multi-behavior RecommendationabstractRapid and accurate prediction of user preferences is the ultimate goal of today’s recommender systems. More and more researchers pay attention to multi-behavior recommender systems which utilize the auxiliary types of user-item interaction data, such as page view and add-to-cart to help estimate user preferences. Recently, graph-based methods were proposed to showcase an advanced capability in representation learning and capturing collaborative signals. However, we argue that these methods ignore the intrinsic difference between the two types of nodes in the bipartite graph and aggregate information from neighboring nodes with the same functions. Besides, these models do not fully explore the collaborative signals implied by the meta-path across different types of behavior, which causes a huge loss of the potential semantic information across behaviors. To address the above limitations, we present a unified graph model named SaGCN (short for Semantic-aware Graph Convolutional Networks). Specifically, we construct separate user-user and item-item graphs by meta-path, and apply separate aggregation and transformation functions to propagate user and item information. To perform better semantic propagation, we design a relation composition function and a semantic propagation architecture for heterogeneous collaborative filtering signals learning. Extensive experiments on two real-world datasets show that SaGCN outperforms a wide range of state-of-the-art methods in multi-behavior scenarios. Daqing Wu, Xiao Luo 0001, Zeyu Ma 0001, Chong Chen 0002, Pengfei Wang 0008, Minghua Deng, Jinwen Ma |
ICDM | 6 |
| 2021 | DNA-GCN: Graph Convolutional Networks for Predicting DNA-Protein Binding
Yuhang Guo 0002, Xiao Luo 0001, Minghua Deng |
ICIC (3) | 4 |
| 2021 | Deep Unsupervised Hashing by Distilled Smooth GuidanceabstractHashing has been widely used in approximate nearest neighbor search recently. Deep supervised hashing methods are not widely-used because of the lack of labeled data, especially when the domain is transferred. Meanwhile, unsupervised deep hashing models can hardly achieve satisfactory performance due to the lack of reliable similarity signals. Here, we propose a novel deep unsupervised hashing method, namely Distilled Smooth Guidance (DSG), which can learn a distilled dataset consisting of similarity signals as well as smooth confidence signals. Specifically, we obtain the similarity confidence weights based on the initial noisy similarity signals learned from local structures and construct a priority loss function for smooth similarity-preserving learning. Besides, global information based on clustering is utilized to distill the image pairs by removing contradictory similarity signals. Extensive experiments on three widely used bench-mark datasets show that the proposed DSG consistently out-performs the state-of-the-art search methods. Xiao Luo 0001, Zeyu Ma 0001, Daqing Wu, Huasong Zhong, Chong Chen 0002, Jinwen Ma, Minghua Deng |
ICME | 7 |
| 2021 | Deep Unsupervised Hashing by Global and Local ConsistencyabstractHashing is widely-used in approximate nearest neighbor search for its computational efficiency. Most of the existing unsupervised hashing methods are based on local consistency that the Hamming distance between two images should be small if their features are similar. However, many false similar pairs may be included for the insufficient representation of features. Here we proposed deep unsupervised hashing by Global and Local Consistency (GLC). Specifically, GLC has two components named semantic information generating and semantic consistency learning, and each component is conducted from both global and local views. From local view, GLC introduces reliable graph and penalty graph to capture local signals with high confidence to preserve the semantic structure. From global view, GLC includes a distribution loss to capture the global consistency with cluster signals. Extensive experimental results on three widely-used benchmark datasets show that GLC performs better than existing state- of-the-art methods. Xiao Luo 0001, Daqing Wu, Chong Chen 0002, Jinwen Ma, Minghua Deng |
ICME | 5 |
| 2021 | Deep Supervised Hashing by Classification for Image Retrieval
Xiao Luo 0001, Yuhang Guo 0002, Zeyu Ma 0001, Huasong Zhong, Tao Li 0040, Wei Ju 0001, Chong Chen 0002, Minghua Deng |
ICONIP (4) | 8 |
| 2021 | Concordant Contrastive Learning for Semi-supervised Node Classification on Graph
Daqing Wu, Xiao Luo 0001, Xiangyang Guo, Chong Chen 0002, Minghua Deng, Jinwen Ma |
ICONIP (1) | 5 |
| 2021 | An Interpretation of Convolutional Neural Networks for Motif Finding from the View of ProbabilityabstractThe powerful learning ability of a convolutional neural network to perform functional classification provides valuable clues for the discovery of biological problems and image recognition. However, the mechanism of deep learning still lacks satisfying interpretation and the convolutional neural network is usually regarded as a black-box model. Consequently, it is challenging to design deep learning models and adjust the hyper-parameters for specific tasks without theoretical guidance. We raise a new task of deep learning – 2-dimensional logo detecting for computer vision and develop a novel model of convolutional neural networks to solve the task. Furthermore, we interpret this specific model from the point of view of statistical learning. With the guidance of statistical interpretation, we can design the model structure more efficiently and obtain better performances. Yuhang Guo 0002, Wei Ju 0001, Xiao Luo 0001, Minghua Deng |
ICTAI | 5 |
| 2021 | CIMON: Towards High-quality Hash CodesabstractRecently, hashing is widely used in approximate nearest neighbor search for its storage and computational efficiency. Most of the unsupervised hashing methods learn to map images into semantic similarity-preserving hash codes by constructing local semantic similarity structure from the pre-trained model as the guiding information, i.e., treating each point pair similar if their distance is small in feature space. However, due to the inefficient representation ability of the pre-trained model, many false positives and negatives in local semantic similarity will be introduced and lead to error propagation during the hash code learning. Moreover, few of the methods consider the robustness of models, which will cause instability of hash codes to disturbance. In this paper, we propose a new method named Comprehensive sImilarity Mining and cOnsistency learNing (CIMON). First, we use global refinement and similarity statistical distribution to obtain reliable and smooth guidance. Second, both semantic and contrastive consistency learning are introduced to derive both disturb-invariant and discriminative hash codes. Extensive experiments on several benchmark datasets show that the proposed method outperforms a wide range of state-of-the-art methods in both retrieval performance and robustness. Xiao Luo 0001, Daqing Wu, Zeyu Ma 0001, Chong Chen 0002, Minghua Deng, Jinwen Ma, Zhongming Jin 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
IJCAI | 5 |
| 2021 | ARGO: Modeling Heterogeneity in E-commerce RecommendationabstractWith the increasing scale and diversification of interaction behaviors in E-commerce, more and more researchers pay attention to multi-behavior recommender systems which utilize interaction data of other auxiliary behaviors. However, all existing models ignore two kinds of intrinsic heterogeneity which are helpful to capture the difference of user preferences and the difference of item attributes. First (intra-heterogeneity), each user has multiple social identities with otherness, and these different identities can result in quite different interaction preferences. Second (inter-heterogeneity), each item can transfer an item-specific percentage of score from low-level behavior to high-level behavior for the gradual relationship among multiple behaviors. Thus, the lack of consideration of these heterogeneities damages recommendation rank performance. To model the above heterogeneities, we propose a novel method named intrA- and inteR-heteroGeneity recOmmendation model (ARGO). Specifically, we embed each user into multiple vectors representing the user's identities, and the maximum of identity scores indicates the interaction preference. Besides, we regard the item-specific transition percentage as trainable transition probability between different behaviors. Extensive experiments on two real-world datasets show that ARGO performs much better than the state-of-the-art in multi-behavior scenarios. Daqing Wu, Xiao Luo 0001, Zeyu Ma 0001, Chong Chen 0002, Minghua Deng, Jinwen Ma |
IJCNN | 5 |
| 2021 | A Statistical Approach to Mining Semantic Similarity for Deep Unsupervised HashingabstractThe majority of deep unsupervised hashing methods usually first construct pairwise semantic similarity information and then learn to map images into compact hash codes while preserving the similarity structure, which implies that the quality of hash codes highly depends on the constructed semantic similarity structure. However, since the features of images for each kind of semantics usually scatter in high-dimensional space with unknown distribution, previous methods could introduce a large number of false positives and negatives for boundary points of distributions in the local semantic structure based on pairwise cosine distances. Towards this limitation, we propose a general distribution-based metric to depict the pairwise distance between images. Specifically, each image is characterized by its random augmentations that can be viewed as samples from the corresponding latent semantic distribution. Then we estimate the distances between images by calculating the sample distribution divergence of their semantics. By applying this new metric to deep unsupervised hashing, we come up with Distribution-based similArity sTructure rEconstruction (DATE). DATE can generate more accurate semantic similarity information by using non-parametric ball divergence. Moreover, DATE explores both semantic-preserving learning and contrastive learning to obtain high-quality hash codes. Extensive experiments on several widely-used datasets validate the superiority of our DATE. Xiao Luo 0001, Daqing Wu, Zeyu Ma 0001, Chong Chen 0002, Minghua Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 5 |
| 2021 | Single-cell RNA-seq data semi-supervised clustering and annotation via structural regularized domain adaptationabstractMOTIVATION: The rapid development of single-cell RNA sequencing (scRNA-seq) technologies allows us to explore tissue heterogeneity at the cellular level. The identification of cell types plays an essential role in the analysis of scRNA-seq data, which, in turn, influences the discovery of regulatory genes that induce heterogeneity. As the scale of sequencing data increases, the classical method of combining clustering and differential expression analysis to annotate cells becomes more costly in terms of both labor and resources. Existing scRNA-seq supervised classification method can alleviate this issue through learning a classifier trained on the labeled reference data and then making a prediction based on the unlabeled target data. However, such label transference strategy carries with risks, such as susceptibility to batch effect and further compromise of inherent discrimination of target data. RESULTS: In this article, inspired by unsupervised domain adaptation, we propose a flexible single cell semi-supervised clustering and annotation framework, scSemiCluster, which integrates the reference data and target data for training. We utilize structure similarity regularization on the reference domain to restrict the clustering solutions of the target domain. We also incorporates pairwise constraints in the feature learning process such that cells belonging to the same cluster are close to each other, and cells belonging to different clusters are far from each other in the latent space. Notably, without explicit domain alignment and batch effect correction, scSemiCluster outperforms other state-of-the-art, single-cell supervised classification and semi-supervised clustering annotation algorithms in both simulation and real data. To the best of our knowledge, we are the first to use both deep discriminative clustering and deep generative clustering techniques in the single-cell field. AVAILABILITYAND IMPLEMENTATION: An implementation of scSemiCluster is available from https://github.com/xuebaliang/scSemiCluster. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qiuyan He, Yuyao Zhai, Minghua Deng |
Bioinform. | 4 |
| 2020 | scRMD: imputation for single cell RNA-seq data via robust matrix decompositionabstractMOTIVATION: Single cell RNA-sequencing (scRNA-seq) technology enables whole transcriptome profiling at single cell resolution and holds great promises in many biological and medical applications. Nevertheless, scRNA-seq often fails to capture expressed genes, leading to the prominent dropout problem. These dropouts cause many problems in down-stream analysis, such as significant increase of noises, power loss in differential expression analysis and obscuring of gene-to-gene or cell-to-cell relationship. Imputation of these dropout values can be beneficial in scRNA-seq data analysis. RESULTS: In this article, we model the dropout imputation problem as robust matrix decomposition. This model has minimal assumptions and allows us to develop a computational efficient imputation method called scRMD. Extensive data analysis shows that scRMD can accurately recover the dropout values and help to improve downstream analysis such as differential expression analysis and clustering analysis. AVAILABILITY AND IMPLEMENTATION: The R package scRMD is available at https://github.com/XiDsLab/scRMD. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chong Chen 0002, Changjing Wu, Linjie Wu, Minghua Deng, Ruibin Xi |
Bioinform. | 5 |
| 2020 | Expectation pooling: an effective and interpretable pooling method for predicting DNA-protein bindingabstractMOTIVATION: Convolutional neural networks (CNNs) have outperformed conventional methods in modeling the sequence specificity of DNA-protein binding. While previous studies have built a connection between CNNs and probabilistic models, simple models of CNNs cannot achieve sufficient accuracy on this problem. Recently, some methods of neural networks have increased performance using complex neural networks whose results cannot be directly interpreted. However, it is difficult to combine probabilistic models and CNNs effectively to improve DNA-protein binding predictions. RESULTS: In this article, we present a novel global pooling method: expectation pooling for predicting DNA-protein binding. Our pooling method stems naturally from the expectation maximization algorithm, and its benefits can be interpreted both statistically and via deep learning theory. Through experiments, we demonstrate that our pooling method improves the prediction performance DNA-protein binding. Our interpretable pooling method combines probabilistic ideas with global pooling by taking the expectations of inputs without increasing the number of parameters. We also analyze the hyperparameters in our method and propose optional structures to help fit different datasets. We explore how to effectively utilize these novel pooling methods and show that combining statistical methods with deep learning is highly beneficial, which is promising and meaningful for future studies in this field. AVAILABILITY AND IMPLEMENTATION: All code is public in https://github.com/gao-lab/ePooling. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiao Luo 0001, Xinming Tu, Ge Gao 0004, Minghua Deng |
Bioinform. | 5 |
| 2019 | Compositional data network analysis via lasso penalized D-trace lossabstractMOTIVATION: With the development of high-throughput sequencing techniques for 16S-rRNA gene profiling, the analysis of microbial communities is becoming more and more attractive and reliable. Inferring the direct interaction network among microbial communities helps in the identification of mechanisms underlying community structure. However, the analysis of compositional data remains challenging by the relative information conveyed by such data, as well as its high dimensionality. RESULTS: In this article, we first propose a novel loss function for compositional data called CD-trace based on D-trace loss. A sparse matrix estimator for the direct interaction network is defined as the minimizer of lasso penalized CD-trace loss under positive-definite constraint. An efficient alternating direction algorithm is developed for numerical computation. Simulation results show that CD-trace compares favorably to gCoda and that it is better than sparse inverse covariance estimation for ecological association inference (SPIEC-EASI) (hereinafter S-E) in network recovery with compositional data. Finally, we test CD-trace and compare it to the other methods noted above using mouse skin microbiome data. AVAILABILITY AND IMPLEMENTATION: The CD-trace is open source and freely available from https://github.com/coamo2/CD-trace under GNU LGPL v3. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huili Yuan, Shun He, Minghua Deng |
Bioinform. | 3 |
| 2019 | Predicting protein inter-residue contacts using composite likelihood maximization and deep learningabstractBACKGROUND: Accurate prediction of inter-residue contacts of a protein is important to calculating its tertiary structure. Analysis of co-evolutionary events among residues has been proved effective in inferring inter-residue contacts. The Markov random field (MRF) technique, although being widely used for contact prediction, suffers from the following dilemma: the actual likelihood function of MRF is accurate but time-consuming to calculate; in contrast, approximations to the actual likelihood, say pseudo-likelihood, are efficient to calculate but inaccurate. Thus, how to achieve both accuracy and efficiency simultaneously remains a challenge. RESULTS: In this study, we present such an approach (called clmDCA) for contact prediction. Unlike plmDCA using pseudo-likelihood, i.e., the product of conditional probability of individual residues, our approach uses composite-likelihood, i.e., the product of conditional probability of all residue pairs. Composite likelihood has been theoretically proved as a better approximation to the actual likelihood function than pseudo-likelihood. Meanwhile, composite likelihood is still efficient to maximize, thus ensuring the efficiency of clmDCA. We present comprehensive experiments on popular benchmark datasets, including PSICOV dataset and CASP-11 dataset, to show that: i) clmDCA alone outperforms the existing MRF-based approaches in prediction accuracy. ii) When equipped with deep learning technique for refinement, the prediction accuracy of clmDCA was further significantly improved, suggesting the suitability of clmDCA for subsequent refinement procedure. We further present a successful application of the predicted contacts to accurately build tertiary structures for proteins in the PSICOV dataset. CONCLUSIONS: Composite likelihood maximization algorithm can efficiently estimate the parameters of Markov Random Fields and can improve the prediction accuracy of protein inter-residue contacts. Haicang Zhang, Fusong Ju, Jianwei Zhu, Yujuan Gao, Ziwei Xie, Minghua Deng, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 7 |
| 2019 | Correction to: Predicting protein inter-residue contacts using composite likelihood maximization and deep learningabstractFollowing publication of the original article [1], the author explained that there are several errors in the original article. Haicang Zhang, Fusong Ju, Jianwei Zhu, Yujuan Gao, Ziwei Xie, Minghua Deng, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 7 |
| 2019 | Network Clustering Analysis Using Mixture Exponential-Family Random Graph Models and Its Application in Genetic Interaction DataabstractMOTIVATION: Epistatic miniarrary profile (EMAP) studies have enabled the mapping of large-scale genetic interaction networks and generated large amounts of data in model organisms. It provides an incredible set of molecular tools and advanced technologies that should be efficiently understanding the relationship between the genotypes and phenotypes of individuals. However, the network information gained from EMAP cannot be fully exploited using the traditional statistical network models. Because the genetic network is always heterogeneous, for example, the network structure features for one subset of nodes are different from those of the left nodes. Exponential-family random graph models (ERGMs) are a family of statistical models, which provide a principled and flexible way to describe the structural features (e.g., the density, centrality, and assortativity) of an observed network. However, the single ERGM is not enough to capture this heterogeneity of networks. In this paper, we consider a mixture ERGM (MixtureEGRM) networks, which model a network with several communities, where each community is described by a single EGRM. RESULTS: EM algorithm is a classical method to solve the mixture problem, however, it will be very slow when the data size is huge in the numerous applications. We adopt an efficient novel online graph clustering algorithm to classify the graph nodes and estimate the ERGM parameters for the MixtureERGM. In comparison studies, the MixtureERGM outperforms the role analysis for the network cluster in which the mixture of exponential-family random graph model is developed for many ego-network according to their roles. One genetic interaction network of yeast and two real social networks (provided as supplemental materials, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/10.1109/TCBB.2017.2743711) show the wide potential application of the MixtureERGM. Yishu Wang 0003, Huaying Fang, Dejie Yang, Hongyu Zhao 0003, Minghua Deng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2018 | RaptorX-Angle: real-value prediction of protein backbone dihedral angles through a hybrid method of clustering and deep learningabstractBACKGROUND: Protein dihedral angles provide a detailed description of protein local conformation. Predicted dihedral angles can be used to narrow down the conformational space of the whole polypeptide chain significantly, thus aiding protein tertiary structure prediction. However, direct angle prediction from sequence alone is challenging. RESULTS: In this article, we present a novel method (named RaptorX-Angle) to predict real-valued angles by combining clustering and deep learning. Tested on a subset of PDB25 and the targets in the latest two Critical Assessment of protein Structure Prediction (CASP), our method outperforms the existing state-of-art method SPIDER2 in terms of Pearson Correlation Coefficient (PCC) and Mean Absolute Error (MAE). Our result also shows approximately linear relationship between the real prediction errors and our estimated bounds. That is, the real prediction error can be well approximated by our estimated bounds. CONCLUSIONS: Our study provides an alternative and more accurate prediction of dihedral angles, which may facilitate protein structure prediction and functional study. Yujuan Gao, Sheng Wang 0001, Minghua Deng, Jinbo Xu |
BMC Bioinform. | 3 |
| 2017 | VCNet: vector-based gene co-expression network construction and its application to RNA-seq dataabstractMOTIVATION: Building gene co-expression network (GCN) from gene expression data is an important field of bioinformatic research. Nowadays, RNA-seq data provides high dimensional information to quantify gene expressions in term of read counts for individual exons of genes. Such an increase in the dimension of expression data during the transition from microarray to RNA-seq era made many previous co-expression analysis algorithms based on simple univariate correlation no longer applicable. Recently, two vector-based methods, SpliceNet and RNASeqNet, have been proposed to build GCN. However, they failed to work when sample size is less than the number of exons. RESULTS: We develop an algorithm called VCNet to construct GCN from RNA-seq data to overcome this dimensional problem. VCNet performs a new statistical hypothesis test based on the correlation matrix of a gene-gene pair using the Frobenius norm. The asymptotic distribution of the new test is obtained under the null model. Simulation studies demonstrate that VCNet outperforms SpliceNet and RNASeqNet for detecting edges of GCN. We also apply VCNet to two expression datasets from TCGA database: the normal breast tissue and kidney tumour tissue, and the results show that the GCNs constructed by VCNet contain more biologically meaningful interactions than existing methods. CONCLUSION: VCNet is a useful tool to construct co-expression network. AVAILABILITY AND IMPLEMENTATION: VCNet is open source and freely available from https://github.com/wangzengmiao/VCNet under GNU LGPL v3. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zengmiao Wang, Huaying Fang, Nelson L. S. Tang, Minghua Deng |
Bioinform. | 4 |
| 2017 | SVmine improves structural variation detection by integrative mining of predictions from multiple algorithmsabstractMOTIVATION: Structural variation (SV) is an important class of genomic variations in human genomes. A number of SV detection algorithms based on high-throughput sequencing data have been developed, but they have various and often limited level of sensitivity, specificity and breakpoint resolution. Furthermore, since overlaps between predictions of algorithms are low, SV detection based on multiple algorithms, an often-used strategy in real applications, has little effect in improving the performance of SV detection. RESULTS: We develop a computational tool called SVmine for further mining of SV predictions from multiple tools to improve the performance of SV detection. SVmine refines SV predictions by performing local realignment and assess quality of SV predictions based on likelihoods of the realignments. The local realignment is performed against a set of sequences constructed from the reference sequence near the candidate SV by incorporating nearby single nucleotide variations, insertions and deletions. A sandwich alignment algorithm is further used to improve the accuracy of breakpoint positions. We evaluate SVmine on a set of simulated data and real data and find that SVmine has superior sensitivity, specificity and breakpoint estimation accuracy. We also find that SVmine can significantly improve overlaps of SV predictions from other algorithms. AVAILABILITY AND IMPLEMENTATION: SVmine is available at https://github.com/xyc0813/SVmine. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuchao Xia, Minghua Deng, Ruibin Xi |
Bioinform. | 3 |
| 2017 | Pysim-sv: a package for simulating structural variation data with GC-biasesabstractBACKGROUND: Structural variations (SVs) are wide-spread in human genomes and may have important implications in disease-related and evolutionary studies. High-throughput sequencing (HTS) has become a major platform for SV detection and simulation serves as a powerful and cost-effective approach for benchmarking SV detection algorithms. Accurate performance assessment by simulation requires the simulator capable of generating simulation data with all important features of real data, such GC biases in HTS data and various complexities in tumor data. However, no available package has systematically addressed all issues in data simulation for SV benchmarking. RESULTS: Pysim-sv is a package for simulating HTS data to evaluate performance of SV detection algorithms. Pysim-sv can introduce a wide spectrum of germline and somatic genomic variations. The package contains functionalities to simulate tumor data with aneuploidy and heterogeneous subclones, which is very useful in assessing algorithm performance in tumor studies. Furthermore, Pysim-sv can introduce GC-bias, the most important and prevalent bias in HTS data, in the simulated HTS data. CONCLUSIONS: Pysim-sv provides an unbiased toolkit for evaluating HTS-based SV detection algorithms. Yuchao Xia, Minghua Deng, Ruibin Xi |
BMC Bioinform. | 3 |
| 2016 | Inference of Markovian properties of molecular sequences from NGS data and applications to comparative genomicsabstractMOTIVATION: Next-generation sequencing (NGS) technologies generate large amounts of short read data for many different organisms. The fact that NGS reads are generally short makes it challenging to assemble the reads and reconstruct the original genome sequence. For clustering genomes using such NGS data, word-count based alignment-free sequence comparison is a promising approach, but for this approach, the underlying expected word counts are essential.A plausible model for this underlying distribution of word counts is given through modeling the DNA sequence as a Markov chain (MC). For single long sequences, efficient statistics are available to estimate the order of MCs and the transition probability matrix for the sequences. As NGS data do not provide a single long sequence, inference methods on Markovian properties of sequences based on single long sequences cannot be directly used for NGS short read data. RESULTS: Here we derive a normal approximation for such word counts. We also show that the traditional Chi-square statistic has an approximate gamma distribution ,: using the Lander-Waterman model for physical mapping. We propose several methods to estimate the order of the MC based on NGS reads and evaluate those using simulations. We illustrate the applications of our results by clustering genomic sequences of several vertebrate and tree species based on NGS reads using alignment-free sequence dissimilarity measures. We find that the estimated order of the MC has a considerable effect on the clustering results ,: and that the clustering results that use a N: MC of the estimated order give a plausible clustering of the species. AVAILABILITY AND IMPLEMENTATION: Our implementation of the statistics developed here is available as R package 'NGS.MC' at http://www-rcf.usc.edu/∼fsun/Programs/NGS-MC/NGS-MC.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jie Ren 0006, Minghua Deng, Gesine Reinert, Charles H. Cannon, Fengzhu Sun |
Bioinform. | 3 |
| 2015 | CCLasso: correlation inference for compositional data through LassoabstractMOTIVATION: Direct analysis of microbial communities in the environment and human body has become more convenient and reliable owing to the advancements of high-throughput sequencing techniques for 16S rRNA gene profiling. Inferring the correlation relationship among members of microbial communities is of fundamental importance for genomic survey study. Traditional Pearson correlation analysis treating the observed data as absolute abundances of the microbes may lead to spurious results because the data only represent relative abundances. Special care and appropriate methods are required prior to correlation analysis for these compositional data. RESULTS: In this article, we first discuss the correlation definition of latent variables for compositional data. We then propose a novel method called CCLasso based on least squares with [Formula: see text] penalty to infer the correlation network for latent variables of compositional data from metagenomic data. An effective alternating direction algorithm from augmented Lagrangian method is used to solve the optimization problem. The simulation results show that CCLasso outperforms existing methods, e.g. SparCC, in edge recovery for compositional data. It also compares well with SparCC in estimating correlation network of microbe species from the Human Microbiome Project. AVAILABILITY AND IMPLEMENTATION: CCLasso is open source and freely available from https://github.com/huayingfang/CCLasso under GNU LGPL v3. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huaying Fang, Chengcheng Huang, Hongyu Zhao 0003, Minghua Deng |
Bioinform. | 4 |
| 2014 | New developments of alignment-free sequence comparison: measures, statistics and next-generation sequencingabstractWith the development of next-generation sequencing (NGS) technologies, a large amount of short read data has been generated. Assembly of these short reads can be challenging for genomes and metagenomes without template sequences, making alignment-based genome sequence comparison difficult. In addition, sequence reads from NGS can come from different regions of various genomes and they may not be alignable. Sequence signature-based methods for genome comparison based on the frequencies of word patterns in genomes and metagenomes can potentially be useful for the analysis of short reads data from NGS. Here we review the recent development of alignment-free genome and metagenome comparison based on the frequencies of word patterns with emphasis on the dissimilarity measures between sequences, the statistical power of these measures when two sequences are related and the applications of these measures to NGS data. Jie Ren 0006, Gesine Reinert, Minghua Deng, Michael S. Waterman, Fengzhu Sun |
Briefings Bioinform. | 4 |
| 2013 | Multiple alignment-free sequence comparisonabstractMOTIVATION: Recently, a range of new statistics have become available for the alignment-free comparison of two sequences based on k-tuple word content. Here, we extend these statistics to the simultaneous comparison of more than two sequences. Our suite of statistics contains, first, C(*)1 and C(S)1, extensions of statistics for pairwise comparison of the joint k-tuple content of all the sequences, and second, C(*)2, C(S)2 and C(geo)2, averages of sums of pairwise comparison statistics. The two tasks we consider are, first, to identify sequences that are similar to a set of target sequences, and, second, to measure the similarity within a set of sequences. RESULTS: Our investigation uses both simulated data as well as cis-regulatory module data where the task is to identify cis-regulatory modules with similar transcription factor binding sites. We find that although for real data, all of our statistics show a similar performance, on simulated data the Shepp-type statistics are in some instances outperformed by star-type statistics. The multiple alignment-free statistics are more sensitive to contamination in the data than the pairwise average statistics. AVAILABILITY: Our implementation of the five statistics is available as R package named 'multiAlignFree' at be http://www-rcf.usc.edu/∼fsun/Programs/multiAlignFree/multiAlignFreemain.html. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jie Ren 0006, Fengzhu Sun, Minghua Deng, Gesine Reinert |
Bioinform. | 4 |
| 2012 | Alignment-Free Sequence Comparison Based on Next Generation Sequencing Reads: Extended Abstract
Jie Ren 0006, Zhiyuan Zhai, Minghua Deng, Fengzhu Sun |
RECOMB | 5 |
| 2011 | Modular analysis of the probabilistic genetic interaction networkabstractMOTIVATION: Epistatic Miniarray Profiles (EMAP) has enabled the mapping of large-scale genetic interaction networks; however, the quantitative information gained from EMAP cannot be fully exploited since the data are usually interpreted as a discrete network based on an arbitrary hard threshold. To address such limitations, we adopted a mixture modeling procedure to construct a probabilistic genetic interaction network and then implemented a Bayesian approach to identify densely interacting modules in the probabilistic network. RESULTS: Mixture modeling has been demonstrated as an effective soft-threshold technique of EMAP measures. The Bayesian approach was applied to an EMAP dataset studying the early secretory pathway in Saccharomyces cerevisiae. Twenty-seven modules were identified, and 14 of those were enriched by gold standard functional gene sets. We also conducted a detailed comparison with state-of-the-art algorithms, hierarchical cluster and Markov clustering. The experimental results show that the Bayesian approach outperforms others in efficiently recovering biologically significant modules. Minping Qian, Dong Li 0018, Chao Tang 0003, Minghua Deng |
Bioinform. | 7 |
| 2011 | A Lasso regression model for the construction of microRNA-target regulatory networksabstractMOTIVATION: MicroRNAs have recently emerged as a major class of regulatory molecules involved in a broad range of biological processes and complex diseases. Construction of miRNA-target regulatory networks can provide useful information for the study and diagnosis of complex diseases. Many sequence-based and evolutionary information-based methods have been developed to identify miRNA-mRNA targeting relationships. However, as the amount of available miRNA and gene expression data grows, a more statistical and systematic method combining sequence-based binding predictions and expression-based correlation data becomes necessary for the accurate identification of miRNA-mRNA pairs. RESULTS: We propose a Lasso regression model for the identification of miRNA-mRNA targeting relationships that combines sequence-based prediction information, miRNA co-regulation, RISC availability and miRNA/mRNA abundance data. By comparing this modelling approach with two other known methods applied to three different datasets, we found that the Lasso regression model has considerable advantages in both sensitivity and specificity. The regression coefficients in the model can be used to determine the true regulatory efficacies in tissues and was demonstrated using the miRNA target site type data. Finally, by constructing the miRNA regulatory networks in two stages of prostate cancer (PCa), we found the several significant miRNA-hubbed network modules associated with PCa metastasis. In conclusion, the Lasso regression model is a robust and informative tool for constructing the miRNA regulatory networks for diagnosis and treatment of complex diseases. AVAILABILITY: The R program for predicting miRNA-mRNA targeting relationships using the Lasso regression model is freely available, along with the described datasets and resulting regulatory network, at http://biocompute.bmi.ac.cn/CZlab/alarmnet/. The source code is open for modification and application to other miRNA/mRNA expression datasets. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wubin Qu, Minghua Deng, Chenggang Zhang |
Bioinform. | 4 |
| 2009 | Searching for bidirectional promoters in Arabidopsis thalianaabstractBACKGROUND: A "bidirectional gene pair" is defined as two adjacent genes which are located on opposite strands of DNA with transcription start sites (TSSs) not more than 1000 base pairs apart and the intergenic region between two TSSs is commonly designated as a putative "bidirectional promoter". Individual examples of bidirectional gene pairs have been reported for years, as well as a few genome-wide analyses have been studied in mammalian and human genomes. However, no genome-wide analysis of bidirectional genes for plants has been done. Furthermore, the exact mechanism of this gene organization is still less understood. RESULTS: We conducted comprehensive analysis of bidirectional gene pairs through the whole Arabidopsis thaliana genome and identified 2471 bidirectional gene pairs. The analysis shows that bidirectional genes are often coexpressed and tend to be involved in the same biological function. Furthermore, bidirectional gene pairs associated with similar functions seem to have stronger expression correlation. We pay more attention to the regulatory analysis on the intergenic regions between bidirectional genes. Using a hierarchical stochastic language model (HSL) (which is developed by ourselves), we can identify intergenic regions enriched of regulatory elements which are essential for the initiation of transcription. Finally, we picked 27 functionally associated bidirectional gene pairs with their intergenic regions enriched of regulatory elements and hypothesized them to be regulated by bidirectional promoters, some of which have the same orthologs in ancient organisms. More than half of these bidirectional gene pairs are further supported by sharing similar functional categories as these of handful experimental verified bidirectional genes. CONCLUSION: Bidirectional gene pairs are concluded also prevalent in plant genome. Promoter analyses of the intergenic regions between bidirectional genes could be a new way to study the bidirectional gene structure, which may provide a important clue for further analysis. Such a method could be applied to other genomes. Quan Wang 0004, Dayong Li, Lihuang Zhu, Minping Qian, Minghua Deng |
BMC Bioinform. | 6 |
| 2009 | REMAS: a new regression model to identify alternative splicing events from exon array dataabstractBACKGROUND: Alternative splicing (AS) is an important regulatory mechanism for gene expression and protein diversity in eukaryotes. Previous studies have demonstrated that it can be causative for, or specific to splicing-related diseases. Understanding the regulation of AS will be helpful for diagnostic efforts and drug discoveries on those splicing-related diseases. As a novel exon-centric microarray platform, exon array enables a comprehensive analysis of AS by investigating the expression of known and predicted exons. Identifying of AS events from exon array has raised much attention, however, new and powerful algorithms for exon array data analysis are still absent till now. RESULTS: Here, we considered identifying of AS events in the framework of variable selection and developed a regression method for AS detection (REMAS). Firstly, features of alternatively spliced exons were scaled by reasonably defined variables. Secondly, we designed a hierarchical model which can represent gene structure and transcriptional influence to exons, and the lasso type penalties were introduced in calculation because of huge variable size. Thirdly, an iterative two-step algorithm was developed to select alternatively spliced genes and exons. To avoid negative effects introduced by small sample size, we ranked genes as parameters indicating their AS capabilities in an iterative manner. After that, both simulation and real data evaluation showed that REMAS could efficiently identify potential AS events, some of which had been validated by RT-PCR or supported by literature evidence. CONCLUSION: As a new lasso regression algorithm based on hierarchical model, REMAS has been demonstrated as a reliable and effective method to identify AS events from exon array data. Xingyi Hang, Minping Qian, Wubin Qu, Chenggang Zhang, Minghua Deng |
BMC Bioinform. | 7 |
| 2006 | Using a Stochastic AdaBoost Algorithm to Discover Interactome Motif Pairs from Sequences
Huan Yu 0012, Minping Qian, Minghua Deng |
ICIC (3) | 3 |
| 2006 | Nonequilibrium Model for Yeast Cell Cycle
Huan Yu 0012, Minghua Deng, Minping Qian |
ICIC (3) | 3 |
| 2006 | An integrated approach to the prediction of domain-domain interactionsabstractBACKGROUND: The development of high-throughput technologies has produced several large scale protein interaction data sets for multiple species, and significant efforts have been made to analyze the data sets in order to understand protein activities. Considering that the basic units of protein interactions are domain interactions, it is crucial to understand protein interactions at the level of the domains. The availability of many diverse biological data sets provides an opportunity to discover the underlying domain interactions within protein interactions through an integration of these biological data sets. RESULTS: We combine protein interaction data sets from multiple species, molecular sequences, and gene ontology to construct a set of high-confidence domain-domain interactions. First, we propose a new measure, the expected number of interactions for each pair of domains, to score domain interactions based on protein interaction data in one species and show that it has similar performance as the E-value defined by Riley et al. Our new measure is applied to the protein interaction data sets from yeast, worm, fruitfly and humans. Second, information on pairs of domains that coexist in known proteins and on pairs of domains with the same gene ontology function annotations are incorporated to construct a high-confidence set of domain-domain interactions using a Bayesian approach. Finally, we evaluate the set of domain-domain interactions by comparing predicted domain interactions with those defined in iPfam database that were derived based on protein structures. The accuracy of predicted domain interactions are also confirmed by comparing with experimentally obtained domain interactions from H. pylori. As a result, a total of 2,391 high-confidence domain interactions are obtained and these domain interactions are used to unravel detailed protein and domain interactions in several protein complexes. CONCLUSION: Our study shows that integration of multiple biological data sets based on the Bayesian approach provides a reliable framework to predict domain interactions. By integrating multiple data sources, the coverage and accuracy of predicted domain interactions can be significantly increased. Minghua Deng, Fengzhu Sun, Ting Chen 0006 |
BMC Bioinform. | 2 |
| 2004 | An Improvement for Eulerian Path of Assembly
Shenglan Peng, Minghua Deng, Minping Qian |
SNPD | 2 |
| 2004 | Mapping gene ontology to proteins based on protein-protein interaction dataabstractMOTIVATION: Gene Ontology (GO) consortium provides structural description of protein function that is used as a common language for gene annotation in many organisms. Large-scale techniques have generated many valuable protein-protein interaction datasets that are useful for the study of protein function. Combining both GO and protein-protein interaction data allows the prediction of function for unknown proteins. RESULT: We apply a Markov random field method to the prediction of yeast protein function based on multiple protein-protein interaction datasets. We assign function to unknown proteins with a probability representing the confidence of this prediction. The functions are based on three general categories of cellular component, molecular function and biological process defined in GO. The yeast proteins are defined in the Saccharomyces Genome Database (SGD). The protein-protein interaction datasets are obtained from the Munich Information Center for Protein Sequences (MIPS), including physical interactions and genetic interactions. The efficiency of our prediction is measured by applying the leave-one-out validation procedure to a functional path matching scheme, which compares the prediction with the GO description of a protein's function from the abstract level to the detailed level along the GO structure. For biological process, the leave-one-out validation procedure shows 52% precision and recall of our method, much better than that of the simple guilty-by-association methods. Minghua Deng, Zhidong Tu, Fengzhu Sun, Ting Chen 0006 |
Bioinform. | 1 |
| 2003 | An integrated probabilistic model for functional prediction of proteinsabstractWe develop an integrated probabilistic model to combine protein physical interactions, genetic interactions, highly correlated gene expression network, protein complex data, and domain structures of individual proteins to predict protein functions. The model is an extension of our previous model for protein function prediction based on Markovian random field theory. The model is flexible in that other protein pairwise relationship information and features of individual proteins can be easily incorporated. Two features distinguish the integrated approach from other available methods for protein function prediction. One is that the integrated approach uses all available sources of information with different weights for different sources of data. It is a global approach that takes the whole network into consideration. The second feature is that the posterior probability that a protein has the function of interest is assigned. The posterior probability indicates how confident we are about assigning the function to the protein. We apply our integrated approach to predict functions of yeast proteins based upon MIPS protein function classifications and upon the interaction networks based on MIPS physical and genetic interactions, gene expression profiles, Tandem Affinity Purification (TAP) protein complex data, and protein domain information. We study the sensitivity and specificity of the integrated approach using different sources of information by the leave-one-out approach. In contrast to using MIPS physical interactions only, the integrated approach combining all of the information increases the sensitivity from 57% to 87% when the specificity is set at 57%-an increase of 30%. It should also be noted that enlarging the interaction network greatly increases the number of proteins whose functions can be predicted. Minghua Deng, Ting Chen 0006, Fengzhu Sun |
RECOMB | 1 |
| 2002 | Inferring domain-domain interactions from protein-protein interactionsabstractProtein-protein interactions are important events in cellular and biochemical processes within a cell. Several researchers have undertaken the task of analyzing protein-protein interactions covering all genes of an organism by using yeast two-hybrid assays. Protein-protein interactions involve physical interactions between protein domains. Therefore, understanding protein interactions at the domain level gives a global view of the protein interaction network, and possibly extends functions of proteins. In this study, we present a Maximum Likelihood approach to infer domain-domain interactions from the 5719 yeast protein-protein interactions obtained in the high throughput two-hybrid experiments by Uetz et al., 2000 and Ito et al., 2001. The accuracies of our predictions are measured at the protein level. Our study includes the following three results: (1) using the inferred domain-domain interactions, we predict interactions between proteins and achieve 39.0% specificity and 79.7% sensitivity; (2) our predicted protein-protein interactions have a significant overlap with the MIPS(http://mips.gfs.de) protein-protein interactions obtained by methods other than the two-hybrid systems; and (3) the mean correlation coefficient of the gene expression profiles for our predicted interacting pairs is significantly higher than that for random pairs as well as that of interacting pairs in Uetz's and Ito's experimental data. Our method has shown robustness in analyzing incomplete data sets and dealing with various experimental errors. We find several novel protein-protein interactions such as RPS0A interacting with APG17 and TAF40 interacting with SPT3, which are consistent with the functions of the proteins. Minghua Deng, Shipra Mehta, Fengzhu Sun, Ting Chen 0006 |
RECOMB | 1 |
| 2000 | A Novel Algorithm for Handwritten Chinese Character Recognition
Minghua Deng, Minping Qian, Xueqing Zhu |
ICMI | 2 |