EDBT 2026 Demo / reviewers in the wild / expert
Juan Liu 0007
dblp:16/3621-7 · also Juliet Juan Liu
· DBLP profile ↗
82ranked-venue papers
0as first author
49since 2021 · last 2026
0000-0001-9344-7415ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 55 · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 16 since 2021Artificial intelligence and machine learning · 15 · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Geometric Equivariant Network for Full-Atom Antibody DesignabstractAntibody design is critically important in biomedical and therapeutic contexts but remains extremely challenging due to the complexity of antibody sequence–structure relationships and stringent antigen specificity requirements. Traditional computational approaches rely on multi-stage pipelines and often overlook full-atom details (e.g., side-chain conformations) as well as fine-grained geometric features, resulting in limited effectiveness. To overcome these limitations, we propose Dynamic Geometric Equivariant Network (DGENet), an end-to-end full-atom antibody design model that integrates a geometric-kinematic equivariant dynamic optimization module (GK-EDO) with an full-atom E(3)-equivariant message-passing architecture. This framework enables iterative optimization of antibody structures under explicit geometric and kinematic constraints, generating complete antibody structures (including backbone and side chains) and simultaneously jointly optimizing the sequences and 3D structures of the complementarity-determining regions (CDRs). DGENet also introduces a novel virtual anchor docking mechanism that employs an adaptive PNet-Kabsch module to explicitly guide antibody–antigen binding and achieve precise bound conformations. Evaluations on multiple benchmark datasets demonstrate that DGENet exhibits outstanding performance in antibody structure and sequence generation as well as in designing high-affinity antibodies, underscoring its reliability as an advanced antibody design model. Qiang Zhang 0031, Juan Liu 0007 |
AAAI | 4 |
| 2026 | Domain-Aware Multi-View Contrastive Representation Learning for Protein Subcellular Localization PredictionabstractProtein subcellular localization prediction is essential for understanding protein function and cellular organization. However, existing methods exhibit two major limitations: (1) they overlook the critical role of evolutionarily conserved protein domains, which are fundamental functional and structural units that significantly influence functions and subcellular localization, and (2) they rarely learn residue order and backbone coordinates simultaneously, neglecting the complementary information inherent in multi-modal representations. In this paper, we propose a novel Domain-Aware Multi-View Contrastive Representation Learning for Protein Subcellular Localization prediction, named DMVCL. Firstly, it devises domain-sequence/structure attention modules, which identify functionally significant regions in protein structures/sequences that critically determine subcellular localization. Secondly, it introduces a multi-view contrastive learning framework that unites inter-view and intra-view objectives. Inter-view contrastive learning aligns protein sequences with their corresponding structures by maximizing mutual information, thereby capturing the consistency of protein residue order and backbone coordinates. Intra-view contrastive learning enhances the representation discriminability of each modality by explicitly separating proteins with no common location and attracting those with any shared localization. Extensive experiments demonstrate that DMVCL significantly outperforms existing baselines. Ablation studies and visualizations further highlight the contributions of domain-sequence/structure attention and multi-view contrastive learning in achieving superior predictive performance. Qiang Zhang 0031, Jing Feng 0005, Juan Liu 0007 |
AAAI | 5 |
| 2026 | A Dependency-Aware Generative Model with Bi-decoders for Super-Resolving Spatial Transcriptome Data from Histology Images
Yuqi Chen 0007, Peng Jiang 0025, Juan Liu 0007 |
ISBRA (1) | 4 |
| 2026 | Multimodal laryngoscopic video analysis for assisted diagnosis of vocal fold paralysis
Yucong Zhang, Jinshan Yang, Juan Liu 0007, Faya Liang, Ming Li 0026 |
Comput. Speech Lang. | 5 |
| 2025 | CrossStateECG-Lite: Lightweight Adaptive Thresholding Network for Dual-State ECG BiometricsabstractECG-based biometric identification faces significant challenges in cross-state scenarios where physiological states differ between enrollment and authentication. Existing methods typically assume identical states, leading to substantial performance degradation when applied to real-world dual-state conditions. We propose CrossStateECG-Lite, an adaptive thresholding attention network specifically designed for robust cross-state ECG identification. The network integrates multi-scale feature learning, deep residual connections, and efficient attention mechanisms to capture state-invariant cardiac patterns. To address the inherent uncertainty in cross-state identification, we develop a personalized decision-making mechanism using Bayesian adaptive thresholding, which customizes authentication boundaries based on individual physiological variability. Evaluated on 45 subjects with paired rest-exercise recordings, CrossStateECG- Lite achieves balanced accuracies of 95.26 % (rest-to-post-exercise) and 96.05% (post-exercise-to-rest). The lightweight architecture with only 797K parameters demonstrates that efficient design combined with personalized decision-making provides an effective solution for real-world ECG biometrics where physiological states cannot be controlled. Comprehensive ablation studies confirm the synergistic contributions of the architectural components, with deep convolutional layers being particularly effective in learning state-invariant representations. Dan Zheng, Jing Feng 0005, Juan Liu 0007 |
BIBM | 3 |
| 2025 | Lip Enhancement and Multi-View Simulation for Robust Visual Speech Recognition in MAVSR 2025abstractIn this paper, we present our work for Visual Speech Recognition (VSR) in the Mandarin Audio-Visual Speech Recognition (MAVSR) Challenge 2025, with a particular focus on improving lipreading under challenging visual conditions. The proposed system leverages cross-modal knowledge transfer and employs a progressive training strategy based on large-scale speech and visual speech datasets. Furthermore, we introduce LIPER, a visual enhancement module designed to generate improved lip-region visual data under conditions such as low resolution, poor illumination, and color distortion. LIPER further facilitates the synthesis of multi-view lip movements through lip pose estimation and 3D reconstruction. These enhancements significantly improve the robustness of the VSR system under low-quality visual conditions. Experimental results show that the proposed approach achieves relative character error rate (CER) reductions of 16.1% on the MOV20-Test set, compared to the official baseline system in track 1, and achieves second place among submitted systems in the challenge. The code is available at https://github.com/yaku122/RVSR. Cancan Li, Juan Liu 0007 |
FG | 3 |
| 2025 | Align-AV-HuBERT: AV-HuBERT with Audio-Visual Temporal AlignmentabstractLip movement and acoustic information complement each other in speech recognition tasks, and audio-visual models have significantly advanced in recent years. However, most existing audio-visual models assume that audio and video inputs are temporally aligned. In practical applications, there are often scenarios where time deviations occur between audio and video, leading to a substantial decline in the performance of previous models when alignment is disrupted. In this paper, we propose Align-AV-HuBERT, a model based on Audio-Visual Hidden Unit BERT (AV-HuBERT) that simultaneously achieves audio-visual temporal alignment and self-supervised learning by training with both alignment loss and masked prediction loss. When evaluated on aligned audio and video datasets, Align-AV-HuBERT matches with the AV-HuBERT baseline. Moreover, when tested on unaligned audio and video datasets, Align-AV-HuBERT maintains stable performance while the Word Error Rate (WER) of the AV-HuBERT baseline increases dramatically. Our code and models are available at https://github.com/lilycalre/Align-AV-HuBERT. Cancan Li, Juan Liu 0007 |
ICME | 3 |
| 2025 | Enhanced Self-Supervised Multi-View Representations with Modality-Missing Robustness for Audio-Visual Speech RecognitionabstractAudio-Visual Speech Recognition (AVSR) leverages visual information to enhance speech understanding. However, current models assume stable, frontal viewpoints, suffering significant performance drops with non-frontal angles or when visual input is missing. Our approach employs a multi-view data generation strategy using 3D head avatar reconstruction, synthesizing viewpoint-diverse data to handle varying head poses. We introduce a self-supervised multi-view representation learning model (MVL), ensuring viewpoint-invariant and domain-agnostic embeddings. Moreover, a Unified Modality Adapter (UMA) enables the model to revert to audio-only performance levels when visual inputs are unavailable. Experimental results demonstrate that, compared with the baseline AV-HuBERT model, our approach improves lip-reading accuracy by 3.0% under large pose deviations and yields a 12.3% overall gain in AVSR performance. Moreover, the system outperforms the baseline on real data, confirming the generalizability of our approach under challenging multi-view and modality-missing scenarios. Code and Data: https://yakumostudio.github.io/yakumo.github.io/mvss/mvss.html. Cancan Li, Juan Liu 0007 |
ICME | 3 |
| 2025 | Multi-scale Scanning Network for Machine Anomalous Sound Detection
Yucong Zhang, Juan Liu 0007, Ming Li 0026 |
ICONIP (3) | 2 |
| 2025 | Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge
Ming Cheng 0005, Cancan Li, Juan Liu 0007, Ming Li 0026 |
INTERSPEECH | 4 |
| 2025 | SMIIP-NV: A Multi-Annotation Non-Verbal Expressive Speech Corpus in Mandarin for LLM-Based Speech SynthesisabstractIn natural language communication, emotions are often conveyed through non-verbal sounds (NVs), such as laughter, crying, cough and so on. However, most existing text-to-speech (TTS) corpora lack annotations for these non-verbal sounds, leading to a scarcity of systems capable of generating them. To address this gap, we introduce SMIIP-NV, a non-verbal speech synthesis corpus annotated with both emotions and non-verbal sounds, including laughter, crying, and cough. To the best of our knowledge, SMIIP-NV is the largest publicly available open-source expressive speech corpus that includes non-verbal speech and rich annotations. It comprises 33 hours of speech data, covering five distinct emotions and three types of non-verbal sounds, with detailed transcriptions and precise timestamps for each occurrence of non-verbal sounds. Additionally, the corpus provides annotations for speech segments that contain laughter or crying. To demonstrate the utility of this dataset, we establish a baseline for non-verbal speech synthesis by employing a lightweight large language model (LLM). The SMIIP-NV dataset and static audio demonstrations are publicly available at https://axunyii.github.io/SMIIP-NV. The interactive real-time demonstrations can be accessed at https://huggingface.co/spaces/xunyi/SMIIP-NV_Finetuned_CosyVoice2. Zhuojun Wu, Dong Liu 0028, Juan Liu 0007, Yechen Wang, Hui Bu, Pengyuan Zhang, Ming Li 0026 |
ACM Multimedia | 3 |
| 2025 | CrossStateECG: Multi-scale Deep Convolutional Network with Attention for Rest-Exercise ECG Biometrics
Dan Zheng, Jing Feng 0005, Juan Liu 0007 |
PRCV (15) | 3 |
| 2025 | Deconvolution of spatial transcriptomics data via graph contrastive learning and partial least square regressionabstractDeciphering the cellular abundance in spatial transcriptomics (ST) is crucial for revealing the spatial architecture of cellular heterogeneity within tissues. However, some of the current spatial sequencing technologies are in low resolutions, leading to each spot having multiple heterogeneous cells. Additionally, current spatial deconvolution methods lack the ability to utilize multi-modality information such as gene expression and chromatin accessibility from single-cell multi-omics data. In this study, we introduce a graph Contrastive Learning and Partial Least Squares regression-based method, CLPLS, to deconvolute ST data. CLPLS is a flexible method that it can be extended to integrate ST data and single-cell multi-omics data, enabling the exploration of the spatially epigenomic heterogeneity. We applied CLPLS to both simulated and real datasets coming from different platforms. Benchmark analyses with other methods on these datasets show the superior performance of CLPLS in deconvoluting spots in single cell level. Yuanyuan Mo, Juan Liu 0007 |
Briefings Bioinform. | 2 |
| 2025 | Predicting breast cancer molecular subtypes from H &E-stained histopathological images using a spatial-transcriptomics-based patch filter
Yuqi Chen 0007, Juan Liu 0007, Peng Jiang 0025, Dehua Cao |
Multim. Tools Appl. | 2 |
| 2025 | TRGOA: Topological-Aware Residue-Gene Ontology Attention Network for Protein Function PredictionabstractProtein function prediction is one of the most important biological problems in the field of bioinformatics. The functions of proteins are generally described by a series of Gene Ontology (GO) terms that have hierarchical relationships. Two factors hinder the effective prediction of protein functions using current methods: 1) they cannot well model and learn the topological semantic similarity between residues and GO terms, resulting in a huge semantic gap; 2) they predict the functions of proteins by calculating the semantic similarity between protein-level embeddings and GO terms, which does not effectively learn the protein-function relationship. To address the above issues, we propose the Topological-aware Residue-Gene Ontology Attention Network (TRGOA) for protein function prediction. First, a topological-aware attention module is designed to leverage attention scores within this joint semantic space allowing for modeling the fine-grained semantic similarity between residues and GO terms, thereby narrowing the semantic gap. Second, a multi-head aggregator is proposed, which adeptly captures the functions relevant fine-grained semantic similarity and filters out function-irrelevant components, which effectively reveal protein-function relationships, thereby enhancing generality and robustness. Finally, TRGOA has demonstrated promising outcomes, revealing our model can understand the protein-function relationship in deep insights. Qiang Zhang 0031, Jing Feng 0005, Juan Liu 0007 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | DCOAT: Dynamic Core-attachment based Protein Complex DetectionabstractIdentifying protein complexes from protein-protein interaction (PPI) networks is essential for understanding cellular organization and biological processes. Most of the existing works treat PPI networks as static and focus on detecting densely connected regions. Several recent works have explored the dynamics of PPIs in the network, but may identify complexes with false-positive edges of low confidence. In this paper, we propose a novel Dynamic COre-ATtachment based protein complex detection algorithm (DCOAT). First, we derive co-expressed gene biclusters from the gene expression data by a memetic algorithm to enable the construction of dynamic subnetworks from a static PPI network. Next, we evaluate the confidence of edges within each dynamic subnetwork by integrating both topological features and biological information to filter false-positive edges. Then, we detect complexes from each weighted subnetwork based on the core-attachment structure. Finally, we build an overlapping graph and merge redundant complexes using maximum weight matching. The experimental results validate that DCOAT achieves state-of-the-art performance compared to the nine baseline methods based on static and dynamic networks across multiple evaluation measures, especially in terms of F1and precision. Moreover, the detected stable and temporal complexes are consistent with real biological scenarios and have biological significance. Our code and datasets are publicly available at https://github.com/LiGuojing194/DCOAT. Yaoran Chen, Difu Feng, Yuanyuan Zhu 0001, Guojing Li, Ming Zhong 0002, Tieyun Qian, Juan Liu 0007 |
BIBM | 7 |
| 2024 | Deconvolution of spatial transcriptomics data based on multi-head dynamic GAT and optimal transportabstractThe majority of spatial transcriptomics datasets are characterized by low resolution, wherein each spot generally encompasses multiple cells. This limitation poses challenges for exploring biological insights at the cellular level. Consequently, the development and application of robust deconvolution methods for spatial transcriptomics data are imperative to address this challenge. Addressing the limitations of previous deconvolution methods—such as the lack of consideration cell type labels from single-cell sequencing data and the inability to adaptively capture local relationship among points—we propose a novel spatial transcriptomics data deconvolution model based on label-guided Multi-Head Dynamic Graph Attention Networks with Optimal Transport(MHDGATOT). Our approach leverages an advanced multi-head dynamic graph attention network to adaptively capture inter-data relationships and generate effective low-dimensional embeddings. Subsequently, we employ optimal transport based on fused gromov-wasserstein to derive the transport matrix between spatial transcriptomics data and single-cell sequencing data, facilitating the accurate deconvolution of spatial transcriptomics datasets. Experimental validation substantiates the effectiveness of our model. Yuqi Chen 0007, Peng Jiang 0025, Qiang Zhang 0031, Juan Liu 0007 |
BIBM | 6 |
| 2024 | Subgraph-based Self-Supervised Learning Framework for Enzymatic Reaction Feasibility PredictionabstractEnzymatic Reaction feasibility prediction is used to determine whether the reaction generated by computational methods can actually occur, which can effectively reduce the complexity of synthetic pathway design. Existing methods often use SMILES or molecular fingerprints to represent molecules, resulting in a lack of molecular structural information. Although some GNNs have been leveraged to solve this problem, the complexity of the intra and inter-substructures interactions of molecules in microorganisms makes it difficult for traditional GNNs to accurately model it. To address these problems, we propose a subgraph-based self-supervised learning framework to predict the feasibility of enzymatic reactions. Specifically, we first propose a subgraph-based two-branch graph neural network. This network leverages the atom graph and substructure graph of a molecule to thoroughly capture its structural and semantic information. Besides, a subgraph interaction module is designed to facilitate the full integration of features. Subsequently, we propose a domain knowledge-guided self-supervised learning task, utilizing molecular fingerprints and substructures to capture the consistency between them effectively. The experimental results show that the proposed method outperforms existing state-ofthe-art methods significantly on all datasets. Juan Liu 0007, Qiang Zhang 0031, Jianghang Liu, Guangsheng Wu |
BIBM | 2 |
| 2024 | A Dual-Path Framework with Frequency-and-Time Excited Network for Anomalous Sound DetectionabstractIn contrast to human speech, machine-generated sounds of the same type often exhibit consistent frequency characteristics and discernible temporal periodicity. However, leveraging these dual attributes in anomaly detection remains relatively under-explored. In this paper, we propose an automated dual-path framework that learns prominent frequency and temporal patterns for diverse machine types. One pathway uses a novel Frequency-and-Time Excited Network (FTE-Net) to learn the salient features across frequency and time axes of the spectrogram. It incorporates a Frequency-and-Time Chunkwise Encoder (FTC-Encoder) and an excitation network. The other pathway uses a 1D convolutional network for utterance-level spectrum. Experimental results on the DCASE 2023 task 2 dataset show the state-of-the-art performance of our proposed method. Moreover, visualizations of the intermediate feature maps in the excitation network are provided to illustrate the effectiveness of our method. Yucong Zhang, Juan Liu 0007, Ming Li 0026 |
ICASSP | 2 |
| 2024 | BS2CL: Balanced Self-supervised Contrastive Learning for Thyroid Cytology Whole Slide Image Multi-classification
Wensi Duan, Juan Liu 0007, Peng Jiang 0025, Dehua Cao |
ICIC (7) | 2 |
| 2024 | Audio-Visual Wake-up Word Spotting Under Noisy and Multi-person Scenarios
Cancan Li, Juan Liu 0007 |
ICPR (33) | 3 |
| 2024 | Variable-Length Promoter Strength Prediction Based on Graph Convolution
Tianqi Teng, Qiang Zhang 0031, Juan Liu 0007 |
ISBRA (1) | 4 |
| 2024 | A dual-scale fused hypergraph convolution-based hyperedge prediction model for predicting missing reactions in genome-scale metabolic networksabstractGenome-scale metabolic models (GEMs) are powerful tools for predicting cellular metabolic and physiological states. However, there are still missing reactions in GEMs due to incomplete knowledge. Recent gaps filling methods suggest directly predicting missing responses without relying on phenotypic data. However, they do not differentiate between substrates and products when constructing the prediction models, which affects the predictive performance of the models. In this paper, we propose a hyperedge prediction model that distinguishes substrates and products based on dual-scale fused hypergraph convolution, DSHCNet, for inferring the missing reactions to effectively fill gaps in the GEM. First, we model each hyperedge as a heterogeneous complete graph and then decompose it into three subgraphs at both homogeneous and heterogeneous scales. Then we design two graph convolution-based models to, respectively, extract features of the vertices in two scales, which are then fused via the attention mechanism. Finally, the features of all vertices are further pooled to generate the representative feature of the hyperedge. The strategy of graph decomposition in DSHCNet enables the vertices to engage in message passing independently at both scales, thereby enhancing the capability of information propagation and making the obtained product and substrate features more distinguishable. The experimental results show that the average recovery rate of missing reactions obtained by DSHCNet is at least 11.7% higher than that of the state-of-the-art methods, and that the gap-filled GEMs based on our DSHCNet model achieve the best prediction performance, demonstrating the superiority of our method. Qiang Zhang 0031, Juan Liu 0007 |
Briefings Bioinform. | 4 |
| 2024 | Interpretable detector for cervical cytology using self-attention and cell origin group guidance
Peng Jiang 0025, Juan Liu 0007, Jing Feng 0005, Yuqi Chen 0007, Dehua Cao |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | A multi-stream network for retrosynthesis prediction
Qiang Zhang 0031, Juan Liu 0007, Wen Zhang 0008 |
Frontiers Comput. Sci. | 2 |
| 2023 | MSCCNet: Multi-Scale Convolution-Capsule Network for Cervical Cell ClassificationabstractCervical cancer is one of the fastest growing and most dangerous cancers, seriously threatening women’s health and lives. Cervical cytopathology image classification is a very important approach for diagnosing cervical cancer. The advent of the automatic computer-aided diagnosis system can tackles this issue. However, cervical cell images of different classes exhibit similar appearances, posing a challenge for accurate classification. To address this challenge, this work proposes a framework named MSCCNet. In our MSCCNet, the cross-layer attention-based feature fusion module is used to obtain multi-scale discriminative features. Meanwhile, the spatial relationship modeling module is utilized to encode the relative relationship between objects and capture more slight differences between cervical cells, further strengthening the representation ability of features. We also introduce the joint loss to enhance the penalty for misclassified samples. The model training and evaluation are performed on our developed DSCC dataset and publicly available SIPaKMeD datasets. The proposed MSCCNet achieves overall accuracies of 87.88% and 97.90% on these two datasets, respectively, outperforming several existing classification methods. Juan Liu 0007, Peng Jiang 0025, Jing Feng 0005, Dehua Cao |
BIBM | 2 |
| 2023 | AMTL-RFC:A multi-task learning based method for evaluating the feasibility of enzymatic reactionsabstractIn the field of metabolic engineering, evaluating the feasibility of newly generated enzymatic reactions in retrobiosynthesis is a crucial process that helps biologists to efficiently screen out infeasible reactions. However, existing methods overlook the significance of sequence features in molecular SMILES and the ability of model to comprehensively mine and extract features needs to be strengthened. To address these issues, our work propose a novel attention-based multi-task learning (MTL) reaction feasibility checker, named AMTL-RFC, for enzymatic reaction feasibility classification. The model consists of two branches: a Transformer network and a 1-D convolutional neural network (CNN) that extracts SMILES sequence features and spatial structure features of the substrate and product in the reactant pair, respectively. Moreover, a multi-task learning strategy is employed to further enhance the model’s performance. Experimental results demonstrate that AMTL-RFC achieves a classification accuracy of 92.27% on the primary test set, which is highly competitive in the task of classifying the feasibility of enzyme reactions. Jianghang Liu, Juan Liu 0007, Qiang Zhang 0031 |
BIBM | 2 |
| 2023 | SLPFA: Protein Structure-Label Embedding Attention Network for Protein Function AnnotationabstractGene Ontology (GO) is a framework that utilizes a series of GO terms in a Directed Acyclic Graph (DAG) to describe protein functions. Proteins are typically annotated with several or dozens of GO terms. However, existing methods often struggle to simultaneously annotate multiple relevant GO terms with hierarchical dependencies to proteins, as they solely rely on protein sequences or structures. To better utilize the hierarchical information of GO terms and improve protein function annotation performance, we propose the Protein Structure-Label Embedding Attention Network for Protein Function Annotation (SLPFA). SLPFA embeds proteins and GO terms into a joint latent space using attention mechanisms to bridge the semantic gap between them. Specifically, we employ a soft-mask GNN to learn the topological structure of proteins, allowing simultaneous focus on key nodes while remaining invariant to irrelevant parts. Additionally, we encode the ancestral information for each GO term in its embedding and utilize a learnable matrix to capture the hierarchical dependencies. Finally, SLPFA employs protein structure-label embedding attention to project the protein structure and label embedding together into a joint latent space. This enables the model to learn the high-level semantics of proteins and hierarchical GO terms, resulting in a reduced semantic gap between proteins and their functions. Experimental results demonstrate that SLPFA outperforms state-of-the-art deep learning-based methods on the PDB-cdhit dataset, which yields Fmax of 0.604, 0.478, 0.524 and the AUPRC of 0.630, 0.357, 0.452 for the MF, BP, CC ontology domains, respectively. Furthermore, when the training and testing proteins have less than 15% sequence identity, SLPFA also achieves competitive results in the MF, BP, and CC ontology domains. Qiang Zhang 0031, Juan Liu 0007, Jing Feng 0005 |
BIBM | 2 |
| 2023 | Classifying Pathological Images Based on Multi-Instance Learning and End-to-End Attention PoolingabstractIn order to address the issue that previous deep learning methods for classifying pathological images cannot adaptively learn features, we propose an end-to-end attention pooling method based on a multi-instance learning patch scoring model. Our method integrates feature extraction and classification into a unified framework that is conducive to extracting the most valuable features. In this model, a patch scoring method is constructed by a multi-instance learning method firstly and then the partial patches selected by the patches scoring model are classified using an end-to-end classification model that incorporates an attention pooled mechanism. To make the pathological image classification mechanism more compatible with the pathologist diagnosis method, we use the squared average normalization function instead of the softmax function to optimize the feature extraction and fusion process, so that the high score patches in positive pathological images receive more attention weights, thus giving better interpretability to the classification results. Experiments on publicly available datasets TCGA_BRCA show a significant improvement in the performance of our approach over other work. Yuqi Chen 0007, Juan Liu 0007, Zhiqun Zuo, Peng Jiang 0025, Guangsheng Wu |
ICASSP | 2 |
| 2023 | MASKED-AP: Attention Pyramid Convolutional Neural Network with Mask for Cervical Cell ClassificationabstractThe automatic and effective cervical cell classification technique is critical for cervical cytology screening and cervical cancer prevention. We notice that cervical cell classification is a fine-grained classification task. The difference between classes is small while the difference within a class is large, so it is difficult to capture the discriminative features between different classes of cells for classification. To address this problem, this paper proposes an attention pyramid model (Masked-AP) used for cervical cell classification. Our Masked-AP effectively combines high-level semantic features with low-level detailed features extracted from images, which are both important for classification. Further, we also use the attention mask to drive our model to focus more on the cell nuclei containing substantial discriminative information. For model evaluation, we built a cervical cell dataset named LDCC including 17476 images from 507 subjects. LDCC is available at http://ldcc3.biodwhu.cn/LDCC3/. Our method yielded 74.11% accuracy, 74.11% recall, 74.19% precision, and 74.07% F1-score, and outperforms the state-of-the-art methods. Juan Liu 0007, Wensi Duan, Dehua Cao |
ICASSP | 2 |
| 2023 | DDN: Dynamic Aggregation Enhanced Dual-Stream Network for Medical Image ClassificationabstractConvolutional Neural Networks (CNNs) have become the de facto approach for medical image classification in recent years. However, the deficiency of convolutional operations in extracting global features has limited the further improvement of this task. Vision Transformers (ViTs) can model long-range dependencies via self-attention mechanism but unfortunately lose local feature details. In this paper, we propose a dynamic aggregation enhanced dual-stream network termed DDN to take the advantage of ViT and CNN to enrich the feature representation of medical images. Specifically, our proposed DDN is built by stacking several Dynamic Dual-stream Units (DDU). In DDU, local and global features are learned by the CNN branch and Transformer branch respectively whilst complementing each other via a bi-directional propagation strategy, then features of both branches are aggregated in a dynamic manner and the integrated information is used to enhance the feature representations of the two branches simultaneously. Extensive experiments show that our proposed DDN performs best compared with other state-of-the-art models on the public Kvasir dataset and ISIC2018 dataset. Juan Liu 0007, Peng Jiang 0025, Dehua Cao |
ICASSP | 2 |
| 2023 | LGVIT: Local-Global Vision Transformer for Breast Cancer Histopathological Image ClassificationabstractBreast cancer histopathological image classification has made great progress with the use of Convolutional Neural Networks (CNNs). However, due to the limited receptive field, CNNs have difficulty in learning the global information of breast cancer histopathological images, hindering the further improvement of this task. To solve this problem, we reasonably apply self-attention mechanism to this task and propose a new network called Local-Global Vision Transformer (LGViT) which utilizes CNNs to capture local features and self-attention mechanism to learn global features of histopathological images. LGViT has several advantages: (1) We propose Local-Global Multi-head Self-attention, a new mechanism that models long-range dependencies with low computational cost. In this mechanism, self-attention is first performed separately within each window. Then, Multiple Instance Learning scheme is utilized to obtain a representative token for each window. Finally, we compute self-attention among these representative tokens to capture global information. (2) We propose Ghost Feed-forward Network, which compensates for the deficiency of Vision Transformer in capturing local features via a locality mechanism. (3) We use a CNN stem to effectively capture low-level information. Experiments on the PatchCamelyon dataset show that LGViT is better than other state-of-the-art methods. Juan Liu 0007, Peng Jiang 0025, Dehua Cao |
ICASSP | 2 |
| 2023 | DeepMAT: Predicting Metabolic Pathways of Compounds Using a Message Passing and Attention-Based Neural Networks
Hayat Ali Shah, Juan Liu 0007, Jing Feng 0005 |
ICIC (3) | 2 |
| 2023 | Pyramid Shape-Aware Semi-supervised Learning for Thyroid Nodules Segmentation in Ultrasound Images
Juan Liu 0007 |
PRCV (5) | 2 |
| 2023 | An adaptive multi-modal hybrid model for classifying thyroid nodules by combining ultrasound and infrared thermal imagesabstractBACKGROUND: Two types of non-invasive, radiation-free, and inexpensive imaging technologies that are widely employed in medical applications are ultrasound (US) and infrared thermography (IRT). The ultrasound image obtained by ultrasound imaging primarily expresses the size, shape, contour boundary, echo, and other morphological information of the lesion, while the infrared thermal image obtained by infrared thermography imaging primarily describes its thermodynamic function information. Although distinguishing between benign and malignant thyroid nodules requires both morphological and functional information, present deep learning models are only based on US images, making it possible that some malignant nodules with insignificant morphological changes but significant functional changes will go undetected. RESULTS: Given the US and IRT images present thyroid nodules through distinct modalities, we proposed an Adaptive multi-modal Hybrid (AmmH) classification model that can leverage the amalgamation of these two image types to achieve superior classification performance. The AmmH approach involves the construction of a hybrid single-modal encoder module for each modal data, which facilitates the extraction of both local and global features by integrating a CNN module and a Transformer module. The extracted features from the two modalities are then weighted adaptively using an adaptive modality-weight generation network and fused using an adaptive cross-modal encoder module. The fused features are subsequently utilized for the classification of thyroid nodules through the use of MLP. On the collected dataset, our AmmH model respectively achieved 97.17% and 97.38% of F1 and F2 scores, which significantly outperformed the single-modal models. The results of four ablation experiments further show the superiority of our proposed method. CONCLUSIONS: The proposed multi-modal model extracts features from various modal images, thereby enhancing the comprehensiveness of thyroid nodules descriptions. The adaptive modality-weight generation network enables adaptive attention to different modalities, facilitating the fusion of features using adaptive weights through the adaptive cross-modal encoder. Consequently, the model has demonstrated promising classification performance, indicating its potential as a non-invasive, radiation-free, and cost-effective screening tool for distinguishing between benign and malignant thyroid nodules. The source code is available at https://github.com/wuliZN2020/AmmH . Juan Liu 0007, Wensi Duan, Ziling Wu, Zhaohui Cai |
BMC Bioinform. | 2 |
| 2023 | FragDPI: a novel drug-protein interaction prediction model based on fragment understanding and unified coding
Juan Liu 0007, Xuekai Zhu, Qiang Zhang 0031, Hayat Ali Shah |
Frontiers Comput. Sci. | 2 |
| 2022 | A Context-Guided Attention Method for Integrating Features of Histopathological PatchesabstractLots of researchers have studied for classifying histopathological whole slide images (WSIs). Since a WSI is too large to be processed directly, researchers usually cut it into many small-sized patches and then integrate the discriminative features extracted from the patches to obtain a slide-level feature of the WSI. The integration strategy generating the slide-level features is crucial for the WSI classification model. Lots of attention-based methods have been proposed for such purpose. However, most attention-based methods do not take the patches relationship into consideration, which affects the classification performance of the models. In this work, we propose a novel Context-Guided attention (CGattention) method to integrate the patch-level features, which constructs a context vector to simulate the global context information of the whole WSI and implicitly characterizes the relationship between patches in the WSI. When evaluated on two publicly available datasets, the CGattention based model obtained the better performance than other attention-based models. Yuqi Chen 0007, Juan Liu 0007, Peng Jiang 0025, Jing Feng 0005, Dehua Cao |
BIBM | 2 |
| 2022 | Predicting Tumor Mutation Burden of TNBC Based on Nuclei Scores of Histopathological ImagesabstractTumor mutation burden(TMB) is a biomarker for predicting immunotherapy responses, which can be used to filter out Triple-Negative Breast Cancer(TNBC) patients who benefit from immunotherapy. It is generally measured by whole-exome sequencing (WES) in clinical practice. However, WES has the disadvantage of being expensive, time-consuming, and operational complexity so that it is not available in most hospitals. To solve these issues, we developed a machine learning algorithm that predicts the TMB of TNBC based on the nuclei score of histopathological images, which can obtain high accuracy without manually labeling tumor regions. We verify the effectiveness of patches filtered by nuclei score to TMB classifier by comparing the performance of the model that trained with all patches and trained with selected patches. Experiments results show that the accuracy of the model trained by patches selected with nuclei score is 87.5% and F1 is 80%, which are much higher than training with all patches(87.5% vs 56.25%, 80% vs 58.82%). The time of testing a sample using our approach is only 1 in 26, compared with the test time with all patches. To the best of our knowledge, this is the first research to predict TMB from TNBC histopathological images. The proposed approach has the potential to provide immunotherapy to a much broader subset of patients with TNBC. Yuqi Chen 0007, Juan Liu 0007, Peng Jiang 0025, Dehua Cao |
BIBM | 2 |
| 2022 | Classifying Cervical Histopathological Whole Slide Images via Deep Multi-Instance Transfer LearningabstractThe cervical histopathology analysis result is the gold standard for cervical cancer diagnosis. Conventional histopathological examination depends on pathologists’ observation under microscope, which is notoriously labor-intensive and subjective. The popularization of digital pathology technology makes the collection of the cervical histopathological whole slide images (WSIs) more convenient, so it has become possible to develop computer-aided diagnosis methods for cervical cancer. In this work, we first collected the cervical histopathological WSIs from 917 patients with pathological diagnosis through a retrospective study, of which 286 WSIs contained annotations of several lesion areas that were manually outlined by the pathologists. Then we proposed a method for classifying cervical histopathological WSIs by combining deep multi-instance transfer learning (DMITL) and support vector machine (SVM). The DMITL aimed for learning the representations of the WSIs, and the SVM was used for building the classification model of the WSIs. We generated the training and test sets based on our collected WSIs to train and evaluate our method. The validation results have shown that the good performance of our proposed method. Peng Jiang 0025, Juan Liu 0007, Jing Feng 0005, Dehua Cao |
BIBM | 2 |
| 2022 | Cross-Attention Based Multi-Scale Feature Fusion Vision Transformer For Breast Ultrasound Image ClassificationabstractBreast cancer has become one of the most common cancers in the world, and it is also the most lethal cancer in women. As a non-invasive imaging modality, ultrasonography can diagnose the degree of breast lesions and be used for large-scale screening. However, since the lesions in breast ultrasound(BUS) images are morphologically diverse, accompanied by relatively low contrast and complex textures, BUS image recognition faces greater challenges than natural images. In this study, We propose a novel network architecture that combines convolutional neural network(CNN) with vision transformer(ViT) to aggregate local feature details and long-range feature dependencies. Moreover, in order to perform multi-scale feature fusion, we introduce cross attention between the deep feature map and the shallow feature map in the network block to carry out the interaction between the deep feature and the shallow feature information. To verify the effectiveness of the model, we constructed a large-scale dataset and conducted extensive experiments. The results show that our method achieves an accuracy of 85.33%, under the comparable parameter complexity, which outperforms most convolutional neural networks(CNNs) and vision transformers (ViTs). Lele Li, Ziling Wu, Juan Liu 0007, Peng Jiang 0025, Jing Feng 0005 |
BIBM | 3 |
| 2022 | A Novel IoMT System for Pathological Diagnosis Based on Intelligent Mobile Scanner and Whole Slide Image Stitching Method
Peng Jiang 0025, Juan Liu 0007, Zongjie Hao, Dehua Cao |
ICIC (3) | 2 |
| 2022 | Channel Spatial Collaborative Attention Network for Fine-Grained Classification of Cervical Cells
Peng Jiang 0025, Juan Liu 0007, Dehua Cao |
ICONIP (6) | 2 |
| 2022 | Accurate classification of white blood cells by coupling pre-trained ResNet and DenseNet with SCAM mechanismabstractBACKGROUND: Via counting the different kinds of white blood cells (WBCs), a good quantitative description of a person's health status is obtained, thus forming the critical aspects for the early treatment of several diseases. Thereby, correct classification of WBCs is crucial. Unfortunately, the manual microscopic evaluation is complicated, time-consuming, and subjective, so its statistical reliability becomes limited. Hence, the automatic and accurate identification of WBCs is of great benefit. However, the similarity between WBC samples and the imbalance and insufficiency of samples in the field of medical computer vision bring challenges to intelligent and accurate classification of WBCs. To tackle these challenges, this study proposes a deep learning framework by coupling the pre-trained ResNet and DenseNet with SCAM (spatial and channel attention module) for accurately classifying WBCs. RESULTS: In the proposed network, ResNet and DenseNet enables information reusage and new information exploration, respectively, which are both important and compatible for learning good representations. Meanwhile, the SCAM module sequentially infers attention maps from two separate dimensions of space and channel to emphasize important information or suppress unnecessary information, further enhancing the representation power of our model for WBCs to overcome the limitation of sample similarity. Moreover, the data augmentation and transfer learning techniques are used to handle the data of imbalance and insufficiency. In addition, the mixup approach is adopted for modeling the vicinity relation across training samples of different categories to increase the generalizability of the model. By comparing with five representative networks on our developed LDWBC dataset and the publicly available LISC, BCCD, and Raabin WBC datasets, our model achieves the best overall performance. We also implement the occlusion testing by the gradient-weighted class activation mapping (Grad-CAM) algorithm to improve the interpretability of our model. CONCLUSION: The proposed method has great potential for application in intelligent and accurate classification of WBCs. Juan Liu 0007, Chunbing Hua, Jing Feng 0005, Dehua Cao |
BMC Bioinform. | 2 |
| 2022 | A novel hybrid framework for metabolic pathways prediction based on the graph attention networkabstractBACKGROUND: Making clear what kinds of metabolic pathways a drug compound involves in can help researchers understand how the drug is absorbed, distributed, metabolized, and excreted. The characteristics of a compound such as structure, composition and so on directly determine the metabolic pathways it participates in. METHODS: We developed a novel hybrid framework based on the graph attention network (GAT) to predict the metabolic pathway classes that a compound involves in, named HFGAT, by making use of its global and local characteristics. The framework mainly consists of a two-branch feature extracting layer and a fully connected (FC) layer. In the two-branch feature extracting layer, one branch is responsible to extract global features of the compound; and the other branch introduces a GAT consisting of two graph attention layers to extract local structural features of the compound. Both the global and the local features of the compound are then integrated into the FC layer which outputs the predicted result of metabolic pathway categories that the compound belongs to. RESULTS: We compared the multi-class classification performance of HFGAT with six other representative methods, including five classic machine learning methods and one graph convolutional network (GCN) based deep learning method, on the benchmark dataset containing 6999 compounds belonging to 11 pathway categories. The results showed that the deep learning-based methods (HFGAT, GCN-based method) outperformed the traditional machine learning methods in the prediction of metabolic pathways and our proposed HFGAT method performed better than the GCN-based method. Moreover, HFGAT achieved higher [Formula: see text] scores on 8 of 11 classes than the GCN-based method. CONCLUSIONS: Our proposed HFGAT makes use of both the global and local information of the compounds to predict their metabolic pathway categories and has achieved a significant performance. Compared with the GCN model, the introduction of the GAT can help our model pay more attention to substructures of the compound that are useful for the prediction task. The study provided a potential method for drug discovery with all types of metabolic reactions that may be involved in the decomposition and synthesis of pharmaceutical compounds in the organism. Juan Liu 0007, Hayat Ali Shah, Jing Feng 0005 |
BMC Bioinform. | 2 |
| 2022 | CNN-based two-branch multi-scale feature extraction network for retrosynthesis predictionabstractBACKGROUND: Retrosynthesis prediction is the task of deducing reactants from reaction products, which is of great importance for designing the synthesis routes of the target products. The product molecules are generally represented with some descriptors such as simplified molecular input line entry specification (SMILES) or molecular fingerprints in order to build the prediction models. However, most of the existing models utilize only one molecular descriptor and simply consider the molecular descriptors in a whole rather than further mining multi-scale features, which cannot fully and finely utilizes molecules and molecular descriptors features. RESULTS: We propose a novel model to address the above concerns. Firstly, we build a new convolutional neural network (CNN) based feature extraction network to extract multi-scale features from the molecular descriptors by utilizing several filters with different sizes. Then, we utilize a two-branch feature extraction layer to fusion the multi-scale features of several molecular descriptors to perform the retrosynthesis prediction without expert knowledge. The comparing result with other models on the benchmark USPTO-50k chemical dataset shows that our model surpasses the state-of-the-art model by 7.4%, 10.8%, 11.7% and 12.2% in terms of the top-1, top-3, top-5 and top-10 accuracies. Since there is no related work in the field of bioretrosynthesis prediction due to the fact that compounds in metabolic reactions are much more difficult to be featured than those in chemical reactions, we further test the feasibility of our model in task of bioretrosynthesis prediction by using the well-known MetaNetX metabolic dataset, and achieve top-1, top-3, top-5 and top-10 accuracies of 45.2%, 67.0%, 73.6% and 82.2%, respectively. CONCLUSION: The comparison result on USPTO-50k indicates that our proposed model surpasses the existing state-of-the-art model. The evaluation result on MetaNetX dataset indicates that the models used for retrosynthesis prediction can also be used for bioretrosynthesis prediction. Juan Liu 0007, Qiang Zhang 0031 |
BMC Bioinform. | 2 |
| 2021 | TransMixNet: An Attention Based Double-Branch Model for White Blood Cell Classification and Its Training with the Fuzzified Training DataabstractWhite blood cells (WBCs) play a critical part in the human immune system, protecting human from bacteria and viruses. In general, WBCs can be divided into five categories: basophil, monocyte, eosinophil, neutrophil, and lymphocyte. The proportion of different kinds of WBCs is closely related to human health, thus analyzing the numbers and percentages of WBCs is usually included in the blood routine examination. Therefore, accurate identification of different types of WBCs has become an important step in the statistical analysis. However, the insufficiency and imbalance of WBC samples in medical computer vision domains is still a challenge to classify WBCs intelligently and accurately. This paper presents an attention based double-branch network, called TransMixNet, to build the classification model for WBCs recognition. During the model construction, we adopt transfer learning and data augmentation strategies to overcome the problems of insufficient and unbalanced samples. In addition, mixup strategy is used to model the domain relationship between different types of training samples to improve the generalization performance of the model. The evaluation experiment results show that our model is promising and applicable to clinical applications. Juan Liu 0007, Chunbing Hua, Zhiqun Zuo, Jing Feng 0005 |
BIBM | 2 |
| 2021 | COMNA: Core-attachment based protein complex detection via multiple network alignmentabstractProtein complexes are important functional units for performing biological functions and revealing cellular organization principles, and its detection is a fundamental problem in computational biology. In this paper, we present a novel algorithm COMNA by using orthology relationship across different species to complement the complex detection of a single protein-protein interaction (PPI) network, which can identify protein complexes from multiple PPI networks simultaneously. First, we propose a novel measure, high-order edge closure coefficient, to assess the reliability of interactions and then identify core-attachment based complexes from each PPI network. Second, we cooperate with the orthology information from multiple PPI network alignment to mine more complementary complexes. Finally, we merge the highly overlapping complexes by constructing an overlapping graph. We compare COMNA with eight state-of-the-art protein complex identification algorithms on three datasets, and the experimental results show that COMNA can detect complexes from PPI networks with higher accuracy. Yaoran Chen, Yuanyuan Zhu 0001, Ming Zhong 0002, Juan Liu 0007 |
BIBM | 4 |
| 2021 | CytoBrain: Cervical Cancer Screening System Based on Deep Learning Technology
Juan Liu 0007, Qing-Man Wen, Zhiqun Zuo, Jiasheng Liu, Jing Feng 0005 |
J. Comput. Sci. Technol. | 2 |
| 2021 | Group-Sparse SVD Models via $L_1$L1- and $L_0$L0-norm Penalties and their Applications in Biological DataabstractSparse Singular Value Decomposition (SVD) models have been proposed for biclustering high dimensional gene expression data to identify block patterns with similar expressions. However, these models do not take into account prior group effects upon variable selection. To this end, we first propose group-sparse SVD models with group Lasso (GL1-SVD) and group L0-norm penalty (GL0-SVD) for non-overlapping group structure of variables. However, such group-sparse SVD models limit their applicability in some problems with overlapping structure. Thus, we also propose two group-sparse SVD models with overlapping group Lasso (OGL1-SVD) and overlapping group L0-norm penalty (OGL0-SVD). We first adopt an alternating iterative strategy to solve GL1-SVD based on a block coordinate descent method, and GL0-SVD based on a projection method. The key of solving OGL1-SVD is a proximal operator with overlapping group Lasso penalty. We employ an alternating direction method of multipliers (ADMM) to solve the proximal operator. Similarly, we develop an approximate method to solve OGL0-SVD. Applications of these methods and comparison with competing ones using simulated data demonstrate their effectiveness. Extensive applications of them onto several real gene expression data with gene prior group knowledge identify some biologically interpretable gene modules. Wenwen Min, Juan Liu 0007 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | GPPIAL: A New Global PPI Network Aligner Based on OrthologsabstractProtein-protein interaction (PPI) network alignment aims to find a node mapping between nodes across networks for different specifies such that the mapped nodes are topologically and biologically similar. Such alignment can help to reveal functionally similar subnetworks and predict previously undiscovered orthologs of different species. Although many aligners have been proposed, few can find global alignments with both high topological and biological quality. In this paper, we propose a new pairwise global PPI network aligner (GPPIAL) based on orthologs. We first construct a score matrix to evaluate the matching possibility of node pairs by combining multiple sources of information, including the biological (sequence and functional), topological (degree and neighborhood), and interaction information. Then we select orthologs with high similarity scores in the Inparanoid database as anchor pairs, and gradually build a global alignment. We compare GPPIAL with the stateof-the-art pairwise global network aligners on multiple PPI network pairs from the BioGRID dataset. The results show that GPPIAL significantly outperforms existing global aligners in terms of biological quality without losing the topological quality. Moreover, GPPIAL can also detect more orthologous proteins than existing global aligners, and thus can be applied to predicting the function of unannotated proteins. Our code and datasets are available at https://github.com/whuyrc/GPPIAL. Yaoran Chen, Yuanyuan Zhu 0001, Ming Zhong 0002, Rong Peng, Juan Liu 0007 |
BIBM | 5 |
| 2020 | Multi-Class Metabolic Pathway Prediction by Graph Attention-Based Deep Learning MethodabstractExploring the relationship between molecular structure and metabolic pathways plays a significant role in researching the metabolism and pharmacokinetic effects of drugs. Therefore, one usually want to know what kinds of metabolic pathways that a new synthetic drug compound may involve in. In this paper, we propose a novel deep learning framework based on GAT (Graph Attention neTwork) for such purpose. The framework mainly consists of a two-branch feature extractor and a FC (Fully Connected) layer. In the two-branch feature extractor, one is used to generate three kinds global features of the compounds, and the other one is used to learn the local structural features via two GAT layers. All features are embedded into the FC layer to output multiple probabilities, each of which corresponds to the probability of one kind of pathway that the compound belongs to. The comparing results to other five state-of-the-art representative methods on the KEGG (Kyoto Encyclopedia of Genes and Genomes) data set have shown that our method can achieve the highest prediction accuracy, illustrating that it is a promising tool that can be helpful to analyze the metabolic pathways of the drug compounds. Juan Liu 0007, Jing Feng 0005 |
BIBM | 2 |
| 2020 | DAAT: A New Method to Train Convolutional Neural Network on Atrial Fibrillation Detection
Juan Liu 0007, Pei-Fang Li, Jing Feng 0005 |
ICIC (3) | 2 |
| 2019 | A Robust Approach to Locate HER2 and CEN17 Signals in Varied FISH ImagesabstractFluorescence in situ hybridization (FISH) is a well-recognized technique for evaluating the human epidermal growth factor receptor 2 (HER2) gene status in breast cancer, where HER2 genes are seen as the red stained spots, and the chromosome centromere (CEN17) on which HER2 is located are visible as the green stained spots under a fluorescent microscope. The traditional manual analysis of FISH images is low efficient and has great subjective differences. Many researches focus on automatically calculating the average ratio of red/green spots in a cell and have achieved good performance. However, there are still a need to improve the segmentation accuracy of cells and HER2/CEN17 signals in FISH images with poor quality. In this paper, we propose a robust approach to locate HER2 and CEN17 signals based on the image standardization of varied FISH images. First we propose to use offset operation to reduce the fluorescence scattering, and to use exposure ratio map estimation and brightness transformation to enhance the weak fluorescence. Then we detect red/green spots by sequentially using top-hat filtering, thresholding and post-processing operations. The experimental results have shown that our proposed method can work well for the clinical FISH images with different quality. Jiasheng Liu, Juan Liu 0007, Yuqi Chen 0007, Chunbing Hua, Zhuoyu Li, Zhiqun Zuo |
BIBM | 2 |
| 2019 | Predicting Drug-Disease Treatment Associations Based on Topological Similarity and Singular Value DecompositionabstractTo assist drug development, many computational methods have been proposed to identify potential drug-disease treatment associations before wet experiments. Based on the assumption that similar drugs may treat similar diseases, most methods need the similarities of drugs and diseases, and they will not work if the biological or chemical features for computing similarities are missing. Besides, being lack of validated negative samples in the drug-disease associations data, most methods simply select some unlabeled samples as negative ones, which may introduce noises. Herein, we propose a new method (TS-SVD) which only uses those known drug-protein, disease-protein and drug-disease interactions to predict the potential drug-disease associations. In a constructed drug-protein-disease heterogeneous network, we consider the common neighbors of drugs and diseases to obtain the topological similarity. Then the topological similarity matrix of drugs (diseases) will be used to get the low dimensional embedding representations of drug-disease pairs. Finally, a Random Forest classifier is trained to do the prediction. To train a more reasonable model, we select out some reliable negative samples based on the k-step neighbors relationships between drugs and diseases. Compared with some state-of-the-art methods, we use less information but achieve better or comparable performance. Meanwhile, our strategy for selecting reliable negative samples can improve the performances of these methods. Guangsheng Wu, Juan Liu 0007 |
BIBM | 2 |
| 2019 | Gene Functional Module Discovery via Integrating Gene Expression and PPI Network Data
Juan Liu 0007, Wenwen Min |
ICIC (2) | 2 |
| 2019 | Prediction of drug-disease associations based on ensemble meta paths and singular value decompositionabstractBACKGROUND: In the field of drug repositioning, it is assumed that similar drugs may treat similar diseases, therefore many existing computational methods need to compute the similarities of drugs and diseases. However, the calculation of similarity depends on the adopted measure and the available features, which may lead that the similarity scores vary dramatically from one to another, and it will not work when facing the incomplete data. Besides, supervised learning based methods usually need both positive and negative samples to train the prediction models, whereas in drug-disease pairs data there are only some verified interactions (positive samples) and a lot of unlabeled pairs. To train the models, many methods simply treat the unlabeled samples as negative ones, which may introduce artificial noises. Herein, we propose a method to predict drug-disease associations without the need of similarity information, and select more likely negative samples. RESULTS: In the proposed EMP-SVD (Ensemble Meta Paths and Singular Value Decomposition), we introduce five meta paths corresponding to different kinds of interaction data, and for each meta path we generate a commuting matrix. Every matrix is factorized into two low rank matrices by SVD which are used for the latent features of drugs and diseases respectively. The features are combined to represent drug-disease pairs. We build a base classifier via Random Forest for each meta path and five base classifiers are combined as the final ensemble classifier. In order to train out a more reliable prediction model, we select more likely negative ones from unlabeled samples under the assumption that non-associated drug and disease pair have no common interacted proteins. The experiments have shown that the proposed EMP-SVD method outperforms several state-of-the-art approaches. Case studies by literature investigation have found that the proposed EMP-SVD can mine out many drug-disease associations, which implies the practicality of EMP-SVD. CONCLUSIONS: The proposed EMP-SVD can integrate the interaction data among drugs, proteins and diseases, and predict the drug-disease associations without the need of similarity information. At the same time, the strategy of selecting more reliable negative samples will benefit the prediction. Guangsheng Wu, Juan Liu 0007, Xiang Yue |
BMC Bioinform. | 2 |
| 2018 | Edge-group sparse PCA for network-guided high dimensional data analysisabstractMotivation: Principal component analysis (PCA) has been widely used to deal with high-dimensional gene expression data. In this study, we proposed an Edge-group Sparse PCA (ESPCA) model by incorporating the group structure from a prior gene network into the PCA framework for dimension reduction and feature interpretation. ESPCA enforces sparsity of principal component (PC) loadings through considering the connectivity of gene variables in the prior network. We developed an alternating iterative algorithm to solve ESPCA. The key of this algorithm is to solve a new k-edge sparse projection problem and a greedy strategy has been adapted to address it. Here we adopted ESPCA for analyzing multiple gene expression matrices simultaneously. By incorporating prior knowledge, our method can overcome the drawbacks of sparse PCA and capture some gene modules with better biological interpretations. Results: We evaluated the performance of ESPCA using a set of artificial datasets and two real biological datasets (including TCGA pan-cancer expression data and ENCODE expression data), and compared their performance with PCA and sparse PCA. The results showed that ESPCA could identify more biologically relevant genes, improve their biological interpretations and reveal distinct sample characteristics. Availability and implementation: An R package of ESPCA is available at http://page.amss.ac.cn/shihua.zhang/. Supplementary information: Supplementary data are available at Bioinformatics online. Wenwen Min, Juan Liu 0007 |
Bioinform. | 2 |
| 2018 | Network-Regularized Sparse Logistic Regression Models for Clinical Risk Prediction and Biomarker DiscoveryabstractMolecular profiling data (e.g., gene expression) has been used for clinical risk prediction and biomarker discovery. However, it is necessary to integrate other prior knowledge like biological pathways or gene interaction networks to improve the predictive ability and biological interpretability of biomarkers. Here, we first introduce a general regularized Logistic Regression (LR) framework with regularized term , which can reduce to different penalties, including Lasso, elastic net, and network-regularized terms with different . This framework can be easily solved in a unified manner by a cyclic coordinate descent algorithm which can avoid inverse matrix operation and accelerate the computing speed. However, if those estimated and have opposite signs, then the traditional network-regularized penalty may not perform well. To address it, we introduce a novel network-regularized sparse LR model with a new penalty to consider the difference between the absolute values of the coefficients. We develop two efficient algorithms to solve it. Finally, we test our methods and compare them with the related ones using simulated and real data to show their efficiency. Wenwen Min, Juan Liu 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | GMAlign: A new network aligner for revealing large conserved functional componentsabstractThe alignment of protein-protein interaction (PPI) networks is an effective approach to uncover the functionally conserved sub-structure between networks. A wealth of approaches have been developed for global PPI network alignment in recent years. However, due to the computational intractability caused by its NP-completeness, global PPI network alignment remains challenging in finding large conserved components stably for various PPI network pairs. In this paper, we introduce a novel global network aligner based on graph matching method called GMAlign. We assess the outperformance of GMAlign over the state-of-the-art network aligners on various PPI network pairs from the largest BioGRID dataset. It is shown that GMAlign not only can produce larger size alignment, but also can find bigger and denser common connected subgraphs robustly for the first time. Moreover, we shows that GMAlign can produce both structurally and functionally meaningful results in detecting large conserved biological pathways between species. The GMAlign software, datasets and supplementary experimental results can be downloaded at https://github.com/yzlwhu/GMAlign. Yuanyuan Zhu 0001, Yuezhi Li, Juan Liu 0007, Lu Qin 0001, Jeffrey Xu Yu |
BIBM | 3 |
| 2017 | Pattern fusion analysis by adaptive alignment of multiple heterogeneous omics dataabstractMOTIVATION: Integrating different omics profiles is a challenging task, which provides a comprehensive way to understand complex diseases in a multi-view manner. One key for such an integration is to extract intrinsic patterns in concordance with data structures, so as to discover consistent information across various data types even with noise pollution. Thus, we proposed a novel framework called 'pattern fusion analysis' (PFA), which performs automated information alignment and bias correction, to fuse local sample-patterns (e.g. from each data type) into a global sample-pattern corresponding to phenotypes (e.g. across most data types). In particular, PFA can identify significant sample-patterns from different omics profiles by optimally adjusting the effects of each data type to the patterns, thereby alleviating the problems to process different platforms and different reliability levels of heterogeneous data. RESULTS: To validate the effectiveness of our method, we first tested PFA on various synthetic datasets, and found that PFA can not only capture the intrinsic sample clustering structures from the multi-omics data in contrast to the state-of-the-art methods, such as iClusterPlus, SNF and moCluster, but also provide an automatic weight-scheme to measure the corresponding contributions by data types or even samples. In addition, the computational results show that PFA can reveal shared and complementary sample-patterns across data types with distinct signal-to-noise ratios in Cancer Cell Line Encyclopedia (CCLE) datasets, and outperforms over other works at identifying clinically distinct cancer subtypes in The Cancer Genome Atlas (TCGA) datasets. AVAILABILITY AND IMPLEMENTATION: PFA has been implemented as a Matlab package, which is available at http://www.sysbio.ac.cn/cb/chenlab/images/PFApackage_0.1.rar . CONTACT: [email protected] , [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chuanchao Zhang, Minrui Peng, Xiangtian Yu, Tao Zeng 0003, Juan Liu 0007, Luonan Chen |
Bioinform. | 6 |
| 2017 | Comparative network stratification analysis for identifying functional interpretable network biomarkersabstractBACKGROUND: A major challenge of bioinformatics in the era of precision medicine is to identify the molecular biomarkers for complex diseases. It is a general expectation that these biomarkers or signatures have not only strong discrimination ability, but also readable interpretations in a biological sense. Generally, the conventional expression-based or network-based methods mainly capture differential genes or differential networks as biomarkers, however, such biomarkers only focus on phenotypic discrimination and usually have less biological or functional interpretation. Meanwhile, the conventional function-based methods could consider the biomarkers corresponding to certain biological functions or pathways, but ignore the differential information of genes, i.e., disregard the active degree of particular genes involved in particular functions, thereby resulting in less discriminative ability on phenotypes. Hence, it is strongly demanded to develop elaborate computational methods to directly identify functional network biomarkers with both discriminative power on disease states and readable interpretation on biological functions. RESULTS: In this paper, we present a new computational framework based on an integer programming model, named as Comparative Network Stratification (CNS), to extract functional or interpretable network biomarkers, which are of strongly discriminative power on disease states and also readable interpretation on biological functions. In addition, CNS can not only recognize the pathogen biological functions disregarded by traditional Expression-based/Network-based methods, but also uncover the active network-structures underlying such dysregulated functions underestimated by traditional Function-based methods. To validate the effectiveness, we have compared CNS with five state-of-the-art methods, i.e. GSVA, Pathifier, stSVM, frSVM and AEP on four datasets of different complex diseases. The results show that CNS can enhance the discriminative power of network biomarkers, and further provide biologically interpretable information or disease pathogenic mechanism of these biomarkers. A case study on type 1 diabetes (T1D) demonstrates that CNS can identify many dysfunctional genes and networks previously disregarded by conventional approaches. CONCLUSION: Therefore, CNS is actually a powerful bioinformatics tool, which can identify functional or interpretable network biomarkers with both discriminative power on disease states and readable interpretation on biological functions. CNS was implemented as a Matlab package, which is available at http://www.sysbio.ac.cn/cb/chenlab/images/CNSpackage_0.1.rar . Chuanchao Zhang, Juan Liu 0007, Tao Zeng 0003, Luonan Chen |
BMC Bioinform. | 2 |
| 2017 | Differential function analysis: identifying structure and activation variations in dysregulated pathways
Chuanchao Zhang, Juan Liu 0007, Tao Zeng 0003, Luonan Chen |
Sci. China Inf. Sci. | 2 |
| 2016 | Semi-supervised graph cut algorithm for drug repositioning by integrating drug, disease and genomic associationsabstractDrug discovery is a cost expensive and time consuming process. Approved drugs have favorable or validated pharmacokinetic properties and toxicological profiles. Therefore repurposing approved drugs to new diseases can potentially avoid expensive costs associated with the early-stage testing. In this work, we propose to integrate drug-drug, disease-disease, gene-gene and drug-disease associations to reposition the approved drugs. Firstly, multiple sources of data are integrated into three layers: the chemical/phenotype layer, the gene/machanism layer and the treatment layer. Secondly, the drug-drug and disease-disease similarities in three layers are respectively computed and then combined together. Finally, based on the hypothesis that similar drugs may treat similar diseases, we model the drug repositioning problem as an optimal problem and propose a semi-supervised graph cut (SSGC) algorithm to solve the problem. The experimental results show that integrating multiple sources of data can achieve better performances than only considering single kind of data, and our method outperforms three representative approaches. Moreover, the predicted top-ranked repositioned relations have been reported in literature, illustrating the usefulness of our method in practice. The predicted results are available at https://github.com/wgs666/SSGC.git. Guangsheng Wu, Juan Liu 0007, Caihua Wang |
BIBM | 2 |
| 2016 | Integration of multiple heterogeneous omics dataabstractIntegration of different genomic profiles is challenging to understand complex diseases in a multi-view manner. Computational method is needed to preserve useful information of data types as well as correct bias. Thus, we proposed a novel framework pattern fusion analysis (PFA), to fuse the local sample patterns into a global pattern of patients with respect to the underlying data, by adaptively aligning the information in each type of biological data. In particular, PFA can adjust the distinct data types and achieve more robust sample pattern within different profiles. To validate the effectiveness of PFA, we tested PFA on various synthetic datasets and found that PFA is able to effectively capture the intrinsic clustering structure than the state-of-the-art integrative methods, such as moCluster, iClusterPlus and SNF. Moreover, in a case study on kidney cancer, PFA not only identified the multi-way feature modules among the prior-known disease associated genes, methylations and miRNAs, but also outperformed in cancer subtypes identification and could get effective clinical prognosis prediction. Totally, PFA not only provides new insights on the more holistic & systems-level sample pattern, but also supplies a new way for selecting more informative types of biological data. Chuanchao Zhang, Juan Liu 0007, Xiangtian Yu, Tao Zeng 0003, Luonan Chen |
BIBM | 2 |
| 2016 | BioSynther: a customized biosynthetic potential explorerabstractMOTIVATION: One of the most promising applications of biosynthetic methods is to produce chemical products of high value from the ready-made chemicals. To explore the biosynthetic potentials of a chemical as a synthesis precursor, biosynthetic databases and related chemoinformatics tools are urgently needed. In the present work, a web-based tool, BioSynther, is developed to explore the biosynthetic potentials of precursor chemicals using BKM-react, Rhea, and more than 50,000 in-house RxnFinder reactions manually curated. BioSynther allows researchers to explore biosynthetic potentials, through so far known biochemical reactions, step by step interactively, which could be used as a useful tool in metabolic engineering and synthetic biology. AVAILABILITY AND IMPLEMENTATION: BioSynther is available at: http://www.lifemodules.org/BioSynther/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Weizhong Tu, Juan Liu 0007, Qian-Nan Hu |
Bioinform. | 3 |
| 2016 | Multi-fields model for predicting target-ligand interaction
Caihua Wang, Juan Liu 0007, Fei Luo 0004, Qian-Nan Hu |
Neurocomputing | 2 |
| 2015 | A novel two-stage method for identifying microRNA-gene regulatory modules in breast cancerabstractIn this paper, we propose a two-stage method for identifying miRNA-gene regulatory modules by integrating miRNA/mRNA expression profiles and miRNA genomic cluster data. We first adopt a Multiple-output Sparse Group Lasso (MSGL) regression model to predict the miRNA-gene regulatory network. Further, we propose a L0-penalized Singular Value Decomposition (L0-SVD) model to identify modules from the predicted network. We apply this method to miRNA and mRNA expression profiles of the breast cancer data from TCGA databases and identify ten miRNA-gene regulatory modules. We find that (1) the modules are significantly associated in a predicted miRNA-gene regulatory network; (2) the modules are significantly enriched in GO biological processes and KEGG pathways, respectively; (3) many miRNAs and genes in the modules are related with breast cancer. On average, 51% of the miRNAs and 30% of the genes are related with breast cancer. The results demonstrate that miRNA-gene regulatory modules provide insights into the mechanisms of the combinatorial regulation between miRNAs and genes. Wenwen Min, Juan Liu 0007, Fei Luo 0004 |
BIBM | 2 |
| 2015 | Identification of phenotypic networks based on whole transcriptome by comparative network decompositionabstractComplex diseases are usually caused by the dysfunctions of the molecular system or molecular network rather than individual molecules. Generally, the conventional methods first obtain a disease-associated network based on expression data and then study its biological functions. However, such a network may be only a part of the system facilitating a biological function or may involve in multiple functions. In this paper, we present a computational framework based on an integer programming model, named as comparative network decomposition (CND), to jointly identify optimal structures of significant and moderate phenotypic functions/networks and their optimal combination by integrating gene expression, gene network and gene ontology together. Particularly, CND makes full use of dysfunctional information, e.g. both strong and weak changes on gene expressions and correlations, to extract various phenotypic networks, where one phenotypic network just corresponds to a specific biological function. A synthetic example clearly suggests that CND can identify multiple types of the disease-related phenotypic networks, rather than conventional approaches only exact significant phenotypic networks. As a proof-of-concept study to real data, CND is further used to identify the significant and moderate phenotypic networks for discriminating two different but associated diseases, e.g. subtypes of diabetes. In the comparison of type 1 and type 2 diabetes, the moderate and significant phenotypic networks can capture the disease-related biological functions and their corresponding networks. Therefore, CND is actually a powerful bioinformatics tool, which can investigate phenotype-associated genes and networks in a whole transcriptome and function-centered manner, and the comparative study of complex diseases with other works also demonstrates its effectiveness. Chuanchao Zhang, Juan Liu 0007, Tao Zeng 0003, Luonan Chen |
BIBM | 2 |
| 2014 | MMSE: A generalized coherence measure for identifying linear patternsabstractBiclustering is very useful in bioinformatics, information retrieval, electoral data analysis, dimension reduction, and so on. It is usually formulated as an optimization problem of searching maximal subsets of rows and columns satisfying some coherence criteria. The found submatrices are called as biclusters. There are several quantitative coherence measurements for linear patterns proposed. However, they are either lack the capability of properly evaluating all subtypes of linear patterns, or sensitive to the noise. In this paper, we propose a coherence measurement for the general linear patterns, the minimal mean squared error (MMSE). By using MMSE, the biclustering algorithms are expected to identify all types of linear patterns, including shifting (additive), scaling (multiplicative), and the general linear (the mixed form of shifting and scaling) ones, if they are presented in the data. Our comparative experimental results have highlighted that MMSE can actually help to identify significant general linear biclusters in artificial and real application data. Shuhua Chen, Juan Liu 0007, Tao Zeng 0003 |
BIBM | 2 |
| 2014 | Pairwise input neural network for target-ligand interaction predictionabstractPrediction the interactions between proteins (targets) and small molecules (ligands) is a critical task for the drug discovery in silico. In this work, we consider the target binding site instead of the whole target and propose a pairwise input neural network (PINN) for constructing the site-ligand interaction prediction model. Different with the ordinary artificial neural network (ANN) with one vector as input, the proposed PINN can accept a pair of vectors as the input, corresponding to a binding site and a ligand respectively. The 5-CV evaluation results show that PINN outperforms other representative target-ligand interaction prediction methods. Caihua Wang, Juan Liu 0007, Fei Luo 0004, Yafang Tan, Zixin Deng, Qian-Nan Hu |
BIBM | 2 |
| 2013 | Predicting immunogenic T-cell epitopes by combining various sequence-derived featuresabstractThe prediction of T-cell epitopes is of great help for facilitating vaccine design and understanding the immune system. In the bioinformatics, the MHC-binding peptides are defined as the T-cell epitopes, which will trigger the immune response to the antigens. However, binding peptides cannot necessarily activate the immune response, namely non-immunogenic. Until now, little attention has been paid to the immunogenic epitopes. Therefore, the recognition of immunogenic epitopes is a challenging task of the practical value. This paper systematically evaluates a wide variety of sequenced-derived features, which have been ever used for epitope prediction or similar tasks, and reveals their relationship with epitope immunogenicity. Then, we consider how to effectively exploit various features for the computational prediction of immunogenic epitopes. Subsequently, the random forest is adopted as the classification engine, and an ensemble model is developed by using the average scores of individual feature-based predictors. Compared with the previously published methods (POPI, POPISK and PAAQD), our models produce better performance on the benchmark datasets. Evaluated by t-test, the improvements of our method against existing methods are statistically significant (P<;0.01), showing the promise for the immunogenic epitope prediction. At present, only one MHC allele (HLA-A2) has sufficient data for the immunogenicity study. In the near future, with the increasing availability of immunogenic epitopes, we will carry out computational experiments on more MHC alleles. The source code for the ensemble model is available at: http://bcell.whu.edu.cn/sourcecode.html. Wen Zhang 0008, Juan Liu 0007, Yi Xiong 0002, Meng Ke |
BIBM | 2 |
| 2013 | Integrating peptides' sequence and energy of contact residues information improves prediction of peptide and HLA-I binding with unknown allelesabstractBACKGROUND: The HLA (human leukocyte antigen) class I is a kind of molecule encoded by a large family of genes and is characteristic of high polymorphism. Now the number of the registered HLA-I molecules has exceeded 3000. Slight differences in the amino acid sequences of HLAs would make them bind to different sets of peptides. In the past decades, although many methods have been proposed to predict the binding between peptides and HLA-I molecules and achieved good performance, most experimental data used by them is limited to the HLAs with a small number of alleles. Thus they are inclined to obtain high prediction accuracy only for data with similar alleles. Because the peptides and HLAs together determine the binding, it's necessary to consider their contribution meanwhile. RESULTS: By taking into account the features of the peptides sequence and the energy of contact residues, in this paper a method based on the artificial neural network is proposed to predict the binding of peptides and HLA-I even when the HLAs' potential alleles are unknown. Two experiments in the allele-specific and super-type cases are performed respectively to validate our method. In the first case, we collect 14 HLA-A and 14 HLA-B molecules on Bjoern Peters dataset, and compare our method with the ARB, SMM, NetMHC and other 16 online methods. Our method gets the best average AUC (Area under the ROC) value as 0.909. In the second one, we use leave one out cross validation on MHC-peptide binding data that has different alleles but shares the common super-type. Compared to gold standard methods like NetMHC and NetMHCpan, our method again achieves the best average AUC value as 0.847. CONCLUSIONS: Our method achieves satisfactory results. Whenever it's tested on the HLA-I with single definite gene or with super-type gene locus, it gets better classification accuracy. Especially, when the training set is small, our method still works better than the other methods in the comparison. Therefore, we could make a conclusion that by combining the peptides' information, HLAs amino acid residues' interaction information and contact energy, our method really could improve prediction of the peptide HLA-I binding even when there aren't the prior experimental dataset for HLAs with various alleles. Fei Luo 0004, Yangyang Gao, Yongqiong Zhu, Juan Liu 0007 |
BMC Bioinform. | 4 |
| 2012 | Predicting Binding-Peptide of HLA-I on Unknown Alleles by Integrating Sequence Information and Energies of Contact Residues
Fei Luo 0004, Yangyang Gao, Yongqiong Zhu, Juan Liu 0007 |
ICIC (3) | 4 |
| 2011 | Prediction of Heme Binding Sites in Heme Proteins Using an Integrative Sequence Profile Coupling Evolutionary Information with Physicochemical PropertiesabstractHeme-protein interactions are essential for various biological processes such as electron transfer, catalysis, signal transduction and the control of gene expression. The knowledge of heme binding residues can provide crucial clues to understand the mechanism of heme- protein interactions and aid in functional annotation. In the present work, we propose a sequence-based approach for the accurate prediction of heme binding residues by a novel integrative sequence profile coupling position specific scoring matrices with heme specific physicochemical properties. Particularly, we design an intuitive feature selection scheme for informative physicochemical properties. As shown in the primary results, our integrative sequence profile approach for prediction of heme binding residues outperforms the conventional methods using amino acid and evolutionary information on the 5-fold cross validation and the independent test. Yi Xiong 0002, Wen Zhang 0008, Tao Zeng 0003, Juan Liu 0007 |
BIBM | 4 |
| 2011 | A novel computational framework for simultaneous integration of multiple types of genomic data to identify microRNA-gene regulatory modulesabstractMOTIVATION: It is well known that microRNAs (miRNAs) and genes work cooperatively to form the key part of gene regulatory networks. However, the specific functional roles of most miRNAs and their combinatorial effects in cellular processes are still unclear. The availability of multiple types of functional genomic data provides unprecedented opportunities to study the miRNA-gene regulation. A major challenge is how to integrate the diverse genomic data to identify the regulatory modules of miRNAs and genes. RESULTS: Here we propose an effective data integration framework to identify the miRNA-gene regulatory comodules. The miRNA and gene expression profiles are jointly analyzed in a multiple non-negative matrix factorization framework, and additional network data are simultaneously integrated in a regularized manner. Meanwhile, we employ the sparsity penalties to the variables to achieve modular solutions. The mathematical formulation can be effectively solved by an iterative multiplicative updating algorithm. We apply the proposed method to integrate a set of heterogeneous data sources including the expression profiles of miRNAs and genes on 385 human ovarian cancer samples, computationally predicted miRNA-gene interactions, and gene-gene interactions. We demonstrate that the miRNAs and genes in 69% of the regulatory comodules are significantly associated. Moreover, the comodules are significantly enriched in known functional sets such as miRNA clusters, GO biological processes and KEGG pathways, respectively. Furthermore, many miRNAs and genes in the comodules are related with various cancers including ovarian cancer. Finally, we show that comodules can stratify patients (samples) into groups with significant clinical characteristics. AVAILABILITY: The program and supplementary materials are available at http://zhoulab.usc.edu/SNMNMF/. CONTACT: [email protected]; [email protected] Qingjiao Li, Juan Liu 0007, Xianghong Jasmine Zhou |
Bioinform. | 3 |
| 2011 | Prediction of conformational B-cell epitopes from 3D structures by random forest with a distance-based featureabstractBACKGROUND: Antigen-antibody interactions are key events in immune system, which provide important clues to the immune processes and responses. In Antigen-antibody interactions, the specific sites on the antigens that are directly bound by the B-cell produced antibodies are well known as B-cell epitopes. The identification of epitopes is a hot topic in bioinformatics because of their potential use in the epitope-based drug design. Although most B-cell epitopes are discontinuous (or conformational), insufficient effort has been put into the conformational epitope prediction, and the performance of existing methods is far from satisfaction. RESULTS: In order to develop the high-accuracy model, we focus on some possible aspects concerning the prediction performance, including the impact of interior residues, different contributions of adjacent residues, and the imbalanced data which contain much more non-epitope residues than epitope residues. In order to address above issues, we take following strategies. Firstly, a concept of 'thick surface patch' instead of 'surface patch' is introduced to describe the local spatial context of each surface residue, which considers the impact of interior residue. The comparison between the thick surface patch and the surface patch shows that interior residues contribute to the recognition of epitopes. Secondly, statistical significance of the distance distribution difference between non-epitope patches and epitope patches is observed, thus an adjacent residue distance feature is presented, which reflects the unequal contributions of adjacent residues to the location of binding sites. Thirdly, a bootstrapping and voting procedure is adopted to deal with the imbalanced dataset. Based on the above ideas, we propose a new method to identify the B-cell conformational epitopes from 3D structures by combining conventional features and the proposed feature, and the random forest (RF) algorithm is used as the classification engine. The experiments show that our method can predict conformational B-cell epitopes with high accuracy. Evaluated by leave-one-out cross validation (LOOCV), our method achieves the mean AUC value of 0.633 for the benchmark bound dataset, and the mean AUC value of 0.654 for the benchmark unbound dataset. When compared with the state-of-the-art prediction models in the independent test, our method demonstrates comparable or better performance. CONCLUSIONS: Our method is demonstrated to be effective for the prediction of conformational epitopes. Based on the study, we develop a tool to predict the conformational epitopes from 3D structures, available at http://code.google.com/p/my-project-bpredictor/downloads/list. Wen Zhang 0008, Yi Xiong 0002, Hua Zou 0002, Xinghuo Ye, Juan Liu 0007 |
BMC Bioinform. | 6 |
| 2010 | Discovering negative correlated gene sets from integrative gene expression data for cancer prognosisabstractAlong with the emergence and development of translational biomedicine, more and more genetic information has been applied in clinical practice. In recent decade, the discovery of genetic biomarkers for cancer prognosis obtains increasing attentions and many methods have been developed. The "element" methods use one or two independent genes to judge the Boolean status of disease. The "set" methods use general genetic biomarkers to classify patients into different risks as a whole. And the advanced "sets" methods use a group of different gene sets as biomarkers. However, the existing methods always concern positive correlations among genes ignoring negative correlations. Whereas the negative regulation, negative feedback, and functional repression are actually the important clues in cancer expression profiles. Therefore, in this paper, we propose to mine negative correlated gene sets (NCGSs) from multiple datasets, and use them along with the pure positive correlated gene sets for prognosis classification. The exploring experimental results have shown the encouraging promotion of cancer prognosis accuracy with NCGSs. Tao Zeng 0003, Juan Liu 0007 |
BIBM | 3 |
| 2010 | Mixture classification model based on clinical markers for breast cancer prognosis
Tao Zeng 0003, Juan Liu 0007 |
Artif. Intell. Medicine | 2 |
| 2010 | Quantitative prediction of MHC-II binding affinity using particle swarm optimization
Wen Zhang 0008, Juan Liu 0007, Yanqing Niu |
Artif. Intell. Medicine | 2 |
| 2009 | Personalized SCORM Learning Experience Based on Rating Scale Model
Ayad R. Abbas, Juan Liu 0007 |
ISNN (1) | 2 |
| 2009 | Supporting E-Learning System with Modified Bayesian Rough Set Model
Ayad R. Abbas, Juan Liu 0007 |
ISNN (2) | 2 |
| 2009 | Quantitative prediction of MHC-II peptide binding affinity using relevance vector machine
Wen Zhang 0008, Juan Liu 0007, Yanqing Niu |
Appl. Intell. | 2 |