EDBT 2026 Demo / reviewers in the wild / expert
Dong Xu 0002
dblp:09/3493-2
· DBLP profile ↗
119ranked-venue papers
4as first author
41since 2021 · last 2026
0000-0002-4809-0514ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 93 · 3 first-author · 30 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 7 since 2021Systems, architecture and hardware · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Theory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CryoFSL: an annotation-efficient, few-shot learning framework for robust protein particle picking in cryo-electron microscopy micrographsabstractAccurate identification of protein particles in cryo-electron microscopy (cryo-EM) micrographs is crucial for high-resolution structure determination, but remains challenging due to the heavy reliance on extensive annotated datasets and the difficulty of ensuring robustness under low signal-to-noise ratio (SNR) conditions. Current approaches require large annotations and exhibit poor generalization to new protein targets. We present CryoFSL (Cryo-EM Few Shot-Learning), a novel few-shot learning framework built on Segment Anything Model 2 with lightweight adapters, enabling robust particle picking with as few as five labeled micrographs and significantly reducing the annotation burden. The framework's hierarchical adapter design supports dynamic feature modulation for low-SNR and heterogeneous conditions, resolving the trade-off between annotation burden and performance. CryoFSL surpasses both traditional template-based methods and state-of-the-art deep learning models across diverse proteins in the few-shot learning setting, achieving superior recall, precision, and 3D reconstruction resolution with minimal supervision. It maintains stability across heterogeneous micrographs and consistently detects high-quality particles with fewer false-positives. Notably, CryoFSL achieves competitive resolution in density map reconstruction with just a fraction of the particles picked by other methods, redefining efficiency and quality in cryo-EM analysis. This work paves the way for scalable, generalizable, and annotation-efficient particle-picking pipelines. The code is available at https://github.com/biplabpoudel25/CryoFSL. Biplab Poudel, Rajan Gyawali, Ashwin Dhakal, Jianlin Cheng, Dong Xu 0002 |
Briefings Bioinform. | 5 |
| 2026 | A masked generative graph representation learning framework empowering precise spatial domain identificationabstractMOTIVATION: Spatial transcriptomics (ST) enables the measurement of gene expression while preserving the spatial context of tissues. However, the sparsity of ST data leads to poor usage of gene expression and spatial information, resulting in the embeddings that are not well represented and challenging for downstream analyses. RESULTS: Here, we introduced GSG, a generative self-supervised representation learning framework for ST data that leverages a masking mechanism to learn informative representations. For spatial domain identification, GSG consistently outperformed state-of-the-art methods across benchmarking datasets, regardless of sequencing platforms. In addition, we applied GSG to an in-house human fetal heart dataset, revealing anatomically coherent spatial domains and identifying APCDD1 as an endocardial-specific marker potentially involved in congenital heart disease. Our results showcase GSG's superiority and underscore its valuable contributions to advancing ST analysis. AVAILABILITY AND IMPLEMENTATION: Our software package is available at https://github.com/keaml-Guan/GSG. Chuyao Wang, Tongdong Zhang, Shuo Liang, Meirong Du, Yanchun Liang 0001, Xin Gao 0001, Dong Xu 0002, Xiaoyue Feng, An Zeng, Renchu Guan |
Bioinform. | 11 |
| 2026 | TOM: An open-source tongue segmentation method with multi-teacher distillation and task-specific data augmentation
Biplab Poudel, Congyu Guo, Guanghui An, Xiaoting Tang, Lening Zhao, Dong Xu 0002 |
Expert Syst. Appl. | 10 |
| 2026 | Annotating Spatial Multi-Omics Spot-Level Niche Types Using Bi-View Retrieval-Augmented Generation With SpotTypeLLMabstractSpatial multi-omics techniques generate extensive spot-level profiles without accompanying spot-type labels, forcing biologists into labor-intensive manual annotation. Although large language models (LLMs) promise automated annotation, they are poorly equipped to handle high-dimensional numeric inputs, struggle to convert complex spatial-omics structures into interpretable text, and lack the specialized biological knowledge needed to avoid hallucinations. Moreover, the intrinsic sparsity and heterogeneity of spatial omics data undermine robust feature extraction and accurate spot-level niche label assignment. To address these challenges, we propose SpotTypeLLM, a framework for annotating spatial multi-omics spot-level niche types using a Bi-view Retrieval-Augmented Generation (BiRAG) tailored for LLMs. Specifically, SpotTypeLLM encodes spatial multi-omics data through scLLM-based embeddings, graph convolutional networks (GCNs), and multi-view attention module, and then selects spot-specific genes to construct the prompts for LLMs. We validate the accuracy and practical utility of our framework by comparing it with ten state-of-the-art methods on labeled spatial multi-omics datasets. The results demonstrate the effectiveness of SpotTypeLLM in overcoming the limitations of current annotation methods and enhancing spatial multi-omics analysis. Longyi Li, Liyan Dong, Decheng Li, Hanbo Liu, Yuheng Zhu, Xiyuan Mei, Hao Zhang 0064, Dong Xu 0002 |
IEEE Trans. Comput. Biol. Bioinform. | 9 |
| 2026 | Applications of Large Language Models and Prompt Optimization for Knowledge Extraction From Biological Pathway FiguresabstractRecent developments in Large Language Models (LLMs) have demonstrated remarkable capabilities for image comprehension. This study aims to automate and enhance the extraction of gene interactions from biological pathway images by integrating LLMs and a Genetic Algorithm (GA). A dataset of 200 tumor signaling pathway figures from the recent biological literature was employed to assess the performance of four AI chatbots: GPT-4oV, Claude-3.5V, Gemini-1.5V, and Llama-3.2V, with GA used to optimize prompts for each model. Model performance was evaluated on both directional and non-directional gene relationship extraction. GA-optimized prompts significantly improved extraction accuracies across all LLMs, with GPT 4oV achieving an F1-score of 0.645 (±0.055) and Llama-3.2V achieving an F1-score of 0.616 (±0.068). For non-directional interactions, GPT-4oV outperformed other models, reaching a precision of 0.805, a recall of 0.695, and an F1 score of 0.757, followed by Llama-3.2V and Claude-3.5V with F1-scores of 0.702 and 0.697, respectively, while Gemini-1.5V lagged with 0.612. In directional interaction predictions, all models performed lower, with GPT-4oV leading at 0.687 F1-score, followed by Llama-3.2V at 0.656, Claude-3.5V at 0.641, and Gemini-1.5V at 0.573. While these results demonstrate substantial improvements over traditional OCR-based approaches, further advances in model accuracy and explainability are needed for widespread adoption in critical biomedical applications. Nevertheless, these findings provide a valuable benchmark for the research community and a foundation for future development of specialized, fine-tuned models and scalable multimodal AI frameworks in biomedical data analysis. The source code is publicly available on https://github.com/Muh-aza/LLM_GPV. Hasanain Aldihis, Mihail Popescu, Dong Xu 0002 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Tokenizing RNA Structures: A Discrete Generative Approach to Represent the RNA Folding LandscapesabstractRNA's diverse biological functions, including gene regulation, enzymatic catalysis, signal transduction, and even therapeutic activity, ultimately depend on its three-dimensional (3D) structure. Departing from the classical task of inferring structure directly from sequence, this study investigates how to efficiently encode and generate RNA 3D structures, thereby enabling systematic exploration of the molecules' vast folding landscape. We propose a generative framework, RNA-FrameEncoder (RNA-FE). First, a Vector-Quantised Variational AutoEncoder compresses continuous RNA backbones into sequences of discrete structural tokens, “verbalizing” the geometry while suppressing high-frequency noise. Next, a Transformer language model is trained on these token sequences to capture their joint distribution, and a decoder subsequently projects the tokens back into Cartesian coordinates. This design sidesteps the complexity of explicit equivariance modelling, lowers computational cost, and markedly enhances the representational capacity. Comprehensive benchmarks show that RNA-FE can unconditionally generate realistic, novel, and designable RNA structures without any sequence or structural priors. Overall, RNA-FE provides a scalable approach for efficiently charting RNA 3D structural space and opens a data-efficient avenue for a range of downstream applications. The code and pretrained models are publicly available at: https://github.com/users/Hanbo24-bit/RNA-FE. Hanbo Liu, Hao Zhang 0064, Longyi Li, Xiyuan Mei, Enshuang Zhao, Yuheng Zhu, Yinfei Dai, Jingxun Cao, Dong Xu 0002 |
BIBM | 9 |
| 2025 | TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese MedicineabstractTraditional Chinese Medicine (TCM), as an effective alternative medicine, has been receiving increasing attention. In recent years, the rapid development of large language models (LLMs) tailored for TCM has highlighted the urgent need for an objective and comprehensive evaluation framework to assess their performance on real-world tasks. However, existing evaluation datasets are limited in scope and primarily text-based, lacking a unified and standardized multimodal question-answering (QA) benchmark. To address this issue, we introduce TCM-Ladder, the first comprehensive multimodal QA dataset specifically designed for evaluating large TCM language models. The dataset covers multiple core disciplines of TCM, including fundamental theory, diagnostics, herbal formulas, internal medicine, surgery, pharmacognosy, and pediatrics. In addition to textual content, TCM-Ladder incorporates various modalities such as images and videos. The dataset was constructed using a combination of automated and manual filtering processes and comprises over 52,000 questions. These questions include single-choice, multiple-choice, fill-in-the-blank, diagnostic dialogue, and visual comprehension tasks. We trained a reasoning model on TCM-Ladder and conducted comparative experiments against nine state-of-the-art general domain and five leading TCM-specific LLMs to evaluate their performance on the dataset. Moreover, we propose Ladder-Score, an evaluation method specifically designed for TCM question answering that effectively assesses answer quality in terms of terminology usage and semantic expression. To the best of our knowledge, this is the first work to systematically evaluate mainstream general domain and TCM-specific LLMs on a unified multimodal benchmark. The datasets and leaderboard are publicly available at https://tcmladder.com and will be continuously updated. The source code is available at https://github.com/orangeshushu/TCM-Ladder. Jiaxuan He, Ayush Vasireddy, Xiaoting Tang, Congyu Guo, Lening Zhao, Congcong Jing, Guanghui An, Dong Xu 0002 |
NeurIPS | 12 |
| 2025 | MULoc-target: Targeting peptide classification and detection using a protein language modelabstractProtein targeting, often guided by targeting peptides, is a critical biological process that directs proteins to their specific cellular destinations, ensuring proper cellular functionality and organization. Accurate classification and detection of targeting peptides are fundamental to understanding protein sorting mechanisms. This study introduces MULoc-Target, a novel deep-learning method designed to detect and classify targeting peptides in eukaryotic proteins. To support its development and evaluation, we curated a benchmark dataset comprising eight types of eukaryotic targeting peptides with manually curated annotations. Comprehensive evaluations on this dataset and external datasets from the literature demonstrate that MULoc-Target achieves state-of-the-art or competitive performance in detecting and classifying targeting peptides. Additionally, it enables the extraction of enriched motif patterns, offering valuable insights into their properties and the underlying targeting mechanisms. The identified motifs align closely with established biological features, further validating MULoc-Target's capabilities. A web server for MULoc-Target is integrated into our MULocDeep localization suite as a new toolkit, publicly accessible at https://mu-loc.org/MULoc-Target, and the inference code is available at https://github.com/yuexujiang/MULoc-Target. Yuexu Jiang, Duolin Wang, Mahdi Pourmirzaei, Negin Manshour, Farzaneh Esmaili, Weinan Zhang 0002, Ian Max Møller, Dong Xu 0002 |
Briefings Bioinform. | 11 |
| 2025 | spaLLM: enhancing spatial domain analysis in multi-omics data through large language model integrationabstractSpatial multi-omics technologies provide valuable data on gene expression from various omics in the same tissue section while preserving spatial information. However, deciphering spatial domains within spatial omics data remains challenging due to the sparse gene expression. We propose spaLLM, the first multi-omics spatial domain analysis method that integrates large language models to enhance data representation. Our method combines a pre-trained single-cell language model (scGPT) with graph neural networks and multi-view attention mechanisms to compensate for limited gene expression information in spatial omics while improving sensitivity and resolution within modalities. SpaLLM processes multiple spatial modalities, including RNA, chromatin, and protein data, potentially adapting to emerging technologies and accommodating additional modalities. Benchmarking against eight state-of-the-art methods across four different datasets and platforms demonstrates that our model consistently outperforms other advanced methods across multiple supervised evaluation metrics. The source code for spaLLM is freely available at https://github.com/liiilongyi/spaLLM. Longyi Li, Liyan Dong, Hao Zhang 0064, Dong Xu 0002 |
Briefings Bioinform. | 4 |
| 2025 | scBSP: a fast and accurate tool for identifying spatially variable features from high-resolution spatial omics dataabstractMOTIVATION: Emerging spatial omics technologies empower comprehensive exploration of biological systems from multi-omics perspectives in their native tissue location in 2D and 3D space. However, the limited sequencing depth, increasing spatial resolution, and growing spatial spots in spatial omics technologies present significant computational challenges in identifying biologically meaningful molecules with variable spatial distributions across various omics modalities. RESULTS: We introduce scBSP, an open-source, versatile, and user-friendly package for identifying spatially variable features in large-scale spatial omics data. scBSP demonstrates significantly enhanced computational efficiency, processing high-resolution spatial omics data within seconds, and exhibits robust cross-platform performance by consistently identifying spatially variable features with high reproducibility across various sequencing platforms. AVAILABILITY AND IMPLEMENTATION: scBSP is available for download from R CRAN at https://cran.r-project.org/web/packages/scBSP/index.html and PyPI at https://pypi.org/project/scbsp/. Jinpu Li, Mauminah Raina, Ricardo Melo Ferreira, Michael Eadon, Qin Ma 0003, Juexin Wang, Dong Xu 0002 |
Bioinform. | 11 |
| 2025 | Relation equivariant graph neural networks to explore the mosaic-like tissue architecture of kidney diseases on spatially resolved transcriptomicsabstractMOTIVATION: Chronic kidney disease (CKD) and acute kidney injury (AKI) are prominent public health concerns affecting more than 15% of the global population. The ongoing development of spatially resolved transcriptomics (SRT) technologies presents a promising approach for discovering the spatial distribution patterns of gene expression within diseased tissues. However, existing computational tools are predominantly calibrated and designed on the ribbon-like structure of the brain cortex, presenting considerable computational obstacles in discerning highly heterogeneous mosaic-like tissue architectures in the kidney. Consequently, timely and cost-effective acquisition of annotation and interpretation in the kidney remains a challenge in exploring the cellular and morphological changes within renal tubules and their interstitial niches. RESULTS: We present an empowered graph deep learning framework, REGNN (Relation Equivariant Graph Neural Networks), designed for SRT data analyses on heterogeneous tissue structures. To increase expressive power in the SRT lattice using graph modeling, REGNN integrates equivariance to handle n-dimensional symmetries of the spatial area, while additionally leveraging Positional Encoding to strengthen relative spatial relations of the nodes uniformly distributed in the lattice. Given the limited availability of well-labeled spatial data, this framework implements both graph autoencoder and graph self-supervised learning strategies. On heterogeneous samples from different kidney conditions, REGNN outperforms existing computational tools in identifying tissue architectures within the 10× Visium platform. This framework offers a powerful graph deep learning tool for investigating tissues within highly heterogeneous expression patterns and paves the way to pinpoint underlying pathological mechanisms that contribute to the progression of complex diseases. AVAILABILITY AND IMPLEMENTATION: REGNN is publicly available at https://github.com/Mraina99/REGNN. Mauminah Raina, Ricardo Melo Ferreira, Treyden Stansfield, Chandrima Modak, Ying-Hua Cheng, Hari Naga Sai Kiran Suryadevara, Dong Xu 0002, Michael Eadon, Qin Ma 0003, Juexin Wang |
Bioinform. | 8 |
| 2025 | Reinforcement Learning-Based Nonautoregressive Solver for Traveling Salesman ProblemsabstractThe traveling salesman problem (TSP) is a well-known combinatorial optimization problem (COP) with broad real-world applications. Recently, neural networks (NNs) have gained popularity in this research area because as shown in the literature, they provide strong heuristic solutions to TSPs. Compared to autoregressive neural approaches, nonautoregressive (NAR) networks exploit the inference parallelism to elevate inference speed but suffer from comparatively low solution quality. In this article, we propose a novel NAR model named NAR4TSP, which incorporates a specially designed architecture and an enhanced reinforcement learning (RL) strategy. To the best of our knowledge, NAR4TSP is the first TSP solver that successfully combines RL and NAR networks. The key lies in the incorporation of NAR network output decoding into the training process. NAR4TSP efficiently represents TSP-encoded information as rewards and seamlessly integrates it into RL strategies, while maintaining consistent TSP sequence constraints during both training and testing phases. Experimental results on both synthetic and real-world TSPs demonstrate that NAR4TSP outperforms five state-of-the-art (SOTA) models in terms of solution quality, inference speed, and generalization to unseen scenarios. Yubin Xiao, Di Wang 0004, Boyang Li 0001, Huanhuan Chen 0001, Wei Pang 0001, Xuan Wu 0004, Dong Xu 0002, Yanchun Liang 0001, You Zhou 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | LungX-Net: Lung Cancer Diagnosis from CT and Histopathological Images via Attention Based Multi-Level Feature Fusion NetworkabstractLung cancer ranks among the top causes of cancer mortality worldwide. However, early detection can significantly boost survival rates, potentially increasing them by as much as 70%. While deep learning algorithms have shown promise in detecting lung cancer from CT scans, current methods are often inefficient, and single-model approaches struggle with complex medical images. Leveraging diverse model features can enhance accuracy and simplify detection. This paper introduces LungX-Net, an innovative and cost-effective deep learning model aimed at improving lung cancer diagnostic accuracy using CT and histopathological images. The proposed model unifies two networks into a unified, trainable feature extraction architecture, employing model compression to preserve essential core blocks. This approach enhances the model’s ability to propagate strong and robust features while significantly reducing computational resource requirements. The model also integrates a Convolutional Block Attention Module (CBAM) to enhance feature extraction by emphasizing the most important regions, which further improves overall accuracy. The integration of these techniques optimizes diagnostic performance while strategically minimizing complexity, making LungX-Net a powerful tool for efficient and precise lung cancer diagnosis. The proposed LungX-Net was evaluated on both CT and histopathological datasets, achieving a recognition accuracy of 99.09% on the CT dataset and 99.30% on the histopathological dataset. LungX-Net demonstrated superior performance compared to previous approaches, achieving greater accuracy with fewer parameters and faster testing times, highlighting its efficiency and cost-effectiveness. Sohaib Asif, Vicky Yang Wang, Dong Xu 0002 |
BIBM | 3 |
| 2024 | Evaluation and Integration of Advanced AI Chatbots for Biological Pathway CurationabstractThe rapid expansion of biological literature presents significant challenges in manually curating pathway knowledge from images for biological and medical research. Recent advancements in AI, particularly multimodal AI chatbots like GPT-4V, Claude, and Gemini, offer promising capabilities for image understanding and extraction of biological interactions. This study evaluates the effectiveness of these AI chatbots in curating gene interactions from 50 curated tumor signaling pathway images. Our results highlight the superior performance of GPT-4V, achieving a precision of 0.778, a recall of 0.664, an F1-score of 0.716 in non-directional analysis, and a precision of 0.566, a recall of 0.471, an F1-score of 0.512 in directional analysis. Claude followed with a precision of 0.735, recall of 0.594, and F1-score of 0.6571 in non-directional analysis, and a precision of 0.500, recall of 0.400, and F1-score of 0.445 in directional analysis. Gemini recorded the lowest performance metrics across both analyses. We further developed a model fusion integrating Claude-3 chatbots and our pathway curation pipeline, achieving notable improvements. Krpns Santosh, Dong Xu 0002, Mihail Popescu |
BIBM | 4 |
| 2024 | Predicting Gene Relations with a Graph Transformer Network Integrating DNA, Protein, and Descriptive DataabstractGene relation prediction is crucial for understanding cancer pathways and developing targeted treatments. This study proposes a novel Graph Transformer Network to predict gene relations by integrating DNA sequences, amino acids sequences, and gene descriptions. Utilizing the KEGG pathway database, our model aggregates diverse gene information to enhance gene regulatory pathway predictions. We encoded DNA sequences using DNA-BERT2, protein sequences using ESM2, and generated gene descriptions using ChatGPT, subsequently encoded by Bio-BERT. These embeddings were integrated and processed through a graph transformer with multi-head and single-head attention layers. Our results demonstrate that incorporating protein and description embeddings significantly improves predictive performance, while including DNA sequences shows marginal impact. The proposed model outperforms traditional machine learning and graph neural network-based methods, achieving an F-1 score of 0.893 on identifying whether two genes interact. This study highlights the potential of integrating multi-faceted genomic data to advance gene relation prediction, offering valuable insights for precision medicine and bioinformatics. Dong Xu 0002, Richard D. Hammer, Mihail Popescu |
BIBM | 2 |
| 2024 | PNESR-DDI: An Effective Drug-Drug Interaction Prediction Model Based on Pretraining Method and Enhanced Subgraph ReconstructionabstractDrug-Drug Interaction (DDI) task plays a crucial role in clinical treatment and drug development. Recently, deep learning methods have been successfully applied for DDI prediction. However, training deep learning models always need large amount of data, while known DDIs are scarce. To address this challenge, a graph neural network-based DDI prediction model named PNESR-DDI is proposed, which compensates for the lack of DDIs by enriching drug representations. First, to obtain initial node representations that incorporate rich semantic information from the biomedical knowledge graph (KG), a link prediction pre-training method on external KG is proposed in the node embedding pre-training module. Then, considering the large scale of the KG, subgraph extraction for the target drug pairs is introduced to reduce noise and decrease computational complexity in the subgraph anchoring module. After that, the subgraph is updated, and node similarities are propagated in the subgraph reconstruction module. Based on the node similarity scores, the subgraph is pruned and reconstructed, which adjusts node representations to be more conducive to DDI prediction. Finally, the drug embeddings, subgraph representations, and drug fingerprint features are concatenated to predict DDIs. PNESRDDI is evaluated on two benchmark DDI datasets: DrugBank and TWOSIDES. Experiment results show that PNESR-DDI achieves better performance than baselines. Ablation results validate the effectiveness of the pre-training method and the adaptive subgraph reconstruction strategy. Xiaosong Han, Yanchun Liang 0001, Dong Xu 0002, Renchu Guan |
BIBM | 5 |
| 2024 | Prompt-Based Learning on Large Protein Language Models Improves Signal Peptide Prediction
Duolin Wang, Dong Xu 0002 |
RECOMB | 4 |
| 2024 | MSI-DTI: predicting drug-target interaction based on multi-source information and multi-head self-attentionabstractIdentifying drug-target interactions (DTIs) holds significant importance in drug discovery and development, playing a crucial role in various areas such as virtual screening, drug repurposing and identification of potential drug side effects. However, existing methods commonly exploit only a single type of feature from drugs and targets, suffering from miscellaneous challenges such as high sparsity and cold-start problems. We propose a novel framework called MSI-DTI (Multi-Source Information-based Drug-Target Interaction Prediction) to enhance prediction performance, which obtains feature representations from different views by integrating biometric features and knowledge graph representations from multi-source information. Our approach involves constructing a Drug-Target Knowledge Graph (DTKG), obtaining multiple feature representations from diverse information sources for SMILES sequences and amino acid sequences, incorporating network features from DTKG and performing an effective multi-source information fusion. Subsequently, we employ a multi-head self-attention mechanism coupled with residual connections to capture higher-order interaction information between sparse features while preserving lower-order information. Experimental results on DTKG and two benchmark datasets demonstrate that our MSI-DTI outperforms several state-of-the-art DTIs prediction methods, yielding more accurate and robust predictions. The source codes and datasets are publicly accessible at https://github.com/KEAML-JLU/MSI-DTI. Wenchuan Zhao, Guosheng Liu, Yanchun Liang 0001, Dong Xu 0002, Xiaoyue Feng, Renchu Guan |
Briefings Bioinform. | 5 |
| 2024 | Mapping brain development against neurological disorder using contrastive sharing
Jieqiong Lin, Ahmed Ameen Fateh, Yijang Zhuang, Guojun Yun, Adnan Zeb, Dong Xu 0002, Hongwu Zeng |
Expert Syst. Appl. | 7 |
| 2024 | Deep learning model for human-intuitive shoeprint reconstruction
Yan Wang 0028, Di Wang 0004, Wei Pang 0001, Daixi Li, You Zhou 0008, Dong Xu 0002, Sami Ur Rahman, Amin ur Rahman, Ahmed Ameen Fateh, Peiwu Qin |
Expert Syst. Appl. | 7 |
| 2024 | SEOE: an option graph based semantically embedding method for prenatal depression detection
Xiaosong Han, Mengchen Cao, Dong Xu 0002, Xiaoyue Feng, Yanchun Liang 0001, Xiaoduo Lang, Renchu Guan |
Frontiers Comput. Sci. | 3 |
| 2024 | Neural Architecture Search for Text Classification With Limited Computing Resources Using Efficient Cartesian Genetic ProgrammingabstractCartesian Genetic Programming (CGP) has often been applied for Neural Architecture Search (NAS). However, the performance of CGP is less than ideal when searching for architectures with limited computing resources. To better facilitate NAS with limited computing resources, this paper proposes a crossover operator, a light-weighted age mechanism, and two adaptive mutation operators as the novel components in our Efficient Cartesian Genetic Programming (ECGP) method. To assess the performance of ECGP, we conduct extensive experiments on three text classification task datasets. The experimental results demonstrate that ECGP outperforms other NAS methods, requiring only hundreds of fitness evaluations to find architectures with competitive accuracy compared with human-designed models. Additionally, the ECGP-evolved architectures are shown as converging fast and stably, and having high-level transferability with merely a 1-2% accuracy drop. Ablation studies demonstrate the effectiveness of the proposed operators and age mechanism, and identify GRU as the most critical function in the text classification task. Finally, we summarize three design principles observed from the ECGP-evolved architectures that are in line with human-design strategies. To the best of our knowledge, this work introduces the first attention-derived NAS benchmark for the text classification task. Xuan Wu 0004, Di Wang 0004, Huanhuan Chen 0001, Lele Yan, Yubin Xiao, Chunyan Miao, Hong-Wei Ge, Dong Xu 0002, Yanchun Liang 0001, Kangping Wang, Chunguo Wu, You Zhou 0008 |
IEEE Trans. Evol. Comput. | 8 |
| 2024 | pathCLIP: Detection of Genes and Gene Relations From Biological Pathway Figures Through Image-Text Contrastive LearningabstractIn biomedical literature, biological pathways are commonly described through a combination of images and text. These pathways contain valuable information, including genes and their relationships, which provide insight into biological mechanisms and precision medicine. Curating pathway information across the literature enables the integration of this information to build a comprehensive knowledge base. While some studies have extracted pathway information from images and text independently, they often overlook the correspondence between the two modalities. In this paper, we present a pathway figure curation system named pathCLIP for identifying genes and gene relations from pathway figures. Our key innovation is the use of an image-text contrastive learning model to learn coordinated embeddings of image snippets and text descriptions of genes and gene relations, thereby improving curation. Our validation results, using pathway figures from PubMed, showed that our multimodal model outperforms models using only a single modality. Additionally, our system effectively curates genes and gene relations from multiple literature sources. Two case studies on extracting pathway information from literature of non-small cell lung cancer and Alzheimer's disease further demonstrate the usefulness of our curated pathway information in enhancing related pathways in the KEGG database. Fei He 0003, Richard D. Hammer, Dong Xu 0002, Mihail Popescu |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Semi-supervised active learning hypothesis verification for improved geometric expression in three-dimensional object recognition
Tingyuan Nie, Dong Xu 0002 |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Large AI Models in Health Informatics: Applications, Challenges, and the FutureabstractLarge AI models, or foundation models, are models recently emerging with massive scales both parameter-wise and data-wise, the magnitudes of which can reach beyond billions. Once pretrained, large AI models demonstrate impressive performance in various downstream tasks. A prime example is ChatGPT, whose capability has compelled people's imagination about the far-reaching influence that large AI models can have and their potential to transform different domains of our lives. In health informatics, the advent of large AI models has brought new paradigms for the design of methodologies. The scale of multi-modal data in the biomedical and health domain has been ever-expanding especially since the community embraced the era of deep learning, which provides the ground to develop, validate, and advance large AI models for breakthroughs in health-related areas. This article presents a comprehensive review of large AI models, from background to their applications. We identify seven key sectors in which large AI models are applicable and might have substantial influence, including: 1) bioinformatics; 2) medical diagnosis; 3) medical imaging; 4) medical informatics; 5) medical education; 6) public health; and 7) medical robotics. We examine their challenges, followed by a critical discussion about potential future directions and pitfalls of large AI models in transforming the field of health informatics. Jianing Qiu, Lin Li 0070, Jiankai Sun, Jiachuan Peng, Peilun Shi, Ruiyang Zhang, Yinzhao Dong, Kyle Lam, Frank P.-W. Lo, Bo Xiao 0002, Wu Yuan 0001, Ningli Wang, Dong Xu 0002, Benny P. L. Lo |
IEEE J. Biomed. Health Informatics | 13 |
| 2023 | Domain-Specific Topic Model for Knowledge Discovery in Computational and Data-Intensive Scientific CommunitiesabstractShortened time to knowledge discovery and adapting prior domain knowledge is a challenge for computational and data-intensive communities such as e.g., bioinformatics and neuroscience. The challenge for a domain scientist lies in the actions to obtain guidance through query of massive information from diverse text corpus comprising of a wide-ranging set of topics when: investigating new methods, developing new tools, or integrating datasets. In this paper, we propose a novel "domain-specific topic model" (DSTM) to discover latent knowledge patterns about relationships among research topics, tools and datasets from exemplary scientific domains. Our DSTM is a generative model that extends the Latent Dirichlet Allocation (LDA) model and uses the Markov chain Monte Carlo (MCMC) algorithm to infer latent patterns within a specific domain in an unsupervised manner. We apply our DSTM to large collections of data from bioinformatics and neuroscience domains that include more than 25,000 of papers over the last ten years, featuring hundreds of tools and datasets that are commonly used in relevant studies. Evaluation experiments based on generalization and information retrieval metrics show that our model has better performance than the state-of-the-art baseline models for discovering highly-specific latent topics within a domain. Lastly, we demonstrate applications that benefit from our DSTM to discover intra-domain, cross-domain and trend knowledge patterns. Yuanxun Zhang, Prasad Calyam, Trupti Joshi, Satish S. Nair, Dong Xu 0002 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Comprehensive Assessment of OCR Tools for Gene Name Recognition in Biological Pathway FiguresabstractOptical Character Recognition (OCR) is becoming more and more effective in text detection in images. However, OCR’s performance in special applications may vary. In particular, OCR in visual representations of complex processes known as pathway figures in the biomedical literature is challenging. The information depicted in a pathway graphic usually represents the article’s most important conclusions. Still, the huge number of pathway figures cannot be automatically processed for large-scale search, data mining, and downstream analysis. Assisted by recent developments in OCR, we have developed a method to extract gene names from pathway images. For usage in the method, we thoroughly evaluated and compared major available OCR tools using 563 genes from 45 pathway images, 1000 images of alphanumeric characters and gene names from HUGO, and KEGG data with 20 random pathway genes curated routes. Our study showed that Google Cloud Vision and MMOCR are best suitable for gene name recognition in pathway figures. Stuart Aldrich, Micheal Olaolu Arowolo, Fei He 0003, Mihail Popescu, Dong Xu 0002 |
BIBM | 5 |
| 2022 | Integration of Gene Regulatory Pathways Found in the Literature in a Graph DatabaseabstractGene regulatory pathways plays a significant role in personalized medicine. Because pathways often branch or converge, graphs are better at describing them. In this work, we introduced Neo4j, a graph database, to represent gene pathways collected from the KEGG Pathway database and from 217 non-small cell lung cancer (NSCLC) articles retrieved from PubMed. We found that the graph representation of disease pathways was able to display regulatory relations between two genes even though they were not mentioned in the same article. Besides, contradictory pathway relations, for example “EGFR activates RAS” and “EGFR inhibits RAS”, self-activation or self-inhibition like “AKT activates AKT”, and pathways with different direction like “EGFR activates RAS” and “RAS activates EGFR” may also be investigated. This contradictory information may come from different conditions of the relations, different scientific opinions or they be mistakes made by our deep learning pipeline that extracts information from the literature. Overall, we concluded that graph representations of disease pathways can give us a panoramic view of the pathway knowledge obtained from various kinds of sources. Although the representation may not manifest real underlying mechanisms, it helps provide research focus and propel the development of personalized medicine. Dong Xu 0002, Mihail Popescu |
BIBM | 2 |
| 2022 | Extraction of Gene Regulatory Relation Using BioBERTabstractRelation Extraction (RE) is a critical task typically carried out after Named Entity recognition for identifying gene-gene association from scientific publication. Current state-of the-art tools have limited capacity as most of them only extract entity relations from abstract texts. The retrieved gene-gene relations typically do not cover gene regulatory relations. There still exists a lot of room for improvement. In this work, we propose GeREx, a transformer based RE tool for identifying gene regulatory relations from full texts by incorporating labeling from pathway figures. GeREx achieved an F1-Score of 83.45% when evaluated on an independent test dataset. Clement Essien, Fei He 0003, Mark Hannink, Mihail Popescu, Dong Xu 0002 |
BIBM | 5 |
| 2022 | Machine learning development environment for single-cell sequencing data analysesabstractMachine learning (ML) is transforming single-cell sequencing data analysis; however, the barriers of technology complexity and biology knowledge remain challenging for the involvement of the ML community in single-cell data analysis. Here we present an ML development environment for single-cell sequencing data analyses, together with a diverse set of realistic and accessible ML-Ready benchmark datasets. A cloud-based platform is built to dynamically scale workflows for collecting, processing, and managing various single-cell sequencing data to make them ML-ready. In addition, benchmarks for each problem formulation and a code-level and web-interface IDE for single-cell analysis method development are provided. These efforts provide an automated end-to-end single-cell analysis ML pipeline that simplifies and standardizes the process of single-cell data formatting, loading, model development, and model evaluation. Yuexu Jiang, Cankun Wang, Clement Essien, Juexin Wang, Anjun Ma, Qin Ma 0003, Dong Xu 0002 |
BIBM | 8 |
| 2022 | Evaluating template-based and template-free protein-peptide complex structure prediction using AlphaFold2abstractProtein-peptide interactions play a crucial role in a wide range of biological processes, such as cellular regulation, immune responses, and signal transduction. Understanding the details of these interactions and predicting their complex structures is a vital key to peptide-based drug design. Predicting protein-peptide docking structure has experienced impressive scientific momentum over the past few years, which plays a crucial role in designing and developing peptide drugs. A wide range of computational algorithms has been developed for protein-peptide-docking predictions, but almost all of them require experimental protein structures, which are expensive and time-consuming. Benefiting from Alphafold2 and RoseTTAFold tools, most protein structures can be accurately deciphered based on protein sequences only. Formulating protein-peptide binding/docking as a protein complex folding problem, we can use Alphafold2 to generate a series of bounded protein-peptide conformations. We have designed a pipeline for predicting protein-peptide complexes and scoring the predicted models using AI-based tools. We benchmarked the pipeline by a set of non-redundant protein-peptide complex structures derived from the most recent-released complexes in PDB during the 2021 and 2022 years. Both template-based and template-free methods of Alphafold2 were specifically analyzed in protein-peptide complex structure predictions, and the strengths and weaknesses of each method were identified. We showed that the near-native complex structures were often not obtained by both of them and suggested that integrating the results of both template-based and template-free methods could be a good strategy. We compared the results with several protein-peptide tools, such as RoseTTAFold. We also evaluated several scoring schemes, including our in-house method based on a graph neural network, in ranking protein-peptide binding conformations. Negin Manshour, Wenyuan Qin, Fei He 0003, Duolin Wang, Dong Xu 0002 |
BIBM | 6 |
| 2022 | Predicting Compound-Protein Interaction by Deepening the Systemic Background via Molecular Network Feature EmbeddingabstractIdentifying compound-protein interactions (CPI) is crucial for drug screening, drug repurposing, and combination therapy studies. The performance of CPI prediction depends heavily on the features extracted from compounds and target proteins. The existing prediction methods use different feature combinations, but both molecular-based and network-based models have the problem of incomplete feature representations. Therefore, completely integrating the relevant features of CPI would be an effective way to solve the existing problem. This study proposed a novel model named MCPI, which integrated the PPI (protein-protein interaction) network, CCI (compound-compound interaction) network, and structure features of CPI to improve prediction performance. We compared our model with other existing methods for predicting CPI on public datasets. The experimental results showed that MCPI outperformed the peer methods. In addition, in response to the SARS-CoV-2 pandemic, we applied the model to search for potential inhibitors among FDA-approved drugs and validated the prediction results through the literature. This work may also provide potential guidance for drug development. Han Wang 0028, Hangxu Zhu, Ming Liu 0024, Dong Xu 0002 |
BIBM | 6 |
| 2022 | Discovering trends and hotspots of biosafety and biosecurity research via machine learningabstractCoronavirus disease 2019 (COVID-19) has infected hundreds of millions of people and killed millions of them. As an RNA virus, COVID-19 is more susceptible to variation than other viruses. Many problems involved in this epidemic have made biosafety and biosecurity (hereafter collectively referred to as 'biosafety') a popular and timely topic globally. Biosafety research covers a broad and diverse range of topics, and it is important to quickly identify hotspots and trends in biosafety research through big data analysis. However, the data-driven literature on biosafety research discovery is quite scant. We developed a novel topic model based on latent Dirichlet allocation, affinity propagation clustering and the PageRank algorithm (LDAPR) to extract knowledge from biosafety research publications from 2011 to 2020. Then, we conducted hotspot and trend analysis with LDAPR and carried out further studies, including annual hot topic extraction, a 10-year keyword evolution trend analysis, topic map construction, hot region discovery and fine-grained correlation analysis of interdisciplinary research topic trends. These analyses revealed valuable information that can guide epidemic prevention work: (1) the research enthusiasm over a certain infectious disease not only is related to its epidemic characteristics but also is affected by the progress of research on other diseases, and (2) infectious diseases are not only strongly related to their corresponding microorganisms but also potentially related to other specific microorganisms. The detailed experimental results and our code are available at https://github.com/KEAML-JLU/Biosafety-analysis. Renchu Guan, Haoyu Pang, Yanchun Liang 0001, Zhongjun Shao, Xin Gao 0001, Dong Xu 0002, Xiaoyue Feng |
Briefings Bioinform. | 6 |
| 2022 | Assessing deep learning methods in cis-regulatory motif finding based on genomic sequencing dataabstractIdentifying cis-regulatory motifs from genomic sequencing data (e.g. ChIP-seq and CLIP-seq) is crucial in identifying transcription factor (TF) binding sites and inferring gene regulatory mechanisms for any organism. Since 2015, deep learning (DL) methods have been widely applied to identify TF binding sites and predict motif patterns, with the strengths of offering a scalable, flexible and unified computational approach for highly accurate predictions. As far as we know, 20 DL methods have been developed. However, without a clear and systematic assessment, users will struggle to choose the most appropriate tool for their specific studies. In this manuscript, we evaluated 20 DL methods for cis-regulatory motif prediction using 690 ENCODE ChIP-seq, 126 cancer ChIP-seq and 55 RNA CLIP-seq data. Four metrics were investigated, including the accuracy of motif finding, the performance of DNA/RNA sequence classification, algorithm scalability and tool usability. The assessment results demonstrated the high complementarity of the existing DL methods. It was determined that the most suitable model should primarily depend on the data size and type and the method's outputs. Shuangquan Zhang, Anjun Ma, Dong Xu 0002, Qin Ma 0003, Yan Wang 0028 |
Briefings Bioinform. | 4 |
| 2022 | scGNN 2.0: a graph neural network tool for imputation and clustering of single-cell RNA-Seq dataabstractMOTIVATION: Gene expression imputation has been an essential step of the single-cell RNA-Seq data analysis workflow. Among several deep-learning methods, the debut of scGNN gained substantial recognition in 2021 for its superior performance and the ability to produce a cell-cell graph. However, the implementation of scGNN was relatively time-consuming and its performance could still be optimized. RESULTS: The implementation of scGNN 2.0 is significantly faster than scGNN thanks to a simplified close-loop architecture. For all eight datasets, cell clustering performance was increased by 85.02% on average in terms of adjusted rand index, and the imputation Median L1 Error was reduced by 67.94% on average. With the built-in visualizations, users can quickly assess the imputation and cell clustering results, compare against benchmarks and interpret the cell-cell interaction. The expanded input and output formats also pave the way for custom workflows that integrate scGNN 2.0 with other scRNA-Seq toolkits on both Python and R platforms. AVAILABILITY AND IMPLEMENTATION: scGNN 2.0 is implemented in Python (as of version 3.8) with the source code available at https://github.com/OSU-BMBL/scGNN2.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Haocheng Gu, Anjun Ma, Yang Li 0089, Juexin Wang, Dong Xu 0002, Qin Ma 0003 |
Bioinform. | 6 |
| 2022 | Restorable-inpainting: A novel deep learning approach for shoeprint restoration
Yan Wang 0028, Di Wang 0004, Wei Pang 0001, Kangping Wang, Daixi Li, You Zhou 0008, Dong Xu 0002 |
Inf. Sci. | 8 |
| 2022 | An Improved Topology Prediction of Alpha-Helical Transmembrane Protein Based on Deep Multi-Scale Convolutional Neural NetworkabstractAlpha-helical proteins ( αTMPs) are essential in various biological processes. Despite their tertiary structures are crucial for revealing complex functions, experimental structure determination remains challenging and costly. In the past decades, various sequence-based topology prediction methods have been developed to bridge the gap between the sequences and structures by characterizing the structural features, but significant improvements are still required. Deep learning brings a great opportunity for its powerful representation learning capability from limited original data. In this work, we improved our αTMP topology prediction method DMCTOP using deep learning, which composed of two deep convolutional blocks to simultaneously extract local and global contextual features. Consequently, the inputs were simplified to reflect the original features of the sequence, including a protein sequence feature and an evolutionary conservation feature. DMCTOP can efficiently and accurately identify all topological types and the N-terminal orientation for an αTMP sequence. To validate the effectiveness of our method, we benchmarked DMCTOP against 13 peer methods according to the whole sequence, the transmembrane segment and the traditional criterion in testing experiments. All the results reveal that our method achieved the highest prediction accuracy and outperformed all the previous methods. The method is available at https://icdtools.nenu.edu.cn/dmctop. Jiawen Yu, Zhe Liu 0030, Han Wang 0028, Zhiqiang Ma 0003, Dong Xu 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2021 | Identifying Genes and Their Interactions from Pathway Figures and Text in Biomedical ArticlesabstractMany high-quality biological pathways are presented in figures and text in biomedical literature. They are great resources for studies of biological mechanisms and precision medicine practices. These pathways need to be carefully curated, reconciled, and transformed into a computable form. Current manual curation approaches are inadequate in keeping up with the pace of the literature growth. New bio-curation approaches are needed to streamline the identification of gene interactions from pathway figures and text. This paper proposes a pathway curation approach for identifying genes and their interactions using both figures and text of biomedical articles. Our method integrates deep learning-based object detection models with a Google optical character recognition service to extract genes and their interactions from pathway figures. Our pipeline was evaluated on the figures from PubMed publications with manual annotations. The results demonstrated that our model could effectively retrieve genes and their interactions in pathway figures. The proposed pipeline may accelerate various applications of the latest biomedical discoveries. We also developed a web server at http://pathwaydeep.top to provide the gene interaction curation on uploaded pathway figures and corresponding articles. Fei He 0005, Joshua Thompson, Ziting Mao, Yijie Ren, Yulia I. Nussbaum, Olha Kholod, Dmitriy Shin, Mark Hannink, Mihail Popescu, Dong Xu 0002 |
BIBM | 10 |
| 2021 | Acupuncture and Tuina Knowledge Graph for Ancient Literature of Traditional Chinese MedicineabstractThe Traditional Chinese Medicine’s ancient literature recorded the massive medical theories and abundant medical experiences. To better understand and utilize, the knowledge from the literature, the Acupuncture and Tuina Knowledge Graph is proposed in this paper. Meanwhile, a deep learning network is established for acupuncture and tuina-related entity recognition and entity-relationship extraction. Finally, the trained network is able to reach an 82%+ F1-score for NER and 70%+ F1-score for relationship extraction. Xiaosong Han, Yanchun Liang 0001, Dong Xu 0002, Renchu Guan |
BIBM | 5 |
| 2021 | Recommender-as-a-service with chatbot guided domain-science knowledge discovery in a science gatewayabstractScientists in disciplines such as neuroscience and bioinformatics are increasingly relying on science gateways for experimentation on voluminous data, as well as analysis and visualization in multiple perspectives. Though current science gateways provide easy access to computing resources, datasets and tools specific to the disciplines, scientists often use slow and tedious manual efforts to perform knowledge discovery to accomplish their research/education tasks. Recommender systems can provide expert guidance and can help them to navigate and discover relevant publications, tools, data sets, or even automate cloud resource configurations suitable for a given scientific task. To realize the potential of integration of recommenders in science gateways in order to spur research productivity, we present a novel "OnTimeRecommend" recommender system. The OnTimeRecommend comprises of several integrated recommender modules implemented as microservices that can be augmented to a science gateway in the form of a recommender-as-a-service. The guidance for use of the recommender modules in a science gateway is aided by a chatbot plug-in viz., Vidura Advisor. To validate our OnTimeRecommend, we integrate and show benefits for both novice and expert users in domain-specific knowledge discovery within two exemplar science gateways, one in neuroscience (CyNeuro) and the other in bioinformatics (KBCommons). Komal Bhupendra Vekaria, Prasad Calyam, Sai Swathi Sivarathri, Songjie Wang, Yuanxun Zhang, Dong Xu 0002, Trupti Joshi, Satish S. Nair |
Concurr. Comput. Pract. Exp. | 8 |
| 2021 | Multi-Cloud Performance and Security Driven Federated Workflow ManagementabstractFederated multi-cloud resource allocation for data-intensive application workflows is generally performed based on performance or quality of service (i.e., QSpecs) considerations. At the same time, end-to-end security requirements of these workflows across multiple domains are considered as an afterthought due to lack of standardized formalization methods. Consequently, diverse/heterogenous domain resource and security policies cause inter-conflicts between application's security and performance requirements that lead to sub-optimal resource allocations. In this paper, we present a joint performance and security-driven federated resource allocation scheme for data-intensive scientific applications. In order to aid joint resource brokering among multi-cloud domains with diverse/heterogenous security postures, we first define and characterize a data-intensive application's security specifications (i.e., SSpecs). Then we describe an alignment technique inspired by Portunes Algebra to homogenize the various domain resource policies (i.e., RSpecs) along an application's workflow lifecycle stages. Using such formalization and alignment, we propose a near optimal cost-aware joint QSpecs-SSpecs-driven, RSpecs-compliant resource allocation algorithm for multi-cloud computing resource domain/location selection as well as network path selection. We implement our security formalization, alignment, and allocation scheme as a framework, viz., “OnTimeURB” and validate it in a multi-cloud environment with exemplar data-intensive application workflows involving distributed computing and remote instrumentation use cases with different performance and security requirements. Matthew Dickinson, Saptarshi Debroy, Prasad Calyam, Samaikya Valluripally, Yuanxun Zhang, Ronny Bazan Antequera, Trupti Joshi, Tommi A. White, Dong Xu 0002 |
IEEE Trans. Cloud Comput. | 9 |
| 2020 | MUFold-SSW: a new web server for predicting protein secondary structures, torsion angles and turnsabstractMOTIVATION: Protein secondary structure and backbone torsion angle prediction can provide important information for predicting protein 3D structures and protein functions. Our new methods MUFold-SS, MUFold-Angle, MUFold-BetaTurn and MUFold-GammaTurn, developed based on advanced deep neural networks, achieved state-of-the-art performance for predicting secondary structures, backbone torsion angles, beta-turns and gamma-turns, respectively. An easy-to-use web service will provide the community a convenient way to use these methods for research and development. RESULTS: MUFold-SSW, a new web server, is presented. It provides predictions of protein secondary structures, torsion angles, beta-turns and gamma-turns for a given protein sequence. This server implements MUFold-SS, MUFold-Angle, MUFold-BetaTurn and MUFold-GammaTurn, which performed well for both easy targets (proteins with weak sequence similarity in PDB) and hard targets (proteins without detectable similarity in PDB) in various experimental tests, achieving results better than or comparable with those of existing methods. AVAILABILITY AND IMPLEMENTATION: MUFold-SSW is accessible at http://mufold.org/mufold-ss-angle. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dong Xu 0002, Yi Shang |
Bioinform. | 3 |
| 2020 | A dynamic programing approach to integrate gene expression data and network information for pathway model generationabstractMOTIVATION: As large amounts of biological data continue to be rapidly generated, a major focus of bioinformatics research has been aimed toward integrating these data to identify active pathways or modules under certain experimental conditions or phenotypes. Although biologically significant modules can often be detected globally by many existing methods, it is often hard to interpret or make use of the results toward pathway model generation and testing. RESULTS: To address this gap, we have developed the IMPRes algorithm, a new step-wise active pathway detection method using a dynamic programing approach. IMPRes takes advantage of the existing pathway interaction knowledge in Kyoto Encyclopedia of Genes and Genomes. Omics data are then used to assign penalties to genes, interactions and pathways. Finally, starting from one or multiple seed genes, a shortest path algorithm is applied to detect downstream pathways that best explain the gene expression data. Since dynamic programing enables the detection one step at a time, it is easy for researchers to trace the pathways, which may lead to more accurate drug design and more effective treatment strategies. The evaluation experiments conducted on three yeast datasets have shown that IMPRes can achieve competitive or better performance than other state-of-the-art methods. Furthermore, a case study on human lung cancer dataset was performed and we provided several insights on genes and mechanisms involved in lung cancer, which had not been discovered before. AVAILABILITY AND IMPLEMENTATION: IMPRes visualization tool is available via web server at http://digbio.missouri.edu/impres. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuexu Jiang, Yanchun Liang 0001, Duolin Wang, Dong Xu 0002, Trupti Joshi |
Bioinform. | 4 |
| 2020 | Two New Heuristic Methods for Protein Model Quality AssessmentabstractProtein tertiary structure prediction is an important open challenge in bioinformatics and requires effective methods to accurately evaluate the quality of protein 3-D models generated computationally. Many quality assessment (QA) methods have been proposed over the past three decades. However, the accuracy or robustness is unsatisfactory for practical applications. In this paper, two new heuristic QA methods are proposed: MUfoldQA_S and MUfoldQA_C. The MUfoldQA_S is a quasi-single-model QA method that assesses the model quality based on the known protein structures with similar sequences. This algorithm can be directly applied to protein fragments without the necessity of building a full structural model. A BLOSUM-based heuristic is also introduced to help differentiate accurate templates from poor ones. In MUfoldQA_C, the ideas from MUfoldQA_S were combined with the consensus approach to create a multi-model QA method that could also utilize information from existing reference models and have demonstrated improved performance. Extensive experimental results of these two methods have shown significant improvement over existing methods. In addition, both methods have been blindly tested in the CASP12 world-wide competition in the protein structure prediction field and ranked as top performers in their respective categories. Wenbo Wang 0009, Dong Xu 0002, Yi Shang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Fuzzified Image Enhancement for Deep Learning in Iris RecognitionabstractDeep learning techniques such as convolutional neural network and capsule network have attained good results in iris recognition. However, due to the influence of eyelashes, skin, and background noises, the model often needs many iterations to retrieve informative iris patterns. Also because of some nonideal situations, such as reflection of glasses and facula on the eyeball, it is hard to detect the boundary of pupil and iris perfectly. Under such a circumstance, discarding the rest parts beyond the boundary may cause losing useful information. Hence, we use Gaussian, triangular fuzzy average, and triangular fuzzy median smoothing filters to preprocess the image by fuzzifying the region beyond the boundary to improve the signal-to-noise ratios. We applied the enhanced images through fuzzy operations to train deep learning methods, which speeds up the process of convergence and also increases the recognition accuracy rate. The saliency maps show that fuzzified image filters make the images more informative for deep learning. The proposed fuzzy operation of images may be a robust technique in many other deep-learning applications of image processing, analysis, and prediction. Ming Liu 0024, Zhiqian Zhou, Penghui Shang, Dong Xu 0002 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2019 | An Informatics Framework for Knowledge Representation and Reconciliation of Disease Pathways
Iuliia Innokenteva, Olha Kholod, Fei He 0005, Duolin Wang, Richard D. Hammer, Dong Xu 0002, Dmitriy Shin |
AMIA | 6 |
| 2019 | Capsule Network for Predicting Zinc Binding Sites in MetalloproteinsabstractZinc is an important cofactor for various biological functions in plants and animals, which are usually associated with proteins. Zinc also plays an important role in protein structures to which it binds. Hence, it is important to predict the Zinc binding sites in these proteins to better understand the structures and functions of these proteins. Most of the existing tools developed in this domain are structure-based predictors implementing Support Vector Machines on datasets that are more than a decade old. As there is little work done to explore the use of deep learning frameworks in this problem, we propose ZinCaps, a framework based on the capsule network for predicting zinc binding site using sequence-only information on more recently compiled datasets. ZinCaps outperforms previous tools. Its source codes is freely available for download at https://github.com/clemEssien/ActiveSitePrediction. Clement Essien, Duolin Wang, Dong Xu 0002 |
BIBM | 3 |
| 2019 | Protein Ubiquitylation and Sumoylation Site Prediction Based on Ensemble and Transfer LearningabstractUbiquitylation, a typical post-translational modification (PTM), plays an important role in signal transduction, apoptosis and cell proliferation. A ubiquitylation like PTM, sumoylation also may affect gene mapping, expression and genomic replication. Over the past two decades, machine learning has been widely employed in protein ubiquitylation and sumoylation site prediction tools. These existing tools require feature engineering, but failed to provide general interpretable features and probably underutilized the growing amount of data. This prompted us to propose a deep learning-based model that integrates multiple convolution and fully-connected layers of seven supervised learning sub-models to extract deep representations from protein sequences and physico-chemical properties (PCPs). Especially, we divided PCPs into 6 clusters and customized deep networks accordingly for handling the high correlations among one cluster. A stacking ensemble strategy was applied to combine these deep representations to make prediction. Furthermore, with the advantage of transfer learning, our deep learning model can work well on protein sumoylation site prediction as well after fine-tuning. On the high-quality annotated database Swiss-Prot, our model outperformed several well-known ubiquitylation and sumoylation site prediction tools. Our code is freely available at https://github.com/ruiwcoding/DeepUbiSumoPre. Fei He 0003, Yanxin Gao, Duolin Wang, Dong Xu 0002, Xiaowei Zhao 0004 |
BIBM | 6 |
| 2019 | Extracting Molecular Entities and Their Interactions from Pathway Figures Based on Deep LearningabstractMany new molecular mechanisms of genomics, pharmacogenomics, immunology, and other fields are reflected in pathway figures and need to be curated for various applications, especially in precision medicine. Current manual curation approaches are inadequate in keeping up with the pace of biomedical literature growth. Compared with textual representations, pathway figures in the biomedical literature often contain more direct representations of the mechanisms. However, no systematic method for curating pathway figures exists in publications due to date. Here, we propose a pathway curation pipeline, which integrates a deep learning model with an optical character recognition method and an image processing strategy to capture the locations, names, and interactions of pathway entities in the figure. Our pipeline was evaluated on the figures from PubMed publications. The results demonstrate that our model can effectively retrieve molecular entities and their interactions from pathway figures at a large scale. The proposed pipeline provides a complementary approach to text-mining in biological literature mining. In future work, we will combine our method with text-mining tools to enrich extracted information and reconstruct pathway mechanisms more precisely. Fei He 0005, Duolin Wang, Iuliia Innokenteva, Olha Kholod, Dmitriy Shin, Dong Xu 0002 |
BIBM | 6 |
| 2019 | DMCTOP: Topology Prediction of Alpha-Helical Transmembrane Protein Based on Deep Multi-Scale Convolutional Neural NetworkabstractAlpha-helical transmembrane proteins ($\alpha \text{TMPs}$) belong to an important category of integral membranes. Their structures are highly valuable in relevant research, but costly to solve experimentally. Sequence-based topology prediction provides a practical computational approach to characterize the structure features. Although much progress had been made in the past decade, there is significant room for improvement in predicting the topology structure. Deep learning brings a great opportunity for its capability of mining new features from data. In this work, we propose a novel$\alpha \text{TMP}$topology prediction method DMCTOP using a Deep Multi-Scale Convolutional Neural Network (DMCNN), which composes of two deep convolutional blocks to extract local and global contextual features. Consequently, the inputs of DMCTOP is simplified to a protein sequence feature and an evolutionary conservation feature. DMCTOP can efficiently and accurately identify all topological types and the N-terminal orientation for an$\alpha \text{TMP}$sequence. In the testing experiments, the prediction accuracy was calculated according to the whole sequence, the transmembrane segments and the traditional criterion. Our state-of-the-art method achieved the highest prediction accuracy compared to all the previous methods. The standalone tool is available at https://github.com/NENUBioCompute/DMCTOP. Han Wang 0028, Jiawen Yu, Dong Xu 0002 |
BIBM | 6 |
| 2019 | Capsule network for protein post-translational modification site predictionabstractMOTIVATION: Computational methods for protein post-translational modification (PTM) site prediction provide a useful approach for studying protein functions. The prediction accuracy of the existing methods has significant room for improvement. A recent deep-learning architecture, Capsule Network (CapsNet), which can characterize the internal hierarchical representation of input data, presents a great opportunity to solve this problem, especially using small training data. RESULTS: We proposed a CapsNet for predicting protein PTM sites, including phosphorylation, N-linked glycosylation, N6-acetyllysine, methyl-arginine, S-palmitoyl-cysteine, pyrrolidone-carboxylic-acid and SUMOylation sites. The CapsNet outperformed the baseline convolutional neural network architecture MusiteDeep and other well-known tools in most cases and provided promising results for practical use, especially in learning from small training data. The capsule length also gives an accurate estimate for the confidence of the PTM prediction. We further demonstrated that the internal capsule features could be trained as a motif detector of phosphorylation sites when no kinase-specific phosphorylation labels were provided. In addition, CapsNet generates robust representations that have strong discriminant power in distinguishing kinase substrates from different kinase families. Our study sheds some light on the recognition mechanism of PTMs and applications of CapsNet on other bioinformatic problems. AVAILABILITY AND IMPLEMENTATION: The codes are free to download from https://github.com/duolinwang/CapsNet_PTM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Duolin Wang, Yanchun Liang 0001, Dong Xu 0002 |
Bioinform. | 3 |
| 2019 | Prediction of Protein Backbone Torsion Angles Using Deep Residual Inception Neural NetworksabstractPrediction of protein backbone torsion angles (Psi and Phi) can provide important information for protein structure prediction and sequence alignment. Existing methods for Psi-Phi angle prediction have significant room for improvement. In this paper, a new deep residual inception network architecture, called DeepRIN, is proposed for the prediction of Psi-Phi angles. The input to DeepRIN is a feature matrix representing a composition of physico-chemical properties of amino acids, a 20-dimensional position-specific substitution matrix (PSSM) generated by PSI-BLAST, a 30-dimensional hidden Markov Model sequence profile generated by HHBlits, and predicted eight-state secondary structure features. DeepRIN is designed based on inception networks and residual networks that have performed well on image classification and text recognition. The architecture of DeepRIN enables effective encoding of local and global interatcions between amino acids in a protein sequence to achieve accruacte prediction. Extensive experimental results show that DeepRIN outperformed the best existing tools significantly. Compared to the recently released state-of-the-art tool, SPIDER3, DeepRIN reduced the Psi angle prediction error by more than 5 degrees and the Phi angle prediction error by more than 2 degrees on average. The executable tool of DeepRIN is available for download at http://dslsrv8.cs.missouri.edu/~cf797/MUFoldAngle/. Yi Shang, Dong Xu 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2019 | New Deep Learning Methods for Protein Loop ModelingabstractComputational protein structure prediction is a long-standing challenge in bioinformatics. In the process of predicting protein 3D structures, it is common that parts of an experimental structure are missing or parts of a predicted structure need to be remodeled. The process of predicting local protein structures of particular regions is called loop modeling. In this paper, five new loop modeling methods based on machine learning techniques, called NearLooper, ConLooper, ResLooper, HyLooper1, and HyLooper2 are proposed. NearLooper is based on the nearest neighbor technique. ConLooper applies deep convolutional neural networks to predict ${\mathrm{C}}_{{{\alpha }}}$Cα atoms distance matrix as an orientation-independent representation of protein structure. ResLooper uses residual neural networks instead of deep convolutional neural networks. HyLooper1 combines the results of NearLooper and ConLooper while HyLooper2 combines NearLooper and ResLooper. Three commonly used benchmarks for loop modeling are used to compare the performance between these methods and existing state-of-the-art methods. The experiment results show promising performance in which our best method improves existing state-of-the-art methods by 28 and 54 percent of average RMSD on two datasets while being comparable on the other one. Son P. Nguyen, Dong Xu 0002, Yi Shang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | A Data Mining Approach for Biomarker Discovery Using Transcriptomics in Endometriosis
Sadia Akter, Dong Xu 0002, Susan C. Nagel, Trupti Joshi |
BIBM | 2 |
| 2018 | Integrating Gene Expression Data and Pathway Knowledge for In Silico Hypothesis Generation with IMPRes
Yuexu Jiang, Duolin Wang, Dong Xu 0002, Trupti Joshi |
BIBM | 3 |
| 2018 | Knowledge Base Commons (KBCommons) v1.0: A multi OMICS' web-based data integration framework for biological discoveries
Zhen Lyu, Siva Ratna Kumari Narisetti, Dong Xu 0002, Trupti Joshi |
BIBM | 4 |
| 2018 | Domain-specific Topic Model for Knowledge Discovery through Conversational Agents in Data Intensive Scientific CommunitiesabstractMachine learning techniques underlying Big Data analytics have the potential to benefit data intensive communities in e.g., bioinformatics and neuroscience domain sciences. Today's innovative advances in these domain communities are increasingly built upon multi-disciplinary knowledge discovery and cross-domain collaborations. Consequently, shortened time to knowledge discovery is a challenge when investigating new methods, developing new tools, or integrating datasets. The challenge for a domain scientist particularly lies in the actions to obtain guidance through query of massive information from diverse text corpus comprising of a wide-ranging set of topics. In this paper, we propose a novel "domain-specific topic model" (DSTM) that can drive conversational agents for users to discover latent knowledge patterns about relationships among research topics, tools and datasets from exemplar scientific domains. The goal of DSTM is to perform data mining to obtain meaningful guidance via a chatbot for domain scientists to choose the relevant tools or datasets pertinent to solving a computational and data intensive research problem at hand. Our DSTM is a Bayesian hierarchical model that extends the Latent Dirichlet Allocation (LDA) model and uses a Markov chain Monte Carlo algorithm to infer latent patterns within a specific domain in an unsupervised manner. We apply our DSTM to large collections of data from bioinformatics and neuroscience domains that include hundreds of papers from reputed journal archives, hundreds of tools and datasets. Through evaluation experiments with a perplexity metric, we show that our model has better generalization performance within a domain for discovering highly specific latent topics. Yuanxun Zhang, Prasad Calyam, Trupti Joshi, Satish S. Nair, Dong Xu 0002 |
IEEE BigData | 5 |
| 2018 | CarbonylDB: a curated data-resource of protein carbonylation sitesabstractMotivation: Oxidative stress and protein damage have been associated with over 200 human ailments including cancer, stroke, neuro-degenerative diseases and aging. Protein carbonylation, a chemically diverse oxidative post-translational modification, is widely considered as the biomarker for oxidative stress and protein damage. Despite their importance and extensive studies, no database/resource on carbonylated proteins/sites exists. As such information is very useful to research in biology/medicine, we have manually curated a data-resource (CarbonylDB) of experimentally-confirmed carbonylated proteins/sites. Results: The CarbonylDB currently contains 1495 carbonylated proteins and 3781 sites from 21 species, with human, rat and yeast as the top three species. We have made further analyses of these carbonylated proteins/sites and presented their occurrence and occupancy patterns. Carbonylation site data on serum albumin, in particular, provides a fine model system to understand the dynamics of oxidative protein modifications/damage. Availability and implementation: The CarbonylDB is available as a web-resource and for download at http://digbio.missouri.edu/CarbonylDB/. Supplementary information: Supplementary data are available at Bioinformatics online. R. Shyama Prasad Rao, Ning Zhang 0006, Dong Xu 0002, Ian Max Møller |
Bioinform. | 3 |
| 2018 | G2S: a web-service for annotating genomic variants on 3D protein structuresabstractMotivation: Accurately mapping and annotating genomic locations on 3D protein structures is a key step in structure-based analysis of genomic variants detected by recent large-scale sequencing efforts. There are several mapping resources currently available, but none of them provides a web API (Application Programming Interface) that supports programmatic access. Results: We present G2S, a real-time web API that provides automated mapping of genomic variants on 3D protein structures. G2S can align genomic locations of variants, protein locations, or protein sequences to protein structures and retrieve the mapped residues from structures. G2S API uses REST-inspired design and it can be used by various clients such as web browsers, command terminals, programming languages and other bioinformatics tools for bringing 3D structures into genomic variant analysis. Availability and implementation: The webserver and source codes are freely available at https://g2s.genomenexus.org. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Juexin Wang, Robert Sheridan, Selçuk Onur Sümer, Nikolaus Schultz, Dong Xu 0002, Jianjiong Gao |
Bioinform. | 5 |
| 2018 | A content-based recommender system for computer science publicationsabstractAs computer science and information technology are making broad and deep impacts on our daily lives, more and more papers are being submitted to computer science journals and conferences. To help authors decide where they should submit their manuscripts, we present the Content-based Journals & Conferences Recommender System on computer science, as well as its web service at http://www.keaml.cn/prs/ . This system recommends suitable journals or conferences with a priority order based on the abstract of a manuscript. To follow the fast development of computer science and technology, a web crawler is employed to continuously update the training set and the learning model. To achieve interactive online response, we propose an efficient hybrid model based on chi-square feature selection and softmax regression. Our test results show that, the system can achieve an accuracy of 61.37% and suggest the best journals or conferences in about 5 s on average. Yanchun Liang 0001, Dong Xu 0002, Xiaoyue Feng, Renchu Guan |
Knowl. Based Syst. | 3 |
| 2018 | ADON: Application-Driven Overlay Network-as-a-Service for Data-Intensive ScienceabstractCampuses are increasingly adopting hybrid cloud architectures for supporting data-intensive science applications that require “on-demand” resources, which are not always available locally on-site. Policies at the campus edge for handling multiple such applications competing for remote resources can cause bottlenecks across applications. These bottlenecks can be proactively avoided with pertinent profiling, monitoring and control of application flows using software-defined networking and pertinent selection of local or remote compute resources. In this paper, we present an “application-driven overlay network-as-a-service” (ADON) that manages the hybrid cloud requirements of multiple applications in a scalable and extensible manner by allowing users to specify requirements of the application that are translated into the underlying network and compute provisioning requirements. Our solution involves scheduling transit selection, a cost optimized selection of site(s) for computation and traffic engineering at the campus-edge based upon real-time policy control that ensures prioritized application performance delivery for multi-tenant traffic profiles. We validate our ADON approach through an emulation study and through a wide-area overlay network testbed implementation across two campuses. Our workflow orchestration results show the ADON effectiveness in handling temporal behavior of multi-tenant traffic burst arrivals using profiles from a diverse set of actual data-intensive applications. Ronny Bazan Antequera, Prasad Calyam, Saptarshi Debroy, Longhai Cui, Sripriya Seetharam, Matthew Dickinson, Trupti Joshi, Dong Xu 0002, Tsegereda Beyene |
IEEE Trans. Cloud Comput. | 8 |
| 2017 | A multimodal deep architecture for large-scale protein ubiquitylation site predictionabstractIn eukaryotes, protein ubiquitylation is an important type of post-translation modification, in which the ubiquitin conjugates to a substrate protein. To have a better insight of the mechanisms underlying ubiquitylation, a key step is to identify protein ubiquitylation sites. Many existing computational methods are based on feature engineering, which may lead to biased and incomplete features. Deep learning provides multiple-layer networks and non-linear mapping operations to detect potential complex patterns in a data-driven way, especially for large-scale data. It provides a promising new method to predict ubiquitylation sites. In this paper, we proposed a multimodal deep architecture for protein ubiquitylation sites prediction. First, we designed different multiple layers to extract hidden informative patterns from three modalities, namely protein fragments, physico-chemical properties, and sequence profiles. Then, the deep representations corresponding to three modalities were merged to implement the classification. On the available largest scale protein ubiquitylation site database PLMD, the performance of our proposed method was measured with 66.7% sensitivity, 66.4% specificity, 66.43% accuracy, and 0.221 MCC value. A range of comparative experiments also showed that our proposed architecture outperformed several popular protein ubiquitylation site prediction tools. Our source code is freely available at https://github.com/jiagenlee/deepUbiquitylation. Fei He 0003, Lingling Bao, Jiagen Li, Dong Xu 0002, Xiaowei Zhao 0004 |
BIBM | 5 |
| 2017 | IMPRes: Integrative MultiOmics pathway resolution algorithm and toolabstractA central goal of systems biology is to uncover the underlying functional architecture of the cell and study its mechanisms. To this end, large amounts of omics data are being rapidly generated, and a focus of bioinformatics research has been towards integrating these data to identify active pathways or modules under certain conditions. Many bioinformatics algorithms include optimization methods, statistical methods, and methods using interaction network topology attributes have been applied for this. Although biologically significant modules can often be detected globally by these methods, it is hard to interpret or make use of the results towards in silico hypothesis generation and testing. We propose a step-wise active pathway detection method (IMPRes) using a dynamic programming approach. First, we take advantage of the existing pathway interaction knowledge in KEGG to build a background network, and then starting from one or multiple receptors of a certain perturbation, we use transcriptomics data collected under these conditions to detect paths that best explain the variations of genes downstream. More other omics data will be integrated in the future. Since dynamic programming enables the detection one step a time, it is easy for biomedical researchers to trace the pathway and finally lead to more accurate drug design and more effective treatment strategies. Additionally, by adding protein-protein interactions in our method, the hypotheses that we generate do not merely utilize existing knowledge, but have potential to discover new knowledge. We have evaluated our method on a dataset of cell wall stress in yeast. The path we found highly agrees with the Cell Wall Integrity (CWI) pathway, which is the main signaling pathway involved in the regulation of cell wall stress responses. We have also compared with other methods on a yeast high osmolality stress dataset and achieved an overall better performance than some other methods. More experiments have been done on human cancer datasets and mouse datasets. Finally, the IMPRes web server is established to offer a simple interface for applying IMPRes. Users can upload their own data and obtain an interactive visualization of the resulting pathway map. Users can further filter or highlight interactions according to pathway information or relation types. All genes in the pathway map are listed with detailed annotations. The IMPRes web server is available at http://gene.rnet.missouri.edu/soykb_dev/IMPRes/. Yuexu Jiang, Yanchun Liang 0001, Duolin Wang, Dong Xu 0002, Trupti Joshi |
BIBM | 4 |
| 2017 | Effects of evolutionary pressure on histone modificationsabstractWith the advent of next-generation sequencing technologies, a considerable effort has been put into sequencing the epigenomes of different species. The efforts such as “Encode” and “Roadmap” epigenomics projects provide an opportunity to compare epigenomes across species (especially between human and mouse). This study is an effort to understand how different histone modifications vary/co-appear between orthologous regions of the two species. In this work, we have used various measures of orthologous similarity between each pair of orthologous genes and explore how histone modifications are conserved with respect to changes in these similarity measures. These measures of similarity “codon usage frequency similarity” (CUFS), Ka/Ks ratio and gene expression similarity. Our simulation indicates that evolutionary selection pressure of an orthologous pair (Ka/Ks ratio) is more strongly correlated with its histone modification than any other similarity measure. We also found that genes with low Ka/Ks have more similar histone profiles across species than the ones with high Ka/Ks, suggesting more differential regulation for genes with higher selection pressure. Saad M. Khan, Gavin C. Conant, Dong Xu 0002 |
BIBM | 3 |
| 2017 | A deep-learning framework for amidation site predictionabstractAmidation plays an important role in a variety of pathological processes and serious diseases like neural dysfunction and hypertension. However, identification of protein amidation sites through traditional experimental methods is time consuming and expensive. Existing computational methods for predicting amidation sites are based on feature extraction, which may result in incomplete or biased features. To address these issues, we proposed a novel deep-learning predictor for amidation sites. Deep learning as the cutting-edge machine learning method has the ability to automatically discover complex representations of amidation patterns from the raw protein sequences, and hence it provides a powerful tool for improvement of amidation site prediction. Penghui Shang, Duolin Wang, Dongpeng Liu, Dong Xu 0002 |
BIBM | 4 |
| 2017 | A new method for disease-related gene prioritizationabstractPrioritizing genes according to their association with a disease allows researchers to explore genes in more informed ways. Although some useful algorithms have been developed, they are based on single gene importance, gene interaction networks, or gene modules with little consideration of relative gene importance in the context of modules. In this paper, we propose to prioritize genes considering both individual genes and their affiliated modules, and utilize Gene Ontology (GO) based fuzzy measure value as well as known disease genes as heuristics. The performance of our method is comprehensively validated by using both simulated and real datasets. Results show that our method outperforms other methods in terms of disease-related gene prioritization. This work will aid researchers in the understanding of the genetic architecture of complex diseases, and improve the accuracy of diagnosis and the effectiveness of therapy. Lingtao Su, Dong Xu 0002, Guixia Liu |
BIBM | 2 |
| 2017 | SoyTSN: A web-based prediction tool for soybean tissue specific network within SoyKBabstractSoybean tissue-specific network helps identify and visualize the gene-gene relationships in various tissues [1]. We have built SoyTSN, a web-based tool for tissue-specific network prediction in soybean using 14 tissues RNA-Seq datasets including flower, root, nodule, leaf, stem, seed, etc. SoyTSN first combines multiple tissue specific RNA-Seq studies, and later Cross-Conditions Cluster Detection (C3D) algorithm [2] was applied to detect modules based on co-expression relationships across all the tissues. Following this soybean tissue-specific interactomes were inferred by combining tissue-specific expression and protein-protein interaction from the STRING database [3]. All these relationships between various soybean genes were collected and stored in Soybean Knowledge Base (SoyKB)[4] and can be queried using SoyTSN. For every query gene, SoyTSN computes and visualizes any of the 14 tissue specific networks both at the expression and interactome level. Users can compare the gene-gene relationship differences at different confidence levels across all the soybean tissues. Juexin Wang, Zhen Lyu, Shakhawat Hossain, Gary Stacey, Dong Xu 0002, Trupti Joshi |
BIBM | 5 |
| 2017 | Computational prediction of ubiquitination protein using evolutionary profiles and functional domainsabstractUbiquitination, as a post-translational modification, is a crucial biological process presented in cell signaling, death and localization. Identification of ubiquitination protein is of fundamental importance for understanding molecular mechanisms in biological systems and diseases. Although high-throughput experimental studies using mass spectrometry have identified many ubiquitination proteins and ubiquitination sites, the vast majority of ubiquitination proteins remain undiscovered, even in well studied model organisms. To reduce experimental costs, computational (in silico) methods have been introduced to predict ubiquitination sites. If we can predict whether a query protein can be ubiquitinated or not, it is meaningful by itself and helpful for predicting ubiquitination sites. However, all the computational methods so far only predict ubiquitination sites, with unsatisfactory accuracy. In this study, we developed the first computational method for predicting ubiquitination proteins without relying on ubiquitination site prediction. The method extracts features from sequence conservation information via a grey system model, as well as functional domain annotation and subcellular localization. Together with the detailed feature analysis and application of the Relief feature selection algorithm, the results of 5-fold cross-validation on three datasets achieved a high accuracy of 0.8981, with the Matthew's correlation coefficient 0.7963. Our study may guide the related experimental design and provide useful insights for studying the mechanisms and modulation of ubiquitination pathways. Wangren Qiu, Dong Xu 0002 |
BIBM | 4 |
| 2017 | A New Deep Neighbor Residual Network for Protein Secondary Structure PredictionabstractA protein secondary structure defines the local conformation of the protein's polypeptide backbone, which provides important information for protein 3D structure prediction and protein functions. In this study, a new deep neural network, the deep neighbor residual network (DeepNRN), is proposed for protein secondary structure predictions. The network takes three types of inputs, namely protein sequence features, profile features generated by PSI-BLAST, and profile features generated by HHBlits, and predicts the protein secondary structure in either one of eight states (Q8) or one of three states (Q3). The basic building block of the network, the neighbor residual unit, is designed with two types of short-cut connections that are more general and expressive than residual units in existing residual deep neural networks, yet can still be computed efficiently. In addition, the prediction result of DeepNRN can be refined by a Struct2Struct network to make the result more protein-like. Extensive experimental results on multiple widely used benchmark data sets show that the new DeepNRN-based method outperformed existing methods and obtained the best results across multiple data sets. Yi Shang, Dong Xu 0002 |
ICTAI | 3 |
| 2017 | Protein Loop Modeling Using Deep Generative Adversarial NetworkabstractBiology and medicine have a long-standing interest in computational structure prediction and modeling of proteins. There are often missing regions or regions that need to be remodeled in protein structures. The process of predicting particular missing regions in a protein structure is called loop modeling. In this paper, we propose a generative adversarial network (GAN) in deep learning for loop modeling using the idea of image inpainting. The generative network is to capture the context of the loop region and predict the missing area. The adversarial network is to make the prediction look real and provide gradients to the generative network. The proposed network was evaluated on a common benchmark for loop modeling. Experiments show that our method can successfully predict the loop region and has achieved better performance than the state-of-the-art tools. To our knowledge, this work represents the first attempt of using GAN for any bioinformatics studies. Son P. Nguyen, Dong Xu 0002, Yi Shang |
ICTAI | 3 |
| 2017 | BioJava-ModFinder: identification of protein modifications in 3D structures from the Protein Data BankabstractSUMMARY: We developed a new software tool, BioJava-ModFinder, for identifying protein modifications observed in 3D structures archived in the Protein Data Bank (PDB). Information on more than 400 types of protein modifications were collected and curated from annotations in PDB, RESID, and PSI-MOD. We divided these modifications into three categories: modified residues, attachment modifications, and cross-links. We have developed a systematic method to identify these modifications in 3D protein structures. We have integrated this package with the RCSB PDB web application and added protein modification annotations to the sequence diagram and structure display. By scanning all 3D structures in the PDB using BioJava-ModFinder, we identified more than 30 000 structures with protein modifications, which can be searched, browsed, and visualized on the RCSB PDB website. AVAILABILITY AND IMPLEMENTATION: BioJava-ModFinder is available as open source (LGPL license) at ( https://github.com/biojava/biojava/tree/master/biojava-modfinder ). The RCSB PDB can be accessed at http://www.rcsb.org . CONTACT: [email protected]. Jianjiong Gao, Andreas Prlic, Chunxiao Bi, Wolfgang Bluhm, Dong Xu 0002, Philip E. Bourne, Peter W. Rose |
Bioinform. | 6 |
| 2017 | MusiteDeep: a deep-learning framework for general and kinase-specific phosphorylation site predictionabstractMOTIVATION: Computational methods for phosphorylation site prediction play important roles in protein function studies and experimental design. Most existing methods are based on feature extraction, which may result in incomplete or biased features. Deep learning as the cutting-edge machine learning method has the ability to automatically discover complex representations of phosphorylation patterns from the raw sequences, and hence it provides a powerful tool for improvement of phosphorylation site prediction. RESULTS: We present MusiteDeep, the first deep-learning framework for predicting general and kinase-specific phosphorylation sites. MusiteDeep takes raw sequence data as input and uses convolutional neural networks with a novel two-dimensional attention mechanism. It achieves over a 50% relative improvement in the area under the precision-recall curve in general phosphorylation site prediction and obtains competitive results in kinase-specific prediction compared to other well-known tools on the benchmark data. AVAILABILITY AND IMPLEMENTATION: MusiteDeep is provided as an open-source tool available at https://github.com/duolinwang/MusiteDeep. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Duolin Wang, Wangren Qiu, Yanchun Liang 0001, Trupti Joshi, Dong Xu 0002 |
Bioinform. | 7 |
| 2016 | End-to-End Security Formalization and Alignment for Federated Workflow ManagementabstractTraditionally, the allocation and dynamic adaptation of federated cyberinfrastructure resources residing across multiple domains for data-intensive application workflows have been performance or quality of service-centric (i.e., QSpecs), often compromising the end-to-end security requirements of scientific workflows. Lack of standardized formalization methods of the workflows' end-to-end security requirements, and diverse/heterogenous domain resource and security policies make inter-conflict characterization between application's security and performance requirements non-trivial, and leads to sub-optimal resource allocation. In this paper, we present a joint security and performance-driven federated resource allocation and adaptation scheme to define and characterize a data-intensive scientific application's security specifications (i.e., SSpecs). In order to aid security-driven resource brokering among domains with diverse security postures, we describe an alignment technique inspired by Portunes Algebra to combine domain-specific resource policies (i.e., RSpecs) along the application workflow life cycle. We use standardized guidelines that help in compute/storage resource domain/location selection as well as network path selection based on both application QSpecs and SSpecs. We implement our security formalization and alignment methods as a framework, viz., "OnTimeURB" and apply it on an exemplar Distributed Computing workflow to show the benefits of joint QSpecs-SSpecs-driven, RSpecs-compliant federated workflow management. Matthew Dickinson, Saptarshi Debroy, Prasad Calyam, Samaikya Valluripally, Yuanxun Zhang, Trupti Joshi, Dong Xu 0002 |
CLOUD | 7 |
| 2016 | Complexity Reduction and Visualization of RDF Knowledge Networks for Precision medicine
Zainab Al-Taie, Nattaphon Thanintorn, Ilker Ersoy, Richard D. Hammer, Dong Xu 0002, Trupti Joshi, Dmitriy Shin |
AMIA | 5 |
| 2016 | Classification of tongue images based on doublet and color space dictionaryabstractRecently, pathological diagnosis plays a crucial role in many areas of medicine, and some researchers have proposed many models and algorithms for improving classification accuracy by extracting excellent feature or modifying the classifier. They have also achieved excellent results on pathological diagnosis using tongue images. However, pixel values can't express intuitive features of tongue images and different classifiers for training samples have different adaptability. Accordingly, this paper presents a robust approach to infer the pathological characteristics by observing tongue images. Our proposed method makes full use of the local information and similarity of tongue images. Firstly, tongue images in RGB color space are converted to Lab. Then, we compute tongue statistics information. In the calculation process, Lab space dictionary is created at first, through it, we compute statistic value for each dictionary value. After that, a method based on Doublets is taken for feature optimization. At last, we use XGBOOST classifier to predict the categories of tongue images. We achieve classification accuracy of 95.39% using statistics feature and the improved classifier, which is helpful for TCM (Traditional Chinese Medicine) diagnosis. Guitao Cao, Ye Duan, Liping Tu, Jiatuo Xu, Dong Xu 0002 |
BIBM | 6 |
| 2016 | A deep tongue image features analysis model for medical applicationabstractWith the improvement of people's living standards, there is no doubt that people are paying more and more attention to their health. However, shortage of medical resources is a critical global problem. As a result, an intelligent prognostics system has a great potential to play important roles in computer aided diagnosis. Numerous papers reported that tongue features have been closely related to a human's state. Among them, the majority of the existing tongue image analyses and classification methods are based on the low-level features, which may not provide a holistic view of the tongue. Inspired by a deep convolutional neural network (CNN), we propose a deep tongue image feature analysis system to extract unbiased features and reduce human labor for tongue diagnosis. With the unbalanced sample distribution, it is hard to form a balanced classification model based on feature representations obtained by existing low-level and high-level methods. Our proposed deep tongue image feature analysis model learns high-level features and provide more classification information during training time, which may result in higher accuracy when predicting testing samples. We tested the proposed system on a set of 267 gastritis patients, and a control group of 48 healthy volunteers (labeled according to Western medical practices). Test results show that the proposed deep tongue image feature analysis model can classify a given tongue image into healthy and diseased state with an average accuracy of 91.49%, which demonstrates the relationship between human body's state and its deep tongue image features. Guitao Cao, Ye Duan, Minghua Zhu, Liping Tu, Jiatuo Xu, Dong Xu 0002 |
BIBM | 7 |
| 2016 | FALCON@home: a high-throughput protein structure prediction server based on remote homologue recognitionabstractSUMMARY: The protein structure prediction approaches can be categorized into template-based modeling (including homology modeling and threading) and free modeling. However, the existing threading tools perform poorly on remote homologous proteins. Thus, improving fold recognition for remote homologous proteins remains a challenge. Besides, the proteome-wide structure prediction poses another challenge of increasing prediction throughput. In this study, we presented FALCON@home as a protein structure prediction server focusing on remote homologue identification. The design of FALCON@home is based on the observation that a structural template, especially for remote homologous proteins, consists of conserved regions interweaved with highly variable regions. The highly variable regions lead to vague alignments in threading approaches. Thus, FALCON@home first extracts conserved regions from each template and then aligns a query protein with conserved regions only rather than the full-length template directly. This helps avoid the vague alignments rooted in highly variable regions, improving remote homologue identification. We implemented FALCON@home using the Berkeley Open Infrastructure of Network Computing (BOINC) volunteer computing protocol. With computation power donated from over 20,000 volunteer CPUs, FALCON@home shows a throughput as high as processing of over 1000 proteins per day. In the Critical Assessment of protein Structure Prediction (CASP11), the FALCON@home-based prediction was ranked the 12th in the template-based modeling category. As an application, the structures of 880 mouse mitochondria proteins were predicted, which revealed the significant correlation between protein half-lives and protein structural factors. AVAILABILITY AND IMPLEMENTATION: FALCON@home is freely available at http://protein.ict.ac.cn/FALCON/. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Haicang Zhang, Wei-Mou Zheng, Dong Xu 0002, Jianwei Zhu, Kang Ning 0001, Shiwei Sun, Shuaicheng Li 0001, Dongbo Bu |
Bioinform. | 4 |
| 2016 | PGen: large-scale genomic variations analysis workflow and browser in SoyKBabstractBACKGROUND: With the advances in next-generation sequencing (NGS) technology and significant reductions in sequencing costs, it is now possible to sequence large collections of germplasm in crops for detecting genome-scale genetic variations and to apply the knowledge towards improvements in traits. To efficiently facilitate large-scale NGS resequencing data analysis of genomic variations, we have developed "PGen", an integrated and optimized workflow using the Extreme Science and Engineering Discovery Environment (XSEDE) high-performance computing (HPC) virtual system, iPlant cloud data storage resources and Pegasus workflow management system (Pegasus-WMS). The workflow allows users to identify single nucleotide polymorphisms (SNPs) and insertion-deletions (indels), perform SNP annotations and conduct copy number variation analyses on multiple resequencing datasets in a user-friendly and seamless way. RESULTS: We have developed both a Linux version in GitHub ( https://github.com/pegasus-isi/PGen-GenomicVariations-Workflow ) and a web-based implementation of the PGen workflow integrated within the Soybean Knowledge Base (SoyKB), ( http://soykb.org/Pegasus/index.php ). Using PGen, we identified 10,218,140 single-nucleotide polymorphisms (SNPs) and 1,398,982 indels from analysis of 106 soybean lines sequenced at 15X coverage. 297,245 non-synonymous SNPs and 3330 copy number variation (CNV) regions were identified from this analysis. SNPs identified using PGen from additional soybean resequencing projects adding to 500+ soybean germplasm lines in total have been integrated. These SNPs are being utilized for trait improvement using genotype to phenotype prediction approaches developed in-house. In order to browse and access NGS data easily, we have also developed an NGS resequencing data browser ( http://soykb.org/NGS_Resequence/NGS_index.php ) within SoyKB to provide easy access to SNP and downstream analysis results for soybean researchers. CONCLUSION: PGen workflow has been optimized for the most efficient analysis of soybean data using thorough testing and validation. This research serves as an example of best practices for development of genomics data analysis workflows by integrating remote HPC resources and efficient data management with ease of use for biological users. PGen workflow can also be easily customized for analysis of data in other species. Saad M. Khan, Juexin Wang, Mats Rynge, Yuanxun Zhang, Shiyuan Chen, João V. Maldonado dos Santos, Babu Valliyodan, Prasad Calyam, Nirav C. Merchant, Henry T. Nguyen, Dong Xu 0002, Trupti Joshi |
BMC Bioinform. | 13 |
| 2015 | Guest Editors Introduction to the Special Issue on Software and DatabasesabstractThe papers in this special section focus on software and databases that are central in bioinformatics and computational biology.. These programs are playing more and more important roles in biology and medical research. These papers cover a broad range of topics, including computational genomics and transcriptomics, analysis of biological networks and interactions, drug design, biomedical signal/image analysis, biomedical text mining and ontologies, biological data mining, visualization and integration, and high performance computing application in bioinformatics. Dong Xu 0002, Kun Huang 0001, Jeanette Schmidt |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2014 | Multi-dimensional scaling and MODELLER-based evolutionary algorithms for protein model refinementabstractProtein structure prediction, i.e., computationally predicting the three-dimensional structure of a protein from its primary sequence, is one of the most important and challenging problems in bioinformatics. Model refinement is a key step in the prediction process, where improved structures are constructed based on a pool of initially generated models. Since the refinement category was added to the biennial Critical Assessment of Structure Prediction (CASP) in 2008, CASP results show that it is a challenge for existing model refinement methods to improve model quality consistently. This paper presents three evolutionary algorithms for protein model refinement, in which multidimensional scaling(MDS), the MODELLER software, and a hybrid of both are used as crossover operators, respectively. The MDS-based method takes a purely geometrical approach and generates a child model by combining the contact maps of multiple parents. The MODELLER-based method takes a statistical and energy minimization approach, and uses the remodeling module in MODELLER program to generate new models from multiple parents. The hybrid method first generates models using the MDS-based method and then run them through the MODELLER-based method, aiming at combining the strength of both. Promising results have been obtained in experiments using CASP datasets. The MDS-based method improved the best of a pool of predicted models in terms of the global distance test score (GDT-TS) in 9 out of 16test targets. Yi Shang, Dong Xu 0002 |
IEEE Congress on Evolutionary Computation | 3 |
| 2014 | DL-PRO: A novel deep learning method for protein model quality assessmentabstractComputational protein structure prediction is very important for many applications in bioinformatics. In the process of predicting protein structures, it is essential to accurately assess the quality of generated models. Although many single-model quality assessment (QA) methods have been developed, their accuracy is not high enough for most real applications. In this paper, a new approach based on C-α atoms distance matrix and machine learning methods is proposed for single-model QA and the identification of native-like models. Different from existing energy/scoring functions and consensus approaches, this new approach is purely geometry based. Furthermore, a novel algorithm based on deep learning techniques, called DL-Pro, is proposed. For a protein model, DL-Pro uses its distance matrix that contains pairwise distances between two residues' C-α atoms in the model, which sometimes is also called contact map, as an orientation-independent representation. From training examples of distance matrices corresponding to good and bad models, DL-Pro learns a stacked autoencoder network as a classifier. In experiments on selected targets from the Critical Assessment of Structure Prediction (CASP) competition, DL-Pro obtained promising results, outperforming state-of-the-art energy/scoring functions, including OPUS-CA, DOPE, DFIRE, and RW. Son P. Nguyen, Yi Shang, Dong Xu 0002 |
IJCNN | 3 |
| 2014 | Identifying critical transitions of complex diseases based on a single sampleabstractMOTIVATION: Unlike traditional diagnosis of an existing disease state, detecting the pre-disease state just before the serious deterioration of a disease is a challenging task, because the state of the system may show little apparent change or symptoms before this critical transition during disease progression. By exploring the rich interaction information provided by high-throughput data, the dynamical network biomarker (DNB) can identify the pre-disease state, but this requires multiple samples to reach a correct diagnosis for one individual, thereby restricting its clinical application. RESULTS: In this article, we have developed a novel computational approach based on the DNB theory and differential distributions between the expressions of DNB and non-DNB molecules, which can detect the pre-disease state reliably even from a single sample taken from one individual, by compensating insufficient samples with existing datasets from population studies. Our approach has been validated by the successful identification of pre-disease samples from subjects or individuals before the emergence of disease symptoms for acute lung injury, influenza and breast cancer. Rui Liu 0009, Xiangtian Yu, Xiaoping Liu 0002, Dong Xu 0002, Kazuyuki Aihara, Luonan Chen |
Bioinform. | 4 |
| 2013 | Soybean knowledge base (SoyKB): Bridging the gap between soybean translational genomics and breedingabstractMany genome-scale data are available in soybean including genomic sequence, transcriptomics (microarray, RNA-seq), proteomics and metabolomics datasets, together with growing knowledge of soybean in gene, microRNAs, pathways, and phenotypes. This represents rich and resourceful information which can provide valuable insights, if mined in an innovative and integrative manner and thus, the need for informatics resources to achieve that. Towards this we have developed Soybean Knowledge Base (SoyKB), a comprehensive all-inclusive web resource for soybean translational genomics and breeding. SoyKB handles the management and integration of soybean genomics and multi-omics data along with gene function annotations, biological pathway and trait information. It has many useful tools including Affymetrix probelD search, gene family search, multiple gene/metabolite analysis, motif analysis tool, protein 3D structure viewer and download/upload capacity for experimental data and annotations. It has a user-friendly web interface together with genome browser and pathway viewer, which display data in an intuitive manner to the soybean researchers, breeders and consumers. SoyKB has new innovative tools for soybean breeding including a graphical chromosome visualizer targeted towards ease of navigation for breeders. It integrates QTLs, traits, germplasm information along with genomic variation data such as single nucleotide polymorphisms (SNPs) and genome-wide association studies (GWAS) data from multiple genotypes, cultivars and G. soja. QTLs for multiple traits can be queried and visualized in the chromosome visualizer simultaneously and overlaid on top of the genes and other molecular markers as well as multi-omics experimental data for meaningful inferences. SoyKB can be publicly accessed at http://soykb.org. Trupti Joshi, Michael R. Fitzpatrick, Shiyuan Chen, Ryan Z. Endacott, Eric C. Gaudiello, Gary Stacey, Henry T. Nguyen, Dong Xu 0002 |
BIBM | 10 |
| 2013 | NOA: a cytoscape plugin for network ontology analysisabstractSUMMARY: The Network Ontology Analysis (NOA) plugin for Cytoscape implements the NOA algorithm for network-based enrichment analysis, which extends Gene Ontology annotations to network links, or edges. The plugin facilitates the annotation and analysis of one or more networks in Cytoscape according to user-defined parameters. In addition to tables, the NOA plugin also presents results in the form of heatmaps and overview networks in Cytoscape, which can be exported for publication figures. AVAILABILITY: The NOA plugin is an open source, Java program for Cytoscape version 2.8 available via the Cytoscape App Store (http://apps.cytoscape.org/apps/noa) and plugin manager. A detailed user manual is available at http://nrnb.org/tools/noa. .ucsf.edu Chao Zhang 0032, Kristina Hanspers, Dong Xu 0002, Luonan Chen, Alexander R. Pico |
Bioinform. | 4 |
| 2012 | Protein Structure Prediction and Clustering Using MUFOLD
Jingfen Zhang, Dong Xu 0002 |
ISBRA | 2 |
| 2012 | Mosaic: making biological sense of complex networksabstractUNLABELLED: We present a Cytoscape plugin called Mosaic to support interactive network annotation, partitioning, layout and coloring based on gene ontology or other relevant annotations. AVAILABILITY: Mosaic is distributed for free under the Apache v2.0 open source license and can be downloaded via the Cytoscape plugin manager. A detailed user manual is available on the Mosaic web site (http://nrnb.org/tools/mosaic). Chao Zhang 0032, Kristina Hanspers, Allan Kuchinsky, Nathan Salomonis, Dong Xu 0002, Alexander R. Pico |
Bioinform. | 5 |
| 2012 | Computational Challenges in Characterization of Bacteria and Bacteria-Host Interactions Based on Genomic Data
Chao Zhang 0032, Guolu Zheng, Shunfu Xu, Dong Xu 0002 |
J. Comput. Sci. Technol. | 4 |
| 2011 | A Hybrid Consensus and Clustering Method for Protein Structure SelectionabstractIn protein tertiary structure prediction, a crucial step is to select near-native structures from a large number of predicted structural models. Over the years, many methods have been proposed for the protein structure selection problem. Despite significant advances, the discerning power of current approaches is still unsatisfactory. In this paper, we propose a new algorithm, CC-Select, that combines consensus with clustering techniques. Given a set of predicted models, CC-Select first calculates a consensus score for each structure based on its average pair wise structural similarity to other models. Then, similar structures are grouped into clusters using multidimensional scaling and clustering algorithms. In each cluster, the one with the highest consensus score is selected as a candidate model. Using extensive benchmark sets of a large collection of predicted models, we compare CC-Select with existing state-of-the-art quality assessment methods and show significant improvement. Qingguo Wang, Yi Shang, Dong Xu 0002 |
ICTAI | 3 |
| 2011 | Improving a Consensus Approach for Protein Structure Selection by Removing RedundancyabstractIn protein tertiary structure prediction, a crucial step is to select near-native structures from a large number of predicted structural models. Over the years, extensive research has been conducted for the protein structure selection problem with most approaches focusing on developing more accurate energy or scoring functions. Despite significant advances in this area, the discerning power of current approaches is still unsatisfactory. In this paper, we propose a novel consensus-based algorithm for the selection of predicted protein structures. Given a set of predicted models, our method first removes redundant structures to derive a subset of reference models. Then, a structure is ranked based on its average pairwise similarity to the reference models. Using the CASP8 data set containing a large collection of predicted models for 122 targets, we compared our method with the best CASP8 quality assessment (QA) servers, which are all consensus based, and showed that our QA scores correlate better with the GDT-TSs than those of the CASP8 QA servers. We also compared our method with the state-of-the-art scoring functions and showed its improved performance for near-native model selection. The GDT-TSs of the top models picked by our method are on average more than 8 percent better than the ones selected by the best performing scoring function. Qingguo Wang, Yi Shang, Dong Xu 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2010 | SoyMetDB: The soybean metabolome databaseabstractSoyMetDB is a metabolomic database for soybean, developed to target the growing needs of the soybean community. The goal is to provide a one-stop web resource for integrating, mining and visualizing soybean metabolomic data, including identification and expression of various metabolites across different experiments and time courses. It incorporates GC-MS and LC-MS based metabolite-profiling data dynamically linked to metabolite information from other public metabolomic databases, including HMDB and Knapsack. SoyMetDB includes Arabidopsis metabolomic data for cross-species comparisons and can retrieve information including the expression patterns of various experiments for complete or partial metabolite name queries. It also incorporates a pathway viewer tool integrating the data from various experimental conditions and presenting them on the pathways to highlight the expressed metabolite, and identifies the most highly represented pathways for multiple metabolite queries. SoyMetDB can be accessed at http://soymetdb.org. Trupti Joshi, Qiuming Yao, D. Franklin Levi, Laurent Brechenmacher, Babu Valliyodan, Gary Stacey, Henry T. Nguyen, Dong Xu 0002 |
BIBM | 8 |
| 2010 | Detection and application of CagA sequence markers for assessing risk factor of gastric cancer caused by Helicobacter pyloriabstractAs a marker of Helicobacter pylori, Cytotoxin-associated gene A (CagA) has been revealed to be the major virulence factor to cause gastroduodenal diseases. However, the molecular mechanisms that underlie the development of different gastroduodenal diseases caused by cagA-positive H. pylori infection remain unknown. Current studies are mainly limited to the relationship between EPIYA motifs in the CagA strain and diseases, but such a relationship is insufficient to explain the diversity of diseases. We propose a new and systematic method to analyze the relationship between the whole CagA sequence patterns and diseases. For this purpose, we introduced entropy calculation to detect key residues of CagA as the gastric cancer biomarkers, and then employed a supervised learning procedure to classify the cancer and non-cancer related CagA strains by using the key residues. We achieved 76% and 71% classification accuracy for Western and East Asian subtypes, respectively. Our study may help establish H. pylori biomarkers for predicting gastroduodenal disease outcome. Chao Zhang 0032, Shunfu Xu, Dong Xu 0002 |
BIBM | 3 |
| 2010 | Protein structure selection based on consensusabstractIn protein tertiary structure prediction, a crucial step is to select near-native structures from a large number of predicted structural models. Over the years, extensive research has been conducted for the protein structure selection problem with most approaches focusing on developing more accurate energy or scoring functions. Despite significant advances in this area, the discerning power of current approaches is still unsatisfactory. In this paper, we propose a novel consensus-based method that outperforms state-of-the-art scoring functions. Given a set of predicted models, the method first removes redundant models to derive a subset of reference models. Then a model is ranked based on its average pairwise similarity score to the set of reference models. Using the CASP8 data set containing a large collections of predicted models for 122 targets, we compared our method with an extensive collection of state-of-the-art scoring functions and showed its superior performance. The top models selected by our method are on average more than 8% better than the ones selected by the best competing method. Qingguo Wang, Yi Shang, Dong Xu 0002 |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | PMirP: A pre-microRNA prediction method based on structure-sequence hybrid features
Dongyu Zhao, Yan Wang 0028, Xiaohu Shi, Liupu Wang, Dong Xu 0002, Yanchun Liang 0001 |
Artif. Intell. Medicine | 6 |
| 2010 | The Musite open-source framework for phosphorylation-site predictionabstractBACKGROUND: With the rapid accumulation of phosphoproteomics data, phosphorylation-site prediction is becoming an increasingly active research area. More than a dozen phosphorylation-site prediction tools have been released in the past decade. However, there is currently no open-source framework specifically designed for phosphorylation-site prediction except Musite. RESULTS: Here we present the Musite open-source framework for building applications to perform machine learning based phosphorylation-site prediction. Musite was implemented with six modules loosely coupled with each other. With its well-designed Java application programming interface (API), Musite can be easily extended to integrate various sources of biological evidence for phosphorylation-site prediction. CONCLUSIONS: Released under the GNU GPL open source license, Musite provides an open and extensible framework for phosphorylation-site prediction. The software with its source code is available at http://musite.sourceforge.net. Jianjiong Gao, Dong Xu 0002 |
BMC Bioinform. | 2 |
| 2010 | Prediction of novel miRNAs and associated target genes in Glycine maxabstractBACKGROUND: Small non-coding RNAs (21 to 24 nucleotides) regulate a number of developmental processes in plants and animals by silencing genes using multiple mechanisms. Among these, the most conserved classes are microRNAs (miRNAs) and small interfering RNAs (siRNAs), both of which are produced by RNase III-like enzymes called Dicers. Many plant miRNAs play critical roles in nutrient homeostasis, developmental processes, abiotic stress and pathogen responses. Currently, only 70 miRNA have been identified in soybean. METHODS: We utilized Illumina's SBS sequencing technology to generate high-quality small RNA (sRNA) data from four soybean (Glycine max) tissues, including root, seed, flower, and nodules, to expand the collection of currently known soybean miRNAs. We developed a bioinformatics pipeline using in-house scripts and publicly available structure prediction tools to differentiate the authentic mature miRNA sequences from other sRNAs and short RNA fragments represented in the public sequencing data. RESULTS: The combined sequencing and bioinformatics analyses identified 129 miRNAs based on hairpin secondary structure features in the predicted precursors. Out of these, 42 miRNAs matched known miRNAs in soybean or other species, while 87 novel miRNAs were identified. We also predicted the putative target genes of all identified miRNAs with computational methods and verified the predicted cleavage sites in vivo for a subset of these targets using the 5' RACE method. Finally, we also studied the relationship between the abundance of miRNA and that of the respective target genes by comparison to Solexa cDNA sequencing data. CONCLUSION: Our study significantly increased the number of miRNAs known to be expressed in soybean. The bioinformatics analysis provided insight on regulation patterns between the miRNAs and their predicted target genes expression. We also deposited the data in a soybean genome browser based on the UCSC Genome Browser architecture. Using the browser, we annotated the soybean data with miRNA sequences from four tissues and cDNA sequencing data. Overlaying these two datasets in the browser allows researchers to analyze the miRNA expression levels relative to that of the associated target genes. The browser can be accessed at http://digbio.missouri.edu/soybean_mirna/. Trupti Joshi, Marc Libault, Dong-Hoon Jeong, Sunhee Park, Pamela J. Green, D. Janine Sherrier, Andrew Farmer, Greg May, Blake C. Meyers, Dong Xu 0002, Gary Stacey |
BMC Bioinform. | 11 |
| 2010 | SeqRate: sequence-based protein folding type classification and rates predictionabstractBACKGROUND: Protein folding rate is an important property of a protein. Predicting protein folding rate is useful for understanding protein folding process and guiding protein design. Most previous methods of predicting protein folding rate require the tertiary structure of a protein as an input. And most methods do not distinguish the different kinetic nature (two-state folding or multi-state folding) of the proteins. Here we developed a method, SeqRate, to predict both protein folding kinetic type (two-state versus multi-state) and real-value folding rate using sequence length, amino acid composition, contact order, contact number, and secondary structure information predicted from only protein sequence with support vector machines. RESULTS: We systematically studied the contributions of individual features to folding rate prediction. On a standard benchmark dataset, the accuracy of folding kinetic type classification is 80%. The Pearson correlation coefficient and the mean absolute difference between predicted and experimental folding rates (sec-1) in the base-10 logarithmic scale are 0.81 and 0.79 for two-state protein folders, and 0.80 and 0.68 for three-state protein folders. SeqRate is the first sequence-based method for protein folding type classification and its accuracy of fold rate prediction is improved over previous sequence-based methods. Its performance can be further enhanced with additional information, such as structure-based geometric contacts, as inputs. CONCLUSIONS: Both the web server and software of predicting folding rate are publicly available at http://casp.rnet.missouri.edu/fold_rate/index.html. Guan Ning Lin, Zheng Wang 0025, Dong Xu 0002, Jianlin Cheng |
BMC Bioinform. | 3 |
| 2009 | Sequence-Based Prediction of Protein Folding Rates Using Contacts, Secondary Structures and Support Vector MachinesabstractPredicting protein folding rate is useful for understanding protein folding process and guiding protein design. Most previous methods of predicting folding rate require the tertiary structure of a protein as an input. And most methods do not distinguish the different kinetic natures (two-state folding and multi-state folding) of the proteins. Here we developed a method, SeqRate, to predict both protein folding kinetic type (two-state versus multi-state) and real-value folding rate using features extracted from only protein sequence with support vector machines. On a standard benchmark dataset, the accuracy of folding kinetic type classification is 80%. The Pearson correlation coefficient and the mean absolute difference between predicted and experimental folding rates (sec-1) in the base-10 logarithmic scale are 0.81 and 0.79 for two-state protein folders, and 0.80 and 0.68 for three-state protein folders. SeqRate is the first sequence-based method for protein folding type classification and its accuracy of fold rate prediction is improved over previous sequence-based methods. Both the Web server and software of predicting folding rate are publicly available at http://casp.rnet.missouri.edu/fold_rate/index.html. Guan Ning Lin, Zheng Wang 0025, Dong Xu 0002, Jianlin Cheng |
BIBM | 3 |
| 2009 | ComPhy: prokaryotic composite distance phylogenies inferred from whole-genome gene setsabstractBACKGROUND: With the increasing availability of whole genome sequences, it is becoming more and more important to use complete genome sequences for inferring species phylogenies. We developed a new tool ComPhy, 'Composite Distance Phylogeny', based on a composite distance matrix calculated from the comparison of complete gene sets between genome pairs to produce a prokaryotic phylogeny. RESULTS: The composite distance between two genomes is defined by three components: Gene Dispersion Distance (GDD), Genome Breakpoint Distance (GBD) and Gene Content Distance (GCD). GDD quantifies the dispersion of orthologous genes along the genomic coordinates from one genome to another; GBD measures the shared breakpoints between two genomes; GCD measures the level of shared orthologs between two genomes. The phylogenetic tree is constructed from the composite distance matrix using a neighbor joining method. We tested our method on 9 datasets from 398 completely sequenced prokaryotic genomes. We have achieved above 90% agreement in quartet topologies between the tree created by our method and the tree from the Bergey's taxonomy. In comparison to several other phylogenetic analysis methods, our method showed consistently better performance. CONCLUSION: ComPhy is a fast and robust tool for genome-wide inference of evolutionary relationship among genomes. It can be downloaded from http://digbio.missouri.edu/ComPhy. Guan Ning Lin, Zhipeng Cai 0001, Guohui Lin, Sounak Chakraborty, Dong Xu 0002 |
BMC Bioinform. | 5 |
| 2008 | A new clustering-based method for protein structure selectionabstractIn protein tertiary structure prediction, it is a crucial step to select near-native structures from a large number of candidate structural models. Despite much effort to tackle the problem of protein structure selection, the discerning power of current scoring functions is still unsatisfactory. In this paper, we developed a new clustering-based method for selecting near-native protein structures. Our method consists of three phases: filtering, clustering and cluster reduction, and centroid construction. Given a set of Cαprotein structures, we apply one or multiple existing scoring functions to filter out bad structures. Then, we group the remaining structures into clusters based on pair-wise similarity measured by RMSD. Each cluster is reduced iteratively to remove outliers and bad structures. Finally, we construct a centroid for each cluster by applying multi-dimensional scale techniques. The centroids are the final models. In experiments, we applied our method to a test set of representative proteins and obtained significant improvement over existing methods. Qingguo Wang, Yi Shang, Dong Xu 0002 |
IJCNN | 3 |
| 2008 | Wanted: unique names for unique atom positions. PDB-wide analysis of diastereotopic atom names of small molecules containing diphosphateabstractBACKGROUND: Biological chemistry is very stereospecific. Nonetheless, the diastereotopic oxygen atoms of diphosphate-containing molecules in the Protein Data Bank (PDB) are often given names that do not uniquely distinguish them from each other due to the lack of standardization. This issue has largely not been addressed by the protein structure community. RESULTS: Of 472 diastereotopic atom pairs studied from the PDB, 118 were found to have names that are not uniquely assigned. Among the molecules identified with these inconsistencies were many cofactors of enzymatic processes such as mononucleotides (e.g. ADP, ATP, GTP), dinucleotide cofactors (e.g. FAD, NAD), and coenzyme A. There were no overall trends in naming conventions, though ligand-specific trends were prominent. CONCLUSION: The lack of standardized naming conventions for diastereotopic atoms of small molecules has left the ad hoc names assigned to many of these atoms non-unique, which may create problems in data-mining of the PDB. We suggest a naming convention to resolve this issue. The in-house software used in this study is available upon request.A version of the software used for the analyses described in this paper is available at our web site: http://digbio.missouri.edu/ddan/DDAN.htm. Christopher A. Bottoms, Dong Xu 0002 |
BMC Bioinform. | 2 |
| 2007 | GoFuzzKegg: Mapping Genes to KEGG Pathways Using an Ontological Fuzzy Rule SystemabstractIn this paper we present a method for finding the main pathways represented in a set of genes (say obtained from a microarray experiment). The method is based on a fuzzy mapping between genes represented as sets of gene ontology terms and KEGG pathways using a new type of fuzzy rule system called ontological fuzzy rule system (OFRS). As opposed to a crisp mapping, the fuzzy mapping produces a nonzero value even if the gene name is not explicitly listed in a given KEGG pathway. An OFRS is a fuzzy rule system in which the rule memberships are obtained using similarity measures between objects computed based on the gene ontology (GO) annotations. To test our approach, we randomly selected without replacement 10 sets of Arabidopsis thaliana genes from KEGG (each set had 15 genes from 3 different pathways) and tried to predict the pathways they were selected from. Our method was able to find, 90% of the right pathways with a 65% false alarm rate at a p-value of 0.01. The high false alarm rate is due in part to the experimental setting. In a pilot dataset of 526 Arabidopsis thaliana genes we identified 8 clusters which proved to be linked to important pathways such as ATP synthesis and transcription factor Mihail Popescu, Dong Xu 0002, Erik Taylor |
CIBCB | 2 |
| 2007 | Mapping Genes to Pathways Using Ontological Fuzzy Rule SystemsabstractIn this paper we present a novel algorithm for mapping genes to pathways. The approach is based on the concept of ontological fuzzy rule system (OFRS) that, we believe, represents a step closer toward Zadeh's "computing with words" paradigm. An OFRS is a fuzzy rule system that uses ontological mapping between objects according to their representation as sets of terms from an ontology. In our mapping approach the left-hand-side contains genes objects represented using the gene ontology (GO), and the right-hand-side consists in pathways described using a pathway ontology (KEGG). The question of mapping a set of genes to pathways often arises in microarray experiments where one would like to know what are the pathways that can explain the observed gene expression patterns. To compare various mapping approaches, we use a pilot dataset extracted from KEGG that consisted of 10 sets of 15 genes taken from 3 pathways (5 gene/pathway). We conclude that the best matching strategy consists in two steps: in the first step the crisp approach is used to find the pathways involved, and in the second step the fuzzy (ontological) approach is employed to map the genes that could not be found in KEGG. Mihail Popescu, Dong Xu 0002 |
FUZZ-IEEE | 2 |
| 2006 | Bioinformatics and Fuzzy LogicabstractMany biological systems and objects are intrinsically fuzzy. Fuzzy set theory and fuzzy logic are ideal frameworks for describing some biological systems/objects and providing suitable computational methods for a widely range of bioinformatics problems. In this paper, we present two examples of using fuzzy set theory in bioinformatics, one in fuzzy measurement of ontological similarity and its application in bioinformatics, and the other in the application of the fuzzy k-nearest neighbor algorithm in protein secondary structure prediction. We also review other "fuzzy" methods for bioinformatics applications. Dong Xu 0002, Rajkumar Bondugula, Mihail Popescu, James Keller 0001 |
FUZZ-IEEE | 1 |
| 2006 | Supervised Inference of Gene Regulatory Networks by Linear Programming
Yong Wang 0001, Trupti Joshi, Dong Xu 0002, Xiang-Sun Zhang, Luonan Chen |
ICIC (3) | 3 |
| 2006 | Inferring gene regulatory networks from multiple microarray datasetsabstractMOTIVATION: Microarray gene expression data has increasingly become the common data source that can provide insights into biological processes at a system-wide level. One of the major problems with microarrays is that a dataset consists of relatively few time points with respect to a large number of genes, which makes the problem of inferring gene regulatory network an ill-posed one. On the other hand, gene expression data generated by different groups worldwide are increasingly accumulated on many species and can be accessed from public databases or individual websites, although each experiment has only a limited number of time-points. RESULTS: This paper proposes a novel method to combine multiple time-course microarray datasets from different conditions for inferring gene regulatory networks. The proposed method is called GNR (Gene Network Reconstruction tool) which is based on linear programming and a decomposition procedure. The method theoretically ensures the derivation of the most consistent network structure with respect to all of the datasets, thereby not only significantly alleviating the problem of data scarcity but also remarkably improving the prediction reliability. We tested GNR using both simulated data and experimental data in yeast and Arabidopsis. The result demonstrates the effectiveness of GNR in terms of predicting new gene regulatory relationship in yeast and Arabidopsis. AVAILABILITY: The software is available from http://zhangorup.aporc.org/bioinfo/grninfer/, http://digbio.missouri.edu/grninfer/ and http://intelligent.eic.osaka-sandai.ac.jp or upon request from the authors. Yong Wang 0001, Trupti Joshi, Xiang-Sun Zhang, Dong Xu 0002, Luonan Chen |
Bioinform. | 4 |
| 2006 | A fast SCOP fold classification system using content-based E-Predict algorithmabstractBACKGROUND: Domain experts manually construct the Structural Classification of Protein (SCOP) database to categorize and compare protein structures. Even though using the SCOP database is believed to be more reliable than classification results from other methods, it is labor intensive. To mimic human classification processes, we develop an automatic SCOP fold classification system to assign possible known SCOP folds and recognize novel folds for newly-discovered proteins. RESULTS: With a sufficient amount of ground truth data, our system is able to assign the known folds for newly-discovered proteins in the latest SCOP v1.69 release with 92.17% accuracy. Our system also recognizes the novel folds with 89.27% accuracy using 10 fold cross validation. The average response time for proteins with 500 and 1409 amino acids to complete the classification process is 4.1 and 17.4 seconds, respectively. By comparison with several structural alignment algorithms, our approach outperforms previous methods on both the classification accuracy and efficiency. CONCLUSION: In this paper, we build an advanced, non-parametric classifier to accelerate the manual classification processes of SCOP. With satisfactory ground truth data from the SCOP database, our approach identifies relevant domain knowledge and yields reasonably accurate classifications. Our system is publicly accessible at http://ProteinDBS.rnet.missouri.edu/E-Predict.php. Pin-Hao Chi, Chi-Ren Shyu, Dong Xu 0002 |
BMC Bioinform. | 3 |
| 2005 | Profiles and fuzzy K-nearest neighbor algorithm for protein secondary structure prediction
Rajkumar Bondugula, Ognen Duzlevski, Dong Xu 0002 |
APBC | 3 |
| 2003 | Towards Automated Derivation of Biological Pathways Using High-Throughput Biological DataabstractCharacterizing biological pathways at the genome scale is one of the most important and challenging tasks in the post genomic era. To address this challenge, we have developed a computational method to systematically and automatically derive partial biological pathways in yeast using high-throughput biological data, including yeast two hybrid data, protein complexes identified from mass spectroscopy, genetics interactions, and microarray gene expression data in yeast Saccharomyces cerevisiae. The inputs of the method are the upstream starting protein (e.g., a sensor of a signal) and the downstream terminal protein (e.g., a transcriptional factor that induces genes to respond the signal); the output of the method is the protein interaction chain between the two proteins. The high-throughput data are coded into a graph of interaction network, where each node represents a protein. The weight of an edge between two nodes models the "closeness" of the two represented proteins in the interaction network and it is defined by a rule-based formula according to the high-throughput data and modified by the protein function classification and subcellular localization information. The protein interaction cascade pathway in vivo is predicted as the shortest path identified from the graph of the interaction network using Dijkstra's algorithm. We have also developed a web server of this method (http://compbio.ornl.gov/structure/pathway) for public use. To our knowledge, our method is the first automated method to generally construct partial biological pathways using a suite of high-throughput biological data. This work demonstrates the proof of principle using computational approaches for discoveries of biological pathways with high-throughput data and biological annotation data. Trupti Joshi, Ying Xu 0001, Dong Xu 0002 |
BIBE | 4 |
| 2003 | A Computational Pipeline for Protein Structure Prediction and Analysis at Genome ScaleabstractTraditionally, protein 3D structures are solved using experimental techniques, like X-ray crystallography or nuclear magnetic resonance (NMR). While these experimental techniques have been the main workhorse for protein structure studies in the past few decades, it is becoming increasingly apparent that they alone cannot keep up with the production rate of protein sequences. Fortunately, computational techniques for protein structure predictions have matured to such a level that they can complement the existing experimental techniques. In this paper, we present an automated pipeline for protein structure prediction. The centerpiece of the pipeline is a threading-based protein structure prediction system, called PROSPECT, which we have been developing for the past few years. The pipeline consists of seven logical phases, utilizing a dozen tools. The pipeline has been implemented to run in a heterogeneous computational environment as a client/server system with a web interface. A number of genome-scale applications have been carried out on microbial genomes. Here we present one genome-scale application on Caenorhabditis elegans. Manesh J. Shah, Sergei Passovets, Dongsup Kim, Kyle Ellrott, Li Wang 0008, Inna Vokler, Philip F. LoCascio, Dong Xu 0002, Ying Xu 0001 |
BIBE | 8 |
| 2003 | More Reliable Protein NMR Peak Assignment via Improved 2-Interval Scheduling
Zhi-Zhong Chen, Tao Jiang 0001, Guohui Lin, Romeo Rizzi, Jianjun Wen, Dong Xu 0002, Ying Xu 0001 |
ESA | 6 |
| 2003 | A computational pipeline for protein structure prediction and analysis at genome scaleabstractMOTIVATION: Experimental techniques alone cannot keep up with the production rate of protein sequences, while computational techniques for protein structure predictions have matured to such a level to provide reliable structural characterization of proteins at large scale. Integration of multiple computational tools for protein structure prediction can complement experimental techniques. RESULTS: We present an automated pipeline for protein structure prediction. The centerpiece of the pipeline is our threading-based protein structure prediction system PROSPECT. The pipeline consists of a dozen tools for identification of protein domains and signal peptide, protein triage to determine the protein type (membrane or globular), protein fold recognition, generation of atomic structural models, prediction result validation, etc. Different processing and prediction branches are determined automatically by a prediction pipeline manager based on identified characteristics of the protein. The pipeline has been implemented to run in a heterogeneous computational environment as a client/server system with a web interface. Genome-scale applications on Caenorhabditis elegans, Pyrococcus furiosus and three cyanobacterial genomes are presented. AVAILABILITY: The pipeline is available at http://compbio.ornl.gov/proteinpipeline/ Manesh J. Shah, Sergei Passovets, Dongsup Kim, Kyle Ellrott, Li Wang 0008, Inna Vokler, Philip F. LoCascio, Dong Xu 0002, Ying Xu 0001 |
Bioinform. | 8 |
| 2003 | Approximation algorithms for NMR spectral peak assignment
Zhi-Zhong Chen, Tao Jiang 0001, Guohui Lin, Jianjun Wen, Dong Xu 0002, Jinbo Xu, Ying Xu 0001 |
Theor. Comput. Sci. | 5 |
| 2002 | Improved Approximation Algorithms for NMR Spectral Peak Assignment
Zhi-Zhong Chen, Tao Jiang 0001, Guohui Lin, Jianjun Wen, Dong Xu 0002, Ying Xu 0001 |
WABI | 5 |
| 2002 | PRIMEGENS: robust and efficient design of gene-specific probes for microarray analysisabstractMOTIVATION: DNA microarray is a powerful high-throughput tool for studying gene function and regulatory networks. Due to the problem of potential cross hybridization, using full-length genes for microarray construction is not appropriate in some situations. A bioinformatic tool, PRIMEGENS, has recently been developed for the automatic design of PCR primers using DNA fragments that are specific to individual open reading frames (ORFs). RESULTS: PRIMEGENS first carries out a BLAST search for each target ORF against all other ORFs of the genome to quickly identify possible homologous sequences. Then it performs optimal sequence alignment between the target ORF and each of its homologous ORFs using dynamic programming. PRIMEGENS uses the sequence alignments to select gene- specific fragments, and then feeds the fragments to the Primer3 program to design primer pairs for PCR amplification. PRIMEGENS can be run from the command line on Unix/Linux platforms as a stand-alone package or it can be used from a Web interface. The program runs efficiently, and it takes a few seconds per sequence on a typical workstation. PCR primers specific to individual ORFs from Shewanella oneidensis MR-1 and Deinococcus radiodurans R1 have been designed. The PCR amplification results indicate that this method is very efficient and reliable for designing specific probes for microarray analysis. Dong Xu 0002, Guangshan Li, Liyou Wu, Jizhong Zhou, Ying Xu 0001 |
Bioinform. | 1 |
| 2002 | Clustering gene expression data using a graph-theoretic approach: an application of minimum spanning treesabstractMOTIVATION: Gene expression data clustering provides a powerful tool for studying functional relationships of genes in a biological process. Identifying correlated expression patterns of genes represents the basic challenge in this clustering problem. RESULTS: This paper describes a new framework for representing a set of multi-dimensional gene expression data as a Minimum Spanning Tree (MST), a concept from the graph theory. A key property of this representation is that each cluster of the expression data corresponds to one subtree of the MST, which rigorously converts a multi-dimensional clustering problem to a tree partitioning problem. We have demonstrated that though the inter-data relationship is greatly simplified in the MST representation, no essential information is lost for the purpose of clustering. Two key advantages in representing a set of multi-dimensional data as an MST are: (1) the simple structure of a tree facilitates efficient implementations of rigorous clustering algorithms, which otherwise are highly computationally challenging; and (2) as an MST-based clustering does not depend on detailed geometric shape of a cluster, it can overcome many of the problems faced by classical clustering algorithms. Based on the MST representation, we have developed a number of rigorous and efficient clustering algorithms, including two with guaranteed global optimality. We have implemented these algorithms as a computer software EXpression data Clustering Analysis and VisualizATiOn Resource (EXCAVATOR). To demonstrate its effectiveness, we have tested it on three data sets, i.e. expression data from yeast Saccharomyces cerevisiae, expression data in response of human fibroblasts to serum, and Arabidopsis expression data in response to chitin elicitation. The test results are highly encouraging. AVAILABILITY: EXCAVATOR is available on request from the authors. Ying Xu 0001, Victor Olman, Dong Xu 0002 |
Bioinform. | 3 |
| 2000 | Protein structure determination using protein threading and sparse NMR data (extended abstract)abstractIt is well known that the NMR method for protein structure determination applies to small proteins and that its effectiveness decreases very rapidly as the molecular weight increases beyond about 30 kD. We have recently developed a method for protein structure determination that can fully utilize partial NMR data as calculation constraints. The core of the method is a threading algorithm that guarantees to find a globally optimal alignment between a query sequence and a template structure, under distance constraints specified by NMR/NOE data. Our preliminary tests have demonstrated that a small number of NMR/NOE distance restraints can significantly improve threading performance in both fold recognition and threading-alignment accuracy, and can possibly extend threading's scope of applicability from structural homologs to structural analogs. An accurate backbone structure generated by NMR-constrained threading can then provide a significant amount of structural information, equivalent to that provided by the NMR method with many NMR/NOE restraints; and hence can greatly reduce the amount of NMR data typically required for accurate structure determination. Our prelimenary study suggest that a small number of NOE restraints may suffice to determine adequately the all-atom structure when those restraints are incorporated in a procedure combining threading, modeling of loops and sidechains, and molecular dynamics simulation. Potentially, this new technique can expand NMR's capability to larger proteins. Ying Xu 0001, Dong Xu 0002, Oakley H. Crawford, J. Ralph Einstein, Engin Serpersu |
RECOMB | 2 |
| 2000 | Sequence-structure specificity of a knowledge based energy function at the secondary structure levelabstractMOTIVATION: This paper investigates the sequence-structure specificity of a representative knowledge based energy function by applying it to threading at the level of secondary structures of proteins. Assessing the strengths and weaknesses of an energy function at this fundamental level provides more detailed and insightful information than at the tertiary structure level and the results obtained can be useful in tertiary level threading. RESULTS: We threaded each of the 293 non-redundant proteins onto the secondary structures contained in its respective native protein (host template). We also used 68 pairs of proteins with similar folds and low sequence identity. For each pair, we threaded the sequence of one protein onto the secondary structures of the other protein. The discerning power of the total energy function and its one-body, pairwise, and mutation components is studied. We then applied our energy function to a recent study which demonstrated how a designed 11-amino acid sequence can replace distinct segments (one segment is an alpha-helix, the other is a beta-sheet) of a protein without changing its fold. We conducted random mutations of the designed sequence to determine the patterns for favorable mutations. We also studied the sequence-structure specificity at the boundaries of a secondary structure. Finally, we demonstrated how to speed up tertiary level threading by filtering out alignments found to be energetically unfavorable during the secondary structure threading. AVAILABILITY: The program is available on request from the authors. CONTACT: [email protected] Dong Xu 0002, Michael A. Unseren, Ying Xu 0001, Edward C. Uberbacher |
Bioinform. | 1 |
| 2000 | Protein domain decomposition using a graph-theoretic approachabstractMOTIVATION: Automatic decomposition of a multi-domain protein into individual domains represents a highly interesting and unsolved problem. As the number of protein structures in PDB is growing at an exponential rate, there is clearly a need for more reliable and efficient methods for protein domain decomposition simply to keep the domain databases up-to-date. RESULTS: We present a new algorithm for solving the domain decomposition problem, using a graph-theoretic approach. We have formulated the problem as a network flow problem, in which each residue of a protein is represented as a node of the network and each residue--residue contact is represented as an edge with a particular capacity, depending on the type of the contact. A two-domain decomposition problem is solved by finding a bottleneck (or a minimum cut) of the network, which minimizes the total cross-edge capacity, using the classical Ford--Fulkerson algorithm. A multi-domain decomposition problem is solved through repeatedly solving a series of two-domain problems. The algorithm has been implemented as a computer program, called DomainParser. We have tested the program on a commonly used test set consisting of 55 proteins. The decomposition results are 78.2% in agreement with the literature on both the number of decomposed domains and the assignments of residues to each domain, which compares favorably to existing programs. On the subset of two-domain proteins (20 in number), the program assigned 96.7% of the residues correctly when we require that the number of decomposed domains is two. Ying Xu 0001, Dong Xu 0002, Harold N. Gabow |
Bioinform. | 2 |
| 1998 | A new method for modeling and solving the protein fold recognition problem (extended abstract)abstractComputational recognition of native-like folds from a protein fold database is considered to be a promising alternative approach to the ab initio fold prediction.We present a new and egective method forprotein fold recognition through optimally aligning (threading) an amino acid sequence and a protein fold (template).A protein fold, in our database, is represented as a se953 of core secondary structures, and the alignment quality is cletermined by three factors.They are (I) the fitness between each amino acid and the environment of its assigned (aligned) template position; (2) pairwise interaction preferences between amino acids that are spatially close,-and (3) alignment gap penalties.Our threading algorithm cons&ucts an optimum alignment between an amino acid sequence of size n and a protein fold template of size m in O((m-l-n1+o-5cMlog(n))nc+1) time and O(nm+nc+2) space, where M is the number of core secondary structures in the fold, and C is a (small) nonnegative integer, determined by a mathematical property of the pairwise interactions in the fold.C is less than or equal to 4 for about 75% of the 296 unique folds in our .database,when pairwise interactions are restricted to amino acids 5 7A apart (measured between their beta carbon atoms).An approximation scheme is developed for fold templates with C > 4, when threading requires too much memo y and time to be practical on a typical workstation.Permission to make digitalibrud copies of all or pat ofthis material for personal orclassroomuseisgranteduithoutfeeprovidedthatthecopies are not made or diiuted for profit or commercial advantage, the wpy-riShtnotice,thetitleofthepublicationanditsdateappear, andnoticek giventhatcopylightkbype -on ofthe AChL Inc To copy othenvise, to republish, to post on servers or to rediiiuteto lists, requires specific pamission ardor fee. Ying Xu 0001, Dong Xu 0002, Edward C. Uberbacher |
RECOMB | 2 |