EDBT 2026 Demo / reviewers in the wild / expert
Ting Chen 0006
dblp:19/1766-6
· DBLP profile ↗
55ranked-venue papers
0as first author
19since 2021 · last 2026
0000-0002-3228-9166ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 38 · 7 since 2021Artificial intelligence and machine learning · 14 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RareAgents: Autonomous Multi-disciplinary Team for Rare Disease Diagnosis and TreatmentabstractRare diseases, despite their low individual incidence, collectively impact around 300 million people worldwide due to the vast number of diseases. The involvement of multiple organs and systems, and the shortage of specialized doctors with relevant experience, make diagnosing and treating rare diseases more challenging than common diseases. Recently, agents powered by large language models (LLMs) have demonstrated notable applications across various domains. In the medical field, some agent methods have outperformed direct prompts in question-answering tasks from medical examinations. However, current agent frameworks are not well-adapted to real-world clinical scenarios, especially those involving the complex demands of rare diseases. To bridge this gap, we introduce RareAgents, the first LLM-driven multi-disciplinary team decision-support tool designed specifically for the complex clinical context of rare diseases. RareAgents integrates advanced Multidisciplinary Team (MDT) coordination, memory mechanisms, and medical tools utilization, leveraging Llama-3.1-8B/70B as the base model. Experimental results show that RareAgents outperforms state-of-the-art domain-specific models, GPT-4o, and current agent frameworks in diagnosis and treatment for rare diseases. Furthermore, we contribute a novel rare disease dataset, MIMIC-IV-Ext-Rare, to facilitate further research in this field. Xuanzhong Chen, Xiaohao Mao, Ting Chen 0006 |
AAAI | 6 |
| 2025 | HeMeNet: Heterogeneous Multichannel Equivariant Network for Protein Multi-task LearningabstractUnderstanding and leveraging the 3D structures of proteins is central to a variety of biological and drug discovery tasks. While deep learning has been applied successfully for structure-based protein function prediction tasks, current methods usually employ distinct training for each task. However, each of the tasks is of small size, and such a single-task strategy hinders the models' performance and generalization ability. As some labeled 3D protein datasets are biologically related, combining multi-source datasets for larger-scale multi-task learning is one way to overcome this problem. In this paper, we propose a neural network model to address multiple tasks jointly upon the input of 3D protein structures. In particular, we first construct a standard structure-based multi-task benchmark called Protein-MT, consisting of 6 biologically relevant tasks, including affinity prediction and property prediction, integrated from 4 public datasets. Then, we develop a novel graph neural network for multi-task learning, dubbed Heterogeneous Multichannel Equivariant Network (HeMeNet), which is E(3) equivariant and able to capture heterogeneous relationships between different atoms. Besides, HeMeNet can achieve task-specific learning via the task-aware readout mechanism. Extensive evaluations of our benchmark verify the effectiveness of multi-task learning, and our model generally surpasses state-of-the-art models. Wenbing Huang 0001, Lingxiao Luo, Xinyan Han, Ting Chen 0006 |
AAAI | 8 |
| 2025 | CoPRA: Bridging Cross-domain Pretrained Sequence Models with Complex Structures for Protein-RNA Binding Affinity PredictionabstractAccurately measuring protein-RNA binding affinity is crucial in many biological processes and drug design. Previous computational methods for protein-RNA binding affinity prediction rely on either sequence or structure features, unable to capture the binding mechanisms comprehensively. The recent emerging pre-trained language models trained on massive unsupervised sequences of protein and RNA have shown strong representation ability for various in-domain downstream tasks, including binding site prediction. However, applying different-domain language models collaboratively for complex-level tasks remains unexplored. In this paper, we propose CoPRA to bridge pre-trained language models from different biological domains via Complex structure for Protein-RNA binding Affinity prediction. We demonstrate for the first time that cross-biological modal language models can collaborate to improve binding affinity prediction. We propose a Co-Former to combine the cross-modal sequence and structure information and a bi-scope pre-training strategy for improving Co-Former's interaction understanding. Meanwhile, we build the largest protein-RNA binding affinity dataset PRA310 for performance evaluation. We also test our model on a public dataset for mutation effect prediction. CoPRA reaches state-of-the-art performance on all the datasets. We provide extensive analyses and verify that CoPRA can (1) accurately predict the protein-RNA binding affinity; (2) understand the binding affinity change caused by mutations; and (3) benefit from scaling data and model size. Xiaohong Liu 0007, Tong Pan, Jing Xu 0008, Xiaoyu Wang 0016, Wuyang Lan, Jiangning Song, Ting Chen 0006 |
AAAI | 11 |
| 2025 | VividMed: Vision Language Model with Versatile Visual Grounding for MedicineabstractLingxiao Luo, Bingda Tang, Xuanzhong Chen, Rong Han, Ting Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Lingxiao Luo, Bingda Tang, Xuanzhong Chen, Ting Chen 0006 |
NAACL (Long Papers) | 5 |
| 2025 | Rag2Mol: Structure-Based Drug Design Based on Retrieval Augmented Generation
Xingang Peng, Ting Chen 0006, Jianzhu Ma |
RECOMB | 4 |
| 2025 | Rag2Mol: structure-based drug design based on retrieval augmented generationabstractArtificial intelligence (AI) has brought tremendous progress to drug discovery, yet identifying hit and lead compounds with optimal physicochemical and pharmacological properties remains a significant challenge. Structure-based drug design (SBDD) has emerged as a promising paradigm, but the inherent data biases and ignorance of synthetic accessibility render SBDD models disconnected from practical drug discovery. In this work, we explore two methodologies, Rag2Mol-G and Rag2Mol-R, both based on retrieval-augmented generation to design small molecules to fit a 3D pocket. These two methods involve searching for similar small molecules that are purchasable in the database based on the generated ones or creating new molecules from those in the database that can fit into a 3D pocket. Experimental results demonstrate that Rag2Mol methods consistently produce drug candidates with superior binding affinities and drug-likeness. We find that Rag2Mol-R provides a broader coverage of the chemical landscapes and more precise targeting capability than advanced virtual screening models. Notably, both workflows identified promising inhibitors for the challenging target protein tyrosine phosphatases PTPN2, which was used to be considered undruggable and still lacks inhibitors that have completed full clinical trials. Our highly extensible framework can integrate diverse SBDD methods, marking a significant advancement in AI-driven SBDD. Xingang Peng, Ting Chen 0006, Jianzhu Ma |
Briefings Bioinform. | 4 |
| 2025 | MyClone: rapid and precise reconstruction of clonal population structures for tumorsabstractBACKGROUND: Understanding tumor heterogeneity is essential for advancing cancer treatment. Clonal reconstruction methods play a pivotal role in deciphering this heterogeneity. Our goal is to develop a clonal reconstruction approach that is clinically applicable, easy to implement, and capable of delivering both high-speed performance and excellent reconstruction accuracy. RESULTS: We present MyClone, a probabilistic method designed to reconstruct the clonal composition of tumors using deep sequencing genomic data. MyClone processes read counts and copy number information of single nucleotide variants derived from deep sequencing data, enabling it to determine the mutational composition of clones and the cancer cell fractions of these mutations. Compared to existing clonal reconstruction methods, MyClone enhances clustering accuracy and cancer cell fraction prediction when applied to deep-targeted sequencing data and bulk tumor sequencing data with deep sequencing coverage. Additionally, MyClone achieves a substantial improvement in computational speed. We rigorously validated MyClone's performance using both simulated and real clinical datasets and applied it to analyze a circulating tumor DNA sequencing dataset from 139 metastatic breast cancer patients. In this analysis, we explored mechanisms of drug resistance in metastatic breast cancer and identified 10 mutated genes potentially associated with drug resistance or sensitivity. CONCLUSIONS: For deeply sequenced data, MyClone outperforms existing methods on both targeted sequencing and bulk tumor data. Its high computational efficiency and reconstruction accuracy position MyClone as a promising tool for broad clinical application in cancer treatment. The source code is publicly available at https://github.com/Hansen0413/myclone-code . Qihan Guo, Dan Lv, Haili Qian, Ting Chen 0006 |
BMC Bioinform. | 6 |
| 2025 | Escaping the Drug-Bias Trap: Using Debiasing Design to Improve Interpretability and Generalization of Drug-Target Interaction PredictionabstractConsidering the high cost associated with determining reaction affinities through in vitro experiments, virtual screening of potential drugs bound to specific protein pockets from vast compounds is critical in AI-assisted drug discovery. Deep-learning approaches have been proposed for predicting Drug-Target Interactions (DTIs). However, they have shown an overestimated accuracy due to the drug-bias trap, a challenge where traditional multimodal models overly rely on the drug branch while underutilizing protein information. This raises doubts about the interpretability and generalizability of existing DTI models. Therefore, we introduce UdanDTI, an innovative deep-learning architecture explicitly designed for predicting drug-protein interactions. UdanDTI applies an unbalanced dual-branch system and an attentive aggregation module to enhance interpretability from a biological perspective. Across various public datasets, UdanDTI demonstrates outstanding performance, outperforming state-of-the-art models under in-domain, cross-domain, and structural interpretability settings. Notably, it demonstrates exceptional accuracy in predicting drug responses of two crucial subgroups of Epidermal Growth Factor Receptor (EGFR) mutations associated with non-small cell lung cancer, consistent with experimental results. Meanwhile, UdanDTI could complement the advanced molecular docking software DiffDock. Jianzhu Ma, Ting Chen 0006 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | DGPO: Discovering Multiple Strategies with Diversity-Guided Policy OptimizationabstractMost reinforcement learning algorithms seek a single optimal strategy that solves a given task. However, it can often be valuable to learn a diverse set of solutions, for instance, to make an agent's interaction with users more engaging, or improve the robustness of a policy to an unexpected perturbance. We propose Diversity-Guided Policy Optimization (DGPO), an on-policy algorithm that discovers multiple strategies for solving a given task. Unlike prior work, it achieves this with a shared policy network trained over a single run. Specifically, we design an intrinsic reward based on an information-theoretic diversity objective. Our final objective alternately constraints on the diversity of the strategies and on the extrinsic reward. We solve the constrained optimization problem by casting it as a probabilistic inference task and use policy iteration to maximize the derived lower bound. Experimental results show that our method efficiently discovers diverse strategies in a wide variety of reinforcement learning tasks. Compared to baseline methods, DGPO achieves comparable rewards, while discovering more diverse strategies, and often with better sample efficiency. Wentse Chen, Shiyu Huang 0001, Yuan Chiang, Tim Pearce, Wei-Wei Tu, Ting Chen 0006, Jun Zhu 0001 |
AAAI | 6 |
| 2024 | RareBench: Can LLMs Serve as Rare Diseases Specialists?abstractGeneralist Large Language Models (LLMs), such as GPT-4, have shown considerable promise in various domains, including medical diagnosis. Rare diseases, affecting approximately 300 million people worldwide, often have unsatisfactory clinical diagnosis rates primarily due to a lack of experienced physicians and the complexity of differentiating among many rare diseases. In this context, recent news such as "ChatGPT correctly diagnosed a 4-year-old's rare disease after 17 doctors failed" underscore LLMs' potential, yet underexplored, role in clinically diagnosing rare diseases. To bridge this research gap, we introduce RareBench, a pioneering benchmark designed to systematically evaluate the capabilities of LLMs on 4 critical dimensions within the realm of rare diseases. Meanwhile, we have compiled the largest open-source dataset on rare disease patients, establishing a benchmark for future studies in this domain. To facilitate differential diagnosis of rare diseases, we develop a dynamic few-shot prompt methodology, leveraging a comprehensive rare disease knowledge graph synthesized from multiple knowledge bases, significantly enhancing LLMs' diagnostic performance. Moreover, we present an exhaustive comparative study of GPT-4's diagnostic capabilities against those of specialist physicians. Our experimental findings underscore the promising potential of integrating LLMs into the clinical diagnostic process for rare diseases. This paves the way for exciting possibilities in future advancements in this field. Xuanzhong Chen, Xiaohao Mao, Qihan Guo, Ting Chen 0006 |
KDD | 6 |
| 2023 | DeSTSeg: Segmentation Guided Denoising Student-Teacher for Anomaly DetectionabstractVisual anomaly detection, an important problem in computer vision, is usually formulated as a one-class classification and segmentation task. The student-teacher (S- T) framework has proved to be effective in solving this chal-lenge. However, previous works based on S-T only empirically applied constraints on normal data and fused multilevel information. In this study, we propose an improved model called DeS TSeg, which integrates a pre-trained teacher network, a denoising student encoder-decoder, and a segmentation network into one framework. First, to strengthen the constraints on anomalous data, we intro-duce a denoising procedure that allows the student net-work to learn more robust representations. From synthet-ically corrupted normal images, we train the student net-work to match the teacher network feature of the same images without corruption. Second, to fuse the multi-level S-T features adaptively, we train a segmentation network with rich supervision from synthetic anomaly masks, achieving a substantial performance improvement. Experiments on the industrial inspection benchmark dataset demonstrate that our method achieves state-of-the-art performance, 98.6% on image-level AUC, 75.8% on pixel-level average precision, and 76.4% on instance-level average precision. Xi Li 0010, Jiulong Shan, Ting Chen 0006 |
CVPR | 6 |
| 2023 | Improving artificial intelligence pipeline for liver malignancy diagnosis using ultrasound images and video framesabstractRecent developments of deep learning methods have demonstrated their feasibility in liver malignancy diagnosis using ultrasound (US) images. However, most of these methods require manual selection and annotation of US images by radiologists, which limit their practical application. On the other hand, US videos provide more comprehensive morphological information about liver masses and their relationships with surrounding structures than US images, potentially leading to a more accurate diagnosis. Here, we developed a fully automated artificial intelligence (AI) pipeline to imitate the workflow of radiologists for detecting liver masses and diagnosing liver malignancy. In this pipeline, we designed an automated mass-guided strategy that used segmentation information to direct diagnostic models to focus on liver masses, thus increasing diagnostic accuracy. The diagnostic models based on US videos utilized bi-directional convolutional long short-term memory modules with an attention-boosted module to learn and fuse spatiotemporal information from consecutive video frames. Using a large-scale dataset of 50 063 US images and video frames from 11 468 patients, we developed and tested the AI pipeline and investigated its applications. A dataset of annotated US images is available at https://doi.org/10.5281/zenodo.7272660. Yiming Xu 0010, Xiaohong Liu 0007, Jinxiu Ju, Shi-jie Wang, Yufan Lian, Tong Liang, Ye Sang, Rui Jiang 0001, Ting Chen 0006 |
Briefings Bioinform. | 14 |
| 2022 | VMAPD: Generate Diverse Solutions for Multi-Agent Games with Recurrent Trajectory DiscriminatorsabstractRecent algorithms designed for multi-agent tasks focus on finding a single optimal solution for all the agents. However, in many tasks (e.g., matrix games and transportation dispatching), there may exist more than one optimal solution, while previous algorithms can only converge to one of them. In many practical applications, it is important to develop reasonable agents with diverse behaviors. In this paper, we propose ”variational multi-agent policy diversification” (VMAPD), an on-policy framework for discovering diverse policies for coordination patterns of multiple agents. By taking advantage of latent variables and exploiting the connection between variational inference and multi-agent reinforcement learning, we derive a tractable evidence lower bound (ELBO) on the trajectories of all agents. Our algorithm uses policy iteration to maximize the derived lower bound and can be simply implemented by adding a pseudo reward during centralized learning. And the trained agents do not need to access the pseudo reward during decentralized execution. We demonstrate the effectiveness of our algorithm on several popular multi-agent testbeds. Experimental results show that VMAPD finds more solutions with similar sample complexity compared with other baselines. Shiyu Huang 0001, Chao Yu 0005, Bin Wang 0034, Dong Li 0016, Yu Wang 0002, Ting Chen 0006, Jun Zhu 0001 |
CoG | 6 |
| 2022 | Yolo-SG: Salience-Guided Detection Of Small Objects In Medical ImagesabstractObject detection, a crucial component of medical image analysis, provides physicians with an interpretable auxiliary diagnostic basis. Although existing object detection models have had great success with natural images, the growing resolution of medical images makes the problem especially challenging because of the increased expectations to exploit the image details and discover small targets in images. For instance, lesions are occasionally diminutive relative to high-resolution medical images. To address this problem, we present YOLO-SG, a salience-guided (SG) deep learning model that improves small object detection by attending to detailed regions via a generated salience map. YOLO-SG performs two rounds of detection: coarse detection and salience-guided detection. In the first round of coarse detection, YOLO-SG detects objects using a deep convolutional detection model and proposes a salience map utilizing the context surrounding objects to guide the subsequent round of detection. In the second round, YOLO-SG extracts salient regions from the original input image based on the generated salience map and combines local detail with global context information to improve the object detection performance. The experimental results demonstrate that YOLO-SG outperforms the state-of-the-art models, especially when detecting small objects. Xiaohong Liu 0007, Ting Chen 0006 |
ICIP | 3 |
| 2022 | Scalable Online Disease Diagnosis via Multi-Model-Fused Actor-Critic Reinforcement LearningabstractFor those seeking healthcare advice online, AI based dialogue agents capable of interacting with patients to perform automatic disease diagnosis are a viable option. This application necessitates efficient inquiry of relevant disease symptoms in order to make accurate diagnosis recommendations. This can be formulated as a problem of sequential feature (symptom) selection and classification for which reinforcement learning (RL) approaches have been proposed as a natural solution. They perform well when the feature space is small, that is, the number of symptoms and diagnosable disease categories is limited, but they frequently fail in assignments with a large number of features. To address this challenge, we propose a Multi-Model-Fused Actor-Critic (MMF-AC) RL framework that consists of a generative actor network and a diagnostic critic network. The actor incorporates a Variational AutoEncoder (VAE) to model the uncertainty induced by partial observations of features, thereby facilitating in making appropriate inquiries. In the critic network, a supervised diagnosis model for disease predictions is involved to precisely estimate the state-value function. Furthermore, inspired by the medical concept of differential diagnosis, we combine the generative and diagnosis models to create a novel reward shaping mechanism to address the sparse reward problem in large search spaces. We conduct extensive experiments on both synthetic and real-world datasets for empirical evaluations. The results demonstrate that our approach outperforms state-of-the-art methods in terms of diagnostic accuracy and interaction efficiency while also being more effectively scalable to large search spaces. Besides, our method is adaptable to both categorical and continuous features, making it ideal for online applications. Ting Chen 0006 |
KDD | 2 |
| 2022 | BSODA: A Bipartite Scalable Framework for Online Disease DiagnosisabstractA growing number of people are seeking healthcare advice online. Usually, they diagnose their medical conditions based on the symptoms they are experiencing, which is also known as self-diagnosis. From the machine learning perspective, online disease diagnosis is a sequential feature (symptom) selection and classification problem. Reinforcement learning (RL) methods are the standard approaches to this type of tasks. Generally, they perform well when the feature space is small, but frequently become inefficient in tasks with a large number of features, such as the self-diagnosis. To address the challenge, we propose a non-RL Bipartite Scalable framework for Online Disease diAgnosis, called BSODA. BSODA is composed of two cooperative branches that handle symptom-inquiry and disease-diagnosis, respectively. The inquiry branch determines which symptom to collect next by an information-theoretic reward. We employ a Product-of-Experts encoder to significantly improve the handling of partial observations of a large number of features. Besides, we propose several approximation methods to substantially reduce the computational cost of the reward to a level that is acceptable for online services. Additionally, we leverage the diagnosis model to estimate the reward more precisely. For the diagnosis branch, we use a knowledge-guided self-attention model to perform predictions. In particular, BSODA determines when to stop inquiry and output predictions using both the inquiry and diagnosis models. We demonstrate that BSODA outperforms the state-of-the-art methods on several public datasets. Moreover, we propose a novel evaluation method to test the transferability of symptom checking methods from synthetic to real-world tasks. Compared to existing RL baselines, BSODA is more effectively scalable to large search spaces. Xiaohao Mao, Chao Ma 0019, José Miguel Hernández-Lobato, Ting Chen 0006 |
WWW | 6 |
| 2022 | Explainable Dynamic Multimodal Variational Autoencoder for the Prediction of Patients With Suspected Central Precocious PubertyabstractCentral precocious puberty (CPP) is the most common type of precocious puberty and has a significant effect on children. A gonadotropin-releasing hormone (GnRH)-stimulation test is the gold standard for confirming CPP. This test, however, is costly and unpleasant for patients. Therefore, it is critical to developing alternative methods for CPP diagnosis in order to alleviate patient suffering. This study aims to develop an artificial intelligence (AI) diagnostic system for predicting response to the GnRH-stimulation test using data from laboratory tests, electronic health records (EHRs), and pelvic ultrasonography and left-hand radiography reports. The challenges are in integrating these multimodal features into a comprehensive deep learning model in order to achieve an accurate diagnosis while also accounting for the missing or incomplete modalities. To begin, we developed a dynamic multimodal variational autoencoder (DMVAE) that can exploit intrinsic correlations between different modalities to impute features for missing modalities. Next, we combined features from all modalities to predict the outcome of a CPP diagnosis. The experimental results (AUROC 0.9086) demonstrate that our DMVAE model is superior to standard methods. Additionally, we showed that by setting appropriate operating thresholds, clinicians could diagnose about two-thirds of patients with confidence (1.0 specificity). Only about one-third of patients require confirmation of their diagnoses using GnRH (or GnRH analog)-stimulation tests. To interpret the results, we implemented an explainer Shapley additive explanation (SHAP) to analyze the local and global feature attributions. Yiming Xu 0010, Xiaohong Liu 0007, Liyan Pan, Xiaojian Mao, Huiying Liang, Ting Chen 0006 |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Deep Active Learning For Fibrosis Segmentation Of Chest CT Scans From Covid-19 PatientsabstractDuring the ongoing COVID-19 outbreak, it is critical to assess patients’ disease progression with COVID-19 pneumonia by computed tomography (CT). As most of the works focused on ground-glass opacity and consolidation segmentation of COVID-19 on CT images, lung fibrosis is relatively undervalued and less studied. Automatic segmentation and accurate measurement of lung fibrosis can potentially aid treatment planning for patients of post-COVID-19 pneumonia. However, the lack of sufficient training data hinders the fibrosis segmentation of CT images. Also, redundancy among CT images can reduce annotating efficiency. To address these issues, we propose deep active learning (AL) framework, which consists of a segmentation model called UNet-RGD, and a novel acquisition method named DeepRISS. The segmentation model consists of improved structures of residual blocks, channel gates, and dropout layers. The deep learning-based acquisition method combines uncertainty estimation and clustering for selecting representative and informative samples. Experimental results show that the AL framework can achieve state-of-the-art performance and effectively reduce the number of selected samples, saving the annotation cost by 25% to 44% compared to the non-selective approach. Xiaohong Liu 0007, Kai Wang 0100, Ting Chen 0006 |
ICIP | 3 |
| 2021 | Off-Policy Training for Truncated TD(λ) Boosted Soft Actor-Critic
Shiyu Huang 0001, Bin Wang 0034, Hang Su 0006, Dong Li 0016, Jianye Hao, Jun Zhu 0001, Ting Chen 0006 |
PRICAI (3) | 7 |
| 2020 | Attention U-net for Interpretable Classification on Chest X-ray ImageabstractConvolutional neural network (CNN) plays a vital role in numerous classification tasks; however, its lack of interpretability limits its application in medical image diagnosis. To tackle this issue, we propose Attention U-net, an interpretable classification model that can generate high-resolution localization maps for the predicted class. The novelty of our model is to adopt an upsampling-concatenating-convolution structure to create a fine-grained segmentation map and use attention pooling over the prior mask for bridging segmentation with classification. Since the relationship between segmentation and classification is equivalent to the formulation of the multiple instance learning (MIL), the attention pooling can be viewed as a MIL pooling function. In the attention pooling, the attention weights can be seen as a localization map, and thus provide evidence of classification. We integrate our model with grad-CAM (class activation mapping), a widely used method for CNN localization, and we prove that our attention-based localization map is highly correlated to the grad-CAM-integrated localization map. We apply our proposed model to the automatic diagnosis of lung diseases with Chest X-ray. Experimental results show that our model can reach high performance on both classification and interpretability simultaneously. Ting Chen 0006 |
BIBM | 2 |
| 2020 | SVQN: Sequential Variational Soft Q-Learning Networks
Shiyu Huang 0001, Hang Su 0006, Jun Zhu 0001, Ting Chen 0006 |
ICLR | 4 |
| 2020 | Automatic Emergency Diagnosis with Knowledge-Based Tree DecodingabstractAutomatic diagnosis based on clinical notes is critical especially in the emergency department, where a fast and professional result is vital in assuring proper and timely treatment. Previous works formalize this task as plain text classification and fail to utilize the medically significant tree structure of International Classification of Diseases (ICD) coding system. Besides, external medical knowledge is rarely used before, and we explore it by extracting relevant materials from Wikipedia or Baidupedia. In this paper, we propose a knowledge-based tree decoding model (K-BTD), and the inference procedure is a top-down decoding process from the root node to leaf nodes. The stepwise inference procedure enables the model to give support for decision at each step, which visualizes the diagnosis procedure and adds to the interpretability of final predictions. Experiments on real-world data from the emergency department of a large-scale hospital indicate that the proposed model outperforms all baselines in both micro-F1 and macro-F1, and reduce the semantic distance dramatically. Xuyan Chen, Ning Chen 0002, Ting Chen 0006 |
IJCAI | 4 |
| 2020 | Joint Medical Ontology Representation Learning for Healthcare PredictionsabstractHealthcare predictions aim at predicting diseases of the next visit to hospital with historical Electronic Health Records (EHR), which is a key research field in personalized healthcare. Previous research has demonstrated that learning meaningful medical ontology representations within the healthcare prediction model can alleviate the data insufficiency problem and thus is beneficial to this task. There are two main pathways of learning medical ontology representations. The first is through pre-defined knowledge graph such as the ICD tree, and the second is through the co-occurrence of diseases within each visit. Majority of existing works formalize their model under only one pathway, and fail to utilize the mutual benefits between them. To exploit these benefits, we propose JMRL, an end-to-end and accurate model for healthcare predictions with Joint Medical ontology Representation Learning. JMRL not only utilizes the joint information from both knowledge graph and co-occurrence statistics, but also make use of the mutual benefits between them in an advanced way with two explicit feedback strategies. Experimental results on the MIMIC-III dataset demonstrate the superiority of our model over all existing state-of-the-art approaches. Ning Chen 0002, Ting Chen 0006 |
IJCNN | 3 |
| 2020 | KISEG: A Three-Stage Segmentation Framework for Multi-level Acceleration of Chest CT Scans from COVID-19 Patients
Xiaohong Liu 0007, Kai Wang 0100, Ting Chen 0006 |
MICCAI (4) | 4 |
| 2020 | Anterior Segment Eye Lesion Segmentation with Advanced Fusion Strategies and Auxiliary Tasks
Xiaohong Liu 0007, Ting Chen 0006 |
MICCAI (5) | 4 |
| 2019 | Combo-Action: Training Agent For FPS Game with Auxiliary TasksabstractDeep reinforcement learning (DRL) has achieved surpassing human performance on Atari games, using raw pixels and rewards to learn everything. However, first-person-shooter (FPS) games in 3D environments contain higher levels of human concepts (enemy, weapon, spatial structure, etc.) and a large action space. In this paper, we explore a novel method which can plan on temporally-extended action sequences, which we refer as Combo-Action to compress the action space. We further train a deep recurrent Q-learning network model as a high-level controller, called supervisory network, to manage the Combo-Actions. Our method can be boosted with auxiliary tasks (enemy detection and depth prediction), which enable the agent to extract high-level concepts in the FPS games. Extensive experiments show that our method is efficient in training process and outperforms previous stateof-the-art approaches by a large margin. Ablation study experiments also indicate that our method can boost the performance of the FPS agent in a reasonable way. Shiyu Huang 0001, Hang Su 0006, Jun Zhu 0001, Ting Chen 0006 |
AAAI | 4 |
| 2019 | Bone Age Assessment by Deep Convolutional Neural Networks Combined with Clinical TW3-RUSabstractBone age assessment is critical to diagnosis of various growth disorders in children, such as endocrine, nutritional disorders and dysplasia. X-rays of hand and wrist are the most common modality used to calculate bone age. In this paper, we propose a novel approach called DeepTW3 for automatic bone age assessment from X-ray images. DeepTW3 integrates Convolutional Neural Networks (CNNs) with expertise knowledge of TW3(Tanner-Whitehouse 3nd edition)-RUS(radius, ulna and short bones) bone age assessment system. The proposed method is tested on a dataset containing 1,100 hand bone X-ray images, all of which were manually annotated with selected region of interests(ROIs). Our method achieved mean absolute errors (MAE) of 0.2685, outperforming all state-of-the-art methods. For the task of grading skeletal maturity, our method using continuous stage distribution is complementary to using the clinical TW3-RUS categorical stages when interpreting critical cases of intermediate bone stage. Xiaohong Liu 0007, Yiming Xu 0010, Ning Chen 0002, Ting Chen 0006 |
BIBM | 5 |
| 2019 | DeepTriager: A Neural Attention Model for Emergency Triage with Electronic Health RecordsabstractAs the first pass for emergency patients, triage is the most important factor affecting emergency department (ED) overcrowding. So it is crucial to develop a data-driven and evidence-based triage method to quickly identify acute and severe patients, and prevent the limited emergency resources from over-diagnosis. To address these challenges, we propose an attention based deep learning framework, named DeepTriager. Trained and tested on 70,918 clinical records, DeepTriager achieved highly accurate performance on assessment of acuity level I (endangered patients), with AUC of 0.98, which was 0.16 higher than the clinical scale method MEWS and NEWS, and 0.04 higher than traditional machine learning methods. In summary, we presented a new approach for clinical evidence based discovery using a cohort of Electronic Health Records (EHRs). This approach not only outperforms the traditional word segmentation methods but also provides evidence for interpreting the results. Xiaohong Liu 0007, Ken Xie, Ning Chen 0002, Ting Chen 0006 |
BIBM | 5 |
| 2019 | How Robust is Your Automatic Diagnosis Model?abstractAutomatic diagnosis based on clinical notes has become a popular research field recently, and many proposed deep learning models have achieved competitive performance in diseases inference. However, previous research reveals that deep learning models are susceptible to negligibly perturbed inputs named adversarial examples, which contradicts with the safety and reliability requirements of the medical domain. To analyze the vulnerability and robustness of current automatic diagnosis models, we investigate in the generation of adversarial text examples. The main challenges for generating adversarial text examples are divided into three parts. First, the word embedding space is discrete, which makes it hard to perturb as small as adversarial image examples generation. Second, previous adversarial example generation methods focus mainly on multi-class classification models, while automatic diagnosis is a multi-label classification task. Third, the semantic and medical meaning of clinical notes are vital in disease inference, and even small perturbations can change them to a large extent. In this paper, we address the three main challenges and propose Clinical-Attacker, a general framework for both white-box and black-box adversarial text examples generation against automatic diagnosis models. Experimental results on MIMIC-III dataset demonstrate that our framework can easily alter the predictions of automatic diagnosis models with the semantic and medical meaning preserved. Ning Chen 0002, Ting Chen 0006 |
BIBM | 4 |
| 2019 | DCMN: Double Core Memory Network for Patient Outcome Prediction with Multimodal DataabstractMore and more healthcare data are becoming readily available nowadays. These data can help the healthcare professionals and patient themselves to better understand the patient status and potentially lead to improved care quality. However, the analysis of these data are challenging because they are large-scale and heterogeneous, high-dimensional and sparse, temporal but irregularly sampled. In this paper, we propose a method called Double Core Memory Networks (DCMN) to integrate information from different modalities of the longitudinal patient data and learn a joint patient representation effective for downstream analytical tasks such as risk prediction. DCMN is designed not only to disentangle the temporal and non-linear intra-modal dependencies for the data within each modality but also to capture the long-term inter-modal interactions. DCMN models are the end-to-end memory networks with two external memory cores where each modality of data is compressed and stored. Each memory core has an information-flow controller named query to interact with an external memory module. In addition, we incorporate a gating mechanism into basic DCMN model to perform dynamic regulation of memory interaction. DCMN models have multiple computational layers (hops) allowing data of different modalities interacting with each other recurrently along with a mechanism of alternating access of external memory for each memory core hop-by-hop. We evaluate DCMN models on two outcome prediction tasks, including a mortality prediction on the public Medical Information Mart for Intensive Care III (MIMIC-III) database and a cost prediction on the Hospital Quality Monitoring System (HQMS) dataset. Experimental results demonstrate that our DCMN models are more competitive over the baseline methods in the multimodal prediction setting. Yujuan Feng, Ning Chen 0002, Ting Chen 0006, Fei Wang 0001 |
ICDM | 6 |
| 2019 | DeepShape: estimating isoform-level ribosome abundance and distribution with Ribo-seq dataabstractBACKGROUND: Ribosome profiling brings insight to the process of translation. A basic step in profile construction at transcript level is to map Ribo-seq data to transcripts, and then assign a huge number of multiple-mapped reads to similar isoforms. Existing methods either discard the multiple mapped-reads, or allocate them randomly, or assign them proportionally according to transcript abundance estimated from RNA-seq data. RESULTS: Here we present DeepShape, an RNA-seq free computational method to estimate ribosome abundance of isoforms, and simultaneously compute their ribosome profiles using a deep learning model. Our simulation results demonstrate that DeepShape can provide more accurate estimations on both ribosome abundance and profiles when compared to state-of-the-art methods. We applied DeepShape to a set of Ribo-seq data from PC3 human prostate cancer cells with and without PP242 treatment. In the four cell invasion/metastasis genes that are translationally regulated by PP242 treatment, different isoforms show very different characteristics of translational efficiency and regulation patterns. Transcript level ribosome distributions were analyzed by "Codon Residence Index (CRI)" proposed in this study to investigate the relative speed that a ribosome moves on a codon compared to its synonymous codons. We observe consistent CRI patterns in PC3 cells. We found that the translation of several codons could be regulated by PP242 treatment. CONCLUSION: In summary, we demonstrate that DeepShape can serve as a powerful tool for Ribo-seq data analysis. Hongfei Cui, Hailin Hu 0002, Jianyang Zeng 0001, Ting Chen 0006 |
BMC Bioinform. | 4 |
| 2018 | Ontology-based Venous Thromboembolism Risk Factors Mining and Model Developing from Medical Records
Xin Wang 0226, Ning Chen 0002, Juhong Shi, Ting Chen 0006 |
BIBM | 6 |
| 2018 | Dropout training for SVMs with data augmentation
Ning Chen 0002, Jun Zhu 0001, Jianfei Chen 0001, Ting Chen 0006 |
Frontiers Comput. Sci. | 4 |
| 2017 | Patient outcome prediction via convolutional neural networks based on multi-granularity medical concept embeddingabstractThe large availability of biomedical data brings opportunities and challenges to health care. Representation of medical concepts has been well studied in many applications, such as medical informatics, cohort selection, risk prediction, and health care quality measurement. In this paper, we propose an efficient multichannel convolutional neural network (CNN) model based on multi-granularity embeddings of medical concepts named MG-CNN, to examine the effect of individual patient characteristics including demographic factors and medical comorbidities on total hospital costs and length of stay (LOS) by using the Hospital Quality Monitoring System (HQMS) data. The proposed embedding method leverages prior medical hierarchical ontology and improves the quality of embedding for rare medical concepts. The embedded vectors are further visualized by the t-Distributed Stochastic Neighbor Embedding (t-SNE) technique to demonstrate the effectiveness of grouping related medical concepts. Experimental results demonstrate that our MG-CNN model outperforms traditional regression methods based on the one-hot representation of medical concepts, especially in the outcome prediction tasks for patients with low-frequency medical events. In summary, MG-CNN model is capable of mining potential knowledge from the clinical data and will be broadly applicable in medical research and inform clinical decisions. Yujuan Feng, Xu Min, Ning Chen 0002, Xiaolei Xie, Ting Chen 0006 |
BIBM | 7 |
| 2017 | Integrating embeddings of multiple gene networks to prioritize complex disease-associated genesabstractGenome-wide association study (GWAS), as one primary approach for genetic studies, has been successfully applied to a variety of complex diseases, leading to the discovery of substantial disease-associated loci. These discovered associations provide unprecedented opportunities for deepening our understanding of complex diseases, such as disease-associated risk variants, genes, and pathways. However, it is non-trivial to extract biological knowledge from the GWAS data due to the existence of several non-negligible factors. For example, the majority of associated loci fall into noncoding regions without certain links to any genes, complicating its functional characterization. Network-based GWAS gene prioritization, aiming to integrate gene networks with GWAS data, emerges as one promising direction towards solving these challenges and has attracted much attention recently. However, gene networks are usually sparse and noisy, and existing methods do not explicitly consider these properties, leading to suboptimal performance. In this paper, we proposed a novel method called REGENT for integrating multiple gene networks with GWAS data to prioritize complex disease-associated genes. Specifically, we leveraged the network representation learning, a recently developed technique for analyzing social networks, to learn compact and robust embeddings from multiple gene networks. To integrate these learned embeddings of genes with GWAS data, we developed a hierarchical statistical model and derived an efficient inference algorithm for model estimation and prediction. Applying to GWAS data of six complex diseases, we demonstrated that REGENT outperformed existing methods regarding the identification of known disease-associated genes. Also, pathway analysis showed that REGENT helped discover disease-associated pathways. Therefore, our method is expected to be a useful tool for post-GWAS analysis. Mengmeng Wu, Wanwen Zeng, Yi-Jia Zhang 0001, Ting Chen 0006, Rui Jiang 0001 |
BIBM | 5 |
| 2017 | DACE: a scalable DP-means algorithm for clustering extremely large sequence dataabstractMotivation: Advancements in next-generation sequencing technology have produced large amounts of reads at low cost in a short time. In metagenomics, 16S and 18S rRNA gene have been widely used as marker genes to profile diversity of microorganisms in environmental samples. Through clustering of sequencing reads we can determine both number of OTUs and their relative abundance. In many applications, clustering of very large sequencing data with high efficiency and accuracy is essential for downstream analysis. Results: Here, we report a scalable D irichlet Process Means (DP-means) a lgorithm for c lustering e xtremely large sequencing data, termed . With an efficient random projection partition strategy for parallel clustering, DACE can cluster billions of sequences within a couple of hours. Experimental results show that DACE runs between 6 and 80 times faster than state-of-the-art programs, while maintaining overall better clustering accuracy. Using 80 cores, DACE clustered the Lake Taihu 16S rRNA gene sequencing data (∼316M reads, 30 GB) in 25 min, and the Ocean TARA Eukaryotic 18S rRNA gene sequencing data (∼500M reads, 88 GB) into ∼100 000 clusters within an hour. When applied to the IGC gene catalogs in human gut microbiome (∼10M genes), DACE produced 9.8M clusters with 52K redundant genes in 1.5 hours of running time. Availability and Implementation: DACE is available at https://github.com/tinglab/DACE . Contacts: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Linhao Jiang, Yichao Dong, Ning Chen 0002, Ting Chen 0006 |
Bioinform. | 4 |
| 2017 | COCACOLA: binning metagenomic contigs using sequence COmposition, read CoverAge, CO-alignment and paired-end read LinkAgeabstractMotivation: The advent of next-generation sequencing technologies enables researchers to sequence complex microbial communities directly from the environment. Because assembly typically produces only genome fragments, also known as contigs, instead of an entire genome, it is crucial to group them into operational taxonomic units (OTUs) for further taxonomic profiling and down-streaming functional analysis. OTU clustering is also referred to as binning. We present COCACOLA, a general framework automatically bin contigs into OTUs based on sequence composition and coverage across multiple samples. Results: The effectiveness of COCACOLA is demonstrated in both simulated and real datasets in comparison with state-of-art binning approaches such as CONCOCT, GroopM, MaxBin and MetaBAT. The superior performance of COCACOLA relies on two aspects. One is using L 1 distance instead of Euclidean distance for better taxonomic identification during initialization. More importantly, COCACOLA takes advantage of both hard clustering and soft clustering by sparsity regularization. In addition, the COCACOLA framework seamlessly embraces customized knowledge to facilitate binning accuracy. In our study, we have investigated two types of additional knowledge, the co-alignment to reference genomes and linkage of contigs provided by paired-end reads, as well as the ensemble of both. We find that both co-alignment and linkage information further improve binning in the majority of cases. COCACOLA is scalable and faster than CONCOCT, GroopM, MaxBin and MetaBAT. Availability and implementation: The software is available at https://github.com/younglululu/COCACOLA . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Yang Young Lu, Ting Chen 0006, Jed A. Fuhrman, Fengzhu Sun |
Bioinform. | 2 |
| 2017 | Chromatin accessibility prediction via convolutional long short-term memory networks with k-mer embeddingabstractMOTIVATION: Experimental techniques for measuring chromatin accessibility are expensive and time consuming, appealing for the development of computational approaches to predict open chromatin regions from DNA sequences. Along this direction, existing methods fall into two classes: one based on handcrafted k -mer features and the other based on convolutional neural networks. Although both categories have shown good performance in specific applications thus far, there still lacks a comprehensive framework to integrate useful k -mer co-occurrence information with recent advances in deep learning. RESULTS: We fill this gap by addressing the problem of chromatin accessibility prediction with a convolutional Long Short-Term Memory (LSTM) network with k -mer embedding. We first split DNA sequences into k -mers and pre-train k -mer embedding vectors based on the co-occurrence matrix of k -mers by using an unsupervised representation learning approach. We then construct a supervised deep learning architecture comprised of an embedding layer, three convolutional layers and a Bidirectional LSTM (BLSTM) layer for feature learning and classification. We demonstrate that our method gains high-quality fixed-length features from variable-length sequences and consistently outperforms baseline methods. We show that k -mer embedding can effectively enhance model performance by exploring different embedding strategies. We also prove the efficacy of both the convolution and the BLSTM layers by comparing two variations of the network architecture. We confirm the robustness of our model to hyper-parameters by performing sensitivity analysis. We hope our method can eventually reinforce our understanding of employing deep learning in genomic studies and shed light on research regarding mechanisms of chromatin accessibility. AVAILABILITY AND IMPLEMENTATION: The source code can be downloaded from https://github.com/minxueric/ismb2017_lstm . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online. Xu Min, Wanwen Zeng, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001 |
Bioinform. | 4 |
| 2017 | Predicting enhancers with deep convolutional neural networksabstractBACKGROUND: With the rapid development of deep sequencing techniques in the recent years, enhancers have been systematically identified in such projects as FANTOM and ENCODE, forming genome-wide landscapes in a series of human cell lines. Nevertheless, experimental approaches are still costly and time consuming for large scale identification of enhancers across a variety of tissues under different disease status, making computational identification of enhancers indispensable. RESULTS: To facilitate the identification of enhancers, we propose a computational framework, named DeepEnhancer, to distinguish enhancers from background genomic sequences. Our method purely relies on DNA sequences to predict enhancers in an end-to-end manner by using a deep convolutional neural network (CNN). We train our deep learning model on permissive enhancers and then adopt a transfer learning strategy to fine-tune the model on enhancers specific to a cell line. Results demonstrate the effectiveness and efficiency of our method in the classification of enhancers against random sequences, exhibiting advantages of deep learning over traditional sequence-based classifiers. We then construct a variety of neural networks with different architectures and show the usefulness of such techniques as max-pooling and batch normalization in our method. To gain the interpretability of our approach, we further visualize convolutional kernels as sequence logos and successfully identify similar motifs in the JASPAR database. CONCLUSIONS: DeepEnhancer enables the identification of novel enhancers using only DNA sequences via a highly accurate deep learning model. The proposed computational framework can also be applied to similar problems, thereby prompting the use of machine learning methods in life sciences. Xu Min, Wanwen Zeng, Shengquan Chen, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001 |
BMC Bioinform. | 5 |
| 2017 | Reading the Underlying Information From Massive Metagenomic Sequencing DataabstractMicroorganisms are everywhere. Recent studies showed that the mixture of microbes or the microbiome on the human body plays important roles in human physiology and diseases. Metagenomic sequencing is a key technology for studying microbiomes. It produces massive amounts of data in the form of short sequencing reads. A single metagenomic sample can contain 107to 108reads of about 100-nucleotide (nt) length each in a typical shotgun metagenomic sequencing study. They contain rich information about microbiomes and their functions, but reading out those information from the huge highly fragmented data has multiple challenges for mathematical models, bioinformatics methods, and computer algorithms. In this paper, we review the basic bioinformatics tasks and existing methods in processing and analyzing metagenomic data, and discuss remaining open challenges and practical observations. The aim of the paper is to provide readers a whole picture of metagenomic data processing and analysis, and a reference and perspective to start with for computational scientists who are interested in this exciting field. Xuegong Zhang, Shansong Liu, Hongfei Cui, Ting Chen 0006 |
Proc. IEEE | 4 |
| 2016 | DeepEnhancer: Predicting enhancers by convolutional neural networksabstractEnhancers are crucial to the understanding of mechanisms underlying gene transcriptional regulation. Although having been successfully applied in such projects as ENCODE and Roadmap to generate landscape of enhancers in human cell lines, high-throughput biological experimental techniques are still costly and time consuming for even larger scale identification of enhancers across a variety of tissues under different disease status, making computational identification of enhancers indispensable. In this paper, we propose a computational framework, named DeepEnhancer, to classify enhancers from background genomic sequences. We construct convolutional neural networks of various architectures and compare the classification performance with traditional sequence-based classifiers. We first train the deep learning model on the FANTOM5 permissive enhancer dataset, and then fine-tune the model on ENCODE cell type-specific enhancer datasets by adopting the transfer learning strategy. Experimental results demonstrate that DeepEnhancer has superior efficiency and effectiveness in classification tasks, and the use of max-pooling and batch normalization is beneficial to higher accuracy. To make our approach more understandable, we propose a strategy to visualize the convolutional kernels as sequence logos and compare them against the JASPAR database using TOMTOM. In summary, DeepEnhancer allows researchers to train highly accurate deep models and will be broadly applicable in computational biology. Xu Min, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001 |
BIBM | 3 |
| 2016 | mLDM: A New Hierarchical Bayesian Statistical Model for Sparse Microbial Association Discovery
Ning Chen 0002, Ting Chen 0006 |
RECOMB | 3 |
| 2016 | Global inference of disease-causing single nucleotide variants from exome sequencing dataabstractBACKGROUND: Whole exome sequencing (WES) has recently emerged as an effective approach for identifying genetic variants underlying human diseases. However, considerable time and labour is needed for careful investigation of candidate variants. Although filtration based on population frequencies and functional prediction scores could effectively remove common and neutral variants, hundreds or even thousands of rare deleterious variants still remain. In addition, current WES platforms also provide variant information in flanking noncoding regions, such as promoters, introns and splice sites. Despite of being recognized to harbour causal variants, these regions are usually ignored by current analysis pipelines. RESULTS: We present a novel computational method, called Glints, to overcome the above limitations. Glints is capable of identifying disease-causing SNVs in both coding and flanking noncoding regions from exome sequencing data. The principle behind Glints is that disease-causing variants should manifest their effect at both variant and gene levels. Specifically, Glints integrates 14 types of functional scores, including predictions for both coding and noncoding variants, and 9 types of association scores, which help identifying disease relevant genes. We conducted a large-scale simulation studies based on 1000 Genomes Project data and demonstrated the effectiveness of our method in both coding and flanking noncoding regions. We also applied Glints in two real exome sequencing and demonstrated its effectiveness for uncovering disease-causing SNVs. Both standalone software and web server are available at our website http://bioinfo.au.tsinghua.edu.cn/jianglab/glints . CONCLUSIONS: Glints is effective for uncovering disease-causing SNVs in coding and flanking noncoding regions, which is supported by both simulation and real case studies. Glints is expected to be a useful tool for human genetics research based on exome sequencing data. Mengmeng Wu, Ting Chen 0006, Rui Jiang 0001 |
BMC Bioinform. | 2 |
| 2015 | Global optimization-based inference of chemogenomic features from drug-target interactionsabstractMOTIVATION: Gaining insight into chemogenomic drug-target interactions, such as those involving the substructures of synthetic drugs and protein domains, is important in fragment-based drug discovery and drug repositioning. Previous studies evaluated the interactions locally, thereby ignoring the competitive effects of different substructures or domains, but this could lead to high false-positive estimation, calling for a computational method that presents more predictive power. RESULTS: A statistical model, termed Global optimization-based InFerence of chemogenomic features from drug-Target interactions, or GIFT, is proposed herein to evaluate substructure-domain interactions globally such that all substructure-domain contributions to drug-target interaction are analyzed simultaneously. Combinations of different chemical substructures were included since they may function as one unit. When compared to previous methods, GIFT showed better interpretive performance, and performance for the recovery of drug-target interactions was good. Among 53 known drug-domain interactions, 81% were accurately predicted by GIFT. Eighteen of the top 100 predicted combined substructure-domain interactions had corresponding drug-target structures in the Protein Data Bank database, and 15 out of the 18 had been proved. GIFT was then implemented to predict substructure-domain interactions based on drug repositioning. For example, the anticancer activities of tazarotene, adapalene, acitretin and raloxifene were identified. In summary, GIFT is a global chemogenomic inference approach and offers fresh insight into drug-target interactions. Songpeng Zu, Ting Chen 0006, Shao Li |
Bioinform. | 2 |
| 2015 | Integrative Data Analysis of Multi-Platform Cancer Data with a Multimodal Deep Learning ApproachabstractIdentification of cancer subtypes plays an important role in revealing useful insights into disease pathogenesis and advancing personalized therapy. The recent development of high-throughput sequencing technologies has enabled the rapid collection of multi-platform genomic data (e.g., gene expression, miRNA expression, and DNA methylation) for the same set of tumor samples. Although numerous integrative clustering approaches have been developed to analyze cancer data, few of them are particularly designed to exploit both deep intrinsic statistical properties of each input modality and complex cross-modality correlations among multi-platform input data. In this paper, we propose a new machine learning model, called multimodal deep belief network (DBN), to cluster cancer patients from multi-platform observation data. In our integrative clustering framework, relationships among inherent features of each single modality are first encoded into multiple layers of hidden variables, and then a joint latent model is employed to fuse common features derived from multiple input modalities. A practical learning algorithm, called contrastive divergence (CD), is applied to infer the parameters of our multimodal DBN model in an unsupervised manner. Tests on two available cancer datasets show that our integrative data analysis approach can effectively extract a unified representation of latent features to capture both intra- and cross-modality correlations, and identify meaningful disease subtypes from multi-platform cancer data. In addition, our approach can identify key genes and miRNAs that may play distinct roles in the pathogenesis of different cancer subtypes. Among those key miRNAs, we found that the expression level of miR-29a is highly correlated with survival time in ovarian cancer patients. These results indicate that our multimodal DBN based data analysis approach may have practical applications in cancer pathogenesis studies and provide useful guidelines for personalized cancer therapy. Muxuan Liang, Ting Chen 0006, Jianyang Zeng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2014 | Integrative approaches for predicting protein function and prioritizing genes for complex phenotypes using protein interaction networksabstractWith the rapid development of biotechnologies, many types of biological data including molecular networks are now available. However, to obtain a more complete understanding of a biological system, the integration of molecular networks with other data, such as molecular sequences, protein domains and gene expression profiles, is needed. A key to the use of networks in biological studies is the definition of similarity among proteins over the networks. Here, we review applications of similarity measures over networks with a special focus on the following four problems: (i) predicting protein functions, (ii) prioritizing genes related to a phenotype given a set of seed genes that have been shown to be related to the phenotype, (iii) prioritizing genes related to a phenotype by integrating gene expression profiles and networks and (iv) identification of false positives and false negatives from RNAi experiments. Diffusion kernels are demonstrated to give superior performance in all these tasks, leading to the suggestion that diffusion kernels should be the primary choice for a network similarity metric over other similarity measures such as direct neighbors and shortest path distance. Xiaotu Ma, Ting Chen 0006, Fengzhu Sun |
Briefings Bioinform. | 2 |
| 2006 | Searching for interpretable rules for disease mutations: a simulated annealing bump hunting strategyabstractBACKGROUND: Understanding how amino acid substitutions affect protein functions is critical for the study of proteins and their implications in diseases. Although methods have been developed for predicting potential effects of amino acid substitutions using sequence, three-dimensional structural, and evolutionary properties of proteins, the applications are limited by the complication of the features and the availability of protein structural information. Another limitation is that the prediction results are hard to be interpreted with physicochemical principles and biological knowledge. RESULTS: To overcome these limitations, we proposed a novel feature set using physicochemical properties of amino acids, evolutionary profiles of proteins, and protein sequence information. We applied the support vector machine and the random forest with the feature set to experimental amino acid substitutions occurring in the E. coli lac repressor and the bacteriophage T4 lysozyme, as well as to annotated amino acid substitutions occurring in a wide range of human proteins. The results showed that the proposed feature set was superior to the existing ones. To explore physicochemical principles behind amino acid substitutions, we designed a simulated annealing bump hunting strategy to automatically extract interpretable rules for amino acid substitutions. We applied the strategy to annotated human amino acid substitutions and successfully extracted several rules which were either consistent with current biological knowledge or providing new insights for the understanding of amino acid substitutions. When applied to unclassified data, these rules could cover a large portion of samples, and most of the covered samples showed good agreement with predictions made by either the support vector machine or the random forest. CONCLUSION: The prediction methods using the proposed feature set can achieve larger AUC (the area under the ROC curve), smaller BER (the balanced error rate), and larger MCC (the Matthews' correlation coefficient) than those using the published feature sets, suggesting that our feature set is superior to the existing ones. The rules extracted by the simulated annealing bump hunting strategy have comparable coverage and accuracy but much better interpretability as those extracted by the patient rule induction method (PRIM), revealing that the strategy is more effective in inducing interpretable rules. Rui Jiang 0001, Fengzhu Sun, Ting Chen 0006 |
BMC Bioinform. | 4 |
| 2006 | An integrated approach to the prediction of domain-domain interactionsabstractBACKGROUND: The development of high-throughput technologies has produced several large scale protein interaction data sets for multiple species, and significant efforts have been made to analyze the data sets in order to understand protein activities. Considering that the basic units of protein interactions are domain interactions, it is crucial to understand protein interactions at the level of the domains. The availability of many diverse biological data sets provides an opportunity to discover the underlying domain interactions within protein interactions through an integration of these biological data sets. RESULTS: We combine protein interaction data sets from multiple species, molecular sequences, and gene ontology to construct a set of high-confidence domain-domain interactions. First, we propose a new measure, the expected number of interactions for each pair of domains, to score domain interactions based on protein interaction data in one species and show that it has similar performance as the E-value defined by Riley et al. Our new measure is applied to the protein interaction data sets from yeast, worm, fruitfly and humans. Second, information on pairs of domains that coexist in known proteins and on pairs of domains with the same gene ontology function annotations are incorporated to construct a high-confidence set of domain-domain interactions using a Bayesian approach. Finally, we evaluate the set of domain-domain interactions by comparing predicted domain interactions with those defined in iPfam database that were derived based on protein structures. The accuracy of predicted domain interactions are also confirmed by comparing with experimentally obtained domain interactions from H. pylori. As a result, a total of 2,391 high-confidence domain interactions are obtained and these domain interactions are used to unravel detailed protein and domain interactions in several protein complexes. CONCLUSION: Our study shows that integration of multiple biological data sets based on the Bayesian approach provides a reliable framework to predict domain interactions. By integrating multiple data sources, the coverage and accuracy of predicted domain interactions can be significantly increased. Minghua Deng, Fengzhu Sun, Ting Chen 0006 |
BMC Bioinform. | 4 |
| 2005 | HapBlock: haplotype block partitioning and tag SNP selection software using a set of dynamic programming algorithmsabstractUNLABELLED: Recent studies have revealed that linkage disequilibrium (LD) patterns vary across the human genome with some regions of high LD interspersed with regions of low LD. Such LD patterns make it possible to select a set of single nucleotide polymorphism (SNPs; tag SNPs) for genome-wide association studies. We have developed a suite of computer programs to analyze the block-like LD patterns and to select the corresponding tag SNPs. Compared to other programs for haplotype block partitioning and tag SNP selection, our program has several notable features. First, the dynamic programming algorithms implemented are guaranteed to find the block partition with minimum number of tag SNPs for the given criteria of blocks and tag SNPs. Second, both haplotype data and genotype data from unrelated individuals and/or from general pedigrees can be analyzed. Third, several existing measures/criteria for haplotype block partitioning and tag SNP selection have been implemented in the program. Finally, the programs provide flexibility to include specific SNPs (e.g. non-synonymous SNPs) as tag SNPs. AVAILABILITY: The HapBlock program and its supplemental documents can be downloaded from the website http://www.cmb.usc.edu/~msms/HapBlock. Zhaohui S. Qin, Ting Chen 0006, Jun S. Liu, Michael S. Waterman, Fengzhu Sun |
Bioinform. | 3 |
| 2005 | Selecting additional tag SNPs for tolerating missing data in genotypingabstractBACKGROUND: Recent studies have shown that the patterns of linkage disequilibrium observed in human populations have a block-like structure, and a small subset of SNPs (called tag SNPs) is sufficient to distinguish each pair of haplotype patterns in the block. In reality, some tag SNPs may be missing, and we may fail to distinguish two distinct haplotypes due to the ambiguity caused by missing data. RESULTS: We show there exists a subset of SNPs (referred to as robust tag SNPs) which can still distinguish all distinct haplotypes even when some SNPs are missing. The problem of finding minimum robust tag SNPs is shown to be NP-hard. To find robust tag SNPs efficiently, we propose two greedy algorithms and one linear programming relaxation algorithm. The experimental results indicate that (1) the solutions found by these algorithms are quite close to the optimal solution; (2) the genotyping cost saved by using tag SNPs can be as high as 80%; and (3) genotyping additional tag SNPs for tolerating missing data is still cost-effective. CONCLUSION: Genotyping robust tag SNPs is more practical than just genotyping the minimum tag SNPs if we can not avoid the occurrence of missing data. Our theoretical analysis and experimental results show that the performance of our algorithms is not only efficient but the solution found is also close to the optimal solution. Yao-Ting Huang, Ting Chen 0006, Kun-Mao Chao |
BMC Bioinform. | 3 |
| 2004 | Approximation Algorithms for the Selection of Robust Tag SNPs
Yao-Ting Huang, Ting Chen 0006, Kun-Mao Chao |
WABI | 3 |
| 2004 | Mapping gene ontology to proteins based on protein-protein interaction dataabstractMOTIVATION: Gene Ontology (GO) consortium provides structural description of protein function that is used as a common language for gene annotation in many organisms. Large-scale techniques have generated many valuable protein-protein interaction datasets that are useful for the study of protein function. Combining both GO and protein-protein interaction data allows the prediction of function for unknown proteins. RESULT: We apply a Markov random field method to the prediction of yeast protein function based on multiple protein-protein interaction datasets. We assign function to unknown proteins with a probability representing the confidence of this prediction. The functions are based on three general categories of cellular component, molecular function and biological process defined in GO. The yeast proteins are defined in the Saccharomyces Genome Database (SGD). The protein-protein interaction datasets are obtained from the Munich Information Center for Protein Sequences (MIPS), including physical interactions and genetic interactions. The efficiency of our prediction is measured by applying the leave-one-out validation procedure to a functional path matching scheme, which compares the prediction with the GO description of a protein's function from the abstract level to the detailed level along the GO structure. For biological process, the leave-one-out validation procedure shows 52% precision and recall of our method, much better than that of the simple guilty-by-association methods. Minghua Deng, Zhidong Tu, Fengzhu Sun, Ting Chen 0006 |
Bioinform. | 4 |
| 2003 | An integrated probabilistic model for functional prediction of proteinsabstractWe develop an integrated probabilistic model to combine protein physical interactions, genetic interactions, highly correlated gene expression network, protein complex data, and domain structures of individual proteins to predict protein functions. The model is an extension of our previous model for protein function prediction based on Markovian random field theory. The model is flexible in that other protein pairwise relationship information and features of individual proteins can be easily incorporated. Two features distinguish the integrated approach from other available methods for protein function prediction. One is that the integrated approach uses all available sources of information with different weights for different sources of data. It is a global approach that takes the whole network into consideration. The second feature is that the posterior probability that a protein has the function of interest is assigned. The posterior probability indicates how confident we are about assigning the function to the protein. We apply our integrated approach to predict functions of yeast proteins based upon MIPS protein function classifications and upon the interaction networks based on MIPS physical and genetic interactions, gene expression profiles, Tandem Affinity Purification (TAP) protein complex data, and protein domain information. We study the sensitivity and specificity of the integrated approach using different sources of information by the leave-one-out approach. In contrast to using MIPS physical interactions only, the integrated approach combining all of the information increases the sensitivity from 57% to 87% when the specificity is set at 57%-an increase of 30%. It should also be noted that enlarging the interaction network greatly increases the number of proteins whose functions can be predicted. Minghua Deng, Ting Chen 0006, Fengzhu Sun |
RECOMB | 2 |
| 2003 | Dynamic programming algorithms for haplotype block partitioning: applications to human chromosome 21 haplotype dataabstractRecent studies have shown that the human genome has a haplotype block structure such that it can be divided into discrete blocks of limited haplotype diversity. Patil et al. [6] and Zhang et al. [12] developed algorithms to partition haplotypes into blocks with minimum number of tag SNPs for the entire chromosome. However, it is not clear how to partition haplotypes into blocks with restricted number of SNPs when only limited resources are available. In this paper, we first formulated this problem as finding a block partition with a fixed number of tag SNPs that can cover the maximal percentage of a genome. Then we solved it by two dynamic programming algorithms, which are fairly flexible to take into account the knowledge of functional polymorphism. We applied our algorithms to the published SNP data of human chromosome 21 combining with the functional information of these SNPs and demonstrated the effectiveness of them. Statistical investigation of the relationship between the starting points of a block partition and the coding and non-coding regions illuminated that the SNPs at these starting points are not significantly enriched in coding regions. We also developed an efficient algorithm to find all possible long local maximal haplotypes across a subset of samples. After applying this algorithm to the human chromosome 21 haplotype data, we found that samples with long local haplotypes are not necessarily globally similar. Fengzhu Sun, Michael S. Waterman, Ting Chen 0006 |
RECOMB | 4 |
| 2002 | Inferring domain-domain interactions from protein-protein interactionsabstractProtein-protein interactions are important events in cellular and biochemical processes within a cell. Several researchers have undertaken the task of analyzing protein-protein interactions covering all genes of an organism by using yeast two-hybrid assays. Protein-protein interactions involve physical interactions between protein domains. Therefore, understanding protein interactions at the domain level gives a global view of the protein interaction network, and possibly extends functions of proteins. In this study, we present a Maximum Likelihood approach to infer domain-domain interactions from the 5719 yeast protein-protein interactions obtained in the high throughput two-hybrid experiments by Uetz et al., 2000 and Ito et al., 2001. The accuracies of our predictions are measured at the protein level. Our study includes the following three results: (1) using the inferred domain-domain interactions, we predict interactions between proteins and achieve 39.0% specificity and 79.7% sensitivity; (2) our predicted protein-protein interactions have a significant overlap with the MIPS(http://mips.gfs.de) protein-protein interactions obtained by methods other than the two-hybrid systems; and (3) the mean correlation coefficient of the gene expression profiles for our predicted interacting pairs is significantly higher than that for random pairs as well as that of interacting pairs in Uetz's and Ito's experimental data. Our method has shown robustness in analyzing incomplete data sets and dealing with various experimental errors. We find several novel protein-protein interactions such as RPS0A interacting with APG17 and TAF40 interacting with SPT3, which are consistent with the functions of the proteins. Minghua Deng, Shipra Mehta, Fengzhu Sun, Ting Chen 0006 |
RECOMB | 4 |