Nan Du 0001

dblp:86/4539-1 · DBLP profile ↗
← Back
41ranked-venue papers
7as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 23 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2024 FIT-graph: A multi-grained evolutionary graph based framework for disease diagnosis
abstract
Early assessment, with the help of machine learning methods, can aid clinicians in optimizing the diagnosis and treatment process, allowing patients to receive critical treatment time. Due to the advantages of effective information organization and interpretable reasoning, knowledge graph-based methods have become one of the most widely used machine learning algorithms for this task. However, due to a lack of effective organization and use of multi-granularity and temporal information, current knowledge graph-based approaches are hard to fully and comprehensively exploit the information contained in medical records, restricting their capacity to make superior quality diagnoses. To address these challenges, we examine and study disease diagnosis applications in-depth, and propose a novel disease diagnosis framework named FIT-Graph. With novel medical multi-grained evolutionary graphs, FIT-Graph efficiently organizes the extracted information from various granularities and time stages, maximizing the retention of valuable information for disease inference and ensuring the comprehensiveness and validity of the final disease inference. We compare FIT-Graph with two real-world clinical datasets from cardiology and respiratory departments with the baseline. The experimental results show that its effect is better than the baseline model, and the baseline performance of the task is improved by about 5% in multiple indices.
Zizhu Liu, Nan Du 0001, Huizhen Shu, Erheng Zhong, Nan Jiang 0021, Qiaoran Chen, Ying Shen 0001
Artif. Intell. Medicine3
2022 Modeling path information for knowledge graph completion
Ying Shen 0001, Dagang Li 0001, Nan Du 0001
Neural Comput. Appl.3
2021 Knowledge-Guided Efficient Representation Learning for Biomedical Domain
abstract
Pre-trained concept representations are essential to many biomedical text mining and natural language processing tasks. As such, various representation learning approaches have been proposed in the literature. More recently, contextualized embedding approaches (i.e., BERT based models) that capture the implicit semantics of concepts at a granular level have significantly outperformed the conventional word embedding approaches (i.e., Word2Vec/GLoVE based models). Despite significant accuracy gains achieved, these approaches are often computationally expensive and memory inefficient. To address this issue, we propose a new representation learning approach that efficiently adapts the concept representations to the newly available data. Specifically, the proposed approach develops a knowledge-guided continual learning strategy wherein the accurate/stable context-information present in human-curated knowledge-bases is exploited to continually identify and retrain the representations of those concepts whose corpus-based context evolved coherently over time. Different from previous studies that mainly leverage the curated knowledge to improve the accuracy of embedding models, the proposed research explores the usefulness of semantic knowledge from the perspective of accelerating the training efficiency of embedding models. Comprehensive experiments under various efficiency constraints demonstrate that the proposed approach significantly improves the computational performance of biomedical word embedding models.
Kishlay Jha, Guangxu Xun, Nan Du 0001, Aidong Zhang 0001
KDD3
2021 Practical fine-grained learning based anomaly classification for ECG image
Nan Du 0001, Ming Zuo, Jingsheng Lin, Nathan Liu, Erheng Zhong, Zizhu Liu, Qiaoran Chen, Ying Shen 0001
Artif. Intell. Medicine2
2021 FM-ECG: A fine-grained multi-label framework for ECG image classification
Nan Du 0001, Nathan Liu, Erheng Zhong, Zizhu Liu, Ying Shen 0001
Inf. Sci.1
2021 Tracking Community Consistency in Dynamic Networks: An Influence-Based Approach
abstract
The dynamic network data have become ubiquitous with the rapid development of Internet and smart devices. To effectively manage the involved vertices in networks, it is crucial to track the special community patterns and analyze the relationships among vertices. In this paper, we propose a new method to measure the coherence strength, also referred to as community consistency, of a community over a specific observation period. The measurement of community consistency is especially challenging given the dynamic community structure over time, i.e., vertices can leave their original communities and join new communities. In order to interpret the causes of evolving community structure and model the influence of evolving community structure on community consistency, we introduce an influence propagation process having a causal relation with the community consistency. Specifically, a generative model is proposed to combine the influence propagation and the network topological structure at each time step. The proposed influence-based approach for modeling evolution can be instantiated in a variety of real-world network data. The comprehensive experiments on both synthetic and real-world datasets demonstrate the superiority of the proposed framework in estimating the community consistency. Besides, we conduct a case study to show the effectiveness of the proposed method in real-world applications.
Xiaowei Jia, Nan Du 0001, Yuan Zhang 0028, Vishrawas Gopalakrishnan, Guangxu Xun, Aidong Zhang 0001
IEEE Trans. Knowl. Data Eng.3
2020 On the Generation of Medical Question-Answer Pairs
abstract
Question answering (QA) has achieved promising progress recently. However, answering a question in real-world scenarios like the medical domain is still challenging, due to the requirement of external knowledge and the insufficient quantity of high-quality training data. In the light of these challenges, we study the task of generating medical QA pairs in this paper. With the insight that each medical question can be considered as a sample from the latent distribution of questions given answers, we propose an automated medical QA pair generation framework, consisting of an unsupervised key phrase detector that explores unstructured material for validity, and a generator that involves a multi-pass decoder to integrate structural knowledge for diversity. A series of experiments have been conducted on a real-world dataset collected from the National Medical Licensing Examination of China. Both automatic evaluation and human annotation demonstrate the effectiveness of the proposed method. Further investigation shows that, by incorporating the generated QA pairs for training, significant improvement in terms of accuracy can be achieved for the examination QA system. 1
Sheng Shen 0001, Yaliang Li, Nan Du 0001, Xian Wu 0001, Yusheng Xie, Shen Ge, Tao Yang 0012, Xingzheng Liang, Wei Fan 0001
AAAI3
2020 Entity Synonym Discovery via Multipiece Bilateral Context Matching
abstract
Being able to automatically discover synonymous entities in an open-world setting benefits various tasks such as entity disambiguation or knowledge graph canonicalization. Existing works either only utilize entity features, or rely on structured annotations from a single piece of context where the entity is mentioned. To leverage diverse contexts where entities are mentioned, in this paper, we generalize the distributional hypothesis to a multi-context setting and propose a synonym discovery framework that detects entity synonyms from free-text corpora with considerations on effectiveness and robustness. As one of the key components in synonym discovery, we introduce a neural network model SynonymNet to determine whether or not two given entities are synonym with each other. Instead of using entities features, SynonymNet makes use of multiple pieces of contexts in which the entity is mentioned, and compares the context-level similarity via a bilateral matching schema. Experimental results demonstrate that the proposed model is able to detect synonym sets that are not observed during training on both generic and domain-specific datasets: Wiki+Freebase, PubMed+UMLS, and MedBook+MKG, with up to 4.16% improvement in terms of Area Under the Curve and 3.19% in terms of Mean Average Precision compared to the best baseline method.
Yaliang Li, Nan Du 0001, Wei Fan 0001, Philip S. Yu
IJCAI3
2020 Extracting Medical Knowledge from Crowdsourced Question Answering Website
abstract
The medical crowdsourced question answering (Q&A) websites are booming in recent years, and an increasingly large amount of patients and doctors are involved. The valuable information from these medical crowdsourced Q&A websites can benefit patients, doctors and the society. One key to unleash the power of these Q&A websites is to extract medical knowledge from the noisy question-answer pairs and filter out unrelated or even incorrect information. Facing the daunting scale of information generated on medical Q&A websites everyday, it is unrealistic to fulfill this task via supervised method due to the expensive annotation cost. In this paper, we propose a Medical Knowledge Extraction (MKE) system that can automatically provide high-quality knowledge triples extracted from the noisy question-answer pairs, and at the same time, estimate expertise for the doctors who give answers on these Q&A websites. The MKE system is built upon a truth discovery framework, where we jointly estimate trustworthiness of answers and doctor expertise from the data without any supervision. We further tackle three unique challenges in the medical knowledge extraction task, namely representation of noisy input, multiple linked truths, and the long-tail phenomenon in the data. The MKE system is applied to real-world datasets crawled fromxywy.com, one of the most popular medical crowdsourced Q&A websites. Both quantitative evaluation and case studies demonstrate that the proposed MKE system can successfully provide useful medical knowledge and accurate doctor expertise. We further demonstrate a real-world application,Ask A Doctor, which can automatically give patients suggestions to their questions.
Yaliang Li, Chaochun Liu, Nan Du 0001, Wei Fan 0001, Qi Li 0012, Jing Gao 0004
IEEE Trans. Big Data3
2019 Multi-Task Learning with Multi-View Attention for Answer Selection and Knowledge Base Question Answering
abstract
Answer selection and knowledge base question answering (KBQA) are two important tasks of question answering (QA) systems. Existing methods solve these two tasks separately, which requires large number of repetitive work and neglects the rich correlation information between tasks. In this paper, we tackle answer selection and KBQA tasks simultaneously via multi-task learning (MTL), motivated by the following motivations. First, both answer selection and KBQA can be regarded as a ranking problem, with one at text-level while the other at knowledge-level. Second, these two tasks can benefit each other: answer selection can incorporate the external knowledge from knowledge base (KB), while KBQA can be improved by learning contextual information from answer selection. To fulfill the goal of jointly learning these two tasks, we propose a novel multi-task learning scheme that utilizes multi-view attention learned from various perspectives to enable these tasks to interact with each other as well as learn more comprehensive sentence representations. The experiments conducted on several real-world datasets demonstrate the effectiveness of the proposed method, and the performance of answer selection and KBQA is improved. Also, the multi-view attention scheme is proved to be effective in assembling attentive information from different representational perspectives.
Yang Deng 0002, Yuexiang Xie, Yaliang Li, Min Yang 0007, Nan Du 0001, Wei Fan 0001, Kai Lei, Ying Shen 0001
AAAI5
2019 Multi-grained Named Entity Recognition
abstract
This paper presents a novel framework, MGNER, for Multi-Grained Named Entity Recognition where multiple entities or entity mentions in a sentence could be nonoverlapping or totally nested.Different from traditional approaches regarding NER as a sequential labeling task and annotate entities consecutively, MGNER detects and recognizes entities on multiple granularities: it is able to recognize named entities without explicitly assuming non-overlapping or totally nested structures.MGNER consists of a Detector that examines all possible word segments and a Classifier that categorizes entities.In addition, contextual information and a self-attention mechanism are utilized throughout the framework to improve the NER performance.Experimental results show that MGNER outperforms current state-of-the-art baselines up to 4.4% in terms of the F1 score among nested/non-overlapping NER tasks.* Work was done when the author Yaliang Li was at Tencent America.
Congying Xia, Tao Yang 0012, Yaliang Li, Nan Du 0001, Xian Wu 0001, Wei Fan 0001, Fenglong Ma, Philip S. Yu
ACL (1)5
2019 Joint Slot Filling and Intent Detection via Capsule Neural Networks
abstract
Being able to recognize words as slots and detect the intent of an utterance has been a keen issue in natural language understanding. The existing works either treat slot filling and intent detection separately in a pipeline manner, or adopt joint models which sequentially label slots while summarizing the utterance-level intent without explicitly preserving the hierarchical relationship among words, slots, and intents. To exploit the semantic hierarchy for effective modeling, we propose a capsule-based neural network model which accomplishes slot filling and intent detection via a dynamic routing-by-agreement schema. A re-routing schema is proposed to further synergize the slot filling performance using the inferred intent representation. Experiments on two real-world datasets show the effectiveness of our model when compared with other alternative model architectures, as well as existing natural language understanding services.
Yaliang Li, Nan Du 0001, Wei Fan 0001, Philip S. Yu
ACL (1)3
2019 MedTruth: A Semi-supervised Approach to Discovering Knowledge Condition Information from Multi-Source Medical Data
abstract
Knowledge Graph (KG) contains entities and the relations between entities. Due to its representation ability, KG has been successfully applied to support many medical/healthcare tasks. However, in the medical domain, knowledge holds under certain conditions. Such conditions for medical knowledge are crucial for decision-making in various medical applications, which is missing in existing medical KGs. In this paper, we aim to discovery medical knowledge conditions from texts to enrich KGs. Electronic Medical Records (EMRs) are systematized collection of clinical data and contain detailed information about patients, thus EMRs can be a good resource to discover medical knowledge conditions. Unfortunately, the amount of available EMRs is limited due to reasons such as regularization. Meanwhile, a large amount of medical question answering (QA) data is available, which can greatly help the studied task. However, the quality of medical QA data is quite diverse, which may degrade the quality of the discovered medical knowledge conditions. In the light of these challenges, we propose a new truth discovery method, MedTruth, for medical knowledge condition discovery, which incorporates prior source quality information into the source reliability estimation procedure, and also utilizes the knowledge triple information for trustworthy information computation. We conduct series of experiments on real-world medical datasets to demonstrate that the proposed method can discover meaningful and accurate conditions for medical knowledge by leveraging both EMR and QA data. Further, the proposed method is tested on synthetic datasets to validate its effectiveness under various scenarios.
Yang Deng 0002, Yaliang Li, Ying Shen 0001, Nan Du 0001, Wei Fan 0001, Min Yang 0007, Kai Lei
CIKM4
2019 Path-based Attribute-aware Representation Learning for Relation Prediction
abstract
Knowledge graphs (KGs) have been applied to many semantic-driven applications, including knowledge interchange and semantic inference. However, most KGs are far from complete and are growing rapidly. Although significant progress has been made in the symbolic representation learning of KGs with structural information, the textual knowledge that plays a crucial role in relation prediction is underutilized, and the issues of redundancy and noise path remain to be settled. In this paper, a Path-based Attribute-aware Representation Learning model (PARL) has been proposed to perform path denoising and path representation learning for the relation prediction task. We develop a novel text-enhanced relation prediction architecture, which interactively learns KG structural and textual representations to vary the sparsity and reliability of KG. Moreover, a path denoising algorithm is presented to emphasize paths with rich information and reduce the impact of redundancy and noise path. Experiments on a public dataset demonstrate that PARL consistently outperforms state-of-the-art methods on relation prediction and KG completion tasks.
Ying Shen 0001, Desi Wen, Yaliang Li, Nan Du 0001, Hai-Tao Zheng 0002, Min Yang 0007
SDM4
2019 MCVAE: Margin-based Conditional Variational Autoencoder for Relation Classification and Pattern Generation
abstract
Relation classification is a basic yet important task in natural language processing. Existing relation classification approaches mainly rely on distant supervision, which assumes that a bag of sentences mentioning a pair of entities and extracted from a given corpus should express the same relation type of this entity pair. The training of these models needs a lot of high-quality bag-level data. However, in some specific domains, such as medical domain, it is difficult to obtain sufficient and high-quality sentences in a text corpus that mention two entities with a certain medical relation between them. In such a case, it is hard for existing discriminative models to capture the representative features (i.e., common patterns) from diversely expressed entity pairs with a given relation. Thus, the classification performance cannot be guaranteed when limited features are obtained from the corpus. To address this challenge, in this paper, we propose to employ a generative model, called conditional variational autoencoder (CVAE), to handle the pattern sparsity. We define that each relation has an individually learned latent distribution from all possible sentences expressing this relation. As these distributions are learned based on the purpose of input reconstruction, the model's classification ability may not be strong enough and should be improved. By distinguishing the differences among different relation distributions, a margin-based regularizer is designed, which leads to a margin-based CVAE (MCVAE) that can significantly enhance the classification ability. Besides, MCVAE can automatically generate semantically meaningful patterns that describe the given relations. Experiments on two real-world datasets validate the effectiveness of the proposed MCVAE on the tasks of relation classification and relation-specific pattern generation.
Fenglong Ma, Yaliang Li, Jing Gao 0004, Nan Du 0001, Wei Fan 0001
WWW5
2018 Drug2Vec: Knowledge-aware Feature-driven Method for Drug Representation Learning
Ying Shen 0001, Kaiqi Yuan, Yaliang Li, Buzhou Tang, Min Yang 0007, Nan Du 0001, Kai Lei
BIBM6
2018 Knowledge as A Bridge: Improving Cross-domain Answer Selection with External Knowledge
abstract
Answer selection is an important but challenging task. Significant progresses have been made in domains where a large amount of labeled training data is available. However, obtaining rich annotated data is a time-consuming and expensive process, creating a substantial barrier for applying answer selection models to a new domain which has limited labeled data. In this paper, we propose Knowledge-aware Attentive Network (KAN), a transfer learning framework for cross-domain answer selection, which uses the knowledge base as a bridge to enable knowledge transfer from the source domain to the target domains. Specifically, we design a knowledge module to integrate the knowledge-based representational learning into answer selection models. The learned knowledge-based representations are shared by source and target domains, which not only leverages large amounts of cross-domain data, but also benefits from a regularization effect that leads to more general representations to help tasks in new domains. To verify the effectiveness of our model, we use SQuAD-T dataset as the source domain and three other datasets (i.e., Yahoo QA, TREC QA and InsuranceQA) as the target domains. The experimental results demonstrate that KAN has remarkable applicability and generality, and consistently outperforms the strong competitors by a noticeable margin for cross-domain answer selection.
Yang Deng 0002, Ying Shen 0001, Min Yang 0007, Yaliang Li, Nan Du 0001, Wei Fan 0001, Kai Lei
COLING5
2018 Cooperative Denoising for Distantly Supervised Relation Extraction
abstract
Distantly supervised relation extraction greatly reduces human efforts in extracting relational facts from unstructured texts. However, it suffers from noisy labeling problem, which can degrade its performance. Meanwhile, the useful information expressed in knowledge graph is still underutilized in the state-of-the-art methods for distantly supervised relation extraction. In the light of these challenges, we propose CORD, a novelCOopeRativeDenoising framework, which consists two base networks leveraging text corpus and knowledge graph respectively, and a cooperative module involving their mutual learning by the adaptive bi-directional knowledge distillation and dynamic ensemble with noisy-varying instances. Experimental results on a real-world dataset demonstrate that the proposed method reduces the noisy labels and achieves substantial improvement over the state-of-the-art methods.
Kai Lei, Daoyuan Chen, Yaliang Li, Nan Du 0001, Min Yang 0007, Wei Fan 0001, Ying Shen 0001
COLING4
2018 MuVAN: A Multi-view Attention Network for Multivariate Temporal Data
abstract
Recent advances in attention networks have gained enormous interest in time series data mining. Various attention mechanisms are proposed to soft-select relevant timestamps from temporal data by assigning learnable attention scores. However, many real-world tasks involve complex multivariate time series that continuously measure target from multiple views. Different views may provide information of different levels of quality varied over time, and thus should be assigned with different attention scores as well. Unfortunately, the existing attention-based architectures cannot be directly used to jointly learn the attention scores in both time and view domains, due to the data structure complexity. Towards this end, we propose a novel multi-view attention network, namely MuVAN, to learn fine-grained attentional representations from multivariate temporal data. MuVAN is a unified deep learning model that can jointly calculate the two-dimensional attention scores to estimate the quality of information contributed by each view within different timestamps. By constructing a hybrid focus procedure, we are able to bring more diversity to attention, in order to fully utilize the multi-view information. To evaluate the performance of our model, we carry out experiments on three real-world benchmark datasets. Experimental results show that the proposed MuVAN model outperforms the state-of-the-art deep representation approaches in different real-world tasks. Analytical results through a case study demonstrate that MuVAN can discover discriminative and meaningful attention scores across views over time, which improves the feature representation of multivariate temporal data.
Ye Yuan 0006, Guangxu Xun, Fenglong Ma, Yaqing Wang 0001, Nan Du 0001, Kebin Jia, Lu Su 0001, Aidong Zhang 0001
ICDM5
2018 On the Generative Discovery of Structured Medical Knowledge
abstract
Online healthcare services can provide the general public with ubiquitous access to medical knowledge and reduce medical information access cost for both individuals and societies. However, expanding the scale of high-quality yet structured medical knowledge usually comes with tedious efforts in data preparation and human annotation. To promote the benefits while minimizing the data requirement in expanding medical knowledge, we introduce a generative perspective to study the relational medical entity pair discovery problem. A generative model named Conditional Relationship Variational Autoencoder is proposed to discover meaningful and novel medical entity pairs by purely learning from the expression diversity in the existing relational medical entity pairs. Unlike discriminative approaches where high-quality contexts and candidate medical entity pairs are carefully prepared to be examined by the model, the proposed model generates novel entity pairs directly by sampling from a learned latent space without further data requirement. The proposed model explores the generative modeling capacity for medical entity pairs while incorporating deep learning for hands-free feature engineering. It is not only able to generate meaningful medical entity pairs that are not yet observed, but also can generate entity pairs for a specific medical relationship. The proposed model adjusts the initial representations of medical entities by addressing their relational commonalities. Quantitative and qualitative evaluations on real-world relational medical entity pairs demonstrate the effectiveness of the proposed method in generating relational medical entity pairs that are meaningful and novel.
Yaliang Li, Nan Du 0001, Wei Fan 0001, Philip S. Yu
KDD3
2018 Ontology Evaluation with Path-based Text-aware Entropy Computation
abstract
With the rising importance of knowledge exchange, ontologies have become a key technology in the development of shared knowledge models for semantic-driven applications, such as knowledge interchange and semantic integration. Significant progress has been made in the use of entropy to measure the predictability and redundancy of knowledge bases, particularly ontologies. However, the current entropy applications used to evaluate ontologies consider only single-point connectivity rather than path connectivity, assign equal weights to each entity and path, and assume that vertices are static. To address these deficiencies, the present study proposes a Path-based Text-aware Entropy Computation method, PTEC, by considering the path information between different vertices and the textual information within the path to calculate the connectivity path of the whole network and the different weights between various nodes. Information obtained from structure-based embedding and text-based embedding is multiplied by the connectivity matrix of the entropy computation. An experimental evaluation of three real-world ontologies is performed based on ontology statistical information (data quantity), entropy evaluation (data quality), and a case study (ontology structure and text visualization). These aspects mutually demonstrate the reliability of our method. Experimental results demonstrate that PTEC can effectively evaluate ontologies, particularly those in the medical field.
Ying Shen 0001, Daoyuan Chen, Min Yang 0007, Yaliang Li, Nan Du 0001, Kai Lei
SIGIR5
2018 Knowledge-aware Attentive Neural Network for Ranking Question Answer Pairs
abstract
Ranking question answer pairs has attracted increasing attention recently due to its broad applications such as information retrieval and question answering (QA). Significant progresses have been made by deep neural networks. However, background information and hidden relations beyond the context, which play crucial roles in human text comprehension, have received little attention in recent deep neural networks that achieve the state of the art in ranking QA pairs. In the paper, we propose KABLSTM, a Knowledge-aware Attentive Bidirectional Long Short-Term Memory, which leverages external knowledge from knowledge graphs (KG) to enrich the representational learning of QA sentences. Specifically, we develop a context-knowledge interactive learning architecture, in which a context-guided attentive convolutional neural network (CNN) is designed to integrate knowledge embeddings into sentence representations. Besides, a knowledge-aware attention mechanism is presented to attend interrelations between each segments of QA pairs. KABLSTM is evaluated on two widely-used benchmark QA datasets: WikiQA and TREC QA. Experiment results demonstrate that KABLSTM has robust superiority over competitors and sets state-of-the-art.
Ying Shen 0001, Yang Deng 0002, Min Yang 0007, Yaliang Li, Nan Du 0001, Wei Fan 0001, Kai Lei
SIGIR5
2017 Bringing semantic structures to user intent detection in online medical queries
abstract
The Internet has revolutionized healthcare by offering medical information ubiquitously to patients via the web search. The healthcare status, complex medical information needs of patients are expressed diversely and implicitly in their medical text queries. Aiming to better capture a focused picture of user's medical-related information search and shed insights on their healthcare information access strategies, it is challenging yet rewarding to detect structured user intentions from their diversely expressed medical text queries. We introduce a graph-based formulation to explore structured concept transitions for effective user intent detection in medical queries, where each node represents a medical concept mention and each directed edge indicates a medical concept transition. A deep model based on multi-task learning is introduced to extract structured semantic transitions from user queries, where the model extracts word-level medical concept mentions as well as sentence-level concept transitions collectively. A customized graph-based mutual transfer loss function is designed to impose explicit constraints and further exploit the contribution of mentioning a medical concept word to the implication of a semantic transition. We observe an 8% relative improvement in AUC and 23% relative reduction in coverage error by comparing the proposed model with the best baseline model for the concept transition inference task on real-world medical text queries.
Nan Du 0001, Wei Fan 0001, Yaliang Li, Chun-Ta Lu, Philip S. Yu
IEEE BigData2
2017 Reliable Medical Diagnosis from Crowdsourcing: Discover Trustworthy Answers from Non-Experts
abstract
Nowadays, increasingly more people are receiving medical diagnoses from healthcare-related question answering platforms as people can get diagnoses quickly and conveniently. However, such diagnoses from non-expert crowdsourcing users are noisy or even wrong due to the lack of medical domain knowledge, which can cause serious consequences. To unleash the power of crowdsourcing on healthcare question answering, it is important to identify trustworthy answers and filter out noisy ones from user-generated data. Truth discovery methods estimate user reliability degrees and infer trustworthy information simultaneously, and thus these methods can be adopted to discover trustworthy diagnoses from crowdsourced answers. However, existing truth discovery methods do not take into account the rich semantic meanings of the answers. In the light of this challenge, we propose a method to automatically capture the semantic meanings of answers, where answers are represented as real-valued vectors in the semantic space. To learn such vector representations from noisy user-generated data, we tightly combine the truth discovery and vector learning processes. In this way, the learned vector representations enable truth discovery method to model the semantic relations among answers, and the information trustworthiness inferred by truth discovery can help the procedure of vector representation learning. To demonstrate the effectiveness of the proposed method, we collect a large-scale real-world dataset that involves 219,527 medical diagnosis questions and 23,657 non-expert users. Experimental results show that the proposed method improves the accuracy of identified trustworthy answers due to the successful consideration of answers' semantic meanings. Further, we demonstrate the fast convergence and good scalability of the proposed method, which makes it practical for real-world applications.
Yaliang Li, Nan Du 0001, Chaochun Liu, Yusheng Xie, Wei Fan 0001, Qi Li 0012, Jing Gao 0004, Huan Sun 0001
WSDM2
2016 Influence based analysis of community consistency in dynamic networks
abstract
The development of Internet and social networks has provided more emerging network data which facilitates the dynamic network analysis. In this paper, we propose a new method to measure coherence strength, also referred to as community consistency, of a community under dynamic settings. In order to better interpret the influence of evolving community structure on community consistency, we model the problem as one of influence propagation processes having a causal relation with the community consistency. To this effect a generative model is proposed to combine the influence propagation and the network topological structure at each time stamp. Our comprehensive experiments on both synthetic and real-world datasets demonstrate the superiority of the proposed framework in estimating the community consistency.
Xiaowei Jia, Nan Du 0001, Yuan Zhang 0028, Vishrawas Gopalakrishnan, Guangxu Xun, Aidong Zhang 0001
ASONAM3
2016 Augmented LSTM Framework to Construct Medical Self-Diagnosis Android
abstract
Given a health-related question (such as "I have a bad stomach ache. What should I do?"), a medical self-diagnosis Android inquires further information from the user, diagnoses the disease, and ultimately recommend best solutions. One practical challenge to build such an Android is to ask correct questions and obtain most relevant information, in order to correctly pinpoint the most likely causes of health conditions. In this paper, we tackle this challenge, named "relevant symptom question generation": Given a limited set of patient described symptoms in the initial question (e.g., "stomach ache"), what are the most critical symptoms to further ask the patient, in order to correctly diagnose their potential problems? We propose an augmented long short-term memory (LSTM) framework, where the network architecture can naturally incorporate the inputs from embedding vectors of patient described symptoms and an initial disease hypothesis given by a predictive model. Then the proposed framework generates the most important symptom questions. The generation process essentially models the conditional probability to observe a new and undisclosed symptom, given a set of symptoms from a patient as well as an initial disease hypothesis. Experimental results show that the proposed model obtains improvements over alternative methods by over 30% (both precision and mean ordinal distance).
Chaochun Liu, Huan Sun 0001, Nan Du 0001, Shulong Tan, Hongliang Fei, Wei Fan 0001, Tao Yang 0012, Yaliang Li
ICDM3
2016 Mining User Intentions from Medical Queries: A Neural Network Based Heterogeneous Jointly Modeling Approach
abstract
Text queries are naturally encoded with user intentions. An intention detection task tries to model and discover intentions that user encoded in text queries. Unlike conventional text classification tasks where the label of text is highly correlated with some topic-specific words, words from different topic categories tend to co-occur in medical related queries. Besides the existence of topic-specific words and word order, word correlations and the way words organized into sentence are crucial to intention detection tasks.
Wei Fan 0001, Nan Du 0001, Philip S. Yu
WWW3
2015 Significant Edge Detection in Target Network by Exploring Multiple Auxiliary Networks
abstract
Despite the ability to model many real world settings as a network, one major challenge in analyzing network data is that important and reliable links between objects are usually obscured by noisy information and hence not readily discernible. In this paper, we propose to detect these important and reliable links - significant edges, from a target network by using multiple auxiliary networks and a limited amount of labelled information. In this process, we first abstract the community knowledge learnt across target and auxiliary networks to detect significant patterns. The mined community knowledge captures the key profile of network relationships and thus can be used to determine whether an existing edge indicates a true or false relationship. Experiments on real world network data show that our two staged solution -- a joint matrix factorisation procedure followed by edge significance score ranking, accurately predicts significant edges in target network by jointly exploring the underlying knowledge embedded in both target and auxiliary networks.
Nan Du 0001, Jing Gao 0004, Vishrawas Gopalakrishnan, Xiaowei Jia, Kang Li 0003, Aidong Zhang 0001
ASONAM1
2015 Functional Node Detection on Linked Data
abstract
Networks, which characterize object relationships, are ubiquitous in various domains. One very important problem is to detect the nodes of a specific function in these networks. For example, is a user normal or anomalous in an email network? Does a protein play a key role in a protein-protein interaction network? In many applications, the information we have about the networks usually includes both node characteristics and network structures. Both types of information can contribute to the task of learning functional nodes, and we call the collection of node and link information as linked data. However, existing methods only use a few subjectively selected topological features from network structures to detect functional nodes, thus fail to include highly discriminative and meaningful patterns hidden in linked data. To address this problem, a novel Feature Integration based Functional Node Detection (FIND) algorithm is presented. Specifically, FIND extracts the most discriminative information from both node characteristics and network structures in the form of a unified latent feature representation with the guidance of several labeled nodes. Experiments on two real world data sets validate that the proposed method significantly outperforms the baselines on the detection of three different types of functional nodes.
Kang Li 0003, Jing Gao 0004, Suxin Guo, Nan Du 0001, Aidong Zhang 0001
SDM4
2015 Identifying Affinity Classes of Inorganic Materials Binding Sequences via a Graph-Based Model
abstract
Rapid advances in bionanotechnology have recently generated growing interest in identifying peptides that bind to inorganic materials and classifying them based on their inorganic material affinities. However, there are some distinct characteristics of inorganic materials binding sequence data that limit the performance of many widely-used classification methods when applied to this problem. In this paper, we propose a novel framework to predict the affinity classes of peptide sequences with respect to an associated inorganic material. We first generate a large set of simulated peptide sequences based on an amino acid transition matrix tailored for the specific inorganic material. Then the probability of test sequences belonging to a specific affinity class is calculated by minimizing an objective function. In addition, the objective function is minimized through iterative propagation of probability estimates among sequences and sequence clusters. Results of computational experiments on two real inorganic material binding sequence data sets show that the proposed framework is highly effective for identifying the affinity classes of inorganic material binding sequences. Moreover, the experiments on the structural classification of proteins (SCOP) data set shows that the proposed framework is general and can be applied to traditional protein sequences.
Nan Du 0001, Marc R. Knecht, Mark T. Swihart, Zhenghua Tang, Tiffany R. Walsh, Aidong Zhang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 Tracking Temporal Community Strength in Dynamic Networks
abstract
Community formation analysis of dynamic networks has been a hot topic in data mining which has attracted much attention. Recently, there are many studies which focus on discovering communities successively from consecutive snapshots by considering both the current and historical information. However, these methods cannot provide us with much historical or successive information related to the detected communities. Different from previous studies which focus on community detection in dynamic networks, we define a new problem of tracking the progression of the community strength-a novel measure that reflects the community robustness and coherence throughout the entire observation period. To achieve this goal, we propose a novel framework which formulates the problem as an optimization task. The proposed community strength analysis also provides foundation for a wide variety of related applications such as discovering how the strength of each detected community changes over the entire observation period. To demonstrate that the proposed method provides precise and meaningful evolutionary patterns of communities which are not directly obtainable from traditional methods, we perform extensive experimental studies on one synthetic and five real datasets: Social evolution, tweeting interaction, actor relationships, bibliography, and biological datasets. Experimental results show that the proposed approach is highly effective in discovering the progression of community strengths and detecting interesting communities.
Nan Du 0001, Xiaowei Jia, Jing Gao 0004, Vishrawas Gopalakrishnan, Aidong Zhang 0001
IEEE Trans. Knowl. Data Eng.1
2014 Analysis on Community Variational Trend in Dynamic Networks
abstract
Temporal analysis on dynamic networks has become a popularly discussed topic today, with more and more emerging data over time. In this paper we investigate the problem of detecting and tracking the variational communities within a given time period. We first define a metric to measure the strength of a community, called the normalized temporal community strength. And then, we propose our analysis framework. The community may evolve over time, either split to multiple communities or merge with others. We address the problem of evolutionary clustering with requirement on temporal smoothness and propose a revised soft clustering method based on non-negative matrix factorization. Then we use a clustering matching method to find the soft correspondence between different community distribution structures. This matching establishes the connection between consecutive snapshots. To estimate the variational rate and meanwhile address the smoothness during continuous evolution, we propose an objective function that combines the conformity of current variation and historical variational trend. In addition, we integrate the weights to the objective function to identify the temporal outliers. An iterative coordinate descent method is proposed to solve the optimization framework. We extensively evaluate our method with a synthetic dataset and several real datasets. The experimental results demonstrate the effectiveness of our method, which is greatly superior to the baselines on detection of the communities with significant variation over time.
Xiaowei Jia, Nan Du 0001, Jing Gao 0004, Aidong Zhang 0001
CIKM2
2014 LRBM: A Restricted Boltzmann Machine Based Approach for Representation Learning on Linked Data
abstract
Linked data consist of both node attributes, e.g., Preferences, posts and degrees, and links which describe the connections between nodes. They have been widely used to represent various network systems, such as social networks, biological networks and etc. Knowledge discovery on linked data is of great importance to many real applications. One of the major challenges of learning linked data is how to effectively and efficiently extract useful information from both node attributes and links in linked data. Current studies on this topic either use selected topological statistics to represent network structures, or linearly map node attributes and network structures to a shared latent feature space. However, while approaches based on statistics may miss critical patterns in network structure, approaches based on linear mappings may not be sufficient to capture the non-linear characteristics of nodes and links. To handle the challenge, we propose, to our knowledge, the first deep learning method to learn from linked data. A restricted Boltzmann machine model named LRBM is developed for representation learning on linked data. In LRBM, we aim to extract the latent feature representation of each node from both node attributes and network structures, non-linearly map each pair of nodes to the links, and use hidden units to control the mapping. The details of how to adapt LRBM for link prediction and node classification on linked data have also been presented. In the experiments, we test the performance of LRBM as well as other baselines on link prediction and node classification. Overall, the extensive experimental evaluations confirm the effectiveness of the proposed LRBM model in mining linked data.
Kang Li 0003, Jing Gao 0004, Suxin Guo, Nan Du 0001, Aidong Zhang 0001
ICDM4
2014 A Deep Learning Approach to Link Prediction in Dynamic Networks
abstract
Time varying problems usually have complex underlying structures represented as dynamic networks where entities and relationships appear and disappear over time. The problem of efficiently performing dynamic link inference is extremely challenging due to the dynamic nature in massive evolving networks especially when there exist sparse connectivities and nonlinear transitional patterns. In this paper, we propose a novel deep learning framework, i.e., Conditional Temporal Restricted Boltzmann Machine (ctRBM), which predicts links based on individual transition variance as well as influence introduced by local neighbors. The proposed model is robust to noise and have the exponential capability to capture nonlinear variance. We tackle the computational challenges by developing an efficient algorithm for learning and inference of the proposed model. To improve the efficiency of the approach, we give a faster approximated implementation based on a proposed Neighbor Influence Clustering algorithm. Extensive experiments on simulated as well as real-world dynamic networks show that the proposed method outperforms existing algorithms in link inference on dynamic networks.
Nan Du 0001, Kang Li 0003, Jing Gao 0004, Aidong Zhang 0001
SDM2
2013 Detecting mutual functional gene clusters from multiple related diseases
abstract
Discovering functional gene clusters based on gene expression data has been a widely-used method that offers a tremendous opportunity for understanding the functional genomics of a specific disease. Due to its strong power of comprehending and interpreting mass of genes, plenty of studies have been done on detecting and analyzing the gene clusters for various diseases. However, more and more evidence suggest that human diseases are not isolated from each other. Therefore, it's significant and interesting to detect the common functional gene clusters driving the core mechanisms among multiple related diseases. There are mainly two challenges for this task: first, the gene expression from each disease may contain noise; second, the common factors underlying multiple diseases are hard to detect. To address these challenges, we propose a novel deep architecture to discover the mutual functional gene clusters across multiple types of diseases. To demonstrate that the proposed method can discover precise and meaningful gene clusters which are not directly obtainable from traditional methods, we perform extensive experimental studies on both synthetic and real datasets - public gene-expression data of three types of cancers. Experimental results show that the proposed approach is highly effective in discovering the mutual functional gene clusters.
Nan Du 0001, Yuan Zhang 0030, Aidong Zhang 0001
BIBM1
2013 Critical protein detection in dynamic PPI networks with multi-source integrated deep belief nets
abstract
Critical node detection in dynamic networks is of great value in many areas, such as the evolving of friendship in social networks, the development of epidemics, molecular pathogenesis of diseases and so on. As for detecting critical nodes in dynamic Protein-Protein Interaction Networks (PPINs), there are mainly two challenges: the first is to construct the dynamic PPINs that are not available directly from biological experiments in laboratories; and the second is how to identify the most critical units that are responsible for the dynamic processes. This paper proposes effective framework to tackle these two problems. First of all, this paper proposes to construct the dynamic PPINs by simultaneously modeling the activity of proteins and assembling the dynamic co-regulation protein network at each time point. As result, more comprehensive dynamic PPINs are built. Besides, a novel critical protein detection method that integrates multiple PPI networks into a Deep Belief Network model (referred to as MIDBN) is developed. The integrated model is trained to get hierarchical common representations of multiple sources which are used to reconstruct the original data. The variabilities of the reconstruction errors across the time courses are ranked to finally get the top proteins that have significantly different evolving structural patterns than the other nodes in the dynamic networks. We evaluated our network construction method by comparing the functional representations of the derived networks with that of two other traditional construction methods, and our method achieved superior function analysis results. The ranking results of critical proteins from MIDBN were compared with results from two baseline methods and the comparison results showed that MIDBN had better reconstruction rate and identified more proteins of critical value to yeast cell cycle process.
Yuan Zhang 0030, Nan Du 0001, Kang Li 0003, Jinchao Feng, Kebin Jia, Aidong Zhang 0001
BIBM2
2013 Progression Analysis of Community Strengths in Dynamic Networks
abstract
Community formation analysis of dynamic networks has been a hot topic in data mining which has attracted much attention. Recently, there are many studies which focus on discovering communities successively from each snapshot by considering both current and historical information. However, the detected communities are isolated at a certain snapshot, because these approaches ignore important historical or successive information. Different from previous studies which focus on community detection in dynamic networks, we define a new problem of tracking the progression of the community strength - a novel measure that reflects the community robustness and coherence throughout the entire observation period. The proposed community strength analysis provides significant insights into entity properties and relationships in a wide variety of applications. To tackle this problem, we propose a novel two-stage framework: we first identify communities via non-negative matrix factorization, and then calculate the strength of each detected community corresponding to each specific snapshot by solving an optimization problem. Experimental results show that the proposed approach is highly effective in discovering the progression of community strengths and detecting interesting communities.
Nan Du 0001, Jing Gao 0004, Aidong Zhang 0001
ICDM1
2013 Learning, Analyzing and Predicting Object Roles on Dynamic Networks
abstract
Dynamic networks are structures with objects and links between the objects that vary in time. Temporal information in dynamic networks can be used to reveal many important phenomena such as bursts of activities in social networks and human communication patterns in email networks. In this area, one very important problem is to understand dynamic patterns of object roles. For instance, will a user become a peripheral node in a social network? Could a website become a hub on the Internet? Will a gene be highly expressed in gene-gene interaction networks in the later stage of a cancer? In this paper, we propose a novel approach that identifies the role of each object, tracks the changes of object roles over time, and predicts the evolving patterns of the object roles in dynamic networks. In particular, a probability model is proposed to extract latent features of object roles from dynamic networks. The extracted latent features are discriminative in learning object roles and are capable of characterizing network structures. The probability model is then extended to learn the dynamic patterns and make predictions on object roles. We assess our method on two data sets on the tasks of exploring how users' importance and political interests evolve as time progresses on dynamic networks. Overall, the extensive experimental evaluations confirm the effectiveness of our approach for identifying, analyzing and predicting object roles on dynamic networks.
Kang Li 0003, Suxin Guo, Nan Du 0001, Jing Gao 0004, Aidong Zhang 0001
ICDM3
2012 De-noise biological network from heterogeneous sources via link propagation
abstract
Lots of recent bioinformatics works have focused on the inference of various types of biological networks, such as gene coexpression networks, protein-protein interaction networks, signal transduction networks, etc. Unfortunately, these raw biological network data often contain much noise, especially the false positive predictions which in many cases hinder accurate reconstruction of biological networks. In addition, since the labeled data is scarce and expensive, we hope that the knowledge from other domains can help handle this lack of labeled data problem. In order to construct a more robust and reliable biological network, we propose a novel link propagation based algorithm to de-noise false positives from the target biological network through propagating information from few labeled samples and a set of auxiliary domain networks. While comparing with many current state-of-the-art algorithms, our proposed approach has shown good performance in de-noising biological network.
Nan Du 0001, Jing Gao 0004, Vishrawas Gopalakrishnan, Aidong Zhang 0001
BIBM1
2012 A link prediction based unsupervised rank aggregation algorithm for informative gene selection
abstract
Informative Gene Selection is the process of identifying relevant genes that are significantly and differentially expressed in biological procedures. The microarray experiments conducted for this purpose usually implement only less than a hundred of samples to rank the relevance of over thousands of genes. Many irrelevant genes thus may gain statistical importance due to the randomness caused by the small sample problem, while relevant genes may lose focus in the same way. Overcoming such a problem goes beyond what a single microarray dataset can offer and stresses the use of multiple experiment results, which is defined as rank aggregation. In this paper, we propose a novel link prediction based rank aggregation algorithm for the purpose of informative gene selection. Each rank is transferred into a fully connected and weighted network, in which the nodes represent genes and the weights of links stand for priorities between connected nodes (genes). The integration of multiple gene ranks is then formulated as an optimization problem of link prediction on multiple networks, with criterion function favoring the maximization of weighted consensus among each network. We solve the problem through iterative estimation of weights and maximization of consensus among them. In the experimental evaluation, we demonstrate our method on the Prostate Cancer Dataset and compare it with other baseline methods. The results show that our link prediction based rank aggregation method remarkably outperforms all the compared methods, which proves the effectiveness of our framework in finding informative genes from multiple microarray experimental results.
Kang Li 0003, Nan Du 0001, Aidong Zhang 0001
BIBM2
2011 Finding Informative Genes from Multiple Microarray Experiments: A Graph-based Consensus Maximization Model
abstract
With the rapid advancement of biology technology, many microarray experiments are conducted towards the same problem of finding informative genes. Therefore, it is important to find a set of informative genes integrating multiple microarray experiments that achieves maximal consensus. Most previous re- searches formulated this problem as a rank aggregation problem. In this paper, we propose a novel Graph-based Consensus Maximization (GCM) model to estimate the conditional probability of each gene being informative, then the genes are ranked by this probability. The estimation of the probabilities is formulated as an optimization problem on a bipartite graph, where the criterion function favors the smoothness of the prediction over the graph and penalizes deviations from the initial input ranked lists from microarray experiments. We solve this problem through iterative propagation of probability estimates among neighboring nodes. In addition, when certain genes have already been identified to be informative, it has never been explored in the literature how to take advantage of such information to improve the consensus result. Our proposed GCM model can be naturally extended to incorporate such information, thus increasing the quality of the predicted result. In the experimental evaluation, we conducted experiments on the five prostate cancer microarray studies. The results showed that our model outperformed other baseline methods in finding informative genes. Furthermore, by adding only one piece of information that some gene is informative, our model yielded a significantly better result. The experimental evaluation demonstrates that the proposed GCM model is effective and superior in finding informative genes from multiple microarray experiments.
Nan Du 0001, Aidong Zhang 0001
BIBM2