EDBT 2026 Demo / reviewers in the wild / expert
Luchen Liu
dblp:217/2275
· DBLP profile ↗
25ranked-venue papers
7as first author
18since 2021 · last 2025
0009-0009-3817-6267ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient EvaluationabstractJinsheng Huang, Liang Chen, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan, Haozhe Zhao, Zhihui Guo, Yichi Zhang, Jingyang Yuan, Wei Ju, Luchen Liu, Tianyu Liu, Baobao Chang, Ming Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jinsheng Huang, Liang Chen 0024, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan 0016, Haozhe Zhao, Zhihui Guo, Yichi Zhang 0010, Jingyang Yuan, Wei Ju 0001, Luchen Liu, Tianyu Liu 0001, Baobao Chang, Ming Zhang 0004 |
NAACL (Long Papers) | 13 |
| 2024 | Self-supervised Graph-level Representation Learning with Adversarial Contrastive LearningabstractThe recently developed unsupervised graph representation learning approaches apply contrastive learning into graph-structured data and achieve promising performance. However, these methods mainly focus on graph augmentation for positive samples, while the negative mining strategies for graph contrastive learning are less explored, leading to sub-optimal performance. To tackle this issue, we propose a Graph Adversarial Contrastive Learning (GraphACL) scheme that learns a bank of negative samples for effective self-supervised whole-graph representation learning. Our GraphACL consists of (i) a graph encoding branch that generates the representations of positive samples and (ii) an adversarial generation branch that produces a bank of negative samples. To generate more powerful hard negative samples, our method minimizes the contrastive loss during encoding updating while maximizing the contrastive loss adversarially over the negative samples for providing the challenging contrastive task. Moreover, the quality of representations produced by the adversarial generation branch is enhanced through the regularization of carefully designed bank divergence loss and bank orthogonality loss. We optimize the parameters of the graph encoding branch and adversarial generation branch alternately. Extensive experiments on 14 real-world benchmarks on both graph classification and transfer learning tasks demonstrate the effectiveness of the proposed approach over existing graph self-supervised representation learning methods. Xiao Luo 0001, Wei Ju 0001, Yiyang Gu, Zhengyang Mao, Luchen Liu, Yuhui Yuan, Ming Zhang 0004 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2024 | Criterion-based Heterogeneous Collaborative Filtering for Multi-behavior Implicit RecommendationabstractRecent years have witnessed the explosive growth of interaction behaviors in multimedia information systems, where multi-behavior recommender systems have received increasing attention by leveraging data from various auxiliary behaviors such as tip and collect. Among various multi-behavior recommendation methods, non-sampling methods have shown superiority over negative sampling methods. However, two observations are usually ignored in existing state-of-the-art non-sampling methods based on binary regression: (1) users have different preference strengths for different items, so they cannot be measured simply by binary implicit data; (2) the dependency across multiple behaviors varies for different users and items. To tackle the above issue, we propose a novel non-sampling learning framework namedCriterion-guidedHeterogeneousCollaborativeFiltering (CHCF). CHCF introduces both upper and lower thresholds to indicate selection criteria, which will guide user preference learning. Besides, CHCF integrates criterion learning and user preference learning into a unified framework, which can be trained jointly for the interaction prediction of the target behavior. We further theoretically demonstrate that the optimization of Collaborative Metric Learning can be approximately achieved by the CHCF learning framework in a non-sampling form effectively. Extensive experiments on three real-world datasets show the effectiveness of CHCF in heterogeneous scenarios. Xiao Luo 0001, Daqing Wu, Yiyang Gu, Chong Chen 0002, Luchen Liu, Jinwen Ma, Ming Zhang 0004, Minghua Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2024 | Structure-Aware Prototypical Neural Process for Few-Shot Graph ClassificationabstractGraph classification plays an important role in a wide range of applications from biological prediction to social analysis. Traditional graph classification models built on graph kernels are hampered by the challenge of poor generalization as they are heavily dependent on the dedicated design of handcrafted features. Recently, graph neural networks (GNNs) become a new class of tools for analyzing graph data and have achieved promising performance. However, it is necessary to collect a large number of labeled graph data for training an accurate GNN, which is often unaffordable in real-world applications. Therefore, it is an open question to build GNNs under the condition of few-shot learning where only a few labeled graphs are available. In this article, we introduce a new Structure-aware Prototypical Neural Process (SPNP for short) for a few-shot graph classification. Specifically, at the encoding stage, SPNP first employs GNNs to capture graph structure information. Then, SPNP incorporates such structural priors into the latent path and the deterministic path for representing stochastic processes. At the decoding stage, SPNP uses a new prototypical decoder to define a metric space where unseen graphs can be predicted effectively. The proposed decoder, which contains a self-attention mechanism to learn the intraclass dependence between graphs, can enhance the class-level representations, especially for new classes. Furthermore, benefited from such a flexible encoding-decoding architecture, SPNP can directly map the context samples to a predictive distribution without any complicated operations used in previous methods. Extensive experiments demonstrate that SPNP achieves consistent and significant improvements over state-of-the-art methods. Further discussions are provided toward model efficiency and more detailed analysis. Xixun Lin, Zhao Li 0007, Peng Zhang 0001, Luchen Liu, Chuan Zhou 0001, Bin Wang 0004, Zhihong Tian 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Redundancy-Free Self-Supervised Relational Learning for Graph ClusteringabstractGraph clustering, which learns the node representations for effective cluster assignments, is a fundamental yet challenging task in data analysis and has received considerable attention accompanied by graph neural networks (GNNs) in recent years. However, most existing methods overlook the inherent relational information among the nonindependent and nonidentically distributed nodes in a graph. Due to the lack of exploration of relational attributes, the semantic information of the graph-structured data fails to be fully exploited which leads to poor clustering performance. In this article, we propose a novel self-supervised deep graph clustering method named relational redundancy-free graph clustering (R2FGC) to tackle the problem. It extracts the attribute- and structure-level relational information from both global and local views based on an autoencoder (AE) and a graph AE (GAE). To obtain effective representations of the semantic information, we preserve the consistent relationship among augmented nodes, whereas the redundant relationship is further reduced for learning discriminative embeddings. In addition, a simple yet valid strategy is used to alleviate the oversmoothing issue. Extensive experiments are performed on widely used benchmark datasets to validate the superiority of our R2FGC over state-of-the-art baselines. Our codes are available at https://github.com/yisiyu95/R2FGC. Siyu Yi, Wei Ju 0001, Yifang Qin, Xiao Luo 0001, Luchen Liu, Ming Zhang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Toward Effective Semi-supervised Node Classification with Hybrid Curriculum Pseudo-labelingabstractSemi-supervised node classification is a crucial challenge in relational data mining and has attracted increasing interest in research on graph neural networks (GNNs). However, previous approaches merely utilize labeled nodes to supervise the overall optimization, but fail to sufficiently explore the information of their underlying label distribution. Even worse, they often overlook the robustness of models, which may cause instability of network outputs to random perturbations. To address the aforementioned shortcomings, we develop a novel framework termed Hybrid Curriculum Pseudo-Labeling (HCPL) for efficient semi-supervised node classification. Technically, HCPL iteratively annotates unlabeled nodes by training a GNN model on the labeled samples and any previously pseudo-labeled samples, and repeatedly conducts this process. To improve the model robustness, we introduce a hybrid pseudo-labeling strategy that incorporates both prediction confidence and uncertainty under random perturbations, therefore mitigating the influence of erroneous pseudo-labels. Finally, we leverage the idea of curriculum learning to start from annotating easy samples, and gradually explore hard samples as the iteration grows. Extensive experiments on a number of benchmarks demonstrate that our HCPL beats various state-of-the-art baselines in diverse settings. Xiao Luo 0001, Wei Ju 0001, Yiyang Gu, Yifang Qin, Siyu Yi, Daqing Wu, Luchen Liu, Ming Zhang 0004 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | Deep spatial-temporal structure learning for rumor detection on Twitter
Chuan Zhou 0001, Jia Wu 0001, Luchen Liu, Bin Wang 0004 |
Neural Comput. Appl. | 4 |
| 2023 | Towards Long-Tailed Recognition for Graph Classification via Collaborative ExpertsabstractGraph classification, aiming at learning the graph-level representations for effective class assignments, has received outstanding achievements, which heavily relies on high-quality datasets that have balanced class distribution. In fact, most real-world graph data naturally presents a long-tailed form, where the head classes occupy much more samples than the tail classes, it thus is essential to study the graph-level classification over long-tailed data while still remaining largely unexplored. However, most existing long-tailed learning methods in visions fail to jointly optimize the representation learning and classifier training, as well as neglect the mining of the hard-to-classify classes. Directly applying existing methods to graphs may lead to sub-optimal performance, since the model trained on graphs would be more sensitive to the long-tailed distribution due to the complex topological characteristics. Hence, in this paper, we propose a novel long-tailed graph-level classification framework viaCollaborativeMulti-expert Learning (CoMe) to tackle the problem. To equilibrate the contributions of head and tail classes, we first develop balanced contrastive learning from the view of representation learning, and then design an individual-expert classifier training based on hard class mining. In addition, we execute gated fusion and disentangled knowledge distillation among the multiple experts to promote the collaboration in a multi-expert framework. Comprehensive experiments are performed on seven widely-used benchmark datasets to demonstrate the superiority of our method CoMe over state-of-the-art baselines. Siyu Yi, Zhengyang Mao, Wei Ju 0001, Luchen Liu, Xiao Luo 0001, Ming Zhang 0004 |
IEEE Trans. Big Data | 5 |
| 2022 | Learning Common Dependency Structure for Unsupervised Cross-Domain NerabstractUnsupervised cross-domain NER task aims to solve the issues when data in a new domain are fully-unlabeled. It leverages labeled data from source domain to predict entities in unlabeled target domain. Since training models on large domain corpus is time-consuming, in this paper, we consider an alternative way by introducing syntactic dependency structure. Such information is more accessible and can be shared between sentences from different domains. We propose a novel framework with dependency-aware GNN (DGNN) to learn these common structures from source domain and adapt them to target domain, alleviating the data scarcity issue and bridging the domain gap. Experimental results show that our method outperforms state-of-the-art methods. Luchen Liu, Xixun Lin, Peng Zhang 0001, Lei Zhang 0119, Bin Wang 0004 |
ICASSP | 1 |
| 2022 | Retrieval Bias Aware Ensemble Model for Conditional Sentence GenerationabstractConditional sentence generation aims to generate proper target sentences with the given condition, and has shown great promise in many text generation applications such as dialogue systems and poetry generation. The ensemble of retrieval and generation-based models retrieve texts according to the input condition to assist the generation-based model. Those approaches obtain great performance tasks as they can absorb both merits to generate informative and coherent sentences. However, the input condition and its retrieved results are usually not highly consistent due to the quality of retrieval. It leads to a retrieval bias between the condition and its retrieved result, and then text generation augmented by such results becomes unreliable. To fix this issue, we propose RBAEM, a Retrieval Bias Aware Ensemble Model. RBAEM employs two CVAEs (Conditional variational Auto-encoder) to represent the retrieved target and the ground truth with latent vectors, and then diminishes the bias by decreasing the distance of two corresponding distributions. The extensive experiments on two tasks show that the proposed methods excel the existing state-of-the-art generation models. Yiping Song, Luchen Liu, Ming Zhang 0004, Zhiliang Tian |
ICASSP | 4 |
| 2022 | Multi-Modal Adversarial Adaptive Network for Misinformation Detection on Social MediaabstractMost existing methods for misinformation detection are mainly based on multi-modal contents, while extract multi-modal features with simple encoders and coarse interactions. What's more, they ignore the complex local structure underlying the event distributions when faced with emerging events. To tackle these issues, we propose a Multi-modal Adversarial Adaptive Network (MAAN) for misinformation detection, extracting fine-grained event-invariant cross-modal features. Specifically, MAAN consists of two components: 1) a multi-modal feature extractor models dense inter-modal interactions, removes noise information, and captures discriminative multi-modal information; 2) an adaptive misinformation detector adaptively evaluates the global and local event distribution discrepancies, and dynamically weights them when disentangling event-specific features from multi-modal representations. Extensive experiments demonstrate that MAAN can outperforms the state-of-the-art approaches on two public datasets. Luchen Liu |
ICME | 4 |
| 2022 | Building Conversational Diagnosis Systems for Fine-Grained Diseases Using Few Annotated Data
Yiping Song, Wei Ju 0001, Zhiliang Tian, Luchen Liu, Ming Zhang 0004 |
ICONIP (3) | 4 |
| 2021 | Improving Cross-Domain Slot Filling with Common Syntactic StructureabstractCross-domain slot filling is a challenging task in spoken language understanding due to the differences in text genre across domains. In this paper, we attempt to solve this task by exploiting the syntactic structures of user utterances, because these syntactic structures are actually accessible and can be shared between utterances from different domains. To this end, we propose a novel Syntactic Structure Encoder (SSE) module and incorporate it into a detection-prediction framework. SSE introduces graph convolutional network (GCN) to learn the common structures from multiple source domains, which are helpful to better adaptation on the target domain. Experimental results conducted on SNIPS dataset show that our model significantly outperforms the state-of-the-art approach in cross-domain slot filling. Specifically, our model outperforms the best model by ∼4% and ∼5% F1-scores under the 20-example and 50-example settings, respectively. Luchen Liu, Xixun Lin, Peng Zhang 0001, Bin Wang 0004 |
ICASSP | 1 |
| 2021 | Multiphish: Multi-Modal Features Fusion Networks for Phishing DetectionabstractPhishing is an increasingly serious cybercrime. Phishers create phishing websites by mimicking legitimate websites to confuse users and steal their personal information. The proliferation of phishing websites and more advanced camouflage techniques are problems faced by most existing methods. In this paper, we propose a features fusion networks (MultiPhish) which is the first study on fusing multi-modal features with neural networks for the phishing detection task. In this end-to-end network, the domain and favicon of the website are represented via deep neural networks, and the representation of the website identity is obtained through multi-modal features fusion. In addition, the variation autoencoder (VAE) is introduced to optimize the representation. In the phishing detection module, we incorporate URL features to improve situations where phishing websites cannot be detected only by estimating whether the website identity is disguised. Based on the latest collected dataset, we have carried out extensive experiments and proved that our model is superior to the relevant methods. In addition, MultiPhish is a completely language-independent strategy, so it can perform phishing detection regardless of the text language. Luchen Liu, Jianlong Tan |
ICASSP | 3 |
| 2021 | HIP Network: Historical Information Passing Network for Extrapolation Reasoning on Temporal Knowledge GraphabstractIn recent years, temporal knowledge graph (TKG) reasoning has received significant attention. Most existing methods assume that all timestamps and corresponding graphs are available during training, which makes it difficult to predict future events. To address this issue, recent works learn to infer future events based on historical information. However, these methods do not comprehensively consider the latent patterns behind temporal changes, to pass historical information selectively, update representations appropriately and predict events accurately. In this paper, we propose the Historical Information Passing (HIP) network to predict future events. HIP network passes information from temporal, structural and repetitive perspectives, which are used to model the temporal evolution of events, the interactions of events at the same time step, and the known events respectively. In particular, our method considers the updating of relation representations and adopts three scoring functions corresponding to the above dimensions. Experimental results on five benchmark datasets show the superiority of HIP network, and the significant improvements on Hits@1 prove that our method can more accurately predict what is going to happen. Yongquan He, Peng Zhang 0001, Luchen Liu, Qi Liang 0002, Wenyuan Zhang 0002 |
IJCAI | 3 |
| 2021 | DCEN: A Decoupled Context Enhanced Network For Few-shot Slot TaggingabstractFew-shot slot tagging is an important task in developing dialogue system. Most previous few-shot slot tagging models classify an item according to its similarity to the representation of each class. These models leverage context information implicitly through each words' contextual embedding. However, the entangled language features of words may interfere with context information, misleading the utilization of crucial slot features in few-shot scenario. To tackle these problems, we propose the Decoupled Context Enhanced Network (DCEN) for few-shot slot tagging. Different from previous models, we extract decoupled context explicitly to make full use of slot features contained in the context. Decoupled context includes two parts, local and global decoupled context information. We introduce a local extractor to extract local decoupled context by integrating information from adjacent words, and a global extractor based on transformer to extract global decoupled information by orthogonalization. Experimental results on SNIPS show that our model achieves the state-of-the-art performance with considerable improvements. Youliang Yuan, Jiaxin Pan 0002, Luchen Liu, Min Peng 0002 |
IJCNN | 4 |
| 2021 | Bipartite Graph Embedding via Mutual Information MaximizationabstractBipartite graph embedding has recently attracted much attention due to the fact that bipartite graphs are widely used in various application domains. Most previous methods, which adopt random walk-based or reconstruction-based objectives, are typically effective to learn local graph structures. However, the global properties of bipartite graph, including community structures of homogeneous nodes and long-range dependencies of heterogeneous nodes, are not well preserved. In this paper, we propose a bipartite graph embedding called BiGI to capture such global properties by introducing a novel local-global infomax objective. Specifically, BiGI first generates a global representation which is composed of two prototype representations. BiGI then encodes sampled edges as local representations via the proposed subgraph-level attention mechanism. Through maximizing the mutual information between local and global representations, BiGI enables nodes in bipartite graph to be globally relevant. Our model is evaluated on various benchmark datasets for the tasks of top-K recommendation and link prediction. Extensive experiments demonstrate that BiGI achieves consistent and significant improvements over state-of-the-art baselines. Detailed analyses verify the high effectiveness of modeling the global properties of bipartite graph. Jiangxia Cao, Xixun Lin, Luchen Liu, Tingwen Liu, Bin Wang 0004 |
WSDM | 4 |
| 2021 | Exploiting Subspace Relation in Semantic Labels for Cross-Modal HashingabstractHashing methods have been extensively applied to efficient multimedia data indexing and retrieval on account of the explosion of multimedia data. Cross-modal hashing usually learns binary codes by mapping multi-modal data into a common Hamming space. Most supervised methods utilize relation information like class labels as pairwise similarities of cross-modal data pair to narrow intra-modal and inter-modal gap. In this paper, we propose a novel supervised cross-modal hashing method dubbed Subspace Relation Learning for Cross-modal Hashing (SRLCH), which exploits relation information of labels in semantic space to make similar data from different modalities closer in the low-dimension Hamming subspace. SRLCH preserves the modality relationships, the discrete constraints and nonlinear structures, while admitting a closed-form binary codes solution, which effectively enhances the training efficiency. An iterative alternative optimization algorithm is developed to simultaneously learn both hash functions and unified binary codes. With these binary codes and hash functions, we can index multimedia data and search them in an efficient way. Evaluations in two cross-modal retrieval tasks on several widely-used datasets show that the proposed SRLCH outperforms most cross-modal hashing methods. Theoretical analysis also illustrates reasons for our method’s promotion in subspace relation learning. Heng Tao Shen, Luchen Liu, Yang Yang 0002, Xing Xu 0001, Zi Huang, Fumin Shen, Richang Hong |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Multi-task Learning via Adaptation to Similar Tasks for Mortality Prediction of Diverse Rare Diseases
Luchen Liu, Zequn Liu, Haoxian Wu, Zichang Wang, Jianhao Shen, Yiping Song, Ming Zhang 0004 |
AMIA | 1 |
| 2019 | Learning Hierarchical Representations of Electronic Health Records for Clinical Outcome Prediction
Luchen Liu, Haoran Li 0003, Zhiting Hu, Haoran Shi 0001, Zichang Wang, Jian Tang 0005, Ming Zhang 0004 |
AMIA | 1 |
| 2019 | Predictive Multi-level Patient Representations from Electronic Health RecordsabstractThe advent of the Internet era has led to an explosive growth in the Electronic Health Records (EHR) in the past decades. The EHR data can be regarded as a collection of clinical events, including laboratory results, medication records, physiological indicators, etc, which can be used for clinical outcome prediction tasks to support constructions of intelligent health systems. Learning patient representation from these clinical events for the clinical outcome prediction is an important but challenging step. Most related studies transform EHR data of a patient into a sequence of clinical events in temporal order and then use sequential models to learn patient representations for outcome prediction. However, clinical event sequence contains thousands of event types and temporal dependencies. We further make an observation that clinical events occurring in a short period are not constrained by any temporal order but events in a long term are influenced by temporal dependencies. The multi-scale temporal property makes it difficult for traditional sequential models to capture the short-term co-occurrence and the long-term temporal dependencies in clinical event sequences. In response to the above challenges, this paper proposes a Multilevel Representation Model (MRM). MRM first uses a sparse attention mechanism to model the short-term co-occurrence, then uses interval-based event pooling to remove redundant information and reduce sequence length and finally predicts clinical outcomes through Long Short-Term Memory (LSTM). Experiments on real-world datasets indicate that our proposed model largely improves the performance of clinical outcome prediction tasks using EHR data. Zichang Wang, Haoran Li 0003, Luchen Liu, Haoxian Wu, Ming Zhang 0004 |
BIBM | 3 |
| 2019 | Decoupling Category-wise Independence and Relevance with Self-attention for Multi-label Image ClassificationabstractMulti-label image classification has achieved remarkable progress thanks to deep convolutional neural networks (CNNs). In this paper, we propose a Decouple Network (DecoupleNet) which is an end-to-end CNN-based framework able to trade off class-level feature independence and relevance during training. The proposed DecoupleNet is able to decouple category-wise independence and relevance with image-level supervision. We design a category-wise space-to-depth module with a spatial pooling strategy to exploit more meaningful convolutional features. They are integrated with class-wise correlated information which is automatically learned via a new self-attention mechanism. We conduct extensive experiments on two large-scale benchmarks: the MS-COCO and the NUS-WIDE, where the proposed DecoupleNet obtains impressive performance compared favorably against the state-of-the-art methods on multi-label image classification. Luchen Liu, Sheng Guo 0005, Matthew R. Scott |
ICASSP | 1 |
| 2018 | Learning the Joint Representation of Heterogeneous Temporal Events for Clinical Endpoint PredictionabstractThe availability of a large amount of electronic health records (EHR) provides huge opportunities to improve health care service by mining these data. One important application is clinical endpoint prediction, which aims to predict whether a disease, a symptom or an abnormal lab test will happen in the future according to patients' history records. This paper develops deep learning techniques for clinical endpoint prediction, which are effective in many practical applications. However, the problem is very challenging since patients' history records contain multiple heterogeneous temporal events such as lab tests, diagnosis, and drug administrations. The visiting patterns of different types of events vary significantly, and there exist complex nonlinear relationships between different events. In this paper, we propose a novel model for learning the joint representation of heterogeneous temporal events. The model adds a new gate to control the visiting rates of different events which effectively models the irregular patterns of different events and their nonlinear correlations. Experiment results with real-world clinical data on the tasks of predicting death and abnormal lab tests prove the effectiveness of our proposed approach over competitive baselines. Luchen Liu, Jianhao Shen, Ming Zhang 0004, Zichang Wang, Jian Tang 0005 |
AAAI | 1 |
| 2018 | Index and Retrieve Multimedia Data: Cross-Modal Hashing by Learning Subspace Relation
Luchen Liu, Yang Yang 0002, Mengqiu Hu, Xing Xu 0001, Fumin Shen, Ning Xie 0003, Zi Huang |
DASFAA (2) | 1 |
| 2017 | Topic medical concept embedding: Multi-sense representation learning for medical conceptabstractRepresentation learning algorithm in medical area maps high dimensional real world medical concepts to low dimensional vector space, encodes rich medical knowledge, and has brought improvement to various machine learning applications in medical area. However, previous representation learning models in medical area failed to consider the multi-sense characteristic of medical concept. Moreover, the inner relationships between representations learned by previous model is implicit and can only be explained according to visualization, which means poor interpretability. In this paper, we propose Topic Medical Concept Embedding (TMCE), a generative embedding model to address above two problems. TMCE is able to learn multi-sense representations for a single medical concept, and TMCE can also improve interpretability by modeling relationships between each concept explicitly. In TMCE, multi-sense concept representations are influenced by its contexts and its topics. In addition, dosage information which is ignored by previous work are also utilized in TMCE. A MCMC method is presented to jointly learn the two-layer topic embeddings and multi-sense concept embeddings. Experimental results show that representations learned by TMCE outperforms those learned by other strong baselines by a large margin in a multi-label diagnose classification tasks. Several case studies further show that TMCE can learn medically correct multi-sense representations with better interpretability than other strong baselines. Chengyue Gong, Luchen Liu, Lei Sha, Ming Zhang 0004 |
BIBM | 3 |