EDBT 2026 Demo / reviewers in the wild / expert
Kang Liu 0001
dblp:42/4903
· DBLP profile ↗
18ranked-venue papers in the field
2as first author
8since 2021 · last 2026
0000-0002-6083-8433ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 15 (1 first)Database Systems & Data Management · 2 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Legal-AP: A Framework to Enhance LLM's Legal Reasoning via Knowledge Augmentation and Adapter-Wise Parametric Fusion
Ao Chang, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
DASFAA (3) | 3 |
| 2026 | ASDE: Low-budget text classification via active semi-supervised learning with debiasing training mechanism
Yubo Chen 0001, Tong Zhou 0014, Daojian Zeng, Kang Liu 0001, Jun Zhao 0001 |
Inf. Process. Manag. | 5 |
| 2024 | Does Knowledge Localization Hold True? Surprising Differences Between Entity and Relation Perspectives in Language Models
Yifan Wei 0001, Yixuan Weng, Huanhuan Ma, Yuanzhe Zhang, Jun Zhao 0001, Kang Liu 0001 |
CIKM | 7 |
| 2024 | Information bottleneck based knowledge selection for commonsense reasoning
Zhao Yang 0004, Yuanzhe Zhang, Cao Liu, Jiansong Chen, Jun Zhao 0001, Kang Liu 0001 |
Inf. Sci. | 7 |
| 2022 | PEMP: Leveraging Physics Properties to Enhance Molecular Property PredictionabstractMolecular property prediction is essential for drug discovery. In recent years, deep learning methods have been introduced to this area and achieved state-of-the-art performances. However, most of existing methods ignore the intrinsic relations between molecular properties which can be utilized to improve the performances of corresponding prediction tasks. In this paper, we propose a new approach, namely Physics properties Enhanced Molecular Property prediction (PEMP), to utilize relations between molecular properties revealed by previous physics theory and physical chemistry studies. Specifically, we enhance the training of the chemical and physiological property predictors with related physics property prediction tasks. We design two different methods for PEMP, respectively based on multi-task learning and transfer learning. Both methods include a model-agnostic molecule representation module and a property prediction module. In our implementation, we adopt both the state-of-the-art molecule embedding models under the supervised learning paradigm and the pretraining paradigm as the molecule representation module of PEMP, respectively. Experimental results on public benchmark MoleculeNet show that the proposed methods have the ability to outperform corresponding state-of-the-art models. Yuancheng Sun, Weizhi Ma, Wenhao Huang 0001, Kang Liu 0001, Zhiming Ma, Wei-Ying Ma, Yanyan Lan |
CIKM | 5 |
| 2021 | Uncertainty-Aware Self-Training for Semi-Supervised Event Temporal Relation ExtractionabstractExtracting event temporal relations is an important task for natural language understanding. Many works have been proposed for supervised event temporal relation extraction, which typically requires a large amount of human-annotated data for model training. However, the data annotation for this task is very time-consuming and challenging. To this end, we study the problem of semi-supervised event temporal relation extraction. Self-training as a widely used semi-supervised learning method can be utilized for this problem. However, it suffers from the noisy pseudo-labeling problem. In this paper, we propose the use of uncertainty-aware self-training framework (UAST) to quantify the model uncertainty for coping with pseudo-labeling errors. Specifically, UAST utilizes (1) Uncertainty Estimation module to compute the model uncertainty for pseudo-labeling unlabeled data; (2) Sample Selection with Exploration module to select informative samples based on uncertainty estimates; and (3) Uncertainty-Aware Learning module to explicitly incorporate the model uncertainty into the self-training process. Experimental results indicate that our approach significantly outperforms previous state-of-the-art methods. Xinyu Zuo, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Wei Bi |
CIKM | 4 |
| 2021 | Multi-Sentence Argument Linking via An Event-Aware Hierarchical EncoderabstractMulti-sentence argument linking aims at detecting implicit event arguments across sentences, which is indispensable when textual events span across multiple sentences in a document. Previous studies suffer from the inherent limitations of error propagation and lack the explicit modeling of the local and non-local interactions in a textual event. In this paper, we propose an event-aware hierarchical encoder for multi-sentence argument linking. Specifically, we introduce a hierarchical encoder to explicitly capture the local and global interactions in a textual event. Furthermore, we introduce an auxiliary task to predict the event-relevant context in a manner of multi-task learning, which can implicitly benefit the argument linking model to be aware of the event-relevant context. The empirical results on the widely used argument linking dataset show that our model significantly outperforms the baselines, which demonstrates the effectiveness of our proposed method. Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Taifeng Wang |
CIKM | 3 |
| 2021 | Multi-Task Self-Supervised Learning for Script Event PredictionabstractMost existing approaches to script event prediction rely on manually labeled data heavily, which is often expensive to obtain. To cope with the training data bottleneck, we investigate methods of combining multiple self-supervised tasks, i.e. tasks where models are explicitly trained with automatically generated labels. We propose two self-supervised pre-training tasks:one is End Identification and the other is Contrastive Scoring. Multi-task learning framework is then leveraged to combine these two tasks to jointly train the model. The pre-trained model is then fine-tuned using human-annotated script event prediction training data. Experimental results on the commonly used dataset show that our approach can achieve competitive performance compared to the previous models which are trained with the whole dataset by using just 10% of the training data, and our model trained on the whole dataset outperforms previous models significantly. Bo Zhou 0024, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Jiexin Xu, Xiaojian Jiang |
CIKM | 3 |
| 2019 | Document Gated Reader for Open-Domain Question AnsweringabstractOpen-domain question answering focuses on using diverse information resources to answer any types of question. Recent years, with the development of large-scale data set and various deep neural networks models, some recent advances in open domain question answering system first utilize the distantly supervised dataset as the knowledge resource, then apply deep learning based machine comprehension techniques to generate the right answers, which achieves impressive results compared with traditional feature-based pipeline methods. Bingning Wang, Jingfang Xu, Zhixing Tian, Kang Liu 0001, Jun Zhao 0001 |
SIGIR | 6 |
| 2018 | Deep Semantic Hashing with Multi-Adversarial TrainingabstractWith the amount of data has been rapidly growing over recent decades, binary hashing has become an attractive approach for fast search over large databases, in which the high-dimensional data such as image, video or text is mapped into a low-dimensional binary code. Searching in this hamming space is extremely efficient which is independent of the data size. A lot of methods have been proposed to learn this binary mapping. However, to make the binary codes conserves the input information, previous works mostly resort to mean squared error, which is prone to lose a lot of input information [11]. On the other hand, most of the previous works adopt the norm constraint or approximation on the hidden representation to make it as close as possible to binary, but the norm constraint is too strict that harms the expressiveness and flexibility of the code. Bingning Wang, Kang Liu 0001, Jun Zhao 0001 |
CIKM | 2 |
| 2015 | Learning to Represent Knowledge Graphs with Gaussian EmbeddingabstractThe representation of a knowledge graph (KG) in a latent space recently has attracted more and more attention. To this end, some proposed models (e.g., TransE) embed entities and relations of a KG into a "point" vector space by optimizing a global loss function which ensures the scores of positive triplets are higher than negative ones. We notice that these models always regard all entities and relations in a same manner and ignore their (un)certainties. In fact, different entities and relations may contain different certainties, which makes identical certainty insufficient for modeling. Therefore, this paper switches to density-based embedding and propose KG2E for explicitly modeling the certainty of entities and relations, which learn the representations of KGs in the space of multi-dimensional Gaussian distributions. Each entity/relation is represented by a Gaussian distribution, where the mean denotes its position and the covariance (currently with diagonal covariance) can properly represent its certainty. In addition, compared with the symmetric measures used in point-based methods, we employ the KL-divergence for scoring triplets, which is a natural asymmetry function for effectively modeling multiple types of relations. We have conducted extensive experiments on link prediction and triplet classification with multiple benchmark datasets (WordNet and Freebase). Our experimental results demonstrate that our method can effectively model the (un)certainties of entities and relations in a KG, and it significantly outperforms state-of-the-art methods (including TransH and TransR). Shizhu He, Kang Liu 0001, Guoliang Ji, Jun Zhao 0001 |
CIKM | 2 |
| 2015 | Large-scale Knowledge Base Completion: Inferring via Grounding Network Sampling over Selected InstancesabstractConstructing large-scale knowledge bases has attracted much attention in recent years, for which Knowledge Base Completion (KBC) is a key technique. In general, inferring new facts in a large-scale knowledge base is not a trivial task. The large number of inferred candidate facts has resulted in the failure of the majority of previous approaches. Inference approaches can achieve high precision for formulas that are accurate, but they are required to infer candidate instances one by one, and extremely large candidate sets bog them down in expensive calculations. In contrast, the existing embedding-based methods can easily calculate similarity-based scores for each candidate instance as opposed to using inference, so they can handle large-scale data. However, this type of method does not consider explicit logical semantics and usually has unsatisfactory precision. To resolve the limitations of the above two types of methods, we propose an approach through Inferring via Grounding Network Sampling over Selected Instances. We first employ an embedding-based model to make the instance selection and generate much smaller candidate sets for subsequent fact inference, which not only narrows the candidate sets but also filters out part of the noise instances. Then, we only make inferences within these candidate sets by running a data-driven inference algorithm on the Markov Logic Network (MLN), which is called Inferring via Grounding Network Sampling (INS). In this process, we especially incorporate the similarity priori generated by embedding-based models into INS to promote the inference precision. The experimental results show that our approach improved [email protected] from 32.911% to 71.692% on the FB15K dataset and achieved much better [email protected] evaluations than state-of-the-art methods. Zhuoyu Wei, Jun Zhao 0001, Kang Liu 0001, Zhenyu Qi 0003, Zhengya Sun, Guanhua Tian |
CIKM | 3 |
| 2015 | Co-Extracting Opinion Targets and Opinion Words from Online Reviews Based on the Word Alignment ModelabstractMining opinion targets and opinion words from online reviews are important tasks for fine-grained opinion mining, the key component of which involves detecting opinion relations among words. To this end, this paper proposes a novel approach based on the partially-supervised alignment model, which regards identifying opinion relations as an alignment process. Then, a graph-based co-ranking algorithm is exploited to estimate the confidence of each candidate. Finally, candidates with higher confidence are extracted as opinion targets or opinion words. Compared to previous methods based on the nearest-neighbor rules, our model captures opinion relations more precisely, especially for long-span relations. Compared to syntax-based methods, our word alignment model effectively alleviates the negative effects of parsing errors when dealing with informal online texts. In particular, compared to the traditional unsupervised alignment model, the proposed model obtains better precision because of the usage of partial supervision. In addition, when estimating candidate confidence, we penalize higher-degree vertices in our graph-based co-ranking algorithm to decrease the probability of error generation. Our experimental results on three corpora with different sizes and languages show that our approach effectively outperforms state-of-the-art methods. Kang Liu 0001, Liheng Xu, Jun Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Exploring the existing category hierarchy to automatically label the newly-arising topics in cQAabstractThis work investigates selecting concise labels for the newly-arising topics in community question answer. Previous methods of generating labels do not take the information of the existing category hierarchy into consideration. The main motivation of our paper is to utilize this information into the label generation process. We propose a general framework to address this problem. Firstly, we map the questions into Wikipedia concept sets, which are more meaningful than terms. Secondly, important concepts are identified to represent the main focus of the newly-arising topics. Thirdly, candidate labels are extracted from Wikipedia category graph. Finally, candidate labels are filtered and reranked by combination of structure information of existing category hierarchy and Wikipedia category graph. The experiments show that in our test collections, about 80% "correct" labels appear in the top ten labels recommended by our system. Guangyou Zhou, Kang Liu 0001, Jun Zhao 0001 |
CIKM | 3 |
| 2012 | Topic-sensitive probabilistic model for expert finding in question answer communitiesabstractIn this paper, we address the problem of expert finding in community question answering (CQA). Most of the existing approaches attempt to find experts in CQA by means of link analysis techniques. However, these traditional techniques only consider the link structure while ignore the topical similarity among users (askers and answerers) and user expertise and user reputation. In this study, we propose a topic-sensitive probabilistic model, which is an extension of PageRank algorithm to find experts in CQA. Compared to the traditional link analysis techniques, our proposed method is more effective because it finds the experts by taking into account both the link structure and the topical similarity among users. We conduct experiments on real world data set from Yahoo! Answers. Experimental results show that our proposed method significantly outperforms the traditional link analysis techniques and achieves the state-of-the-art performance for expert finding in CQA. Guangyou Zhou, Siwei Lai, Kang Liu 0001, Jun Zhao 0001 |
CIKM | 3 |
| 2012 | Joint relevance and answer quality learning for question routing in community QAabstractCommunity question answering (cQA) has become a popular service for users to ask and answer questions. In recent years, the efficiency of cQA service is hindered by a sharp increase of questions in the community. This paper is concerned with the problem of question routing. Question routing in cQA aims to route new questions to the eligible answerers who can give high quality answers. However, the traditional methods suffer from the following two problems: (1) word mismatch between the new questions and the users' answering history; (2) high variance in perceived answer quality. To solve the above two problems, this paper proposes a novel joint learning method by taking both word mismatch and answer quality into a unified framework for question routing. We conduct experiments on large-scale real world data set from Yahoo! Answers. Experimental results show that our proposed method significantly outperforms the traditional query likelihood language model (QLLM) as well as state-of-the-art cluster-based language model (CBLM) and category-sensitive query likelihood language model (TCSLM). Guangyou Zhou, Kang Liu 0001, Jun Zhao 0001 |
CIKM | 2 |
| 2011 | Large-scale question classification in cQA by leveraging Wikipedia semantic knowledgeabstractWith the flourishing of community-based question answering (cQA) services like Yahoo! Answers, more and more web users seek their information need from these sites. Understanding user's information need expressed through their search questions is crucial to information providers. Question classification in cQA is studied for this purpose. However, there are two main difficulties in applying traditional methods (question classification in TREC QA and text classification) to cQA: (1) Traditional methods confine themselves to classify a text or question into two or a few predefined categories. While in cQA, the number of categories is much larger, such as Yahoo! Answers, there contains 1,263 categories. Our empirical results show that with the increasing of the number of categories to moderate size, the performance of the classification accuracy dramatically decreases. (2) Unlike the normal texts, questions in cQA are very short, which cannot provide sufficient word co-occurrence or shared information for a good similarity measure due to the data sparseness. In this paper, we propose a two-stage approach for question classification in cQA that can tackle the difficulties of the traditional methods. In the first stage, we preform a search process to prune the large-scale categories to focus our classification effort on a small subset. In the second stage, we enrich questions by leveraging Wikipedia semantic knowledge to tackle the data sparseness. As a result, the classification model is trained on the enriched small subset. We demonstrate the performance of our proposed method on Yahoo! Answers with 1,263 categories. The experimental results show that our proposed method significantly outperforms the baseline method (with error reductions of 23.21%). Guangyou Zhou, Kang Liu 0001, Jun Zhao 0001 |
CIKM | 3 |
| 2009 | Cross-domain sentiment classification using a two-stage methodabstractIn this paper, we give out a two-stage approach for domain adaptation problem in sentiment classification. In the first stage, based on our observation that customers often use different words to comment on the similar topics in the different domains, we regard these common topics as the bridge to link the different domain-specific features. We propose a novel topic model named Transfer-PLSA to extract the topic knowledge between different domains. Through these common topics, the features in the source domain are corresponded to the target features, so that those domain-specific knowledge can be transferred across different domains. In the second step, we use the classifier trained on the labeled examples in the source domain to pick up some informative examples in the target domain. Then we retrain the classifier on these selected examples, so that the classifier is adapted for the target domain. Experimental results on sentiment classification in four different domains indicate that our method outperforms other traditional methods. Kang Liu 0001, Jun Zhao 0001 |
CIKM | 1 |