VLDB 2026 Research / reviewers in the wild / expert
Jun Zhao 0001
dblp:47/2026-1
· DBLP profile ↗
24ranked-venue papers in the field
0as first author
9since 2021 · last 2026
0000-0003-3370-2263ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 19Database Systems & Data Management · 3Data Mining & Knowledge Discovery · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Legal-AP: A Framework to Enhance LLM's Legal Reasoning via Knowledge Augmentation and Adapter-Wise Parametric Fusion
Ao Chang, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
DASFAA (3) | 4 |
| 2026 | ASDE: Low-budget text classification via active semi-supervised learning with debiasing training mechanism
Yubo Chen 0001, Tong Zhou 0014, Daojian Zeng, Kang Liu 0001, Jun Zhao 0001 |
Inf. Process. Manag. | 6 |
| 2026 | Enhancing Event Causality Extraction With Mention-Level Causal Evidence and Global Causal Graph ReasoningabstractEvent Causality Extraction (ECE) aims to extract causal event pairs from text. Existing methods overlook the interplay between causal event pairs and their corresponding textual evidence (e.g., causal event mention pairs), and fail to effectively leverage global causal dependency information. To address these issues, we propose a Mention-Level Causal Evidence and Global Causal Graph Reasoning (MLCE-GCGR) framework to enhance ECE. First, we introduce an auxiliary Event Mention Causality Extraction (EMCE) task, which extracts causal event mention pairs, to provide evidence for the main ECE task, and design a Dual-Level Interaction Enhancement (DLIE) strategy to enhance the bidirectional interplay between event-level and mention-level causality. Second, we develop a Global Causal Graph Reasoning (GCGR) module that simulates human-like multi-turn reasoning, aiming to progressively refine the causal graph by capturing global dependencies among event mentions, types, and arguments. Experiments on four benchmark datasets show that our method outperforms state-of-the-art approaches. Moreover, by extracting causal event mention pairs as supporting evidence, our approach improves the interpretability of structured causality extraction. Ruili Pu, Yang Li 0074, Jun Zhao 0001, Suge Wang, Xiaoli Li 0001, Deyu Li 0001, Jian Liao 0005, Jianxing Zheng, Bin Liang 0004, Kam-Fai Wong |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Prompt robust large language model for Chinese medical named entity recognition
Yubo Chen 0001, Baoli Zhang, Zhuoran Jin, Zhengyuan Cai, Yingzheng Wang, Delai Qiu, Shengping Liu, Jun Zhao 0001 |
Inf. Process. Manag. | 9 |
| 2024 | Does Knowledge Localization Hold True? Surprising Differences Between Entity and Relation Perspectives in Language Models
Yifan Wei 0001, Yixuan Weng, Huanhuan Ma, Yuanzhe Zhang, Jun Zhao 0001, Kang Liu 0001 |
CIKM | 6 |
| 2024 | Information bottleneck based knowledge selection for commonsense reasoning
Zhao Yang 0004, Yuanzhe Zhang, Cao Liu, Jiansong Chen, Jun Zhao 0001, Kang Liu 0001 |
Inf. Sci. | 6 |
| 2021 | Uncertainty-Aware Self-Training for Semi-Supervised Event Temporal Relation ExtractionabstractExtracting event temporal relations is an important task for natural language understanding. Many works have been proposed for supervised event temporal relation extraction, which typically requires a large amount of human-annotated data for model training. However, the data annotation for this task is very time-consuming and challenging. To this end, we study the problem of semi-supervised event temporal relation extraction. Self-training as a widely used semi-supervised learning method can be utilized for this problem. However, it suffers from the noisy pseudo-labeling problem. In this paper, we propose the use of uncertainty-aware self-training framework (UAST) to quantify the model uncertainty for coping with pseudo-labeling errors. Specifically, UAST utilizes (1) Uncertainty Estimation module to compute the model uncertainty for pseudo-labeling unlabeled data; (2) Sample Selection with Exploration module to select informative samples based on uncertainty estimates; and (3) Uncertainty-Aware Learning module to explicitly incorporate the model uncertainty into the self-training process. Experimental results indicate that our approach significantly outperforms previous state-of-the-art methods. Xinyu Zuo, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Wei Bi |
CIKM | 5 |
| 2021 | Multi-Sentence Argument Linking via An Event-Aware Hierarchical EncoderabstractMulti-sentence argument linking aims at detecting implicit event arguments across sentences, which is indispensable when textual events span across multiple sentences in a document. Previous studies suffer from the inherent limitations of error propagation and lack the explicit modeling of the local and non-local interactions in a textual event. In this paper, we propose an event-aware hierarchical encoder for multi-sentence argument linking. Specifically, we introduce a hierarchical encoder to explicitly capture the local and global interactions in a textual event. Furthermore, we introduce an auxiliary task to predict the event-relevant context in a manner of multi-task learning, which can implicitly benefit the argument linking model to be aware of the event-relevant context. The empirical results on the widely used argument linking dataset show that our model significantly outperforms the baselines, which demonstrates the effectiveness of our proposed method. Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Taifeng Wang |
CIKM | 4 |
| 2021 | Multi-Task Self-Supervised Learning for Script Event PredictionabstractMost existing approaches to script event prediction rely on manually labeled data heavily, which is often expensive to obtain. To cope with the training data bottleneck, we investigate methods of combining multiple self-supervised tasks, i.e. tasks where models are explicitly trained with automatically generated labels. We propose two self-supervised pre-training tasks:one is End Identification and the other is Contrastive Scoring. Multi-task learning framework is then leveraged to combine these two tasks to jointly train the model. The pre-trained model is then fine-tuned using human-annotated script event prediction training data. Experimental results on the commonly used dataset show that our approach can achieve competitive performance compared to the previous models which are trained with the whole dataset by using just 10% of the training data, and our model trained on the whole dataset outperforms previous models significantly. Bo Zhou 0024, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Jiexin Xu, Xiaojian Jiang |
CIKM | 4 |
| 2019 | Document Gated Reader for Open-Domain Question AnsweringabstractOpen-domain question answering focuses on using diverse information resources to answer any types of question. Recent years, with the development of large-scale data set and various deep neural networks models, some recent advances in open domain question answering system first utilize the distantly supervised dataset as the knowledge resource, then apply deep learning based machine comprehension techniques to generate the right answers, which achieves impressive results compared with traditional feature-based pipeline methods. Bingning Wang, Jingfang Xu, Zhixing Tian, Kang Liu 0001, Jun Zhao 0001 |
SIGIR | 7 |
| 2018 | Deep Semantic Hashing with Multi-Adversarial TrainingabstractWith the amount of data has been rapidly growing over recent decades, binary hashing has become an attractive approach for fast search over large databases, in which the high-dimensional data such as image, video or text is mapped into a low-dimensional binary code. Searching in this hamming space is extremely efficient which is independent of the data size. A lot of methods have been proposed to learn this binary mapping. However, to make the binary codes conserves the input information, previous works mostly resort to mean squared error, which is prone to lose a lot of input information [11]. On the other hand, most of the previous works adopt the norm constraint or approximation on the hidden representation to make it as close as possible to binary, but the norm constraint is too strict that harms the expressiveness and flexibility of the code. Bingning Wang, Kang Liu 0001, Jun Zhao 0001 |
CIKM | 3 |
| 2015 | Learning to Represent Knowledge Graphs with Gaussian EmbeddingabstractThe representation of a knowledge graph (KG) in a latent space recently has attracted more and more attention. To this end, some proposed models (e.g., TransE) embed entities and relations of a KG into a "point" vector space by optimizing a global loss function which ensures the scores of positive triplets are higher than negative ones. We notice that these models always regard all entities and relations in a same manner and ignore their (un)certainties. In fact, different entities and relations may contain different certainties, which makes identical certainty insufficient for modeling. Therefore, this paper switches to density-based embedding and propose KG2E for explicitly modeling the certainty of entities and relations, which learn the representations of KGs in the space of multi-dimensional Gaussian distributions. Each entity/relation is represented by a Gaussian distribution, where the mean denotes its position and the covariance (currently with diagonal covariance) can properly represent its certainty. In addition, compared with the symmetric measures used in point-based methods, we employ the KL-divergence for scoring triplets, which is a natural asymmetry function for effectively modeling multiple types of relations. We have conducted extensive experiments on link prediction and triplet classification with multiple benchmark datasets (WordNet and Freebase). Our experimental results demonstrate that our method can effectively model the (un)certainties of entities and relations in a KG, and it significantly outperforms state-of-the-art methods (including TransH and TransR). Shizhu He, Kang Liu 0001, Guoliang Ji, Jun Zhao 0001 |
CIKM | 4 |
| 2015 | Large-scale Knowledge Base Completion: Inferring via Grounding Network Sampling over Selected InstancesabstractConstructing large-scale knowledge bases has attracted much attention in recent years, for which Knowledge Base Completion (KBC) is a key technique. In general, inferring new facts in a large-scale knowledge base is not a trivial task. The large number of inferred candidate facts has resulted in the failure of the majority of previous approaches. Inference approaches can achieve high precision for formulas that are accurate, but they are required to infer candidate instances one by one, and extremely large candidate sets bog them down in expensive calculations. In contrast, the existing embedding-based methods can easily calculate similarity-based scores for each candidate instance as opposed to using inference, so they can handle large-scale data. However, this type of method does not consider explicit logical semantics and usually has unsatisfactory precision. To resolve the limitations of the above two types of methods, we propose an approach through Inferring via Grounding Network Sampling over Selected Instances. We first employ an embedding-based model to make the instance selection and generate much smaller candidate sets for subsequent fact inference, which not only narrows the candidate sets but also filters out part of the noise instances. Then, we only make inferences within these candidate sets by running a data-driven inference algorithm on the Markov Logic Network (MLN), which is called Inferring via Grounding Network Sampling (INS). In this process, we especially incorporate the similarity priori generated by embedding-based models into INS to promote the inference precision. The experimental results show that our approach improved [email protected] from 32.911% to 71.692% on the FB15K dataset and achieved much better [email protected] evaluations than state-of-the-art methods. Zhuoyu Wei, Jun Zhao 0001, Kang Liu 0001, Zhenyu Qi 0003, Zhengya Sun, Guanhua Tian |
CIKM | 2 |
| 2015 | Co-Extracting Opinion Targets and Opinion Words from Online Reviews Based on the Word Alignment ModelabstractMining opinion targets and opinion words from online reviews are important tasks for fine-grained opinion mining, the key component of which involves detecting opinion relations among words. To this end, this paper proposes a novel approach based on the partially-supervised alignment model, which regards identifying opinion relations as an alignment process. Then, a graph-based co-ranking algorithm is exploited to estimate the confidence of each candidate. Finally, candidates with higher confidence are extracted as opinion targets or opinion words. Compared to previous methods based on the nearest-neighbor rules, our model captures opinion relations more precisely, especially for long-span relations. Compared to syntax-based methods, our word alignment model effectively alleviates the negative effects of parsing errors when dealing with informal online texts. In particular, compared to the traditional unsupervised alignment model, the proposed model obtains better precision because of the usage of partial supervision. In addition, when estimating candidate confidence, we penalize higher-degree vertices in our graph-based co-ranking algorithm to decrease the probability of error generation. Our experimental results on three corpora with different sizes and languages show that our approach effectively outperforms state-of-the-art methods. Kang Liu 0001, Liheng Xu, Jun Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | Towards faster and better retrieval models for question searchabstractCommunity question answering (cQA) has become an important service due to the popularity of cQA archives on the web. This paper is concerned with the problem of question search. Question search in cQA aims to find the historical questions that are semantically equivalent or similar to the queried questions. In this paper, we propose a faster and better retrieval model for question search by leveraging user chosen category. After introducing the question category, we can filter certain amount of irrelevant historical questions under a wide range of leaf categories. Experimental results conducted on real cQA data set demonstrate that the proposed techniques are more effective and efficient than a variety of baseline methods. Guangyou Zhou, Yubo Chen 0001, Daojian Zeng, Jun Zhao 0001 |
CIKM | 4 |
| 2012 | Exploring the existing category hierarchy to automatically label the newly-arising topics in cQAabstractThis work investigates selecting concise labels for the newly-arising topics in community question answer. Previous methods of generating labels do not take the information of the existing category hierarchy into consideration. The main motivation of our paper is to utilize this information into the label generation process. We propose a general framework to address this problem. Firstly, we map the questions into Wikipedia concept sets, which are more meaningful than terms. Secondly, important concepts are identified to represent the main focus of the newly-arising topics. Thirdly, candidate labels are extracted from Wikipedia category graph. Finally, candidate labels are filtered and reranked by combination of structure information of existing category hierarchy and Wikipedia category graph. The experiments show that in our test collections, about 80% "correct" labels appear in the top ten labels recommended by our system. Guangyou Zhou, Kang Liu 0001, Jun Zhao 0001 |
CIKM | 4 |
| 2012 | Topic-sensitive probabilistic model for expert finding in question answer communitiesabstractIn this paper, we address the problem of expert finding in community question answering (CQA). Most of the existing approaches attempt to find experts in CQA by means of link analysis techniques. However, these traditional techniques only consider the link structure while ignore the topical similarity among users (askers and answerers) and user expertise and user reputation. In this study, we propose a topic-sensitive probabilistic model, which is an extension of PageRank algorithm to find experts in CQA. Compared to the traditional link analysis techniques, our proposed method is more effective because it finds the experts by taking into account both the link structure and the topical similarity among users. We conduct experiments on real world data set from Yahoo! Answers. Experimental results show that our proposed method significantly outperforms the traditional link analysis techniques and achieves the state-of-the-art performance for expert finding in CQA. Guangyou Zhou, Siwei Lai, Kang Liu 0001, Jun Zhao 0001 |
CIKM | 4 |
| 2012 | Joint relevance and answer quality learning for question routing in community QAabstractCommunity question answering (cQA) has become a popular service for users to ask and answer questions. In recent years, the efficiency of cQA service is hindered by a sharp increase of questions in the community. This paper is concerned with the problem of question routing. Question routing in cQA aims to route new questions to the eligible answerers who can give high quality answers. However, the traditional methods suffer from the following two problems: (1) word mismatch between the new questions and the users' answering history; (2) high variance in perceived answer quality. To solve the above two problems, this paper proposes a novel joint learning method by taking both word mismatch and answer quality into a unified framework for question routing. We conduct experiments on large-scale real world data set from Yahoo! Answers. Experimental results show that our proposed method significantly outperforms the traditional query likelihood language model (QLLM) as well as state-of-the-art cluster-based language model (CBLM) and category-sensitive query likelihood language model (TCSLM). Guangyou Zhou, Kang Liu 0001, Jun Zhao 0001 |
CIKM | 3 |
| 2011 | Large-scale question classification in cQA by leveraging Wikipedia semantic knowledgeabstractWith the flourishing of community-based question answering (cQA) services like Yahoo! Answers, more and more web users seek their information need from these sites. Understanding user's information need expressed through their search questions is crucial to information providers. Question classification in cQA is studied for this purpose. However, there are two main difficulties in applying traditional methods (question classification in TREC QA and text classification) to cQA: (1) Traditional methods confine themselves to classify a text or question into two or a few predefined categories. While in cQA, the number of categories is much larger, such as Yahoo! Answers, there contains 1,263 categories. Our empirical results show that with the increasing of the number of categories to moderate size, the performance of the classification accuracy dramatically decreases. (2) Unlike the normal texts, questions in cQA are very short, which cannot provide sufficient word co-occurrence or shared information for a good similarity measure due to the data sparseness. In this paper, we propose a two-stage approach for question classification in cQA that can tackle the difficulties of the traditional methods. In the first stage, we preform a search process to prune the large-scale categories to focus our classification effort on a small subset. In the second stage, we enrich questions by leveraging Wikipedia semantic knowledge to tackle the data sparseness. As a result, the classification model is trained on the enriched small subset. We demonstrate the performance of our proposed method on Yahoo! Answers with 1,263 categories. The experimental results show that our proposed method significantly outperforms the baseline method (with error reductions of 23.21%). Guangyou Zhou, Kang Liu 0001, Jun Zhao 0001 |
CIKM | 4 |
| 2011 | Collective entity linking in web text: a graph-based methodabstractEntity Linking (EL) is the task of linking name mentions in Web text with their referent entities in a knowledge base. Traditional EL methods usually link name mentions in a document by assuming them to be independent. However, there is often additional interdependence between different EL decisions, i.e., the entities in the same document should be semantically related to each other. In these cases, Collective Entity Linking, in which the name mentions in the same document are linked jointly by exploiting the interdependence between them, can improve the entity linking accuracy. Xianpei Han, Le Sun 0001, Jun Zhao 0001 |
SIGIR | 3 |
| 2010 | Topic-driven web search result organization by leveraging wikipedia semantic knowledgeabstractEffective organization of web search results can greatly improve the utility of search engine and enhance the quality of search results. However, the organization of search results is difficult because the sub-topics of a query are usually not explicitly given. In this paper, we propose a novel topic-driven search result organization method, which can first detect the sub-topics of a query by finding the coherent Wikipedia concept groups from its search results; then organize these results using a topic-driven clustering algorithm; in the end we score and rank the topics using the support vector regression model. Empirical results show that our method can achieve competitive performance. Xianpei Han, Jun Zhao 0001 |
CIKM | 2 |
| 2009 | Named entity disambiguation by leveraging wikipedia semantic knowledgeabstractName ambiguity problem has raised an urgent demand for efficient, high-quality named entity disambiguation methods. The key problem of named entity disambiguation is to measure the similarity between occurrences of names. The traditional methods measure the similarity using the bag of words (BOW) model. The BOW, however, ignores all the semantic relations such as social relatedness between named entities, associative relatedness between concepts, polysemy and synonymy between key terms. So the BOW cannot reflect the actual similarity. Some research has investigated social networks as background knowledge for disambiguation. Social networks, however, can only capture the social relatedness between named entities, and often suffer the limited coverage problem. Xianpei Han, Jun Zhao 0001 |
CIKM | 2 |
| 2009 | Cross-domain sentiment classification using a two-stage methodabstractIn this paper, we give out a two-stage approach for domain adaptation problem in sentiment classification. In the first stage, based on our observation that customers often use different words to comment on the similar topics in the different domains, we regard these common topics as the bridge to link the different domain-specific features. We propose a novel topic model named Transfer-PLSA to extract the topic knowledge between different domains. Through these common topics, the features in the source domain are corresponded to the target features, so that those domain-specific knowledge can be transferred across different domains. In the second step, we use the classifier trained on the labeled examples in the source domain to pick up some informative examples in the target domain. Then we retrain the classifier on these selected examples, so that the classifier is adapted for the target domain. Experimental results on sentiment classification in four different domains indicate that our method outperforms other traditional methods. Kang Liu 0001, Jun Zhao 0001 |
CIKM | 2 |
| 2007 | Probabilistic Models for Action-Based Chinese Dependency Parsing
Xiangyu Duan, Jun Zhao 0001, Bo Xu 0002 |
ECML | 2 |