VLDB 2026 Research / reviewers in the wild / expert
Kun Zhang 0041
dblp:96/3115-41
· DBLP profile ↗
9ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0002-7984-1082ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FactVerse: A Benchmark for Factual Consistency in Interleaved Image-Text GenerationabstractYubo Shan, Kun Zhang, Qiming Xu, Liping Cao, Yingying Cao, Jian Zhang, Yu Wang, Jingyuan Li, Yuanzhuo Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yubo Shan, Kun Zhang 0041, Liping Cao, Yingying Cao, Jingyuan Li 0002, Yuanzhuo Wang |
ACL (1) | 2 |
| 2025 | GenR1-Searcher: Curriculum Reinforcement Learning for Dynamic Retrieval and Document Generation
Renrui Duan, Jingyuan Li 0002, Yuanzhuo Wang, Kun Zhang 0041 |
CIKM | 6 |
| 2025 | Retriever-generator-verification: A novel approach to enhancing factual coherence in open-domain question answering
Shiqi Sun 0003, Kun Zhang 0041, Jingyuan Li 0002, Min Yu 0001, Kun Hou, Yuanzhuo Wang, Xueqi Cheng 0001 |
Inf. Process. Manag. | 2 |
| 2024 | Tree-of-Reasoning Question Decomposition for Complex Question Answering with Large Language ModelsabstractLarge language models (LLMs) have recently demonstrated remarkable performance across various Natual Language Processing tasks. In the field of multi-hop reasoning, the Chain-of-thought (CoT) prompt method has emerged as a paradigm, using curated stepwise reasoning demonstrations to enhance LLM's ability to reason and produce coherent rational pathways. To ensure the accuracy, reliability, and traceability of the generated answers, many studies have incorporated information retrieval (IR) to provide LLMs with external knowledge. However, existing CoT with IR methods decomposes questions into sub-questions based on a single compositionality type, which limits their effectiveness for questions involving multiple compositionality types. Additionally, these methods suffer from inefficient retrieval, as complex questions often contain abundant information, leading to the retrieval of irrelevant information inconsistent with the query's intent. In this work, we propose a novel question decomposition framework called TRQA for multi-hop question answering, which addresses these limitations. Our framework introduces a reasoning tree (RT) to represent the structure of complex questions. It consists of four components: the Reasoning Tree Constructor (RTC), the Question Generator (QG), the Retrieval and LLM Interaction Module (RAIL), and the Answer Aggregation Module (AAM). Specifically, the RTC predicts diverse sub-question structures to construct the reasoning tree, allowing a more comprehensive representation of complex questions. The QG generates sub-questions for leaf-node in the reasoning tree, and we explore two methods for QG: prompt-based and T5-based approaches. The IR module retrieves documents aligned with sub-questions, while the LLM formulates answers based on the retrieved information. Finally, the AAM aggregates answers along the reason tree, producing a definitive response from bottom to top. Kun Zhang 0041, Jiali Zeng, Fandong Meng, Yuanzhuo Wang, Shiqi Sun 0003, Long Bai 0002, Huawei Shen, Jie Zhou 0016 |
AAAI | 1 |
| 2024 | A novel feature integration method for named entity recognition model in product titlesabstractAbstract Entity recognition of product titles is essential for retrieving and recommending product information. Due to the irregularity of product title text, such as informal sentence structure, a large number of professional attribute words, a large number of unrelated independent entities of various combinations, the existing general named entity recognition model is limited in the e‐commerce field of product title entity recognition. Most of the current studies focus on only one of the two challenges instead of considering the two challenges together. Our approach proposes NEZHA‐CNN‐GlobalPointer architecture with the addition of label semantic network, and uses multigranularity contextual and label semantic information to fully capture the internal structure and category information of words and texts to improve the entity recognition accuracy. Through a series of experiments, we proved the efficiency of our approach over a dataset of Chinese product titles from JD.com, improving the F1‐value by 5.98%, when compared to the BERT‐LSTM‐CRF model on the product title corpus. Shiqi Sun 0003, Jingyuan Li 0002, Kun Zhang 0041, Xinghang Sun, Jianhe Cen, Yuanzhuo Wang |
Comput. Intell. | 3 |
| 2023 | CFGL-LCR: A Counterfactual Graph Learning Framework for Legal Case RetrievalabstractLegal case retrieval, which aims to find relevant cases based on a short case description, serves as an important part of modern legal systems. Despite the success of existing retrieval methods based on Pretrained Language Models, there are still two issues in legal case retrieval that have not been well considered before. First, existing methods underestimate the semantics associations among legal elements, e.g., law articles and crimes, which played an essential role in legal case retrieval. These methods only adopt the pre-training language model to encode the whole legal case, instead of distinguishing different legal elements in the legal case. They randomly split a legal case into different segments, which may break the completeness of each legal element. Second, due to the difficulty in annotating the relevant labels of similar cases, legal case retrieval inevitably faces the problem of lacking training data. In this paper, we propose a counterfactual graph learning framework for legal case retrieval. Concretely, to overcome the above challenges, we transform the legal case document into a graph and model the semantics of the legal elements through a graph neural network. To alleviate the low resource and learn the causal relationship between the semantics of legal elements and relevance, a counterfactual data generator is designed to augment counterfactual data and enhance legal case representation. Extensive experiments based on two publicly available legal benchmarks demonstrate that our CFGL-LCR can significantly outperform previous state-of-the-art methods in legal case retrieval. Kun Zhang 0041, Chong Chen 0001, Yuanzhuo Wang, Qi Tian 0001, Long Bai 0002 |
KDD | 1 |
| 2022 | Meta-CQG: A Meta-Learning Framework for Complex Question Generation over Knowledge BasesabstractComplex question generation over knowledge bases (KB) aims to generate natural language questions involving multiple KB relations or functional constraints. Existing methods train one encoder-decoder-based model to fit all questions. However, such a one-size-fits-all strategy may not perform well since complex questions exhibit an uneven distribution in many dimensions, such as question types, involved KB relations, and query structures, resulting in insufficient learning for long-tailed samples under different dimensions. To address this problem, we propose a meta-learning framework for complex question generation. The meta-trained generator can acquire universal and transferable meta-knowledge and quickly adapt to long-tailed samples through a few most related training samples. To retrieve similar samples for each input query, we design a self-supervised graph retriever to learn distributed representations for samples, and contrastive learning is leveraged to improve the learned representations. We conduct experiments on both WebQuestionsSP and ComplexWebQuestion, and results on long-tailed samples of different dimensions have been significantly improved, which demonstrates the effectiveness of the proposed framework. Kun Zhang 0041, Yunqi Qiu, Yuanzhuo Wang, Long Bai 0002, Wei Li 0176, Xuhui Jiang, Huawei Shen, Xueqi Cheng 0001 |
COLING | 1 |
| 2020 | Hierarchical Query Graph Generation for Complex Question Answering over Knowledge GraphabstractKnowledge Graph Question Answering aims to automatically answer natural language questions via well-structured relation information between entities stored in knowledge graphs. When faced with a complex question with compositional semantics, query graph generation is a practical semantic parsing-based method. But existing works rely on heuristic rules with limited coverage, making them impractical on more complex questions. This paper proposes a Director-Actor-Critic framework to overcome these challenges. Through options over a Markov Decision Process, query graph generation is formulated as a hierarchical decision problem. The Director determines which types of triples the query graph needs, the Actor generates corresponding triples by choosing nodes and edges, and the Critic calculates the semantic similarity between the generated triples and the given questions. Moreover, to train from weak supervision, we base the framework on hierarchical Reinforcement Learning with intrinsic motivation. To accelerate the training process, we pre-train the Critic with high-reward trajectories generated by hand-crafted rules, and leverage curriculum learning to gradually increase the complexity of questions during query graph generation. Extensive experiments conducted over widely-used benchmark datasets demonstrate the effectiveness of the proposed framework. Yunqi Qiu, Kun Zhang 0041, Yuanzhuo Wang, Xiaolong Jin 0001, Long Bai 0002, Saiping Guan, Xueqi Cheng 0001 |
CIKM | 2 |
| 2020 | Stepwise Reasoning for Multi-Relation Question Answering over Knowledge Graph with Weak SupervisionabstractKnowledge Graph Question Answering aims to automatically answer natural language questions via well-structured relation information between entities stored in knowledge graphs. When faced with a multi-relation question, existing embedding-based approaches take the whole topic-entity-centric subgraph into account, resulting in high time complexity. Meanwhile, due to the high cost for data annotations, it is impractical to exactly show how to answer a complex question step by step, and only the final answer is labeled, as weak supervision. To address these challenges, this paper proposes a neural method based on reinforcement learning, namely Stepwise Reasoning Network, which formulates multi-relation question answering as a sequential decision problem. The proposed model performs effective path search over the knowledge graph to obtain the answer, and leverages beam search to reduce the number of candidates significantly. Meanwhile, based on the attention mechanism and neural networks, the policy network can enhance the unique impact of different parts of a given question over triple selection. Moreover, to alleviate the delayed and sparse reward problem caused by weak supervision, we propose a potential-based reward shaping strategy, which can accelerate the convergence of the training algorithm and help the model perform better. Extensive experiments conducted over three benchmark datasets well demonstrate the effectiveness of the proposed model, which outperforms the state-of-the-art approaches. Yunqi Qiu, Yuanzhuo Wang, Xiaolong Jin 0001, Kun Zhang 0041 |
WSDM | 4 |