VLDB 2026 Research / reviewers in the wild / expert
Jun Zhao 0019
dblp:47/2026-19
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data ContaminationabstractReasoning in large language models has long been a central research focus, and recent studies employing reinforcement learning (RL) have introduced diverse methods that yield substantial performance gains with minimal or even no external supervision. Surprisingly, some studies even suggest that random or incorrect reward signals can enhance performance. However, these breakthroughs are predominantly observed for the mathematically strong Qwen2.5 series on benchmarks such as MATH-500, AMC, and AIME, and seldom transfer to models like Llama, which warrants a more in-depth investigation. In this work, our empirical analysis reveals that pre-training on massive web-scale corpora leaves Qwen2.5 susceptible to data contamination in widely used benchmarks. Consequently, conclusions derived from contaminated benchmarks on Qwen2.5 series may be unreliable. To obtain trustworthy evaluation results, we introduce a generator that creates fully clean arithmetic problems of arbitrary length and difficulty, dubbed RandomCalculation. Using this leakage-free dataset, we show that only accurate reward signals yield steady improvements that surpass the base model’s performance boundary in mathematical reasoning, whereas random or incorrect rewards do not. Moreover, we conduct more fine-grained analyses to elucidate the factors underlying the different performance observed on the MATH-500 and RandomCalculation benchmarks. Consequently, we recommend that future studies evaluate models on uncontaminated benchmarks and, when feasible, test various model series to ensure trustworthy conclusions about RL and related methods. Mingqi Wu, Zhihao Zhang 0002, Qiaole Dong, Zhiheng Xi, Jun Zhao 0019, Senjie Jin, Xiaoran Fan, Yuhao Zhou 0005, Huijie Lv, Ming Zhang 0030, Yanwei Fu 0001, Qin Liu 0010, Songyang Zhang 0001, Qi Zhang 0001 |
AAAI | 5 |
| 2026 | Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative AlignmentabstractYuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuming Yang 0001, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao 0019, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 8 |
| 2025 | Understanding Parametric and Contextual Knowledge Reconciliation within Large Language ModelsabstractRetrieval-Augmented Generation (RAG) provides additional contextual knowledge to complement the parametric knowledge in Large Language Models (LLMs). These two knowledge interweave to enhance the accuracy and timeliness of LLM responses. However,
the internal mechanisms by which LLMs utilize these knowledge remain unclear. We propose modeling the forward propagation of knowledge as an entity flow, employing this framework to trace LLMs' internal behaviors when processing mixed-source knowledge. Linear probing utilizes a trainable linear classifier to detect specific attributes in hidden layers. However, once trained, a probe cannot adapt to dynamically specified entities. To address this challenge, we construct an entity-aware probe, which introduces special tokens to mark probing targets and employs a small trainable rank-8 lora update to process these special markers. We first verify this approach through an attribution experiment, demonstrating that it can accurately detect information about ad-hoc entities from complex hidden states. Next, we trace entity flows across layers to understand how LLMs reconcile conflicting knowledge internally. Our probing results reveal that contextual and parametric knowledge are routed between tokens through distinct sets of attention heads, supporting attention competition only within knowledge types. While conflicting knowledge maintains a residual presence across layers, aligned knowledge from multiple sources gradually accumulates, with the magnitude of this accumulation directly determining its influence on final outputs. Jun Zhao 0019, Yongzhuo Yang, Jingqi Tong, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
NeurIPS | 1 |
| 2024 | Unveiling Linguistic Regions in Large Language ModelsabstractLarge Language Models (LLMs) have demonstrated considerable cross-lingual alignment and generalization ability.Current research primarily focuses on improving LLMs' crosslingual generalization capabilities.However, there is still a lack of research on the intrinsic mechanisms of how LLMs achieve crosslingual alignment.From the perspective of region partitioning, this paper conducts several investigations on the linguistic competence of LLMs.We discover a core region in LLMs that corresponds to linguistic competence, accounting for approximately 1% of the total model parameters.Removing this core region by setting parameters to zero results in a significant performance decrease across 30 different languages.Furthermore, this core region exhibits significant dimensional dependence, perturbations to even a single parameter on specific dimensions leading to a loss of linguistic competence.Moreover, we discover that distinct monolingual regions exist for different languages, and disruption to these specific regions substantially reduces the LLMs' proficiency in those corresponding languages.Our research also indicates that freezing the core linguistic region during further pre-training can mitigate the issue of catastrophic forgetting (CF), a common phenomenon observed during further pre-training of LLMs.Overall, exploring the LLMs' functional regions provides insights into the foundation of their intelligence 1 . * Equal contributions.† Corresponding authors. 1 Our code is released in https://github.com/ zzhang0179/Unveiling-Linguistic-Regions-in-LLMs. Zhihao Zhang 0002, Jun Zhao 0019, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
ACL (1) | 2 |
| 2024 | LONGAGENT: Achieving Question Answering for 128k-Token-Long Documents through Multi-Agent CollaborationabstractLarge language models (LLMs) have achieved tremendous success in understanding language and processing text.However, questionanswering (QA) on lengthy documents faces challenges of resource constraints and a high propensity for errors, even for the most advanced models such as GPT-4 and Claude2.In this paper, we introduce LONGAGENT, a multi-agent collaboration method that enables efficient and effective QA over 128k-tokenlong documents.LONGAGENT adopts a divideand-conquer strategy, breaking down lengthy documents into shorter, more manageable text chunks.A leader agent comprehends the user's query and organizes the member agents to read their assigned chunks, reasoning a final answer through multiple rounds of discussion.Due to members' hallucinations, it's difficult to guarantee that every response provided by each member is accurate.To address this, we develop an inter-member communication mechanism that facilitates information sharing, allowing for the detection and mitigation of hallucinatory responses.Experimental results show that a LLaMA-2 7B driven by LONGAGENT can effectively support QA over 128k-token documents, achieving 16.42% and 1.63% accuracy gains over GPT-4 on singlehop and multi-hop QA settings, respectively. Jun Zhao 0019, Can Zu, Xu Hao, Wei He 0024, Yiwen Ding, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 1 |
| 2024 | TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer CapabilitiesabstractMing Zhang, Caishuang Huang, Yilong Wu, Shichun Liu, Huiyuan Zheng, Yurui Dong, Yujiong Shen, Shihan Dou, Jun Zhao, Junjie Ye, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ming Zhang 0030, Caishuang Huang, Yilong Wu, Shichun Liu, Huiyuan Zheng, Yurui Dong 0001, Yujiong Shen, Shihan Dou, Jun Zhao 0019, Junjie Ye 0005, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
EMNLP | 9 |
| 2024 | Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning Through Trap ProblemsabstractHuman cognition exhibits systematic compositionality, the algebraic ability to generate infinite novel combinations from finite learned components, which is the key to understanding and reasoning about complex logic.In this work, we investigate the compositionality of large language models (LLMs) in mathematical reasoning.Specifically, we construct a new dataset MATHTRAP ‡ by introducing carefully designed logical traps into the problem descriptions of MATH and GSM8K.Since problems with logical flaws are quite rare in the real world, these represent "unseen" cases to LLMs.Solving these requires the models to systematically compose (1) the mathematical knowledge involved in the original problems with (2) knowledge related to the introduced traps.Our experiments show that while LLMs possess both components of requisite knowledge, they do not spontaneously combine them to handle these novel cases.We explore several methods to mitigate this deficiency, such as natural language prompts, few-shot demonstrations, and fine-tuning.Additionally, we test the recently released OpenAI o1 model and find that human-like 'slow thinking' helps improve the compositionality of LLMs.Overall, systematic compositionality remains an open challenge for large language models. Jun Zhao 0019, Jingqi Tong, Yurong Mou, Ming Zhang 0030, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 1 |
| 2023 | Actively Supervised Clustering for Open Relation ExtractionabstractJun Zhao, Yongxin Zhang, Qi Zhang, Tao Gui, Zhongyu Wei, Minlong Peng, Mingming Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jun Zhao 0019, Qi Zhang 0001, Tao Gui, Zhongyu Wei, Minlong Peng, Mingming Sun 0001 |
ACL (1) | 1 |
| 2023 | Open Set Relation Extraction via Unknown-Aware TrainingabstractJun Zhao, Xin Zhao, WenYu Zhan, Qi Zhang, Tao Gui, Zhongyu Wei, Yun Wen Chen, Xiang Gao, Xuanjing Huang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jun Zhao 0019, Wenyu Zhan, Qi Zhang 0001, Tao Gui, Zhongyu Wei, Yun Wen Chen, Xiang Gao 0017, Xuanjing Huang 0001 |
ACL (1) | 1 |
| 2023 | RE-Matching: A Fine-Grained Semantic Matching Method for Zero-Shot Relation ExtractionabstractJun Zhao, WenYu Zhan, Xin Zhao, Qi Zhang, Tao Gui, Zhongyu Wei, Junzhe Wang, Minlong Peng, Mingming Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jun Zhao 0019, Wenyu Zhan, Qi Zhang 0001, Tao Gui, Zhongyu Wei, Junzhe Wang 0001, Minlong Peng, Mingming Sun 0001 |
ACL (1) | 1 |
| 2022 | Read Extensively, Focus Smartly: A Cross-document Semantic Enhancement Method for Visual Documents NERabstractThe introduction of multimodal information and pretraining technique significantly improves entity recognition from visually-rich documents. However, most of the existing methods pay unnecessary attention to irrelevant regions of the current document while ignoring the potentially valuable information in related documents. To deal with this problem, this work proposes a cross-document semantic enhancement method, which consists of two modules: 1) To prevent distractions from irrelevant regions in the current document, we design a learnable attention mask mechanism, which is used to adaptively filter redundant information in the current document. 2) To further enrich the entity-related context, we propose a cross-document information awareness technique, which enables the model to collect more evidence across documents to assist in prediction. The experimental results on two documents understanding benchmarks covering eight languages demonstrate that our method outperforms the SOTA methods. Jun Zhao 0019, Wenyu Zhan, Tao Gui, Qi Zhang 0001, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu |
COLING | 1 |
| 2021 | A Relation-Oriented Clustering Method for Open Relation ExtractionabstractThe clustering-based unsupervised relation discovery method has gradually become one of the important methods of open relation extraction (OpenRE).However, high-dimensional vectors can encode complex linguistic information which leads to the problem that the derived clusters cannot explicitly align with the relational semantic classes.In this work, we propose a relationoriented clustering model and use it to identify the novel relations in the unlabeled data.Specifically, to enable the model to learn to cluster relational data, our method leverages the readily available labeled data of pre-defined relations to learn a relationoriented representation.We minimize distance between the instance with same relation by gathering the instances towards their corresponding relation centroids to form a cluster structure, so that the learned representation is cluster-friendly.To reduce the clustering bias on predefined classes, we optimize the model by minimizing a joint objective on both labeled and unlabeled data.Experimental results show that our method reduces the error rate by 29.2% and 15.7%, on two datasets respectively, compared with current SOTA methods. Jun Zhao 0019, Tao Gui, Qi Zhang 0001, Yaqian Zhou 0001 |
EMNLP (1) | 1 |