VLDB 2026 Research / reviewers in the wild / expert
Jinhao Jiang
dblp:261/6942
· DBLP profile ↗
17ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vertical Federated K-Means for Multi-View Data Guided by a K-Means Cost Bound after ProjectionabstractMulti-view data is widely present in the real world. Multi-view clustering is an unsupervised method for capturing the grouping structure of such data. However, multi-view clustering struggles to meet the requirements of real-world scenarios, such as distributed storage of different views and data protection needs. These requirements align with the setting of vertical federated clustering. However, vertical federated clustering still faces two challenges: (1) Under the constraints of privacy protection mechanisms, how to theoretically analyze the clustering consistency between the data uploaded by clients to the server and the original client data is challenging. (2) The feature space differences among different clients make cross-view information sharing and fusion difficult. To address the first challenge, we provide a theoretical analysis of the upper bound of the loss of k-means for transformation matrix mapping, revealing the relationship between the k-means loss of the transformed data and the original data. We then propose a vertical federated clustering method (V-HDKM). In this method, clients handle the second challenge by transposing the feature matrix. Guided by the projected k-means loss bound, we expand the feature space and perform k-means clustering to obtain feature cluster centers, which are then uploaded to the server. The server aggregates the global centers and feeds back the optimized results, achieving cross-view knowledge fusion through iterative interactions. Experimental results show that V-HDKM significantly improves local clustering performance and performances better than other seven vertical federated mthods on 20 multi-view datasets. Furthermore, sensitivity analysis on 8 UCI datasets with respect to the number of clients demonstrates the stability of the method. The code is available at https://github.com/jiangjh/V-HDKM. Feijiang Li, Jinhao Jiang, Jieting Wang, Liang Du 0003 |
KDD (1) | 2 |
| 2026 | A Survey of Large Language ModelsabstractAbstract The rapid evolution of large language models (LLMs) has driven a transformative shift in artificial intelligence (AI), reshaping both research paradigms and practical applications. Distinguished from their predecessors by unprecedented scale and advanced capabilities, LLMs necessitate new frameworks for understanding their development, behavior, and societal impact. This survey systematically reviews recent advancements in LLM techniques across four key dimensions: (1) pre-training methodologies, which establish core model capabilities through large-scale self-supervised training, architectural innovations, and data curation strategies; (2) post-training techniques, including supervised fine-tuning and reinforcement learning, which adapt foundational models to downstream tasks and enhance their alignment and safety; (3) utilization strategies, such as in-context learning, prompt engineering, and agentic reasoning, that optimize real-world deployment and enable effective interaction with external environments; and (4) evaluation methods, encompassing benchmarks for key ability dimensions such as core language capabilities, reasoning, and safety, which support comprehensive and reliable assessment of model performance. Additionally, we identify critical research issues, including those concerning theoretical foundations, efficient scaling, alignment, and agentic capability, and highlight the open challenges they present. By synthesizing state-of-the-art insights and emerging trends, this survey aims to provide a systematic and comprehensive framework for understanding the trajectory, current limitations, and future directions of LLM progress. Wayne Xin Zhao, Kun Zhou 0002, Junyi Li 0001, Zican Dong, Yupeng Hou, Beichen Zhang 0003, Yingqian Min, Junjie Zhang 0009, Peiyu Liu 0002, Xiaolei Wang 0005, Yifan Du 0002, Chen Yang 0032, Zhipeng Chen 0001, Jinhao Jiang, Ruiyang Ren, Yifan Li 0009, Xinyu Tang 0004, Zikang Liu 0001, Jian-Yun Nie, Ji-Rong Wen |
Frontiers Comput. Sci. | 16 |
| 2026 | Coupling Structural Descriptors With a Novel Semantic Graph Matching Approach for LiDAR Loop DetectionabstractOutdoor loop closure detection is essential for correcting odometry drift and constructing a globally consistent map. Semantic-graph-based approaches effectively model object-level topology and achieve strong loop closure performance; however, their effectiveness degrades in background-dominated scenes with few distinctive objects, and establishing accurate injective node correspondences remains challenging. In contrast, structural descriptor methods, though offering stronger environmental generality through spatial-distribution modeling, remain susceptible to LiDAR noise and the discriminative power of point-level features. These limitations motivate the need for a more robust method that combines adaptability with enhanced descriptive power. We propose a novel loop-closure detection framework, SAGE, that integrates highly adaptable point-cloud shape-distribution features and generally reliable semantic graph topology, adaptively combining their similarity measures to improve detection performance. Specifically, we design a semantic graph matching module with dual constraints, local graph feature consistency and global spatial consistency, to achieve more accurate injective node correspondences. In addition, we extract point-cloud shape-distribution features and introduce a fusion mechanism that integrates them with the semantic graph module, assessing reliability and adaptively weighting their contributions. Extensive loop closure detection and pose estimation experiments on various datasets demonstrate that SAGE achieves superior performance over strong baselines. We provide the code at https://github.com/SAGE-11/SAGE. Meiling Wang 0002, Sibo Zuo, Chengxi Yang, Jinhao Jiang, Xieyuanli Chen, Yufeng Yue |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Towards Effective and Efficient Continual Pre-training of Large Language ModelsabstractContinual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. In this paper, we comprehensively study its key designs to balance the new abilities while retaining the original abilities, and present an effective CPT method that can greatly improve the Chinese language ability and scientific reasoning ability of LLMs. To achieve it, we design specific data mixture and curriculum strategies based on existing datasets and synthetic high-quality data. Concretely, we synthesize multidisciplinary scientific QA pairs based on related web pages to guarantee the data quality, and also devise the performance tracking and data mixture adjustment strategy to ensure the training stability. For the detailed designs, we conduct preliminary studies on a relatively small model, and summarize the findings to help optimize our CPT method. Extensive experiments on a number of evaluation benchmarks show that our approach can largely improve the performance of Llama-3 (8B), including both the general abilities (+8.81 on C-Eval and +6.31 on CMMLU) and the scientific reasoning abilities (+12.00 on MATH and +4.13 on SciEval). Our model, data, and codes are available at https://github.com/RUC-GSAI/Llama-3-SynE. Jie Chen 0007, Zhipeng Chen 0001, Kun Zhou 0002, Yutao Zhu 0001, Jinhao Jiang, Yingqian Min, Wayne Xin Zhao, Zhicheng Dou, Jiaxin Mao, Yankai Lin 0001, Ruihua Song, Jun Xu 0001, Xu Chen 0017, Rui Yan 0001, Zhewei Wei, Di Hu 0001, Wenbing Huang 0001, Ji-Rong Wen |
ACL (1) | 6 |
| 2025 | LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration DistillationabstractZican Dong, Junyi Li, Jinhao Jiang, Mingyu Xu, Xin Zhao, Bingning Wang, Weipeng Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zican Dong, Jinhao Jiang, Bingning Wang, Weipeng Chen |
ACL (1) | 3 |
| 2025 | YuLan-Mini: Pushing the Limits of Open Data-efficient Language ModelabstractDue to the immense resource demands and the involved complex techniques, it is still challenging for successfully pre-training a large language models (LLMs) with state-of-the-art performance. In this paper, we explore the key bottlenecks and designs during pre-training, and make the following contributions: (1) a comprehensive investigation into the factors contributing to training instability; (2) a robust optimization approach designed to mitigate training instability effectively; (3) an elaborate data pipeline that integrates data synthesis, data curriculum, and data selection. By integrating the above techniques, we create a rather low-cost training recipe and use it to pre-train YuLan-Mini, a fully-open base model with 2.4B parameters on 1.08T tokens. Remarkably, YuLan-Mini achieves top-tier performance among models of similar parameter scale, with comparable performance to industry-leading models that require significantly more data. To facilitate reproduction, we release the full details of training recipe and data composition. Project details can be accessed at the following link: https://anonymous.4open.science/r/YuLan-Mini/README.md. Huatong Song, Jie Chen 0007, Kun Zhou 0002, Yutao Zhu 0001, Jinhao Jiang, Zican Dong, Xu Miao, Wayne Xin Zhao, Ji-Rong Wen |
ACL (1) | 8 |
| 2025 | KG-Agent: An Efficient Autonomous Agent Framework for Complex Reasoning over Knowledge GraphabstractIn this paper, we aim to improve the reasoning ability of large language models (LLMs) over knowledge graphs (KGs) to answer complex questions.Inspired by existing methods that design the interaction strategy between LLMs and KG, we propose an autonomous LLM-based agent framework, called KG-Agent, which enables a small LLM to actively make decisions until finishing the reasoning process over KGs.In KG-Agent, we integrate the LLM, multifunctional toolbox, KG-based executor, and knowledge memory, and develop an iteration mechanism that autonomously selects the tool and then updates the memory for reasoning over KG.To guarantee the effectiveness, we leverage program language to formulate the multi-hop reasoning process over the KG and synthesize a code-based instruction dataset to fine-tune the base LLM.Extensive experiments demonstrate that only using 10K samples for tuning LLaMA2-7B can outperform competitive methods using larger LLMs or more data, on both in-domain and out-domain datasets.Our code and data will be publicly released. Jinhao Jiang, Kun Zhou 0002, Wayne Xin Zhao, Yang Song 0021, Chen Zhu 0003, Hengshu Zhu, Ji-Rong Wen |
ACL (1) | 1 |
| 2025 | Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling FrameworkabstractLarge reasoning models (LRMs) have exhibited strong performance on complex reasoning tasks, with further gains achievable through increased computational budgets at inference.However, current test-time scaling methods predominantly rely on redundant sampling, ignoring the historical experience utilization, thereby limiting computational efficiency.To overcome this limitation, we propose Sticker-TTS, a novel test-time scaling framework that coordinates three collaborative LRMs to iteratively explore and refine solutions guided by historical attempts.At the core of our framework are distilled key conditions-termed stickers-which drive the extraction, refinement, and reuse of critical information across multiple rounds of reasoning.To further enhance the efficiency and performance of our framework, we introduce a two-stage optimization strategy that combines imitation learning with self-improvement, enabling progressive refinement.Extensive evaluations on three challenging mathematical reasoning benchmarks, including AIME-24, AIME-25, and Olym-MATH, demonstrate that Sticker-TTS consistently surpasses strong baselines, including self-consistency and advanced reinforcement learning approaches, under comparable inference budgets.These results highlight the effectiveness of sticker-guided historical experience utilization. Jie Chen 0007, Jinhao Jiang, Yingqian Min, Zican Dong, Wayne Xin Zhao, Ji-Rong Wen |
EMNLP | 2 |
| 2025 | CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA CapabilityabstractAdvancements in Large Language Models (LLMs) have extended their input context length, yet they still struggle with retrieval and reasoning in long-context inputs.Existing methods propose to utilize the prompt strategy and Retrieval-Augmented Generation (RAG) to alleviate this limitation.However, they still face challenges in balancing retrieval precision and recall, impacting their efficacy in answering questions.To address this, we introduce CAFE, a two-stage coarse-to-fine method to enhance multi-document questionanswering capacities.By gradually eliminating the negative impacts of background and distracting documents, CAFE makes the responses more reliant on the evidence documents.Initially, a coarse-grained filtering method leverages retrieval heads to identify and rank relevant documents.Then, a fine-grained steering method guides attention to the most relevant content.Experiments across benchmarks show that CAFE outperforms baselines, achieving an average SubEM improvement of up to 22.1% and 13.7% over SFT and RAG methods, respectively, across three different models.Our code is available at https://github. com/RUCAIBox/CAFE. Jinhao Jiang, Zican Dong, Wayne Xin Zhao |
EMNLP | 2 |
| 2025 | Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format AlignmentabstractAdapting large language models (LLMs) to specialized domains typically requires domain-specific corpora for continual pre-training to facilitate knowledge memorization and related instructions for fine-tuning to apply this knowledge.
However, this method may lead to inefficient knowledge memorization due to a lack of awareness of knowledge utilization during the continual pre-training and demands LLMs to simultaneously learn knowledge utilization and format alignment with divergent training objectives during the fine-tuning.
To enhance the domain adaptation of LLMs, we revise this process and propose a new domain adaptation framework including domain knowledge learning and general format alignment, called \emph{Mix-CPT}. Specifically, we first conduct a knowledge mixture continual pre-training that concurrently focuses on knowledge memorization and utilization. To avoid catastrophic forgetting, we further propose a logit swap self-distillation constraint. By leveraging the knowledge and capabilities acquired during continual pre-training, we then efficiently perform instruction tuning and alignment with a few general training samples to achieve format alignment.
Extensive experiments show that our proposed \emph{Mix-CPT} framework can simultaneously improve the task-solving capabilities of LLMs on the target and general domains. Jinhao Jiang, Junyi Li 0001, Wayne Xin Zhao, Yang Song 0021, Tao Zhang 0070, Ji-Rong Wen |
ICLR | 1 |
| 2025 | RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and RefinementabstractJinhao Jiang, Jiayi Chen, Junyi Li, Ruiyang Ren, Shijie Wang, Xin Zhao, Yang Song, Tao Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jinhao Jiang, Jiayi Chen 0005, Junyi Li 0001, Ruiyang Ren, Wayne Xin Zhao, Yang Song 0021, Tao Zhang 0070 |
NAACL (Long Papers) | 1 |
| 2025 | LLM-based Search Assistant with Holistically Guided MCTS for Intricate Information SeekingabstractIn the era of vast digital information, the sheer volume and heterogeneity of available information present significant challenges for intricate information seeking. Users frequently face multistep web search tasks that involve navigating vast and varied data sources. This complexity demands every step remains comprehensive, accurate, and relevant. However, traditional search methods often struggle to balance the need for localized precision with the broader context required for holistic understanding, leaving critical facets of intricate queries underexplored. In this paper, we introduce an LLM-based search assistant that adopts a new information seeking paradigm with holistically guided Monte Carlo tree search (HG-MCTS). We reformulate the task as a progressive information collection process with a knowledge memory and unite an adaptive checklist with multi-perspective reward modeling in MCTS. The adaptive checklist provides explicit sub-goals to guide the MCTS process toward comprehensive coverage of complex user queries. Simultaneously, our multi-perspective reward modeling offers both exploration and retrieval rewards, along with progress feedback that tracks completed and remaining sub-goals, refining the checklist as the tree search progresses. By striking a balance between localized tree expansion and global guidance, HG-MCTS reduces redundancy in search paths and ensures that all crucial aspects of an intricate query are properly addressed. Extensive experiments on real-world intricate information seeking tasks demonstrate that HG-MCTS acquires thorough knowledge collections and delivers more accurate final responses compared with existing baselines. Ruiyang Ren, Yuhao Wang 0007, Junyi Li 0001, Jinhao Jiang, Wayne Xin Zhao, Wenjie Wang 0007, Tat-Seng Chua |
SIGIR | 4 |
| 2023 | StructGPT: A General Framework for Large Language Model to Reason over Structured DataabstractIn this paper, we aim to improve the reasoning ability of large language models (LLMs) over structured data in a unified way.Inspired by the studies on tool augmentation for LLMs, we develop an Iterative Reading-then-Reasoning (IRR) framework to solve question answering tasks based on structured data, called StructGPT.In this framework, we construct the specialized interfaces to collect relevant evidence from structured data (i.e., reading), and let LLMs concentrate on the reasoning task based on the collected information (i.e., reasoning).Specially, we propose an invokinglinearization-generation procedure to support LLMs in reasoning on the structured data with the help of the interfaces.By iterating this procedure with provided interfaces, our approach can gradually approach the target answers to a given query.Experiments conducted on three types of structured data show that StructGPT greatly improves the performance of LLMs, under the few-shot and zero-shot settings.Our codes and data are publicly available at https://github.com/RUCAIBox/StructGPT. Jinhao Jiang, Kun Zhou 0002, Zican Dong, Keming Ye, Wayne Xin Zhao, Ji-Rong Wen |
EMNLP | 1 |
| 2023 | ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge GraphabstractQuestion Answering over Knowledge Graph (KGQA) aims to seek answer entities for the natural language question from a largescale Knowledge Graph (KG).To better perform reasoning on KG, recent work typically adopts a pre-trained language model (PLM) to model the question, and a graph neural network (GNN) based module to perform multihop reasoning on the KG.Despite the effectiveness, due to the divergence in model architecture, the PLM and GNN are not closely integrated, limiting the knowledge sharing and finegrained feature interactions.To solve it, we aim to simplify the above two-module approach, and develop a more capable PLM that can directly support subgraph reasoning for KGQA, namely ReasoningLM.In our approach, we propose a subgraph-aware self-attention mechanism to imitate the GNN for performing structured reasoning, and also adopt an adaptation tuning strategy to adapt the model parameters with 20,000 subgraphs with synthesized questions.After adaptation, the PLM can be parameter-efficient fine-tuned on downstream tasks.Experiments show that ReasoningLM surpasses state-of-the-art models by a large margin, even with fewer updated parameters and less training data.Our codes and data are publicly available at https://github.com/ RUCAIBox/ReasoningLM. Jinhao Jiang, Kun Zhou 0002, Wayne Xin Zhao, Yaliang Li, Ji-Rong Wen |
EMNLP | 1 |
| 2023 | UniKGQA: Unified Retrieval and Reasoning for Solving Multi-hop Question Answering Over Knowledge Graph
Jinhao Jiang, Kun Zhou 0002, Wayne Xin Zhao, Ji-Rong Wen |
ICLR | 1 |
| 2023 | Complex Knowledge Base Question Answering: A SurveyabstractKnowledge base question answering (KBQA) aims to answer a question over a knowledge base (KB). Early studies mainly focused on answering simple questions over KBs and achieved great success. However, their performances on complex questions are still far from satisfactory. Therefore, in recent years, researchers propose a large number of novel methods, which looked into the challenges of answering complex questions. In this survey, we review recent advances in KBQA with the focus on solving complex questions, which usually contain multiple subjects, express compound relations, or involve numerical operations. In detail, we begin with introducing the complex KBQA task and relevant background. Then, we present two mainstream categories of methods for complex KBQA, namely semantic parsing-based (SP-based) methods and information retrieval-based (IR-based) methods. Specifically, we illustrate their procedures with flow designs and discuss their difference and similarity. Next, we summarize the challenges that these two categories of methods encounter when answering complex questions, and explicate advanced solutions as well as techniques used in existing work. After that, we discuss the potential impact of pre-trained language models (PLMs) on complex KBQA. To help readers catch up with SOTA methods, we also provide a comprehensive evaluation and resource about complex KBQA task. Finally, we conclude and discuss several promising directions related to complex KBQA for future research. Yunshi Lan, Gaole He, Jinhao Jiang, Jing Jiang 0001, Wayne Xin Zhao, Ji-Rong Wen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | A Survey on Complex Knowledge Base Question Answering: Methods, Challenges and SolutionsabstractKnowledge base question answering (KBQA) aims to answer a question over a knowledge base (KB). Recently, a large number of studies focus on semantically or syntactically complicated questions. In this paper, we elaborately summarize the typical challenges and solutions for complex KBQA. We begin with introducing the background about the KBQA task. Next, we present the two mainstream categories of methods for complex KBQA, namely semantic parsing-based (SP-based) methods and information retrieval-based (IR-based) methods. We then review the advanced methods comprehensively from the perspective of the two categories. Specifically, we explicate their solutions to the typical challenges. Finally, we conclude and discuss some promising directions for future research. Yunshi Lan, Gaole He, Jinhao Jiang, Jing Jiang 0001, Wayne Xin Zhao, Ji-Rong Wen |
IJCAI | 3 |