Jialong Wu 0007

dblp:73/498-7 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Nested Browser-Use Learning for Agentic Information Seeking
abstract
Baixuan Li, Jialong Wu, Wenbiao Yin, Kuan Li, Zhongwang Zhang, Huifeng Yin, Zhengwei Tao, Liwen Zhang, Pengjun Xie, Jingren Zhou, Yong Jiang, Wentao Zhang, Zhiqiang Gao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Baixuan Li, Jialong Wu 0007, Wenbiao Yin, Kuan Li, Zhongwang Zhang, Huifeng Yin, Zhengwei Tao, Pengjun Xie, Jingren Zhou 0001, Yong Jiang 0005, Wentao Zhang 0001
ACL (1)2
2025 Causal Prompting: Debiasing Large Language Model Prompting Based on Front-Door Adjustment
abstract
Despite the notable advancements of existing prompting methods, such as In-Context Learning and Chain-of-Thought for Large Language Models (LLMs), they still face challenges related to various biases. Traditional debiasing methods primarily focus on the model training stage, including approaches based on data augmentation and reweighting, yet they struggle with the complex biases inherent in LLMs. To address such limitations, the causal relationship behind the prompting methods is uncovered using a structural causal model, and a novel causal prompting method based on front-door adjustment is proposed to effectively mitigate LLMs biases. In specific, causal intervention is achieved by designing the prompts without accessing the parameters and logits of LLMs. The chain-of-thought generated by LLM is employed as the mediator variable and the causal effect between input prompts and output answers is calculated through front-door adjustment to mitigate model biases. Moreover, to accurately represent the chain-of-thoughts and estimate the causal effects, contrastive learning is used to fine-tune the encoder of chain-of-thought by aligning its space with that of the LLM. Experimental results show that the proposed causal prompting approach achieves excellent performance across seven natural language processing datasets on both open-source and closed-source LLMs.
Congzhi Zhang, Linhai Zhang, Jialong Wu 0007, Yulan He 0001
AAAI3
2025 SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation
abstract
Key-Value (KV) cache has become a bottleneck of LLMs for long-context generation.Despite the numerous efforts in this area, the optimization for the decoding phase is generally ignored.However, we believe such optimization is crucial, especially for long-output generation tasks based on the following two observations: (i) Excessive compression during the prefill phase which requires specific full context, impairs the comprehension of the reasoning task; (ii) Deviation of heavy hitters 1 occurs in the reasoning tasks with long outputs.Therefore, SCOPE, a simple yet efficient framework that separately performs KV cache optimization during the prefill and decoding phases, is introduced.Specifically, the KV cache during the prefill phase is preserved to maintain the essential information, while a novel strategy based on sliding is proposed to select essential heavy hitters for the decoding phase.Memory usage and memory transfer are further optimized using adaptive and discontinuous strategies.Extensive experiments on LONGGENBENCH show the effectiveness and generalization of SCOPE and its compatibility as a plug-in to other prefill-only KV compression methods. 2
Jialong Wu 0007, Zhenglin Wang, Linhai Zhang, Yilong Lai, Yulan He 0001
ACL (1)1
2025 WebWalker: Benchmarking LLMs in Web Traversal
abstract
Retrieval-augmented generation (RAG) demonstrates remarkable performance across tasks in open-domain question-answering. However, traditional search engines may retrieve shallow content, limiting the ability of LLMs to handle complex, multi-layered information. To address this, we introduce WebWalkerQA, a benchmark designed to assess the ability of LLMs to perform web traversal. It evaluates the capacity of LLMs to traverse a website’s subpages to extract high-quality data systematically. We propose WebWalker, which is a multi-agent framework that mimics human-like web navigation through an explore-critic paradigm. Extensive experimental results show that WebWalkerQA is challenging and demonstrates the effectiveness of RAG combined with WebWalker, through this horizontal and vertical integration in real-world scenarios.
Jialong Wu 0007, Wenbiao Yin, Yong Jiang 0005, Zhenglin Wang, Zekun Xi, Runnan Fang, Linhai Zhang, Yulan He 0001, Pengjun Xie, Fei Huang 0002
ACL (1)1
2025 PROPER: A Progressive Learning Framework for Personalized Large Language Models with Group-Level Adaptation
abstract
Personalized large language models (LLMs) aim to tailor their outputs to user preferences.Recent advances in parameter-efficient finetuning (PEFT) methods have highlighted the effectiveness of adapting population-level LLMs to personalized LLMs by fine-tuning userspecific parameters with user history.However, user data is typically sparse, making it challenging to adapt LLMs to specific user patterns.To address this challenge, we propose PROgressive PERsonalization (PROPER), a novel progressive learning framework inspired by meso-level theory in social science.PROPER bridges population-level and user-level models by grouping users based on preferences and adapting LLMs in stages.It combines a Mixture-of-Experts (MoE) structure with Low Ranked Adaptation (LoRA), using a user-aware router to assign users to appropriate groups automatically.Additionally, a LoRA-aware router is proposed to facilitate the integration of individual user LoRAs with group-level LoRAs.Experimental results show that PROPER significantly outperforms SOTA models across multiple tasks, demonstrating the effectiveness of our approach.Our code is available at https://github.com/callanwu/PROPER.
Linhai Zhang, Jialong Wu 0007, Yulan He 0001
ACL (1)2
2025 AdaCQR: Enhancing Query Reformulation for Conversational Search via Sparse and Dense Retrieval Alignment
abstract
Conversational Query Reformulation (CQR) has significantly advanced in addressing the challenges of conversational search, particularly those stemming from the latent user intent and the need for historical context. Recent works aimed to boost the performance of CQR through alignment. However, they are designed for one specific retrieval system, which potentially results in sub-optimal generalization. To overcome this limitation, we present a novel framework AdaCQR. By aligning reformulation models with both term-based and semantic-based retrieval systems, AdaCQR enhances the generalizability of information-seeking queries among diverse retrieval environments through a two-stage training strategy. Moreover, two effective approaches are proposed to obtain superior labels and diverse input candidates, boosting the efficiency and robustness of the framework. Experimental results on the TopiOCQA, QReCC and TREC CAsT datasets demonstrate that AdaCQR outperforms the existing methods in a more efficient framework, offering both quantitative and qualitative improvements in conversational query reformulation.
Yilong Lai, Jialong Wu 0007, Congzhi Zhang
COLING2
2025 SEED: Accelerating Reasoning Tree Construction via Scheduled Speculative Decoding
abstract
Large Language Models (LLMs) demonstrate remarkable emergent abilities across various tasks, yet fall short of complex reasoning and planning tasks. The tree-search-based reasoning methods address this by encouraging the exploration of intermediate steps, surpassing the capabilities of chain-of-thought prompting. However, significant inference latency is introduced due to the systematic exploration and evaluation of multiple thought paths. This paper introduces SEED, a novel and efficient inference framework to improve both runtime speed and GPU memory management concurrently. Based on a scheduled speculative execution, SEED efficiently handles multiple iterations for thought generation and state evaluation, leveraging a rounds-scheduled strategy to manage draft model dispatching. Extensive experimental evaluations on three reasoning datasets demonstrate the superior speedup performance of SEED.
Zhenglin Wang, Jialong Wu 0007, Yilong Lai, Congzhi Zhang
COLING2
2025 AdaRewriter: Unleashing the Power of Prompting-based Conversational Query Reformulation via Test-Time Adaptation
abstract
Prompting-based conversational query reformulation has emerged as a powerful approach for conversational search, refining ambiguous user queries into standalone search queries.Bestof-N reformulation over the generated candidates via prompting shows impressive potential scaling capability.However, both the previous tuning methods (training time) and adaptation approaches (test time) can not fully unleash their benefits.In this paper, we propose AdaRewriter, a novel framework for query reformulation using an outcome-supervised reward model via test-time adaptation.By training a lightweight reward model with contrastive ranking loss, AdaRewriter selects the most promising reformulation during inference.Notably, it can operate effectively in black-box systems, including commercial LLM APIs.Experiments on five conversational search datasets show that AdaRewriter significantly outperforms the existing methods across most settings, demonstrating the potential of test-time adaptation for conversational query reformulation. 1
Yilong Lai, Jialong Wu 0007, Zhenglin Wang
EMNLP2
2025 OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking
abstract
Zekun Xi, Wenbiao Yin, Jizhan Fang, Jialong Wu, Runnan Fang, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen, Ningyu Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zekun Xi, Wenbiao Yin, Jizhan Fang, Jialong Wu 0007, Runnan Fang, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen, Ningyu Zhang 0001
EMNLP4
2025 EvolveSearch: An Iterative Self-Evolving Search Agent
abstract
Ding-Chu Zhang, Yida Zhao, Jialong Wu, Liwen Zhang, Baixuan Li, Wenbiao Yin, Yong Jiang, Yu-Feng Li, Kewei Tu, Pengjun Xie, Fei Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Dingchu Zhang, Yida Zhao, Jialong Wu 0007, Baixuan Li, Wenbiao Yin, Yong Jiang 0005, Kewei Tu, Pengjun Xie, Fei Huang 0002
EMNLP3
2025 WebDancer: Towards Autonomous Information Seeking Agency
abstract
Addressing intricate real-world problems necessitates in-depth information seeking and multi-step reasoning. Recent progress in agentic systems, exemplified by Deep Research, underscores the potential for autonomous multi-step research. In this work, we present a cohesive paradigm for building end-to-end agentic information seeking agents from a data-centric and training-stage perspective. Our approach consists of four key stages: (1) browsing data construction, (2) trajectories sampling, (3) supervised fine-tuning for effective cold start, and (4) reinforcement learning for enhanced generalisation. We instantiate this framework in a web agent based on the ReAct format, WebDancer. Empirical evaluations on the challenging GAIA and WebWalkerQA benchmarks demonstrate the strong performance of WebDancer, achieving considerable results and highlighting the efficacy of our training paradigm. Further analysis of agent training provides valuable insights and actionable, systematic pathways for developing more capable agentic models.
Jialong Wu 0007, Baixuan Li, Runnan Fang, Wenbiao Yin, Zhenglin Wang, Zhengwei Tao, Dingchu Zhang, Zekun Xi, Robert Tang, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Jingren Zhou 0001
NeurIPS1