Zhong Zhang 0004

dblp:28/1568-4 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0003-1349-9755ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub
abstract
Bohan Lyu, Xin Cong, Heyang Yu, Pan Yang, Cheng Qian, Zihe Wang, Yujia Qin, Yining Ye, Yaxi Lu, Chen Qian, Zhong Zhang, Yukun Yan, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Bohan Lyu 0001, Xin Cong, Heyang Yu, Pan Yang 0022, Cheng Qian 0008, Yujia Qin, Yining Ye, Yaxi Lu, Zhong Zhang 0004, Yukun Yan, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)11
2025 Learning to Generate Structured Output with Schema Reinforcement Learning
abstract
Yaxi Lu, Haolun Li, Xin Cong, Zhong Zhang, Yesai Wu, Yankai Lin, Zhiyuan Liu, Fangming Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yaxi Lu, Haolun Li 0003, Xin Cong, Zhong Zhang 0004, Yesai Wu, Yankai Lin 0001, Zhiyuan Liu 0001, Fangming Liu, Maosong Sun 0001
ACL (1)4
2025 AgentRM: Enhancing Agent Generalization with Reward Modeling
abstract
Yu Xia, Jingru Fan, Weize Chen, Siyu Yan, Xin Cong, Zhong Zhang, Yaxi Lu, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jingru Fan, Weize Chen, Xin Cong, Zhong Zhang 0004, Yaxi Lu, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)6
2025 Learning Evolving Tools for Large Language Models
abstract
Tool learning enables large language models (LLMs) to interact with external tools and APIs, greatly expanding the application scope of LLMs. However, due to the dynamic nature of external environments, these tools and APIs may become outdated over time, preventing LLMs from correctly invoking tools. Existing research primarily focuses on static environments and overlooks this issue, limiting the adaptability of LLMs in real-world applications. In this paper, we propose ToolEVO, a novel framework designed to enhance the adaptive and reflective capabilities of LLMs against tool variability. By leveraging Monte Carlo Tree Search, ToolEVO facilitates active exploration and interaction of LLMs within dynamic environments, allowing for autonomous self-reflection and self-updating of tool usage based on environmental feedback. Additionally, we introduce ToolQA-D, a benchmark specifically designed to evaluate the impact of tool variability. Extensive experiments demonstrate the effectiveness and stability of our approach, highlighting the importance of adaptability to tool variability for effective tool learning.
Guoxin Chen, Zhong Zhang 0004, Xin Cong, Fangda Guo, Yesai Wu, Yankai Lin 0001, Wenzheng Feng, Yasheng Wang
ICLR2
2025 WorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language Models
abstract
Recent advancements in large language models (LLMs) have driven a revolutionary paradigm shift in process automation from Robotic Process Automation to Agentic Process Automation by automating the workflow orchestration procedure based on LLMs. However, existing LLMs (even the advanced OpenAI GPT-4o) are confined to achieving satisfactory capability in workflow orchestration. To address this limitation, we present WorkflowLLM, a data-centric framework elaborately designed to enhance the capability of LLMs in workflow orchestration. It first constructs a large-scale fine-tuning dataset WorkflowBench with 106, 763 samples, covering 1, 503 APIs from 83 applications across 28 categories. Specifically, the construction process can be divided into three phases: (1) Data Collection: we collect real-world workflow data from Apple Shortcuts and RoutineHub, transcribing them into Python-style code. We further equip them with generated hierarchical thought via GPT-4o-mini. (2) Query Expansion: we prompt GPT-4o-mini to generate more task queries to enrich the diversity and complexity of workflows. (3) Workflow Generation: we leverage an annotator model trained on collected data to generate workflows for synthesized queries. Finally, we merge the synthetic samples that pass quality confirmation with the collected samples to obtain the WorkflowBench. Based on WorkflowBench, we fine-tune Llama-3.1-8B to obtain WorkflowLlama. Our experiments show that WorkflowLlama demonstrates a strong capacity to orchestrate complex workflows, while also achieving notable generalization performance on previously unseen APIs. Additionally, WorkflowBench exhibits robust zero-shot generalization capabilities on an out-of-distribution task planning dataset, T-Eval. Our data and code are available at https://github.com/OpenBMB/WorkflowLLM.
Shengda Fan, Xin Cong, Yuepeng Fu, Zhong Zhang 0004, Yuanwei Liu, Yesai Wu, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR4
2025 Proactive Agent: Shifting LLM Agents from Reactive Responses to Active Assistance
abstract
Agents powered by large language models have shown remarkable abilities in solving complex tasks. However, most agent systems remain reactive, limiting their effectiveness in scenarios requiring foresight and autonomous decision-making. In this paper, we tackle the challenge of developing proactive agents capable of anticipating and initiating tasks without explicit human instructions. We propose a novel data-driven approach for this problem. Firstly, we collect real-world human activities to generate proactive task predictions. These predictions are then labeled by human annotators as either accepted or rejected. The labeled data is used to train a reward model that simulates human judgment and serves as an automatic evaluator of the proactiveness of LLM agents. Building on this, we develop a comprehensive data generation pipeline to create a diverse dataset, ProactiveBench, containing 6,790 events. Finally, we demonstrate that fine-tuning models with the proposed ProactiveBench can significantly elicit the proactiveness of LLM agents. Experimental results show that our fine-tuned model achieves an F1-Score of 66.47% in proactively offering assistance, outperforming all open-source and close-source models. These results highlight the potential of our method in creating more proactive and effective agent systems, paving the way for future advancements in human-agent collaboration.
Yaxi Lu, Shenzhi Yang, Cheng Qian 0008, Guirong Chen, Qinyu Luo, Yesai Wu, Xin Cong, Zhong Zhang 0004, Yankai Lin 0001, Weiwen Liu, Yasheng Wang, Zhiyuan Liu 0001, Fangming Liu, Maosong Sun 0001
ICLR9
2025 Generalizing Experience for Language Agents with Hierarchical MetaFlows
abstract
Recent efforts to employ large language models (LLMs) as agents have demonstrated promising results in a wide range of multi-step agent tasks. However, existing agents lack an effective experience reuse approach to leverage historical completed tasks. In this paper, we propose a novel experience reuse framework MetaFlowLLM, which constructs a hierarchical experience tree from historically completed tasks. Each node in this experience tree is presented as a MetaFlow which contains static execution workflow and subtask required by agents to complete dynamically. Then, we propose a Hierarchical MetaFlow Merging algorithm to construct the hierarchical experience tree. When accomplishing a new task, MetaFlowLLM can first retrieve the most relevant MetaFlow node from the experience tree and then execute it accordingly. To effectively generate valid MetaFlows from historical data, we further propose a reinforcement learning pipeline to train the MetaFlowGen. Extensive experimental results on AppWorld and WorkBench demonstrate that integrating with MetaFlowLLM, existing agents (e.g., ReAct, Reflexion) can gain substantial performance improvement with reducing execution costs. Notably, MetaFlowLLM achieves an average success rate improvement of 32.3% on AppWorld and 6.2% on WorkBench, respectively.
Shengda Fan, Xin Cong, Zhong Zhang 0004, Yuepeng Fu, Yesai Wu, Hao Wang 0139, Enrui Hu, Yankai Lin 0001
NeurIPS3
2024 Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
abstract
Cheng Qian, Bingxiang He, Zhong Zhuang, Jia Deng, Yujia Qin, Xin Cong, Zhong Zhang, Jie Zhou, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Cheng Qian 0008, Bingxiang He, Zhong Zhuang, Yujia Qin, Xin Cong, Zhong Zhang 0004, Jie Zhou 0016, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)7
2023 Fine-tuning Happens in Tiny Subspaces: Exploring Intrinsic Task-specific Subspaces of Pre-trained Language Models
abstract
Pre-trained language models (PLMs) are known to be overly parameterized and have significant redundancy, indicating a small degree of freedom of the PLMs.Motivated by the observation, in this paper, we study the problem of re-parameterizing and fine-tuning PLMs from a new perspective: Discovery of intrinsic task-specific subspace.Specifically, by exploiting the dynamics of the fine-tuning process for a given task, the parameter optimization trajectory is learned to uncover its intrinsic task-specific subspace.A key finding is that PLMs can be effectively fine-tuned in the subspace with a small number of free parameters.Beyond, we observe some outlier dimensions emerging during fine-tuning in the subspace.Disabling these dimensions degrades the model performance significantly.This suggests that these dimensions are crucial to induce task-specific knowledge to downstream tasks.
Zhong Zhang 0004, Junming Shao
ACL (1)1
2023 Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive Recommendation
abstract
Offline reinforcement learning (RL), a technology that offline learns a policy from logged data without the need to interact with online environments, has become a favorable choice in decision-making processes like interactive recommendation. Offline RL faces the value overestimation problem. To address it, existing methods employ conservatism, e.g., by constraining the learned policy to be close to behavior policies or punishing the rarely visited state-action pairs. However, when applying such offline RL to recommendation, it will cause a severe Matthew effect, i.e., the rich get richer and the poor get poorer, by promoting popular items or categories while suppressing the less popular ones. It is a notorious issue that needs to be addressed in practical recommender systems. In this paper, we aim to alleviate the Matthew effect in offline RL-based recommendation. Through theoretical analyses, we find that the conservatism of existing methods fails in pursuing users' long-term satisfaction. It inspires us to add a penalty term to relax the pessimism on states with high entropy of the logging policy and indirectly penalizes actions leading to less diverse states. This leads to the main technical contribution of the work: Debiased model-based Offline RL (DORL) method. Experiments show that DORL not only captures user interests well but also alleviates the Matthew effect. The implementation is available via https://github.com/chongminggao/DORL-codes.
Chongming Gao, Jiawei Chen 0007, Yuan Zhang 0024, Biao Li 0002, Peng Jiang 0002, Shiqi Wang 0018, Zhong Zhang 0004, Xiangnan He 0001
SIGIR8
2022 Mixhead: Breaking the low-rank bottleneck in multi-head attention language models
Zhong Zhang 0004, Nian Shao, Chongming Gao, Qinli Yang, Junming Shao
Knowl. Based Syst.1
2022 Pixel-wise triplet learning for enhancing boundary discrimination in medical image segmentation
Leiting Chen, Yu Deng 0005, Zhong Zhang 0004, Chuan Zhou 0004
Knowl. Based Syst.4
2020 Semantic trajectory representation and retrieval via hierarchical embedding
Chongming Gao, Zhong Zhang 0004, Hongzhi Yin, Qinli Yang, Junming Shao
Inf. Sci.2
2020 Structured subspace embedding on attributed networks
Zhongjing Yu, Zhong Zhang 0004, Junming Shao
Inf. Sci.2
2019 Towards Robust Arbitrarily Oriented Subspace Clustering
Zhong Zhang 0004, Chongming Gao, Chongzhi Liu, Qinli Yang, Junming Shao
DASFAA (1)1
2019 Community Detection and Link Prediction via Cluster-driven Low-rank Matrix Completion
abstract
Community detection and link prediction are highly dependent since knowing cluster structure as a priori will help identify missing links, and in return, clustering on networks with supplemented missing links will improve community detection performance. In this paper, we propose a Cluster-driven Low-rank Matrix Completion (CLMC), for performing community detection and link prediction simultaneously in a unified framework. To this end, CLMC decomposes the adjacent matrix of a target network as three additive matrices: clustering matrix, noise matrix and supplement matrix. The community-structure and low-rank constraints are imposed on the clustering matrix, such that the noisy edges between communities are removed and the resulting matrix is an ideal block-diagonal matrix. Missing edges are further learned via low-rank matrix completion. Extensive experiments show that CLMC achieves state-of-the-art performance.
Junming Shao, Zhong Zhang 0004, Zhongjing Yu, Qinli Yang
IJCAI2
2018 Graph Clustering with Local Density-Cut
Junming Shao, Qinli Yang, Zhong Zhang 0004, Jinhu Liu, Stefan Kramer 0001
DASFAA (1)3
2018 Multi-view Discriminative Learning via Joint Non-negative Matrix Factorization
Zhong Zhang 0004, Zhili Qin, Peiyan Li 0002, Qinli Yang, Junming Shao
DASFAA (2)1