Zezhong Wang 0004

dblp:217/9660-4 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0003-4079-0097ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool learning
abstract
Tool learning, which allows Large Language Models (LLMs) to leverage external tools for solving complex user tasks, has emerged as a promising avenue for extending model capabilities. However, existing approaches primarily focus on data synthesis for fine-tuning LLMs to invoke tools effectively, largely ignoring how to fully stimulate the potential of the model. In this paper, we propose ToolACE-R, a novel framework that includes both model-aware iterative training and adaptive refinement for tool learning. ToolACE-R features a model-aware iterative training procedure that progressively adjust training samples based on the model’s evolving capabilities to maximize its potential. Additionally, it incorporates self-refinement training corpus which emphasizes LLM's ability to iteratively refine their tool calls, optimizing performance without requiring external feedback. Furthermore, we introduce adaptive self-refinement for efficient test-time scaling, where the trained model can autonomously determine when to stop the process based on iterative self-refinement. We conduct extensive experiments across several benchmark datasets, showing that ToolACE-R achieves competitive performance compared to advanced LLMs. The performance can be further improved efficiently through adaptive self-refinement. These results highlight the effectiveness and generalizability of ToolACE-R, offering a promising direction for more efficient and scalable tool learning.
Xingshan Zeng, Weiwen Liu, Xu Huang 0008, Zezhong Wang 0004, Lingzhi Wang 0001, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Qun Liu 0001
AAAI4
2026 Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors
abstract
Zhiwei Zhang, Fei Zhao, Rui Wang, Zezhong Wang, Bin Liang, Jiakang Wang, Yao Hu, Shaosheng Cao, Kam-Fai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Fei Zhao 0012, Rui Wang 0092, Zezhong Wang 0004, Bin Liang 0004, Jiakang Wang, Yao Hu 0002, Shaosheng Cao, Kam-Fai Wong
ACL (1)4
2026 Guaranteeing Knowledge Integration with Joint Decoding for Retrieval-Augmented Generation
abstract
Zhengyi Zhao, Shubo Zhang, Zezhong Wang, Yuxi Zhang, Huimin Wang, Yutian Zhao, Yefeng Zheng, Binyang Li, Kam-Fai Wong, Xian Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengyi Zhao 0001, Shubo Zhang, Zezhong Wang 0004, Yutian Zhao, Yefeng Zheng 0001, Binyang Li, Kam-Fai Wong, Xian Wu 0001
ACL (1)3
2026 Bridging Auxiliary Constraints to Resolve Instruction Following in Large Reasoning Models
abstract
Abstract Large Reasoning Models (LRMs) have demonstrated impressive capabilities in many tasks, yet they struggle with reliably following multiple instructions, either by failing to satisfy individual constraints or by struggling to balance competing constraints simultaneously. We formalize this challenge as the Constraint Adherence Problem (CAP). This paper introduces a novel framework that addresses CAP by representing instructions as a structured knowledge graph of constraints. Our approach, Constraint Relationship Graph Completion (CRGC), explicitly models relationships between constraints, identifies adherence challenges, and discovers “bridge constraints” that help the model better focus on and reconcile requirements. Bridge constraints act as auxiliary instructions that make primary constraints more salient and compatible. Unlike existing approaches that enhance instruction following through general training methods, CRGC specifically improves constraint satisfaction by leveraging the model’s own knowledge to create better pathways for generation. Experiments across three popular instruction following datasets demonstrate that our approach reduces constraint violations by 39% compared to standard prompting while maintaining reasoning abilities of large reasoning models.
Zhengyi Zhao 0001, Shubo Zhang, Zezhong Wang 0004, Yutian Zhao, Yefeng Zheng 0001, Binyang Li, Yulan He 0001, Kam-Fai Wong, Xian Wu 0001
Trans. Assoc. Comput. Linguistics4
2025 Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
abstract
Zezhong Wang, Xingshan Zeng, Weiwen Liu, Yufei Wang, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Yufei Wang 0005, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
EMNLP1
2025 T2: An Adaptive Test-Time Scaling Strategy for Contextual Question Answering
abstract
Zhengyi Zhao, Shubo Zhang, Zezhong Wang, Huimin Wang, Yutian Zhao, Bin Liang, Yefeng Zheng, Binyang Li, Kam-Fai Wong, Xian Wu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zhengyi Zhao 0001, Shubo Zhang, Zezhong Wang 0004, Yutian Zhao, Bin Liang 0004, Yefeng Zheng 0001, Binyang Li, Kam-Fai Wong, Xian Wu 0001
EMNLP3
2025 MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models
abstract
Zhengyi Zhao, Shubo Zhang, Yuxi Zhang, Yanxi Zhao, Yifan Zhang, Zezhong Wang, Huimin Wang, Yutian Zhao, Bin Liang, Yefeng Zheng, Binyang Li, Kam-Fai Wong, Xian Wu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zhengyi Zhao 0001, Shubo Zhang, Yanxi Zhao, Yifan Zhang 0004, Zezhong Wang 0004, Yutian Zhao, Bin Liang 0004, Yefeng Zheng 0001, Binyang Li, Kam-Fai Wong, Xian Wu 0001
EMNLP6
2025 ToolACE: Winning the Points of LLM Function Calling
abstract
Function calling significantly extends the application boundary of large language models (LLMs), where high-quality and diverse training data is critical for unlocking this capability. However, collecting and annotating real function-calling data is challenging, while synthetic data from existing pipelines often lack coverage and accuracy. In this paper, we present ToolACE, an automatic agentic pipeline designed to generate accurate, complex, and diverse tool-learning data, specifically tailored to the capabilities of LLMs. ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs. Dialogs are further generated through the interplay among multiple agents, under the guidance of a complexity evaluator. To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks. We demonstrate that models trained on our synthesized data---even with only 8B parameters---achieve state-of-the-art performance, comparable to the latest GPT-4 models. Our model and a subset of the data are publicly available at https://huggingface.co/Team-ACE.
Weiwen Liu, Xu Huang 0008, Xingshan Zeng, Xinlong Hao, Dexun Li, Shuai Wang 0020, Weinan Gan, Zhengying Liu, Yuanqing Yu, Zezhong Wang 0004, Yuxian Wang, Wu Ning, Yutai Hou, Bin Wang 0004, Chuhan Wu, Yong Liu 0020, Yasheng Wang, Duyu Tang, Dandan Tu, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Defu Lian, Qun Liu 0001, Enhong Chen
ICLR11
2025 ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
abstract
Zezhong Wang, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
NAACL (Long Papers)1
2024 JoTR: A Joint Transformer and Reinforcement Learning Framework for Dialogue Policy Learning
abstract
Dialogue policy learning (DPL) aims to determine an abstract representation (also known as action) to guide what the response should be. Typically, DPL is cast as a sequential decision problem across a series of predefined action candidates. However, such static and narrow actions can limit response diversity and impede the dialogue agent’s adaptability to new scenarios and edge cases. To overcome these challenges, we introduce a novel Joint Transformer Reinforcement Learning framework, coined as JoTR, where a text-to-text Transformer-based model is employed to directly generate dialogue actions. More concretely, JoTR formulates a token-grained policy, facilitating more dynamic and adaptable dialogue action generation without the need for predefined action candidates. This method not only enhances the diversity of responses but also significantly improves the system’s capability to manage unfamiliar scenarios. Furthermore, JoTR utilizes Reinforcement Learning with a reward-shaping mechanism to efficiently fine-tune the token-grained policy. This allows the model to evolve through interactions, thereby enhancing its performance over time. Our extensive evaluation demonstrates that JoTR surpasses previous state-of-the-art models, showing improvements of 9% and 13% in success rate, and 34% and 37% in the diversity of dialogue actions across two benchmark dialogue modeling tasks respectively. These results have been validated by both user simulators and human evaluators. Code and data are available at ://github.com/KwanWaiChung/JoTR.
Wai-Chung Kwan, Hongru Wang 0003, Zezhong Wang 0004, Bin Liang 0004, Xian Wu 0001, Yefeng Zheng 0001, Kam-Fai Wong
LREC/COLING4
2024 SELF-GUARD: Empower the LLM to Safeguard Itself
abstract
Zezhong Wang, Fangkai Yang, Lu Wang, Pu Zhao, Hongru Wang, Liang Chen, Qingwei Lin, Kam-Fai Wong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zezhong Wang 0004, Fangkai Yang, Lu Wang 0029, Pu Zhao 0004, Hongru Wang 0003, Liang Chen 0001, Qingwei Lin, Kam-Fai Wong
NAACL-HLT1
2023 MCML: A Novel Memory-based Contrastive Meta-Learning Method for Few Shot Slot Tagging
abstract
Hongru Wang, Zezhong Wang, Wai Chung Kwan, Kam-Fai Wong. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hongru Wang 0003, Zezhong Wang 0004, Wai Chung Kwan, Kam-Fai Wong
IJCNLP (1)2
2022 Integrating Pretrained Language Model for Dialogue Policy Evaluation
abstract
Reinforcement Learning (RL) has been witnessed its potential for training a dialogue policy agent towards maximizing the accumulated rewards given from users. However, the reward can be very sparse for it is usually only provided at the end of a dialog session, which causes unaffordable interaction requirements for an acceptable dialog agent. Distinguished from many efforts dedicated to optimizing the policy and recovering the reward alternatively which suffers from easily getting stuck in local optima and model collapse, we decompose the adversarial training into two steps: 1) we integrate a pre-trained language model as a discriminator to judge whether the current system action is good enough for the last user action (i.e., next action prediction); 2) the discriminator gives and extra local dense reward to guide the agent’s exploration. The experimental result demonstrates that our method significantly improves the complete rate (4.4%) and success rate ( 8.0%) of the dialogue system.
Hongru Wang 0003, Zezhong Wang 0004, Kam-Fai Wong
ICASSP3
2022 GSAN: Graph Self-Attention Network for Learning Spatial-Temporal Interaction Representation in Autonomous Driving
abstract
Modeling interactions among vehicles is critical in improving the efficiency and safety of autonomous driving since complex interactions are ubiquitous in many traffic scenarios. To model interactions under different traffic scenarios, most existing works consider interaction information implicitly in their specific tasks with hand-crafted features and predefined maneuvers. Extracting interaction representation, which can be commonly used among different downstream tasks, is not explored. In this article, we propose a general and novel graph self-attention network (GSAN) to learn the spatial–temporal interaction representation among vehicles by a framework consisting of pretraining and fine-tuning. Specifically, in the pretraining step, we construct the GSAN module based on a graph self-attention layer and a gated recurrent unit layer, and use trajectory autoregression to learn the interaction information among vehicles. In the fine-tuning step, we propose two different adaptation schemes to utilize the learned interaction information in various downstream tasks and fine-tune the entire model with only a few steps. To illustrate the effectiveness and generality of our spatial–temporal interaction model, we conduct extensive experiments on two typical interaction-related tasks, namely, lane-changing classification and trajectory prediction. The experiment results demonstrate that our approach significantly outperforms the state-of-the-art solutions of these two tasks. We also visualize the impact of surrounding vehicles on the ego vehicle in different interaction scenes. The visualization offers an intuitive explanation on how our model captures the dynamic changing interactions among vehicles and makes good predictions in various interaction-related tasks.
Luyao Ye, Zezhong Wang 0004, Xinhong Chen 0003, Jianping Wang 0001, Kui Wu 0001, Kejie Lu
IEEE Internet Things J.2
2020 GSAN: Graph Self-Attention Network for Interaction Measurement in Autonomous Driving
abstract
Modeling the interactions among vehicles has been considered essential in improving efficiency and safety in autonomous driving, since the real traffic scenarios, such as merging lanes, intersection, and lane change, are full of complex interactions. In the literature, interaction is considered implicitly in individual tasks, which makes it hard to extract the interactions for other related downstream tasks. In this paper, we propose a novel Graph Self-Attention Network (GSAN) to quickly capture and quantify the influence of interactions among vehicles from historical trajectories, which can be used as a tool to introduce the impact of interactions into different downstream tasks and further analyze the dominating features affecting the interactions among vehicles. We conduct experiments on the trajectory prediction task as one example to illustrate how to use the spatial-temporal interaction vector to improve the performance of interaction related tasks. The experiment results demonstrate that the GSAN module outperforms the state-of-the-art solutions in terms of the trajectory prediction accuracy. Also, we visualize the effects from all surrounding vehicles on the ego vehicle by heat maps using the trained attention values from the GSAN module.
Luyao Ye, Zezhong Wang 0004, Xinhong Chen 0003, Jianping Wang 0001, Kui Wu 0001, Kejie Lu
MASS2