Tong Zhu 0002

dblp:36/1469-2 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-5433-8504ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 15 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
abstract
Zhenhua Liu, Lijun Li, Ruizhe Chen, Yuxian Jiang, Tong Zhu, Zhaochen Su, Wenliang Chen, Jing Shao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ruizhe Chen, Yuxian Jiang, Tong Zhu 0002, Zhaochen Su, Wenliang Chen
ACL (1)5
2025 NesTools: A Dataset for Evaluating Nested Tool Learning Abilities of Large Language Models
abstract
Large language models (LLMs) combined with tool learning have gained impressive results in real-world applications. During tool learning, LLMs may call multiple tools in nested orders, where the latter tool call may take the former response as its input parameters. However, current research on the nested tool learning capabilities is still under-explored, since the existing benchmarks lack relevant data instances. To address this problem, we introduce NesTools to bridge the current gap in comprehensive nested tool learning evaluations. NesTools comprises a novel automatic data generation method to construct large-scale nested tool calls with different nesting structures. With manual review and refinement, the dataset is in high quality and closely aligned with real-world scenarios. Therefore, NesTools can serve as a new benchmark to evaluate the nested tool learning abilities of LLMs. We conduct extensive experiments on 22 LLMs, and provide in-depth analyses with NesTools, which shows that current LLMs still suffer from the complex nested tool learning task.
Tong Zhu 0002, Mengsong Wu, Wenliang Chen
COLING2
2025 Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
abstract
Large language models (LLMs) exhibit remarkable capabilities in understanding and generating natural language. However, these models can inadvertently memorize private information, posing significant privacy risks. This study addresses the challenge of enabling LLMs to protect specific individuals’ private data without the need for complete retraining. We propose RETURN, a Real-world pErsonal daTa UnleaRNing dataset, comprising 2,492 individuals from Wikipedia with associated QA pairs, to evaluate machine unlearning (MU) methods for protecting personal data in a realistic scenario. Additionally, we introduce the Name-Aware Unlearning Framework (NAUF) for Privacy Protection, which enables the model to learn which individuals’ information should be protected without affecting its ability to answer questions related to other unrelated individuals. Our extensive experiments demonstrate that NAUF achieves a state-of-the-art average unlearning score, surpassing the best baseline method by 5.65 points, effectively protecting target individuals’ personal data while maintaining the model’s general capabilities.
Tong Zhu 0002, Chuanyuan Tan, Wenliang Chen
COLING2
2025 CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
abstract
Contrastive Language-Image Pre-training (CLIP) has become a cornerstone in multimodal intelligence.However, recent studies discovered that CLIP can only encode one aspect of the feature space, leading to substantial information loss and indistinctive features.To mitigate this issue, this paper introduces a novel strategy that fine-tunes a series of complementary CLIP models and transforms them into a CLIP-MoE.Specifically, we propose a model-agnostic Diversified Multiplet Upcycling (DMU) framework for CLIP.Instead of training multiple CLIP models from scratch, DMU leverages a pre-trained CLIP and fine-tunes it into a diverse set with highly cost-effective multistage contrastive learning, thus capturing distinct feature subspaces efficiently.To fully exploit these fine-tuned models while minimizing computational overhead, we transform them into a CLIP-MoE, which dynamically activates a subset of CLIP experts, achieving an effective balance between model capacity and computational cost.Comprehensive experiments demonstrate the superior performance of CLIP-MoE across various zero-shot retrieval, zero-shot image classification tasks, and downstream Multimodal Large Language Model (MLLM) benchmarks when used as a vision encoder.Code is available at https: //github.com/OpenSparseLLMs/CLIP-MoE.
Jihai Zhang 0002, Xiaoye Qu, Tong Zhu 0002, Yu Cheng 0001
EMNLP3
2025 Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-Experts
abstract
Tong Zhu, Daize Dong, Xiaoye Qu, Jiacheng Ruan, Wenliang Chen, Yu Cheng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Tong Zhu 0002, Daize Dong, Xiaoye Qu, Jiacheng Ruan, Wenliang Chen, Yu Cheng 0001
NAACL (Long Papers)1
2025 Parallel Task Planning via Model Collaboration
Tong Zhu 0002, Mengsong Wu, Wenliang Chen
NLPCC (4)2
2024 Probing Language Models for Pre-training Data Detection
abstract
Large Language Models (LLMs) have shown their impressive capabilities, while also raising concerns about the data contamination problems due to privacy issues and leakage of benchmark datasets in the pre-training phase.Therefore, it is vital to detect the contamination by checking whether an LLM has been pre-trained on the target texts.Recent studies focus on the generated texts and compute perplexities, which are superficial features and not reliable.In this study, we propose to utilize the probing technique for pre-training data detection by examining the model's internal activations.Our method is simple yet effective and leads to more trustworthy pre-training data detection.Additionally, we propose ArxivMIA, a new challenging benchmark comprising arxiv abstracts from Computer Science and Mathematics categories.Our experiments demonstrate that our method outperforms all the baselines, and achieves state-of-the-art performance on both WikiMIA and ArxivMIA, with additional experiments confirming its efficacy 1 .
Tong Zhu 0002, Chuanyuan Tan, Haonan Lu, Wenliang Chen
ACL (1)2
2024 Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?
abstract
Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu, Xiaoye Qu, Pan Zhou, Yan Bowen, Yu Cheng, Min Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhaochen Su, Juntao Li 0005, Jun Zhang 0069, Tong Zhu 0002, Xiaoye Qu, Pan Zhou 0001, Yan Bowen, Yu Cheng 0001, Min Zhang 0005
ACL (1)4
2024 MoPE: Mixture of Prefix Experts for Zero-Shot Dialogue State Tracking
abstract
Zero-shot dialogue state tracking (DST) transfers knowledge to unseen domains, reducing the cost of annotating new datasets. Previous zero-shot DST models mainly suffer from domain transferring and partial prediction problems. To address these challenges, we propose Mixture of Prefix Experts (MoPE) to establish connections between similar slots in different domains, which strengthens the model transfer performance in unseen domains. Empirical results demonstrate that MoPE-DST achieves the joint goal accuracy of 57.13% on MultiWOZ2.1 and 55.4.
Tianwen Tang, Tong Zhu 0002, Yin Bai, Wenliang Chen
LREC/COLING2
2024 LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-Training
abstract
Mixture-of-Experts (MoE) has gained increasing popularity as a promising framework for scaling up large language models (LLMs).However, training MoE from scratch in a largescale setting still suffers from data-hungry and instability problems.Motivated by this limit, we investigate building MoE models from existing dense large language models.Specifically, based on the well-known LLaMA-2 7B model, we obtain an MoE model by: (1) Expert Construction, which partitions the parameters of original Feed-Forward Networks (FFNs) into multiple experts; (2) Continual pretraining, which further trains the transformed MoE model and additional gate networks.In this paper, we comprehensively explore different methods for expert construction and various data sampling strategies for continual pretraining.After these stages, our LLaMA-MoE models could maintain language abilities and route the input tokens to specific experts with part of the parameters activated.Empirically, by training 200B tokens, LLaMA-MoE-3.5Bmodels significantly outperform dense models that contain similar activation parameters.
Tong Zhu 0002, Xiaoye Qu, Daize Dong, Jiacheng Ruan, Jingqi Tong, Conghui He, Yu Cheng 0001
EMNLP1
2024 ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLMs
Zhaochen Su, Jun Zhang 0069, Xiaoye Qu, Tong Zhu 0002, Yanshu Li, Jiashuo Sun, Juntao Li 0005, Min Zhang 0005, Yu Cheng 0001
NeurIPS4
2024 Seal-Tools: Self-instruct Tool Learning Dataset for Agent Tuning and Detailed Benchmark
Mengsong Wu, Tong Zhu 0002, Chuanyuan Tan, Wenliang Chen
NLPCC (2)2
2023 Mirror: A Universal Framework for Various Information Extraction Tasks
abstract
Tong Zhu, Junfei Ren, Zijian Yu, Mengsong Wu, Guoliang Zhang, Xiaoye Qu, Wenliang Chen, Zhefeng Wang, Baoxing Huai, Min Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Tong Zhu 0002, Junfei Ren, Zijian Yu, Mengsong Wu, Xiaoye Qu, Wenliang Chen, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005
EMNLP1
2023 CED: Catalog Extraction from Documents
Tong Zhu 0002, Zechang Li, Zijian Yu, Junfei Ren, Mengsong Wu, Zhefeng Wang 0001, Baoxing Huai, Pingfu Chao, Wenliang Chen
ICDAR (3)1
2022 Efficient Document-level Event Extraction via Pseudo-Trigger-aware Pruned Complete Graph
abstract
Most previous studies of document-level event extraction mainly focus on building argument chains in an autoregressive way, which achieves a certain success but is inefficient in both training and inference. In contrast to the previous studies, we propose a fast and lightweight model named as PTPCG. In our model, we design a novel strategy for event argument combination together with a non-autoregressive decoding algorithm via pruned complete graphs, which are constructed under the guidance of the automatically selected pseudo triggers. Compared to the previous systems, our system achieves competitive results with 19.8% of parameters and much lower resource consumption, taking only 3.8% GPU hours for training and up to 8.5 times faster for inference. Besides, our model shows superior compatibility for the datasets with (or without) triggers and the pseudo triggers can be the supplements for annotated triggers to make further improvements. Codes are available at https://github.com/Spico197/DocEE .
Tong Zhu 0002, Xiaoye Qu, Wenliang Chen, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan, Min Zhang 0005
IJCAI1
2020 Improving Relation Extraction with Relational Paraphrase Sentences
abstract
Supervised models for Relation Extraction (RE) typically require human-annotated training data.Due to the limited size, the human-annotated data is usually incapable of covering diverse relation expressions, which could limit the performance of RE.To increase the coverage of relation expressions, we may enlarge the labeled data by hiring annotators or applying Distant Supervision (DS).However, the human-annotated data is costly and non-scalable while the distantly supervised data contains many noises.In this paper, we propose an alternative approach to improve RE systems via enriching diverse expressions by relational paraphrase sentences.Based on an existing labeled data, we first automatically build a task-specific paraphrase data.Then, we propose a novel model to learn the information of diverse relation expressions.In our model, we try to capture this information on the paraphrases via a joint learning framework.Finally, we conduct experiments on a widely used dataset and the experimental results show that our approach is effective to improve the performance on relation extraction, even compared with a strong baseline.
Tong Zhu 0002, Wenliang Chen, Wei Zhang 0027, Min Zhang 0005
COLING2
2020 Towards Accurate and Consistent Evaluation: A Dataset for Distantly-Supervised Relation Extraction
abstract
In recent years, distantly-supervised relation extraction has achieved a certain success by using deep neural networks.Distant Supervision (DS) can automatically generate large-scale annotated data by aligning entity pairs from Knowledge Bases (KB) to sentences.However, these DSgenerated datasets inevitably have wrong labels that result in incorrect evaluation scores during testing, which may mislead the researchers.To solve this problem, we build a new dataset NYT-H, where we use the DS-generated data as training data and hire annotators to label test data.Compared with the previous datasets, NYT-H has a much larger test set and then we can perform more accurate and consistent evaluation.Finally, we present the experimental results of several widely used systems on NYT-H.The experimental results show that the ranking lists of the comparison systems on the DS-labelled test data and human-annotated test data are different.This indicates that our human-annotated data is necessary for evaluation of distantly-supervised relation extraction.
Tong Zhu 0002, Haitao Wang 0019, Xiabing Zhou, Wenliang Chen, Wei Zhang 0027, Min Zhang 0005
COLING1