Zichun Yu

dblp:274/8681 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-4423-9657ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 23% Efficient and distributed learning · 20% Reinforcement learning · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
synthetic data generation
1.922026
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search · ACL (1) 2026
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning · ICLR 2025
Machine learning › Reinforcement learning
LLM agent training
1.012026
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search · ACL (1) 2026
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.012026
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search · ACL (1) 2026
Knowledge, reasoning and agents › Multi-agent systems › multi-agent learning
multi-agent training
1.012026
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search · ACL (1) 2026
Machine learning › Trustworthy machine learning › Data-centric AI
data influence
0.912025
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning · ICLR 2025
Natural language and speech › Language models and text generation
instruction tuning
0.912025
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning · ICLR 2025
Natural language and speech › Language models and text generation › large language model training
pretraining data selection
0.912025
Group-Level Data Selection for Efficient Pretraining · NeurIPS 2025
Machine learning › Generative modeling
synthetic training data
0.912025
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning · ICLR 2025
Machine learning › Efficient and distributed learning
data selection
0.812024
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models · NeurIPS 2024
Machine learning › Efficient and distributed learning › efficient training
efficient pre-training
0.812024
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models · NeurIPS 2024
Machine learning › Representation and self-supervised learning
pre-training
0.812024
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models · NeurIPS 2024
Information retrieval
retrieval augmentation
0.712023
Augmentation-Adapted Retriever Improves Generalization of Language Models as Generic Plug-In · ACL (1) 2023
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.312026
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search · ACL (1) 2026
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.212023
Augmentation-Adapted Retriever Improves Generalization of Language Models as Generic Plug-In · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

data influence model · 1.6source LM preference learning · 1.3augmentation-adapted retriever · 1.3reinforcement learning · 1.0monte carlo tree search · 1.0influence score estimation · 1.0direct preference optimization · 0.9data influence estimation · 0.9clustering · 0.9oracle data influence · 0.8
YearPublicationVenuePosition
2026 Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search
abstract
Large Language Model (LLM) based multiagent systems (MAS) show strong potential for tackling complex tasks through collaborative intelligence.Monte Carlo Tree Search (MCTS) based methods provide promising approaches for enhancing MAS self-training by generating synthetic data, using Q-values to estimate agent contributions.However, relying solely on Q-values may misalign with the goal of selecting data most beneficial for MAS improvement.To address this discrepancy, we propose Data Influence-oriented Tree Search (DITS), a novel framework that incorporates influence scores to guide both tree search and data selection in data synthesis.By leveraging influence scores, we effectively identify the most impactful data for MAS improvement, thereby enhancing model performance.Furthermore, we derive a novel influence score estimation method tailored for non-differentiable metrics, significantly reducing computational overhead by calculating performance changes on the validation set.Extensive experiments on three different multi-agent tasks demonstrate the robustness and effectiveness of the proposed methods.Notably, our findings reveal that allocating more resources to estimate influence scores, rather than Q-values, during data synthesis can more effectively and efficiently enhance model training.The code is available at https://github.com/swt-user/DITS.
Wentao Shi 0002, Zichun Yu, Fuli Feng, Xiangnan He 0001, Chenyan Xiong
ACL (1)2
2025 Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning
abstract
Synthetic data has been widely used to train large language models, but their generative nature inevitably introduces noisy, non-informative, and misleading learning signals. In this paper, we propose Montessori-Instruct, a novel data synthesis framework that tailors the data synthesis ability of the teacher language model toward the student language model's learning process. Specifically, we utilize local data influence of synthetic training data points on students to characterize students' learning preferences. Then, we train the teacher model with Direct Preference Optimization (DPO) to generate synthetic data tailored toward student learning preferences. Experiments with Llama3-8B-Instruct (teacher) and Llama3-8B (student) on Alpaca Eval and MT-Bench demonstrate that Montessori-Instruct significantly outperforms standard synthesis methods by 18.35\% and 46.24\% relatively. Our method also beats data synthesized by a stronger teacher model, GPT-4o. Further analysis confirms the benefits of teacher's learning to generate more influential training data in the student's improved learning, the advantages of local data influence in accurately measuring student preferences, and the robustness of Montessori-Instruct across different student models. Our code and data are open-sourced at https://github.com/cxcscmu/Montessori-Instruct.
Xiaochuan Li 0003, Zichun Yu, Chenyan Xiong
ICLR2
2025 Group-Level Data Selection for Efficient Pretraining
abstract
The efficiency and quality of language model pretraining are largely determined by the way pretraining data are selected. In this paper, we introduce *Group-MATES*, an efficient group-level data selection approach to optimize the speed-quality frontier of language model pretraining. Specifically, Group-MATES parameterizes costly group-level selection with a relational data influence model. To train this model, we sample training trajectories of the language model and collect oracle data influences alongside. The relational data influence model approximates the oracle data influence by weighting individual influence with relationships among training data. To enable efficient selection with our relational data influence model, we partition the dataset into small clusters using relationship weights and select data within each cluster independently. Experiments on DCLM 400M-4x, 1B-1x, and 3B-1x show that Group-MATES achieves 3.5\%-9.4\% relative performance gains over random selection across 22 downstream tasks, nearly doubling the improvements achieved by state-of-the-art individual data selection baselines. Furthermore, Group-MATES reduces the number of tokens required to reach a certain downstream performance by up to 1.75x, substantially elevating the speed-quality frontier. Further analyses highlight the critical role of relationship weights in the relational data influence model and the effectiveness of our cluster-based inference. Our code is open-sourced at https://github.com/facebookresearch/Group-MATES.
Zichun Yu, Arnold Overwijk, Scott Yih, Chenyan Xiong
NeurIPS1
2024 MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
abstract
Pretraining data selection has the potential to improve language model pretraining efficiency by utilizing higher-quality data from massive web data corpora. Current data selection methods, which rely on either hand-crafted rules or larger reference models, are conducted statically and do not capture the evolving data preferences during pretraining. In this paper, we introduce *model-aware data selection with data influence models (MATES)*, where a data influence model continuously adapts to the evolving data preferences of the pretraining model and then selects the data most effective for the current pretraining progress. Specifically, we collect oracle data influence by locally probing the pretraining model and fine-tune a small data influence model to approximate it accurately. The data influence model then predicts data influence over the whole pretraining corpus and selects the most influential data for the next pretraining stage. Experiments of pretraining 410M and 1B models on the C4 dataset demonstrate that MATES significantly outperforms random data selection on extensive downstream tasks. It doubles the gains achieved by the state-of-the-art data selection approach that leverages larger reference models and reduces the total FLOPs required to reach certain performances by half. Further analyses validate the effectiveness of the locally probed oracle data influence and the approximation with data influence models. Our code is open-sourced at https://github.com/cxcscmu/MATES.
Zichun Yu, Spandan Das, Chenyan Xiong
NeurIPS1
2023 Augmentation-Adapted Retriever Improves Generalization of Language Models as Generic Plug-In
abstract
Retrieval augmentation can aid language models (LMs) in knowledge-intensive tasks by supplying them with external information.Prior works on retrieval augmentation usually jointly fine-tune the retriever and the LM, making them closely coupled.In this paper, we explore the scheme of generic retrieval plug-in: the retriever is to assist target LMs that may not be known beforehand or are unable to be fine-tuned together.To retrieve useful documents for unseen target LMs, we propose augmentation-adapted retriever (AAR), which learns LM's preferences obtained from a known source LM.Experiments on the MMLU and PopQA datasets demonstrate that our AAR trained with a small source LM is able to significantly improve the zero-shot generalization of larger target LMs ranging from 250M Flan-T5 to 175B InstructGPT.Further analysis indicates that the preferences of different LMs overlap, enabling AAR trained with a single source LM to serve as a generic plug-in for various target LMs.Our code is open-sourced at https://github.com/OpenMatch/Augmentation- Adapted-Retriever.
Zichun Yu, Chenyan Xiong, Shi Yu 0001, Zhiyuan Liu 0001
ACL (1)1
2022 Automatic Label Sequence Generation for Prompting Sequence-to-sequence Models
abstract
Prompting, which casts downstream applications as language modeling tasks, has shown to be sample efficient compared to standard fine-tuning with pre-trained models. However, one pitfall of prompting is the need of manually-designed patterns, whose outcome can be unintuitive and requires large validation sets to tune. To tackle the challenge, we propose AutoSeq, a fully automatic prompting method: (1) We adopt natural language prompts on sequence-to-sequence models, enabling free-form generation and larger label search space; (2) We propose label sequences – phrases with indefinite lengths to verbalize the labels – which eliminate the need of manual templates and are more expressive than single label words; (3) We use beam search to automatically generate a large amount of label sequence candidates and propose contrastive re-ranking to get the best combinations. AutoSeq significantly outperforms other no-manual-design methods, such as soft prompt tuning, adapter tuning, and automatic search on single label words; the generated label sequences are even better than curated manual ones on a variety of tasks. Our method reveals the potential of sequence-to-sequence models in few-shot learning and sheds light on a path to generic and automatic prompting. The source code of this paper can be obtained from https://github.com/thunlp/Seq2Seq-Prompt.
Zichun Yu, Tianyu Gao 0001, Zhengyan Zhang, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016
COLING1