Guangtao Zeng

dblp:264/9714 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 35% Generative modeling · 18% Trustworthy machine learning · 13%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 21 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search · ICML 2025
Machine learning › Representation and self-supervised learning › pre-training
data mixture optimization
0.912025
RegMix: Data Mixture as Regression for Language Model Pre-training · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.912025
Scaling up Masked Diffusion Models on Text · ICLR 2025
Natural language and speech › Language models and text generation
language modeling
0.912025
Scaling up Masked Diffusion Models on Text · ICLR 2025
Natural language and speech › Language models and text generation › large language model training
language model pretraining
0.912025
RegMix: Data Mixture as Regression for Language Model Pre-training · ICLR 2025
Machine learning › Generative modeling › diffusion model › discrete diffusion model
masked diffusion language model
0.912025
Scaling up Masked Diffusion Models on Text · ICLR 2025
Machine learning › Generative modeling › diffusion model › discrete diffusion model
masked diffusion model
0.912025
Scaling up Masked Diffusion Models on Text · ICLR 2025
Natural language and speech › Language models and text generation › large language model training
post-training
0.912025
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search · ICML 2025
Natural language and speech › Language models and text generation
self-improvement
0.912025
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search · ICML 2025
Program synthesis and code generation
code generation with language models
0.912025
EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning · ICML 2025
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.712023
Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models · EMNLP 2023
Machine learning › Transfer learning and domain adaptation
parameter-efficient transfer learning
0.712023
One Network, Many Masks: Towards More Parameter-Efficient Transfer Learning · ACL (1) 2023
Machine learning › Trustworthy machine learning › model security
model intellectual property protection
0.612022
Unsupervised Non-transferable Text Classification · EMNLP 2022
Machine learning › Trustworthy machine learning › model security
non-transferable learning
0.612022
Unsupervised Non-transferable Text Classification · EMNLP 2022
Natural language and speech › Information extraction and text analysis
text classification
0.612022
Unsupervised Non-transferable Text Classification · EMNLP 2022
Natural language and speech › Question answering and dialogue systems
medical dialogue systems
0.412020
MedDialog: Large-scale Medical Dialogue Datasets · EMNLP (1) 2020
Robotics › Motion planning and robot control
robot learning
0.312026
Tailored Primitive Initialization is the Secret Key to Reinforcement Learning · ACL (1) 2026
Program synthesis and code generation › code generation with language models
fine-tuning for code generation
0.312025
EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning · ICML 2025
Natural language and speech › Language models and text generation › large language model reasoning
multi-step reasoning
0.212023
Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models · EMNLP 2023
Natural language and speech › Language models and text generation › large language model › large language model adaptation
pre-trained language model fine-tuning
0.212023
One Network, Many Masks: Towards More Parameter-Efficient Transfer Learning · ACL (1) 2023
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.112020
MedDialog: Large-scale Medical Dialogue Datasets · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.9primitive initialization · 1.0unsupervised classifier-free guidance · 0.9self-reflection · 0.9self-exploration · 0.9scaling laws · 0.9scaling law analysis · 0.9regression · 0.9masked diffusion · 0.9large language model fine-tuning · 0.9execution-based code selection · 0.9chain-of-action-thought · 0.9
YearPublicationVenuePosition
2026 Tailored Primitive Initialization is the Secret Key to Reinforcement Learning
abstract
Yihang Yao, Guangtao Zeng, Raina Wu, Yang Zhang, Ding Zhao, Zhang-Wei Hong, Chuang Gan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yihang Yao, Guangtao Zeng, Raina Wu, Yang Zhang 0001, Ding Zhao, Zhang-Wei Hong, Chuang Gan 0001
ACL (1)2
2025 RegMix: Data Mixture as Regression for Language Model Pre-training
abstract
The data mixture for large language model pre-training significantly impacts performance, yet how to determine an effective mixture remains unclear. We propose RegMix to automatically identify a high-performing data mixture by formulating it as a regression task. RegMix trains many small models on diverse data mixtures, uses regression to predict performance of unseen mixtures, and applies the best predicted mixture to train a large-scale model with orders of magnitude more compute. To empirically validate RegMix, we train 512 models with 1M parameters for 1B tokens to fit the regression model and predict the best data mixture. Using this mixture we train a 1B parameter model for 25B tokens (i.e. 1000× larger and 25× longer) which we find performs best among 64 candidate 1B parameter models with other mixtures. Furthermore, RegMix consistently outperforms human selection in experiments involving models up to 7B models trained on 100B tokens, while matching or exceeding DoReMi using just 10% of the computational resources. Our experiments also show that (1) Data mixtures significantly impact performance; (2) Web corpora rather than data perceived as high-quality like Wikipedia have the strongest positive correlation with downstream performance; (3) Domains interact in complex ways often contradicting common sense, thus automatic approaches like RegMix are needed; (4) Data mixture effects transcend scaling laws. Our code is available at https://github.com/sail-sg/regmix.
Qian Liu 0033, Xiaosen Zheng, Niklas Muennighoff, Guangtao Zeng, Longxu Dou, Tianyu Pang, Jing Jiang 0001
ICLR4
2025 Scaling up Masked Diffusion Models on Text
abstract
Masked diffusion models (MDMs) have shown promise in language modeling, yet their scalability and effectiveness in core language tasks, such as text generation and language understanding, remain underexplored. This paper establishes the first scaling law for MDMs, demonstrating a scaling rate comparable to autoregressive models (ARMs) and a relatively small compute gap. Motivated by their scalability, we train a family of MDMs with up to 1.1 billion (B) parameters to systematically evaluate their performance against ARMs of comparable or larger sizes. Fully leveraging the probabilistic formulation of MDMs, we propose a simple yet effective *unsupervised classifier-free guidance* that effectively exploits large-scale unpaired data, boosting performance for conditional inference. In language understanding, the 1.1B MDM outperforms the 1.1B TinyLlama model trained on the same data across four of eight zero-shot benchmarks. Notably, it achieves competitive math reasoning ability with the 7B Llama-2 model on the GSM8K dataset. In text generation, MDMs with 16 times more pre-training time offer a flexible trade-off against ARMs with the accelerated sampling technique KV-Cache: MDMs match ARMs in performance while being 1.4 times faster during sampling. Moreover, MDMs address challenging tasks for ARMs by effectively handling bidirectional reasoning and adapting to temporal shifts in data. Notably, a 1.1B MDM breaks the *reverse curse* encountered by much larger ARMs with significantly more data and computation, such as 13B Llama-2 and 175B GPT-3. Our code is available at https://github.com/ML-GSAI/SMDM.
Shen Nie, Fengqi Zhu, Tianyu Pang, Qian Liu 0033, Guangtao Zeng, Chongxuan Li
ICLR6
2025 EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning
abstract
As large language models (LLMs) play an increasingly important role in code generation, enhancing both correctness and efficiency has become crucial. Current methods primarily focus on correctness, often overlooking efficiency. To address this gap, we introduce SWIFTCODE to improve both aspects by fine-tuning LLMs on a high-quality dataset comprising correct and efficient code samples. Our methodology involves leveraging multiple LLMs to generate diverse candidate code solutions for various tasks across different programming languages. We then evaluate these solutions by directly measuring their execution time and memory usage through local execution. The code solution with the lowest execution time and memory consumption is selected as the final output for each task. Experimental results demonstrate significant improvements when fine-tuning with SWIFTCODE. For instance, Qwen2.5-Coder-7B-Instruct’s pass@1 score increases from 44.8% to 57.7%, while the average execution time for correct tasks decreases by 48.4%. SWIFTCODE offers a scalable and effective solution for advancing AI-driven code generation, benefiting both software development and computational problem-solving.
Dong Huang 0005, Guangtao Zeng, Jianbo Dai, Meng Luo 0010, Han Weng, Yuhao Qing, Heming Cui, Zhijiang Guo, Jie Zhang 0050
ICML2
2025 Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
abstract
Large language models (LLMs) have demonstrated remarkable reasoning capabilities across diverse domains. Recent studies have shown that increasing test-time computation enhances LLMs' reasoning capabilities. This typically involves extensive sampling at inference time guided by an external LLM verifier, resulting in a two-player system. Despite external guidance, the effectiveness of this system demonstrates the potential of a single LLM to tackle complex tasks. Thus, we pose a new research problem: *Can we internalize the searching capabilities to fundamentally enhance the reasoning abilities of a single LLM?* This work explores an orthogonal direction focusing on post-training LLMs for autoregressive searching (*i.e.,* an extended reasoning process with self-reflection and self-exploration of new strategies). To achieve this, we propose the Chain-of-Action-Thought (COAT) reasoning and a two-stage training paradigm: 1) a small-scale format tuning stage to internalize the COAT reasoning format and 2) a large-scale self-improvement stage leveraging reinforcement learning. Our approach results in Satori, a 7B LLM trained on open-source models and data. Extensive empirical evaluations demonstrate that Satori achieves state-of-the-art performance on mathematical reasoning benchmarks while exhibits strong generalization to out-of-domain tasks. Code, data, and models are fully open-sourced.
Maohao Shen, Guangtao Zeng, Zhenting Qi, Zhang-Wei Hong, Zhenfang Chen, Gregory W. Wornell, Subhro Das, David D. Cox, Chuang Gan 0001
ICML2
2023 One Network, Many Masks: Towards More Parameter-Efficient Transfer Learning
abstract
Fine-tuning pre-trained language models for multiple tasks tends to be expensive in terms of storage.To mitigate this, parameter-efficient transfer learning (PETL) methods have been proposed to address this issue, but they still require a significant number of parameters and storage when being applied to broader ranges of tasks.To achieve even greater storage reduction, we propose PROPETL, a novel method that enables efficient sharing of a single PETL module which we call prototype network (e.g., adapter, LoRA, and prefix-tuning) across layers and tasks.We then learn binary masks to select different sub-networks from the shared prototype network and apply them as PETL modules into different layers.We find that the binary masks can determine crucial information from the network, which is often ignored in previous studies.Our work can also be seen as a type of pruning method, where we find that overparameterization also exists in the seemingly small PETL modules.We evaluate PROPETL on various downstream tasks and show that it can outperform other PETL methods with approximately 10% of the parameter storage required by the latter. 1
Guangtao Zeng, Peiyuan Zhang, Wei Lu 0011
ACL (1)1
2023 Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models
abstract
Yifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, Mrinmaya Sachan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Jiaoda Li, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, Mrinmaya Sachan
EMNLP6
2022 Unsupervised Non-transferable Text Classification
abstract
Training a good deep learning model requires substantial data and computing resources, which makes the resulting neural model a valuable intellectual property.To prevent the neural network from being undesirably exploited, non-transferable learning has been proposed to reduce the model generalization ability in specific target domains.However, existing approaches require labeled data for the target domain which can be difficult to obtain.Furthermore, they do not have the mechanism to still recover the model's ability to access the target domain.In this paper, we propose a novel unsupervised non-transferable learning method for the text classification task that does not require annotated target domain data.We further introduce a secret key component in our approach for recovering the access to the target domain, where we design both an explicit and an implicit method for doing so.Extensive experiments demonstrate the effectiveness of our approach.
Guangtao Zeng, Wei Lu 0011
EMNLP1
2020 MedDialog: Large-scale Medical Dialogue Datasets
abstract
Guangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang, Sicheng Wang, Ruisi Zhang, Meng Zhou, Jiaqi Zeng, Xiangyu Dong, Ruoyu Zhang, Hongchao Fang, Penghui Zhu, Shu Chen, Pengtao Xie. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Guangtao Zeng, Wenmian Yang, Zeqian Ju, Sicheng Wang 0001, Ruisi Zhang, Jiaqi Zeng, Xiangyu Dong 0002, Hongchao Fang, Penghui Zhu, Pengtao Xie
EMNLP (1)1
2020 BETA-Rec: Build, Evaluate and Tune Automated Recommender Systems
abstract
The field of recommender systems has rapidly evolved over the last few years, with significant advances made due to the in-flux of deep learning techniques. However, as a result of this rapid progress, escalating barriers-to-entry for new researchers is emerging. In particular, state-of-the-art approaches have fragmented into a large number of code-bases, often requiring different input formats, pre-processing stages and evaluating with different metric packages. Hence, it is time-consuming for new researchers to reach the point of having both an effective baseline set and a sound comparative environment. As a step towards elevating this problem, we have developed BETA-Rec, an open source project for Building, Evaluating and Tuning Automated Recommender Systems. BETA-Rec aims to provide a practical data toolkit for building end-to-end recommendation systems in a standardized way. It provides means for dataset preparation and splitting using common strategies, a generalized model engine for implementing recommender models using Pytorch with 9 models available out-of-the-box, as well as a unified training, validation, tuning and testing pipeline. Furthermore, BETA-Rec is designed to be both modular and extensible, enabling new models to be quickly added to the framework. It is deployable in a wide range of environments via pre-built docker containers and supports distributed parameter tuning using Ray. In this demo, we will illustrate the deployment and use of BETA-Rec for researchers and practitioners on a number of standard recommendation datasets. The source code of the project is available at github: https://github.com/beta-team/beta-recsys.
Zaiqiao Meng, Richard McCreadie, Craig Macdonald, Iadh Ounis, Siwei Liu 0001, Yaxiong Wu 0001, Xi Wang 0012, Shangsong Liang, Yucheng Liang, Guangtao Zeng, Junhua Liang, Qiang Zhang 0026
RecSys10