Pengyu Wang 0006

dblp:14/3832-6 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Case2Code: Scalable Synthetic Data for Code Generation
abstract
Large Language Models (LLMs) have shown outstanding breakthroughs in code generation. Recent work improves code LLMs by training on synthetic data generated by some powerful LLMs, which can be challenging to scale due to the dependence on a teacher model and high generation costs. In this paper, we focus on synthesizing code data at scale and propose a Case2Code task by exploiting the expressiveness and correctness of programs. Case2Code is an inductive inference task that aims to infer underlying code implementations by observing input-output examples or program behaviors, By incorporating LLMs to generate program inputs, and executing the program with these inputs to obtain the program outputs, we can synthesize diverse and high-quality Case2Code data at scale for training and evaluating code LLMs. Experimental results show that case-to-code induction is challenging for current representative LLMs if they are untrained. Models trained with Case2Code improve performance not only on distribution case-to-code induction but also various coding-generation tasks, demonstrating the great potential of large-scale synthetic data and inductive learning.
Yunfan Shao, Linyang Li, Yichuan Ma, Peiji Li, Demin Song, Qinyuan Cheng, Pengyu Wang 0006, Qipeng Guo, Hang Yan 0001, Xipeng Qiu, Xuanjing Huang 0001, Dahua Lin
COLING9
2025 Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
abstract
Wei Wang, Zhaowei Li, Qi Xu, Linfeng Li, YiQing Cai, Botian Jiang, Hang Song, Xingcan Hu, Pengyu Wang, Li Xiao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Wei Wang 0378, Yiqing Cai, Botian Jiang, Xingcan Hu, Pengyu Wang 0006, Li Xiao 0002
EMNLP9
2025 UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
abstract
Pengyu Wang, Shaojun Zhou, Chenkun Tan, Xinghao Wang, Wei Huang, Zhen Ye, Zhaowei Li, Botian Jiang, Dong Zhang, Xipeng Qiu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Pengyu Wang 0006, Shaojun Zhou, Chenkun Tan, Zhen Ye 0006, Botian Jiang, Xipeng Qiu
EMNLP1
2025 BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments
abstract
Large language models (LLMs) have revolutionized numerous applications, yet their deployment remains challenged by memory constraints on local devices. While scaling laws have enhanced LLM capabilities, the primary bottleneck has shifted from $\textit{capability}$ to $\textit{availability}$, emphasizing the need for efficient memory management. Traditional compression methods, such as quantization, often require predefined compression ratios and separate compression processes for each setting, complicating deployment in variable memory environments. In this paper, we introduce $\textbf{BitStack}$, a novel, training-free weight compression approach that enables megabyte-level trade-offs between memory usage and model performance. By leveraging weight decomposition, BitStack can dynamically adjust the model size with minimal transmission between running memory and storage devices. Our approach iteratively decomposes weight matrices while considering the significance of each parameter, resulting in an approximately 1-bit per parameter residual block in each decomposition iteration. These blocks are sorted and stacked in storage as basic transmission units, with different quantities loaded based on current memory availability. Extensive experiments across a wide range of tasks demonstrate that, despite offering fine-grained size control, BitStack consistently matches or surpasses strong quantization baselines, particularly at extreme compression ratios. To the best of our knowledge, this is the first decomposition-based method that effectively bridges the gap to practical compression techniques like quantization. Code is available at https://github.com/xinghaow99/BitStack.
Pengyu Wang 0006, Bo Wang 0084, Yunhua Zhou, Xipeng Qiu
ICLR2
2024 DenoSent: A Denoising Objective for Self-Supervised Sentence Representation Learning
abstract
Contrastive-learning-based methods have dominated sentence representation learning. These methods regularize the representation space by pulling similar sentence representations closer and pushing away the dissimilar ones and have been proven effective in various NLP tasks, e.g., semantic textual similarity (STS) tasks. However, it is challenging for these methods to learn fine-grained semantics as they only learn from the inter-sentence perspective, i.e., their supervision signal comes from the relationship between data samples. In this work, we propose a novel denoising objective that inherits from another perspective, i.e., the intra-sentence perspective. By introducing both discrete and continuous noise, we generate noisy sentences and then train our model to restore them to their original form. Our empirical evaluations demonstrate that this approach delivers competitive results on both semantic textual similarity (STS) and a wide range of transfer tasks, standing up well in comparison to contrastive-learning-based methods. Notably, the proposed intra-sentence denoising objective complements existing inter-sentence contrastive methodologies and can be integrated with them to further enhance performance. Our code is available at https://github.com/xinghaow99/DenoSent.
Junliang He, Pengyu Wang 0006, Yunhua Zhou, Tianxiang Sun, Xipeng Qiu
AAAI3
2024 The Open-World Lottery Ticket Hypothesis for OOD Intent Classification
abstract
Most existing methods of Out-of-Domain (OOD) intent classification rely on extensive auxiliary OOD corpora or specific training paradigms. However, they are underdeveloped in the underlying principle that the models should have differentiated confidence in In- and Out-of-domain intent. In this work, we shed light on the fundamental cause of model overconfidence on OOD and demonstrate that calibrated subnetworks can be uncovered by pruning the overparameterized model. Calibrated confidence provided by the subnetwork can better distinguish In- and Out-of-domain, which can be a benefit for almost all post hoc methods. In addition to bringing fundamental insights, we also extend the Lottery Ticket Hypothesis to open-world scenarios. We conduct extensive experiments on four real-world datasets to demonstrate our approach can establish consistent improvements compared with a suite of competitive baselines.
Yunhua Zhou, Pengyu Wang 0006, Peiju Liu, Yuxin Wang 0005, Xipeng Qiu
LREC/COLING2
2024 InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
abstract
Pengyu Wang, Dong Zhang, Linyang Li, Chenkun Tan, Xinghao Wang, Mozhi Zhang, Ke Ren, Botian Jiang, Xipeng Qiu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Pengyu Wang 0006, Linyang Li, Chenkun Tan, Mozhi Zhang, Botian Jiang, Xipeng Qiu
EMNLP1
2024 SpeechAlign: Aligning Speech Generation to Human Preferences
abstract
Speech language models have significantly advanced in generating realistic speech, with neural codec language models standing out. However, the integration of preference optimization to align speech outputs to human preferences is often neglected. This paper addresses this gap by first analyzing the distribution gap in codec language models, highlighting how it leads to discrepancies between the training and inference phases, which negatively affects performance. Then we explore leveraging preference optimization to bridge the distribution gap. We introduce SpeechAlign, an iterative self-improvement strategy that aligns speech language models to human preferences. SpeechAlign involves constructing a preference codec dataset contrasting golden codec tokens against synthetic tokens, followed by preference optimization to improve the codec language model. This cycle of improvement is carried out iteratively to steadily convert weak models to strong ones. Through both subjective and objective evaluations, we show that SpeechAlign can bridge the distribution gap and facilitating continuous self-improvement of the speech language model. Moreover, SpeechAlign exhibits robust generalization capabilities and works for smaller models. Demos are available at https://0nutation.github.io/SpeechAlign.github.io/.
Pengyu Wang 0006, Yaqian Zhou 0001, Xipeng Qiu
NeurIPS5
2023 Two Birds One Stone: Dynamic Ensemble for OOD Intent Classification
abstract
Out-of-domain (OOD) intent classification is an active field of natural language understanding, which is of great practical significance for intelligent devices such as the Task-Oriented Dialogue System.It mainly contains two challenges: it requires the model to know what it knows and what it does not know.This paper investigates "overthinking" in the openworld scenario and its impact on OOD intent classification.Inspired by this, we propose a two-birds-one-stone method, which allows the model to decide whether to make a decision on OOD classification early during inference and can ensure accuracy and accelerate inference.At the same time, to adapt to the behavior of dynamic inference, we also propose a training method based on ensemble methods.In addition to bringing certain theoretical insights, we also conduct detailed experiments on three real-world intent datasets.Compared with the previous baselines, our method can not only improve inference speed, but also achieve significant performance improvements.Code is publicly available.
Yunhua Zhou, Jianqiang Yang, Pengyu Wang 0006, Xipeng Qiu
ACL (1)3
2023 SeqXGPT: Sentence-Level AI-Generated Text Detection
abstract
Widely applied large language models (LLMs) can generate human-like content, raising concerns about the abuse of LLMs.Therefore, it is important to build strong AI-generated text (AIGT) detectors.Current works only consider document-level AIGT detection, therefore, in this paper, we first introduce a sentence-level detection challenge by synthesizing a dataset that contains documents that are polished with LLMs, that is, the documents contain sentences written by humans and sentences modified by LLMs.Then we propose Sequence X (Check) GPT, a novel method that utilizes log probability lists from white-box LLMs as features for sentence-level AIGT detection.These features are composed like waves in speech processing and cannot be studied by LLMs.Therefore, we build SeqXGPT based on convolution and self-attention networks.We test it in both sentence and document-level detection challenges.Experimental results show that previous methods struggle in solving sentence-level AIGT detection, while our method not only significantly surpasses baseline methods in both sentence and document-level detection challenges but also exhibits strong generalization capabilities.1
Pengyu Wang 0006, Linyang Li, Botian Jiang, Xipeng Qiu
EMNLP1