Pengyu Cheng

dblp:223/6048 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DesireKV: Decoupling Sensitivity and Importance for Reasoning-Aware KV Cache Compression
Pengyu Cheng, Xiaofeng Hou, Jiacheng Liu 0001
AAAI1
2026 AdaReason: Progressive Training of Multi-LoRA Adapters for Budget-Adaptive Language Reasoning Models
abstract
Large reasoning models (LRMs) have demonstrated remarkable capabilities in solving complex problems through extended chain-of-thought reasoning. However, existing approaches face a fundamental trade-off between computational efficiency and reasoning accuracy. Current methods either lack support for user-specified computational budgets or require maintaining multiple independent models, leading to significant resource overhead. In this paper, we present AdaReason, a unified framework that trains a single base model to support arbitrary user-defined computational budgets through dynamic adapter composition. Our approach introduces three key innovations: (1) a length-adaptive step reward function that stabilizes training across diverse budget constraints, (2) a progressive training strategy that gradually tightens computational bounds while maintaining model performance, and (3) a runtime adapter merging mechanism that dynamically interpolates between different computational preferences. Unlike existing methods that suffer from training instability in large context windows, AdaReason achieves stable convergence through careful reward shaping and progressive constraint tightening. Additionally, we provide a rigorous theoretical analysis, establishing a performance bound for our merged model. Experiments on different reasoning benchmarks demonstrate that AdaReason establishes a new state-of-the-art in the performance-efficiency trade-off and enables flexible runtime budget adaptation.
Pengyu Cheng, Xiaofeng Hou, Jiacheng Liu 0001
AAAI3
2026 Adaptive Spatial and Temporal Redundancy Optimization for Efficient Reasoning in Large Language Models
abstract
Large Language Models (LLMs) have achieved exceptional performance in complex reasoning via Chain-of-Thought (CoT), yet the associated computational costs remain prohibitive. CoT reasoning contains significant untapped efficiency potential across two dimensions: temporal redundancy, where reasoning steps may be unnecessary, and spatial redundancy, where computations can be performed at reduced precision. While current optimization techniques often necessitate resource-intensive fine-tuning or data curation, we introduce ASTRO (Adaptive Spatial and Temporal Redundancy Optimization), a training-free framework that simultaneously addresses both dimensions. ASTRO leverages Dewey’s reflective thinking model to segment reasoning phases, applying a progressive precision reduction strategy coupled with an entropy-based confidence mechanism for adaptive termination. Empirical results across diverse reasoning benchmarks demonstrate that ASTRO achieves up to an 11.3 \times efficiency gain without compromising accuracy, highlighting the advantages of holistic multi-dimensional redundancy management over isolated optimization methods.
Pengyu Cheng, Qiyuan Zhu, Hao Gu 0001, Ruijie Shen, Xiaofeng Hou, Sirui Han, Jiacheng Liu 0001
ACL (1)2
2026 MARCH: Multi-Agent Reinforced Check for Hallucination
abstract
Zhuo Li, Yupeng Zhang, Pengyu Cheng, Jiajun Song, Mengyu Zhou, Hao Li, Shujie Hu, Yu Qin, Erchao.zec, Xiaoxi Jiang, Guanjunjiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pengyu Cheng, Jiajun Song, Mengyu Zhou, Shujie Hu, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang
ACL (1)3
2025 Outlier-Aware Model Merging for Efficient Multitask Inference
abstract
Model merging techniques aim to consolidate multiple fine-tuned models into a single unified model, reducing both storage and computational overhead while retaining task-specific performance. However, existing methods face several limitations: monotonous compression techniques that fail to account for task-specific weight distribution characteristics, weight-magnitude-based compression that fails to consider functional importance revealed by activation patterns, and non-adaptive allocation strategies that ignores task-specific layer importance. To overcome these challenges, we propose OA-Merge, a novel Outlier-Aware Model Merging framework that leverages task activation outliers to enable adaptive compression and resource allocation across tasks. OA-Merge comprises three key components: (1) dynamic hybrid decomposition technique that formulates task vectors as tailored combinations of low-rank and sparse components adapted to task-specific statistical distributions, (2) activation-informed compression methodology that incorporates task-specific activation statistics to prioritize functionally important weights, and (3) task-related allocation that optimizes the distribution of compression resources according to layer-specific importance metrics derived from activation outlier analysis. These hybrid outlier-aware strategies adapt dynamically to each task's intrinsic characteristics, avoiding the pitfalls of one-size-fits-all ways. Extensive experiments on both vision models (e.g., ViT) and language models (e.g., RoBERTa, Qwen) demonstrate that OA-Merge outperforms state-of-the-art baselines, achieving average performance gains of 3.2% on vision tasks and 2.8% on language tasks.
Qiyuan Zhu, Lujun Li 0001, Jiacheng Liu 0001, Pengyu Cheng, Sirui Han, Yike Guo
ACM Multimedia5
2024 Chunk, Align, Select: A Simple Long-sequence Processing Method for Transformers
abstract
Although dominant in natural language processing, transformer-based models still struggle with long-sequence processing, due to the computational costs of their self-attention operations, which increase exponentially as the length of the input sequence grows.To address this challenge, we propose a Simple framework to enhance the long-content processing of off-the-shelf pre-trained transformers via three steps: Chunk, Align, and Select (SimCAS).More specifically, we first divide each longsequence input into a batch of chunks, then align the inter-chunk information during the encoding steps, and finally, select the most representative hidden states from the encoder for the decoding process.With our SimCAS, the computation and memory costs can be reduced to linear complexity.In experiments, we demonstrate the effectiveness of the proposed method on various real-world long-text summarization and reading comprehension tasks, in which SimCAS significantly outperforms prior longsequence processing baselines.The code is at https://github.com/xjw-nlp/SimCAS.
Jiawen Xie, Pengyu Cheng, Yong Dai 0001
ACL (1)2
2024 Self-playing Adversarial Language Game Enhances LLM Reasoning
abstract
We explore the potential of self-play training for large language models (LLMs) in a two-player adversarial language game called Adversarial Taboo. In this game, an attacker and a defender communicate around a target word only visible to the attacker. The attacker aims to induce the defender to speak the target word unconsciously, while the defender tries to infer the target word from the attacker's utterances. To win the game, both players must have sufficient knowledge about the target word and high-level reasoning ability to infer and express in this information-reserved conversation. Hence, we are curious about whether LLMs' reasoning ability can be further enhanced by Self-Playing this Adversarial language Game (SPAG). With this goal, we select several open-source LLMs and let each act as the attacker and play with a copy of itself as the defender on an extensive range of target words. Through reinforcement learning on the game outcomes, we observe that the LLMs' performances uniformly improve on a broad range of reasoning benchmarks. Furthermore, iteratively adopting this self-play process can continuously promote LLMs' reasoning abilities. The code is available at https://github.com/Linear95/SPAG.
Pengyu Cheng, Tianhao Hu, Zhisong Zhang, Yong Dai 0001
NeurIPS1
2023 Estimating Total Correlation with Mutual Information Estimators
abstract
Total correlation (TC) is a fundamental concept in information theory that measures statistical dependency among multiple random variables. Recently, TC has shown noticeable effectiveness as a regularizer in many learning tasks, where the correlation among multiple latent embeddings requires to be jointly minimized or maximized. However, calculating precise TC values is challenging, especially when the closed-form distributions of embedding variables are unknown. In this paper, we introduce a unified framework to estimate total correlation values with sample-based mutual information (MI) estimators. More specifically, we discover a relation between TC and MI and propose two types of calculation paths (tree-like and line-like) to decompose TC into MI terms. With each MI term being bounded, the TC values can be successfully estimated. Further, we provide theoretical analyses concerning the statistical consistency of the proposed TC estimators. Experiments are presented on both synthetic and real-world scenarios, where our estimators demonstrate effectiveness in all TC estimation, minimization, and maximization tasks.
Ke Bai 0001, Pengyu Cheng, Weituo Hao, Ricardo Henao, Larry Carin
AISTATS2
2023 Toward Fairness in Text Generation via Mutual Information Minimization based on Importance Sampling
abstract
Pretrained language models (PLMs), such as GPT- 2, have achieved remarkable empirical performance in text generation tasks. However, pre- trained on large-scale natural language corpora, the generated text from PLMs may exhibit social bias against disadvantaged demographic groups. To improve the fairness of PLMs in text generation, we propose to minimize the mutual information between the semantics in the generated text sentences and their demographic polarity, i.e., the demographic group to which the sentence is referring. In this way, the mentioning of a demographic group (e.g., male or female) is encouraged to be independent from how it is described in the generated text, thus effectively alleviating the so cial bias. Moreover, we propose to efficiently estimate the upper bound of the above mutual information via importance sampling, leveraging a natural language corpus. We also propose a distillation mechanism that preserves the language modeling ability of the PLMs after debiasing. Empirical results on real-world benchmarks demonstrate that the proposed method yields superior performance in term of both fairness and language modeling ability.
Rui Wang 0088, Pengyu Cheng, Ricardo Henao
AISTATS2
2023 PrSpMV: An Efficient Predictable Kernel for SpMV
abstract
Sparse Matrix-Vector Multiplication (SpMV) has been widely applied in scientific computation, industry simulation, and intelligent computation domains, which is the critical algorithm in all these applications. Due to the poor data locality, low cache usage, and extremely irregular branch patterns caused by the highly sparse and random distributions, SpMV optimization has become one of the most challenging problems for modern high-performance processors. In this paper, we study the bottlenecks of SpMV on current out-of-order CPUs and propose a novel SpMV kernel named PrSpMV to improve its performance by pursuing high predictability. Specifically, we improve the memory access regularity and locality by creating serialized access patterns so that the data prefetching efficiency and cache usage are optimized. We also improve pipeline efficiency by creating regular branch patterns to make branch prediction more accurate. Experiment results show that using the above optimization approaches, PrSpMV can eliminate nearly all branch mispredictions. Moreover, it can also significantly reduce the average L2 cache miss rate from 57% to 20% via efficiently leveraging hardware prefetchers. By using PrSpMV, stride prefetcher can be boosted with 1.31× speedup and dedicated irregular prefetcher can be improved with 1.40× speedup. Meanwhile, on commercial high-end Intel processors, it achieves 1.32× speedup against some state-of-the-art SpMV kernels.
Gelin Fu, Tian Xia 0008, Shaoru Qu, Zhongpei Luo, Pengyu Cheng, Runfan Guo, Yitong Ding, Pengju Ren
ICCD6
2023 Short-term load forecasting of multi-scale recurrent neural networks based on residual structure
abstract
Summary Accurate short‐term load forecasting plays an important role in reducing power generation costs, maintaining supply and demand balance, and stabling the power grids operation. In recent years, deep learning models based on recurrent neural networks (RNN) have been widely used in short‐term load forecasting. Nevertheless, RNN cannot extract multi‐scale features of load data, resulting in low forecasting accuracy. A model for short‐term power load forecasting of residual multiscale‐RNN (RM‐RNN) was proposed in this study. RM‐RNN uses the multilayer RNN network structure. Specifically, each layer sets the dilated convolution with different dilated coefficients to extract the multi‐scale features of the load data. Adjacent networks transfer feature information for feature fusion through the residual structures. The experiment used random sampling data training model, and compared RM‐RNN with multiple deep learning models. The experimental results demonstrated that the mean error of RM‐RNN prediction is the lowest, indicating that dilated convolution can effectively extract multi‐scale features of load data. This result verified the effectiveness of residual structure fusion features, and improved the accuracy of short‐term load forecasting.
Jia Zhao 0001, Pengyu Cheng, Jiazhen Hou, Tanghuai Fan, Longzhe Han
Concurr. Comput. Pract. Exp.2
2023 Multiscale Visual-Attribute Co-Attention for Zero-Shot Image Recognition
abstract
Zero-shot image recognition aims to classify data from unseen classes, by exploring the association between visual features and the semantic representations of each class. Most existing approaches focus on learning a shared single-scale embedding space (often at the output layer of the network) for both visual and semantic features, ignoring a fact that different-scale visual features exhibit different semantics. In this article, we propose a multi-scale visual-attribute co-attention (mVACA) model, considering both visual-semantic alignment and visual discrimination at multiple scales. At each scale, a hybrid visual attention is realized by attribute-related attention and visual self-attention. The attribute-related attention is guided by a pseudo attribute vector inferred via a mutual information regularization (MIR). The visual self-attentive features further influence the attribute attention to emphasize visual-associated attributes. Leveraging multiscale visual discrimination, mVACA unifies standard zero-shot learning (ZSL) and generalized ZSL tasks in one framework, achieving state-of-the-art or competitive performance on several commonly used benchmarks of both setups. To better understand the interaction between images and attributes in mVACA, we also provide visualized analysis.
Hao Zhang 0050, Zhengjue Wang, Yishi Xu, Pengyu Cheng, Ke Bai 0001, Bo Chen 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 MACT: A multi-channel anonymous consensus based on Tor
Ziyu Zheng, Pengyu Cheng, Lin You
World Wide Web (WWW)3
2021 FairFil: Contrastive Neural Debiasing Method for Pretrained Text Encoders
Pengyu Cheng, Weituo Hao, Siyang Yuan, Shijing Si, Lawrence Carin
ICLR1
2021 Improving Zero-Shot Voice Style Transfer via Disentangled Representation Learning
Siyang Yuan, Pengyu Cheng, Ruiyi Zhang 0002, Weituo Hao, Zhe Gan, Lawrence Carin
ICLR2
2020 Dynamic Embedding on Textual Networks via a Gaussian Process
abstract
Textual network embedding aims to learn low-dimensional representations of text-annotated nodes in a graph. Prior work in this area has typically focused on fixed graph structures; however, real-world networks are often dynamic. We address this challenge with a novel end-to-end node-embedding model, called Dynamic Embedding for Textual Networks with a Gaussian Process (DetGP). After training, DetGP can be applied efficiently to dynamic graphs without re-training or backpropagation. The learned representation of each node is a combination of textual and structural embeddings. Because the structure is allowed to be dynamic, our method uses the Gaussian process to take advantage of its non-parametric properties. To use both local and global graph structures, diffusion is used to model multiple hops between neighbors. The relative importance of global versus local structure for the embeddings is learned automatically. With the non-parametric nature of the Gaussian process, updating the embeddings for a changed graph structure requires only a forward pass through the learned model. Considering link prediction and node classification, experiments demonstrate the empirical effectiveness of our method compared to baseline approaches. We further show that DetGP can be straightforwardly and efficiently applied to dynamic textual networks.
Pengyu Cheng, Yitong Li 0001, Xinyuan Zhang 0001, Liqun Chen 0001, David E. Carlson, Lawrence Carin
AAAI1
2020 Improving Disentangled Text Representation Learning with Information-Theoretic Guidance
abstract
Pengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon, Yizhe Zhang, Yitong Li, Lawrence Carin. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Pengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon, Yizhe Zhang 0002, Yitong Li 0001, Lawrence Carin
ACL1
2020 CLUB: A Contrastive Log-ratio Upper Bound of Mutual Information
abstract
Mutual information (MI) minimization has gained considerable interests in various machine learning tasks. However, estimating and minimizing MI in high-dimensional spaces remains a challenging problem, especially when only samples, rather than distribution forms, are accessible. Previous works mainly focus on MI lower bound approximation, which is not applicable to MI minimization problems. In this paper, we propose a novel Contrastive Log-ratio Upper Bound (CLUB) of mutual information. We provide a theoretical analysis of the properties of CLUB and its variational approximation. Based on this upper bound, we introduce a MI minimization training scheme and further accelerate it with a negative sampling strategy. Simulation studies on Gaussian distributions show the reliable estimation ability of CLUB. Real-world MI minimization experiments, including domain adaptation and information bottleneck, demonstrate the effectiveness of the proposed method. The code is at https://github.com/Linear95/CLUB.
Pengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 0001, Zhe Gan, Lawrence Carin
ICML1
2019 Improving Textual Network Embedding with Global Attention via Optimal Transport
abstract
Liqun Chen, Guoyin Wang, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang, Wenlin Wang, Yizhe Zhang, Lawrence Carin. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Liqun Chen 0001, Guoyin Wang 0002, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang 0001, Wenlin Wang, Yizhe Zhang 0002, Lawrence Carin
ACL (1)5
2019 Learning Compressed Sentence Representations for On-Device Text Processing
abstract
Dinghan Shen, Pengyu Cheng, Dhanasekar Sundararaman, Xinyuan Zhang, Qian Yang, Meng Tang, Asli Celikyilmaz, Lawrence Carin. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Dinghan Shen, Pengyu Cheng, Dhanasekar Sundararaman, Xinyuan Zhang 0001, Qian Yang 0003, Asli Celikyilmaz, Lawrence Carin
ACL (1)2
2019 Understanding and Accelerating Particle-Based Variational Inference
abstract
Particle-based variational inference methods (ParVIs) have gained attention in the Bayesian inference literature, for their capacity to yield flexible and accurate approximations. We explore ParVIs from the perspective of Wasserstein gradient flows, and make both theoretical and practical contributions. We unify various finite-particle approximations that existing ParVIs use, and recognize that the approximation is essentially a compulsory smoothing treatment, in either of two equivalent forms. This novel understanding reveals the assumptions and relations of existing ParVIs, and also inspires new ParVIs. We propose an acceleration framework and a principled bandwidth-selection method for general ParVIs; these are based on the developed theory and leverage the geometry of the Wasserstein space. Experimental results show the improved convergence by the acceleration framework and enhanced sample accuracy by the bandwidth-selection method.
Chang Liu 0030, Jingwei Zhuo, Pengyu Cheng, Ruiyi Zhang 0002, Jun Zhu 0001
ICML3