Ru Peng

dblp:305/5740 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 W2S: Weak-to-Strong Prompt Correction for Large Language Models
Lirong Gao, Hao Chen 0081, Ru Peng, Qi Zhang 0077, Yiming Zhang 0023, Wentao Ye, Haobo Wang 0001, Junbo Zhao 0002
Mach. Learn.4
2025 DataMan: Data Manager for Pre-training Large Language Models
abstract
The performance emergence of large language models (LLMs) driven by data scaling laws makes the selection of pre-training data increasingly important. However, existing methods rely on limited heuristics and human intuition, lacking comprehensive and clear guidelines. To address this, we are inspired by *``reverse thinking''* -- prompting LLMs to self-identify which criteria benefit its performance. As its pre-training capabilities are related to perplexity (PPL), we derive 14 quality criteria from the causes of text perplexity anomalies and introduce 15 common application domains to support domain mixing. In this paper, we train a **Data** **Man**ager (**DataMan**) to learn quality ratings and domain recognition from pointwise rating, and use it to annotate a 447B token pre-training corpus with 14 quality ratings and domain type. Our experiments validate our approach, using DataMan to select 30B tokens to train a 1.3B-parameter language model, demonstrating significant improvements in in-context learning (ICL), perplexity, and instruction-following ability over the state-of-the-art baseline. The best-performing model, based on the *Overall Score l=5* surpasses a model trained with 50% more data using uniform sampling. We continue pre-training with high-rated, domain-specific data annotated by DataMan to enhance domain-specific ICL performance and thus verify DataMan's domain mixing ability. Our findings emphasize the importance of quality ranking, the complementary nature of quality criteria, and their low correlation with perplexity, analyzing misalignment between PPL and ICL performance. We also thoroughly analyzed our pre-training dataset, examining its composition, the distribution of quality ratings, and the original document sources.
Ru Peng, Kexin Yang 0002, Yawen Zeng, Junyang Lin, Dayiheng Liu, Junbo Zhao 0002
ICLR1
2025 Improving Sample Efficiency Through Stability Enhancement in Deep-Reinforcement Learning
abstract
Prioritizing or reweighting important samples has been recognized as an effective means of improving the efficiency of deep-reinforcement learning (DRL) algorithms. However, many existing techniques encounter stability challenges, limiting efficiency and increasing computational costs and training time. In this study, we aim to improve training efficiency by exploring the intrinsic relationship between sample efficiency and stability. To achieve this, we propose the Stability Contribution Index (SI), which assigns sample priorities based on their impact on stability and employs them to weight the value loss, thereby promoting stable learning and improving efficiency. The effectiveness of our method is validated through comprehensive experiments on two distinct benchmarks: 1) the continuous control domain DMControl and 2) the discrete control environment ProcGen. Compatible with both off-policy and on-policy DRL algorithms, our approach significantly improves sample efficiency and overall performance by fostering greater stability during training. Additionally, experimental results show that our method outperforms well-established sample-efficient reinforcement learning techniques across multiple settings.
Ziru Wang, Wanli Jiang, Ru Peng, Qian Kou, Lipeng Wan 0003, Xuguang Lan
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Embedding and Gradient Say Wrong: A White-Box Method for Hallucination Detection
abstract
In recent years, large language models (LLMs) have achieved remarkable success in the field of natural language generation.Compared to previous small-scale models, they are capable of generating fluent output based on the provided prefix or prompt.However, one critical challenge -the hallucination problem -remains to be resolved.Generally, the community refers to the undetected hallucination scenario where the LLMs generate text unrelated to the input text or facts.In this study, we intend to model the distributional distance between the regular conditional output and the unconditional output, which is generated without a given input text.Based upon Taylor Expansion for this distance at the output probability space, our approach manages to leverage the embedding and first-order gradient information.The resulting approach is plug-and-play that can be easily adapted to any autoregressive LLM.On the hallucination benchmarks HADES and other datasets, our approach achieves state-of-the-art performance.
Xiaomeng Hu, Yiming Zhang 0023, Ru Peng, Chenwei Wu 0010, Gang Chen 0001, Junbo Zhao 0002
EMNLP3
2024 Predicting Rewards Alongside Tokens: Non-disruptive Parameter Insertion for Efficient Inference Intervention in Large Language Model
abstract
Transformer-based large language models (LLMs) exhibit limitations such as generating unsafe responses, unreliable reasoning, etc. Existing inference intervention approaches attempt to mitigate these issues by finetuning additional models to produce calibration signals (such as rewards) that guide the LLM's decoding process.However, this solution introduces substantial time and space overhead due to the separate models required.This work proposes NOn-disruptive parameters insertion (Otter), inserting extra parameters into the transformer architecture to predict calibration signals along with the original LLM output.Otter offers state-of-the-art performance on multiple demanding tasks while saving up to 86.5% extra space and 98.5% extra time.Furthermore, Otter seamlessly integrates with existing inference engines, requiring only a oneline code change, and the original model response remains accessible after the parameter insertion.
Chenhan Yuan, Fei Huang 0005, Ru Peng, Keming Lu, Bowen Yu 0002, Chang Zhou 0005, Jingren Zhou 0001
EMNLP3
2024 Energy-based Automated Model Evaluation
abstract
The conventional evaluation protocols on machine learning models rely heavily on a labeled, i.i.d-assumed testing dataset, which is not often present in real-world applications. The Automated Model Evaluation (AutoEval) shows an alternative to this traditional workflow, by forming a proximal prediction pipeline of the testing performance without the presence of ground-truth labels. Despite its recent successes, the AutoEval frameworks still suffer from an overconfidence issue, substantial storage and computational cost. In that regard, we propose a novel measure --- Meta-Distribution Energy (MDE) that allows the AutoEval framework to be both more efficient and effective. The core of the MDE is to establish a meta-distribution statistic, on the information (energy) associated with individual samples, then offer a smoother representation enabled by energy-based learning. We further provide our theoretical insights by connecting the MDE with the classification loss. We provide extensive experiments across modalities, datasets and different architectural backbones to validate MDE's validity, together with its superiority compared with prior approaches. We also prove MDE's versatility by showing its seamless integration with large-scale models, and easy adaption to learning scenarios with noisy- or imbalanced- labels.
Ru Peng, Heming Zou, Haobo Wang 0001, Yawen Zeng, Zenan Huang, Junbo Zhao 0002
ICLR1
2024 Experience Consistency Distillation Continual Reinforcement Learning for Robotic Manipulation Tasks
abstract
Continual reinforcement learning, which aims to help robots acquire skills without catastrophic forgetting, obviating the need to re-learn all tasks from scratch. In order to enable lifelong acquisition of skills in robots, replay-based continual reinforcement learning has emerged as a promising research direction. These techniques replay data from previous tasks to mitigate forgetting when learning new skills. However, existing replay-based methods store poor representative experience, and the experience utilization of old tasks is inefficient. To address these issues, we propose an experience consistency distillation method for robot continual reinforcement learning to improve the data efficiency of the experience. Specifically, the experience of old tasks are distilled to obtain Markov Decision Process (MDP) data with high compression ratio and information content. To ensure consistent data distributions before and after distillation, we further utilize a Fréchet Inception Distance (FID) loss as a regularization constraint. In order to improve experience utilization efficiency, the policy is then trained using both the distilled data and current task data, with policy distillation performed based on uncertainty metrics. Our method is validated in the continual reinforcement learning simulation platform and real scene with a UR5e robot arm. Experimental results indicate that our method achieves higher success and lower buffer size requirement compared to other methods.
Ru Peng, Xingyu Chen 0001, Xuguang Lan
ICRA3
2024 A causality guided loss for imbalanced learning in scene graph generation
Ru Peng, Xingyu Chen 0001, Ziru Wang, Xuguang Lan
Neurocomputing1
2024 A Stabilized Fast Transversal Equalizer for Single Carrier Underwater Acoustic Communications
abstract
The use of Decision Feedback Equalization (DFE) with Digital Phase-Locked Loop (DPLL) technology has proven effective in enhancing resistance to multipath delay and Doppler spread in underwater acoustic (UWA) communications. However, traditional equalizers face challenges in achieving a balance between convergence speed and computational complexity. To address the balance, this paper proposes a Stabilized Fast Transversal Equalizer (SFTE) for Quadrature Amplitude Modulation (QAM) in UWA communications. The SFTE incorporates four collaborative parallel transversal filters and introduces redundancy in the conversion factor and previous reverse prediction error, resulting in a reliable structure with reduced complexity. Compared to recursive least squares (RLS)-based equalizers with length L, the SFTE achieves comparable complexity of O(L), significantly reducing it from the traditional O(L2). Simulations and UWA communication experiments are conducted to validate the superiority of SFTE in terms of convergence, tracking capability, runtime, and Bit Error Rate (BER) performance.
Fei-Yun Wu, Ru Peng, Yiyang Ni 0003, Yan-Chong Song
IEEE Internet Things J.2
2023 CAME: Contrastive Automated Model Evaluation
abstract
The Automated Model Evaluation (AutoEval) framework entertains the possibility of evaluating a trained machine learning model without resorting to a labeled testing set. Despite the promise and some decent results, the existing AutoEval methods heavily rely on computing distribution shifts between the unlabelled testing set and the training set. We believe this reliance on the training set becomes another obstacle in shipping this technology to real-world ML development. In this work, we propose Contrastive Automatic Model Evaluation (CAME), a novel AutoEval framework that is rid of involving training set in the loop. The core idea of CAME bases on a theoretical analysis which bonds the model performance with a contrastive loss. Further, with extensive empirical validation, we manage to set up a predictable relationship between the two, simply by deducing on the unlabeled/unseen testing set. The resulting framework CAME establishes a new SOTA results for AutoEval by surpassing prior work significantly.1
Ru Peng, Qiuyang Duan, Haobo Wang 0001, Jiachen Ma 0005, Yanbo Jiang, Yongjun Tu, Xiu Jiang, Junbo Zhao 0002
ICCV1
2023 A neighborhood-based multiple orthogonal least square method for sparse signal recovery
Yan-Chong Song, Fei-Yun Wu, Ru Peng
Signal Process.3
2022 Distill The Image to Nowhere: Inversion Knowledge Distillation for Multimodal Machine Translation
abstract
Past works on multimodal machine translation (MMT) elevate bilingual setup by incorporating additional aligned vision information.However, an image-must requirement of the multimodal dataset largely hinders MMT's development -namely that it demands an aligned form of [image, source text, target text].This limitation is generally troublesome during the inference phase especially when the aligned image is not provided as in the normal NMT setup.Thus, in this work, we introduce IKD-MMT, a novel MMT framework to support the image-free inference phase via an inversion knowledge distillation scheme.In particular, a multimodal feature generator is executed with a knowledge distillation module, which directly generates the multimodal feature from (only) source texts as the input.While there have been a few prior works entertaining the possibility to support image-free inference for machine translation, their performances have yet to rival the image-must translation.In our experiments, we identify our method as the first image-free approach to comprehensively rival or even surpass (almost) all image-must frameworks, and achieved the state-of-the-art result on the often-used Multi30k benchmark 1 .
Ru Peng, Yawen Zeng, Jake Zhao
EMNLP1
2022 Deps-SAN: Neural Machine Translation with Dependency-Scaled Self-Attention Network
Ru Peng, Nankai Lin, Shengyi Jiang, Tianyong Hao, Junbo Zhao 0002
ICONIP (3)1
2022 HybridVocab: Towards Multi-Modal Machine Translation via Multi-Aspect Alignment
abstract
Multi-modal machine translation (MMT) aims to augment the linguistic machine translation frameworks by incorporating aligned vision information. As the core research challenge for MMT, how to fuse the image information and further align it with the bilingual data remains critical. Existing works have either focused on a methodological alignment in the space of bilingual text or emphasized the combination of the one-sided text and given image. In this work, we entertain the possibility of a triplet alignment, among the source and target text together with the image instance. In particular, we propose Multi-aspect AlignmenT (MAT) model that augments the MMT tasks to three sub-tasks --- namely cross-language translation alignment, cross-modal captioning alignment and multi-modal hybrid alignment tasks. Core to this model consists of a hybrid vocabulary which compiles the visually depictable entity (nouns) occurrence on both sides of the text as well as the detected object labels appearing in the images. Through this sub-task, we postulate that MAT manages to further align the modalities by casting three instances into a shared domain, as compared against previously proposed methods. Extensive experiments and analyses demonstrate the superiority of our approaches, which achieve several state-of-the-art results on two benchmark datasets of the MMT task.
Ru Peng, Yawen Zeng, Junbo Zhao 0002
ICMR1
2022 TaiSu: A 166M Large-scale High-Quality Dataset for Chinese Vision-Language Pre-training
abstract
Vision-Language Pre-training (VLP) has been shown to be an efficient method to improve the performance of models on different vision-and-language downstream tasks. Substantial studies have shown that neural networks may be able to learn some general rules about language and visual concepts from a large-scale weakly labeled image-text dataset. However, most of the public cross-modal datasets that contain more than 100M image-text pairs are in English; there is a lack of available large-scale and high-quality Chinese VLP datasets. In this work, we propose a new framework for automatic dataset acquisition and cleaning with which we construct a new large-scale and high-quality cross-modal dataset named as TaiSu, containing 166 million images and 219 million Chinese captions. Compared with the recently released Wukong dataset, our dataset is achieved with much stricter restrictions on the semantic correlation of image-text pairs. We also propose to combine texts collected from the web with texts generated by a pre-trained image-captioning model. To the best of our knowledge, TaiSu is currently the largest publicly accessible Chinese cross-modal dataset. Furthermore, we test our dataset on several vision-language downstream tasks. TaiSu outperforms BriVL by a large margin on the zero-shot image-text retrieval task and zero-shot image classification task. TaiSu also shows better performance than Wukong on the image-retrieval task without using image augmentation for training. Results demonstrate that TaiSu can serve as a promising VLP dataset, both for understanding and generative tasks. More information can be referred to https://github.com/ksOAn6g5/TaiSu.
Guibo Zhu, Qi Song 0003, Guojing Ge, Guanhui Qiao, Ru Peng, Lingxiang Wu, Jinqiao Wang
NeurIPS8
2022 A Hybrid Classification to Detect Abstinent Heroin-Addicted Individuals Using EEG Microstates
abstract
Objective: Diagnosis of the severity of heroin addiction with electroencephalography (EEG) signals is a challenging problem. It has been shown that brain microstates are associated with brain status and healthy condition. However, there is no study on how heroin addiction affects brain microstates. Approach: We propose a hybrid classifier based on the microstate features, extracting from resting state EEGs, to objectively and effectively identify abstinent heroin-addicted individuals (AHAIs) and healthy controls (HCs). In addition to the commonly used features such as duration, occurrence, and transition, we calculated three new features. Main Results: The results showed that the support vector machine (SVM), which allows classification of the AHAIs and HCs with a 73% accuracy rate, was an optimal classifier. Moreover, the weight setting-based genetic algorithm (GA) further improved the accuracy rate to 81%. The hybrid classification not only provides direct evidence showing the differences in EEG microstate features between AHAIs and HCs, but also offers a method to distinguish the heroin brain states of people addicted to heroin and healthy individuals and demonstrates that microstate features could serve as potential bio-markers for identifying AHAIs. Significance: our methods and the selected features may provide electrophysiological insights for the assessment of the heroin withdrawal treatment effects.
Ru Peng, Quanying Liu, Hong Peng 0003
IEEE Trans. Comput. Soc. Syst.2
2021 Syntax-aware neural machine translation directed by syntactic dependency degree
Ru Peng, Tianyong Hao
Neural Comput. Appl.1