Yige Yuan

dblp:205/6235 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Personality Generation of LLMs at Decoding-time
abstract
Multi-personality generation for LLMs, enabling simultaneous embodiment of multiple personalization attributes, is a fundamental challenge. Existing retraining-based approaches are costly and poorly scalable, while decoding-time methods often rely on external models or heuristics, limiting flexibility and robustness. In this paper, we propose a novel Multi-Personality Generation (MPG) framework under the decoding-time combination paradigm. It flexibly controls multi-personality without relying on scarce multi-dimensional models or extra training, leveraging implicit density ratios in single-dimensional models as a ''free lunch'' to reformulate the task as sampling from a target strategy aggregating these ratios. To implement MPG efficiently, we design Speculative Chunk-level based Rejection sampling (SCR), which generates responses in chunks and parallelly validates them via estimated thresholds within a sliding window. This significantly reduces computational overhead while maintaining high-quality generation. Experiments on MBTI personality and Role-Playing demonstrate the effectiveness of MPG, showing improvements up to 16%–18%. Code and data are available at https://github.com/Libra117/MPG.
Rongxin Chen, Yige Yuan, Bingbing Xu 0001, Huawei Shen
WSDM3
2025 From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment
abstract
Inference-time alignment methods have gained significant attention for their efficiency and effectiveness in aligning large language models (LLMs) with human preferences.However, existing dominant approaches using rewardguided search (RGS) primarily rely on outcome reward models (ORMs), which suffer from a critical granularity mismatch: ORMs are designed to provide outcome rewards for complete responses, while RGS methods rely on process rewards to guide the policy, leading to inconsistent scoring and suboptimal alignment.To address this challenge, we introduce process reward models (PRMs) into RGS and argue that an ideal PRM should satisfy two objectives: Score Consistency, ensuring coherent evaluation across partial and complete responses, and Preference Consistency, aligning partial sequence assessments with human preferences.Based on these, we propose SP-PRM, a novel dual-consistency framework integrating score consistency-based and preference consistency-based partial evaluation modules without relying on human annotation.Extensive experiments on dialogue, summarization, and reasoning tasks demonstrate that SP-PRM substantially enhances existing RGS methods, achieving a 3.6%-10.3%improvement in GPT-4 evaluation scores across all tasks.Code is publicly available at this link.
Bingbing Xu 0001, Yige Yuan, Shengmao Zhu, Huawei Shen
ACL (1)3
2025 The 1st Workshop on LLM Agents for Social Simulation
abstract
Social simulation has long played a crucial role in exploring the mechanisms underlying human behavior and societal structures. Traditional social simulation relies on rule-based or statistical models, which makes it difficult to capture the complexity and variability of the real world. With the emergence and rapid development of large language model (LLM), new frontiers have been opened toward leveraging LLMs as agent to model human behavior and interactions. This cutting-edge direction has gained significant attention and demonstrated promising results, not only advancing research across a wide range of social science disciplines, but also enabling practical applications in role-playing scenarios. However, this field still faces multiple challenges, such as capturing real-world social phenomena, eliminating bias or ethical considerations, and ensuring usability and reliability. This workshop on LLM Agent for Social Simulation (LASS) aims to bring together researchers and practitioners from diverse backgrounds to foster interdisciplinary collaboration, address key challenges, explore new technologies, and chart promising future directions in this rapidly evolving field.
Yige Yuan, Junkai Zhou, Bingbing Xu 0001, Liang Pang 0001, Du Su, An Zhang 0003, Teng Xiao, Fengli Xu, Zhaochun Ren, Xu Chen 0017
CIKM1
2025 SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters
abstract
Existing preference optimization objectives for language model alignment require additional hyperparameters that must be extensively tuned to achieve optimal performance, increasing both the complexity and time required for fine-tuning large language models. In this paper, we propose a simple yet effective hyperparameter-free preference optimization algorithm for alignment. We observe that promising performance can be achieved simply by optimizing inverse perplexity, which is calculated as the inverse of the exponentiated average log-likelihood of the chosen and rejected responses in the preference dataset. The resulting simple learning objective, SimPER, is easy to implement and eliminates the need for expensive hyperparameter tuning and a reference model, making it both computationally and memory efficient. Extensive experiments on widely used real-world benchmarks, including MT-Bench, AlpacaEval 2, and 10 key benchmarks of the Open LLM Leaderboard with 5 base models, demonstrate that SimPER consistently and significantly outperforms existing approaches—even without any hyperparameters or a reference model. For example, despite its simplicity, SimPER outperforms state-of-the-art methods by up to 5.7 points on AlpacaEval 2 and achieves the highest average ranking across 10 benchmarks on the Open LLM Leaderboard. The source code for SimPER is publicly available at: https://github.com/tengxiao1/SimPER.
Teng Xiao, Yige Yuan, Zhengyu Chen 0001, Mingxiao Li 0004, Shangsong Liang, Zhaochun Ren, Vasant G. Honavar
ICLR2
2025 On a Connection Between Imitation Learning and RLHF
abstract
This work studies the alignment of large language models with preference data from an imitation learning perspective. We establish a close theoretical connection between reinforcement learning from human feedback RLHF and imitation learning (IL), revealing that RLHF implicitly performs imitation learning on the preference data distribution. Building on this connection, we propose DIL, a principled framework that directly optimizes the imitation learning objective. DIL provides a unified imitation learning perspective on alignment, encompassing existing alignment algorithms as special cases while naturally introducing new variants. By bridging IL and RLHF, DIL offers new insights into alignment with RLHF. Extensive experiments demonstrate that DIL outperforms existing methods on various challenging benchmarks.
Teng Xiao, Yige Yuan, Mingxiao Li 0004, Zhengyu Chen 0001, Vasant G. Honavar
ICLR2
2025 Adaptive Multiscale Decomposition Echo State Network for Chaotic Time Series Prediction
abstract
Chaotic time series are characterized by high nonlinearity, multi-scale properties, and frequency overlap. Traditional methods struggle to fully extract multi-scale dynamic information and are susceptible to the effects of random initialization and noise interference. This study proposes an Adaptive Multi-Scale Decomposition Echo State Network (AMDESN) model. First, an HP filter is used to decompose the original time series into sub-sequences of different frequency bands; then, a structurally adaptive multi-Reservoir ESN subnetwork is configured for each sub-sequence, with the Reservoir size and connection structure adjusted according to the dynamic characteristics of the subsequence. Additionally, the AMDESN introduces a new weight initialization strategy and composite activation function, and integrates the states of each subnetwork through a multi-layer Reservoir state concatenation fusion strategy to output the final prediction. Experiments on multiple typical chaotic time series benchmark datasets demonstrate that the proposed AMDESN model significantly outperforms traditional methods in terms of prediction accuracy, result stability, and noise resistance. Even when noise is added, AMDESN maintains a smaller error growth range, exhibiting stronger robustness. These results validate that the method effectively enhances the prediction performance of chaotic time series.
XiaoDan He, Yige Yuan
ICPADS3
2025 MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing
abstract
Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging. Existing methods typically treat them separately, struggling with limited high-quality data and poor generalization. However, both tasks require capturing complex visual variations while maintaining consistency between inputs and outputs. Inspired by this, we propose MIGE, a unified framework that standardizes task representations using multimodal instructions. It first treats subject-driven generation as creation on a blank canvas and instruction-based editing as modification of an existing image, establishing a shared input-output formulation, then introduces a novel multimodal encoder that maps free-form multimodal instructions into a unified vision-language space, integrating visual and semantic features through a feature fusion mechanism. This unification enables joint training of both tasks, providing two key advantages: (1) Cross-Task Enhancement: by leveraging shared visual and semantic representations, joint training improves instruction adherence and visual consistency in both subject-driven generation and instruction-based editing. (2) Generalization: learning in a unified format facilitates cross-task knowledge transfer, enabling MIGE to generalize to novel compositional tasks, including instruction-based subject-driven editing. Experiments show that MIGE excels in both subject-driven generation and instruction-based editing while setting a SOTA in the new task of instruction-based subject-driven editing. Code and model have been publicly available at https://github.com/Eureka-Maggie/MIGE.
Xueyun Tian, Wei Li 0176, Bingbing Xu 0001, Yige Yuan, Yuanzhuo Wang, Huawei Shen
ACM Multimedia4
2025 Inference-time Alignment in Continuous Space
abstract
Aligning large language models with human feedback at inference time has received increasing attention due to its flexibility. Existing methods rely on generating multiple responses from the base policy for search using a reward model, which can be considered as searching in a discrete response space. However, these methods struggle to explore informative candidates when the base policy is weak or the candidate set is small, resulting in limited effectiveness. In this paper, to address this problem, we propose Simple Energy Adaptation ($\textbf{SEA}$), a simple yet effective algorithm for inference-time alignment. In contrast to expensive search over the discrete space, SEA directly adapts original responses from the base policy toward the optimal one via gradient-based sampling in continuous latent space. Specifically, SEA formulates inference as an iterative optimization procedure on an energy function over actions in the continuous space defined by the optimal policy, enabling simple and effective alignment. For instance, despite its simplicity, SEA outperforms the second-best baseline with a relative improvement of up to $ \textbf{77.51\%}$ on AdvBench and $\textbf{16.36\%}$ on MATH. Code is publicly available at [this link](https://github.com/yuanyige/sea).
Yige Yuan, Teng Xiao, Bingbing Xu 0001, Shuchang Tao, Yunqi Qiu, Huawei Shen, Xueqi Cheng 0001
NeurIPS1
2025 InfoNCE is a Free Lunch for Semantically guided Graph Contrastive Learning
abstract
As an important graph pre-training method, Graph Contrastive Learning (GCL) continues to play a crucial role in the ongoing surge of research on graph foundation models or LLM as enhancer for graphs. Traditional GCL optimizes InfoNCE by using augmentations to define self-supervised tasks, treating augmented pairs as positive samples and others as negative. However, this leads to semantically similar pairs being classified as negative, causing significant sampling bias and limiting performance. In this paper, we argue that GCL is essentially a Positive-Unlabeled (PU) learning problem, where the definition of self-supervised tasks should be semantically guided, i.e., augmented samples with similar semantics are considered positive, while others, with unknown semantics, are treated as unlabeled. From this perspective, the key lies in how to extract semantic information. To achieve this, we propose IFL-GCL, using InfoNCE as a "free lunch" to extract semantic information. Specifically, We first prove that under InfoNCE, the representation similarity of node pairs aligns with the probability that the corresponding contrastive sample is positive. Then we redefine the maximum likelihood objective based on the corrected samples, leading to a new InfoNCE loss function. Extensive experiments on both the graph pretraining framework and LLM as an enhancer show significantly improvements of IFL-GCL in both IID and OOD scenarios, achieving up to a 9.05% improvement, validating the effectiveness of semantically guided. Code for IFL-GCL is publicly available at: https://github.com/Camel-Prince/IFL-GCL.
Bingbing Xu 0001, Yige Yuan, Huawei Shen, Xueqi Cheng 0001
SIGIR3
2025 Fact-Level Calibration and Correction for Long-Form Generations
abstract
Large language models (LLMs) have achieved remarkable progress across various domains, yet their tendency to generate hallucinations remains a critical barrier to their practical reliability.Confidence calibration addresses this challenge by aligning a model's confidence with its actual accuracy, improving self-evaluation and trustworthiness.However, traditional confidence calibration, operating at response level, are inadequate for long-form generation, which involve complex outputs composed of multiple atomic facts, each with varying confidence, correctness, and relevance to the query.To overcome this limitation, we propose a fact-level confidence calibration framework that evaluates and adjusts confidence at the granularity of individual facts, incorporating both relevance and correctness.This framework identifies finer-grained calibration discrepancies, reduces overconfidence, and reveals confidence variance.Based on this framework, we introduce CARE (Confidence-Aware Fact Correction), a method that leverages high-confidence facts to iteratively refine and correct low-confidence ones.Experimental results demonstrate that our CARE effectively improves the quality of generated content.Our code is available at this link.
Yige Yuan, Bingbing Xu 0001, Hexiang Tan, Fei Sun 0001, Teng Xiao, Wei Li 0176, Huawei Shen, Xueqi Cheng 0001
SIGIR1
2024 PDE+: Enhancing Generalization via PDE with Adaptive Distributional Diffusion
abstract
The generalization of neural networks is a central challenge in machine learning, especially concerning the performance under distributions that differ from training ones. Current methods, mainly based on the data-driven paradigm such as data augmentation, adversarial training, and noise injection, may encounter limited generalization due to model non-smoothness. In this paper, we propose to investigate generalization from a Partial Differential Equation (PDE) perspective, aiming to enhance it directly through the underlying function of neural networks, rather than focusing on adjusting input data. Specifically, we first establish the connection between neural network generalization and the smoothness of the solution to a specific PDE, namely transport equation. Building upon this, we propose a general framework that introduces adaptive distributional diffusion into transport equation to enhance the smoothness of its solution, thereby improving generalization. In the context of neural networks, we put this theoretical framework into practice as PDE+ (PDE with Adaptive Distributional Diffusion) which diffuses each sample into a distribution covering semantically similar inputs. This enables better coverage of potentially unobserved distributions in training, thus improving generalization beyond merely data-driven methods. The effectiveness of PDE+ is validated through extensive experimental settings, demonstrating its superior performance compared to state-of-the-art methods. Our code is available at https://github.com/yuanyige/pde-add.
Yige Yuan, Bingbing Xu 0001, Fei Sun 0001, Huawei Shen, Xueqi Cheng 0001
AAAI1
2024 TEA: Test-Time Energy Adaptation
abstract
Test Time Adaptation (TTA) aims to improve model generalizability when test data diverges from training distribution, with the distinct advantage of not requiring access to training data and processes, especially valuable in the context of pre-trained models. However, current TTA methods fail to address the fundamental issue: covariate shift, i. e., the decreased generalizability can be attributed to the model's reliance on the marginal distribution of the training data, which may impair model calibration and introduce confirmation bias. To address this, we propose a novel energy-based perspective, enhancing the model's perception of target data distributions without requiring access to training data or processes. Building on this perspective, we introduce Test-time Energy Adaptation (TEA), which transforms the trained classifier into an energy-based model and aligns the model's distribution with the test data's, enhancing its ability to perceive test distributions and thus improving overall generalizability. Extensive experiments across multiple tasks, benchmarks and architectures demonstrate TEA's superior generalization performance against state-of-the-art methods. Further in-depth analyses reveal that TEA can equip the model with a comprehensive perception of test distribution, ultimately paving the way toward improved generalization and calibration11Code is available at https://github.com/yuanyige/tea
Yige Yuan, Bingbing Xu 0001, Fei Sun 0001, Huawei Shen, Xueqi Cheng 0001
CVPR1
2024 History Driven Sampling for Scalable Graph Neural Networks
Yang Li 0202, Bingbing Xu 0001, Fei Sun 0001, Qi Cao 0005, Yige Yuan, Huawei Shen, Xueqi Cheng 0001
DASFAA (6)5
2024 How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective
abstract
This paper introduces a novel generalized selfimitation learning (GSIL) framework, which effectively and efficiently aligns large language models with offline demonstration data.We develop GSIL by deriving a surrogate objective of imitation learning with density ratio estimates, facilitating the use of self-generated data and optimizing the imitation learning objective with simple classification losses.GSIL eliminates the need for complex adversarial training in standard imitation learning, achieving lightweight and efficient fine-tuning for large language models.In addition, GSIL encompasses a family of offline losses parameterized by a general class of convex functions for density ratio estimation and enables a unified view for alignment with demonstration data.Extensive experiments show that GSIL consistently and significantly outperforms baselines in many challenging benchmarks, such as coding (HuamnEval), mathematical reasoning (GSM8K) and instruction-following benchmark (MT-Bench).Code is public available at https://github.com/tengxiao1/GSIL.
Teng Xiao, Mingxiao Li 0004, Yige Yuan, Huaisheng Zhu, Chao Cui, Vasant G. Honavar
EMNLP3
2024 Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
abstract
We study the problem of aligning large language models (LLMs) with human preference data. Contrastive preference optimization has shown promising results in aligning LLMs with available preference data by optimizing the implicit reward associated with the policy. However, the contrastive objective focuses mainly on the relative values of implicit rewards associated with two responses while ignoring their actual values, resulting in suboptimal alignment with human preferences. To address this limitation, we propose calibrated direct preference optimization (Cal-DPO), a simple yet effective algorithm. We show that substantial improvement in alignment with the given preferences can be achieved simply by calibrating the implicit reward to ensure that the learned implicit rewards are comparable in scale to the ground-truth rewards. We demonstrate the theoretical advantages of Cal-DPO over existing approaches. The results of our experiments on a variety of standard benchmarks show that Cal-DPO remarkably improves off-the-shelf methods.
Teng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li 0004, Vasant G. Honavar
NeurIPS2
2024 Negative as Positive: Enhancing Out-of-distribution Generalization for Graph Contrastive Learning
abstract
Graph contrastive learning (GCL), standing as the dominant paradigm in the realm of graph pre-training, has yielded considerable progress. Nonetheless, its capacity for out-of-distribution (OOD) generalization has been relatively underexplored. In this work, we point out that the traditional optimization of InfoNCE in GCL restricts the cross-domain pairs only to be negative samples, which inevitably enlarges the distribution gap between different domains. This violates the requirement of domain invariance under OOD scenario and consequently impairs the model's OOD generalization performance. To address this issue, we propose a novel strategy ''Negative as Positive'', where the most semantically similar cross-domain negative pairs are treated as positive during GCL. Our experimental results, spanning a wide array of datasets, confirm that this method substantially improves the OOD generalization performance of GCL.
Bingbing Xu 0001, Yige Yuan, Huawei Shen, Xueqi Cheng 0001
SIGIR3
2024 Towards generalizable Graph Contrastive Learning: An information theory perspective
Yige Yuan, Bingbing Xu 0001, Huawei Shen, Qi Cao 0005, Keting Cen, Xueqi Cheng 0001
Neural Networks1
2023 Augmentation-Aware Self-Supervision for Data-Efficient GAN Training
abstract
Training generative adversarial networks (GANs) with limited data is challenging because the discriminator is prone to overfitting. Previously proposed differentiable augmentation demonstrates improved data efficiency of training GANs. However, the augmentation implicitly introduces undesired invariance to augmentation for the discriminator since it ignores the change of semantics in the label space caused by data transformation, which may limit the representation learning ability of the discriminator and ultimately affect the generative modeling performance of the generator. To mitigate the negative impact of invariance while inheriting the benefits of data augmentation, we propose a novel augmentation-aware self-supervised discriminator that predicts the augmentation parameter of the augmented data. Particularly, the prediction targets of real data and generated data are required to be distinguished since they are different during training. We further encourage the generator to adversarially learn from the self-supervised discriminator by generating augmentation-predictable real and not fake data. This formulation connects the learning objective of the generator and the arithmetic $-$ harmonic mean divergence under certain assumptions. We compare our method with state-of-the-art (SOTA) methods using the class-conditional BigGAN and unconditional StyleGAN2 architectures on data-limited CIFAR-10, CIFAR-100, FFHQ, LSUN-Cat, and five low-shot datasets. Experimental results demonstrate significant improvements of our method over SOTA methods in training data-efficient GANs.
Qi Cao 0005, Yige Yuan, Songtao Zhao, Chongyang Ma, Siyuan Pan, Pengfei Wan 0001, Huawei Shen, Xueqi Cheng 0001
NeurIPS3
2017 Several Generalized Interval-Valued 2-Tuple Linguistic Interval Distance Measures and Their Application
abstract
Interval-valued linguistic variables are efficient tools to express the decision makers’ uncertain qualitative judgments. Considering the application of interval-valued linguistic variables, this paper proposes an interval distance measure, which is then used to define interval-valued linguistic interval distance measures by combining the 2-tuple linguistic representation model. To reflect the interactions between elements in a set, three correlative interval distance measures on intervalvalued linguistic variables are proposed. Meanwhile, several models designed to obtain the optimal weighting vector are constructed. After that, an approach to pattern recognition and to multi-attribute decision making with interval-valued linguistic information is developed. Meanwhile, associated examples are offered to demonstrate the concrete application of the proposed procedure.
Fanyong Meng 0001, Yige Yuan, Xiaohong Chen 0001
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2