Yongzhe Chang

dblp:238/4188 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0002-9083-5348ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Surrogate-Assisted Evolutionary Multi-Agent Reinforcement Learning with Adaptive Fitness Evaluation
abstract
Deep Multi-Agent Reinforcement Learning (MARL) excels in co-operative tasks but often struggles with local optima in high - dimensional joint action spaces. In contrast, Evolutionary Algorithms (EAs) offer robust global exploration capabilities. Although hybrid approaches seek to combine the strengths of both paradigms, they typically face a critical bottleneck: the prohibitive sample cost of evaluating large populations via environment rollouts. To address this challenge, we propose Surrogate-assisted Evolutionary Multi-Agent Reinforcement Learning (SEMARL), a unified framework that synergizes gradient-based refinement with surrogate-assisted evolutionary search. SEMARL employs a cooperative co-evolutionary architecture to maintain diverse agent policies and injects gradient-refined parameters into the population to accelerate convergence. Crucially, we leverage the centralized critic from the gradient learner as a computationally efficient surrogate for fitness estimation. To prevent misleading guidance from an inaccurate critic, we introduce an adaptive reliability control mechanism based on temporal difference (TD) error, which dynamically regulates the surrogate's influence. Experiments on the Multi-Agent MuJoCo benchmark demonstrate that SEMARL significantly outperforms other algorithms, achieving superior asymptotic performance with substantially higher sample efficiency.
Cong Yu 0018, Zaihui Yang, Haoyu Wang 0018, Junbo Tan, Yongzhe Chang, Tiantian Zhang 0002, Xueqian Wang 0001
GECCO7
2026 Tacit mechanism: Bridging pre-training of individuality to multi-agent adversarial coordination
Shiqing Yao, Jiajun Chai, Haixin Yu, Yongzhe Chang, Tiantian Zhang 0002, Yuanheng Zhu, Xueqian Wang 0001
Neural Networks4
2025 Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences Through f-Divergence Minimization
abstract
Direct Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable interest within the community. However, we have observed that these approaches rely solely on minimizing the reverse Kullback-Leibler divergence during alignment process between the fine-tuned model and the reference model, neglecting incorporation of other divergence constraints. In this study, we focus on extending reverse Kullback-Leibler divergence in the alignment paradigm of text-to-image models to f-divergence, which aims to garner better alignment performance as well as good generation diversity. We provide the generalized formula of text-to-image alignment paradigm under f-divergence condition and thoroughly analyze the impact of different divergence constraints on alignment process from the perspective of gradient fields. We conduct comprehensive evaluation on text-image alignment performance, human value alignment performance and generation diversity performance under different divergence constraints, and the results indicate that text-to-image alignment based on Jensen-Shannon divergence achieves the best trade-off among them. The option of divergence employed for aligning text-to-image models significantly impacts the trade-off between alignment performance (especially human value alignment) and generation diversity, which highlights the necessity of selecting an appropriate divergence for practical applications.
Bo Xia, Yongzhe Chang, Xueqian Wang 0001
AAAI3
2025 Identical Human Preference Alignment Paradigm for Text-to-Image Models
abstract
Implicit reward mechanism of Direct Preference Optimization (DPO) has facilitated its recent applications beyond large language models (LLMs), notably in aligning text-to-image models with human preferences. While promising results have been achieved with algorithms such as Diffusion-DPO, their reliance on the assumptions of the Bradley-Terry model could potentially lead to significant overfitting. In this paper, we propose the Step Identical Preference Alignment (SIPA) method, departing text-to-image alignment from the assumptions of Bradley-Terry preference model. We assess the performance of four models, Diffusion-DPO, SPO, and SIPA, alongside the original model, on the HPS-V2 test set, which focus on three key aspects: text-image alignment, human value alignment, and generation diversity. Experimental results show that SIPA matches or outperforms existing SOTA alignment methods, and even exceeds the original model in terms of generation diversity, which compellingly demonstrates SIPA’s superiority in mitigating alignment overfitting.
Bo Xia, Yongzhe Chang, Xueqian Wang 0001
ICASSP4
2025 Positive Enhanced Preference Alignment for Text-to-Image Models
abstract
Direct Preference Optimization (DPO) has recently expanded its successful application beyond aligning large language models (LLMs), further targeting the alignment of text-to-image models with human preferences. However, traditional DPO approach would inadvertently result in a simultaneous reduction of sampling probabilities for preferred and dispreferred items during the alignment process, thereby potentially diminishing model's generative capacity. In this paper, we firstly undertake a revisit of DPO by grounding our analysis in the framework of contrastive loss. It reveals that DPO only emphasizes the part quantifying dissimilarity between items, while overlooking aspects pertinent to positive items. Hence, we propose the Positive Enhanced Preference Alignment (PEPA). Three enhancement strategies are introduced herein, and after comprehensive empirical evaluation, we recommend implementation of enhancing the log probability of preferred ratio in practice applications, which is distinguished by both stability and effectiveness. Experimental assessments are carried out on the HPS-V2 test set, with results demonstrating that PEPA outperforms or matches current state-of-the-art alignment techniques, thus highlighting PEPA's exceptional practical efficacy.
Bo Xia, Yongzhe Chang, Xueqian Wang 0001
ICASSP4
2025 Entropy-based Activation Function Optimization: A Method on Searching Better Activation Functions
abstract
The success of artificial neural networks (ANNs) hinges greatly on the judicious selection of an activation function, introducing non-linearity into network and enabling them to model sophisticated relationships in data. However, the search of activation functions has largely relied on empirical knowledge in the past, lacking theoretical guidance, which has hindered the identification of more effective activation functions. In this work, we offer a proper solution to such issue. Firstly, we theoretically demonstrate the existence of the worst activation function with boundary conditions (WAFBC) from the perspective of information entropy. Furthermore, inspired by the Taylor expansion form of information entropy functional, we propose the Entropy-based Activation Function Optimization (EAFO) methodology. EAFO methodology presents a novel perspective for designing static activation functions in deep neural networks and the potential of dynamically optimizing activation during iterative training. Utilizing EAFO methodology, we derive a novel activation function from ReLU, known as Correction Regularized ReLU (CRReLU). Experiments conducted with vision transformer and its variants on CIFAR-10, CIFAR-100 and ImageNet-1K datasets demonstrate the superiority of CRReLU over existing corrections of ReLU. Extensive empirical studies on task of large language model (LLM) fine-tuning, CRReLU exhibits superior performance compared to GELU, suggesting its broader potential for practical applications.
Bo Xia, Pu Chang, Zibin Dong, Yifu Yuan, Yongzhe Chang, Xueqian Wang 0001
ICLR7
2025 Learning Pre-Trained Tacit Behavior for Efficient Multi-Agent Adversarial Coordination
Shiqing Yao, Jiajun Chai, Haixin Yu, Yongzhe Chang, Yuanheng Zhu, Xueqian Wang 0001
AAMAS4
2025 Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning
abstract
Diffusion probability models have shown significant promise in offline reinforcement learning by directly modeling trajectory sequences. However, existing approaches primarily focus on time-domain features while overlooking frequency-domain features, leading to frequency shift and degraded performance according to our observation. In this paper, we investigate the RL problem from a new perspective of the frequency domain. We first observe that time-domain-only approaches inadvertently introduce shifts in the low-frequency components of the frequency domain, which results in trajectory instability and degraded performance. To address this issue, we propose Wavelet Fourier Diffuser (WFDiffuser), a novel diffusion-based RL framework that integrates Discrete Wavelet Transform to decompose trajectories into low- and high-frequency components. To further enhance diffusion modeling for each component, WFDiffuser employs Short-Time Fourier Transform and cross attention mechanisms to extract frequency-domain features and facilitate cross-frequency interaction. Extensive experiment results on the D4RL benchmark demonstrate that WFDiffuser effectively mitigates frequency shift, leading to smoother, more stable trajectories and improved decision-making performance over existing methods.
Yifu Luo, Yongzhe Chang, Xueqian Wang 0001
IJCNN2
2025 Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
abstract
Reinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an important alternative: masked generative models. In this work, we propose Mask-GRPO, the first method to incorporate Group Relative Policy Optimization (GRPO)-based RL into this overlooked paradigm. Our core insight is to redefine the transition probability, which is different from current approaches, and formulate the unmasking process as a multi-step decision-making problem. To further enhance our method, we explore several useful strategies, including removing the Kullback–Leibler constraint, applying the reduction strategy, and filtering out low-quality samples. Using Mask-GRPO, we improve a base model, Show-o, with substantial improvements on standard T2I benchmarks and preference alignment, outperforming existing state-of-the-art approaches.
Yifu Luo, Xinhao Hu, Keyu Fan, Bo Xia, Tiantian Zhang 0002, Yongzhe Chang, Xueqian Wang 0001
NeurIPS8
2025 Large Language Model Based Multi-agent Learning for Mixed Cooperative-Competitive Environments
Chenghua He, Qiyue Yin, Yongzhe Chang, Tongtong Yu
NLPCC (1)3
2024 D3D: Conditional Diffusion Model for Decision-Making Under Random Frame Dropping
abstract
The occurrence of frame drops due to issues such as corrupted communications or malfunctioning sensors presents a significant challenge to an agent’s decision-making, especially in remote control scenarios. Classical reinforcement learning (RL) usually assumes a continuous data stream without frame drops and relies heavily on online interactions, which is time-consuming, resource-intensive, and often impractical in certain scenarios. Consequently, the performance of RL may deteriorate significantly in face of non-negligible frame drops. To tackle this challenge caused by frame dropping, We propose Conditional Diffusion Model for Decision-Making under Random Frame Dropping (D3D), an offline algorithm that can effectively enhance performance robustness in frame dropping scenarios. D3D addresses this issue through a two-phase approach: 1) During the policy generation phase, D3D adopts a return-conditional diffusion model for decision making rather than the temporal difference learning, whose policy is derived using offline datasets of return-labeled trajectories without information loss. 2) When frame dropping occurs during evaluation, D3D seamlessly substitutes the missing state with its corresponding prediction in the horizon made by the diffusion model. Extensive experiments are conducted on MuJoCo and Adroit tasks to validate D3D’s robustness and efficiency. The results demonstrate that D3D consistently outperforms state-of-the-art RL algorithms, especially excelling on tasks featuring severe drop rates.
Bo Xia, Yifu Luo, Yongzhe Chang, Bo Yuan 0003, Zhiheng Li 0001, Xueqian Wang 0001
RO-MAN3
2024 Solving time-delay issues in reinforcement learning via transformers
Bo Xia, Zaihui Yang, Minzhi Xie, Yongzhe Chang, Bo Yuan 0003, Zhiheng Li 0001, Xueqian Wang 0001, Bin Liang 0001
Appl. Intell.4
2023 Addressing Delays in Reinforcement Learning via Delayed Adversarial Imitation Learning
Minzhi Xie, Bo Xia, Yalou Yu, Xueqian Wang 0001, Yongzhe Chang
ICANN (3)5
2023 Curriculum-based Co-design of Morphology and Control of Voxel-based Soft Robots
Haobo Fu, Qiang Fu 0016, Tiantian Zhang 0002, Yongzhe Chang, Xueqian Wang 0001
ICLR6
2023 Overcoming Delayed Feedback via Overlook Decision Making
abstract
Reinforcement learning is one of the most general paradigms to solve sequential decision making issues on the assumption that the action selection and environmental feedback are instantaneous, however, unfortunately this assumption is rarely true with regard to such ubiquitous delays in real-world system which could degrade the performance of reinforcement learning algorithms. The most common solution to solve a fixed delay problem is to design a forward dynamic model which is used to predict the newest state by recursively iterating over long steps so that a predicted state can be got and it would be taken as the agent's observation to make the newest decision. However, there exists cumulative errors during the iterative process which make long-term prediction inaccurate and further affect agent's decision. Motivated by the goal to reduce cumulative errors, we propose a new algorithm named Multi-step Prediction model with Delayed Observation(MPDO), aiming at accurately predicting future state at longer horizons for better decision making. Our approach includes two parts: a multi-step prediction model and a strategy training based on proximal policy optimization algorithms(PPO). Our model only needs a small amount of data to conduct dynamic modeling quickly, and the accuracy of prediction and iteration speed are higher than traditional methods. Experiments on Gym and MuJoCo show that MPDO achieves higher performance in such different tasks with different delays compared with other state-of-the-art methods, which verify our method's effectiveness.
Yalou Yu, Bo Xia, Minzhi Xie, Xueqian Wang 0001, Zhiheng Li 0001, Yongzhe Chang
SMC6
2023 EPO-S: A Constrained RL Method to Enhance UAV Safety with Spatial Representation
abstract
Path planning and collision avoidance are critical components of UAV control algorithms that play a crucial role in executing UAV missions. As scenarios become increasingly complex, the traditional control methods just ain't cutting it to meet the requirements. Reinforcement learning is an emerging decision-making control algorithm that attempts to address these issues as an alternative to traditional methods and has made significant advances. Unfortunately, standard RL approaches only aim to maximize rewards, however balancing task performance and safety in completing UAV tasks poses a challenge since these two objectives sometimes conflict, leading to a trade-off often difficult to manage. This paper proposes three techniques to address this problem. First, we model the path planning and collision avoidance issue in a constrained RL framework, eliminating the need for complex reward engineering. Second, we expand our previous work in the UAV setting and introduce an exact penalty optimization (EPO) algorithm to provide stricter constraint guarantees. We also propose a novel spatial information representation method for the UAV scenario to help UAVs better understand environmental information. The experimental results demonstrate the effectiveness of the EPO and spatial representation modules proposed in this paper, through a significant reduction in collisions as well as a strong improvement in the rate of reaching the destination.
Linrui Zhang, Zaihui Yang, Haoyu Wang 0018, Xueqian Wang 0001, Yongzhe Chang
SMC6
2022 A surrogate-assisted controller for expensive evolutionary reinforcement learning
Tiantian Zhang 0002, Yongzhe Chang, Xueqian Wang 0001, Bin Liang 0001, Bo Yuan 0003
Inf. Sci.3
2020 A Dual Input-aware Factorization Machine for CTR Prediction
abstract
Factorization Machines (FMs) refer to a class of general predictors working with real valued feature vectors, which are well-known for their ability to estimate model parameters under significant sparsity and have found successful applications in many areas such as the click-through rate (CTR) prediction. However, standard FMs only produce a single fixed representation for each feature across different input instances, which may limit the CTR model’s expressive and predictive power. Inspired by the success of Input-aware Factorization Machines (IFMs), which aim to learn more flexible and informative representations of a given feature according to different input instances, we propose a novel model named Dual Input-aware Factorization Machines (DIFMs) that can adaptively reweight the original feature representations at the bit-wise and vector-wise levels simultaneously. Furthermore, DIFMs strategically integrate various components including Multi-Head Self-Attention, Residual Networks and DNNs into a unified end-to-end model. Comprehensive experiments on two real-world CTR prediction datasets show that the DIFM model can outperform several state-of-the-art models consistently.
Wantong Lu, Yongzhe Chang, Zhen Wang 0030, Bo Yuan 0003
IJCAI3
2019 Exploring Latent Structure Similarity for Bayesian Nonparameteric Model with Mixture of NHPP Sequence
Yongzhe Chang, Zhidong Li, Ling Luo 0002, Simon Luo, Arcot Sowmya, Yang Wang 0002, Fang Chen 0001
ICONIP (2)1
2019 Recovering DTW Distance Between Noise Superposed NHPP
Yongzhe Chang, Zhidong Li, Bang Zhang, Ling Luo 0002, Arcot Sowmya, Yang Wang 0002, Fang Chen 0001
PAKDD (2)1