Yunjian Xu

dblp:09/3682 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0003-0110-8032ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
abstract
Hallucination remains a major challenge for Large Vision-Language Models (LVLMs). Direct Preference Optimization (DPO) has gained increasing attention as a simple solution to hallucination issues. It directly learns from constructed preference pairs that reflect the severity of hallucinations in responses to the same prompt and image. Nonetheless, different data construction methods in existing works bring notable performance variations. We identify a crucial factor here: outcomes are largely contingent on whether the constructed data aligns on-policy w.r.t the initial (reference) policy of DPO. Theoretical analysis suggests that learning from off-policy data is impeded by the presence of KL-divergence between the updated policy and the reference policy. From the perspective of dataset distribution, we systematically summarize the inherent flaws in existing algorithms that employ DPO to address hallucination issues. To alleviate the problems, we propose On-Policy Alignment (OPA)-DPO framework, which uniquely leverages expert feedback to correct hallucinated responses and aligns both the original and expert-revised responses in an on-policy manner. Notably, with only 4.8k data, OPADPO achieves an additional reduction in the hallucination rate of LLaVA-1.5-7B: 13.26% on the AMBER benchmark and 5.39% on the Object-Hal benchmark, compared to the previous SOTA algorithm trained with 16k samples.
Zhihe Yang, Xufang Luo, Yunjian Xu, Dongsheng Li 0002
CVPR4
2025 Q-Supervised Contrastive Representation: A State Decoupling Framework for Safe Offline Reinforcement Learning
abstract
Safe offline reinforcement learning (RL), which aims to learn the safety-guaranteed policy without risky online interaction with environments, has attracted growing recent attention for safety-critical scenarios. However, existing approaches encounter out-of-distribution problems during the testing phase, which can result in potentially unsafe outcomes. This issue arises due to the infinite possible combinations of reward-related and cost-related states. In this work, we propose State Decoupling with Q-supervised Contrastive representation (SDQC), a novel framework that decouples the global observations into reward- and cost-related representations for decision-making, thereby improving the generalization capability for unfamiliar global observations. Compared with the classical representation learning methods, which typically require model-based estimation (e.g., bisimulation), we theoretically prove that our Q-supervised method generates a coarser representation while preserving the optimal policy, resulting in improved generalization performance. Experiments on DSRL benchmark problems provide compelling evidence that SDQC surpasses other baseline algorithms, especially for its exceptional ability to achieve almost zero violations in more than half of the tasks, while the state-of-the-art algorithm can only achieve the same level of success in a quarter of the tasks. Further, we demonstrate that SDQC possesses superior generalization ability when confronted with unseen environments.
Zhihe Yang, Yunjian Xu
ICML2
2025 ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning
abstract
Real-world datasets collected from sensors or human inputs are prone to noise and errors, posing significant challenges for applying offline reinforcement learning (RL). While existing methods have made progress in addressing corrupted actions and rewards, they remain insufficient for handling corruption in high-dimensional state spaces and for cases where multiple elements in the dataset are corrupted simultaneously. Diffusion models, known for their strong denoising capabilities, offer a promising direction for this problem—but their tendency to overfit noisy samples limits their direct applicability. To overcome this, we propose **A**mbient **D**iffusion-**G**uided Dataset Recovery (**ADG**), a novel approach that pioneers the use of diffusion models to tackle data corruption in offline RL. First, we introduce Ambient Denoising Diffusion Probabilistic Models (DDPM) from approximated distributions, which enable learning on partially corrupted datasets with theoretical guarantees. Second, we use the noise-prediction property of Ambient DDPM to distinguish between clean and corrupted data, and then use the clean subset to train a standard DDPM. Third, we employ the trained standard DDPM to refine the previously identified corrupted data, enhancing data quality for subsequent offline RL training. A notable strength of ADG is its versatility—it can be seamlessly integrated with any offline RL algorithm. Experiments on a range of benchmarks, including MuJoCo, Kitchen, and Adroit, demonstrate that ADG effectively mitigates the impact of corrupted data and improves the robustness of offline RL under various noise settings, achieving state-of-the-art results.
Zeyuan Liu, Zhihe Yang, Rui Yang 0010, Jiafei Lyu, Baoxiang Wang 0001, Yunjian Xu, Xiu Li 0001
NeurIPS7
2025 MOSDT: Self-Distillation-Based Decision Transformer for Multi-Agent Offline Safe Reinforcement Learning
abstract
We introduce MOSDT, the first algorithm designed for multi-agent offline safe reinforcement learning (MOSRL), alongside MOSDB, the first dataset and benchmark for this domain. Different from most existing knowledge distillation-based multi-agent RL methods, we propose policy self-distillation (PSD) with a new global information reconstruction scheme by fusing the observation features of all agents, streamlining training and improving parameter efficiency. We adopt full parameter sharing across agents, significantly slashing parameter count and boosting returns up to 38.4-fold by stabilizing training. We propose a new plug-and-play cost binary embedding (CBE) module, which binarizes cumulative costs as safety signals and embeds the signals into return features for efficient information aggregation. On the strong MOSDB benchmark, MOSDT achieves state-of-the-art (SOTA) returns in 14 out of 18 tasks (across all base environments including MuJoCo, Safety Gym, and Isaac Gym) while ensuring complete safety, with only 65% of the execution parameter count of a SOTA single-agent offline safe RL method CDT. Code, dataset, and results are available at this website: https://github.com/Lucian1115/MOSDT.git
Yuchen Xia, Yunjian Xu
NeurIPS2
2024 DMBP: Diffusion model-based predictor for robust offline reinforcement learning against state observation perturbations
abstract
Offline reinforcement learning (RL), which aims to fully explore offline datasets for training without interaction with environments, has attracted growing recent attention. A major challenge for the real-world application of offline RL stems from the robustness against state observation perturbations, e.g., as a result of sensor errors or adversarial attacks. Unlike online robust RL, agents cannot be adversarially trained in the offline setting. In this work, we propose Diffusion Model-Based Predictor (DMBP) in a new framework that recovers the actual states with conditional diffusion models for state-based RL tasks. To mitigate the error accumulation issue in model-based estimation resulting from the classical training of conventional diffusion models, we propose a non-Markovian training objective to minimize the sum entropy of denoised states in RL trajectory. Experiments on standard benchmark problems demonstrate that DMBP can significantly enhance the robustness of existing offline RL algorithms against different scales of ran- dom noises and adversarial attacks on state observations. Further, the proposed framework can effectively deal with incomplete state observations with random combinations of multiple unobserved dimensions in the test. Our implementation is available at https://github.com/zhyang2226/DMBP.
Zhihe Yang, Yunjian Xu
ICLR2
2023 Short-Term EV Charging Load Predicting Based on Adaptive VMD and LSTM Methods
abstract
The uncoordinated charging of large-scale electric vehicles (EVs) generally deteriorates the peak-valley difference of daily electric demands. To facilitate the operation of charging stations and electric power distributers, this work proposes a charging load prediction algorithm by combining the Variational Mode Decomposition (VMD) and the Long Short-Term Memory (LSTM) methods. The VMD is adopted to extract the EV charging load features at different time scales, obtaining multiple intrinsic mode functions (IMFs). Then the LSTM establishes the dependencies between these IMFs of historical data and the predicted load. To trade-off between the prediction accuracy and the computation overhead, an additional Snake Optimization (SO) technique is applied to adaptively optimize the VMD parameters. Experimental results show that the proposed algorithm outperforms the traditional LSTM alone and the Gate Recurrent Unit alone neural networks in terms of the overall prediction accuracy. The proposed parallel LSTM structure with the optimized VMD further reduces the Root Mean Square Error (RMSE) and the Mean Absolute Error (MAE) significantly by 55.1% and 55.9% with respect to the LSTM method with non-optimized VMD.
Quanxue Guan, Qinhe Liu, Yunjian Xu, Xiaojun Tan
IECON4
2023 Semisupervised Learning-Based Occupancy Estimation for Real-Time Energy Management Using Ambient Data
abstract
Occupancy status information is essential for efficient energy management of deferrable and flexible electricity loads in both residential and commercial sectors. Motivated by the hardness to collect ground-truth occupancy data, we propose a new, semisupervised learning-based occupancy estimator, the label-expanded time-series propensity weighted estimator (LTPWE), to improve the prediction accuracy using privacy-preserving ambient sensing data (in temperature, humidity, light, etc.), especially when the labeling frequency (the ratio of the labeled positive instances) is low. The proposed energy management scheme integrates the LTPWE occupancy estimator into a data-driven deep reinforcement learning algorithm, the soft actor–critic (SAC), to make real-time scheduling decisions without any prior knowledge on the dynamics of random renewable generation and electricity price. Simulation results on real-world data sets show that compared with multiple state-of-the-art baselines, the proposed energy management scheme can reduce the energy cost by 18.79%–55.79% without sacrificing occupants’ thermal comfort.
Liangliang Hao, Yunjian Xu
IEEE Internet Things J.2
2023 Energy and Reserve Sharing Considering Uncertainty and Communication Resources
abstract
In this article, we study the joint energy and reserve sharing problem considering renewable generation uncertainty and limited communication resources. We propose a data-driven distributionally robust energy and reserve sharing model among different agents in electricity markets. We put forward data-driven distributionally robust chance constraints (DRCC) to determine the reserve capacity, which cannot be directly solved. The inner approximation is employed to convert the DRCC into tractable linear constraints. Taking into account the agents in the Internet of Things exchange information by a resource-limited communication network, we develop a communication-censored consensus alternating direction method of multipliers (ADMMs) to utilize the limited communication resources and solve the sharing problem in a fully decentralized manner. We analyze the convergence of the proposed algorithm, and propose an adaptive penalty parameter method to speed up the convergence. Extensive simulations are conducted to verify the effectiveness of the proposed model and theoretical results.
Wenjie Liu 0013, Yunjian Xu, Wenqian Yin, Yunhe Hou, Zaiyue Yang
IEEE Internet Things J.2
2023 Structural Charging and Replenishment Policies for Battery Swapping Charging System Operation Under Uncertainty
abstract
We study the joint battery charging and replenishment scheduling of a battery swapping charging system (BSCS) considering random electric vehicle (EV) arrivals, renewable generation, and electricity prices. We formulate the problem as a Markov decision process with an objective to minimize the expected sum of the operation cost (battery charging and replenishment cost) and the waiting cost of EV customers. The joint scheduling problem is challenging due to the stochasticity in EV arrivals, renewable generation, and electricity prices, as well as the curse of dimensionality in the system state and action spaces. To reduce the dimension of the action space, we propose to integrate structural properties into BSCS operation, i.e., the threshold-charging (TC) and least demand first (LDF) structures into the charging policy, and the$(s, S)$structure into the replenishment policy (when the number of fully-charged batteries at a battery swapping station is below$s$, the inventory is replenished to a higher threshold$S$). Numerical experiments on real-world data show that the proposed SAC+TC+$(s, S)$approach saves 7.16%-78.61% and 6.53%-93.73% of total average cost resulting from various structural charging and replenishment policies and the vanilla soft actor-critic (SAC) algorithm under different settings.
Yan Wang 0067, Quanxue Guan, Yunjian Xu
IEEE Trans. Intell. Transp. Syst.4
2022 Optimal Policy Characterization Enhanced Proximal Policy Optimization for Multitask Scheduling in Cloud Computing
abstract
For a serving system with multiple servers and a public queue, we study the scheduling of multiple tasks with deadlines, under random task arrivals and renewable energy generation. To minimize the weighted sum of the serving cost (associated with the energy consumption) and the delay cost (resulting from deferring the processing of tasks after their deadlines), we formulate the problem as a dynamic program with unknown transition probability. To mitigate the curse of dimensionality, we establish a partial priority rule, the earlier deadline and less demand first (ED-LDF): priority should be given to tasks with earlier deadline and less demand. In the heavy-traffic regime, the established ED-LDF characterization is proved to be optimal under arbitrary system dynamics. We propose a new, scalable ED-LDF-based proximal policy optimization (PPO) approach that integrates our (partial) optimal policy characterizations into the state-of-the-art deep reinforcement learning (DRL) algorithm. Numerical results demonstrate that the proposed ED-LDF-based PPO approach outperforms the classical PPO and three other priority rule-based PPO approaches.
Jiangliang Jin, Yunjian Xu
IEEE Internet Things J.2
2022 Shortest-Path-Based Deep Reinforcement Learning for EV Charging Routing Under Stochastic Traffic Condition and Electricity Prices
abstract
We study the charging routing problem faced by a smart electric vehicle (EV) that looks for an EV charging station (EVCS) to fulfill its battery charging demand. Leveraging the real-time information from both the power and intelligent transportation systems, the EV seeks to minimize the sum of the travel cost (to a selected EVCS) and the charging cost (a weighted sum of electricity cost and waiting time cost at the EVCS). We formulate the problem as a Markov decision process with unknown dynamics of system uncertainties (in traffic conditions, charging prices, and waiting time). To mitigate the curse of dimensionality, we reformulate the deterministic charging routing problem (a mixed-integer program) as a two-level shortest-path (SP)-based optimization problem that can be solved in polynomial time. Its low dimensional solution is input into a state-of-the-art deep reinforcement learning (DRL) algorithm, the advantage actor–critic (A2C) method, to make efficient online routing decisions. Numerical results (on a real-world transportation network) demonstrate that the proposed SP-based A2C approach outperforms the classical A2C method and two alternative SP-based DRL methods.
Jiangliang Jin, Yunjian Xu
IEEE Internet Things J.2
2022 Laxity Differentiated Pricing and Deadline Differentiated Threshold Scheduling for a Public Electric Vehicle Charging Station
abstract
In this article, we study the online pricing and charging scheduling problem for a public electric vehicle (EV) charging station under stochastic electricity prices and renewable generation. We formulate the sequential decision making problem as a partially observable Markov decision process with continuous state and action spaces, with an objective of profit or social welfare maximization. The joint pricing and charging problem is challenging due to the curse of dimensionality (in both the system state and action spaces) and the unknown dynamics of system uncertainties. We propose a novel laxity differentiated pricing (LDP) scheme to tradeoff between electricity cost (associated with EV charging) and opportunity cost (associated with parking infrastructure occupancy). Shown to be optimal under arbitrary system dynamics, the characterized deadline-differentiated threshold charging (DTC) policy is integrated into a model-free soft actor critic (SAC) algorithm to reduce the action dimensionality. Numerical results demonstrate that the proposed SAC + LDP + DTC approach significantly outperforms alternative methods with various pricing and charging schemes.
Liangliang Hao, Jiangliang Jin, Yunjian Xu
IEEE Trans. Ind. Informatics3
2015 Commitment in First-Price Auctions
Yunjian Xu, Katrina Ligett
SAGT1
2014 Spectrum sharing in frequency-selective unlicensed bands: a game theoretic approach
abstract
Power allocation is an important issue for spectrum sharing of unlicensed bands, in which multiple unlicensed systems may coexist and operate. Some recent works have been reported on spectrum sharing in frequency-flat FF unlicensed bands. However, there has not been much work on power allocation for spectrum sharing in frequency-selective FS unlicensed bands. For multiple cooperative systems cooperating on FS interference channels ICs, we study an optimal power allocation strategy, which allows the transmission power density to vary within one subcarrier. By showing the duality of FS and parallel FF channels, we can therefore compute the achievable rate region of the proposed strategy when systems cooperate with each other. For non-cooperative scenarios, we construct a game-theoretical framework for multiple selfish systems on FS ICs. This framework enables us to utilize existing protocols designed for FF ICs to FS scenarios. By numerical results, in both cooperative and non-cooperative scenarios, we show that the proposed strategy achieves a larger rate region than a conventional strategy, where the transmission power density on each subcarrier is set equal. Our work can be regarded as an extension of previous works for FF scenarios. Copyright © 2012 John Wiley & Sons, Ltd.
Yunjian Xu, Wei Chen 0002, Zhigang Cao 0001
Wirel. Commun. Mob. Comput.1
2008 Game-Theoretic Analysis for Power Allocation in Frequency-Selective Unlicensed Bands
abstract
Power allocation is an important issue for spectrum sharing of unlicensed bands, in which multiple unlicensed systems may coexist and operate. Recently some works have been reported on game theoretical analysis for multiple systems cooperating in frequency-flat unlicensed bands. However, there has not been much work on the cooperative and competitive strategic behavior of multiple mutually interfering systems in frequency-selective unlicensed bands. In this paper, we construct a game theoretical framework for multiple selfish systems on frequency-selective Interference Channels (IC). This framework enables us to utilize existing protocols designed for frequency-flat ICs in frequency-selective scenarios and can be regarded as an extension of previous results for frequency-flat scenarios.
Yunjian Xu, Wei Chen 0002, Zhigang Cao 0001, Khaled Ben Letaief
GLOBECOM1
2008 A Distributed Random Access Protocol with Enhanced Routing in Time-Slotted MANETs
abstract
In Mobile Ad hoc NETworks (MANETs), random access and dynamic routing are two critical techniques for mobile nodes to convey information without centralized scheduling. Conventionally, random access and dynamic routing are implemented at the Medium Access Control (MAC) and the network layer, respectively. However, the current MAC protocol Carrier Sense Multiple Access with Collision Avoidance (CSMA/CA) cannot support dynamic routing efficiently. To overcome this limitation, we shall propose a cross-layer protocol which takes dynamic routing into consideration when the mobile nodes contend to access the channel. In the proposed distributed protocol, the routing packets of the network layer are transmitted within the contention period. Since the transmission of routing packets is separated from data transmission in the time-domain, our cross- layer protocol eliminates the collision caused by the transmission of short routing packets. Simulation results will show that our design could significantly improve the system performance at both the MAC and network layers.
Yunjian Xu, Wei Chen 0002, Zhigang Cao 0001, Khaled Ben Letaief
GLOBECOM1