Bo Xia

dblp:76/6557 · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 4 · 3 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences Through f-Divergence Minimization
abstract
Direct Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable interest within the community. However, we have observed that these approaches rely solely on minimizing the reverse Kullback-Leibler divergence during alignment process between the fine-tuned model and the reference model, neglecting incorporation of other divergence constraints. In this study, we focus on extending reverse Kullback-Leibler divergence in the alignment paradigm of text-to-image models to f-divergence, which aims to garner better alignment performance as well as good generation diversity. We provide the generalized formula of text-to-image alignment paradigm under f-divergence condition and thoroughly analyze the impact of different divergence constraints on alignment process from the perspective of gradient fields. We conduct comprehensive evaluation on text-image alignment performance, human value alignment performance and generation diversity performance under different divergence constraints, and the results indicate that text-to-image alignment based on Jensen-Shannon divergence achieves the best trade-off among them. The option of divergence employed for aligning text-to-image models significantly impacts the trade-off between alignment performance (especially human value alignment) and generation diversity, which highlights the necessity of selecting an appropriate divergence for practical applications.
Bo Xia, Yongzhe Chang, Xueqian Wang 0001
AAAI2
2025 Identical Human Preference Alignment Paradigm for Text-to-Image Models
abstract
Implicit reward mechanism of Direct Preference Optimization (DPO) has facilitated its recent applications beyond large language models (LLMs), notably in aligning text-to-image models with human preferences. While promising results have been achieved with algorithms such as Diffusion-DPO, their reliance on the assumptions of the Bradley-Terry model could potentially lead to significant overfitting. In this paper, we propose the Step Identical Preference Alignment (SIPA) method, departing text-to-image alignment from the assumptions of Bradley-Terry preference model. We assess the performance of four models, Diffusion-DPO, SPO, and SIPA, alongside the original model, on the HPS-V2 test set, which focus on three key aspects: text-image alignment, human value alignment, and generation diversity. Experimental results show that SIPA matches or outperforms existing SOTA alignment methods, and even exceeds the original model in terms of generation diversity, which compellingly demonstrates SIPA’s superiority in mitigating alignment overfitting.
Bo Xia, Yongzhe Chang, Xueqian Wang 0001
ICASSP2
2025 Positive Enhanced Preference Alignment for Text-to-Image Models
abstract
Direct Preference Optimization (DPO) has recently expanded its successful application beyond aligning large language models (LLMs), further targeting the alignment of text-to-image models with human preferences. However, traditional DPO approach would inadvertently result in a simultaneous reduction of sampling probabilities for preferred and dispreferred items during the alignment process, thereby potentially diminishing model's generative capacity. In this paper, we firstly undertake a revisit of DPO by grounding our analysis in the framework of contrastive loss. It reveals that DPO only emphasizes the part quantifying dissimilarity between items, while overlooking aspects pertinent to positive items. Hence, we propose the Positive Enhanced Preference Alignment (PEPA). Three enhancement strategies are introduced herein, and after comprehensive empirical evaluation, we recommend implementation of enhancing the log probability of preferred ratio in practice applications, which is distinguished by both stability and effectiveness. Experimental assessments are carried out on the HPS-V2 test set, with results demonstrating that PEPA outperforms or matches current state-of-the-art alignment techniques, thus highlighting PEPA's exceptional practical efficacy.
Bo Xia, Yongzhe Chang, Xueqian Wang 0001
ICASSP2
2025 Entropy-based Activation Function Optimization: A Method on Searching Better Activation Functions
abstract
The success of artificial neural networks (ANNs) hinges greatly on the judicious selection of an activation function, introducing non-linearity into network and enabling them to model sophisticated relationships in data. However, the search of activation functions has largely relied on empirical knowledge in the past, lacking theoretical guidance, which has hindered the identification of more effective activation functions. In this work, we offer a proper solution to such issue. Firstly, we theoretically demonstrate the existence of the worst activation function with boundary conditions (WAFBC) from the perspective of information entropy. Furthermore, inspired by the Taylor expansion form of information entropy functional, we propose the Entropy-based Activation Function Optimization (EAFO) methodology. EAFO methodology presents a novel perspective for designing static activation functions in deep neural networks and the potential of dynamically optimizing activation during iterative training. Utilizing EAFO methodology, we derive a novel activation function from ReLU, known as Correction Regularized ReLU (CRReLU). Experiments conducted with vision transformer and its variants on CIFAR-10, CIFAR-100 and ImageNet-1K datasets demonstrate the superiority of CRReLU over existing corrections of ReLU. Extensive empirical studies on task of large language model (LLM) fine-tuning, CRReLU exhibits superior performance compared to GELU, suggesting its broader potential for practical applications.
Bo Xia, Pu Chang, Zibin Dong, Yifu Yuan, Yongzhe Chang, Xueqian Wang 0001
ICLR3
2025 DeepMF: Deep Motion Factorization for Closed-Loop Safety-Critical Driving Scenario Simulation
abstract
Safety-critical traffic scenarios are of great practical relevance to evaluating the robustness of autonomous driving (AD) systems. Given that these long-tail events are extremely rare in real-world traffic data, there is a growing body of work dedicated to the automatic traffic scenario generation. However, nearly all existing algorithms for generating safety-critical scenarios rely on snippets of previously recorded traffic events, transforming normal traffic flow into accident-prone situations directly. In other words, safety-critical traffic scenario generation is hindsight and not applicable to newly encountered and open-ended traffic events. In this paper, we propose the Deep Motion Factorization (DeepMF) framework, which extends static safety-critical driving scenario generation to closed-loop and interactive adversarial traffic simulation. DeepMF casts safety-critical traffic simulation as a Bayesian factorization that includes the assignment of hazardous traffic participants, the motion prediction of selected opponents, the reaction estimation of autonomous vehicle (AV) and the probability estimation of the accident occur. All the aforementioned terms are calculated using decoupled deep neural networks, with inputs limited to the current observation and historical states. Consequently, DeepMF can effectively and efficiently simulate safety-critical traffic scenarios at any triggered time and for any duration by maximizing the compounded posterior probability of traffic risk. Extensive experiments demonstrate that DeepMF excels in terms of risk management, flexibility, and diversity, showcasing outstanding performance in simulating a wide range of realistic, high-risk traffic scenarios.
Linrui Zhang, Bo Xia, Xueqian Wang 0001, Houde Liu
IJCNN3
2025 Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
abstract
Reinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an important alternative: masked generative models. In this work, we propose Mask-GRPO, the first method to incorporate Group Relative Policy Optimization (GRPO)-based RL into this overlooked paradigm. Our core insight is to redefine the transition probability, which is different from current approaches, and formulate the unmasking process as a multi-step decision-making problem. To further enhance our method, we explore several useful strategies, including removing the Kullback–Leibler constraint, applying the reduction strategy, and filtering out low-quality samples. Using Mask-GRPO, we improve a base model, Show-o, with substantial improvements on standard T2I benchmarks and preference alignment, outperforming existing state-of-the-art approaches.
Yifu Luo, Xinhao Hu, Keyu Fan, Bo Xia, Tiantian Zhang 0002, Yongzhe Chang, Xueqian Wang 0001
NeurIPS6
2025 Dual-path aggregation transformer network for super-resolution with images occlusions and variability
Qinghui Chen, Lunqian Wang, Xinghua Wang 0008, Bo Xia, Hao Ding 0014, Jinglin Zhang 0001
Eng. Appl. Artif. Intell.6
2025 A delay-robust method for enhanced real-time reinforcement learning
Bo Xia, Bo Yuan 0003, Zhiheng Li 0001, Bin Liang 0001, Xueqian Wang 0001
Neural Networks1
2025 Agent-Based Space Teleoperation: Mitigating Time Delays With Deep Reinforcement Learning
abstract
Space teleoperation significantly extends human reach in space missions. However, traditional approaches are constrained by factors, such as the reliance on accurate dynamic models and the risk of operator fatigue during prolonged tasks. Additionally, while data-driven intelligent approaches reduce the need for prior knowledge, they have yet to adequately address the time delay issues inherent in these systems. To overcome these challenges, we introduce the belief state actor-critic (BSAC) method, the first deep reinforcement learning approach tailored for space teleoperation capture tasks within a bilateral control framework. We first establish a generalized agent-based architecture for space teleoperation, shifting decision-making from human operators to autonomous agents. Following a comprehensive analysis of the time delay challenges, we propose the BSAC algorithm, which integrates state augmentation and belief state techniques to mitigate the effects of delays in teleoperated Markov decision processes. Extensive experiments are conducted on the MuJoCo simulation platform, modeling a real hardware system across various scenarios. The learned policies are then successfully transferred and validated in a real-world setup, demonstrating the effectiveness and robustness of BSAC. In summary, our results support the feasibility of agent-based frameworks capable of overcoming time delay challenges in space teleoperation.
Bo Xia, Xianru Tian, Bo Yuan 0003, Chunju Yang, Zhiheng Li 0001, Bin Liang 0001, Xueqian Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2024 Dynamic Modeling for Reinforcement Learning with Random Delay
Yalou Yu, Bo Xia, Minzhi Xie, Zhiheng Li 0001, Xuwqian Wang
ICANN (4)2
2024 D3D: Conditional Diffusion Model for Decision-Making Under Random Frame Dropping
abstract
The occurrence of frame drops due to issues such as corrupted communications or malfunctioning sensors presents a significant challenge to an agent’s decision-making, especially in remote control scenarios. Classical reinforcement learning (RL) usually assumes a continuous data stream without frame drops and relies heavily on online interactions, which is time-consuming, resource-intensive, and often impractical in certain scenarios. Consequently, the performance of RL may deteriorate significantly in face of non-negligible frame drops. To tackle this challenge caused by frame dropping, We propose Conditional Diffusion Model for Decision-Making under Random Frame Dropping (D3D), an offline algorithm that can effectively enhance performance robustness in frame dropping scenarios. D3D addresses this issue through a two-phase approach: 1) During the policy generation phase, D3D adopts a return-conditional diffusion model for decision making rather than the temporal difference learning, whose policy is derived using offline datasets of return-labeled trajectories without information loss. 2) When frame dropping occurs during evaluation, D3D seamlessly substitutes the missing state with its corresponding prediction in the horizon made by the diffusion model. Extensive experiments are conducted on MuJoCo and Adroit tasks to validate D3D’s robustness and efficiency. The results demonstrate that D3D consistently outperforms state-of-the-art RL algorithms, especially excelling on tasks featuring severe drop rates.
Bo Xia, Yifu Luo, Yongzhe Chang, Bo Yuan 0003, Zhiheng Li 0001, Xueqian Wang 0001
RO-MAN1
2024 Solving time-delay issues in reinforcement learning via transformers
Bo Xia, Zaihui Yang, Minzhi Xie, Yongzhe Chang, Bo Yuan 0003, Zhiheng Li 0001, Xueqian Wang 0001, Bin Liang 0001
Appl. Intell.1
2023 Addressing Delays in Reinforcement Learning via Delayed Adversarial Imitation Learning
Minzhi Xie, Bo Xia, Yalou Yu, Xueqian Wang 0001, Yongzhe Chang
ICANN (3)2
2023 ULKNet:Rethinking Large Kernel CNN with UNet-Attention for Remote Sensing Images Semantic Segmentation
abstract
Since the advent of the Transformer, the Vision Transformer (ViT) model has become the optimal solution for extracting semantic information in remote sensing image semantic segmentation tasks. However, recent research has discovered that using larger convolutional kernels to extract global information can achieve performance comparable to, or even better than the ViT models. This finding inspires us to redesign the convolutional neural network structure and innovatively construct a novel parallel global-local module based on large kernel convolution, in order to better extract semantic information and local contextual information. Following this design approach, we propose ULKNet, a pure CNN architecture that achieves performance comparable to the ViT model. Through experiments on the ISPRS Vaihingen and Potsdam datasets, We respectively achieve mIoU 82.7% and 86.0% mIoU scores. This finding demonstrates that a well-designed large kernel CNN network structure can provide a larger and more effective receptive field, thereby better extracting semantic information.
Lunqian Wang, Bo Xia
IECON5
2023 Overcoming Delayed Feedback via Overlook Decision Making
abstract
Reinforcement learning is one of the most general paradigms to solve sequential decision making issues on the assumption that the action selection and environmental feedback are instantaneous, however, unfortunately this assumption is rarely true with regard to such ubiquitous delays in real-world system which could degrade the performance of reinforcement learning algorithms. The most common solution to solve a fixed delay problem is to design a forward dynamic model which is used to predict the newest state by recursively iterating over long steps so that a predicted state can be got and it would be taken as the agent's observation to make the newest decision. However, there exists cumulative errors during the iterative process which make long-term prediction inaccurate and further affect agent's decision. Motivated by the goal to reduce cumulative errors, we propose a new algorithm named Multi-step Prediction model with Delayed Observation(MPDO), aiming at accurately predicting future state at longer horizons for better decision making. Our approach includes two parts: a multi-step prediction model and a strategy training based on proximal policy optimization algorithms(PPO). Our model only needs a small amount of data to conduct dynamic modeling quickly, and the accuracy of prediction and iteration speed are higher than traditional methods. Experiments on Gym and MuJoCo show that MPDO achieves higher performance in such different tasks with different delays compared with other state-of-the-art methods, which verify our method's effectiveness.
Yalou Yu, Bo Xia, Minzhi Xie, Xueqian Wang 0001, Zhiheng Li 0001, Yongzhe Chang
SMC2
2022 Don't Touch What Matters: Task-Aware Lipschitz Data Augmentation for Visual Reinforcement Learning
abstract
One of the key challenges in visual Reinforcement Learning (RL) is to learn policies that can generalize to unseen environments. Recently, data augmentation techniques aiming at enhancing data diversity have demonstrated proven performance in improving the generalization ability of learned policies. However, due to the sensitivity of RL training, naively applying data augmentation, which transforms each pixel in a task-agnostic manner, may suffer from instability and damage the sample efficiency, thus further exacerbating the generalization performance. At the heart of this phenomenon is the diverged action distribution and high-variance value estimation in the face of augmented images. To alleviate this issue, we propose Task-aware Lipschitz Data Augmentation (TLDA) for visual RL, which explicitly identifies the task-correlated pixels with large Lipschitz constants, and only augments the task-irrelevant pixels for stability. We verify the effectiveness of our approach on DeepMind Control suite, CARLA and DeepMind Manipulation tasks. The extensive empirical results show that TLDA improves both sample efficiency and generalization; it outperforms previous state-of-the-art methods across 3 different visual control benchmarks.
Zhecheng Yuan, Guozheng Ma, Yao Mu 0001, Bo Xia, Bo Yuan 0003, Xueqian Wang 0001, Ping Luo 0002, Huazhe Xu
IJCAI4
2021 Real-time indirect illumination by virtual planar area lights
Bo Xia, Guanyu Xing, Yanli Liu 0002, Yanci Zhang
Comput. Graph.2
2018 Distributed Representation of Chinese Collocation
Bo Xia, Endong Xun
PACLIC1
2007 A Waiting-Time Dependent Backoff Algorithm for QoS Enhancement of Voice over WLAN
abstract
As the integration of prevailing Wireless Local Area Network (WLAN) and Voice over Internet Protocol (VoIP) technology, Voice over WLAN (VoWLAN) is expected to experience a dramatic growth in the near future. However, the widely deployed IEEE 802.11 contention-based Media Access Control (MAC) mechanism is unsuitable for voice transmission with strict Quality of Service (QoS) requirements. In this paper, the delay jitter and collision problems in Binary Exponential Backoff (BEB) algorithm are analyzed first. Then, a novel "Waiting-time Dependent Backoff (WDB)" algorithm is proposed to cope with these problems and improve QoS of voice transmission. Simulation results show that WDB can dramatically reduce delay jitter and collision probability, as well as end-to-end delay, while maintaining acceptable drop rate.
Jing Chi, Bo Xia, Ruozhou Lin
PIMRC3
2005 Quasi-cyclic codes from extended difference families
abstract
Quasi-cyclic codes are renowned for their structural design and low-complexity shift register encoding. However, existing design techniques based on difference families have limited code size options due to algebraic constraints. We introduce in this paper a notation of extended difference family (EDF) and provide a systematic code design based on EDFs with a high degree of flexibility in code size. Short/medium-length codes of high rate (/spl ges/ 0.9) can be easily constructed.
Bo Xia
WCNC2
2005 Importance sampling for tanner trees
abstract
This correspondence presents an importance sampling (IS) simulation scheme for the soft iterative decoding on loop-free multiple-layer trees. It is shown that this scheme is asymptotically efficient in that, for an arbitrary tree and a given estimation precision, the required number of samples is inversely proportional to the noise standard deviation. This work has its application in the simulation of low-density parity-check (LDPC) codes.
Bo Xia, William E. Ryan
IEEE Trans. Inf. Theory1
2004 Estimating LDPC codeword error rates via importance sampling
abstract
This paper proposes an importance sampling design for the codeword error rate estimation of low-density parity-check codes. An error set partitioning method is introduced to reduce the error boundary complexity and to improve the simulation efficiency. This is a follow-up to the work in (B. Xia et al., 2003).
Bo Xia, William E. Ryan
ISIT1
2003 On importance sampling for linear block codes
abstract
We introduce an importance sampling scheme for linear block codes with message-passing decoding. This novel scheme overcomes an existing difficulty in the IS practice that requires codebook information. Experiments show large IS gains for single parity-check codes and short-length block codes. For medium-length block codes, IS gains in the order of 10/sup 3/ and higher are observed at high signal-to-noise ratio.
Bo Xia, William E. Ryan
ICC1
2002 Near-optimal convolutionally coded asynchronous CDMA with iterative multiuser detection/decoding
abstract
We consider in this paper a convolutionally coded asynchronous CDMA system with iterative multiuser detection and decoding. Each user's bits are convolutionally encoded, interleaved, and differentially encoded before being transmitted over the multiuser CDMA channel. In contrast to previous work in this area, a differential encoder is inserted to effect an interleaver gain. We view the asynchronous CDMA channel as a periodically time-varying ISI channel. The receiver jointly decodes the differential encoders and the CDMA channel with a combined trellis, and shares soft output information with the convolutional decoders in an iterative (turbo) fashion. Dramatic gains over conventional convolutionally coded systems are demonstrated via simulation. Near single-user turbo performance is obtained with a moderate number of users. We also show that there exists an optimal code rate under a bandwidth constraint. The performance and optimal code rates are also demonstrated via density evolution analyses.
Bo Xia, Williarn E. Ryan
ICC1
2001 Near-optimal convolutionally coded synchronous CDMA with iterative multiuser detection/decoding
abstract
We consider in this paper a convolutionally coded synchronous DS-CDMA system with iterative multiuser detection and decoding. Each user's bits are convolutionally encoded, interleaved, and differentially encoded before being transmitted over the multiuser CDMA channel. In contrast to previous work in this area, a differential encoder is inserted to effect an interleaver gain. The receiver consists of a soft-in soft-out multiuser detector followed by a bank of iterative decoders matched to the convolutional code-differential encoder combination. Dramatic gains over systems which contain no differential encoder are demonstrated. We also show that there exists an optimal code rate under a bandwidth constraint. The gains and optimal code rates are demonstrated via bit error rate simulations and density evolution analyses.
Bo Xia, William E. Ryan
GLOBECOM1