VLDB 2026 Research / reviewers in the wild / expert
Bo Xia
dblp:76/6557
· DBLP profile ↗
25ranked-venue papers
10as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 4 · 3 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences Through f-Divergence MinimizationabstractDirect Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable interest within the community. However, we have observed that these approaches rely solely on minimizing the reverse Kullback-Leibler divergence during alignment process between the fine-tuned model and the reference model, neglecting incorporation of other divergence constraints. In this study, we focus on extending reverse Kullback-Leibler divergence in the alignment paradigm of text-to-image models to f-divergence, which aims to garner better alignment performance as well as good generation diversity. We provide the generalized formula of text-to-image alignment paradigm under f-divergence condition and thoroughly analyze the impact of different divergence constraints on alignment process from the perspective of gradient fields. We conduct comprehensive evaluation on text-image alignment performance, human value alignment performance and generation diversity performance under different divergence constraints, and the results indicate that text-to-image alignment based on Jensen-Shannon divergence achieves the best trade-off among them. The option of divergence employed for aligning text-to-image models significantly impacts the trade-off between alignment performance (especially human value alignment) and generation diversity, which highlights the necessity of selecting an appropriate divergence for practical applications. Bo Xia, Yongzhe Chang, Xueqian Wang 0001 |
AAAI | 2 |
| 2025 | Identical Human Preference Alignment Paradigm for Text-to-Image ModelsabstractImplicit reward mechanism of Direct Preference Optimization (DPO) has facilitated its recent applications beyond large language models (LLMs), notably in aligning text-to-image models with human preferences. While promising results have been achieved with algorithms such as Diffusion-DPO, their reliance on the assumptions of the Bradley-Terry model could potentially lead to significant overfitting. In this paper, we propose the Step Identical Preference Alignment (SIPA) method, departing text-to-image alignment from the assumptions of Bradley-Terry preference model. We assess the performance of four models, Diffusion-DPO, SPO, and SIPA, alongside the original model, on the HPS-V2 test set, which focus on three key aspects: text-image alignment, human value alignment, and generation diversity. Experimental results show that SIPA matches or outperforms existing SOTA alignment methods, and even exceeds the original model in terms of generation diversity, which compellingly demonstrates SIPA’s superiority in mitigating alignment overfitting. Bo Xia, Yongzhe Chang, Xueqian Wang 0001 |
ICASSP | 2 |
| 2025 | Positive Enhanced Preference Alignment for Text-to-Image ModelsabstractDirect Preference Optimization (DPO) has recently expanded its successful application beyond aligning large language models (LLMs), further targeting the alignment of text-to-image models with human preferences. However, traditional DPO approach would inadvertently result in a simultaneous reduction of sampling probabilities for preferred and dispreferred items during the alignment process, thereby potentially diminishing model's generative capacity. In this paper, we firstly undertake a revisit of DPO by grounding our analysis in the framework of contrastive loss. It reveals that DPO only emphasizes the part quantifying dissimilarity between items, while overlooking aspects pertinent to positive items. Hence, we propose the Positive Enhanced Preference Alignment (PEPA). Three enhancement strategies are introduced herein, and after comprehensive empirical evaluation, we recommend implementation of enhancing the log probability of preferred ratio in practice applications, which is distinguished by both stability and effectiveness. Experimental assessments are carried out on the HPS-V2 test set, with results demonstrating that PEPA outperforms or matches current state-of-the-art alignment techniques, thus highlighting PEPA's exceptional practical efficacy. Bo Xia, Yongzhe Chang, Xueqian Wang 0001 |
ICASSP | 2 |
| 2025 | Entropy-based Activation Function Optimization: A Method on Searching Better Activation FunctionsabstractThe success of artificial neural networks (ANNs) hinges greatly on the judicious selection of an activation function, introducing non-linearity into network and enabling them to model sophisticated relationships in data. However, the search of activation functions has largely relied on empirical knowledge in the past, lacking theoretical guidance, which has hindered the identification of more effective activation functions. In this work, we offer a proper solution to such issue. Firstly, we theoretically demonstrate the existence of the worst activation function with boundary conditions (WAFBC) from the perspective of information entropy. Furthermore, inspired by the Taylor expansion form of information entropy functional, we propose the Entropy-based Activation Function Optimization (EAFO) methodology. EAFO methodology presents a novel perspective for designing static activation functions in deep neural networks and the potential of dynamically optimizing activation during iterative training. Utilizing EAFO methodology, we derive a novel activation function from ReLU, known as Correction Regularized ReLU (CRReLU). Experiments conducted with vision transformer and its variants on CIFAR-10, CIFAR-100 and ImageNet-1K datasets demonstrate the superiority of CRReLU over existing corrections of ReLU. Extensive empirical studies on task of large language model (LLM) fine-tuning, CRReLU exhibits superior performance compared to GELU, suggesting its broader potential for practical applications. Bo Xia, Pu Chang, Zibin Dong, Yifu Yuan, Yongzhe Chang, Xueqian Wang 0001 |
ICLR | 3 |
| 2025 | DeepMF: Deep Motion Factorization for Closed-Loop Safety-Critical Driving Scenario SimulationabstractSafety-critical traffic scenarios are of great practical relevance to evaluating the robustness of autonomous driving (AD) systems. Given that these long-tail events are extremely rare in real-world traffic data, there is a growing body of work dedicated to the automatic traffic scenario generation. However, nearly all existing algorithms for generating safety-critical scenarios rely on snippets of previously recorded traffic events, transforming normal traffic flow into accident-prone situations directly. In other words, safety-critical traffic scenario generation is hindsight and not applicable to newly encountered and open-ended traffic events. In this paper, we propose the Deep Motion Factorization (DeepMF) framework, which extends static safety-critical driving scenario generation to closed-loop and interactive adversarial traffic simulation. DeepMF casts safety-critical traffic simulation as a Bayesian factorization that includes the assignment of hazardous traffic participants, the motion prediction of selected opponents, the reaction estimation of autonomous vehicle (AV) and the probability estimation of the accident occur. All the aforementioned terms are calculated using decoupled deep neural networks, with inputs limited to the current observation and historical states. Consequently, DeepMF can effectively and efficiently simulate safety-critical traffic scenarios at any triggered time and for any duration by maximizing the compounded posterior probability of traffic risk. Extensive experiments demonstrate that DeepMF excels in terms of risk management, flexibility, and diversity, showcasing outstanding performance in simulating a wide range of realistic, high-risk traffic scenarios. Linrui Zhang, Bo Xia, Xueqian Wang 0001, Houde Liu |
IJCNN | 3 |
| 2025 | Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image GenerationabstractReinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an important alternative: masked generative models. In this work, we propose Mask-GRPO, the first method to incorporate Group Relative Policy Optimization (GRPO)-based RL into this overlooked paradigm. Our core insight is to redefine the transition probability, which is different from current approaches, and formulate the unmasking process as a multi-step decision-making problem. To further enhance our method, we explore several useful strategies, including removing the Kullback–Leibler constraint, applying the reduction strategy, and filtering out low-quality samples. Using Mask-GRPO, we improve a base model, Show-o, with substantial improvements on standard T2I benchmarks and preference alignment, outperforming existing state-of-the-art approaches. Yifu Luo, Xinhao Hu, Keyu Fan, Bo Xia, Tiantian Zhang 0002, Yongzhe Chang, Xueqian Wang 0001 |
NeurIPS | 6 |
| 2025 | Dual-path aggregation transformer network for super-resolution with images occlusions and variability
Qinghui Chen, Lunqian Wang, Xinghua Wang 0008, Bo Xia, Hao Ding 0014, Jinglin Zhang 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | A delay-robust method for enhanced real-time reinforcement learning
Bo Xia, Bo Yuan 0003, Zhiheng Li 0001, Bin Liang 0001, Xueqian Wang 0001 |
Neural Networks | 1 |
| 2025 | Agent-Based Space Teleoperation: Mitigating Time Delays With Deep Reinforcement LearningabstractSpace teleoperation significantly extends human reach in space missions. However, traditional approaches are constrained by factors, such as the reliance on accurate dynamic models and the risk of operator fatigue during prolonged tasks. Additionally, while data-driven intelligent approaches reduce the need for prior knowledge, they have yet to adequately address the time delay issues inherent in these systems. To overcome these challenges, we introduce the belief state actor-critic (BSAC) method, the first deep reinforcement learning approach tailored for space teleoperation capture tasks within a bilateral control framework. We first establish a generalized agent-based architecture for space teleoperation, shifting decision-making from human operators to autonomous agents. Following a comprehensive analysis of the time delay challenges, we propose the BSAC algorithm, which integrates state augmentation and belief state techniques to mitigate the effects of delays in teleoperated Markov decision processes. Extensive experiments are conducted on the MuJoCo simulation platform, modeling a real hardware system across various scenarios. The learned policies are then successfully transferred and validated in a real-world setup, demonstrating the effectiveness and robustness of BSAC. In summary, our results support the feasibility of agent-based frameworks capable of overcoming time delay challenges in space teleoperation. Bo Xia, Xianru Tian, Bo Yuan 0003, Chunju Yang, Zhiheng Li 0001, Bin Liang 0001, Xueqian Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | Dynamic Modeling for Reinforcement Learning with Random Delay
Yalou Yu, Bo Xia, Minzhi Xie, Zhiheng Li 0001, Xuwqian Wang |
ICANN (4) | 2 |
| 2024 | D3D: Conditional Diffusion Model for Decision-Making Under Random Frame DroppingabstractThe occurrence of frame drops due to issues such as corrupted communications or malfunctioning sensors presents a significant challenge to an agent’s decision-making, especially in remote control scenarios. Classical reinforcement learning (RL) usually assumes a continuous data stream without frame drops and relies heavily on online interactions, which is time-consuming, resource-intensive, and often impractical in certain scenarios. Consequently, the performance of RL may deteriorate significantly in face of non-negligible frame drops. To tackle this challenge caused by frame dropping, We propose Conditional Diffusion Model for Decision-Making under Random Frame Dropping (D3D), an offline algorithm that can effectively enhance performance robustness in frame dropping scenarios. D3D addresses this issue through a two-phase approach: 1) During the policy generation phase, D3D adopts a return-conditional diffusion model for decision making rather than the temporal difference learning, whose policy is derived using offline datasets of return-labeled trajectories without information loss. 2) When frame dropping occurs during evaluation, D3D seamlessly substitutes the missing state with its corresponding prediction in the horizon made by the diffusion model. Extensive experiments are conducted on MuJoCo and Adroit tasks to validate D3D’s robustness and efficiency. The results demonstrate that D3D consistently outperforms state-of-the-art RL algorithms, especially excelling on tasks featuring severe drop rates. Bo Xia, Yifu Luo, Yongzhe Chang, Bo Yuan 0003, Zhiheng Li 0001, Xueqian Wang 0001 |
RO-MAN | 1 |
| 2024 | Solving time-delay issues in reinforcement learning via transformers
Bo Xia, Zaihui Yang, Minzhi Xie, Yongzhe Chang, Bo Yuan 0003, Zhiheng Li 0001, Xueqian Wang 0001, Bin Liang 0001 |
Appl. Intell. | 1 |
| 2023 | Addressing Delays in Reinforcement Learning via Delayed Adversarial Imitation Learning
Minzhi Xie, Bo Xia, Yalou Yu, Xueqian Wang 0001, Yongzhe Chang |
ICANN (3) | 2 |
| 2023 | ULKNet:Rethinking Large Kernel CNN with UNet-Attention for Remote Sensing Images Semantic SegmentationabstractSince the advent of the Transformer, the Vision Transformer (ViT) model has become the optimal solution for extracting semantic information in remote sensing image semantic segmentation tasks. However, recent research has discovered that using larger convolutional kernels to extract global information can achieve performance comparable to, or even better than the ViT models. This finding inspires us to redesign the convolutional neural network structure and innovatively construct a novel parallel global-local module based on large kernel convolution, in order to better extract semantic information and local contextual information. Following this design approach, we propose ULKNet, a pure CNN architecture that achieves performance comparable to the ViT model. Through experiments on the ISPRS Vaihingen and Potsdam datasets, We respectively achieve mIoU 82.7% and 86.0% mIoU scores. This finding demonstrates that a well-designed large kernel CNN network structure can provide a larger and more effective receptive field, thereby better extracting semantic information. Lunqian Wang, Bo Xia |
IECON | 5 |
| 2023 | Overcoming Delayed Feedback via Overlook Decision MakingabstractReinforcement learning is one of the most general paradigms to solve sequential decision making issues on the assumption that the action selection and environmental feedback are instantaneous, however, unfortunately this assumption is rarely true with regard to such ubiquitous delays in real-world system which could degrade the performance of reinforcement learning algorithms. The most common solution to solve a fixed delay problem is to design a forward dynamic model which is used to predict the newest state by recursively iterating over long steps so that a predicted state can be got and it would be taken as the agent's observation to make the newest decision. However, there exists cumulative errors during the iterative process which make long-term prediction inaccurate and further affect agent's decision. Motivated by the goal to reduce cumulative errors, we propose a new algorithm named Multi-step Prediction model with Delayed Observation(MPDO), aiming at accurately predicting future state at longer horizons for better decision making. Our approach includes two parts: a multi-step prediction model and a strategy training based on proximal policy optimization algorithms(PPO). Our model only needs a small amount of data to conduct dynamic modeling quickly, and the accuracy of prediction and iteration speed are higher than traditional methods. Experiments on Gym and MuJoCo show that MPDO achieves higher performance in such different tasks with different delays compared with other state-of-the-art methods, which verify our method's effectiveness. Yalou Yu, Bo Xia, Minzhi Xie, Xueqian Wang 0001, Zhiheng Li 0001, Yongzhe Chang |
SMC | 2 |
| 2022 | Don't Touch What Matters: Task-Aware Lipschitz Data Augmentation for Visual Reinforcement LearningabstractOne of the key challenges in visual Reinforcement Learning (RL) is to learn policies that can generalize to unseen environments. Recently, data augmentation techniques aiming at enhancing data diversity have demonstrated proven performance in improving the generalization ability of learned policies. However, due to the sensitivity of RL training, naively applying data augmentation, which transforms each pixel in a task-agnostic manner, may suffer from instability and damage the sample efficiency, thus further exacerbating the generalization performance. At the heart of this phenomenon is the diverged action distribution and high-variance value estimation in the face of augmented images. To alleviate this issue, we propose Task-aware Lipschitz Data Augmentation (TLDA) for visual RL, which explicitly identifies the task-correlated pixels with large Lipschitz constants, and only augments the task-irrelevant pixels for stability. We verify the effectiveness of our approach on DeepMind Control suite, CARLA and DeepMind Manipulation tasks. The extensive empirical results show that TLDA improves both sample efficiency and generalization; it outperforms previous state-of-the-art methods across 3 different visual control benchmarks. Zhecheng Yuan, Guozheng Ma, Yao Mu 0001, Bo Xia, Bo Yuan 0003, Xueqian Wang 0001, Ping Luo 0002, Huazhe Xu |
IJCAI | 4 |
| 2021 | Real-time indirect illumination by virtual planar area lights
Bo Xia, Guanyu Xing, Yanli Liu 0002, Yanci Zhang |
Comput. Graph. | 2 |
| 2018 | Distributed Representation of Chinese Collocation
Bo Xia, Endong Xun |
PACLIC | 1 |
| 2007 | A Waiting-Time Dependent Backoff Algorithm for QoS Enhancement of Voice over WLANabstractAs the integration of prevailing Wireless Local Area Network (WLAN) and Voice over Internet Protocol (VoIP) technology, Voice over WLAN (VoWLAN) is expected to experience a dramatic growth in the near future. However, the widely deployed IEEE 802.11 contention-based Media Access Control (MAC) mechanism is unsuitable for voice transmission with strict Quality of Service (QoS) requirements. In this paper, the delay jitter and collision problems in Binary Exponential Backoff (BEB) algorithm are analyzed first. Then, a novel "Waiting-time Dependent Backoff (WDB)" algorithm is proposed to cope with these problems and improve QoS of voice transmission. Simulation results show that WDB can dramatically reduce delay jitter and collision probability, as well as end-to-end delay, while maintaining acceptable drop rate. Jing Chi, Bo Xia, Ruozhou Lin |
PIMRC | 3 |
| 2005 | Quasi-cyclic codes from extended difference familiesabstractQuasi-cyclic codes are renowned for their structural design and low-complexity shift register encoding. However, existing design techniques based on difference families have limited code size options due to algebraic constraints. We introduce in this paper a notation of extended difference family (EDF) and provide a systematic code design based on EDFs with a high degree of flexibility in code size. Short/medium-length codes of high rate (/spl ges/ 0.9) can be easily constructed. Bo Xia |
WCNC | 2 |
| 2005 | Importance sampling for tanner treesabstractThis correspondence presents an importance sampling (IS) simulation scheme for the soft iterative decoding on loop-free multiple-layer trees. It is shown that this scheme is asymptotically efficient in that, for an arbitrary tree and a given estimation precision, the required number of samples is inversely proportional to the noise standard deviation. This work has its application in the simulation of low-density parity-check (LDPC) codes. Bo Xia, William E. Ryan |
IEEE Trans. Inf. Theory | 1 |
| 2004 | Estimating LDPC codeword error rates via importance samplingabstractThis paper proposes an importance sampling design for the codeword error rate estimation of low-density parity-check codes. An error set partitioning method is introduced to reduce the error boundary complexity and to improve the simulation efficiency. This is a follow-up to the work in (B. Xia et al., 2003). Bo Xia, William E. Ryan |
ISIT | 1 |
| 2003 | On importance sampling for linear block codesabstractWe introduce an importance sampling scheme for linear block codes with message-passing decoding. This novel scheme overcomes an existing difficulty in the IS practice that requires codebook information. Experiments show large IS gains for single parity-check codes and short-length block codes. For medium-length block codes, IS gains in the order of 10/sup 3/ and higher are observed at high signal-to-noise ratio. Bo Xia, William E. Ryan |
ICC | 1 |
| 2002 | Near-optimal convolutionally coded asynchronous CDMA with iterative multiuser detection/decodingabstractWe consider in this paper a convolutionally coded asynchronous CDMA system with iterative multiuser detection and decoding. Each user's bits are convolutionally encoded, interleaved, and differentially encoded before being transmitted over the multiuser CDMA channel. In contrast to previous work in this area, a differential encoder is inserted to effect an interleaver gain. We view the asynchronous CDMA channel as a periodically time-varying ISI channel. The receiver jointly decodes the differential encoders and the CDMA channel with a combined trellis, and shares soft output information with the convolutional decoders in an iterative (turbo) fashion. Dramatic gains over conventional convolutionally coded systems are demonstrated via simulation. Near single-user turbo performance is obtained with a moderate number of users. We also show that there exists an optimal code rate under a bandwidth constraint. The performance and optimal code rates are also demonstrated via density evolution analyses. Bo Xia, Williarn E. Ryan |
ICC | 1 |
| 2001 | Near-optimal convolutionally coded synchronous CDMA with iterative multiuser detection/decodingabstractWe consider in this paper a convolutionally coded synchronous DS-CDMA system with iterative multiuser detection and decoding. Each user's bits are convolutionally encoded, interleaved, and differentially encoded before being transmitted over the multiuser CDMA channel. In contrast to previous work in this area, a differential encoder is inserted to effect an interleaver gain. The receiver consists of a soft-in soft-out multiuser detector followed by a bank of iterative decoders matched to the convolutional code-differential encoder combination. Dramatic gains over systems which contain no differential encoder are demonstrated. We also show that there exists an optimal code rate under a bandwidth constraint. The gains and optimal code rates are demonstrated via bit error rate simulations and density evolution analyses. Bo Xia, William E. Ryan |
GLOBECOM | 1 |