Bo Xiao 0006

dblp:92/6046-6 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
15since 2021 · last 2027
0000-0003-3392-3293ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2027 RADiff: Retrieval-augmented diffusion with route-destination priors for multi-aircraft trajectory prediction
Benkui Zhang, Jinxiao Wang, Bo Xiao 0006
Expert Syst. Appl.4
2026 HCPS: Heuristic Candidate-Space Policy Search for Agile Earth Observation Satellite Scheduling
Benkui Zhang, Jinxiao Wang, Kangning Du, Bo Xiao 0006
PPSN (1)7
2026 CoDRE-Net: Collaborative dual-domain representation enhancement network for flight trajectory prediction
Benkui Zhang, Ying Chang, Bo Xiao 0006
Neurocomputing4
2026 BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual Grounding
abstract
Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent grounding transfer works have advanced the field through one-tower architectures, they still suffer from two primary limitations: (1) over-entangled multimodal representations that exacerbate deceptive modality biases, and (2) insufficient semantic reasoning that hinders the comprehension of referential cues. In this paper, we propose BARE, a bias-aware and reasoning-enhanced framework for one-tower visual grounding. BARE introduces a mechanism that preserves modality-specific features and constructs referential semantics through three novel modules: (i) language salience modulator, (ii) visual bias correction and (iii) referential relationship enhancement, which jointly mitigate multimodal distractions and enhance referential comprehension. Extensive experimental results on five benchmarks demonstrate that BARE not only achieves state-of-the-art performance but also delivers superior computational efficiency compared to existing approaches. The code is publicly accessible at https://github.com/Marloweeee/BARE.
Hongbing Li, Linhui Xiao, Bo Xiao 0006, Zhanyu Ma
IEEE Trans. Circuits Syst. Video Technol.6
2025 T-Stars-Poster: A Framework for Product-Centric Advertising Image Design
Hongyu Chen 0005, Zihang Lin, Bo Xiao 0006, Tiezheng Ge, Bo Zheng 0007
CIKM7
2025 SALA: Semantic alignment and localization alignment for visual grounding
abstract
Visual grounding focuses on establishing fine-grained alignment between specific regions and query expressions, which is increasingly essential as a cornerstone of visual intelligence. Despite recent success, existing methods often struggle with two main issues. Firstly, using independently pre-trained uni-modal encoders to extract expressive feature embeddings leads to a significant semantic gap between uni-modal features, hindering the effective interaction of visual-linguistic contexts. Secondly, the supervision provided by box annotations is inherently sparse and often underexploited, which limits the model’s ability to capture the fine-grained visual cues necessary to distinguish referent objects from the background, thereby leading to localization ambiguity. In this paper, we propose a Semantic Alignment and Localization Alignment (SALA) framework for visual grounding, which effectively bridges the cross- and uni-modal semantic gap and improves localization performance through patch-level and pixel-level alignment. This contributes to enhancing the consistency of representation before multimodal fusion, thereby improving the localization performance. Extensive experiments show that the proposed method outperforms state-of-the-art methods on five widely used datasets. Codes will be made publicly available after acceptance.
Hongbing Li, Linyi Yang, Bo Xiao 0006
Neurocomputing5
2024 Joint recognition of basic and compound facial expressions by mining latent soft labels
Mei Wang 0001, Bo Xiao 0006, Jiani Hu, Weihong Deng
Pattern Recognit.3
2023 Visual tracking using transformer with a combination of convolution and attention
Zihang Feng, Yuanqing Xia, Bo Xiao 0006
Image Vis. Comput.5
2023 Multi-Task Probabilistic Regression With Overlap Maximization for Visual Tracking
abstract
Recent researches made a breakthrough in visual tracking accuracy. Many trackers benefit from the object state representations and network loss functions, which mine the output space and improve the power of supervision, respectively. Probabilistic regression method models the noises and uncertainties in the annotations. However, advanced trackers with probabilistic regression are not studied sufficiently in the aspect of supervision and the aspect of robustness of evaluation maximization. In this paper, an overlap maximization network in the manner of probabilistic regression is proposed to improve the learning ability of the network and the discriminative ability in the evaluation maximization. Firstly, the probabilistic regression is extended with the intersection over union (IoU) evaluation, which is normalized as a probability density in the regression space. Secondly, the classification probability is added as a branch of the iterative evaluation module to improve the ability of distinguishing objects in the evaluation maximization. Moreover, the two branches are constructed into a joint probabilistic regression task of IoU evaluation, which makes the network learn from two types of ground truth and provide a consistent result with multi-branch outputs. For feature interpretation, the strip pooling network and the space-time memory network are introduced to encode long-range context and provide dynamic features, respectively. Compared to the state-of-the-art probabilistic regression trackers and other advanced trackers, the experiments show that the proposed tracker achieves outstanding performance across the six datasets, including GOT-10k, LaSOT, TrackingNet, UAV123, OTB-100 and VOT2018.
Zihang Feng, Yuanqing Xia, Bo Xiao 0006
IEEE Trans. Circuits Syst. Video Technol.4
2023 On the Comparisons of Decorrelation Approaches for Non-Gaussian Neutral Vector Variables
abstract
-norm equals one. In addition, its neutral properties make it significantly different from the commonly studied vector variables (e.g., the Gaussian vector variables). Due to the aforementioned properties, the conventionally applied linear transformation approaches [e.g., principal component analysis (PCA) and independent component analysis (ICA)] are not suitable for neutral vector variables, as PCA cannot transform a neutral vector variable, which is highly negatively correlated, into a set of mutually independent scalar variables and ICA cannot preserve the bounded property after transformation. In recent work, we proposed an efficient nonlinear transformation approach, i.e., the parallel nonlinear transformation (PNT), for decorrelating neutral vector variables. In this article, we extensively compare PNT with PCA and ICA through both theoretical analysis and experimental evaluations. The results of our investigations demonstrate the superiority of PNT for decorrelating the neutral vector variables.
Zhanyu Ma, Xiaoou Lu, Jiyang Xie 0001, Zhen Yang 0004, Jing-Hao Xue, Zheng-Hua Tan, Bo Xiao 0006, Jun Guo 0002
IEEE Trans. Neural Networks Learn. Syst.7
2022 They Like Comedy, Don't You? A Cluster-Based Meta-Learning for Cold-Start Recommendation
abstract
Cold start has always been a challenging problem due to the sparse user-item interaction. Recently meta-learning models have performed outstandingly in solving this problem, which train an optimal initialization parameter by sharing the knowledge of all users. However, knowledge sharing between users with different preferences is negatively affected. In this paper, we try to share knowledge only among users with similar interests. Specifically, we first propose a user cluster method in the cold start scenario. Then, with the user's cluster information we design a conversion network to transform the global initialization parameters learned by meta learning into the cluster optimal initialization parameters. In this way, we can reduce the negative impact among users with large differences in preferences, and improve the performance of the meta-learning model. Extensive experiments on MovieLens and Yelp demonstrate that our method significantly outperforms the state of the arts in both warm up and cold-start scenarios.
Qianfang Xu, Wenliang Li 0001, Bo Xiao 0006
ICME4
2022 Multisensor fusion estimation of nonlinear systems with intermittent observations and heavy-tailed noises
Bo Xiao 0006, Q. M. Jonathan Wu
Sci. China Inf. Sci.1
2022 Multi-feature fusion tracking algorithm based on peak-context learning
Tayssir Bouraffa, Zihang Feng, Yuanqing Xia, Bo Xiao 0006
Image Vis. Comput.5
2021 A multi-step decision prediction model based on LightGBM
abstract
This paper is based on the 3rd place solution to IEEE BigData Cup 2021: RL based RecSys. Standing on the competition task of game-like multi-step recommendation, we have analyzed user data and task goals, and treated recommendation task as multi-classification. After comparing different models, the LightGBM multi-classification model is finally selected. In addition, we evaluate those features that potentially cause the greatest impact on the prediction of our model, making it possible to predict the users’ purchase sequence accurately.
Qianfang Xu, Wenliang Li 0001, Bo Xiao 0006
IEEE BigData5
2021 Context-Aware Correlation Filter Learning Toward Peak Strength for Visual Tracking
abstract
Recently, the correlation filter (CF) has been catching significant attention in visual tracking for its high efficiency in most state-of-the-art algorithms. However, the tracker easily fails when facing the distractions caused by background clutter, occlusion, and other challenging situations. These distractions commonly exist in the visual object tracking of real applications. Keep tracking under these circumstances is the bottleneck in the field. To improve tracking performance under complex interference, a combination of least absolute shrinkage and selection operator (LASSO) regression and contextual information is introduced to the CF framework through the learning stage in this article to ignore these distractions. Moreover, an elastic net regression is proposed to regroup the features, and an adaptive scale method is implemented to deal with the scale changes during tracking. Theoretical analysis and exhaustive experimental analysis show that the proposed peak strength context-aware (PSCA) CF significantly improves the kernelized CF (KCF) and achieves better performance than other state-of-the-art trackers.
Tayssir Bouraffa, Zihang Feng, Bo Xiao 0006, Q. M. Jonathan Wu, Yuanqing Xia
IEEE Trans. Cybern.4
2020 PA-GGAN: Session-Based Recommendation with Position-Aware Gated Graph Attention Network
abstract
Session-based recommendation aims to predict user behaviors based on anonymous sessions. Recently, session sequences are modeled as graph-structured data. Based on the session graphs, Graph Neural Networks (GNNs) can capture complex transitions of items, compared with previous conventional sequential methods. However, the existing graph-construction approaches have limited power in capturing the position information of items in the session sequences. In addition, GNNs employed in the existing session-based recommendation are not capable to attend over their neighborhoods' features in feature aggregation phase. In this paper, we propose a Position-Aware Gated Graph Attention Network (PAGGAN). Specifically, a reverse-position mechanism is proposed to assign position embeddings to nodes in the session graphs based on the order of items in each session sequence. And we enhance Gated Graph Neural Network (GGNN) by introducing self-attention mechanism when aggregating features from nodes. Experimental results on two real-world datasets show that the PA-GGAN outperforms state-of-the-art methods.
Jinshan Wang, Qianfang Xu, Jiahuan Lei, Chaoqun Lin, Bo Xiao 0006
ICME5
2020 DFH-GAN: A Deep Face Hashing with Generative Adversarial Network
abstract
Face Image retrieval is one of the key research directions in computer vision field. Thanks to the rapid development of deep neural network in recent years, deep hashing has achieved good performance in the field of image retrieval. But for large-scale face image retrieval, the performance needs to be further improved. In this paper, we propose Deep Face Hashing with GAN (DFH-GAN), a novel deep hashing method for face image retrieval, which mainly consists of three components: a generator network for generating synthesized images, a discriminator network with a shared CNN to learn multi-domain face feature, and a hash encoding network to generate compact binary hash codes. The generator network is used to perform data augmentation so that the model could learn from both real images and diverse synthesized images. We adopt a two-stage training strategy. In the first stage, the GAN is trained to generate fake images, while in the second stage, to make the network convergence faster. The model inherits the trained shared CNN of discriminator to train the DFH model by using many different supervised loss functions not only in the last layer but also in the middle layer of the network. Extensive experiments on two widely used datasets demonstrate that DFH-GAN can generate high-quality binary hash codes and exceed the performance of the state-of-the-art model greatly.
Lanxiang Zhou, Bo Xiao 0006, Qianfang Xu
ICPR3
2020 A novel Pooling Block for improving lightweight deep neural networks
Bo Xiao 0006, Chun-Guang Li, Qianfang Xu
Pattern Recognit. Lett.1
2017 Event-triggered multisensor data fusion with correlated noise
abstract
As communication bandwidth and resources are limited in network-based control systems, in order to reduce superfluous waste, it is necessary to design an event-triggered communication mechanism. In this paper, the problem of event-triggered state estimation is studied for fusion of multiple sensors with correlated noise. The noise of different sensors are cross-correlated and coupled with the system noise of the previous step and the same time step. An optimal state estimation algorithm based on iterative estimation of white noise estimator is presented, which makes full use of the observation information effectively. A numerical example is used to illustrate the effectiveness of the presented algorithm.
Lu Jiang 0005, Yuanqing Xia, Qiao Guo, Mengyin Fu, Bo Xiao 0006
FUSION6
2017 Robust scene matching method based on sparse representation and iterative correction
Sai Yang, Bo Xiao 0006, Yuanqing Xia, Mengyin Fu, Yang Liu 0038
Image Vis. Comput.2