Boqiang Xu

dblp:23/7627 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Procedure-Aware Hierarchical Alignment for Open Surgery Video-Language Pretraining
abstract
Recent advances in surgical robotics and computer vision have greatly improved intelligent systems' autonomy and perception in the operating room (OR), especially in endoscopic and minimally invasive surgeries. However, for open surgery, which is still the predominant form of surgical intervention worldwide, there has been relatively limited exploration due to its inherent complexity and the lack of large-scale, diverse datasets. To close this gap, we present OpenSurgery, by far the largest video-text pretraining and evaluation dataset for open surgery understanding. OpenSurgery consists of two subsets: OpenSurgery-Pretrain and OpenSurgery-EVAL. OpenSurgery-Pretrain consists of 843 publicly available open surgery videos for pretraining, spanning 102 hours and encompassing over 20 distinct surgical types. OpenSurgery-EVAL is a benchmark dataset for evaluating model performance in open surgery understanding, comprising 280 training and 120 test videos, totaling 49 hours. Each video in OpenSurgery is meticulously annotated by expert surgeons at three hierarchical levels of video, operation, and frame to ensure both high quality and strong clinical applicability. Next, we propose the Hierarchical Surgical Knowledge Pretraining (HierSKP) framework to facilitate large-scale multimodal representation learning for open surgery understanding. HierSKP leverages a granularity-aware contrastive learning strategy and enhances procedural comprehension by constructing hard negative samples and incorporating a Dynamic Time Warping (DTW)-based loss to capture fine-grained temporal alignment of visual semantics. Extensive experiments show that HierSKP achieves state-of-the-art performance on OpenSurgegy-EVAL across multiple tasks, including operation recognition, temporal action localization, and zero-shot cross-modal retrieval. This demonstrates its strong generalizability for further advances in open surgery understanding.
Boqiang Xu, Jinlin Wu, Jian Liang 0001, Zhenan Sun, Hongbin Liu 0001, Jiebo Luo 0001, Zhen Lei 0001
IEEE Trans. Image Process.1
2026 Multi-View Images Suffice 3D Reasoning Through Chain-of-Thought Selection and Question-Guided Fusion
abstract
3D reasoning is crucial in areas like robotics and autonomous driving. Due to the high cost of 3D data acquisition, some recent methods attempt to enable LLMs to perform 3D reasoning through multi-view images, thereby transferring the powerful 2D reasoning capabilities of LLMs to 3D environments. However, these methods face challenges: either they use redundant views that contain many perspectives irrelevant to the question, or they rely on globally aggregated multi-view representations, losing the fine-grained vision-language correlations. To tackle these challenges, we propose 3DMulti-LLM, which mainly consists of three components: a COT selector, a question-guided fusion block, and pre-trained LLMs. Specifically, first, the COT selector leverages the powerful chain-of-thought reasoning capabilities of LLMs to identify question-related multi-view images. In this way, 3DMulti-LLM can eliminate a substantial amount of interference from unnecessary viewpoints. Then, we propose a question-guided fusion block for integrating multi-view features via question-guided interaction among various viewpoints. Finally, the pre-trained LLMs are utilized to reason in 3D scenes directly through multi-view features. Notably, our approach understands the 3D scene solely through multi-view images, without requiring the input of point cloud information or additional 3D feature extraction. Through our experiments, 3DMulti-LLM achieves impressive performance and surpasses existing 3D-input-free methods by + 12.2% and + 7.1% on ScanQA and 3DMV-VQA datasets, respectively.
Boqiang Xu, Jinlin Wu, Wei Zhang 0255, Chenyang Su, Jian Liang 0001, Zhenan Sun, Zhen Lei 0001
IEEE Trans. Image Process.1
2025 Filtering Before Detection: Objects of Interest Enhanced Anomaly Detection for Surveillance
abstract
Recent advances in video anomaly detection (VAD) have primarily focused on modeling the spatio-temporal characteristics of normal patterns, with an emphasis on understanding background semantics. However, in real-world surveillance scenarios, we observe that anomalies are predominantly triggered by foreground objects, rather than background scenes. Given that surveillance cameras are typically fixed in position, the background remains largely static, making foreground objects the primary source of irregular events. Motivated by this insight, we propose a novel Filtering Before Detection strategy that prioritizes potentially anomalous foreground objects prior to performing anomaly detection. Specifically, we introduce an object-centric VAD framework that identifies and filters objects of interest based on visual appearance, motion dynamics, and semantic cues. These selected object features are then integrated into a general VAD model to enhance detection accuracy. Extensive experiments on five real-world surveillance datasets demonstrate the effectiveness of our approach. Our method achieves consistent and significant performance gains when plugged into state-of-the-art VAD models, offering a more efficient and robust solution for real-world anomaly detection in surveillance videos.
Boqiang Xu
IJCNN2
2025 MedICL: In-Context Learning for Semantically Enhanced AKI Prediction in Cardiac Surgery
Chenyang Su, Yishun Wang, Boqiang Xu, Rong Feng, Hongbin Liu 0001, Gaofeng Meng
MICCAI (11)3
2025 Bayesian-optimized deep learning model for real-time spatial distribution identification of vehicle axle load
Boqiang Xu, Genyu Feng, Xingbao Liu
Expert Syst. Appl.1
2024 Keypoint detection-based and multi-deep learning model integrated method for identifying vehicle axle load spatial-temporal distribution
Boqiang Xu
Adv. Eng. Informatics1
2024 A monocular-based framework for accurate identification of spatial-temporal distribution of vehicle wheel loads under occlusion scenarios
Boqiang Xu, Xingbao Liu, Genyu Feng
Eng. Appl. Artif. Intell.1
2024 Deep Learning Based Occluded Person Re-Identification: A Survey
abstract
Occluded person re-identification (Re-ID) focuses on addressing the occlusion problem when retrieving the person of interest across non-overlapping cameras. With the increasing demand for intelligent video surveillance and the application of person Re-ID technology, the real-world occlusion problem draws considerable interest from researchers. Although a large number of occluded person Re-ID methods have been proposed, there are few surveys that focus on occlusion. To fill this gap and help boost future research, this article provides a systematic survey of occluded person Re-ID. In this work, we review recent deep learning based occluded person Re-ID research. First, we summarize the main issues caused by occlusion as four groups: position misalignment, scale misalignment, noisy information, and missing information. Second, we categorize existing methods into six solution groups: matching, image transformation, multi-scale features, attention mechanism, auxiliary information, and contextual recovery. We also discuss the characteristics of each approach, as well as the issues they address. Furthermore, we present the performance comparison of recent occluded person Re-ID methods on four public datasets: Partial-ReID, Partial-iLIDS, Occluded-ReID, and Occluded-DukeMTMC. We conclude the study with thoughts on promising future research directions.
Yunjie Peng, Jinlin Wu, Boqiang Xu, Chunshui Cao, Xu Liu 0008, Zhenan Sun, Zhiqiang He 0002
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Color-Unrelated Head-Shoulder Networks for Fine-Grained Person Re-identification
abstract
Person re-identification (re-id) attempts to match pedestrian images with the same identity across non-overlapping cameras. Existing methods usually study person re-id by learning discriminative features based on the clothing attributes (e.g., color, texture). However, the clothing appearance is not sufficient to distinguish different persons especially when they are in similar clothes, which is known as the fine-grained (FG) person re-id problem. By contrast, this paper proposes to exploit the color-unrelated feature along with the head-shoulder feature for FG person re-id. Specifically, a color-unrelated head-shoulder network (CUHS) is developed, which is featured in three aspects: (1) It consists of a lightweight head-shoulder segmentation layer for localizing the head-shoulder region and learning the corresponding feature. (2) It exploits instance normalization (IN) for learning color-unrelated features. (3) As IN inevitably reduces inter-class differences, we propose to explore richer visual cues for IN by an attention exploration mechanism to ensure high discrimination. We evaluate our model on the FG-reID, Market1501, and DukeMTMC-reID datasets, and the results show that CUHS surpasses previous methods on both the FG and conventional person re-id problems.
Boqiang Xu, Jian Liang 0001, Lingxiao He, Jinlin Wu, Zhenan Sun
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Mimic Embedding via Adaptive Aggregation: Learning Generalizable Person Re-identification
Boqiang Xu, Jian Liang 0001, Lingxiao He, Zhenan Sun
ECCV (14)1
2022 Learning Feature Recovery Transformer for Occluded Person Re-Identification
abstract
One major issue that challenges person re-identification (Re-ID) is the ubiquitous occlusion over the captured persons. There are two main challenges for the occluded person Re-ID problem, i.e. , the interference of noise during feature matching and the loss of pedestrian information brought by the occlusions. In this paper, we propose a new approach called Feature Recovery Transformer (FRT) to address the two challenges simultaneously, which mainly consists of visibility graph matching and feature recovery transformer. To reduce the interference of the noise during feature matching, we mainly focus on visible regions that appear in both images and develop a visibility graph to calculate the similarity. In terms of the second challenge, based on the developed graph similarity, for each query image, we propose a recovery transformer that exploits the feature sets of its k -nearest neighbors in the gallery to recover the complete features. Extensive experiments across different person Re-ID datasets, including occluded, partial and holistic datasets, demonstrate the effectiveness of FRT. Specifically, FRT significantly outperforms state-of-the-art results by at least 6.2% Rank- 1 accuracy and 7.2% mAP scores on the challenging Occluded-Duke dataset.
Boqiang Xu, Lingxiao He, Jian Liang 0001, Zhenan Sun
IEEE Trans. Image Process.1
2020 A Lightweight Multi-Label Segmentation Network for Mobile Iris Biometrics
abstract
This paper proposes a novel, lightweight deep convolutional neural network specifically designed for iris segmentation of noisy images acquired by mobile devices. Unlike previous studies, which only focused on improving the accuracy of segmentation mask using the popular CNN technology, our method is a complete end-to-end iris segmentation solution, i.e., segmentation mask and parameterized pupillary and limbic boundaries of the iris are obtained simultaneously, which further enables CNN-based iris segmentation to be applied in any regular iris recognition systems. By introducing an intermediate pictorial boundary representation, predictions of iris boundaries and segmentation mask have collectively formed a multi-label semantic segmentation problem, which could be well solved by a carefully adapted stacked hourglass network. Experimental results show that our method achieves competitive or state-of-the-art performance in both iris segmentation and localization on two challenging mobile iris databases.
Caiyong Wang, Yunlong Wang 0003, Boqiang Xu, Yong He 0009, Zhiwei Dong, Zhenan Sun
ICASSP3
2020 Black Re-ID: A Head-shoulder Descriptor for the Challenging Problem of Person Re-Identification
abstract
Person re-identification (Re-ID) aims at retrieving an input person image from a set of images captured by multiple cameras. Although recent Re-ID methods have made great success, most of them extract features in terms of the attributes of clothing (e.g., color, texture). However, it is common for people to wear black clothes or be captured by surveillance systems in low light illumination, in which cases the attributes of the clothing are severely missing. We call this problem the Black Re-ID problem. To solve this problem, rather than relying on the clothing information, we propose to exploit head-shoulder features to assist person Re-ID. The head-shoulder adaptive attention network (HAA) is proposed to learn the head-shoulder feature and an innovative ensemble method is designed to enhance the generalization of our model. Given the input person image, the ensemble method would focus on the head-shoulder feature by assigning a larger weight if the individual insides the image is in black clothing. Due to the lack of a suitable benchmark dataset for studying the Black Re-ID problem, we also contribute the first Black-reID dataset, which contains 1274 identities in training set. Extensive evaluations on the Black-reID, Market1501 and DukeMTMC-reID datasets show that our model achieves the best result compared with the state-of-the-art Re-ID methods on both Black and conventional Re-ID problems. Furthermore, our method is also proved to be effective in dealing with person Re-ID in similar clothing. Our code and dataset are avaliable on https://github.com/xbq1994/.
Boqiang Xu, Lingxiao He, Xingyu Liao, Wu Liu 0005, Zhenan Sun, Tao Mei 0001
ACM Multimedia1
2016 Analysis of fault ride-through of doubly-fed wind power generator based on rotor series resistor
abstract
In the case of power grid failure, the rotor side converter (RSC) of double fed wind generator (DFIG) will be shorted when Crowbar protection is activated, which results in the out of control of DFIG. And the traditional control method of converter can lead to the deterioration of control performance of DFIG. In view of the above problems, this paper improves converter control strategy on the basis of considering the dynamic changes of the stator excitation current and proposes fault ride-through scheme based on rotor series resistor. The proposed scheme is analyzed and its superiority is verified through Matlab/Simulink simulation platform. Simulation results show that the proposed method involving the cooperation between rotor series resistor and improved control strategy can help the DFIG achieve fault ride-through under grid faults condition, which is conducive to the stable operation of the wind power system.
Liling Sun, Nana Meng, Boqiang Xu
IECON3