VLDB 2026 Research / reviewers in the wild / expert
Baojie Fan
dblp:08/7627
· DBLP profile ↗
49ranked-venue papers
26as first author
24since 2021 · last 2026
0000-0002-1627-9726ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 16 first-author · 13 since 2021Artificial intelligence and machine learning · 26 · 13 first-author · 15 since 2021Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MGAF: LiDAR-Camera 3D Object Detection With Multiple Guidance and Adaptive FusionabstractRecent years have witnessed the remarkable progress of 3D multi-modality object detection methods based on the Bird's-Eye-View (BEV) perspective. However, most of them overlook the complementary interaction and guidance between LiDAR and camera. In this work, we propose a novel multi-modality 3D objection detection method, with multi-guided global interaction and LiDAR-guided adaptive fusion, named MGAF. Specifically, we introduce sparse depth guidance (SDG) and LiDAR occupancy guidance (LOG) to generate 3D features with sufficient depth and spatial information. The designed semantic segmentation network captures category and orientation prior information for raw point clouds. In the following, an Adaptive Fusion Dual Transformer (AFDT) is developed to adaptively enhance the interaction of different modal BEV features from both global and bidirectional perspectives. Meanwhile, additional downsampling with sparse height compression and multi-scale dual-path transformer (MSDPT) are designed in order to enlarge the receptive fields of different modal features. Finally, a temporal fusion module is introduced to aggregate features from previous frames. Notably, the proposed AFDT is general, which also shows superior performance on other models. Our framework has undergone extensive experimentation on the large-scale nuScenes dataset, Waymo Open Dataset, and long-range Argoverse2 dataset, consistently demonstrating state-of-the-art performance. Baojie Fan, Caixia Xia, Huijie Fan, Fengyu Xu 0001, Jiandong Tian |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | All-Day Multi-Camera Multi-Target TrackingabstractThe capability of tracking objects in low-light environments like nighttime is crucial for numerous real-world applications. However, previous Multi-Camera Multi-Target(MCMT) tracking methods are primarily focused on tracking during daytime with favorable lighting, overlooking the challenge posed by low-light conditions. The main difficulty of tracking under low-light condition is the lack of detailed visible appearance features. To address this issue, we incorporate the infrared modality into MCMT tracking framework to provide more useful information. We constructed the first Multi-modality (RGBT) Multi-camera Multi-target tracking dataset named M3Track, which contains sequences captured in low-light environments, laying a solid foundation for all-day multi-camera tracking. Based on the proposed dataset, we propose All-Day Multi-Camera Multi-Target tracking network, termed as ADM-CMT. Specifically, we propose an All-Day Mamba Fusion(ADMF) module to fuse information from different modalities adaptively. Within ADMF, the Lighting Guidance Model(LGM) extracts lighting relevant information to guide the fusion process. Furthermore, the Nearby Target Collection(NTC) strategy is designed to enhance tracking accuracy by leveraging information derived from surrounding objects of targets. Experiments conducted on M3Track demonstrate that ADMCMT exhibits strong generalization across different lighting conditions. The code will be released at https://github.com/QTRACKY/ADMCMT. Huijie Fan, Yihao Zhen, Tinghui Zhao, Baojie Fan, Qiang Wang 0015 |
CVPR | 5 |
| 2025 | RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy Prediction
Baojie Fan, Yuyu Jiang, Jiandong Tian, Huijie Fan |
ICCV | 1 |
| 2025 | Progressive Knowledge Learning for Source-Free Domain Adaptive Object Detection
Caiyu Zhang, Baojie Fan, Wenzhang Zhou |
ICIG (3) | 2 |
| 2025 | AdaptiveFusion: LiDAR-Camera Adaptive Fusion for 3D Object DetectionabstractLidar-Camera fusion modules based on bird’s-eye view (BEV) have significantly pushed the state of the art in visual perception for accurate and reliable autonomous driving systems. However, most of them only focus on how to unify the multimodal data into the BEV perspective, and few methods consider the flexible fusion strategy based on the characteristics of BEV features under different modalities. We propose a novel LiDAR-Camera adaptive fusion framework for 3D object detection, named AdaptiveFusion, which efficiently predicts high-quality 3D detection results through more robust BEV fusion features with the help of transformer operations. AdaptiveFusion consists of two novel designs. Firstly, an Adaptive Fusion Encoder is developed to enhance the spatial geometric information in the channel of LiDAR-BEV features through simple global pooling, and enriches the semantic information of sparse Camera-Bev features by mining semantic entities distributed in the feature vector in the form of groups. Secondly, the proposed Dynamic Decoder utilizes two linear projections to interact with the information of each head representing a subspace in standard multi-head attention, and improves the fused BEV representation ability. Extensive experiments show that AdaptiveFusion achieves competitive performance on nuScenes dataset, with 71.4% mAP and 73.8% NDS on 3D object detection, and 63.2% mIoU on BEV map segmentation. Baojie Fan |
ICME | 3 |
| 2025 | Unidirectional Point-Voxel Fusion for Enhanced 3D Single Object TrackingabstractSparse point-based trackers struggle with texture-less and incomplete point clouds. Conversely, dense voxel-based trackers have richer spatial and semantic information, but filtering out interference from complex backgrounds remains a challenge. Additionally, there is still a gap between point and voxel-based trackers in exploiting their complementary strengths. To address these issues, we propose UTracker, which uses unidirectional point-voxel fusion to construct a bridge between point and voxel tracking features, enabling them to complement and enhance each other. Specifically, we design template-enhanced unidirectional attention (TEUA) and historical template fusion (HTF), which enable unidirectional interaction from historical templates to the search area in the point branch, retaining the pure template features. Then, a point-guided adaptive feature transformer (PGAFT) is developed to unidirectionally enhance the interaction between point and voxel features. Extensive experiments demonstrate that UTracker achieves superior performance, reaching an average accuracy of 89.5%, 72.58%, and 63.4% on the KITTI, NuScenes, and Waymo Open Dataset, respectively. Yuyu Jiang, Baojie Fan, Jinrong Du |
IROS | 2 |
| 2024 | GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object DetectionabstractRecent years have witnessed the remarkable progress of 3D multi-modality object detection methods based on the Bird's-Eye-View (BEV) perspective. However, most of them overlook the complementary interaction and guidance be-tween LiDAR and camera. In this work, we propose a novel multi-modality 3D objection detection method, named GA-Fusion, with LiDAR-guided global interaction and adaptive fusion. Specifically, we introduce sparse depth guidance (SDG) and LiDAR occupancy guidance (LOG) to generate 3D features with sufficient depth information. In the following, LiDAR-guided adaptive fusion transformer (LGAFT) is developed to adaptively enhance the interaction of different modal BEV features from a global perspective. Meanwhile, additional downsampling with sparse height compression and multi-scale dual-path transformer (MSDPT) are de-signed to enlarge the receptive fields of different modal features. Finally, a temporal fusion module is introduced to ag-gregate features from previous frames. GAFusion achieves state-of-the-art 3D object detection results with 73.6% mAP and 74.9% NDS on the nuScenes test set. Baojie Fan, Jiandong Tian, Huijie Fan |
CVPR | 2 |
| 2024 | Unbiased Faster R-CNN for Single-source Domain Generalized Object DetectionabstractSingle-source domain generalization (SDG) for object detection is a challenging yet essential task as the distribution bias of the unseen domain degrades the algorithm per-formance significantly. However, existing methods attempt to extract domain-invariant features, neglecting that the bi-ased data leads the network to learn biased features that are non-causal and poorly generalizable. To this end, we pro-pose an Unbiased Faster R-CNN (UFR) for generalizable feature learning. Specifically, we formulate SDG in object detection from a causal perspective and construct a Struc-tural Causal Model (SCM) to analyze the data bias andfeature bias in the task, which are caused by scene confounders and object attribute confounders. Based on the SCM, we de-sign a Global-Local Transformation module for data aug-mentation, which effectively simulates domain diversity and mitigates the data bias. Additionally, we introduce a Causal Attention Learning module that incorporates a designed at-tention invariance loss to learn image-level features that are robust to scene confounders. Moreover, we develop a Causal Prototype Learning module with an explicit instance constraint and an implicit prototype constraint, which fur-ther alleviates the negative impact of object attribute con-founders. Experimental results on five scenes demonstrate the prominent generalization ability of our method, with an improvement of 3.9% mAP on the Night-Clear scene. Shijun Zhou, Xiyao Liu 0002, Chunhui Hao, Baojie Fan, Jiandong Tian |
CVPR | 5 |
| 2024 | Masked Mutual Guidance Transformer TrackingabstractVisual mask learning has received increasing attention in the field of visual object tracking. However, most existing studies merely utilize visual mask learning works as pre-training models without fully exploiting their potential for visual representation. In this paper, we present a novel approach for learning tracking target features, leveraging an encoder-decoder architecture with a masked mutual guidance tracking(MMG). Initially, we perform joint visual feature extraction on both the template and search areas. Subsequently, these features undergo separate self-decoding processes, followed by mutual guidance decoding to reconstruct the original search and template images. This process fosters mutual understanding between the images, facilitating improved learning of object states and shapes across different frames. During the inference phase, we offload the decoder and implement a simple and effective tracker. Experimental results indicate that our proposed method is effective that the mutual guidance strategy can achieve state-of-the-art performance on five tracking datasets. Baojie Fan, Jiajun Ai, Caiyu Zhang |
IROS | 1 |
| 2024 | QO-Net: Query Optimization Underwater Object Detection NetworkabstractUnderwater object detection has attracted increasing interest for its wide application in various underwater tasks. However, due to underwater image quality degradation and the lack of large-scale underwater object datasets, many underwater detectors suffer from low detection performance. To address the issues, we not only propose a novel underwater transformer detector with multi-scale feature enhancement and query optimization, named QO-Net, but also construct a new underwater object detection dataset, called UODD. Specifically, a Conv-Trans Layer is developed as the unit of QO-Net, which effectively learns multi-scale image feature representation through CNN and simultaneously captures the dependencies among different positions in the sequence data through Transformer, enabling QO-Net to process underwater image sequence information over longer distances. An effective combination can enhance the representation of multi-scale features. Then, QO-Net develops a positional query enhancement strategy to optimize the spatial prior of positional queries, thereby speeding up the convergence of the network training. In addition, UODD also contains more than 20,000 underwater images for training and validation, with a variety of rich underwater categories. Extensive experiments on UODD, Brackish, and TrashCan datasets demonstrate that QO-Net presents favorable detection performance against state-of-the-art methods in terms of robustness and accuracy. Jiandong Tian, Hongyang Sun 0003, Baojie Fan, Hongxin Xu |
IROS | 3 |
| 2024 | SDTrack: Spatially decoupled tracker for visual trackingabstractRecent models based on encoder-decoder architecture have shown excellent performance in visual object tracking. The encoder models the global spatiotemporal feature correlation between the template and the search regions, while the decoder learns query embeddings to predict the spatial location of the target. However, in previous methods, decoders are query-shared, which may lead to suboptimal results. We observe that different regions in the visual feature map are suitable for performing different tasks. Salient regions in object provide important information for classification task, while the boundaries around it are more beneficial for box localization task. We therefore propose a spatially decoupled tracker called SDTrack. The tracker contains a query selection module that we carefully design to select appropriate queries for both classification and regression tasks. We divide the cross-attention module in the decoder and add the box-to-pixel relative position offset (BoxRPB) term to the cross-attention, so that the attention is more focused on the respective areas of interest while introducing smaller overhead. Finally, we propose an alignment loss to solve the misalignment problem between accurate classification and precise localization, further improving tracking performance. Through extensive experiments, we demonstrate that SDTrack achieves new SOTA performance on multiple benchmarks compared to previous work, while running at real-time speeds. Zihao Xia, Baojie Fan |
IROS | 3 |
| 2024 | Enhancing 3D Single Object Tracking with Efficient Point Cloud Segmentationabstract3D single object tracking (SOT) based on point cloud has attracted much attention due to its important role in machine vision and autonomous driving. Recently, M2-Track proposes a two-stage tracking structure centered on motion, but they ignore the effect of segmentation errors in sparse point cloud scenarios, which hinder the ability of networks to accurately represent tracking targets. To solve the problems, we propose an efficient 3D single object tracker (Abbr. EST) that can effectively segment point cloud features. Firstly, the proposed fusion segmentation module makes up for the feature loss caused by the downsampling strategy and enhances the ability of the network to recognize foreground points. In addition, the global embedded module is used to further focus on the crucial features of the target. This module provides global information by using residual networks and adding background information. Numerous experiments conducted on KITTI and NuScenes benchmarks show that EST achieves superior point cloud tracking in both performance and efficiency. Baojie Fan, Yuyu Jiang, Wuyang Zhou, Hongxin Xu |
IROS | 2 |
| 2024 | A Novel Image Formation Model for DescatteringabstractIn the field of image descattering, the image formation models employed for restoration approaches are often simplified. In these models, scattering distribution is uniform in homogeneous media when transmission is fixed. Through specifically designed experiments, we discover that scattering exhibits non-uniform characteristics even in homogeneous media. Neglecting non-uniform scattering in these models limits their accuracy in representing scattering distribution, resulting in existing image descattering approaches inadequate. To tackle these issues, this paper proposes a novel image formation model for image descattering, considering more physical parameters, such as zenith angle, azimuth angle, scattering phase function, and camera focal length. Our model describes the light transfer process in scattering media more accurately. For image descattering, we introduce corresponding algorithms for parameter estimation in our model and simultaneous restoration from degraded images. Experimental evaluations demonstrate the effectiveness of our proposed model in various tasks, including physical parameter estimation, pure-scattering removal, image dehazing, and underwater image restoration. In terms of calculating parameters, our results are close to the real values; in terms of underwater image restoration, our work outperforms the state-of-art methods; in terms of image dehazing, our work promotes the performance of existing methods by replacing previous models with our model. Jiandong Tian, Shijun Zhou, Baojie Fan, Hui Zhang 0023 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | HCPVF: Hierarchical Cascaded Point-Voxel Fusion for 3D Object DetectionabstractWith the astonishing development of 3D sensors, point cloud based 3D object detection is attracting increasing attention from both industry and academia, and widely applied in various fields, such as robotics and autonomous driving. However, how to balance the 3D object detecting accuracy and speed is still a challenging problem. In this paper, we study this issue and propose a novel and effective 3D point cloudy object detection network based on hierarchical cascaded point-voxel fusion, called HCPVF. Firstly, a novel bird’s-eye-view(BEV) attention mechanism with linear complexity is developed to improve point cloud feature backbone network, which can be implemented easily to mine the point-to-point similarity in BEV’s view, by two cascaded linear layers and two normalization layers. This operation captures long-range dependencies and reduces the uneven sampling of sparse BEV features, making the extracted point cloudy features more discriminative. Secondly, the proposed HCPVF module is equipped with dual-level hierarchical cascaded detection head, including voxel level and the following point level. The voxel level is composed of coarse Region of interest(RoI) pooling and fine RoI pooling, which are cooperated to aggregate voxel features from different grid divisions and predict relatively coarse detection boxes. In the following, the point level is based on Key Points Transformer. It firstly encodes the spatial context information between the original point and the voxel level box. And then, a novel dual-weighted decoder is developed to enhance the context interaction by weighting the channel and spatial dimensions to obtain more accurate detection results. This design utilizes the voxel based method with high computational efficiency and the point based method with more complete spatial information, fusing low-level voxel features and high-level point features through hierarchical cascaded strategy. Extensive experiments demonstate that the proposed HCPVF achieves state-of-the-art 3D detection performance while maintaining computational efficiency on both the Waymo Open Dataset and the highly-competitive KITTI benchmark. Baojie Fan, Jiandong Tian |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | QueryTrack: Joint-Modality Query Fusion Network for RGBT TrackingabstractExisting RGB-Thermal trackers usually treat intra-modal feature extraction and inter-modal feature fusion as two separate processes, therefore the mutual promotion of extraction and fusion is neglected. Then, the complementary advantages of RGB-T fusion are not fully exploited, and the independent feature extraction is not adaptive to modal quality fluctuation during tracking. To address the limitations, we design a joint-modality query fusion network, in which the intra-modal feature extraction and the inter-modal fusion are coupled together and promote each other via joint-modality queries. The queries are initialized based on the multimodal features of the current frame, making the subsequent fusion adaptive to modal quality fluctuation during tracking. Then the joint-modality query fusion (JQF) utilizes the queries to interact with RGB-T features, allowing the intra-modal enhancement and the inter-modal interactions to be unified for mutual promotion. In this way, JQF can distinguish and enhance the complementary modality features, while filtering out redundant information. For real-time tracking, we propose regional cross-attention for cross-modal interactions to reduce computational cost. Our end-to-end tracker sets a new state-of-the-art performance on multiple RGBT tracking benchmarks including LasHeR, VTUAV, RGBT234 and GTOT, while running at a real-time speed. Huijie Fan, Zhencheng Yu, Qiang Wang 0015, Baojie Fan, Yandong Tang |
IEEE Trans. Image Process. | 4 |
| 2023 | Mitigate the classification ambiguity via localization-classification sequence in object detection
Chang Liu 0082, Shaorong Xie, Xiaomao Li, Jiantao Gao, Weiping Xiao, Baojie Fan, Yan Peng 0001 |
Pattern Recognit. | 6 |
| 2023 | Two-Way Complementary Tracking GuidanceabstractRecently, most impressive Siamese network-based trackers are equipped with two independent branches: tracked object classification and bounding box regression. However, there is no tracking information exchange between them during the tracking optimization process. This may lead to the task-mismatch and accuracy inconsistency between both classification and regression branches during inference. To tackle the problems, we propose a novel Mutual Guidance (MG) strategy for visual object tracking, which constructs the bidirectional and complementary tracking information interaction to maintain the tracked object is well-classified to also be well-localized, between classification and regression branches. Specifically, the classification branch can guide the regression one to pay more attention to the sample with high classified scores, by re-weighting the regression loss with the classification confidence. Similarity, the regression branch also guides the classifier optimization process to focus on samples with larger IoU values. And then, the proposed Mutual Guidance is completed by a series of regularization designs on classification score and regression IoU, which dynamically re-assign the adaptive weights to the losses for each sample during the joint tracking optimization. The developed MG is generic and easy to be plugged into various tracking frameworks such as anchor-based, anchor-free based and transformer based, and boost their performance to some extent with negligible additional cost. In addition, we also develop an adaptive localization(L) branch selection scheme to further assist trackers, which determines proper localization branch for different trackers according to the difference in the way of discriminating positive and negative samples. Extensive experiments verify the effectiveness of MGL and its superiority against the state-of-the-art tracking modules on OTB100, GOT-10K, LaSOT, TrackingNet, UAV123, VOT2018 and VOT2019. Baojie Fan, Guoping Jiang, Jiandong Tian |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Object Tracking Network Based on Deformable Attention Mechanism
Baojie Fan, Xiaobin Guo |
BMVC | 2 |
| 2022 | Dual Aligned Siamese Dense Regression TrackerabstractAnchor or anchor-free based Siamese trackers have achieved the astonishing advancement. However, their parallel regression and classification branches lack the tracked target information link and interaction, and the corresponding independent optimization maybe lead to task-misalignment, such as the reliable classification prediction with imprecisely localization and vice versa. To address this problem, we develop a general Siamese dense regression tracker (SDRT) with both task and feature alignments. It consists of two cooperative and mutual-guidance core branches: dense local regression with RepPoint representation, the global and local multi-classifier fusion with aligned features. They complement and boost each other to constrain the results with well-localized followed to also be well-classified. Specifically, a dense local regression with RepPoint representation, directly estimates and averages multiple dense local bounding box offsets for accurate localization. And then, the refined bounding boxes can be used to learn the global and local affine alignment features for reliable multi-classifier fusion. The classified scores in turn guide the assigned positive bounding boxes for the regression task. The mutual guidance operations can bridge the connection between classification and regression substantially, since the assigned labels of one task depend on the prediction quality of the other task. The proposed tracking module is general, and it can boost both the anchor or anchor-free based Siamese trackers to some extent. The extensive tracking comparisons on six tracking benchmarks verify its favorable and competitive performance over states-of-the-arts tracking modules. Baojie Fan, Hui Zhang 0023, Yang Cong, Yandong Tang, Huijie Fan, Jiandong Tian |
IEEE Trans. Image Process. | 1 |
| 2022 | Discriminative Siamese Complementary Tracker With Flexible UpdateabstractThe offline generative Siamese trackers are equipped with the pre-defined anchors and the fixed target template. They overlook the target-background discriminative information, and lack the flexible target-specific update strategy. To overcome above drawbacks, we propose an adaptive and discriminative Siamese complementary tracking network with flexible update scheme. It consists of three collaborate subnetworks: anchor-free Siamese attention classification and regression subnetwork, online discriminative learning with multi-attention and multi-peak suppression, classifier guided template update subnetwork. All of them are interdependent and complementary to enhance each other for accurate target location. Specifically, an anchor-free multi-attention Siamese tracking subnetwork directly classifies the corresponding image patches with reliability assessment, and cascaded regresses the bounding boxes to progressively refine the predicting accuracy. Its evaluation is flexible and general with both proposal and anchor free in per-pixel prediction manner. Then, we integrate an online discriminative classifier optimizing module as a complementary subnetwork. It introduces spatial-temporal attention mechanism to fully explore multi-view multi-scale target-specific features, and evaluates multi-peak suppression to obtain a single centered peak response map. Its classified results can be fused with Siamese classification branch for accurate target location. Finally, the template update subnetwork is guided by the online discriminative classification scores. Extensive experiments on recent tracking datasets verify its top-ranked tracking accuracy and robustness against some state-of-the-art trackers. Baojie Fan, Jiandong Tian, Yan Peng 0001, Yandong Tang |
IEEE Trans. Multim. | 1 |
| 2022 | MedUCC: Medium-Driven Underwater Camera Calibration for Refractive 3-D ReconstructionabstractUnderwater camera calibration has attracted much attentions due to its significance in high-precision three-dimensional (3-D) pose estimation and scene reconstruction. However, most existing calibration methods focus on calibrating the underwater camera in a single scenario [e.g., air-glass-water], which can not well formulate the geometry constraint and further result in the complex calibration process. Moreover, the calibration precision of these methods is low, since multilayer transparent refractions with unknown layer orientation and distance make the task more difficult than that in air. To address these challenges, we develop a novel and efficient medium-driven method for underwater camera calibration (MedUCC), which can calibrate the underwater camera parameters, including the orientation and position of the transparent glass accurately. Our key idea of this article is to leverage the light-path changes formed by medium refractions between different media to acquire calibration data, which can better formulate the geometry constraint, and estimate the initial value of the underwater camera parameters. To improve the calibration accuracy of the underwater camera system, a quaternion-based solution is developed to refine the underwater camera parameters. To the end, we evaluate the calibration performance on an underwater camera system. Extensive experiment results demonstrate that our proposed method can obtain a better performance in comparison to the existing works. We also validate our proposed MedUCC method on our designed 3-D scanner prototype, which illustrates the superiority of our proposed calibration method. Changjun Gu, Yang Cong, Gan Sun, Yajun Gao, Tao Zhang 0084, Baojie Fan |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2021 | Spatial Graph Regularized Multi-kernel Subtask Cross-correlation TrackerabstractSome impressing multi-kernel or multi-task correlation filter trackers only focus on boosting the discrimination of multi-channel features, or exploiting the interdependence among different tasks. However, the cooperation and complementary of both technologies are missed, and the spatial structure among or inside target regions is also ignored. Therefore, this paper proposes a spatial graph regularized hierarchical subtask multi-kernel cross-correlation tracker (GHMK) via Gaussian process regression view, which enjoys the merits of multi-subtask multi-kernel learning and Gaussian process regression to jointly learn kernel cross-correlation filters, and makes them complement and boost each other. The interdependence and discrimination among multi-kernel multi-subtask are jointly exploited by group structure sparsity, which is also used to evaluate spatial feature selection. The spatial graph is constructed via cross similarity to maintain the geometric structure among or inside hierarchy subtasks. Besides, the developed model is general, and provides a unified solution from GPR for CF trackers without boundary effect. Comprehensive experiments demonstrate its favorable and competitive performance against the state-of-the-art trackers. Baojie Fan |
ICRA | 1 |
| 2021 | Dynamic and reliable subtask tracker with general schatten p-norm regularization
Baojie Fan, Yang Cong, Jiandong Tian, Yandong Tang |
Pattern Recognit. | 1 |
| 2021 | Structured and Consistent Multi-Layer Multi-Kernel Subtask Correction Filter TrackerabstractSome multi-task correlation filter trackers achieve the top-ranked performance in terms of accuracy and robustness. However, they directly fuse multiple types of features into a single kernel space. This operation fails to fully explore the discriminative strength and diversity of different features, and also ignores the structured correspondence of different tasks. To solve these issues, we propose a structured multi-kernel subtask correlation filter tracker with temporal-spatial consistency, which enjoys the merits of both layered multi-kernel subtask learning and structured correlation filter. Specifically, we firstly assign one kernel space to each channel feature. Multi-channel features correspond to multi-kernel spaces to boost their powerful discriminability. And then, we divide the target into multi-layer patches with different sizes, and regard the correlation filter trace of each patch with one channel feature as a subtask. In the following, we incorporate globally and locally structured correlation filters into a unified multi-kernel subtask particle tracking framework. The global and local subtasks complement and enhance each other with similar motion model. The proposed tracker not only exploits the cooperation and complementarity of layered multi-kernel subtask correlation filters, but also mines the underlying geometric structure of global subtasks, and the inner spatial locality correspondences of local subtasks inside the target. This operation is achieved by dual group sparsity regularized terms with mixed-norm lp,q, which decomposes the multi-kernel subtask filter matrix into two collaborative components. They correspond to the adaptive filter feature selection and outlier subtask detection, respectively. Besides, the developed tracking model maintains the temporal coherence and spatial consistency of multi-layer subtask filters via the smooth regularizer. Finally, the tracking formulation is optimized by the accelerated proximal gradient approach (APG). Encouraging analyses on six benchmark datasets, verify the favorable effectiveness and robustness of our method against state-of-the-art trackers. Baojie Fan, Yang Cong, Yandong Tang, Jiandong Tian, Chenliang Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Dual Refinement Underwater Object Detection Network
Baojie Fan, Yang Cong, Jiandong Tian |
ECCV (20) | 1 |
| 2020 | Locally Structured Multi-Task Multi-Kernel TrackerabstractMany impressive correlation filter trackers only construct a single kernel space for combined feature vector, or introduce independently multiple kernels with limit performance improvement, and neglect the spatial structure relationship among and inside target samples. In this paper, we propose a structured multi-task multi-kernel correlation particle filter tracker, which introduces multiple non-linear kernels for multiple channel features to fully take the strength of their powerful discriminability. Meanwhile, the developed tracker exploits the interdependencies and cooperations among different features to learn multi-task multi-kernel correlation filter jointly, and makes the learned multi-kernel filters complement and boost each other. Besides, by introducing cross patch similarity with the spatial constraints, the proposed tracker not only exploits the intrinsic structure information among target samples, but also preserves the spatial layout structure and correspondence of local image patches inside target samples. This is achieved by sparsity-induced mixed norm regulator. Furthermore, a shepherded particle filter tracking framework employs the affine motion modal to cover the states of target among adjacent frames as many as possible, including large scale variances and small move. Experiments on three popular datasets demonstrate that our tracker performs significantly better than other state-of-theart methods. Baojie Fan |
ICME | 1 |
| 2020 | Multi-classifier Guided Discriminative Siamese Tracking Network
Baojie Fan |
PRCV (2) | 2 |
| 2020 | Diverse receptive field network with context aggregation for fast object detection
Shaorong Xie, Chang Liu 0082, Jiantao Gao, Xiaomao Li, Jun Luo 0006, Baojie Fan, Jiahong Chen, Huayan Pu, Yan Peng 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2020 | Reliable Multi-Kernel Subtask Graph Correlation TrackerabstractMany astonishing correlation filter trackers pay limited concentration on the tracking reliability and locating accuracy. To solve the issues, we propose a reliable and accurate cross correlation particle filter tracker via graph regularized multi-kernel multi-subtask learning. Specifically, multiple non-linear kernels are assigned to multi-channel features with reliable feature selection. Each kernel space corresponds to one type of reliable and discriminative features. Then, we define the trace of each target subregion with one feature as a single view, and their multi-view cooperations and interdependencies are exploited to jointly learn multi-kernel subtask cross correlation particle filters, and make them complement and boost each other. The learned filters consist of two complementary parts: weighted combination of base kernels and reliable integration of base filters. The former is associated to feature reliability with importance map, and the weighted information reflects different tracking contribution to accurate location. The second part is to find the reliable target subtasks via the response map, to exclude the distractive subtasks or backgrounds. Besides, the proposed tracker constructs the Laplacian graph regularization via cross similarity of different subtasks, which not only exploits the intrinsic structure among subtasks, and preserves their spatial layout structure, but also maintains the temporal-spatial consistency of subtasks. Comprehensive experiments on five datasets demonstrate its remarkable and competitive performance against state-of-the-art methods. Baojie Fan, Yang Cong, Jiandong Tian, Yandong Tang |
IEEE Trans. Image Process. | 1 |
| 2019 | Novel event analysis for human-machine collaborative underwater exploration
Yang Cong, Baojie Fan, Dongdong Hou, Huijie Fan, Kaizhou Liu, Jiebo Luo 0001 |
Pattern Recognit. | 2 |
| 2019 | Context-aware long-term correlation tracking with hierarchical convolutional features
Baojie Fan, Huizhi Chen |
Pattern Recognit. Lett. | 1 |
| 2019 | Speedup 3-D Texture-Less Object Recognition Against Self-Occlusion for Intelligent ManufacturingabstractRealtime 3-D object detection and 6-DOF pose estimation in clutter background is crucial for intelligent manufacturing, for example, robot feeding and assembly, where robustness and efficiency are the two most desirable goals. Especially for various metal parts with a textless surface, it is hard for most state of the arts to extract robust feature from the clutter background with various occlusions. To overcome this, in this paper, we propose an online 3-D object detection and pose estimation method to overcome self-occlusion for textureless objects. For feature representation, we only adopt the raw 3-D point clouds with normal cues to define our local reference frame and we automatically learn the compact 3-D feature from the simple local normal statistics via autoencoder. For a similarity search, a new basis buffer k-d tree method is designed without suffering branch divergence; therefore, ours can maximize the GPU parallel processing capabilities especially in practice. We then generate the hypothesis candidates via the hough voting, filter the false hypotheses, and refine the pose estimation via the iterative closest point strategy. For the experiments, we build a new 3-D dataset including industrial objects with heavy self-occlusions and conduct various comparisons with the state of the arts to justify the effectiveness and efficiency of our method. Yang Cong, Dongying Tian, Baojie Fan |
IEEE Trans. Cybern. | 4 |
| 2018 | Structured and weighted multi-task low rank tracker
Baojie Fan, Xiaomao Li, Yang Cong, Yandong Tang |
Pattern Recognit. | 1 |
| 2018 | Online Similarity Learning for Big Data with OverfittingabstractIn this paper, we propose a general model to address the overfitting problem in online similarity learning for big data, which is generally generated by two kinds of redundancies: 1) feature redundancy, that is there exists redundant (irrelevant) features in the training data; 2) rank redundancy, that is non-redundant (or relevant) features lie in a low rank space. To overcome these, our model is designed to obtain a simple and robust metric matrix through detecting the redundant rows and columns in the metric matrix and constraining the remaining matrix to a low rank space. To reduce feature redundancy, we employ the group sparsity regularization, i.e., the `2;1 norm, to encourage a sparse feature set. To address rank redundancy, we adopt the low rank regularization, the max norm, instead of calculating the SVD as in traditional models using the nuclear norm. Therefore, our model can not only generate a low rank metric matrix to avoid overfitting, but also achieves feature selection simultaneously. For model optimization, an online algorithm based on the stochastic proximal method is derived to solve this problem efficiently with the complexity of O(d2). To validate the effectiveness and efficiency of our algorithms, we apply our model to online scene categorization and synthesized data and conduct experiments on various benchmark datasets with comparisons to several state-of-the-art methods. Our model is as efficient as the fastest online similarity learning model OASIS, while performing generally as well as the accurate model OMLLR. Moreover, our model can exclude irrelevant / redundant feature dimension simultaneously. Yang Cong, Ji Liu 0002, Baojie Fan, Peng Zeng 0001, Jiebo Luo 0001 |
IEEE Trans. Big Data | 3 |
| 2018 | Dual-Graph Regularized Discriminative Multitask TrackerabstractMultitask and low-rank learning methods have attracted increasing attention for visual tracking. However, most trackers only focus on learning appearance subspace basis or the sparse low rankness of representation and, thus, do not make full use of the structure information among and inside target candidates (or samples). In this paper, we propose a dual-graph regularized discriminative low-rank learning for a multitask tracker, which integrates the discriminative subspace and intrinsic geometric structures among tasks. By constructing dual-graph regulations from two views of multitask observation, the developed model not only exploits the intrinsic relationship among tasks, and preserves the spatial layout structure among the local patches inside each candidate, but also learns the salient features of the target samples. This operation has the benefit of having good target representation and improving the performance of the tracker. Moreover, our developed tracker is a collaborate multitask tracking model and learns the discriminative subspace with adaptive dimension and optimal classifier simultaneously. Then, a collaborate metric is developed to find the best candidate, which integrates both classification reliability and representation accuracy. Encouraging experimental results on a large set of public video sequences justify that our tracker performs favorably against many other state-of-the-art trackers. Baojie Fan, Yang Cong, Yandong Tang |
IEEE Trans. Multim. | 1 |
| 2017 | Consistent multi-layer subtask tracker via hyper-graph regularization
Baojie Fan, Yang Cong |
Pattern Recognit. | 1 |
| 2017 | Layered Multitask Tracker via Spatial-Temporal Laplacian GraphabstractMost multitask trackers define the trace of each candidate as one task, and assume all tasks are equally related. Multitask learning is only evaluated on the current frame. In fact, these assumptions are limited, and ignore the multitask relationship in consecutive frames. In this letter, we propose a discriminative layered multitask tracker via spatial-temporal Laplacian graphs, which defines the layered tasks from a novel view, and naturally incorporates the global and local target information into reverse multitask tracking process. The spatial-temporal Laplacian graphs not only exploit the sequential consistent information of the target, but also make full use of the geometric structure corresponding to the tasks among the adjacent frames. Besides, l0norm constraint and labeling information are used to improve the tracking robustness. Encouraging experimental results on challenging sequences justify that the proposed method performs well both in accuracy and robustness against some related trackers. Baojie Fan, Xiaomao Li, Yang Cong |
IEEE Signal Process. Lett. | 1 |
| 2017 | Multi-Class Latent Concept Pooling for Computer-Aided Endoscopy DiagnosisabstractSuccessful computer-aided diagnosis systems typically rely on training datasets containing sufficient and richly annotated images. However, detailed image annotation is often time consuming and subjective, especially for medical images, which becomes the bottleneck for the collection of large datasets and then building computer-aided diagnosis systems. In this article, we design a novel computer-aided endoscopy diagnosis system to deal with the multi-classification problem of electronic endoscopy medical records (EEMRs) containing sets of frames, while labels of EEMRs can be mined from the corresponding text records using an automatic text-matching strategy without human special labeling. With unambiguous EEMR labels and ambiguous frame labels, we propose a simple but effective pooling scheme called Multi-class Latent Concept Pooling, which learns a codebook from EEMRs with different classes step by step and encodes EEMRs based on a soft weighting strategy. In our method, a computer-aided diagnosis system can be extended to new unseen classes with ease and applied to the standard single-instance classification problem even though detailed annotated images are unavailable. In order to validate our system, we collect 1,889 EEMRs with more than 59K frames and successfully mine labels for 348 of them. The experimental results show that our proposed system significantly outperforms the state-of-the-art methods. Moreover, we apply the learned latent concept codebook to detect the abnormalities in endoscopy images and compare it with a supervised learning classifier, and the evaluation shows that our codebook learning method can effectively extract the true prototypes related to different classes from the ambiguous data. Shuai Wang 0003, Yang Cong, Huijie Fan, Baojie Fan, Lianqing Liu, Yunsheng Yang, Yandong Tang, Huaici Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2016 | Structured low rank tracker with smoothed regularizationabstractIn this paper, we propose a structured low rank learning algorithm with smoothed regularization for robust object tracking, under particle filter framework. Specifically, the relationships among the particles are exploited with structured low rank regularization term, and simultaneously handle the outlier using a group sparsity regularization. The label information from training data is incorporated into the tracking objective function as the classification error term and idea coding regularization term respectively. By the smoothed regularization, the developed structured low rank learning based tracker can be efficiently solved by iterative reweighed least squares algorithm(IRLS), and avoids svd operation. Moreover, the collaborate normalized metric is developed to find the best candidate. Compared with some state-of-the-art tracking methods on 50 challenging sequences, the proposed algorithms perform well in terms of accuracy, robustness. Baojie Fan, Yang Cong, Xiaomao Li, Yandong Tang |
VCIP | 1 |
| 2016 | UDSFS: Unsupervised deep sparse feature selection
Yang Cong, Shuai Wang 0003, Baojie Fan, Yunsheng Yang |
Neurocomputing | 3 |
| 2015 | Speeded Up Low-Rank Online Metric Learning for Object TrackingabstractVisual object tracking can be considered as an online procedure to adaptively measure the foreground object similarity itself. However, many previous works usually adopt a fixed metric or offline metric learning to evaluate this dynamic process; even with some online metric learning (OML) trackers, their models often suffer from overfitting issues. To overcome these deficiencies, we propose a self-supervised tracking method that incorporates adaptive metric learning and semisupervised learning into a unified framework. For similarity measurement, we design a new OML model via low-rank constraint to handle overfitting. In particular, we employ the max norm instead of the trace norm used in our previous work. This not only maintains the low-rank property to overcome overfitting, but also reduces the computational complexity from O(n3) to O(n2), such that the new model is more suitable for object tracking. Moreover, by associating the information from stored training templates with unlabeled testing samples, a bilinear graph is defined accordingly to propagate the label of each sample. High-confidence samples are then collected for self-training the model and updating the templates concurrently to handle large scale. Experiments on various benchmark data sets and comparisons to several state-of-the-art methods demonstrate the effectiveness and efficiency of our algorithm. Yang Cong, Baojie Fan, Ji Liu 0002, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Online discriminative dictionary learning via label information for multi task object trackingabstractIn this paper, a supervised approach to online learn a structured sparse and discriminative representation for object tracking is presented. Label information from training data is incorporated into the dictionary learning process to construct a compact and discriminative dictionary. This is accomplished by adding an ideal-code regularization term and classification error term to the total objective function. By minimizing the total objective function, we learn the high quality dictionary and optimal linear multi-classifier simultaneously. Combined with multi task sparse learning, the learned classifier is employed directly to separate the object from background. As the tracking continues, the proposed algorithm alternates between multi task sparse coding and dictionary updating. Experimental evaluations on the challenging sequences show that the proposed algorithm performs favorably against state-of-the-art methods in terms of effectiveness, accuracy and robustness. Baojie Fan, Yingkui Du, Hao Gao 0005, Baoyun Wang |
ICME | 1 |
| 2014 | A Unified Online Dictionary Learning Framework with Label Information for Robust Object TrackingabstractIn this paper, a supervised approach to online learn a structured sparse and discriminative representation for object tracking is presented. Label information from training data is incorporated into the dictionary learning process to construct a robust and discriminative dictionary. This is accomplished by adding an ideal-code regularization term and classification error term to the unified objective function. By minimizing the unified objective function we learn the high quality dictionary and optimal linear multi-classifier jointly. Combined with robust sparse coding, the learned classifier is employed directly to separate the object from background. As the tracking continues, the proposed algorithm alternates between robust sparse coding and dictionary updating. Experimental evaluations on the challenging sequences show that the proposed algorithm performs favorably against state-of-the-art methods in terms of effectiveness, accuracy and robustness. Baojie Fan, Yang Cong, Yingkui Du |
ICPR | 1 |
| 2014 | Discriminative multi-task objects tracking with active feature selection and drift correction
Baojie Fan, Yang Cong, Yingkui Du |
Pattern Recognit. | 1 |
| 2014 | A Hybrid Particle-Swarm Tabu Search Algorithm for Solving Job Shop Scheduling ProblemsabstractThis paper proposes a method for the job shop scheduling problem (JSSP) based on the hybrid metaheuristic method. This method makes use of the merits of an improved particle swarm optimization (PSO) and a tabu search (TS) algorithm. In this work, based on scanning a valuable region thoroughly, a balance strategy is introduced into the PSO for enhancing its exploration ability. Then, the improved PSO could provide diverse and elite initial solutions to the TS for making a better search in the global space. We also present a new local search strategy for obtaining better results in JSSP. A real-integer encode and decode scheme for associating a solution in continuous space to a discrete schedule solution is designed for the improved PSO and the tabu algorithm to directly apply their solutions for intensifying the search of better solutions. Experimental comparisons with several traditional metaheuristic methods demonstrate the effectiveness of the proposed PSO-TS algorithm. Hao Gao 0005, Sam Kwong, Baojie Fan, Ran Wang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2013 | Novel Dominant Plant Detection Algorithm for Image SequenceabstractDominant plane is an important geometric feature and can be used in a wide range of applications. In this paper, we develop a robust technology for dominant plane detection using the progressive structure from motion and dominant plane fitting algorithms. Firstly, we propose a progressive structure from motion algorithm to reconstruct the real scene from the image sequences, and obtain the information of three-dimensional points in the scene. In the following, based on the least median square (LMedSq) estimation and ransac theory, we present a novel plane fitting algorithm to find the dominant plane near the scene surface from the reconstructive points. Experimental results from different outdoor scenarios show that the proposed reconstruction algorithm obtains dense 3D point cloud, and achieves satisfactory recovery from image sequence. Then, the plane fitting algorithm accurately detects the dominant plane region from the reconstructive points. The detected plane is the approximation for local surface and meets the actual scene. These tests verify that the proposed method is effective, accurate and robust. Baojie Fan, Yingkui Du |
ICIG | 1 |
| 2013 | Robust and accurate online pose estimation algorithm via efficient three-dimensional collinearity modelabstractIn this study, the authors propose a robust and high accurate pose estimation algorithm to solve the perspective‐ N ‐point problem in real time. This algorithm does away with the distinction between coplanar and non‐coplanar point configurations, and provides a unified formulation for the configurations. Based on the inverse projection ray, an efficient collinearity model in object–space is proposed as the cost function. The principle depth and the relative depth of reference points are introduced to remove the residual error of the cost function and to improve the robustness and the accuracy of the authors pose estimation method. The authors solve the pose information and the depth of the points iteratively by minimising the cost function, and then reconstruct their coordinates in camera coordinate system. In the following, the optimal absolute orientation solution gives the relative pose information between the estimated three‐dimensional (3D) point set and the 3D mode point set. This procedure with the above two steps is repeated until the result converges. The experimental results on simulated and real data show that the superior performance of the proposed algorithm: its accuracy is higher than the state‐of‐the‐art algorithms, and has best anti‐noise property and least deviation by the influence of outlier among the tested algorithms. Baojie Fan, Yingkui Du, Yang Cong |
IET Comput. Vis. | 1 |
| 2012 | Active drift correction template tracking algorithmabstractThis paper presents a novel active drift correction template tracking algorithm. Compared to Matthews' algorithm in [8], the proposed algorithm achieves synchronously object tracking and drift correction, and save half running time. For the template drift problem during long sequential object tracking, we introduce the active drift correction term into inverse compositional affine image alignment algorithm. This operation can avoid the template drift before it occurs, or reduce the drift after it happens. The total energy function consists of two terms: the tracking term and the active drift correction term. By minimizing the total energy function with the steepest descent algorithm, the proposed algorithm can decrease the accumulative tracking error, and prevent the drift during the tracking process effectively. Various object tracking experiments show that our method has super performance than the passive drift correction algorithm in [8]. Baojie Fan, Yingkui Du, Yang Cong, Yandong Tang |
ICIP | 1 |
| 2011 | A robust template tracking algorithm with weighted active drift correction
Baojie Fan, Yingkui Du, Yandong Tang |
Pattern Recognit. Lett. | 1 |