Bin Fan 0001

dblp:60/105-1 · DBLP profile ↗
← Back
92ranked-venue papers
15as first author
46since 2021 · last 2026
0000-0002-1155-467XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 7 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 53 · 8 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Group Orthogonal Low-Rank Adaptation for RGB-T Tracking
abstract
Parameter-efficient fine-tuning has emerged as a promising paradigm in RGB-T tracking, enabling downstream task adaptation by freezing pretrained parameters and fine-tuning only a small set of parameters. This set forms a rank space made up of multiple individual ranks, whose expressiveness directly shapes the model's adaptability. However, quantitative analysis reveals low-rank adaptation exhibits significant redundancy in the rank space, with many ranks contributing almost no practical information. This hinders the model's ability to learn more diverse knowledge to address the various challenges in RGB-T tracking. To address this issue, we propose the Group Orthogonal Low-Rank Adaptation (GOLA) framework for RGB-T tracking, which effectively leverages the rank space through structured parameter learning. Specifically, we adopt a rank decomposition partitioning strategy utilizing singular value decomposition to quantify rank importance, freeze crucial ranks to preserve the pretrained priors, and cluster the redundant ranks into groups to prepare for subsequent orthogonal constraints. We further design an inter-group orthogonal constraint strategy. This constraint enforces orthogonality between rank groups, compelling them to learn complementary features that target diverse challenges, thereby alleviating information redundancy. Experimental results demonstrate that GOLA effectively reduces parameter redundancy and enhances feature representation capabilities, significantly outperforming state-of-the-art methods across four benchmark datasets and validating its effectiveness in RGB-T tracking tasks.
Zekai Shao 0002, Yufan Hu, Bin Fan 0001, Hongmin Liu 0001
AAAI4
2026 Towards Accurate 3D Object Detection in Adverse Weather by Leveraging 4D Radar for LiDAR Geometry Enhancement
abstract
3D object detection is a critical component of autonomous driving, yet its performance degrades severely in adverse weather due to the degradation of LiDAR point clouds. While existing LiDAR-4D radar fusion methods enhance robustness by incorporating weather-robust 4D radar data, they often depend on well geometric structures from LiDAR and so struggle to effectively exploit radar data in case of degraded LiDAR data. To tackle this challenge, we propose REL, a novel 4D radar-guided LiDAR geometric enhancement framework. It utilizes 4D radar features to dynamically generate virtual LiDAR points, effectively increasing the density of degraded LiDAR data. Moreover, a Position-Guided Cross Attention (PGCA) module is proposed to enhance the feature representation of virtual points, while an Adaptive Feature Fusion (AFF) module is designed to integrate virtual and real LiDAR features. Extensive experiments on the K-Radar and Vod-Fog datasets demonstrate that REL achieves state-of-the-art 3D object detection performance under diverse adverse weather conditions. Notably, REL improves the overall AP3D by 9.3% on K-Radar and boosts the cyclist class by up to 52.9% 3D mAP under the most severe foggy condition on Vod-Fog.
Tianxu Tong, Xinrun Liu, Hongmin Liu 0001, Bin Fan 0001
AAAI4
2026 From Image to Pixels: Towards Fine-Grained Medical Vision-Language Models
abstract
Multimodal large language models (MLLMs) offer immense potential for biomedical AI, yet current applications remain limited to coarse-grained image understanding and basic textual queries-falling short of the fine-grained reasoning required in clinical contexts. In this work, we present a comprehensive solution spanning data, model, and training innovations to advance pixel-level multimodal intelligence in biomedicine. First, we construct MeCoVQA, a new visual-language benchmark that spans eight medical imaging modalities and four core tasks, supporting both spatially-grounded reasoning and fine-grained diagnostic comprehension. Building on this, we introduce MedPLIB, an end-to-end biomedical MLLM equipped with pixel-level visual understanding. MedPLIB supports diverse multimodal tasks-including VQA, point- and region-based querying, grounding, and segmentation-through unified modeling. To further accommodate the heterogeneous nature of biomedical tasks, we design a task-specialized Mixture-of-Experts (MoE) architecture, where each expert is tailored to a specific task and jointly optimized via unified fine-tuning. This modular design accommodates diverse biomedical tasks while maintaining a unified and efficient architecture. By integrating retrieval-augmented generation (RAG) and in-context learning (ICL), MedPLIB also demonstrates strong generalization on out-of-distribution (OOD) medical image segmentation. Experiments across multiple benchmarks show that MedPLIB sets a new state-of-the-art on biomedical vision-language tasks; notably, it outperforms the best existing small and large models by 19.7 and 15.6 mDice in zero-shot pixel-level grounding, highlighting its clinical utility and generalization strength.
Lingdong Shen, Xiaoshuang Huang, Fangxin Shang, Yehui Yang, Bin Fan 0001, Shiming Xiang
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Unifying RGB and thermal object detection in one detector
Bin Fan 0001, Wei Zhao 0019, Yongjie Chen, Hongmin Liu 0001
Pattern Recognit.1
2026 IIFNet3D: Instance-to-instance fusion with dual attention for indoor RGB-D 3D object detection
Zixin Fan, Bin Fan 0001, Hongmin Liu 0001
Pattern Recognit.3
2026 OV-GT3D: A generalizable open-vocabulary two-stage 3D detector with dual path distillation
Xiuwei Xu, Bin Fan 0001, Jiwen Lu, Hongmin Liu 0001
Pattern Recognit.3
2026 MC-MVSNet: When multi-view stereo meets monocular cues
Xincheng Tang, Mengqi Rong, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen
Pattern Recognit.3
2026 Global-local synergistic fine-tuning of DINOv3 for underwater instance segmentation
Bin Fan 0001
Pattern Recognit. Lett.2
2025 PURA: Parameter Update-Recovery Test-Time Adaption for RGB-T Tracking
abstract
Maintaining stable tracking of objects in domain shift scenarios is crucial for RGB-T tracking, prompting us to explore the use of unlabeled test sample information for effective online model adaptation. However, current Test-Time Adaptation (TTA) methods in RGB-T tracking dramatically change the model’s internal parameters during long-term adaptation. At the same time, the gradient computations involved in the optimization process impose a significant computational burden. To address these challenges, we propose a Parameter Update-Recovery Adaptation (PURA) framework based on parameter decomposition. Firstly, our fast parameter update strategy adjusts model parameters using statistical information from test samples without requiring gradient calculations, ensuring consistency between the model and test data distribution. Secondly, our parameter decomposition recovery employs orthogonal decomposition to identify the principal update direction and recover parameters in this direction, aiding in the retention of critical knowledge. Finally, we leverage the information obtained from decomposition to provide feedback on the momentum during the update phase, ensuring a stable updating process. Experimental results demonstrate that PURA outperforms current state-of-the-art methods across multiple datasets, validating its effectiveness. The project page is at https://melantech.github.io/PURA.
Zekai Shao 0002, Yufan Hu, Bin Fan 0001, Hongmin Liu 0001
CVPR3
2025 Learning Evidential Delta Denoising Scores for Video Editing
Yufan Hu, Junyu Gao 0002, Bin Fan 0001, Hongmin Liu 0001
ACM Multimedia4
2025 Uncertainty Aware Multiple View Stereo Network with Accurate Supervision
abstract
Learning-based multiple view stereo has gained significant attention recently. However, most methods rely on direct network supervision using provided ground-truth depth, which poses three inherent problems: resolution-dependent ground-truth artifacts, excessively challenging training examples (with relatively featureless textures), and use of less-viewed reference pixels for supervision, all of which hinder network optimization. To alleviate these problems, we propose an accurate network supervision paradigm that includes a ground-truth mask, an entropy mask, and a consistency mask, which provide more accurate supervision signals to aid network optimization. Furthermore, we introduce UANet, an uncertainty aware multi-view stereo network, which adaptively determines a pixel-wise search range using a dynamic range sampler (DRS) built upon estimation confidence and learned uncertainty. Experimental results on recent MVS datasets demonstrate the effectiveness of our method.
Xincheng Tang, Mengqi Rong, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen
Comput. Vis. Media3
2025 ProIn: Learning to predict trajectory based on progressive interactions for autonomous driving
Yinke Dong, Haifeng Yuan, Hongkun Liu, Fangzhen Li, Hongmin Liu 0001, Bin Fan 0001
Neurocomputing7
2025 SARFormer: Segmenting Anything Guided Transformer for semantic segmentation
Lixin Zhang 0004, Wenteng Huang, Bin Fan 0001
Neurocomputing3
2025 Bidirectional Agent-Map Interaction Feature Learning Leveraged by Map-Related Tasks for Trajectory Prediction in Autonomous Driving
abstract
Accurate prediction of the future trajectories of surrounding agents is essential for safely autonomous vehicles. However, it is quite challenging due to the dynamics of driving environments caused by frequent agent-agent and agent-map interactions. Most existing methods mainly focus on modeling the dynamics of agents but overlook the inherent dynamics of the usable map at different driving moments. To address this limitation, this paper proposes a bidirectional interaction network (DyMap) that takes into account the bilateral constraints of the map on the agents and the impact of the agents on the map, using a unique bidirectional attention module. By employing multiple stages of bidirectional interaction modules, the features of agents and map nodes are continuously updated to adapt to changing environments. To facilitate network learning, we further incorporate map-related auxiliary tasks into the learning objectives, including traffic flow prediction, map node occupancy prediction, and road direction prediction. These tasks are designed to help the network learn map features that can reflect the dynamics of the driving environment, leading to a comprehensive feature representation of the agents. This representation encodes not only the agent’s historical motions but also the latent map information triggered by other agents. This proposed network offers a new approach to modeling environmental dynamics in driving scenarios, and achieves the state-of-the-art performance on the Argoverse 1 and Argoverse 2 trajectory prediction benchmarks. Note to Practitioners—Autonomous vehicles are of particular interest to practitioners in the field of automation science and technology. Trajectory prediction is of critical importance to the safety of autonomous driving, which aims to forecast the future positions of a given traffic participator (called the agent in this paper, such as a vehicle, pedestrian, etc.) based on the observations of itself and other agents’ historical trajectories as well as the map information. This paper proposes a trajectory prediction network based on a novel bidirectional interaction module that enables bidirectional information flow between the features of map nodes and agents so as to simultaneously learn their feature representations being aware of the timing-vary property of traffic scenes. The proposed DyMap contains three blocks of bidirectional interaction modules, which are interleaved with self-interactions of the agents and map nodes. Three different kinds of map-related tasks are developed to facilitate network training by leveraging on maximizing the map features for predicting traffic flow, occupancy, and lane direction. The superiority of the proposed method has been confirmed by the widely used benchmarks in the community. Besides modelling the complex interaction between agents and map nodes in dynamic traffic scenarios, the proposed bidirectional interaction network can be beneficial to address other problems requiring symmetric modelling of interactions between different components, such as human-robot interaction. The proposed map-related tasks can be directly used in other applications requiring to extract map features while considering the time-varying conditions, including but not limited to robot navigation and embody intelligence.
Bin Fan 0001, Haifeng Yuan, Yinke Dong, Zhengyu Zhu 0002, Hongmin Liu 0001
IEEE Trans Autom. Sci. Eng.1
2025 DMotion: Diverse Modalities Alignment Enhanced Motion Prediction for Autonomous Driving
abstract
In autonomous driving, motion prediction is vital for anticipating the behaviors of surrounding vehicles, pedestrians, and other road users, enabling the system to make accurate decisions and plan driving paths effectively. Current motion prediction models typically utilize an encoder–decoder architecture, and many methods focus on the decoder’s design because the decoder is directly responsible for generating future trajectories. However, this often leads to neglecting the encoder’s capability to represent input information, resulting in a failure to provide the decoder with accurate and semantic prior features, which impacts the overall prediction accuracy. Contrastive learning, as an effective approach for enhancing feature representation through cross-modal alignment, demonstrates strong generalization and reduced reliance on labeled data. Therefore, we propose a motion prediction network named DMotion that leverages contrastive learning to align trajectory features with numeric signals and textual descriptions, improving the model’s representational capacity and enriching contextual priors. To the best of our knowledge, this is the first approach to use textual descriptions as a modality to enhance motion prediction accuracy. We sparsify dense agent attribute labels, such as historical distance and angular variation, to enable these prior features learned by the model through contrastive learning. By investigating the impact of numeric and textual supervision signals on contrastive learning effectiveness, textual supervision achieves superior results compared with numeric signals, benefiting from richer input information and more robust extraction capabilities of the text model. We further apply the low-rank adaptation (LoRA) method to fine-tune the text encoder, improving model performance and preventing catastrophic forgetting with only 0.1M additional trainable parameters. Our experiments demonstrate that DMotion shows competitive performance on the Waymo motion prediction and interaction prediction challenges. Additionally, the contrast learning module of DMotion does not introduce additional parameters or computational overhead during inference, maintaining the efficiency of the original encoder-decoder model.
Hongkun Liu, Hongmin Liu 0001, Bin Fan 0001, Jinglin Xu
IEEE Trans. Comput. Soc. Syst.3
2025 BRTAL: Boundary Refinement Temporal Action Localization via Offset-Driven Diffusion Models
abstract
Temporal Action Localization (TAL) aims to classify and localize all actions within untrimmed videos. Existing TAL methods often struggle with inaccurate boundary predictions due to the similarity of action content and the uncertainty of boundaries between adjacent frames. Many of these methods rely on fixed or global proposal learning strategies, which lack a more refined method to improve localization accuracy. In this paper, we propose BRTAL, a new Boundary Refinement framework for TAL based on an offset-driven diffusion model, specifically designed to enhance action boundary precision through a refined approach iteratively. Unlike traditional TAL methods emphasizing global target predictions, BRTAL adopts a local refinement perspective by leveraging an offset-driven strategy. Specifically, our framework employs diffusion to iteratively generate local offsets between predictions and ground truth, gradually reducing these offsets to achieve better alignment with the ground truth. This refined approach is particularly effective in addressing the challenges of ambiguous boundaries frequently encountered in TAL, enabling BRTAL to achieve more refined boundary localization than existing methods. Furthermore, we introduce a lightweight yet powerful Temporal Context Modeling (TCM) module to enhance temporal information modeling for accurate action localization. TCM features a Temporal Representation Perception (TRP) layer, which captures temporal evolution and long-term contextual dependencies through a squeeze-and-excitation design combined with large convolutional kernels, ensuring robust temporal representation learning. Extensive experiments on THUMOS14, ActivityNet-1.3, and EPIC-KITCHEN 100 datasets highlight the significant advantages of BRTAL. Notably, BRTAL achieves an average mAP of 69.6% on THUMOS14, establishing a new state-of-the-art benchmark and demonstrating its outstanding boundary refinement capability.
Hongmin Liu 0001, Xueli Li, Bin Fan 0001, Jinglin Xu
IEEE Trans. Circuits Syst. Video Technol.3
2025 Dynamic Learnable Label Assignment for Indoor 3D Object Detection
abstract
In this paper, we present a dynamic learnable label assignment (DLLA) method for indoor anchor-free one-stage 3D object detection. Existing methods principally depend on hand-crafted strategies with fixed thresholds, which fail to adapt to the inherent variability in object characteristics such as size, shape, and occlusion levels. This lack of adaptability results in suboptimal sample assignments and unstable detection performance. To address this challenge, we map the features of proposals and ground truths separately into the same embedding space, enabling a dynamic strategy of assigning appropriate positive samples to each instance. Specifically, we first interact with the features of all proposals to effectively integrate information from each proposal in the scene and capture long-range dependencies between different locations. Additionally, to extract more discriminative and generalized features for positive and negative samples, we employ a contrastive learning process to optimize the elemental relationships and distances between proposals and ground truths. Finally, we introduce a denoising task to alleviate the difficulty of the unsupervised learning process in DLLA. Experimental results show that our DLLA outperforms other methods on three popular indoor datasets (ScanNet V2, SUN RGB-D, and ScanNet200).
Xinrun Liu, Linqing Zhao, Bin Fan 0001, Jiwen Lu, Hongmin Liu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Learning Boundary Continuity-Aware Gaussian Encoder for Oriented Object Detection
abstract
Oriented object detection has been crucial for rotation-sensitive tasks and has garnered significant attention. Most existing methods generate angles as detector output vectors, but this strategy can abnormally magnify visually similar differences between two boxes in certain circumstances, termed boundary discontinuity issue. To overcome this limitation, we propose a boundary continuity-aware Gaussian encoder (BCGE). Specifically, BCGE directly predicts target Gaussian distributions for proposals and learns an oriented bounding box as an integrated 2-D matrix, effectively addressing boundary discontinuity issues. We also propose a transformation from Gaussian representation back to boxes and extend this transformation theory to the complex domain to adapt to the learning characteristics of neural networks. Furthermore, BCGE serves as a versatile plug-and-play architectural encoder, directly replacing the standard coding process in various oriented detectors with adaptability. Experimental results on five popular datasets, i.e., DOTA, UCAS-AOD, HRSC2016, SSDD, and HRSID, consistently show the effectiveness of our approach.
Hongmin Liu 0001, Chengyi Zhao, Bin Fan 0001, Yufan Hu
IEEE Trans. Cybern.3
2025 Dual-Level Modality De-Biasing for RGB-T Tracking
abstract
RGB-T tracking aims to effectively leverage the complement ability of visual (RGB) and infrared (TIR) modalities to achieve robust tracking performance in various scenarios. Existing RGB-T tracking methods typically adopt backbone networks pre-trained on large-scale RGB datasets, which can lead to a predisposition toward RGB image patterns. RGB and TIR modalities also exhibit inconsistent responses to regions with diverse properties, resulting in imbalances in tracking decisions. We refer to these issues as feature-level and decision-level biases in the TIR modality. In this paper, we propose a novel dual-level modality de-biasing framework for RGB-T tracking to eliminate the inherent feature and decision-level biases. Specifically, we propose a joint infrared-fusion adapter, comprising an infrared-aware adapter and a cross-fusion adapter, designed to adaptively mitigate feature-level biases and utilize complementary information between the two modalities. In addition to implicit feature-level adjustment, we propose a response-decoupled distillation strategy to explicitly alleviate decision-level biases, aiming to achieve consistently accurate decision-making between the RGB and TIR modalities. Extensive experiments on several popular RGB-T tracking benchmarks validate the effectiveness of our proposed method.
Yufan Hu, Zekai Shao 0002, Bin Fan 0001, Hongmin Liu 0001
IEEE Trans. Image Process.3
2025 ScalableTrack: Scalable One-Stream Tracking via Alternating Learning
abstract
Transformer-based one-stream trackers are widely used to extract features and interact information for visual object tracking. However, the current one-stream tracker has fixed computational dimensions between different stages, which limits the network's ability to learn context clues and global representations, resulting in a decrease in the ability to distinguish between targets and backgrounds. To address this issue, a new scalable one-stream tracking framework, ScalableTrack, is proposed. It unifies feature extraction and information integration by intrastage mutual guidance, leveraging the scalability of target-oriented features to enhance object sensitivity and obtain discriminative global representations. In addition, we bridge interstage contextual cues by introducing an alternating learning strategy and solve the arrangement problem of the two modules. The alternating learning strategy uses alternate stacks of feature extraction and information interaction to focus on tracked objects and prevent catastrophic forgetting of target information between different stages. Experiments on eight challenging benchmarks (TrackingNet, GOT-10k, VOT2020, UAV123, LaSOT, LaSOText, OTB100, and TC128) show that ScalableTrack outperforms state-of-the-art (SOTA) methods with better generalization and global representation ability.
Hongmin Liu 0001, Yuefeng Cai, Bin Fan 0001, Jinglin Xu
IEEE Trans. Neural Networks Learn. Syst.3
2025 A Cascaded Multimodule Image Enhancement Framework for Underwater Visual Perception
abstract
Underwater images usually exhibit severe color cast, hazy appearance, and/or dark regions because of the complex lighting absorption and scattering in water. How to increase the quality of these degraded underwater images has emerged as a key issue for various underwater application tasks. Recent efforts have been made to deal with single type degradation, however, it is still challenging to deal with multiple degradations that usually coexist in an underwater image with a general network. The degradations in underwater images can be divided into medium-agnostic (hazy or low-light which also encountered in in-air images) and medium-specific (color distortion caused by the specific light attenuation property in water) ones. According to this observation, this article proposes a cascaded multimodule underwater image enhancement (UIE) framework to address the coexisted multiple degradations. In the proposed framework, an in-air image enhancement module and a novel proposed adaptive color channel compensation network (AC3Net) are cascaded, in which the former focuses primarily on solving medium-agnostic degradations and the latter is for handling the medium-specific degradation. This framework has good flexibility by cascading different types of in-air image enhancement networks with AC3Net to achieve various UIE. The effectiveness of the proposed framework has been extensively validated on various degraded underwater images as well as different underwater visual perception tasks.
Hongmin Liu 0001, Hui Zeng 0003, Huayan Pu, Jun Luo 0003, Bin Fan 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Defying Imbalanced Forgetting in Class Incremental Learning
abstract
We observe a high level of imbalance in the accuracy of different learned classes in the same old task for the first time. This intriguing phenomenon, discovered in replay-based Class Incremental Learning (CIL), highlights the imbalanced forgetting of learned classes, as their accuracy is similar before the occurrence of catastrophic forgetting. This discovery remains previously unidentified due to the reliance on average incremental accuracy as the measurement for CIL, which assumes that the accuracy of classes within the same task is similar. However, this assumption is invalid in the face of catastrophic forgetting. Further empirical studies indicate that this imbalanced forgetting is caused by conflicts in representation between semantically similar old and new classes. These conflicts are rooted in the data imbalance present in replay-based CIL methods. Building on these insights, we propose CLass-Aware Disentanglement (CLAD) as a means to predict the old classes that are more likely to be forgotten and enhance their accuracy. Importantly, CLAD can be seamlessly integrated into existing CIL methods. Extensive experiments demonstrate that CLAD consistently improves current replay-based methods, resulting in performance gains of up to 2.56%.
Shixiong Xu, Gaofeng Meng, Xing Nie, Bolin Ni, Bin Fan 0001, Shiming Xiang
AAAI5
2024 Point-Voxel Based Geometry-Adaptive Network for 3D Point Cloud Analysis
Tianmeng Zhao, Hui Zeng 0003, Baoqing Zhang, Hong-Min Liu, Bin Fan 0001
J. Comput. Sci. Technol.5
2024 Reusable Architecture Growth for Continual Stereo Matching
abstract
The remarkable performance of recent stereo depth estimation models benefits from the successful use of convolutional neural networks to regress dense disparity. Akin to most tasks, this needs gathering training data that covers a number of heterogeneous scenes at deployment time. However, training samples are typically acquired continuously in practical applications, making the capability to learn new scenes continually even more crucial. For this purpose, we propose to perform continual stereo matching where a model is tasked to 1) continually learn new scenes, 2) overcome forgetting previously learned scenes, and 3) continuously predict disparities at inference. We achieve this goal by introducing a Reusable Architecture Growth (RAG) framework. RAG leverages task-specific neural unit search and architecture growth to learn new scenes continually in both supervised and self-supervised manners. It can maintain high reusability during growth by reusing previous units while obtaining good performance. Additionally, we present a Scene Router module to adaptively select the scene-specific architecture path at inference. Comprehensive experiments on numerous datasets show that our framework performs impressively in various weather, road, and city circumstances and surpasses the state-of-the-art methods in more challenging cross-dataset settings. Further experiments also demonstrate the adaptability of our method to unseen scenes, which can facilitate end-to-end stereo architecture learning and practical deployment.
Chenghao Zhang 0003, Gaofeng Meng, Bin Fan 0001, Zhaoxiang Zhang 0001, Shiming Xiang, Chunhong Pan
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Learning Proposal-Aware Re-Ranking for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (WTAL) aims to localize and classify action instances in untrimmed videos with only video-level labels available. Despite the remarkable success of existing methods, whose generated proposals are commonly far more than the ground-truth action instances, it still makes sense to improve the ranking accuracy of the generated proposals since users in real-world scenarios usually prioritize the action proposals with the highest confidence scores. The inaccuracy of the proposal ranking mainly comes from two aspects: For one thing, the traditional proposal generation manner entirely relies on snippet-level perception, resulting in a significant yet unnoticed gap with the target of proposal-level localization. For another, existing methods commonly employ a hand-crafted proposal generation manner, a post-process that does not participate in model optimization. To address the above issues, we propose an end-to-end trained two-stage method, termed as Learning Proposal-aware Re-ranking (LPR) for WTAL. In the first stage, we design a proposal-aware feature learning module to inject the proposal-aware contextual information into each snippet, and then the enhanced features are utilized for predicting initial proposals. Furthermore, to perform effective and efficient proposal re-ranking, in the second stage, we contrast the proposals attached with high confidence scores with our constructed multi-scale foreground/background prototypes for further optimization. Evaluated by both the vanilla and Top-$k$mAP metrics, results of extensive experiments on two popular benchmarks demonstrate the effectiveness of our proposed method.
Yufan Hu, Jie Fu 0004, Junyu Gao 0002, Jianfeng Dong, Bin Fan 0001, Hongmin Liu 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 DomainFeat: Learning Local Features With Domain Adaptation
abstract
Accurate and efficient keypoint detection and description is a fundamental step in various computer vision tasks. In this paper, we extract robust descriptors and detect accurate keypoints by learning local Features with Domain adaptation (DomainFeat). Specifically, our Domainfeat includes image-level domain invariance supervision, pixel-level domain consistency supervision, Pixel-Adaptive keypoint Detection(PA-Det), and cross-domain dataset with domain stable point supervision. First, we introduce the image-level domain invariance supervision to make the high-level feature distributions from different domains close by fusing domain-invariant representations in the decoder. Furthermore, to compensate for the inconsistency between descriptors corresponding to the keypoints at the pixel level, we propose the pixel-level domain consistency supervision. Then we present the Pixel-Adaptive keypoint Detection to efficiently detect accurate keypoints, which can improve accuracy by enhancing the local consistency of heatmaps. Finally, we propose an efficient approach to construct data and supervision labels in diverse domains, which can tackle complex application scenarios. With these novel modules and supervision methods, our DomainFeat can make feature detectors more accurate and descriptors more robust. Extensive experiments confirm that Domainfeat achieves state-of-the-art performance on benchmarks such as Aachen-Day-Night localization, HPatches image matching, and the challenging DNIM dataset.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Pro2Diff: Proposal Propagation for Multi-Object Tracking via the Diffusion Model
abstract
Multi-object tracking (MOT) aims to estimate the bounding boxes and ID labels of objects in videos. The challenging issue in this task is to alleviate competitive learning between the detection and tracking subtasks, for which, two-stage Tracking-By-Detection (TBD) optimizes the two subtasks individually, and the single-stage Joint Detection and Tracking (JDT) adjusts the complex network architectures finely in an end-to-end pipeline. In this paper, we propose a new MOT method, i.e., Proposal Propagation via Diffusion Models, called Pro2Diff, which integrates a diffusion model into the proposal propagation in multi-object tracking, focusing on the model training process rather than complex network design. Specifically, using a generative approach, Pro2Diff generates a considerable number of noisy proposals for the tracking image sequence in the forward process, and subsequently, Pro2Diff learns the discrepancies between these noisy proposals and the actual bounding boxes of the tracked objects, gradually optimizing these noisy proposals to obtain the initial sequence of real tracked objects. By introducing the denoising diffusion process into multi-object tracking, we have made three further important findings: 1) Generative methods can effectively handle multi-object tracking tasks; 2) Without the need to modify the model structure, we propose self-conditional proposal propagation to enhance model performance effectively during inference; 3) By adjusting the numbers of proposals and iterations appropriately for different tracking sequences, the optimal performance of the model can be achieved. Extensive experimental results on MOT17 and DanceTrack datasets demonstrate that Pro2Diff outperforms current end-to-end multi-object tracking methods. We achieve 61.9 HOTA on DanceTrack and 57.6 HOTA on MOT17, reaching the competitive result of the JDT approach.
Hongmin Liu 0001, Canbin Zhang, Bin Fan 0001, Jinglin Xu
IEEE Trans. Image Process.3
2024 Exploring Rich Semantics for Open-Set Action Recognition
abstract
Open-set action recognition (OSAR) aims to learn a recognition framework capable of both classifying known classes and identifying unknown actions in open-set scenarios. Existing OSAR methods typically reside in a data-driven paradigm, which ignore the rich semantics in both known and unknown categories. In fact, we humans have the capability of leveraging the captured semantic information, i.e., knowledge and experience, to incisively distinguish samples from known and unknown classes. Motivated by this observation, in this paper, we propose a Unified Semantic Exploration (USE) framework for recognizing actions in open-set scenarios. Specifically, we explore the explicit knowledge semantics by simulating the unknown classes with knowledge-guided virtual classes based on an external knowledge graph, which enables the model to simulate open-set perception during model training. Besides, we propose to learn the implicit data semantics by transferring the knowledge structure of action categories to the visual prototype space for semantic structure preservation. Extensive experiments on several action recognition benchmarks validate the effectiveness of our proposed method.
Yufan Hu, Junyu Gao 0002, Jianfeng Dong, Bin Fan 0001, Hongmin Liu 0001
IEEE Trans. Multim.4
2024 Image Enhancement Guided Object Detection in Visually Degraded Scenes
abstract
Object detection accuracy degrades seriously in visually degraded scenes. A natural solution is to first enhance the degraded image and then perform object detection. However, it is suboptimal and does not necessarily lead to the improvement of object detection due to the separation of the image enhancement and object detection tasks. To solve this problem, we propose an image enhancement guided object detection method, which refines the detection network with an additional enhancement branch in an end-to-end way. Specifically, the enhancement branch and detection branch are organized in a parallel way, and a feature guided module is designed to connect the two branches, which optimizes the shallow feature of the input image in the detection branch to be as consistent as possible with that of the enhanced image. As the enhancement branch is frozen during training, such a design plays a role in using the features of enhanced images to guide the learning of object detection branch, so as to make the learned detection branch being aware of both image quality and object detection. When testing, the enhancement branch and feature guided module are removed, and so no additional computation cost is introduced for detection. Extensive experimental results, on underwater, hazy, and low-light object detection datasets, demonstrate that the proposed method can improve the detection performance of popular detection networks (YOLO v3, Faster R-CNN, DetectoRS) significantly in visually degraded scenes.
Hongmin Liu 0001, Hui Zeng 0003, Huayan Pu, Bin Fan 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Building Vision Transformers with Hierarchy Aware Feature Aggregation
abstract
Thanks to the excellent global modeling capability of attention mechanisms, the Vision Transformer has achieved better results than ConvNet in many computer tasks. However, in generating hierarchical feature maps, the Transformer still adopts the ConvNet feature aggregation scheme. This leads to the problem that the semantic information of the grid area of image becomes confused after feature aggregation, making it difficult for attention to accurately model global relationships. To address this, we propose the Hierarchy Aware Feature Aggregation framework (HAFA). HAFA enhances the extraction of local features adaptively in shallow layers where semantic information is weak, while is able to aggregate patches with similar semantics in deep layers. The clear semantic information of the aggregated patches, enables the attention mechanism to more accurately model global information at the semantic level. Extensive experiments show that after using the HAFA framework, significant improvements have been achieved relative to the baseline models in image classification, object detection, and semantic segmentation tasks.
Yongjie Chen, Hongmin Liu 0001, Bin Fan 0001
ICCV4
2023 Attention Weighted Local Descriptors
abstract
Local features detection and description are widely used in many vision applications with high industrial and commercial demands. With large-scale applications, these tasks raise high expectations for both the accuracy and speed of local features. Most existing studies on local features learning focus on the local descriptions of individual keypoints, which neglect their relationships established from global spatial awareness. In this paper, we present AWDesc with a consistent attention mechanism (CoAM) that opens up the possibility for local descriptors to embrace image-level spatial awareness in both the training and matching stages. For local features detection, we adopt local features detection with feature pyramid to obtain more stable and accurate keypoints localization. For local features description, we provide two versions of AWDesc to cope with different accuracy and speed requirements. On the one hand, we introduce Context Augmentation to address the inherent locality of convolutional neural networks by injecting non-local context information, so that local descriptors can "look wider to describe better". Specifically, well-designed Adaptive Global Context Augmented Module (AGCA) and Diverse Surrounding Context Augmented Module (DSCA) are proposed to construct robust local descriptors with context information from global to surrounding. On the other hand, we design an extremely lightweight backbone network coupled with the proposed special knowledge distillation strategy to achieve the best trade-off in accuracy and speed. What is more, we perform thorough experiments on image matching, homography estimation, visual localization, and 3D reconstruction tasks, and the results demonstrate that our method surpasses the current state-of-the-art local descriptors. Code is available at: https://github.com/vignywang/AWDesc.
Changwei Wang 0001, Rongtao Xu, Ke Lu 0002, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 Seeing Through Darkness: Visual Localization at Night via Weakly Supervised Learning of Domain Invariant Features
abstract
Long term visual localization has to conquer the problem of matching images with dramatic photometric changes caused by different seasons, natural and man-made illumination changes, etc. Visual localization at night plays a vital role in many applications like autonomous driving and augmented reality, for which extracting keypoints and descriptors with robustness to day-night illumination changes has became the bottleneck. This paper proposes an adversarial learning based solution to harvest from the weakly domain labels of day and night images, along with the point level correspondences among day time images, to achieve robust local feature extraction and description across day-night images. The key idea is to learn a discriminator to distinguish whether a feature map is generated from the day or night images, and simultaneously to adjust the parameters of feature extraction network so as to fool the discriminator. After adversarial training of the discriminator and feature extraction network, the feature extraction network finally reaches a stable status so that the extracted feature maps are robust to day-night photometric changes, based on which day-night domain invariant keypoints and descriptors can be extracted. Compared to existing local feature learning methods, it only requires an additional set of easily captured night images to improve the domain invariance of learned features. Experiments on two challenging benchmarks show the effectiveness of proposed method. In addition, this paper revisits the widely used image matching metrics on HPatches and finds that recall of different methods is highly related to their relative localization performance.
Bin Fan 0001, Yuzhu Yang, Wensen Feng, Fuchao Wu, Jiwen Lu, Hongmin Liu 0001
IEEE Trans. Multim.1
2022 MTLDesc: Looking Wider to Describe Better
abstract
Limited by the locality of convolutional neural networks, most existing local features description methods only learn local descriptors with local information and lack awareness of global and surrounding spatial context. In this work, we focus on making local descriptors ``look wider to describe better'' by learning local Descriptors with More Than Local information (MTLDesc). Specifically, we resort to context augmentation and spatial attention mechanism to make the descriptors obtain non-local awareness. First, Adaptive Global Context Augmented Module and Diverse Local Context Augmented Module are proposed to construct robust local descriptors with context information from global to local. Second, we propose the Consistent Attention Weighted Triplet Loss to leverage spatial attention awareness in both optimization and matching of local descriptors. Third, Local Features Detection with Feature Pyramid is proposed to obtain more stable and accurate keypoints localization. With the above innovations, the performance of the proposed MTLDesc significantly surpasses the current state-of-the-art local descriptors on HPatches, Aachen Day-Night localization and InLoc indoor localization benchmarks. Our code is available at https://github.com/vignywang/MTLDesc.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
AAAI6
2022 Continual Stereo Matching of Continuous Driving Scenes with Growing Architecture
abstract
The deep stereo models have achieved state-of-the-art performance on driving scenes, but they suffer from severe performance degradation when tested on unseen scenes. Although recent work has narrowed this performance gap through continuous online adaptation, this setup requires continuous gradient updates at inference and can hardly deal with rapidly changing scenes. To address these challenges, we propose to perform continual stereo matching where a model is tasked to 1) continually learn new scenes, 2) overcome forgetting previously learned scenes, and 3) continuously predict disparities at deployment. We achieve this goal by introducing a Reusable Architecture Growth (RAG) framework. RAG leverages task-specific neural unit search and architecture growth for continual learning of new scenes. During growth, it can maintain high reusability by reusing previous neural units while achieving good performance. A module named Scene Router is further introduced to adaptively select the scene-specific architecture path at inference. Experimental results demonstrate that our method achieves compelling performance in various types of challenging driving scenes.
Chenghao Zhang 0003, Bin Fan 0001, Gaofeng Meng, Zhaoxiang Zhang 0001, Chunhong Pan
CVPR3
2022 Out-of-distribution Detection with Boundary Aware Learning
Sen Pei, Xin Zhang 0093, Bin Fan 0001, Gaofeng Meng
ECCV (24)3
2022 Stereo Depth Estimation with Echoes
Chenghao Zhang 0003, Bolin Ni, Gaofeng Meng, Bin Fan 0001, Zhaoxiang Zhang 0001, Chunhong Pan
ECCV (27)5
2022 DOMAINDESC: Learning Local Descriptors With Domain Adaptation
abstract
Robust and efficient local descriptor is crucial in a wide range of applications. In this paper, we propose a novel descriptor DomainDesc which is invariant as much as possible by learning local Descriptor with Domain adaptation. We design the feature-level domain adaptation loss to improve robustness of our DomainDesc by punishing inconsistent high-level feature distributions of different images, while we present the pixel-level cross-domain consistency loss to compensate for the inconsistency between the descriptors corresponding to the keypoints at the pixel level. Besides, we adopt a new architecture to make the descriptor contain as much information as possible, and combine triplet loss and cross-domain consistency loss for descriptor supervision to ensure the distinguished ability of our descriptor. Finally, we give a cross-domain dataset generation strategy to quickly construct our training dataset for diverse domains to adapt to complex application scenarios. Experiments validate that our DomainDesc achieves state-of-the-art performances on HPatches image matching benchmark and Aachen-Day-Night localization benchmark.
Rongtao Xu, Changwei Wang 0001, Bin Fan 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICASSP3
2022 CARR-Net: Leveraging on Subtle Variance of Neighbors for Point Cloud Semantic Segmentation
Mingming Song, Bin Fan 0001, Hongmin Liu 0001
PRCV (3)2
2022 Robust self-supervised monocular visual odometry based on prediction-update pose estimation network
Haixin Xiu, Yiyou Liang, Hui Zeng 0003, Qing Li 0015, Hongmin Liu 0001, Bin Fan 0001
Eng. Appl. Artif. Intell.6
2022 CMT: Cross Mean Teacher Unsupervised Domain Adaptation for VHR Image Semantic Segmentation
abstract
Semantic segmentation of remote sensing images has achieved superior results with the supervised deep learning models. However, their performance to unseen data domains could be very bad due to the domain shift between different domains. Recently, a series of unsupervised domain adaptation (UDA) methods has been developed to solve the domain shift problem in semantic segmentation. Most of them use adversarial learning to achieve global cross-domain alignment and use a self-training (ST) strategy to generate pseudo-labels for classwise alignment. However, these methods ignore the pixels that are not assigned pseudo-labels. Those pixels are mostly at the boundaries, which are vital to the final segmentation results. To solve this problem, this letter proposes a cross mean teacher (CMT) UDA method. The whole framework consists of two parts. On the one hand, the global cross-domain distribution alignment is performed, and then, reliable pseudo-labels are assigned to the target data. On the other hand, a cross teacher–student network (CTSN) is developed to effectively use those pixels with and without pseudo-labels. This network contains two student networks ($S_{1}$and$S_{2}$) and two teacher networks ($T_{1}$and$T_{2}$) for cross-consistency constraints that supervises$S_{2}$(or$S_{1}$) by the prediction results of$T_{1}$(or$T_{2}$). The cross supervision by CTSN is helpful to prevent performance bottlenecks caused by the high coupling of teacher–student network in existing methods. Extensive experiments on three different remote sensing adaptation scenes verify the effectiveness and superiority of the proposed method.
Bin Fan 0001, Shiming Xiang, Chunhong Pan
IEEE Geosci. Remote. Sens. Lett.2
2022 Deep attention aware feature learning for person re-Identification
Xiaolu Sun, Bin Fan 0001, Chu Tang, Hui Zeng 0003
Pattern Recognit.4
2022 Incremental Translation Averaging
abstract
Translation averaging is known to be more difficult than rotation averaging due to scale ambiguity, estimation sensitivity, and solution uncertainty. Existing approaches have exposed their limitations in terms of accuracy, robustness, simplicity, or efficiency. To tackle this tough problem, a simple yet effective translation averaging pipeline, termed as Incremental Translation Averaging (ITA), is proposed in this paper. It combines the advantages of high accuracy and robustness in incremental parameter estimation pipeline and the advantages of high simplicity and efficiency in global motion averaging approach. Unlike the traditional translation averaging methods which estimate all the absolute camera locations simultaneously and suffer from inaccuracy in parameter estimation and incompleteness in scene reconstruction, our ITA computes them novelly in an incremental way with higher accuracy and robustness. Thanks to the introduction of incremental parameter estimation thought into the translation averaging pipeline, 1) our ITA is robust to measurement outliers and accurate in parameter estimation; and 2) our ITA is simple and efficient because of its less dependency on complicated optimization, carefully-designed preprocessing, or additional information. Comprehensive evaluations on the 1DSfM dataset demonstrate the effectiveness of our ITA and its advantages over several state-of-the-art translation averaging approaches.
Xiang Gao 0009, Lingjie Zhu, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen
IEEE Trans. Circuits Syst. Video Technol.3
2022 Learning Semantic-Aware Local Features for Long Term Visual Localization
abstract
Extracting robust and discriminative local features from images plays a vital role for long term visual localization, whose challenges are mainly caused by the severe appearance differences between matching images due to the day-night illuminations, seasonal changes, and human activities. Existing solutions resort to jointly learning both keypoints and their descriptors in an end-to-end manner, leveraged on large number of annotations of point correspondence which are harvested from the structure from motion and depth estimation algorithms. While these methods show improved performance over non-deep methods or those two-stage deep methods, i.e., detection and then description, they are still struggled to conquer the problems encountered in long term visual localization. Since the intrinsic semantics are invariant to the local appearance changes, this paper proposes to learn semantic-aware local features in order to improve robustness of local feature matching for long term localization. Based on a state of the art CNN architecture for local feature learning, i.e., ASLFeat, this paper leverages on the semantic information from an off-the-shelf semantic segmentation network to learn semantic-aware feature maps. The learned correspondence-aware feature descriptors and semantic features are then merged to form the final feature descriptors, for which the improved feature matching ability has been observed in experiments. In addition, the learned semantics embedded in the features can be further used to filter out noisy keypoints, leading to additional accuracy improvement and faster matching speed. Experiments on two popular long term visual localization benchmarks (Aachen Day and Night v1.1, Robotcar Seasons) and one challenging indoor benchmark (InLoc) demonstrate encouraging improvements of the localization accuracy over its counterpart and other competitive methods.
Bin Fan 0001, Wensen Feng, Huayan Pu, Yuzhu Yang, Qingqun Kong, Fuchao Wu, Hongmin Liu 0001
IEEE Trans. Image Process.1
2021 Geometric attentional dynamic graph convolutional neural networks for point cloud analysis
Yiming Cui 0002, Xin Liu 0027, Hongmin Liu 0001, Jiyong Zhang 0001, Alina Zare, Bin Fan 0001
Neurocomputing6
2021 PrGCN: Probability prediction with graph convolutional network for person re-identification
Hongmin Liu 0001, Zhenzhen Xiao, Bin Fan 0001, Hui Zeng 0003, Guoquan Jiang
Neurocomputing3
2021 Deep Unsupervised Binary Descriptor Learning Through Locality Consistency and Self Distinctiveness
abstract
Deep learning has been successfully applied to learn local feature descriptors in recent years. However, most of existing methods are supervised methods relying on a large number of labeled training patches, which are also proposed for learning real valued descriptors. In this paper, we propose a novel unsupervised deep learning method for binary descriptor learning. The binary descriptors are much more compact and efficient than the real valued descriptors and unsupervised leaning is highly required in many applications due to its label-free characteristic as the annotations are sometimes expensive to obtain. The core idea of our method is to explore the locality consistency in the descriptor space as well as to distinguish different patches while maintaining the ability to match a patch with its geometric transformed ones. We also give a theorical analysis about the role of batch normalization in learning effective binary descriptors. Benefited from this analysis, there is no need to append two additional losses on minimizing the quantization error and maximizing the entropy to the final learning objective like previous works did, thus simplifying our network training. Experiments on four benchmarks demonstrate that the proposed method is able to learn binary descriptors significantly outperforming previous unsupervised binary descriptors, even superior to most supervised ones. Especially, it obtains 21.2% of improvement on the UBC Phototour dataset, and 19.8%, 26.7%, 26.0% of improvements for patch verification, matching, retrieval tasks respectively on the HPatches dataset compared to the previous best unsupervised method.
Bin Fan 0001, Hongmin Liu 0001, Hui Zeng 0003, Jiyong Zhang 0001, Xin Liu 0027, Junwei Han 0001
IEEE Trans. Multim.1
2020 AugFPN: Improving Multi-Scale Feature Learning for Object Detection
abstract
Current state-of-the-art detectors typically exploit feature pyramid to detect objects at different scales. Among them, FPN is one of the representative works that build a feature pyramid by multi-scale features summation. However, the design defects behind prevent the multi-scale features from being fully exploited. In this paper, we begin by first analyzing the design defects of feature pyramid in FPN, and then introduce a new feature pyramid architecture named AugFPN to address these problems. Specifically, AugFPN consists of three components: Consistent Supervision, Residual Feature Augmentation, and Soft RoI Selection. AugFPN narrows the semantic gaps between features of different scales before feature fusion through Consistent Supervision. In feature fusion, ratio-invariant context information is extracted by Residual Feature Augmentation to reduce the information loss of feature map at the highest pyramid level. Finally, Soft RoI Selection is employed to learn a better RoI feature adaptively after feature fusion. By replacing FPN with AugFPN in Faster R-CNN, our models achieve 2.3 and 1.6 points higher Average Precision (AP) when using ResNet50 and MobileNet-v2 as backbone respectively. Furthermore, AugFPN improves RetinaNet by 1.6 points AP and FCOS by 0.9 points AP when using ResNet50 as backbone. Codes are available on https://github.com/Gus-Guo/AugFPN.
Chaoxu Guo, Bin Fan 0001, Qian Zhang 0009, Shiming Xiang, Chunhong Pan
CVPR2
2020 PointSpherical: Deep Shape Context for Point Cloud Learning in Spherical Coordinates
abstract
We propose Spherical Hierarchical modeling of 3D point cloud. Inspired by Shape Context, we design a receptive field on each 3D point by placing a spherical coordinate on it. We sample points using the furthest point method and creating overlapping balls of points. We divide the space into radial, polar angular, and azimuthal angular bins on which we form a Spherical Hierarchy for each ball. We apply 1x1 CNN convolution on points to start the initial feature extraction. Repeated 3D CNN and max-pooling over the Spherical bins propagate contextual information until all the information is condensed in the center bin. Extensive experiments on five datasets strongly evidence that our method outperforms current models on various Point Cloud Learning tasks, including 2D/3D shape classification, 3D part segmentation, and 3D semantic segmentation.
Bin Fan 0001, Yongcheng Liu, Yirong Yang, Jianbo Shi, Chunhong Pan, Huiwen Xie
ICPR2
2020 Deep Space Probing for Point Cloud Analysis
abstract
3D points distribute in a continuous 3D space irregularly, thus directly adapting 2D image convolution to 3D points is not an easy job. Previous works often artificially divide the space into regular grids, yet it could be suboptimal to learn geometry. In this paper, we propose SPCNN, namely, Space Probing Convolutional Neural Network, which naturally generalizes image CNN to deal with point clouds. The key idea of SPCNN is learning to probe the 3D space in an adaptive manner. Specifically, we define a pool of learnable convolutional weights, and let each point in the local region learn to choose a suitable convolutional weight from the pool. This is achieved by constructing a geometry guided index-mapping function that implicitly establishes a correspondence between convolutional weights and some local regions in the neighborhood (Fig. 1). In this way, the index-mapping function learns to adaptively partition nearby space for local geometry pattern recognition. With this convolution as a basic operator, SPCNN, a hierarchical architecture can be developed for effective point cloud analysis. Extensive experiments on challenging benchmarks across three tasks demonstrate that SPCNN achieves the state-of-the-art or has competitive performance.
Yirong Yang, Bin Fan 0001, Yongcheng Liu, Jiyong Zhang 0001, Xin Liu 0027, Xinyu Cai, Shiming Xiang, Chunhong Pan
ICPR2
2020 Video object detection for autonomous driving: Motion-aid feature calibration
Dongfang Liu, Yiming Cui 0002, Victor Y. Chen, Jiyong Zhang 0001, Bin Fan 0001
Neurocomputing5
2020 Efficient nearest neighbor search in high dimensional hamming space
Bin Fan 0001, Qingqun Kong, Baoqian Zhang, Hongmin Liu 0001, Chunhong Pan, Jiwen Lu
Pattern Recognit.1
2020 Geometric rectification of document images using adversarial gated unwarping network
Xiyan Liu, Gaofeng Meng, Bin Fan 0001, Shiming Xiang, Chunhong Pan
Pattern Recognit.3
2020 Features Combined Binary Descriptor Based on Voted Ring-Sampling Pattern
abstract
Most existing binary descriptors are only based on intensity and simply compare averaged intensities of one sample pair to obtain binary output. To address this issue, we propose a novel ring-sampling pattern based binary descriptor encoding both intensity and gradient. For intensity coding, a ring-sampling pattern is presented to define a number of sample points and their neighboring points given an interest point. The intensity difference between two sample points is obtained by performing the comparisons of the intensities of their corresponding neighboring points directly. Further, a majority based voting strategy is employed to obtain compact representation (1 bit or 3 bits) based on all these intensity difference among neighboring points. As for gradient coding, the gradient orientation histogram is computed for each sample point, and the gradient difference of two sample points is obtained by comparing the gradient magnitude on each orientation bin. The raw binary descriptor is constructed by concentrating the intensity and gradient differences of all the sample pairs, and two feature selection strategies are proposed to obtain the final compact descriptor, named as Features Combined Binary Descriptor based on Voted Ring-Sampling Pattern (BDVRP). The experimental results on the tasks of object recognition and image matching demonstrate the superiority and the effectiveness of the proposed descriptor, with comparison to the state-of-the-art hand-crafted binary descriptors.
Hongmin Liu 0001, Bin Fan 0001, Zhiheng Wang 0001, Junwei Han 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Triplet Adversarial Domain Adaptation for Pixel-Level Classification of VHR Remote Sensing Images
abstract
Pixel-level classification for very high resolution (VHR) images is a crucial but challenging task in remote sensing. However, since the diverse ways of satellite image acquisition and the distinct structures of various regions, the distributions of the same semantic classes among different data sets are dissimilar. Therefore, the classification model trained on one data set (source domain) may collapse, when it is directly applied to another one (target domain). To solve this problem, many adversarial-based domain adaptation methods have been proposed. However, these methods only consider the source and the target domains independently in the adversarial training, where only the target domain is explicitly contributed to narrow the gap between the distributions of both domains. Unlike previous methods, we propose a triplet adversarial domain adaptation (TriADA) method that jointly considers both domains to learn a domain-invariant classifier by a novel domain similarity discriminator. Specifically, the discriminator takes a triplet of segmentation maps as input, where two segmentation maps from the same domain are to be distinguished from the two maps from the different domains during the adversarial learning. Consequently, it explicitly considers both domains' information to narrow the distribution gap across domains. To enhance the discriminability of the classifier on the target domain, a class-aware self-training strategy, which depends on the output of the discriminator, is proposed to assign pseudo-labels with high adapted confidence on target data to retrain the classifier. Extensive experiments on several VHR pixel-level classification benchmarks demonstrate the effectiveness of our method as well as its superiority to the-state of the art.
Bin Fan 0001, Hongmin Liu 0001, Chunlei Huo, Shiming Xiang, Chunhong Pan
IEEE Trans. Geosci. Remote. Sens.2
2019 Relation-Shape Convolutional Neural Network for Point Cloud Analysis
abstract
Point cloud analysis is very challenging, as the shape implied in irregular points is difficult to capture. In this paper, we propose RS-CNN, namely, Relation-Shape Convolutional Neural Network, which extends regular grid CNN to irregular configuration for point cloud analysis. The key to RS-CNN is learning from relation, i.e., the geometric topology constraint among points. Specifically, the convolutional weight for local point set is forced to learn a high-level relation expression from predefined geometric priors, between a sampled point from this point set and the others. In this way, an inductive local representation with explicit reasoning about the spatial layout of points can be obtained, which leads to much shape awareness and robustness. With this convolution as a basic operator, RS-CNN, a hierarchical architecture can be developed to achieve contextual shape-aware learning for point cloud analysis. Extensive experiments on challenging benchmarks across three tasks verify RS-CNN achieves the state of the arts.
Yongcheng Liu, Bin Fan 0001, Shiming Xiang, Chunhong Pan
CVPR2
2019 SOSNet: Second Order Similarity Regularization for Local Descriptor Learning
abstract
Despite the fact that Second Order Similarity (SOS) has been used with significant success in tasks such as graph matching and clustering, it has not been exploited for learning local descriptors. In this work, we explore the potential of \sos in the field of descriptor learning by building upon the intuition that a positive pair of matching points should exhibit similar distances with respect to other points in the embedding space. Thus, we propose a novel regularization term, named Second Order Similarity Regularization (SOSR), that follows this principle. By incorporating SOSR into training, our learned descriptor achieves state-of-the-art performance on several challenging benchmarks containing distinct tasks ranging from local patch retrieval to structure from motion. Furthermore, by designing a von Mises-Fischer distribution based evaluation method, we link the utilization of the descriptor space to the matching performance, thus demonstrating the effectiveness of our proposed SOSR. Extensive experimental results, empirical evidence, and in-depth analysis are provided, indicating that SOSR can significantly boost the matching performance of the learned descriptor.
Yurun Tian, Xin Yu 0002, Bin Fan 0001, Fuchao Wu, Huub Heijnen, Vassileios Balntas
CVPR3
2019 Guiding the Flowing of Semantics: Interpretable Video Captioning via POS Tag
abstract
Xinyu Xiao, Lingfeng Wang, Bin Fan, Shinming Xiang, Chunhong Pan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xinyu Xiao, Lingfeng Wang 0002, Bin Fan 0001, Shiming Xiang, Chunhong Pan
EMNLP/IJCNLP (1)3
2019 Progressive Sparse Local Attention for Video Object Detection
abstract
Transferring image-based object detectors to the domain of videos remains a challenging problem. Previous efforts mostly exploit optical flow to propagate features across frames, aiming to achieve a good trade-off between accuracy and efficiency. However, introducing an extra model to estimate optical flow can significantly increase the overall model size. The gap between optical flow and high-level features can also hinder it from establishing spatial correspondence accurately. Instead of relying on optical flow, this paper proposes a novel module called Progressive Sparse Local Attention (PSLA), which establishes the spatial correspondence between features across frames in a local region with progressively sparser stride and uses the correspondence to propagate features. Based on PSLA, Recursive Feature Updating (RFU) and Dense Feature Transforming (DenseFT) are proposed to model temporal appearance and enrich feature representation respectively in a novel video object detection framework. Experiments on ImageNet VID show that our method achieves the best accuracy compared to existing methods with smaller model size and acceptable runtime speed.
Chaoxu Guo, Bin Fan 0001, Jie Gu 0002, Qian Zhang 0009, Shiming Xiang, Véronique Prinet, Chunhong Pan
ICCV2
2019 DensePoint: Learning Densely Contextual Representation for Efficient Point Cloud Processing
abstract
Point cloud processing is very challenging, as the diverse shapes formed by irregular points are often indistinguishable. A thorough grasp of the elusive shape requires sufficiently contextual semantic information, yet few works devote to this. Here we propose DensePoint, a general architecture to learn densely contextual representation for point cloud processing. Technically, it extends regular grid CNN to irregular point configuration by generalizing a convolution operator, which holds the permutation invariance of points, and achieves efficient inductive learning of local patterns. Architecturally, it finds inspiration from dense connection mode, to repeatedly aggregate multi-level and multi-scale semantics in a deep hierarchy. As a result, densely contextual information along with rich semantics, can be acquired by DensePoint in an organic manner, making it highly effective. Extensive experiments on challenging benchmarks across four tasks, as well as thorough model analysis, verify DensePoint achieves the state of the arts.
Yongcheng Liu, Bin Fan 0001, Gaofeng Meng, Jiwen Lu, Shiming Xiang, Chunhong Pan
ICCV2
2019 A Performance Evaluation of Local Features for Image-Based 3D Reconstruction
abstract
This paper performs a comprehensive and comparative evaluation of the state-of-the-art local features for the task of image-based 3D reconstruction. The evaluated local features cover the recently developed ones by using powerful machine learning techniques and the elaborately designed handcrafted features. To obtain a comprehensive evaluation, we choose to include both float type features and binary ones. Meanwhile, two kinds of datasets have been used in this evaluation. One is a dataset of many different scene types with groundtruth 3D points, containing images of different scenes captured at fixed positions, for quantitative performance evaluation of different local features in the controlled image capturing situation. The other dataset contains Internet scale image sets of several landmarks with a lot of unrelated images, which is used for qualitative performance evaluation of different local features in the free image collection situation. Our experimental results show that binary features are competent to reconstruct scenes from controlled image sequences with only a fraction of processing time compared to using float type features. However, for the case of a large scale image set with many distracting images, float type features show a clear advantage over binary ones. Currently, the most traditional SIFT is very stable with regard to scene types in this specific task and produces very competitive reconstruction results among all the evaluated local features. Meanwhile, although the learned binary features are not as competitive as the handcrafted ones, learning float type features with CNN is promising but still requires much effort in the future.
Bin Fan 0001, Qingqun Kong, Xinchao Wang, Zhiheng Wang 0001, Shiming Xiang, Chunhong Pan, Pascal Fua
IEEE Trans. Image Process.1
2018 Adversarial Domain Adaptation with a Domain Similarity Discriminator for Semantic Segmentation of Urban Areas
abstract
Existing semantic segmentation models of urban areas have shown to perform well in a supervised setting. However, collecting lots of annotated images from each city to train such models is time-consuming or difficult. In addition, when transferring the segmentation model from the trained city (source domain) to an unseen city (target domain), the performance will largely degrade due to the domain shift. For this reason, we propose a domain adaptation method with a domain similarity discriminator to eliminate such domain shift in the framework of adversarial learning. Contrary to the single-input adversarial network, our domain similarity discriminator, which consists of a Siamese network, is able to measure the similarity of the pairwise-input data. In this way, we can use more information about the pairwise-input to measure the similarity between different distributions so as to address the problem of domain shift. Experimental results demonstrate that our approach outperforms the competing methods on three different cities.
Bin Fan 0001, Shiming Xiang, Chunhong Pan
ICIP2
2018 Cayley- Klein Metric Learning with Shrinkage-Expansion Constraints
abstract
Cayley-Klein metric is a specific kind of non-Euclidean metric in projective space. Recently, it has been introduced into metric learning with encouraging performance when dealing with computer vision tasks. However, the original Cayley-Klein metric learning methods with conventional pairwise and triplet-wise constraints, which are fixed bound based constraints, may not perform well when the intra-and inter-class variations of data distribution become complex. Pairwise constraints restrict the distance between samples of a similar pair to be lower than a fixed upper bound, and the distance between samples of a dissimilar pair higher than a fixed lower bound. Triplet-wise constraints restrict the distance between samples of a similar pair to be smaller than that between a pair of samples from different classes. In this paper, we propose a novel Cayley-Klein metric learning method (CKseML) with adaptive shrinkage-expansion pairwise constraints. CKseML is very effective in learning metric from data with complex distributions. Our experimental results demonstrate that CKseML achieves better performance than the original Cayley-Klein metric learning methods.
Yanhong Bi, Bin Fan 0001, Fuchao Wu
ICPR2
2018 Traffic Sign Recognition Using a Multi-Task Convolutional Neural Network
abstract
Although traffic sign recognition has been studied for many years, most existing works are focused on the symbol-based traffic signs. This paper proposes a new data-driven system to recognize all categories of traffic signs, which include both symbol-based and text-based signs, in video sequences captured by a camera mounted on a car. The system consists of three stages, traffic sign regions of interest (ROIs) extraction, ROIs refinement and classification, and post-processing. Traffic sign ROIs from each frame are first extracted using maximally stable extremal regions on gray and normalized RGB channels. Then, they are refined and assigned to their detailed classes via the proposed multi-task convolutional neural network, which is trained with a large amount of data, including synthetic traffic signs and images labeled from street views. The post-processing finally combines the results in all frames to make a recognition decision. Experimental results have demonstrated the effectiveness of the proposed system.
Hengliang Luo, Bei Tong, Fuchao Wu, Bin Fan 0001
IEEE Trans. Intell. Transp. Syst.5
2018 In Defense of Locality-Sensitive Hashing
abstract
Hashing-based semantic similarity search is becoming increasingly important for building large-scale content-based retrieval system. The state-of-the-art supervised hashing techniques use flexible two-step strategy to learn hash functions. The first step learns binary codes for training data by solving binary optimization problems with millions of variables, thus usually requiring intensive computations. Despite simplicity and efficiency, locality-sensitive hashing (LSH) has never been recognized as a good way to generate such codes due to its poor performance in traditional approximate neighbor search. We claim in this paper that the true merit of LSH lies in transforming the semantic labels to obtain the binary codes, resulting in an effective and efficient two-step hashing framework. Specifically, we developed the locality-sensitive two-step hashing (LS-TSH) that generates the binary codes through LSH rather than any complex optimization technique. Theoretically, with proper assumption, LS-TSH is actually a useful LSH scheme, so that it preserves the label-based semantic similarity and possesses sublinear query complexity for hash lookup. Experimentally, LS-TSH could obtain comparable retrieval accuracy with state of the arts with two to three orders of magnitudes faster training speed.
Kun Ding 0001, Chunlei Huo, Bin Fan 0001, Shiming Xiang, Chunhong Pan
IEEE Trans. Neural Networks Learn. Syst.3
2017 L2-Net: Deep Learning of Discriminative Patch Descriptor in Euclidean Space
abstract
The research focus of designing local patch descriptors has gradually shifted from handcrafted ones (e.g., SIFT) to learned ones. In this paper, we propose to learn high performance descriptor in Euclidean space via the Convolutional Neural Network (CNN). Our method is distinctive in four aspects: (i) We propose a progressive sampling strategy which enables the network to access billions of training samples in a few epochs. (ii) Derived from the basic concept of local patch matching problem, we empha-size the relative distance between descriptors. (iii) Extra supervision is imposed on the intermediate feature maps. (iv) Compactness of the descriptor is taken into account. The proposed network is named as L2-Net since the output descriptor can be matched in Euclidean space by L2 distance. L2-Net achieves state-of-the-art performance on the Brown datasets [16], Oxford dataset [18] and the newly proposed Hpatches dataset [11]. The good generalization ability shown by experiments indicates that L2-Net can serve as a direct substitution of the existing handcrafted descriptors. The pre-trained L2-Net is publicly available.
Yurun Tian, Bin Fan 0001, Fuchao Wu
CVPR2
2017 Context-aware cascade network for semantic labeling in VHR image
abstract
Semantic labeling for the very high resolution (VHR) image of urban areas is challenging, because of many complex manmade objects with different materials and fine-structured objects located together. Under the framework of convolutional neural networks (CNNs), this paper proposes a novel end-to-end network for semantic labeling. Specifically, our network not only improves the labeling accuracy of complex manmade objects by aggregating multiple context semantics with a cascaded architecture, but also refines fine-structured objects by utilizing the low-level detail in shallow layers of CNNs with a hierarchical pyramid structure. Throughout the network, a dedicated residual correction scheme is employed to amend the latent fitting residual. As a result of these specific components, the whole model works in a global-to-local and coarse-to-fine manner. Experimental results show that our network outperforms the state-of-the-art methods on the large-scale ISPRS Vaihingen 2D Semantic Labeling Challenge dataset.
Yongcheng Liu, Bin Fan 0001, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan
ICIP2
2017 Structured binary feature extraction for hyperspectral imagery classification
abstract
In this paper, we propose a novel structured binary feature extraction method for hyperspectral image classification. To pursuit high discriminative ability and low memory cost, we resort to applying the learning to hash technique to the traditional spectral-spatial hyperspectral features. We show how the structured information among different kinds of features and different feature groups can be used to learn discriminative binary features for classification. Experiments on two standard benchmark hyperspectral data sets demonstrate the effectiveness of the proposed method.
Zisha Zhong, Bin Fan 0001, Shiming Xiang, Chunhong Pan
ICIP2
2017 Ordinal pyramid coding for rotation invariant feature extraction
Guoli Wang 0004, Bin Fan 0001, Chunhong Pan
Neurocomputing2
2017 Feature Extraction by Rotation-Invariant Matrix Representation for Object Detection in Aerial Image
abstract
This letter proposes a novel rotation-invariant feature for object detection in optical remote sensing images. Different from previous rotation-invariant features, the proposed rotation-invariant matrix (RIM) can incorporate partial angular spatial information in addition to radial spatial information. Moreover, it can be further calculated between different rings for a redundant representation of the spatial layout. Based on the RIM, we further propose an RIM_FV_RPP feature for object detection. For an image region, we first densely extract RIM features from overlapping blocks; then, these RIM features are encoded into Fisher vectors; finally, a pyramid pooling strategy that hierarchically accumulates Fisher vectors in ring subregions is used to encode richer spatial information while maintaining rotation invariance. Both of the RIM and RIM_FV_RPP are rotation invariant. Experiments on airplane and car detection in optical remote sensing images demonstrate the superiority of our feature to the state of the art.
Guoli Wang 0004, Xinchao Wang, Bin Fan 0001, Chunhong Pan
IEEE Geosci. Remote. Sens. Lett.3
2017 Greedy Batch-Based Minimum-Cost Flows for Tracking Multiple Objects
abstract
Minimum-cost flow algorithms have recently achieved state-of-the-art results in multi-object tracking. However, they rely on the whole image sequence as input. When deployed in real-time applications or in distributed settings, these algorithms first operate on short batches of frames and then stitch the results into full trajectories. This decoupled strategy is prone to errors because the batch-based tracking errors may propagate to the final trajectories and cannot be corrected by other batches. In this paper, we propose a greedy batch-based minimum-cost flow approach for tracking multiple objects. Unlike existing approaches that conduct batch-based tracking and stitching sequentially, we optimize consecutive batches jointly so that the tracking results on one batch may benefit the results on the other. Specifically, we apply a generalized minimum-cost flows (MCF) algorithm on each batch and generate a set of conflicting trajectories. These trajectories comprise the ones with high probabilities, but also those with low probabilities potentially missed by detectors and trackers. We then apply the generalized MCF again to obtain the optimal matching between trajectories from consecutive batches. Our proposed approach is simple, effective, and does not require training. We demonstrate the power of our approach on data sets of different scenarios.
Xinchao Wang, Bin Fan 0001, Shiyu Chang, Zhangyang Wang, Xianming Liu 0005, Dacheng Tao, Thomas S. Huang
IEEE Trans. Image Process.2
2017 Cross-Modal Hashing via Rank-Order Preserving
abstract
Due to the query effectiveness and efficiency, cross-modal similarity search based on hashing has acquired extensive attention in the multimedia community. Most existing methods do not explicitly employ the ranking information when learning hash functions, which is quite important for building practical retrieval systems. To solve this issue, this paper proposes a rank-order preserving hashing (RoPH) method with a novel regression-based rank-order preserving loss that has provable large margin property and is easy to optimize. Moreover, we jointly learn the binary codes and hash functions instead of using any relaxation trick. To solve the induced optimization problem, the alternating descent technique is adopted and each subproblem can be solved conveniently. Specifically, we show that the involved binary quadratic programming subproblem with respect to an introduced auxiliary binary variable satisfies submodularity, enabling us to use the off-the-shelf graph-cut algorithms to solve it exactly and efficiently. Extensive experiments on three benchmarks demonstrate that RoPH significantly improves the ranking quality over the state of the arts.
Kun Ding 0001, Bin Fan 0001, Chunlei Huo, Shiming Xiang, Chunhong Pan
IEEE Trans. Multim.2
2016 Feature matching using guidance-constraint method
abstract
Image distortions and repetitive patterns widely exist in real images, which results in that feature matching is still a challenging problem though great progress has been made recently. This study presents a matching method, called guidance‐constraints method (GCM), which has obvious advantages in resolving problems of matching features on images with image distortions or repetitive patterns. In GCM, feature points are paired and connection compatibility is introduced to describe the relative geometric relations among features, and then potential matches are found by using the defined geometric guidance and are verified by using the defined geometric constraints. Experimental evaluation shows that the proposed GCM can significantly improve both the number of correct matches and correct ratio under various image transformations, especially more effective on images with distortions or containing repetitive patterns.
Zhiheng Wang 0001, Hongmin Liu 0001, Zhanqiang Huo, Bin Fan 0001
IET Comput. Vis.5
2016 Exploring Local and Overall Ordinal Information for Robust Feature Description
abstract
This paper aims to build robust feature descriptors by exploring intensity order information in a patch. To this end, the local intensity order pattern (LIOP) and the overall intensity order pattern (OIOP) are proposed to effectively encode intensity order information of each pixel in different aspects. Specifically, LIOP captures the local ordinal information by using the intensity relationships among all the neighbouring sampling points around a pixel, while OIOP exploits the coarsely quantized overall intensity order of these sampling points. These two kinds of patterns are then separately aggregated into different ordinal bins, leading to two kinds of feature descriptors. Furthermore, as these two kinds of descriptors could encode complementary ordinal information, they are combined together to obtain a discriminative and compact mixed intensity order pattern descriptor. All these descriptors are constructed on the basis of relative relationships of intensities in a rotationally invariant way, making them be inherently invariant to image rotation and any monotonic intensity changes. Experimental results on image matching and object recognition are encouraging, demonstrating the superiorities of our descriptors over the state of the art.
Zhenhua Wang 0002, Bin Fan 0001, Gang Wang 0012, Fuchao Wu
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Efficient Multiple Feature Fusion With Hashing for Hyperspectral Imagery Classification: A Comparative Study
abstract
Due to the complementary properties of different features, multiple feature fusion has a large potential for hyperspectral imagery classification. At the meantime, hashing is promising in representing a high-dimensional float-type feature with extremely low bit binary codes while maintaining the performance. In this paper, we study the possibility of using hashing to fuse multiple features for hyperspectral imagery classification. For this purpose, we propose a multiple feature fusion framework to evaluate the performance of using different hashing methods. For comparison and completeness, we also have an extensive comparison to five subspace-based dimension reduction methods and six fusion-based methods which are popular solutions to deal with multiple features in hyperspectral image classification. Experimental results on four benchmark hyperspectral data sets demonstrate that using hashing to fuse multiple features can achieve comparable or better performance with the traditional subspace-based dimension reduction methods and fusion-based methods. Moreover, the binary features obtained by using hashing need much less storage and are faster to compute distances with the help of machine instructions.
Zisha Zhong, Bin Fan 0001, Kun Ding 0001, Haichang Li, Shiming Xiang, Chunhong Pan
IEEE Trans. Geosci. Remote. Sens.2
2016 Layer-Wise Floorplan Extraction for Automatic Urban Building Reconstruction
abstract
Urban building reconstruction is an important step for urban digitization and realisticvisualization. In this paper, we propose a novel automatic method to recover urban building geometry from 3D point clouds. The proposed method is suitable for buildings composed of planar polygons and aligned with the gravity direction, which are quite common in the city. Our key observation is that the building shapes are usually piecewise constant along the gravity direction and determined by several dominant shapes. Based on this observation, we formulate building reconstruction as an energy minimization problem under the Markov Random Field (MRF) framework. Specifically, point clouds are first cutinto a sequence of slices along the gravity direction. Then, floorplans are reconstructed by extracting boundaries of these slices, among which dominant floorplans are extracted and propagated to other floors via MRF. To guarantee correct propagation, a new distance measurement for floorplans is designed, which first encodes floorplans into strings and then calculates distances between their corresponding strings. Additionally, an image based editing method is also proposed to recover detailed window structures. Experimental results on both synthetic and real data sets have validated the effectiveness of our method.
Wei Sui, Lingfeng Wang 0002, Bin Fan 0001, Hongfei Xiao, Huai-Yu Wu, Chunhong Pan
IEEE Trans. Vis. Comput. Graph.3
2015 10, 000+ Times Accelerated Robust Subset Selection
abstract
Subset selection from massive data with noised information is increasingly popular for various applications. This problem is still highly challenging as current methods are generally slow in speed and sensitive to outliers. To address the above two issues, we propose an accelerated robust subset selection (ARSS) method. Extensive experiments on ten benchmark datasets verify that our method not only outperforms state of the art methods, but also runs 10,000+ times faster than the most related method.
Feiyun Zhu, Bin Fan 0001, Xinliang Zhu, Ying Wang 0008, Shiming Xiang, Chunhong Pan
AAAI2
2015 Beyond Mahalanobis metric: Cayley-Klein metric learning
abstract
Cayley-Klein metric is a kind of non-Euclidean metric suitable for projective space. In this paper, we introduce it into the computer vision community as a powerful metric and an alternative to the widely studied Mahalanobis metric. We show that besides its good characteristic in non-Euclidean space, it is a generalization of Mahalanobis metric in some specific cases. Furthermore, as many Mahalanobis metric learning, we give two kinds of Cayley-Klein metric learning methods: MMC Cayley-Klein metric learning and LMNN Cayley-Klein metric learning. Experiments have shown the superiority of Cayley-Klein metric over Mahalanobis ones and the effectiveness of our Cayley-Klein metric learning methods.
Yanhong Bi, Bin Fan 0001, Fuchao Wu
CVPR2
2015 Ordinal pyramid pooling for rotation invariant object recognition
abstract
Local feature descriptor plays a fundamental role in many visual tasks, and its rotation invariance is a key issue for many recognition and detection problems. This paper proposes a novel rotation invariant descriptor by ordinal pyramid pooling of local Fourier transform features based on their radial gradient orientations. Since both the low-level feature and pooling strategy are rotation invariant, the obtained descriptor is rotation invariant by nature. Pooling based on orders of gradient orientations is not only invariant to in-plane rotation, but also encodes gradient orientation information into descriptor as well as spatial information to some extent. Moreover, these information is enhanced by the proposed pyramid pooling structure. Therefore, our method is naturally rotation invariant and has strong discriminative ability. Experimental results on the aerial car dataset demonstrate the effectiveness of our descriptor.
Guoli Wang 0004, Bin Fan 0001, Chunhong Pan
ICASSP2
2015 kNN Hashing with Factorized Neighborhood Representation
abstract
Hashing is very effective for many tasks in reducing the processing time and in compressing massive databases. Although lots of approaches have been developed to learn data-dependent hash functions in recent years, how to learn hash functions to yield good performance with acceptable computational and memory cost is still a challenging problem. Based on the observation that retrieval precision is highly related to the kNN classification accuracy, this paper proposes a novel kNN-based supervised hashing method, which learns hash functions by directly maximizing the kNN accuracy of the Hamming-embedded training data. To make it scalable well to large problem, we propose a factorized neighborhood representation to parsimoniously model the neighborhood relationships inherent in training data. Considering that real-world data are often linearly inseparable, we further kernelize this basic model to improve its performance. As a result, the proposed method is able to learn accurate hashing functions with tolerable computation and storage cost. Experiments on four benchmarks demonstrate that our method outperforms the state-of-the-arts.
Kun Ding 0001, Chunlei Huo, Bin Fan 0001, Chunhong Pan
ICCV3
2015 Discriminant Tensor Spectral-Spatial Feature Extraction for Hyperspectral Image Classification
abstract
We propose to integrate spectral-spatial feature extraction and tensor discriminant analysis for hyperspectral image classification. First, we apply remarkable spectral-spatial feature extraction approaches in the hyperspectral cube to extract a feature tensor for each pixel. Then, based on class label information, local tensor discriminant analysis is used to remove redundant information for subsequent classification procedure. The approach not only extracts sufficient spectral-spatial features from original hyperspectral images but also gets better feature representation owing to tensor framework. Comparative results on two benchmarks demonstrate the effectiveness of our method.
Zisha Zhong, Bin Fan 0001, Jiangyong Duan, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan
IEEE Geosci. Remote. Sens. Lett.2
2014 Affine Subspace Representation for Feature Description
Zhenhua Wang 0002, Bin Fan 0001, Fuchao Wu
ECCV (7)2
2014 Receptive Fields Selection for Binary Feature Description
abstract
Feature description for local image patch is widely used in computer vision. While the conventional way to design local descriptor is based on expert experience and knowledge, learning-based methods for designing local descriptor become more and more popular because of their good performance and data-driven property. This paper proposes a novel data-driven method for designing binary feature descriptor, which we call receptive fields descriptor (RFD). Technically, RFD is constructed by thresholding responses of a set of receptive fields, which are selected from a large number of candidates according to their distinctiveness and correlations in a greedy way. Using two different kinds of receptive fields (namely rectangular pooling area and Gaussian pooling area) for selection, we obtain two binary descriptors RFDR and RFDG .accordingly. Image matching experiments on the well-known patch data set and Oxford data set demonstrate that RFD significantly outperforms the state-of-the-art binary descriptors, and is comparable with the best float-valued descriptors at a fraction of processing time. Finally, experiments on object recognition tasks confirm that both RFDR and RFDG successfully bridge the performance gap between binary descriptors and their floating-point competitors.
Bin Fan 0001, Qingqun Kong, Tomasz Trzcinski, Zhiheng Wang 0001, Chunhong Pan, Pascal Fua
IEEE Trans. Image Process.1
2014 Spectral Unmixing via Data-Guided Sparsity
abstract
Hyperspectral unmixing, the process of estimating a common set of spectral bases and their corresponding composite percentages at each pixel, is an important task for hyperspectral analysis, visualization, and understanding. From an unsupervised learning perspective, this problem is very challenging-both the spectral bases and their composite percentages are unknown, making the solution space too large. To reduce the solution space, priors. In practice, these priors would easily lead to some unsuitable solution. This is because they are achieved by applying an identical strength of constraints to all the factors, which does not hold in practice. To overcome this limitation, we propose a novel sparsity-based method by learning a data-guided map (DgMap) to describe the individual mixed level of each pixel. Through this DgMap, the l(p) (0 < p < 1) constraint is applied in an adaptive manner. Such implementation not only meets the practical situation, but also guides the spectral bases toward the pixels under highly sparse constraint. What is more, an elegant optimization scheme as well as its convergence proof have been provided in this paper. Extensive experiments on several datasets also demonstrate that the DgMap is feasible, and high quality unmixing results could be obtained by our method.
Feiyun Zhu, Ying Wang 0008, Bin Fan 0001, Shiming Xiang, Gaofeng Meng, Chunhong Pan
IEEE Trans. Image Process.3
2013 FRIF: Fast Robust Invariant Feature
abstract
Establishing robust visual correspondences is a fundamental component of many computer vision applications. However, it is very challenging to obtain high quality features while maintaining a low computational cost. This paper aims to tackle this problem by adopting a novel Fast Robust Invariant Feature (FRIF) for both feature detection and description. The basic idea is to employ a fast approximated LoG detector to select scale-invariant keypoints and incorporate local pattern and inter-pattern information to construct distinctive binary descriptors. A comprehensive evaluation on standard dataset shows that FRIF achieves quite a high performance with a computation time comparable to state-of-the-art real-time features.
Zhenhua Wang 0002, Bin Fan 0001, Fuchao Wu
BMVC2
2013 Learning weighted Hamming distance for binary descriptors
abstract
Local image descriptors are one of the key components in many computer vision applications. Recently, binary descriptors have received increasing interest of the community for its efficiency and low memory cost. The similarity of binary descriptors is measured by Hamming distance which has equal emphasis on all elements of binary descriptors. This paper improves the performance of binary descriptors by learning a weighted Hamming distance for binary descriptors with larger weights assigned to more discriminative elements. What is more, the weighted Hamming distance can be computed as fast as the Hamming distance on the basis of a pre-computed look-up-table. Therefore, the proposed method improves the matching performance of binary descriptors without sacrificing matching speed. Experimental results on two popular binary descriptors (BRIEF [1] and FREAK [2]) validate the effectiveness of the proposed method.
Bin Fan 0001, Qingqun Kong, Xiao-Tong Yuan, Zhiheng Wang 0001, Chunhong Pan
ICASSP1
2013 Registration of Optical and SAR Satellite Images by Exploring the Spatial Relationship of the Improved SIFT
abstract
Although feature-based methods have been successfully developed in the past decades for the registration of optical images, the registration of optical and synthetic aperture radar (SAR) images is still a challenging problem in remote sensing. In this letter, an improved version of the scale-invariant feature transform is first proposed to obtain initial matching features from optical and SAR images. Then, the initial matching features are refined by exploring their spatial relationship. The refined feature matches are finally used for estimating registration parameters. Experimental results have shown the effectiveness of the proposed method.
Bin Fan 0001, Chunlei Huo, Chunhong Pan, Qingqun Kong
IEEE Geosci. Remote. Sens. Lett.1
2012 Rotationally Invariant Descriptors Using Intensity Order Pooling
abstract
This paper proposes a novel method for interest region description which pools local features based on their intensity orders in multiple support regions. Pooling by intensity orders is not only invariant to rotation and monotonic intensity changes, but also encodes ordinal information into a descriptor. Two kinds of local features are used in this paper, one based on gradients and the other on intensities; hence, two descriptors are obtained: the Multisupport Region Order-Based Gradient Histogram (MROGH) and the Multisupport Region Rotation and Intensity Monotonic Invariant Descriptor (MRRID). Thanks to the intensity order pooling scheme, the two descriptors are rotation invariant without estimating a reference orientation, which appears to be a major error source for most of the existing methods, such as Scale Invariant Feature Transform (SIFT), SURF, and DAISY. Promising experimental results on image matching and object recognition demonstrate the effectiveness of the proposed descriptors compared to state-of-the-art descriptors.
Bin Fan 0001, Fuchao Wu, Zhanyi Hu
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Robust line matching through line-point invariants
Bin Fan 0001, Fuchao Wu, Zhanyi Hu
Pattern Recognit.1
2011 Aggregating gradient distributions into intensity orders: A novel local image descriptor
abstract
A novel local image descriptor is proposed in this paper, which combines intensity orders and gradient distributions in multiple support regions. The novelty lies in three aspects: 1) The gradient is calculated in a rotation invariant way in a given support region; 2) The rotation invariant gradients are adaptively pooled spatially based on intensity orders in order to encode spatial information; 3) Multiple support regions are used for constructing descriptor which further improves its discriminative ability. Therefore, the proposed descriptor encodes not only gradient information but also information about relative relationship of intensities as well as spatial information. In addition, it is truly rotation invariant in theory without the need of computing a dominant orientation which is a major error source of most existing methods, such as SIFT. Results on the standard Oxford dataset and 3D objects have shown a significant improvement over the state-of-the-art methods under various image transformations.
Bin Fan 0001, Fuchao Wu, Zhanyi Hu
CVPR1
2011 Local Intensity Order Pattern for feature description
abstract
This paper presents a novel method for feature description based on intensity order. Specifically, a Local Intensity Order Pattern(LIOP) is proposed to encode the local ordinal information of each pixel and the overall ordinal information is used to divide the local patch into subregions which are used for accumulating the LIOPs respectively. Therefore, both local and overall intensity ordinal information of the local patch are captured by the proposed LIOP descriptor so as to make it a highly discriminative descriptor. It is shown that the proposed descriptor is not only invariant to monotonic intensity changes and image rotation but also robust to many other geometric and photometric transformations such as viewpoint change, image blur and JEPG compression. The proposed descriptor has been evaluated on the standard Oxford dataset and four additional image pairs with complex illumination changes. The experimental results show that the proposed descriptor obtains a significant improvement over the existing state-of-the-art descriptors.
Zhenhua Wang 0002, Bin Fan 0001, Fuchao Wu
ICCV2
2011 Towards reliable matching of images containing repetitive patterns
Bin Fan 0001, Fuchao Wu, Zhanyi Hu
Pattern Recognit. Lett.1
2010 Line matching leveraged by point correspondences
abstract
A novel method for line matching is proposed. The basic idea is to use tentative point correspondences, which can be easily obtained by keypoint matching methods, to significantly improve line matching performance, even when the point correspondences are severely contaminated by outliers. When matching a pair of image lines, a group of corresponding points that may be coplanar with these lines in 3D space is firstly obtained from all corresponding image points in the local neighborhoods of these lines. Then given such a group of corresponding points, the similarity between this pair of lines is calculated based on an affine invariant from one line and two points. The similarity is defined on the basis of median statistic in order to handle the problem of inevitable incorrect correspondences in the group of point correspondences. Furthermore, the relationship of rotation between the reference and query images is estimated from all corresponding points to filter out those pairs of lines which are obviously impossible to be matches, hence speeding up the matching process as well as further improving its robustness. Extensive experiments on real images demonstrate the good performance of the proposed method as well as its superiority to the state-of-the-art methods.
Bin Fan 0001, Fuchao Wu, Zhanyi Hu
CVPR1