VLDB 2026 Research / reviewers in the wild / expert
Xiao Wang 0014
dblp:49/67-14
· DBLP profile ↗
78ranked-venue papers
22as first author
66since 2021 · last 2026
0000-0001-6117-6745ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 15 first-author · 39 since 2021Artificial intelligence and machine learning · 39 · 15 first-author · 36 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning ModelabstractLong Chain-of-Thought (CoT) reasoning enhances large reasoning models' performance but suffers from severe inefficiencies, as models often overthink simple problems or underthink complex ones. Current sequence-level optimizations, like length penalties, are too coarse-grained to distinguish core logic from verbose language, precluding the necessary token-level control for efficient reasoning CoT. To overcome these limitations, we introduce Time-Frequency token Advantage Clipping (TFAC), a novel training framework designed to build efficient large reasoning models via token-level interventions. Specifically, TFAC functions along two dimensions: 1) The Frequency Dimension: It discourages inefficient loops and encourages deeper exploration by dynamically reducing the advantage scores of high-entropy tokens that are repeatedly generated within a single reasoning path. 2) The Time Dimension: It reduces excessive overthinking of the system by establishing a historical baseline for the occurrence count of each critical token in previously successful trajectories, and clipping the advantages of tokens that exceed this baseline during training. Crucially, to preserve the model's exploratory capabilities on novel problems, this suppression mechanism is automatically disabled when no historical record of success is available. Experiments conducted on the Deepseek-Distill-32B and Qwen3-8B models show that TFAC outperforms leading baseline methods, improving performance by 2.3 and 3.1 percentage points, respectively, while simultaneously reducing inference costs by 35% and 28% in scenarios where correct answers are generated. These results validate the significant efficacy of TFAC in training large reasoning models that are both powerful and highly efficient. Rong Bao, Bo Wang 0084, Xiao Wang 0014, Hongyu Li 0004, Leszek Rutkowski, Qi Zhang 0001, Liang Ding 0006, Dacheng Tao |
AAAI | 3 |
| 2026 | When Person Re-Identification Meets Event Camera: A Benchmark Dataset and an Attribute-Guided Re-Identification FrameworkabstractRecent researchers have proposed using event cameras for person re-identification (ReID) due to their promising performance and better balance in terms of privacy protection, event camera-based person ReID has attracted significant attention. Currently, mainstream event-based person ReID algorithms primarily focus on fusing visible light and event stream, as well as preserving privacy. Although significant progress has been made, these methods are typically trained and evaluated on small-scale or simulated event camera datasets, making it difficult to assess their real identification performance and generalization ability. To address the issue of data scarcity, this paper introduces a large-scale RGB-event based person ReID dataset, called EvReID. The dataset contains 118,988 image pairs and covers 1200 pedestrian identities, with data collected across multiple seasons, scenes, and lighting conditions. We also evaluate 15 state-of-the-art person ReID algorithms, laying a solid foundation for future research in terms of both data and benchmarking. Based on our newly constructed dataset, this paper further proposes a pedestrian attribute-guided contrastive learning framework to enhance feature learning for person re-identification, termed TriPro-ReID. This framework not only effectively explores the visual features from both RGB frames and event streams, but also fully utilizes pedestrian attributes as mid-level semantic features. Extensive experiments on the EvReID dataset and MARS datasets fully validated the effectiveness of our proposed RGB-Event person ReID framework. Xiao Wang 0014, Shujuan Wu, Bo Jiang 0002, Shiliang Zhang |
AAAI | 1 |
| 2026 | Spatio-temporal side tuning pre-trained foundation models for video-based pedestrian attribute recognition
Xiao Wang 0014, Jiandong Jin, Jun Zhu 0001, Futian Wang, Bo Jiang 0002, Yaowei Wang 0001, Yonghong Tian 0001 |
Comput. Vis. Image Underst. | 1 |
| 2026 | Event Stream based Human Action Recognition: A High-Definition Benchmark Dataset and Algorithms
Xiao Wang 0014, Shiao Wang, Pengpeng Shao, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001 |
Int. J. Comput. Vis. | 1 |
| 2026 | ESTR-CoT: Towards explainable and accurate event stream based scene text recognition with chain-of-thought reasoning
Xiao Wang 0014, Jingtao Jiang, Qiang Chen 0007, Lan Chen 0003, Lin Zhu 0012, Yaowei Wang 0001, Yonghong Tian 0001, Jin Tang 0001 |
Neurocomputing | 1 |
| 2026 | SequencePAR: Understanding pedestrian attributes via a sequence generation paradigm
Jiandong Jin, Xiao Wang 0014, Yin Lin, Chenglong Li 0002, Lili Huang 0006, Aihua Zheng, Jin Tang 0001 |
Pattern Recognit. | 2 |
| 2026 | Semantic change detection of roads and bridges: A fine-grained dataset and multimodal frequency-driven detector
Qing-Ling Shu, Sibao Chen 0001, Xiao Wang 0014, Zhi-Hui You, Wei Lu 0032, Jin Tang 0001, Bin Luo 0001 |
Pattern Recognit. | 3 |
| 2026 | Revisiting color-event based tracking: A unified network, dataset, and metric
Chuanming Tang, Xiao Wang 0014, Ju Huang, Bo Jiang 0002, Lin Zhu 0012, Shifeng Chen, Jianlin Zhang 0001, Yaowei Wang 0001, Yonghong Tian 0001 |
Pattern Recognit. | 2 |
| 2026 | MambaEVT: Event Stream-Based Visual Object Tracking Using State Space ModelabstractEvent camera-based visual tracking has drawn more and more attention in recent years due to the unique imaging principle and advantages of low energy consumption, high dynamic range, and dense temporal resolution. Current event-based tracking algorithms are gradually hitting their performance bottlenecks, due to the utilization of vision Transformer and the static template for target object localization. In this paper, we propose a novel Mamba-based visual tracking framework that adopts the state space model with linear complexity as a backbone network. The search regions and target template are fed into the vision Mamba network for simultaneous feature extraction and interaction. The output tokens of search regions will be fed into the tracking head for target localization. More importantly, we consider introducing a dynamic template update strategy into the tracking framework using the Memory Mamba network. By considering the diversity of samples in the target template library and making appropriate adjustments to the template memory module, a more effective dynamic template can be integrated. The effective combination of dynamic and static templates allows our Mamba-based tracking algorithm to achieve a good balance between accuracy and computational cost on multiple large-scale datasets, including EventVOT, VisEvent, and FE240hz. The source code and checkpoint have been released on https://github.com/Event-AHU/MambaEVT. Xiao Wang 0014, Shiao Wang, Xixi Wang 0005, Zhicheng Zhao 0002, Lin Zhu 0012, Bo Jiang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Vehicle-Centric Perception via Multimodal Structured Pre-TrainingabstractVehicle-centric perception plays a crucial role in many intelligent systems, including large-scale surveillance systems, intelligent transportation, and autonomous driving. Existing approaches typically employ general pre-trained weights to initialize backbone networks, followed by task-specific fine-tuning. However, these models lack effective learning of vehiclerelated knowledge during pre-training, resulting in poor capability for modeling general vehicle perception representations. To handle this problem, we propose VehicleMAE-V2, a novel vehicle-centric pre-trained large model. By exploring and exploiting vehicle-related multimodal structured priors to guide the masked token reconstruction process, our approach can significantly enhance the model’s capability to learn generalizable representations for vehicle-centric perception. Specifically, we design the Symmetry-guided Mask Module (SMM), Contour-guided Representation Module (CRM) and Semantics-guided Representation Module (SRM) to incorporate three kinds of structured priors into token reconstruction including symmetry, contour and semantics of vehicles respectively. SMM utilizes the vehicle symmetry constraints to avoid retaining symmetric patches and can thus select high-quality masked image patches and reduce information redundancy. CRM minimizes the prob23 ability distribution divergence between contour features and reconstructed features and can thus preserve holistic vehicle structure information during pixel-level reconstruction. SRM aligns image-text features through contrastive learning and cross-modal distillation to address the feature confusion caused by insufficient semantic understanding during masked reconstruction. To support the pre-training of VehicleMAE-V2, we construct Autobot4M, a large-scale dataset comprising approximately 4 million vehicle images and 12,693 text descriptions. Extensive experiments on five downstream tasks demonstrate the superior performance of VehicleMAE-V2. The source code, dataset, and pre-trained large models are available on https://github.com/Vehicle-AHU/VehicleMAE. Xiao Wang 0014, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Adversarial Semantic and Label Perturbation Attack for Pedestrian Attribute RecognitionabstractPedestrian Attribute Recognition (PAR) is an indispensable task in human-centered research and has made great progress in recent years with the development of deep neural networks. However, the potential vulnerability and anti-interference ability have still not been fully explored. To bridge this gap, this paper proposes the first adversarial attack and defense framework for pedestrian attribute recognition. Specifically, we exploit both global- and patch-level attacks on the pedestrian images, based on the pre-trained CLIP-based PAR framework. It first divides the input pedestrian image into non-overlapping patches and embeds them into feature embeddings using a projection layer. Meanwhile, the attribute set is expanded into sentences using prompts and embedded into attribute features using a pre-trained CLIP text encoder. A multi-modal Transformer is adopted to fuse the obtained vision and text tokens, and a feed-forward network is utilized for attribute recognition. Based on the aforementioned PAR framework, we adopt the adversarial semantic and label-perturbation to generate the adversarial noise, termed ASL-PAR. We also design a semantic offset defense strategy to suppress the influence of adversarial attacks. Extensive experiments conducted on both digital domains (i.e., PETA, PA100K, MSP60K, RAPv2) and physical domains fully validated the effectiveness of our proposed adversarial attack and defense strategies for the pedestrian attribute recognition. The source code of this paper will be released on https://github.com/Event-AHU/OpenPAR. Weizhe Kong, Xiao Wang 0014, Ruichong Gao, Chenglong Li 0002, Yu Zhang 0091, Xing Yang 0004, Yaowei Wang 0001, Jin Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object DetectionabstractExisting multimodal UAV object detection methods often overlook the impact of semantic gaps between modalities, which makes it difficult to achieve accurate semantic and spatial alignments and ultimately limits detection performance. To address this problem, we propose a Large Language Model (LLM) guided Progressive feature Alignment Network called LPANet, which leverages the semantic features extracted from a large language model to guide the progressive semantic and spatial alignment between modalities for multimodal UAV object detection. To employ the powerful semantic representation of LLM, we generate the fine-grained text descriptions of each object category by ChatGPT and then extract the semantic features using the large language model MPNet, providing high-level semantic priors to guide multimodal alignment. Based on the semantic features, we guide the semantic and spatial alignments in a progressive manner as follows. First, we design the Semantic Alignment Module (SAM) to pull the semantic features and multimodal visual features of each object closer, alleviating the semantic differences of objects between modalities. Second, we design the Explicit Spatial Alignment Module (ESM) by integrating the semantic relations into the estimation of feature-level offsets, alleviating the coarse spatial misalignment between modalities. Finally, we design the Implicit Spatial alignment Module (ISM), which leverages the cross-modal correlations to aggregate key features from neighboring regions to achieve implicit spatial alignment. Comprehensive experiments on two public multimodal UAV object detection datasets demonstrate that our approach outperforms state-of-the-art multimodal UAV object detectors. The source code will be released on https://github.com/Vehicle-AHU/LPANet. Chenglong Li 0002, Xiao Wang 0014, Bin Luo 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Attribute-Guided Semantic Alignment With Pre-Trained Foundation Models for Vehicle DetectionabstractVehicle detection is a fundamental perception task in intelligent transportation systems and plays a crucial role in enabling reliable traffic perception and analysis. Existing vehicle detectors are typically obtained by training conventional object detection models (e.g., YOLO, RCNN, and DETR series) on vehicle images based on pre-trained backbone networks (e.g., ResNet and ViT). Although some studies introduce large-scale foundation models to improve detection performance, these models are not specifically designed for vehicle-centric scenarios and therefore tend to yield sub-optimal results in complex traffic environments. Moreover, most existing methods heavily rely on visual features and pay limited attention to the alignment between vehicle semantic information and visual representations. In this paper, we propose a novel vehicle detection paradigm, termed VFM-Det, which integrates a pre-trained vehicle foundation model (VehicleMAE) with a large language model (T5) to achieve semantically enhanced vehicle detection for intelligent transportation scenarios. Specifically, the proposed method follows a region proposal-based detection framework and employs VehicleMAE to enhance the features of each proposal. More importantly, we introduce a novel VAtt2Vec module to predict the vehicle semantic attributes corresponding to each proposal and transform them into feature vectors, which further enhance visual features through contrastive learning. Extensive experiments on three vehicle detection benchmark datasets thoroughly proved the effectiveness of our vehicle detector. Specifically, our model improves the baseline approach by +6.0%, +8.4% on the$AP_{0.5}$,$AP_{0.75}$metrics, respectively, on the Cityscapes dataset. The source code of this work will be released athttps://github.com/Event-AHU/VFM-Det Fanghua Hong, Xiao Wang 0014, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2026 | Activating Associative Disease-Aware Vision Token Memory for LLM-Based X-Ray Report GenerationabstractX-ray image based medical report generation achieves significant progress in recent years with the help of large language models, however, these models have not fully exploited the effective information in visual image regions, resulting in reports that are linguistically sound but insufficient in describing key diseases. In this paper, we propose a novel associative memory-enhanced X-ray report generation model that effectively mimics the process of professional doctors writing medical reports. It considers both the mining of global and local visual information and associates historical report information to better complete the writing of the current report. Specifically, given an X-ray image, we first utilize a classification model along with its activation maps to accomplish the mining of visual regions highly associated with diseases and the learning of disease query tokens. Then, we employ a visual Hopfield network to establish memory associations for disease-related tokens, and a report Hopfield network to retrieve report memory information. This process facilitates the generation of high-quality reports based on a large language model and achieves state-of-the-art performance on multiple benchmark datasets, including the IU X-ray, MIMIC-CXR, and Chexpert Plus. The source code and pre-trained models of this work have been released on https://github.com/Event-AHU/Medical_Image_Analysis. Xiao Wang 0014, Fuling Wang, Bo Jiang 0002, Chuanfu Li, Yaowei Wang 0001, Yonghong Tian 0001, Jin Tang 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark DatasetabstractObject detection in event streams has emerged as a cutting-edge research area, demonstrating superior performance in low-light conditions, scenarios with motion blur, and rapid movements. Current detectors leverage spiking neural networks, Transformers, or convolutional neural networks as their core architectures, each with its own set of limitations including restricted performance, high computational overhead, or limited local receptive fields. This paper introduces a novel MoE (Mixture of Experts) heat conduction-based object detection algorithm that strikingly balances accuracy and computational efficiency. Initially, we employ a stem network for event data embedding, followed by processing through our innovative MoE-HCO blocks. Each block integrates various expert modules to mimic heat conduction within event streams. Subsequently, an IoU-based query selection module is utilized for efficient token extraction, which is then channeled into a detection head for the final object detection process. Furthermore, we are pleased to introduce EvDET200K, a novel benchmark dataset for event-based object detection. Captured with a high-definition Prophesee EVK4-HD event camera, this dataset encompasses 10 distinct categories, 200,000 bounding boxes, and 10,054 samples, each spanning 2 to 5 seconds. We also provide comprehensive results from over 15 state-of-the-art detectors, offering a solid foundation for future research and comparison. The source code has been released on: https://github.com/Event-AHU/OpenEvDET Xiao Wang 0014, Wei Zhang 0161, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001 |
CVPR | 1 |
| 2025 | CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus DatasetabstractX-ray image-based medical report generation (MRG) is a pivotal area in artificial intelligence that can significantly reduce diagnostic burdens and patient wait times. Despite significant progress, we believe that the task has reached a bottleneck due to the limited benchmark datasets and the existing large models’ insufficient capability enhancements in this specialized domain. Specifically, the recently released CheXpert Plus dataset lacks comparative evaluation algorithms and their results, providing only the dataset itself. This situation makes the training, evaluation, and comparison of subsequent algorithms challenging. Thus, we conduct a comprehensive benchmarking of existing mainstream X-ray report generation models and large language models (LLMs), on the CheXpert Plus dataset. We believe that the proposed benchmark can provide a solid comparative basis for subsequent algorithms and serve as a guide for researchers to quickly grasp the state-of-the-art models in this field. More importantly, we propose a large model for the X-ray image report generation using a multi-stage pre-training strategy, including self-supervised autoregressive generation and Xray-report contrastive learning, and supervised fine-tuning. Extensive experimental results indicate that the autoregressive pre-training based on Mamba effectively encodes X-ray images, and the image-text contrastive pre-training further aligns the feature spaces, achieving better experimental results. Source code can be found on https://github.com/Event-AHU/Medical_Image_Analysis. Xiao Wang 0014, Fuling Wang, Yuehang Li, Qingchuan Ma, Shiao Wang, Bo Jiang 0002, Jin Tang 0001 |
CVPR | 1 |
| 2025 | Mix-Mask Augmentation and Self-Reconstruction for Cross-Domain Few-Shot Hyperspectral Image ClassificationabstractRecently, the metric-based prototypical methods achieves promising performance in few-shot learning (FSL) for hyperspectral image (HSI) classification. However, the existing models are easily affected by the noisy pixels of different categories around the center pixel of the patch, and tend to focus on the most representative features while ignoring other important ones, which cause the overfitting problem. Moreover, the commonly used dimension reduction operation of the feature of source and target domains inevitably results in the loss of valuable spectral information. To address these issues, we propose the mix-mask augmentation and self-reconstruction for cross-domain HSI classification. The pixel mask augmentation is introduced to enhance the sample diversity of query set and suppress the impact of noisy pixels, thus encouraging the model to discover discriminative features on a wider range. The CutMix augmentation is also adopted to generate the mixed support set and mixed prototypes, mitigating the negative impact of confusing prototypes. Furthermore, we develop the self-reconstruction module which can preserve more useful feature information during the dimension reduction for feature representation of the source and target domains. Extensive experiments on three public HSI datasets demonstrate that the proposed method achieves superior performance with fewer computational costs in comparison with the SOTA methods. Qihang Wu, Xiao Wang 0014, Jinpei Liu, Bo Jiang 0002 |
ICASSP | 5 |
| 2025 | EMatch: A Unified Framework for Event-Based Optical Flow and Stereo MatchingabstractEvent cameras have shown promise in vision applications like optical flow estimation and stereo matching, with many specialized architectures leveraging the asynchronous and sparse nature of event data. However, existing works only focus event data within the confines of task-specific domains, overlooking how tasks across the temporal and spatial domains can reinforce each other. In this paper, we reformulate event-based flow estimation and stereo matching as a unified dense correspondence matching problem, enabling us to solve both tasks within a single model by directly matching features in a shared representation space. Specifically, our method utilizes a Temporal Recurrent Network to aggregate event features across temporal or spatial domains, and a Spatial Contextual Attention to enhance knowledge transfer across event flows via temporal or spatial interactions. By utilizing a shared feature similarities module that integrates knowledge from event streams via temporal or spatial interactions, our network performs optical flow estimation from temporal event segment inputs and stereo matching from spatial event segment inputs simultaneously. We demonstrate that our unified model inherently supports multi-task fusion and cross-task transfer. Without the need for retraining for specific task, our model can effectively handle both optical flow and stereo estimation, achieving state-of-the-art performance on both tasks. Pengjie Zhang, Lin Zhu 0012, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ICCV | 3 |
| 2025 | EvFocus: Learning to Reconstruct Sharp Images from Out-of-Focus Event StreamsabstractEvent cameras are innovative sensors that capture brightness changes as asynchronous events rather than traditional intensity frames. These cameras offer substantial advantages over conventional cameras, including high temporal resolution, high dynamic range, and the elimination of motion blur. However, defocus blur, a common image quality degradation resulting from out-of-focus lenses, complicates the challenge of event-based imaging. Due to the unique imaging mechanism of event cameras, existing focusing algorithms struggle to operate efficiently on sparse event data. In this work, we propose EvFocus, a novel architecture designed to reconstruct sharp images from defocus event streams for the first time. Our work includes the development of an event-based out-of-focus camera model and a simulator to generate realistic defocus event streams for robust training and testing. EvDefous integrates a temporal information encoder, a blur-aware two-branch decoder, and a reconstruction and re-defocus module to effectively learn and correct defocus blur. Extensive experiments on both simulated and real-world datasets demonstrate that EvFocus outperforms existing methods across varying lighting conditions and blur sizes, proving its robustness and practical applicability in event-based defocus imaging. Lin Zhu 0012, Xiantao Ma, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ICML | 3 |
| 2025 | Revealing Latent Information: A Physics-inspired Self-supervised Pre-training Framework for Noisy and Sparse Events
Lin Zhu 0012, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ACM Multimedia | 3 |
| 2025 | CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training FrameworkabstractEvent cameras have attracted increasing attention in recent years due to their advantages in high dynamic range, high temporal resolution, low power consumption, and low latency. Some researchers have begun exploring pre-training directly on event data. Nevertheless, these efforts often fail to establish strong connections with RGB frames, limiting their applicability in multi-modal fusion scenarios. To address these issues, we propose a novel CM3AE pre-training framework for the RGB-Event perception. This framework accepts multi-modalities/views of data as input, including RGB images, event images, and event voxels, providing robust support for both event-based and RGB-event fusion based downstream tasks. Specifically, we design a multi-modal fusion reconstruction module that reconstructs the original image from fused multi-modal features, explicitly enhancing the model's ability to aggregate cross-modal complementary information. Additionally, we employ a multi-modal contrastive learning strategy to align cross-modal feature representations in a shared latent space, which effectively enhances the model's capability for multi-modal understanding and capturing global dependencies. We construct a large-scale dataset containing 2,535,759 RGB-Event data pairs for the pre-training. Extensive experiments on five downstream tasks fully demonstrated the effectiveness of CM3AE. Source code and pre-trained models will be released on https://github.com/Event-AHU/CM3AE. Xiao Wang 0014, Chenglong Li 0002, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Qi Liu 0003 |
ACM Multimedia | 2 |
| 2025 | Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionabstractEvent cameras provide asynchronous, low-latency, and high-dynamic-range visual signals, making them ideal for real-time perception tasks such as object detection. However, effectively modeling the temporal dynamics of event streams remains a core challenge. Most existing methods follow frame-based detection paradigms, applying temporal modules only at high-level features, which limits early-stage temporal modeling. Transformer-based approaches introduce global attention to capture long-range dependencies, but often add unnecessary complexity and overlook fine-grained temporal cues. In this paper, we propose a CNN-RNN hybrid framework that rethinks temporal modeling for event-based object detection. Our approach is based on two key insights: (1) introducing recurrent modules at lower spatial scales to preserve detailed temporal information where events are most dense, and (2) utilizing Decoupled Deformable-enhanced Recurrent Layers specifically designed according to the inherent motion characteristics of event cameras to extract multiple spatiotemporal features, and performing independent downsampling at multiple spatiotemporal scales to enable flexible, scale-aware representation learning. These multi-scale features are then fused via a feature pyramid network to produce robust detection outputs. Experiments on Gen1, 1 Mpx and eTram dataset demonstrate that our approach achieves superior accuracy over recent transformer-based models, highlighting the importance of precise temporal feature extraction in early stages. This work offers a new perspective on designing architectures for event-driven vision beyond attention-centric paradigms. Code: https://github.com/BIT-Vision/SATE. Lin Zhu 0012, Tengyu Long, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
NeurIPS | 3 |
| 2025 | Learning Dynamic Batch-Graph Representation for Deep Representation Learning
Xixi Wang 0005, Bo Jiang 0002, Xiao Wang 0014, Bin Luo 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | Continuous-Time Object Segmentation Using High Temporal Resolution Event CameraabstractEvent cameras are novel bio-inspired sensors, where individual pixels operate independently and asynchronously, generating intensity changes as events. Leveraging the microsecond resolution (no motion blur) and high dynamic range (compatible with extreme light conditions) of events, there is considerable promise in directly segmenting objects from sparse and asynchronous event streams in various applications. However, different from the rich cues in video object segmentation, it is challenging to segment complete objects from the sparse event stream. In this paper, we present the first framework for continuous-time object segmentation from event stream. Given the object mask at the initial time, our task aims to segment the complete object at any subsequent time in event streams. Specifically, our framework consists of a Recurrent Temporal Embedding Extraction (RTEE) module based on a novel ResLSTM, a Cross-time Spatiotemporal Feature Modeling (CSFM) module which is a transformer architecture with long-term and short-term matching modules, and a segmentation head. The historical events and masks (reference sets) are recurrently fed into our framework along with current-time events. The temporal embedding is updated as new events are input, enabling our framework to continuously process the event stream. To train and test our model, we construct both real-world and simulated event-based object segmentation datasets, each comprising event streams, APS images, and object annotations. Extensive experiments on our datasets demonstrate the effectiveness of the proposed recurrent architecture. Lin Zhu 0012, Xianzhang Chen, Lizhi Wang 0001, Xiao Wang 0014, Yonghong Tian 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | A complementary dual model for weakly supervised salient object detection
Dawei Zhang 0002, Xiao Wang 0014, Chang-Dong Wang 0001, Zhonglong Zheng |
Pattern Recognit. | 3 |
| 2025 | Semantic-aware frame-event fusion based pattern recognition via large vision-language models
Jiandong Jin, Yanlin Zhong, Yaoyang Wu, Lan Chen 0003, Xiao Wang 0014, Bin Luo 0001 |
Pattern Recognit. | 7 |
| 2025 | Temporal adaptive bidirectional bridging for RGB-D tracking
Ge Ying, Dawei Zhang 0002, Zhou Ou, Xiao Wang 0014, Zhonglong Zheng |
Pattern Recognit. | 4 |
| 2025 | Pedestrian Attribute Recognition via CLIP-Based Prompt Vision-Language FusionabstractExisting pedestrian attribute recognition (PAR) algorithms adopt pre-trained CNN (e.g., ResNet) as their backbone network for visual feature learning, which might obtain sub-optimal results due to the insufficient employment of the relations between pedestrian images and attribute labels. In this paper, we formulate PAR as a vision-language fusion problem and fully exploit the relations between pedestrian images and attribute labels. Specifically, the attribute phrases are first expanded into sentences, and then the pre-trained vision-language model CLIP is adopted as our backbone for feature embedding of visual images and attribute descriptions. The contrastive learning objective connects the vision and language modalities well in the CLIP-based feature space, and the Transformer layers used in CLIP can capture the long-range relations between pixels. Then, a multi-modal Transformer is adopted to fuse the dual features effectively and feed-forward network is used to predict attributes. To optimize our network efficiently, we propose the region-aware prompt tuning technique to adjust very few parameters (i.e., only the prompt vectors and classification heads) and fix both the pre-trained VL model and multi-modal Transformer. Our proposed PAR algorithm only adjusts 0.75% learnable parameters compared with the fine-tuning strategy. It also achieves new state-of-the-art performance on both standard and zero-shot settings for PAR, including RAPv1, RAPv2, WIDER, PA100K, and PETA-ZS, RAP-ZS datasets. The source code and pre-trained models will be released onhttps://github.com/Event-AHU/OpenPAR. Xiao Wang 0014, Jiandong Jin, Chenglong Li 0002, Jin Tang 0001, Cheng Zhang 0010, Wei Wang 0115 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Mask-Guided Frequency Feature Fusion for Visible-Infrared Remote Sensing Object DetectionabstractVisible-infrared remote sensing object detection aims to achieve all-weather object detection by leveraging the complementary information from paired visible and infrared (RGB-IR) images. However, modality differences and weak alignment often limit its performance. Existing methods largely neglect the frequency discrepancies between modalities and require strict alignment, increasing complexity. To address these challenges, this study proposes a novel mask-guided frequency feature fusion (MGFF) method for RGB-IR object detection in remote sensing. Specifically, we develop a feature frequency decomposition and enhancement module using wavelet transform to reduce modality differences between RGB and IR images by restructuring and enhancing their frequency components. Additionally, we introduce a mask-guided feature reconstruction module and a feature-guided consistency loss, ensuring that even under weak alignment, the focus remains on integrating the target features from different modalities. Meanwhile, this loss is used to guide the reconstruction of features from different modalities. Finally, We design a multi-directional perception cross-modality fusion module to achieve deep fusion of multimodal information, which enhances object perception from different directions across modalities. Extensive evaluations on the widely recognized RGB-IR remote sensing benchmarks, including DroneVehicle and VEDAI, as well as the RGB-IR pedestrian dataset KAIST, substantiate the effectiveness of the proposed MGFF method. The results consistently demonstrate that the MGFF achieves a superior performance in terms of detection accuracy and robustness compared to existing state-of-the-art approaches. Xiangqi Chen, Li Zhao 0005, Chengzhuan Yang, Dawei Zhang 0002, Xiao Wang 0014, Xiaowei He 0003, Hua Wang 0002, Zhonglong Zheng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Retain, Blend, and Exchange: A Quality-Aware Spatial-Stereo Fusion Approach for Event Stream RecognitionabstractCurrent event stream-based pattern recognition models typically present the event stream as the point cloud, voxel, image, and the like, and formulate multiple deep neural networks to acquire their features. Although considerable results can be achieved in simple cases, however, the performance of the model might be restricted by monotonous modality expressions, sub-optimal fusion, and readout mechanisms. In this article, we put forward a novel dual-stream framework for event stream-based pattern recognition through differentiated fusion, which is called EFV++. It models two common event representations simultaneously, i.e., event images and event voxels. The spatial and three-dimensional stereo information can be separately learned by making use of Transformer and Graph Neural Network (GNN). We believe the features of each representation still contain both efficient and redundant features and a sub-optimal solution may be obtained if we directly fuse them without differentiation. Thus, we divide each feature into three levels and retain high-quality features, blend medium-quality features, and exchange low-quality features. The enhanced dual features will be provided to the fusion Transformer together with bottleneck features. In addition, we introduce a novel hybrid interaction readout mechanism to enhance the diversity of features as final representations. Comprehensive experiments validate that the framework we have proposed attains cutting-edge performance on a variety of extensively utilized event stream-based classification datasets. Particularly, we have realized a freshly pioneering performance on the Bullying10 k dataset, precisely 90.51%, and this outpaces the runner-up by$+2.21\%$. Lan Chen 0003, Xiao Wang 0014, Pengpeng Shao, Wei Zhang 0161, Yaowei Wang 0001, Yonghong Tian 0001, Jin Tang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | CRSOT: Cross-Resolution Object Tracking Using Unaligned Frame and Event CamerasabstractExisting datasets for RGB-DVS tracking are collected with DVS346 camera and their resolution ($346 \times 260$) is low for practical applications. Actually, only visible cameras are deployed in many practical systems, and the newly designed neuromorphic cameras may have different resolutions. The latest neuromorphic sensors can output high-definition event streams, but it is very difficult to achieve strict alignment between events and frames on both spatial and temporal views. Therefore, how to achieve accurate tracking with unaligned neuromorphic and visible sensors is a valuable but unresearched problem. In this work, we formally propose the task of object tracking using unaligned neuromorphic and visible cameras. We build the first unaligned frame-event dataset CRSOT collected with a specially built data acquisition system, which contains 1,030 high-definition RGB-Event video pairs, 304,974 video frames. In addition, we propose a novel unaligned object tracking framework that can realize robust tracking even using the loosely aligned RGB-Event data. This proposed method utilizes uncertainty perception techniques, which can effectively reduce the negative impact of noise (especially noise in event data) on tracking performance. Specifically, we extract the template and search regions of RGB and Event data and feed them into a unified ViT backbone for feature embedding. Next, we propose uncertainty perception modules to encode the RGB and Event features, respectively, then, we propose a modality uncertainty fusion module to aggregate the two modalities. These three branches are jointly optimized in the training phase. Extensive experiments demonstrate that our tracker can collaborate the dual modalities for high-performance tracking even without strictly temporal and spatial alignment. Yabin Zhu, Xiao Wang 0014, Chenglong Li 0002, Bo Jiang 0002, Lin Zhu 0012, Zhixiang Huang, Yonghong Tian 0001, Jin Tang 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | HARDVS: Revisiting Human Activity Recognition with Dynamic Vision SensorsabstractThe main streams of human activity recognition (HAR) algorithms are developed based on RGB cameras which usually suffer from illumination, fast motion, privacy preservation, and large energy consumption. Meanwhile, the biologically inspired event cameras attracted great interest due to their unique features, such as high dynamic range, dense temporal but sparse spatial resolution, low latency, low power, etc. As it is a newly arising sensor, even there is no realistic large-scale dataset for HAR. Considering its great practical value, in this paper, we propose a large-scale benchmark dataset to bridge this gap, termed HARDVS, which contains 300 categories and more than 100K event sequences. We evaluate and report the performance of multiple popular HAR algorithms, which provide extensive baselines for future works to compare. More importantly, we propose a novel spatial-temporal feature learning and fusion framework, termed ESTF, for event stream based human activity recognition. It first projects the event streams into spatial and temporal embeddings using StemNet, then, encodes and fuses the dual-view representations using Transformer networks. Finally, the dual features are concatenated and fed into a classification head for activity prediction. Extensive experiments on multiple datasets fully validated the effectiveness of our model. Both the dataset and source code will be released at https://github.com/Event-AHU/HARDVS. Xiao Wang 0014, Zongzhen Wu, Bo Jiang 0002, Zhimin Bao, Lin Zhu 0012, Guoqi Li 0002, Yaowei Wang 0001, Yonghong Tian 0001 |
AAAI | 1 |
| 2024 | Structural Information Guided Multimodal Pre-training for Vehicle-Centric PerceptionabstractUnderstanding vehicles in images is important for various applications such as intelligent transportation and self-driving system. Existing vehicle-centric works typically pre-train models on large-scale classification datasets and then fine-tune them for specific downstream tasks. However, they neglect the specific characteristics of vehicle perception in different tasks and might thus lead to sub-optimal performance. To address this issue, we propose a novel vehicle-centric pre-training framework called VehicleMAE, which incorporates the structural information including the spatial structure from vehicle profile information and the semantic structure from informative high-level natural language descriptions for effective masked vehicle appearance reconstruction. To be specific, we explicitly extract the sketch lines of vehicles as a form of the spatial structure to guide vehicle reconstruction. The more comprehensive knowledge distilled from the CLIP big model based on the similarity between the paired/unpaired vehicle image-text sample is further taken into consideration to help achieve a better understanding of vehicles. A large-scale dataset is built to pre-train our model, termed Autobot1M, which contains about 1M vehicle images and 12693 text information. Extensive experiments on four vehicle-based downstream tasks fully validated the effectiveness of our VehicleMAE. The source code and pre-trained models will be released at https://github.com/Event-AHU/VehicleMAE. Xiao Wang 0014, Chenglong Li 0002, Zhicheng Zhao 0001, Zhe Chen 0013, Yukai Shi, Jin Tang 0001 |
AAAI | 1 |
| 2024 | Finding Visual Saliency in Continuous Spike StreamabstractAs a bio-inspired vision sensor, the spike camera emulates the operational principles of the fovea, a compact retinal region, by employing spike discharges to encode the accumulation of per-pixel luminance intensity. Leveraging its high temporal resolution and bio-inspired neuromorphic design, the spike camera holds significant promise for advancing computer vision applications. Saliency detection mimic the behavior of human beings and capture the most salient region from the scenes. In this paper, we investigate the visual saliency in the continuous spike stream for the first time. To effectively process the binary spike stream, we propose a Recurrent Spiking Transformer (RST) framework, which is based on a full spiking neural network. Our framework enables the extraction of spatio-temporal features from the continuous spatio-temporal spike stream while maintaining low power consumption. To facilitate the training and validation of our proposed model, we build a comprehensive real-world spike-based visual saliency dataset, enriched with numerous light conditions. Extensive experiments demonstrate the superior performance of our Recurrent Spiking Transformer framework in comparison to other spike neural network-based methods. Our framework exhibits a substantial margin of improvement in capturing and highlighting visual saliency in the spike stream, which not only provides a new perspective for spike-based saliency segmentation but also shows a new paradigm for full SNN-based transformer models. The code and dataset are available at https://github.com/BIT-Vision/SVS. Lin Zhu 0012, Xianzhang Chen, Xiao Wang 0014, Hua Huang 0001 |
AAAI | 3 |
| 2024 | Event Stream-Based Visual Object Tracking: A High-Resolution Benchmark Dataset and A Novel BaselineabstractTracking with bio-inspired event cameras has garnered increasing interest in recent years. Existing works either utilize aligned RGB and event data for accurate tracking or directly learn an event-based tracker. The former incurs higher inference costs while the latter may be susceptible to the impact of noisy events or sparse spatial resolution. In this paper, we propose a novel hierarchical knowledge distillation framework that can fully utilize multimodal / multi-view information during training to facilitate knowledge transfer, enabling us to achieve high-speed and low-latency visual tracking during testing by using only event signals. Specifically, a teacher Transformer-based multimodal tracking framework is first trained by feeding the RGB frame and event stream simultaneously. Then, we design a new hierarchical knowledge distillation strategy which includes pairwise similarity, feature representation, and response maps-based knowledge distillation to guide the learning of the student Transformer network. In particular, since existing event-based tracking datasets are all low-resolution (346 × 260), we propose the first large-scale high-resolution (1280 × 720) dataset named EventVOT. It contains 1141 videos and covers a wide range of categories such as pedestrians, vehicles, UAVs, ping pong, etc. Ex-tensive experiments on both low-resolution (FE240hz, Vi-sEvent, COESOT), and our newly proposed high-resolution EventVOT dataset fully validated the effectiveness of our proposed method. Xiao Wang 0014, Shiao Wang, Chuanming Tang, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001, Jin Tang 0001 |
CVPR | 1 |
| 2024 | Temporal Residual Guided Diffusion Framework for Event-Driven Video Reconstruction
Lin Zhu 0012, Yunlong Zheng, Yijun Zhang 0003, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ECCV (40) | 4 |
| 2024 | A Reinforced Passage Interactive Retrieval Framework Incorporating Implicit Knowledge for KB-VQA
Hongyan Zheng, Gengchen Liu, Xiao Wang 0014 |
ICIC (12) | 7 |
| 2024 | Mamba-FETrack: Frame-Event Tracking via State Space Model
Ju Huang, Shiao Wang, Zhe Wu 0006, Xiao Wang 0014, Bo Jiang 0002 |
PRCV (12) | 5 |
| 2024 | MutualFormer: Multi-modal Representation Learning via Cross-Diffusion Attention
Xixi Wang 0005, Xiao Wang 0014, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Learning Graph Attentions via Replicator DynamicsabstractGraph Attention (GA) which aims to learn the attention coefficients for graph edges has achieved impressive performance in GNNs on many graph learning tasks. However, existing GAs are usually learned based on edges' (or connected nodes') features which fail to fully capture the rich structural information of edges. Some recent research attempts to incorporate the structural information into GA learning but how to fully exploit them in GA learning is still a challenging problem. To address this challenge, in this work, we propose to leverage a new Replicator Dynamics model for graph attention learning, termed Graph Replicator Attention (GRA). The core of GRA is our derivation of replicator dynamics based sparse attention diffusion which can explicitly learn context-aware and sparse preserved graph attentions via a simple self-supervised way. Moreover, GRA can be theoretically explained from an energy minimization model. This provides a more theoretical justification for the proposed GRA method. Experiments on several graph learning tasks demonstrate the effectiveness and advantages of the proposed GRA method on ten benchmark datasets. Bo Jiang 0002, Sheng Ge, Beibei Wang 0006, Xiao Wang 0014, Jin Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | RGBT Tracking via Progressive Fusion Transformer With Dynamically Guided LearningabstractExisting Transformer-based RGB-Thermal (RGBT) tracking methods either use cross-attention to fuse the two modalities, or use self-attention and cross-attention to model both modality-specific and modality-sharing information. However, the significant appearance gap between modalities limits the feature representation ability of certain modalities during the fusion process. To address this problem, we propose a novel Progressive Fusion Transformer called ProFormer, which progressively integrates single-modality information into the multimodal representation for robust RGBT tracking. In particular, ProFormer first uses a self-attention module to collaboratively extract the multimodal representation. Then, ProFormer introduces two cross-attention modules to interact it with the features of the dual modalities for enhancing modality-specific information in the multimodal representation. In addition, we propose a dynamically guided learning algorithm that adaptively employs the well-performing branches to guide the learning of other branches, to improve the representation ability of each branch. Extensive experiments demonstrate that our proposed ProFormer achieves a new state-of-the-art performance on RGBT210, RGBT234, LasHeR, and VTUAV datasets. Yabin Zhu, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001, Zhixiang Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | VisEvent: Reliable Object Tracking via Collaboration of Frame and Event FlowsabstractDifferent from visible cameras which record intensity images frame by frame, the biologically inspired event camera produces a stream of asynchronous and sparse events with much lower latency. In practice, visible cameras can better perceive texture details and slow motion, while event cameras can be free from motion blurs and have a larger dynamic range which enables them to work well under fast motion and low illumination (LI). Therefore, the two sensors can cooperate with each other to achieve more reliable object tracking. In this work, we propose a large-scale Visible-Event benchmark (termed VisEvent) due to the lack of a realistic and scaled dataset for this task. Our dataset consists of 820 video pairs captured under LI, high speed, and background clutter scenarios, and it is divided into a training and a testing subset, each of which contains 500 and 320 videos, respectively. Based on VisEvent, we transform the event flows into event images and construct more than 30 baseline methods by extending current single-modality trackers into dual-modality versions. More importantly, we further build a simple but effective tracking algorithm by proposing a cross-modality transformer, to achieve more effective feature fusion between visible and event data. Extensive experiments on the proposed VisEvent dataset, FE108, COESOT, and two simulated datasets (i.e., OTB-DVS and VOT-DVS), validated the effectiveness of our model. The dataset and source code have been released on: https://github.com/wangxiao5791509/VisEvent_SOT_Benchmark. Xiao Wang 0014, Jianing Li 0001, Lin Zhu 0012, Zhe Chen 0013, Xin Li 0034, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001 |
IEEE Trans. Cybern. | 1 |
| 2024 | AMatFormer: Efficient Feature Matching via Anchor Matching TransformerabstractLearning based feature matching methods have been commonly studied in recent years. The core issue for learning feature matching is to how to learn (1) discriminative representations for feature points (or regions) within each intra-image and (2) consensus representations for feature points across inter-images. Recently, self- and cross-attention models have been exploited to address this issue. However, in many scenes, features are coming with large-scale, redundant and outliers contaminated. Previous self-/cross-attention models generally conduct message passing on all primal features which thus lead to redundant learning and high computational cost. To mitigate limitations, inspired by recent seed matching methods, in this article, we propose a novel efficient Anchor Matching Transformer (AMatFormer) for the feature matching problem. AMatFormer has two main aspects: First, it mainly conducts self-/cross-attention on some anchor features and leverages these anchor features as message bottleneck to learn the representations for all primal features. Thus, it can be implemented efficiently and compactly. Second, AMatFormer adopts a shared FFN module to further embed the features of two images into the common domain and thus learn the consensus feature representations for the matching problem. Experiments on several benchmarks demonstrate the effectiveness and efficiency of the proposed AMatFormer matching approach. Bo Jiang 0002, Shuxian Luo, Xiao Wang 0014, Chuanfu Li, Jin Tang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Rethinking Batch Sample Relationships for Data Representation: A Batch-Graph Transformer Based ApproachabstractExploring sample relationships within each mini-batch has shown great potential for learning image representations. Existing works generally adopt the regular Transformer to model the visual content relationships, ignoring the cues of semantic/label correlations between samples. Also, they generally adopt the ‘full’ self-attention mechanism which are obviously redundant and also sensitive to the noisy samples. To overcome these issues, in this paper, we design a simple yet flexible Batch-Graph Transformer (BGFormer) for mini-batch sample representations by deeply capturing the relationships of image samples from both visual and semantic perspectives. BGFormer has three main aspects. (1) It employs a flexible graph model, termedBatch Graphto jointly encode both visual and semantic relationships of samples within each mini-batch. (2) It explores the neighborhood relationships of samples by borrowing the idea of sparse graph representation which thus performs robustly, w.r.t., noisy samples. (3) It devises a novel specific Transformer architecture that mainly adoptsdualstructure-constrained self-attention (SSA), together with graph normalization, FFN, etc, to carefully exploit the batch graph information for sample tokens (nodes) representations. As an application, we apply BGFormer to the metric learning tasks. Extensive experiments on four popular datasets demonstrate the effectiveness of the proposed model. Xixi Wang 0005, Bo Jiang 0002, Xiao Wang 0014, Jinhui Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Prompt-Based Learning for Unpaired Image CaptioningabstractUnpaired Image Captioning (UIC) has been developed to learn image descriptions from unaligned vision-language sample pairs. Existing works usually tackle this task using adversarial learning and visual concept reward based on reinforcement learning. However, these existing works were only able to learn limited cross-domain information in vision and language domains, which restrains the captioning performance of UIC. Inspired by the success of Vision-Language Pre-Trained Models (VL-PTMs) in this research, we attempt to infer the cross-domain cue information about a given image from the large VL-PTMs for the UIC task. This research is also motivated by recent successes of prompt learning in many downstream multi-modal tasks, including image-text retrieval and vision question answering. In this work, a semantic prompt is introduced and aggregated with visual features for more accurate caption prediction under the adversarial learning framework. In addition, a metric prompt is designed to select high-quality pseudo image-caption samples obtained from the basic captioning model and refine the model in an iterative manner. Extensive experiments on the COCO and Flickr30 K datasets validate the promising captioning ability of the proposed model. We expect that the proposed prompt-based UIC model will stimulate a new line of research for the VL-PTMs based captioning. Peipei Zhu, Xiao Wang 0014, Lin Zhu 0012, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 2 |
| 2024 | Tiny Object Tracking: A Large-Scale Dataset and a BaselineabstractTiny objects, frequently appearing in practical applications, have weak appearance and features, and receive increasing interests in many vision tasks, such as object detection and segmentation. To promote the research and development of tiny object tracking, we create a large-scale video dataset, which contains 434 sequences with a total of more than 217K frames. Each frame is carefully annotated with a high-quality bounding box. In data creation, we take 12 challenge attributes into account to cover a broad range of viewpoints and scene complexities, and annotate these attributes for facilitating the attribute-based performance analysis. To provide a strong baseline in tiny object tracking, we propose a novel multilevel knowledge distillation network (MKDNet), which pursues three-level knowledge distillations in a unified framework to effectively enhance the feature representation, discrimination, and localization abilities in tracking tiny objects. Extensive experiments are performed on the proposed dataset, and the results prove the superiority and effectiveness of MKDNet compared with state-of-the-art methods. The dataset, the algorithm code, and the evaluation code are available at https://github.com/mmic-lcl/Datasets-and-benchmark-code. Yabin Zhu, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001, Zhixiang Huang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Transformer vision-language tracking via proxy token guided cross-modal fusion
Haojie Zhao, Xiao Wang 0014, Dong Wang 0004, Huchuan Lu, Xiang Ruan |
Pattern Recognit. Lett. | 2 |
| 2023 | Learning Spatial-Frequency Transformer for Visual Object TrackingabstractRecently, some researchers have begun to adopt the Transformer to combine or replace the widely used ResNet as their new backbone network. As the Transformer captures the long-range relations between pixels well using the self-attention scheme, which complements the issues caused by the limited receptive field of CNN. Although their trackers work well in regular scenarios, they simply flatten the 2D features into a sequence to better match the Transformer. We believe these operations ignore the spatial prior of the target object, which may lead to sub-optimal results only. In addition, many works demonstrate that self-attention is actually a low-pass filter, which is independent of input features or keys/queries. That is to say, it may suppress the high-frequency component of the input features and preserve or even amplify the low-frequency information. To handle these issues, in this paper, we propose a unified Spatial-Frequency Transformer that models the Gaussian spatial Prior and High-frequency emphasis Attention (GPHA) simultaneously. To be specific, Gaussian spatial prior is generated using dual Multi-Layer Perceptrons (MLPs) and injected into the similarity matrix produced by multiplying Query and Key features in self-attention. The output will be fed into a softmax layer and then decomposed into two components, i.e., the direct and high-frequency signal. The low- and high-pass branches are rescaled and combined to achieve all-pass, therefore, the high-frequency features will be protected well in stacked self-attention layers. We further integrate the Spatial-Frequency Transformer into the Siamese tracking framework and propose a novel tracking algorithm termed SFTransT. The cross-scale fusion based SwinTransformer is adopted as the backbone, and also a multi-head cross-attention module is used to boost the interaction between search and template features. The output will be fed into the tracking head for target localization. Extensive experiments on short-term and long-term tracking benchmarks all demonstrate the effectiveness of our proposed framework. Source code will be released athttps://github.com/Tchuanm/SFTransT.git. Chuanming Tang, Xiao Wang 0014, Yuanchao Bai, Zhe Wu 0006, Jianlin Zhang 0001, Yongmei Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Few-Shot Learning Meets Transformer: Unified Query-Support Transformers for Few-Shot ClassificationabstractThe goal of Few-shot classification (FSL) is to identify unseen classes with very limited samples has attracted more and more attention. Usually, it is formulated as a metric learning problem. The core issue of few-shot classification is how to learn (1) consistent representations for images in both support and query sets and (2) effective metric learning for images between support and query sets. In this paper, we show that the two challenges can be well modeled simultaneously via a unified Query-Support TransFormer (QSFormer) model. To be specific, the proposed QSFormer involves global query-support sample Transformer (sampleFormer) branch and local patch Transformer (patchFormer) learning branch. sampleFormer aims to capture the dependence of samples in support and query sets for image representation. It adopts the Encoder, QS-Decoder and Cross-Attention to respectively model the Support, Query (image) representation and Metric learning for few-shot classification task. Also, as a complementary to global learning branch, we adopt a local patch Transformer to extract structural representation for each image sample by capturing the long-range dependence of local image patches. In addition, we introduce a novel Cross-scale Interactive Feature Extractor (CIFE) to extract and fuse different scale CNN features as an effective backbone module for the proposed few-shot learning method. We integrate these into a unified framework and train it in an end-to-end way. A large number of experiments are conducted on four popular datasets to validate the superiority and effectiveness of the proposed QSFormer. Xixi Wang 0005, Xiao Wang 0014, Bo Jiang 0002, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | VcT: Visual Change Transformer for Remote Sensing Image Change DetectionabstractGiven two remote sensing images, the goal of visual change detection task is to detect significantly changed areas between them. Existing visual change detectors usually adopt CNNs or Transformers for feature representation learning and focus on learning effective representation for the changed regions between images. Although good performance can be obtained by enhancing the features of the change regions, however, these works are still limited mainly due to the ignorance of mining the unchanged background context information. It is known that one main challenge for change detection is how to obtain the consistent representations for two images involving different variations, such as spatial variation, sunlight intensity, etc. In this work, we demonstrate that carefully mining the common background information provides an important cue to learn the consistent representations for the two images which thus obviously facilitates the visual change detection problem. Based on this observation, we propose a novel Visual change Transformer (VcT) model for visual change detection problem. To be specific, a shared backbone network is first used to extract the feature maps for the given image pair. Then, each pixel of feature map is regarded as a graph node and the graph neural network is proposed to model the structured information for coarse change map prediction. Top-K reliable tokens can be mined from the map and refined by using the clustering algorithm. Then, these reliable tokens are enhanced by first utilizing self/cross-attention schemes and then interacting with original features via an anchor-primary attention learning module. Finally, the prediction head is proposed to get a more accurate change map. Extensive experiments on multiple benchmark datasets validated the effectiveness of our proposed VcT model. The source code and pre-trained models are available at https://github.com/Event-AHU/VcT_Remote_Sensing_Change_Detection. Bo Jiang 0002, Zitian Wang, Xixi Wang 0005, Lan Chen 0003, Xiao Wang 0014, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | MFGNet: Dynamic Modality-Aware Filter Generation for RGB-T TrackingabstractMany RGB-T trackers attempt to attain robust feature representation by utilizing an adaptive weighting scheme (or attention mechanism). Different from these works, we propose a new dynamic modality-aware filter generation module (named MFGNet) to boost the message communication between visible and thermal data by adaptively adjusting the convolutional kernels for various input images in practical tracking. Given the image pairs as input, we first encode their features with the backbone network. Then, we concatenate these feature maps and generate dynamic modality-aware filters with two independent networks. The visible and thermal filters will be used to conduct a dynamic convolutional operation on their corresponding input feature maps respectively. Inspired by residual connection, both the generated visible and thermal feature maps will be summarized with input feature maps. The augmented feature maps will be fed into the RoI align module to generate instance-level features for subsequent classification. To address issues caused by heavy occlusion, fast motion and out-of-view, we propose to conduct a joint local and global search by exploiting a new direction-aware target driven attention mechanism. The spatial and temporal recurrent neural network is used to capture the direction-aware context for accurate global attention prediction. Extensive experiments on three large-scale RGB-T tracking benchmark datasets validated the effectiveness of our proposed algorithm. Xiao Wang 0014, Xiujun Shu, Shiliang Zhang, Bo Jiang 0002, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Unpaired Image Captioning by Image-Level Weakly-Supervised Visual Concept RecognitionabstractThe goal of unpaired image captioning (UIC) is to describe images without using image-caption pairs in the training phase. Although challenging, we expect the task can be accomplished by leveraging images aligned with visual concepts. Most existing studies use off-the-shelf algorithms to obtain the visual concepts because the Bounding Box (BBox) labels or relationship-triplet labels used for training are expensive to acquire. To avoid exhaustive annotations, we propose a novel approach to achieve cost-effective UIC. Specifically, we adopt image-level labels to optimize the UIC model in a weakly-supervised manner. For each image, we assume that only the image-level labels are available without specific locations and numbers. The image-level labels are utilized to train a weakly-supervised object recognition model to extract object information (e.g., instance), and the extracted instances are adopted to infer the relationships among different objects using an enhanced graph neural network (GNN). The proposed approach achieves comparable or even better performance compared with previous methods without expensive annotations. Furthermore, we design an unrecognized object (UnO) loss to improve the alignment of the inferred object and relationship information with the images. It can effectively alleviate the issue encountered by existing UIC models when generating sentences with nonexistent objects. To the best of our knowledge, this is the first attempt to address the problem of Weakly-Supervised visual concept recognition for UIC (WS-UIC) based only on image-level labels. Extensive experiments demonstrate that the proposed method achieves inspiring results on the COCO dataset while significantly reducing the labeling cost. Peipei Zhu, Xiao Wang 0014, Yong Luo 0002, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 2 |
| 2022 | Retinomorphic Object Detection in Asynchronous Visual StreamsabstractDue to high-speed motion blur and challenging illumination, conventional frame-based cameras have encountered an important challenge in object detection tasks. Neuromorphic cameras that output asynchronous visual streams instead of intensity frames, by taking the advantage of high temporal resolution and high dynamic range, have brought a new perspective to address the challenge. In this paper, we propose a novel problem setting, retinomorphic object detection, which is the first trial that integrates foveal-like and peripheral-like visual streams. Technically, we first build a large-scale multimodal neuromorphic object detection dataset (i.e., PKU-Vidar-DVS) over 215.5k spatio-temporal synchronized labels. Then, we design temporal aggregation representations to preserve the spatio-temporal information from asynchronous visual streams. Finally, we present a novel bio-inspired unifying framework to fuse two sensing modalities via a dynamic interaction mechanism. Our experimental evaluation shows that our approach has significant improvements over the state-of-the-art methods with the single-modality, especially in high-speed motion and low-light scenarios. We hope that our work will attract further research into this newly identified, yet crucial research direction. Our dataset can be available at https://www.pkuml.org/resources/pku-vidar-dvs.html. Jianing Li 0001, Xiao Wang 0014, Lin Zhu 0012, Jia Li 0003, Tiejun Huang 0001, Yonghong Tian 0001 |
AAAI | 2 |
| 2022 | Event-based Video Reconstruction via Potential-assisted Spiking Neural NetworkabstractNeuromorphic vision sensor is a new bio-inspired imaging paradigm that reports asynchronous, continuously perpixel brightness changes called ‘events’ with high temporal resolution and high dynamic range. So far, the event-based image reconstruction methods are based on artificial neural networks (ANN) or hand-crafted spatiotemporal smoothing techniques. In this paper, we first implement the image reconstruction work via deep spiking neural network (SNN) architecture. As the bio-inspired neural networks, SNNs operating with asynchronous binary spikes distributed over time, can potentially lead to greater computational efficiency on event-driven hardware. We propose a novel Event-based Video reconstruction framework based on a fully Spiking Neural Network (EVSNN), which utilizes Leaky-Integrate-and-Fire (LIF) neuron and Membrane Potential (MP) neuron. We find that the spiking neurons have the potential to store useful temporal information (memory) to complete such time-dependent tasks. Further-more, to better utilize the temporal information, we propose a hybrid potential-assisted framework (PAEVSNN) using the membrane potential of spiking neuron. The proposed neuron is referred as Adaptive Membrane Potential (AMP) neuron, which adaptively updates the membrane potential according to the input spikes. The experimental results demonstrate that our models achieve comparable performance to ANN-based models on IJRR, MVSEC, and HQF datasets. The energy consumptions of EVSNN and PAEVSNN are$19.36\times$and$7.75\times$more computationally ef-ficient than their ANN architectures, respectively. The code and pretrained model are available at https://sites.google.com/view/evsnn. Lin Zhu 0012, Xiao Wang 0014, Yi Chang 0002, Jianing Li 0001, Tiejun Huang 0001, Yonghong Tian 0001 |
CVPR | 2 |
| 2022 | Pedestrian attribute recognition: A survey
Xiao Wang 0014, Shaofei Zheng, Aihua Zheng, Zhe Chen 0013, Jin Tang 0001, Bin Luo 0001 |
Pattern Recognit. | 1 |
| 2022 | Criteria Comparative Learning for Real-Scene Image Super-ResolutionabstractReal-scene image super-resolution aims to restore real-world low-resolution images into their high-quality versions. A typical RealSR framework usually includes the optimization of multiple criteria which are designed for different image properties, by making the implicit assumption that the ground-truth images can provide a good trade-off between different criteria. However, this assumption could be easily violated in practice due to the inherent contrastive relationship between different image properties. Contrastive learning (CL) provides a promising recipe to relieve this problem by learning discriminative features using the triplet contrastive losses. Though CL has achieved significant success in many computer vision tasks, it is non-trivial to introduce CL to RealSR due to the difficulty in defining valid positive image pairs in this case. Inspired by the observation that the contrastive relationship could also exist between the criteria, in this work, we propose a novel training paradigm for RealSR, named Criteria Comparative Learning (Cria-CL), by developing contrastive losses defined on criteria instead of image patches. In addition, a spatial projector is proposed to obtain a good view for Cria-CL in RealSR. Our experiments demonstrate that compared with the typical weighted regression strategy, our method achieves a significant improvement under similar parameter settings. Yukai Shi, Hao Li 0058, Sen Zhang 0006, Zhijing Yang, Xiao Wang 0014 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Large-Scale Spatio-Temporal Person Re-Identification: Algorithms and BenchmarkabstractPerson re-identification (re-ID) in the scenario with large spatial and temporal spans has not been fully explored. This fact partially occurs because existing benchmark datasets were mainly collected with limited spatial and temporal ranges,e.g.,using videos recorded in a few days by cameras in a specific region of the campus. Such limited spatial and temporal ranges make it hard to simulate the difficulties of person re-ID in real scenarios. In this work, we contribute a novel Large-scale Spatio-Temporal (LaST) person re-ID dataset, including 10,862 identities with more than 228k images. Compared with existing datasets, LaST presents more challenging and high-diversity re-ID settings and significantly larger spatial and temporal ranges. For instance, each person can appear in different cities or countries, and in various time slots from day to evening, and in different seasons from spring to winter. To our best knowledge, LaST is a novel person re-ID dataset with the largest spatio-temporal ranges. Based on LaST, we verified its challenge by conducting a comprehensive performance evaluation of 14 re-ID algorithms. We further propose an easy-to-implement baseline that works well in such challenging re-ID settings. We also verified that models pre-trained on LaST can generalize well on existing datasets with short-term and cloth-changing scenarios. We expect LaST to inspire future works toward more realistic and challenging re-ID tasks. More information about the dataset is available athttps://github.com/shuxjweb/last.git. Xiujun Shu, Xiao Wang 0014, Xianghao Zang, Shiliang Zhang, Yuanqi Chen, Ge Li 0002, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Beyond Greedy Search: Tracking by Multi-Agent Reinforcement Learning-Based Beam SearchabstractTo track the target in a video, current visual trackers usually adopt greedy search for target object localization in each frame, that is, the candidate region with the maximum response score will be selected as the tracking result of each frame. However, we found that this may be not an optimal choice, especially when encountering challenging tracking scenarios such as heavy occlusion and fast motion. In particular, if a tracker drifts, errors will be accumulated and would further make response scores estimated by the tracker unreliable in future frames. To address this issue, we propose to maintain multiple tracking trajectories and apply beam search strategy for visual tracking, so that the trajectory with fewer accumulated errors can be identified. Accordingly, this paper introduces a novel multi-agent reinforcement learning based beam search tracking strategy, termed BeamTracking. It is mainly inspired by the image captioning task, which takes an image as input and generates diverse descriptions using beam search algorithm. Accordingly, we formulate the tracking as a sample selection problem fulfilled by multiple parallel decision-making processes, each of which aims at picking out one sample as their tracking result in each frame. Each maintained trajectory is associated with an agent to perform the decision-making and determine what actions should be taken to update related information. More specifically, using the classification-based tracker as the baseline, we first adopt bi-GRU to encode the target feature, proposal feature, and its response score into a unified state representation. The state feature and greedy search result are then fed into the first agent for independent action selection. Afterwards, the output action and state features are fed into the subsequent agent for diverse results prediction. When all the frames are processed, we select the trajectory with the maximum accumulated score as the tracking result. Extensive experiments on seven popular tracking benchmark datasets validated the effectiveness of the proposed algorithm. Xiao Wang 0014, Zhe Chen 0013, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2022 | Tracking by Joint Local and Global Search: A Target-Aware Attention-Based ApproachabstractTracking-by-detection is a very popular framework for single-object tracking that attempts to search the target object within a local search window for each frame. Although such a local search mechanism works well on simple videos, however, it makes the trackers sensitive to extremely challenging scenarios, such as heavy occlusion and fast motion. In this article, we propose a novel and general target-aware attention mechanism (termed TANet) and integrate it with a tracking-by-detection framework to conduct joint local and global search for robust tracking. Specifically, we extract the features of the target object patch and continuous video frames; then, we concatenate and feed them into a decoder network to generate target-aware global attention maps. More importantly, we resort to adversarial training for better attention prediction. The appearance and motion discriminator networks are designed to ensure its consistency in spatial and temporal views. In the tracking procedure, we integrate target-aware attention with multiple trackers by exploring candidate search regions for robust tracking. Extensive experiments on both short- and long-term tracking benchmark datasets all validated the effectiveness of our algorithm. Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Towards More Flexible and Accurate Object Tracking With Natural Language: Algorithms and BenchmarkabstractTracking by natural language specification is a new rising research topic that aims at locating the target object in the video sequence based on its language description. Compared with traditional bounding box (BBox) based tracking, this setting guides object tracking with high-level semantic information, addresses the ambiguity of BBox, and links local and global search organically together. Those benefits may bring more flexible, robust and accurate tracking performance in practical scenarios. However, existing natural language initialized trackers are developed and compared on benchmark datasets proposed for tracking-by-BBox, which can’t reflect the true power of tracking-by-language. In this work, we propose a new benchmark specifically dedicated to the tracking-by-language, including a large scale dataset, strong and diverse baseline methods. Specifically, we collect 2k video sequences (contains a total of 1,244,340 frames, 663 words) and split 1300/700 for the train/testing respectively. We densely annotate one sentence in English and corresponding bounding boxes of the target object for each video. We also introduce two new challenges into TNL2K for the object tracking task, i.e., adversarial samples and modality switch. A strong baseline method based on an adaptive local-global-search scheme is proposed for future works to compare. We believe this benchmark will greatly boost related researches on natural language guided tracking. Xiao Wang 0014, Xiujun Shu, Bo Jiang 0002, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001 |
CVPR | 1 |
| 2021 | NeuSpike-Net: High Speed Video Reconstruction via Bio-inspired Neuromorphic CamerasabstractNeuromorphic vision sensor is a new bio-inspired imaging paradigm that emerged in recent years, which continuously sensing luminance intensity and firing asynchronous spikes (events) with high temporal resolution. Typically, there are two types of neuromorphic vision sensors, namely dynamic vision sensor (DVS) and spike camera. From the perspective of bio-inspired sampling, DVS only perceives movement by imitating the retinal periphery, while the spike camera was developed to perceive fine textures by simulating the fovea. It is meaningful to explore how to combine two types of neuromorphic cameras to reconstruct high quality image like human vision. In this paper, we propose a NeuSpike-Net to learn both the high dynamic range and high motion sensitivity of DVS and the full texture sampling of spike camera to achieve high-speed and high dynamic image reconstruction. We propose a novel representation to effectively extract the temporal information of spike and event data. By introducing the feature fusion module, the two types of neuromorphic data achieve complementary to each other. The experimental results on the simulated and real datasets demonstrate that the proposed approach is effective to reconstruct high-speed and high dynamic range images via the combination of spike and event data. Lin Zhu 0012, Jianing Li 0001, Xiao Wang 0014, Tiejun Huang 0001, Yonghong Tian 0001 |
ICCV | 3 |
| 2021 | Learn to Match: Automatic Matching Network Design for Visual TrackingabstractSiamese tracking has achieved groundbreaking performance in recent years, where the essence is the efficient matching operator cross-correlation and its variants. Besides the remarkable success, it is important to note that the heuristic matching network design relies heavily on expert experience. Moreover, we experimentally find that one sole matching operator is difficult to guarantee stable tracking in all challenging environments. Thus, in this work, we introduce six novel matching operators from the perspective of feature fusion instead of explicit similarity learning, namely Concatenation, Pointwise-Addition, Pairwise-Relation, FiLM, Simple-Transformer and Transductive-Guidance, to explore more feasibility on matching operator selection. The analyses reveal these operators’ selective adaptability on different environment degradation types, which inspires us to combine them to explore complementary features. To this end, we propose binary channel manipulation (BCM) to search for the optimal combination of these operators. BCM determines to retrain or discard one operator by learning its contribution to other tracking steps. By inserting the learned matching networks to a strong baseline tracker Ocean [47], our model achieves favorable gains by 67.2 → 71.4, 52.6 → 58.3, 70.3 → 76.0 success on OTB100, LaSOT, and TrackingNet, respectively. Notably, Our tracker, dubbed AutoMatch, uses less than half of training data/time than the baseline tracker, and runs at 50 FPS using PyTorch. Code and model are released at https://github.com/JudasDie/SOTS. Yihao Liu 0001, Xiao Wang 0014, Bing Li 0001, Weiming Hu 0004 |
ICCV | 3 |
| 2021 | RGBT tracking via cross-modality message passing
Xiao Wang 0014, Chenglong Li 0002, Jinmin Hu, Jin Tang 0001 |
Neurocomputing | 2 |
| 2021 | Semantic-Guided Pixel Sampling for Cloth-Changing Person Re-IdentificationabstractCloth-changing person re-identification (re-ID) is a new rising research topic that aims at retrieving pedestrians whose clothes are changed. This task is quite challenging and has not been fully studied to date. Current works mainly focus on body shape or contour sketch, but they are not robust enough due to view and posture variations. The key to this task is to exploit cloth-irrelevant cues. This paper proposes a semantic-guided pixel sampling approach for the cloth-changing person re-ID task. We do not explicitly define which feature to extract but force the model to automatically learn cloth-irrelevant cues. Specifically, we firstly recognize the pedestrian's upper clothes and pants, then randomly change them by sampling pixels from other pedestrians. The changed samples retain the identity labels but exchange the pixels of clothes or pants among different pedestrians. Besides, we adopt a loss function to constrain the learned features to keep consistent before and after changes. In this way, the model is forced to learn cues that are irrelevant to upper clothes and pants. We conduct extensive experiments on the latest released PRCC dataset. Our method achieved 65.8% on Rank1 accuracy, which outperforms previous methods with a large margin. The code is available athttps://github.com/shuxjweb/pixel_sampling.git. Xiujun Shu, Ge Li 0002, Xiao Wang 0014, Weijian Ruan, Qi Tian 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Dynamic Attention Guided Multi-Trajectory Analysis for Single Object TrackingabstractMost of the existing single object trackers track the target in a unitary local search window, making them particularly vulnerable to challenging factors such as heavy occlusions and out-of-view movements. Despite the attempts to further incorporate global search, prevailing mechanisms that cooperate local and global search are relatively static, thus are still sub-optimal for improving tracking performance. By further studying the local and global search results, we raise a question: can we allow more dynamics for cooperating both results? In this paper, we propose to introduce more dynamics by devising a dynamic attention-guided multi-trajectory tracking strategy. In particular, we construct dynamic appearance model that contains multiple target templates, each of which provides its own attention for locating the target in the new frame. Guided by different attention, we maintain diversified tracking results for the target to build multi-trajectory tracking history, allowing more candidates to represent the true target trajectory. After spanning the whole sequence, we introduce a multi-trajectory selection network to find the best trajectory that deliver improved tracking performance. Extensive experimental results show that our proposed tracking strategy achieves compelling performance on various large-scale tracking benchmarks. The project page of this paper can be found athttps://sites.google.com/view/mt-track/. Xiao Wang 0014, Zhe Chen 0013, Jin Tang 0001, Bin Luo 0001, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | cmSalGAN: RGB-D Salient Object Detection With Cross-View Generative Adversarial NetworksabstractImage salient object detection (SOD) is an active research topic in computer vision and multimedia area. Fusing complementary information of RGB and depth has been demonstrated to be effective for image salient object detection which is known as RGB-D salient object detection problem. The main challenge for RGB-D salient object detection is how to exploit the salient cues of both intra-modality (RGB, depth) and cross-modality simultaneously which is known as cross-modality detection problem. In this paper, we tackle this challenge by designing a novel cross-modality Saliency Generative Adversarial Network (cmSalGAN). cmSalGAN aims to learn an optimal view-invariant and consistent pixel-level representation for RGB and depth images via a novel adversarial learning framework, which thus incorporates both information of intra-view and correlation information of cross-view images simultaneously for RGB-D saliency detection problem. To further improve the detection results, the attention mechanism and edge detection module are also incorporated into cmSalGAN. The entire cmSalGAN can be trained in an end-to-end manner by using the standard deep neural network framework. Experimental results show that cmSalGAN achieves the new state-of-the-art RGB-D saliency detection performance on several benchmark datasets. Bo Jiang 0002, Zitai Zhou, Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | Multi-modal foreground detection via inter- and intra-modality-consistent low-rank separation
Aihua Zheng, Naipeng Ye, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001 |
Neurocomputing | 4 |
| 2019 | Learning Target-aware Attention for Robust Tracking with Conditional Adversarial Network
Xiao Wang 0014, Bin Luo 0001 |
BMVC | 1 |
| 2019 | Learning Target-Oriented Dual Attention for Robust RGB-T TrackingabstractRGB-Thermal object tracking attempts to locate target object using complementary visual and thermal infrared data. Existing RGB-T trackers fuse different modalities by robust feature representation learning or adaptive modal weighting. However, how to integrate dual attention mechanism for visual tracking is still a subject that has not been studied yet. In this paper, we propose two visual attention mechanisms for robust RGB-T object tracking. Specifically, the local attention is implemented by exploiting the common visual attention of RGB and thermal data to train deep classifiers. We also introduce the global attention, which is a multimodal target-driven attention estimation network. It can provide global proposals for the classifier together with local proposals extracted from previous tracking result. Extensive experiments on two RGB-T benchmark datasets validated the effectiveness of our proposed algorithm. Yabin Zhu, Xiao Wang 0014, Chenglong Li 0002, Jin Tang 0001 |
ICIP | 3 |
| 2019 | Dense Feature Aggregation and Pruning for RGBT TrackingabstractHow to perform effective information fusion of different modalities is a core factor in boosting the performance of RGBT tracking. This paper presents a novel deep fusion algorithm based on the representations from an end-to-end trained convolutional neural network. To deploy the complementarity of features of all layers, we propose a recursive strategy to densely aggregate these features that yield robust representations of target objects in each modality. In different modalities, we propose to prune the densely aggregated features of all modalities in a collaborative way. In a specific, we employ the operations of global average pooling and weighted random selection to perform channel scoring and selection, which could remove redundant and noisy features to achieve more robust feature representation. Experimental results on two RGBT tracking benchmark datasets suggest that our tracker achieves clear state-of-the-art against other RGB and RGBT tracking methods. Yabin Zhu, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001, Xiao Wang 0014 |
ACM Multimedia | 5 |
| 2019 | A Novel Method for Thermal Image Based Electrical-Equipment Detection
Futian Wang, Songjian Hua, Xiao Wang 0014, Zhengzheng Tu, Cheng Zhang 0010, Jin Tang 0001 |
PRCV (1) | 3 |
| 2019 | FMT: fusing multi-task convolutional neural network for person search
Sulan Zhai, Shunqiang Liu, Xiao Wang 0014, Jin Tang 0001 |
Multim. Tools Appl. | 3 |
| 2019 | Quality-aware dual-modal saliency detection via deep reinforcement learning
Xiao Wang 0014, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001 |
Signal Process. Image Commun. | 1 |
| 2018 | SINT++: Robust Visual Tracking via Adversarial Positive Instance GenerationabstractExisting visual trackers are easily disturbed by occlusion, blur and large deformation. We think the performance of existing visual trackers may be limited due to the following issues: i) Adopting the dense sampling strategy to generate positive examples will make them less diverse; ii) The training data with different challenging factors are limited, even through collecting large training dataset. Collecting even larger training dataset is the most intuitive paradigm, but it may still can not cover all situations and the positive samples are still monotonous. In this paper, we propose to generate hard positive samples via adversarial learning for visual tracking. Specifically speaking, we assume the target objects all lie on a manifold, hence, we introduce the positive samples generation network (PSGN) to sampling massive diverse training data through traversing over the constructed target object manifold. The generated diverse target object images can enrich the training dataset and enhance the robustness of visual trackers. To make the tracker more robust to occlusion, we adopt the hard positive transformation network (HPTN) which can generate hard samples for tracking algorithm to recognize. We train this network with deep reinforcement learning to automatically occlude the target object with a negative patch. Based on the generated hard positive samples, we train a Siamese network for visual tracking and our experiments validate the effectiveness of the introduced algorithm. The project page of this paper can be found from the website1. Xiao Wang 0014, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001 |
CVPR | 1 |
| 2018 | Moving object detection via robust background modeling with recurring patterns voting
Chenglong Li 0002, Zhimin Bao, Xiao Wang 0014, Jin Tang 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Deep Co-Space: Sample Mining Across Feature Transformation for Semi-Supervised LearningabstractAiming at improving the performance of visual classification in a cost-effective manner, this paper proposes an incremental semi-supervised learning paradigm called deep co-space (DCS). Unlike many conventional semi-supervised learning methods usually performed within a fixed feature space, our DCS gradually propagates information from labeled samples to unlabeled ones along with deep feature learning. We regard deep feature learning as a series of steps pursuing feature transformation, i.e., projecting the samples from a previous space into a new one, which tends to select the reliable unlabeled samples with respect to this setting. Specifically, for each unlabeled image instance, we measure its reliability by calculating the category variations of feature transformation from two different neighborhood variation perspectives and merged them into a unified sample mining criterion deriving from Hellinger distance. Then, those samples keeping stable correlation to their neighboring samples (i.e., having small category variation in distribution) across the successive feature space transformation are automatically received labels and incorporated into the model for incrementally training in terms of classification. Our extensive experiments on standard image classification benchmarks (e.g., Caltech-256 and SUN-397) demonstrate that the proposed framework is capable of effectively mining from large-scale unlabeled images, which boosts image classification performance and achieves promising results compared with other semi-supervised learning methods. Ziliang Chen 0001, Keze Wang, Xiao Wang 0014, Ebroul Izquierdo, Liang Lin 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Weighted Low-Rank Decomposition for Robust Grayscale-Thermal Foreground DetectionabstractThis paper investigates how to fuse grayscale and thermal video data for detecting foreground objects in challenging scenarios. To this end, we propose an intuitive yet effective method called weighted low-rank decomposition (WELD), which adaptively pursues the cross-modality low-rank representation. Specifically, we form two data matrices by accumulating sequential frames from the grayscale and the thermal videos, respectively. Within these two observing matrices, WELD detects moving foreground pixels as sparse outliers against the low-rank structure background and incorporates the weight variables to make the models of two modalities complementary to each other. The smoothness constraints of object motion are also introduced in WELD to further improve the robustness to noises. For optimization, we propose an iterative algorithm to efficiently solve the low-rank models with three subproblems. Moreover, we utilize an edge-preserving filtering-based method to substantially speed up WELD while preserving its accuracy. To provide a comprehensive evaluation benchmark of grayscale-thermal foreground detection, we create a new data set including 25 aligned grayscale-thermal video pairs with high diversity. Our extensive experiments on both the newly created data set and the public data set OSU3 suggest that WELD achieves superior performance and comparable efficiency against other state-of-the-art approaches. Chenglong Li 0002, Xiao Wang 0014, Lei Zhang 0074, Jin Tang 0001, Hejun Wu, Liang Lin 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Grayscale-Thermal Object Tracking via Multitask Laplacian Sparse RepresentationabstractThis paper studies the problem of object tracking in challenging scenarios by leveraging multimodal visual data. We propose a grayscale-thermal object tracking method in Bayesian filtering framework based on multitask Laplacian sparse representation. Given one bounding box, we extract a set of overlapping local patches within it, and pursue the multitask joint sparse representation for grayscale and thermal modalities. Then, the representation coefficients of the two modalities are concatenated into a vector to represent the feature of the bounding box. Moreover, the similarity between each patch pair is deployed to refine their representation coefficients in the sparse representation, which can be formulated as the Laplacian sparse representation. We also incorporate the modal reliability into the Laplacian sparse representation to achieve an adaptive fusion of different source data. Experiments on two grayscale-thermal datasets suggest that the proposed approach outperforms both grayscale and grayscale-thermal tracking approaches. Chenglong Li 0002, Xiao Wang 0014, Lei Zhang 0074, Jin Tang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |