EDBT 2026 Demo / reviewers in the wild / expert
Xing Wei 0001
dblp:14/4301-1
· DBLP profile ↗
73ranked-venue papers
9as first author
53since 2021 · last 2026
0000-0002-5025-3941ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 6 first-author · 43 since 2021Artificial intelligence and machine learning · 48 · 6 first-author · 34 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Persistent Autoregressive Mapping with Traffic Rules for Autonomous DrivingabstractSafe autonomous driving requires both accurate HD map construction and persistent awareness of traffic rules, even when their associated signs are no longer visible. However, existing methods either focus solely on geometric elements or treat rules as temporary classifications, failing to capture their persistent effectiveness across extended driving sequences. In this paper, we present PAMR (Persistent Autoregressive Mapping with Traffic Rules), a novel framework that performs autoregressive co-construction of lane vectors and traffic rules from visual observations. Our approach introduces two key mechanisms: Map-Rule Co-Construction for processing driving scenes in temporal segments, and Map-Rule Cache for maintaining rule consistency across these segments. To properly evaluate continuous and consistent map generation, we develop MapDRv2, featuring improved lane geometry annotations. Extensive experiments demonstrate that PAMR achieves superior performance in joint vector-rule mapping tasks, while maintaining persistent rule effectiveness throughout extended driving sequences. Shiyi Liang, Xinyuan Chang, Changjie Wu, Huiyuan Yan, Yifan Bai 0001, Yujian Yuan, Shuang Zeng, Mu Xu, Xing Wei 0001 |
AAAI | 11 |
| 2026 | PriorDrive: Enhancing Online HD Mapping with Unified Vector PriorsabstractHigh-Definition Maps (HD maps) are essential for the precise navigation and decision-making of autonomous vehicles, yet their creation and upkeep present significant cost and timeliness challenges. The online construction of HD maps using on-board sensors has emerged as a promising solution; however, these methods can be impeded by incomplete data due to occlusions and inclement weather, while their performance in distant regions remains unsatisfying. This paper proposes PriorDrive to address these limitations by directly harnessing the power of various vectorized prior maps, significantly enhancing the robustness and accuracy of online HD map construction. Our approach integrates a variety of prior maps uniformly, such as OpenStreetMap's Standard Definition Maps (SD maps), outdated HD maps from vendors, and locally constructed maps from historical vehicle data. To effectively integrate such prior information into online mapping models, we introduce a Hybrid Prior Representation (HPQuery) that standardizes the representation of diverse map elements. We further propose a Unified Vector Encoder (UVE), which employs fused prior embedding and a dual encoding mechanism to encode vector data. To improve the UVE's generalizability and performance, we propose a segment-level and point-level pre-training strategy that enables the UVE to learn the prior distribution of vector data. Through extensive testing on the nuScenes, Argoverse 2 and OpenLane-V2, we demonstrate that PriorDrive is highly compatible with various online mapping models and substantially improves map prediction capabilities. The integration of prior maps through PriorDrive offers a robust solution to the challenges of single-perception data, paving the way for more reliable autonomous driving. Shuang Zeng, Xinyuan Chang, Yujian Yuan, Shiyi Liang, Mu Xu, Xing Wei 0001 |
AAAI | 8 |
| 2026 | SCAT: Shared-convolution adaptation tuning
Zelin Yang, Jiashun Chen, Chengshen He, Wencong Zhang, Xing Wei 0001 |
Neurocomputing | 6 |
| 2025 | FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-trainingabstractLanguage-image pre-training faces significant challenges due to limited data in specific formats and the constrained capacities of text encoders. While prevailing methods attempt to address these issues through data augmentation and architecture modifications, they continue to struggle with processing long-form text inputs, and the inherent limitations of traditional CLIP text encoders lead to suboptimal downstream generalization. In this paper, we propose FLAME (Frozen Large lAnguage Models Enable data-efficient language-image pre-training) that leverages frozen large language models as text encoders, naturally processing long text inputs and demonstrating impressive multilingual generalization. FLAME comprises two key components: 1) a multifaceted prompt distillation technique for extracting diverse semantic representations from long captions, which better aligns with the multifaceted nature of images, and 2) a facet-decoupled attention mechanism, complemented by an offline embedding strategy, to ensure efficient computation. Extensive empirical evaluations demonstrate FLAME’s superior performance. When trained on CC3M, FLAME surpasses the previous state-of-the-art by 4.9% in ImageNet top-1 accuracy. On YFCC15M, FLAME surpasses the WIT-400M-trained CLIP by 44.4% in average image-to-text recall@1 across 36 languages, and by 34.6% in text-to-image recall@1 for long-context retrieval on Urban-1k. Code is available at https://github.com/MIV-XJTU/FLAME. Anjia Cao, Xing Wei 0001, Zhiheng Ma |
CVPR | 2 |
| 2025 | Driving by the Rules: A Benchmark for Integrating Traffic Sign Regulations into Vectorized HD MapabstractEnsuring adherence to traffic sign regulations is essential for both human and autonomous vehicle navigation. While current online mapping solutions often prioritize the construction of the geometric and connectivity layers of HD maps, overlooking the construction of the traffic regulation layer within HD maps. Addressing this gap, we introduce MapDR, a novel dataset designed for the extraction of Driving Rules from traffic signs and their association with vectorized, locally perceived HD Maps. MapDR features over 10,000 annotated video clips that capture the intricate correlation between traffic sign regulations and lanes. Built upon this benchmark and the newly defined task of integrating traffic regulations into online HD maps, we provide modular and end-to-end solutions: VLE-MEE and RuleVLM, offering a strong baseline for advancing autonomous driving technology. It fills a critical gap in the integration of traffic sign rules, contributing to the development of reliable autonomous driving systems. Code is available at https://github.com/MIV-XJTU/MapDR. Xinyuan Chang, Maixuan Xue, Xing Wei 0001 |
CVPR | 5 |
| 2025 | Dynamic Integration of Task-Specific Adapters for Class Incremental LearningabstractNon-exemplar Class Incremental Learning (NECIL) enables models to continuously acquire new classes without retraining from scratch and storing old task exemplars, addressing privacy and storage issues. However, the absence of data from earlier tasks exacerbates the challenge of catastrophic forgetting in NECIL. In this paper, we propose a novel framework called Dynamic Integration of task-specific Adapters (DIA), which comprises two key components: Task-Specific Adapter Integration (TSAI) and Patch-Level Model Alignment. TSAI boosts compositionality through a patch-level adapter integration strategy, aggregating richer task-specific information while maintaining low computation costs. Patch-Level Model Alignment maintains feature consistency and accurate decision boundaries via two specialized mechanisms: Patch-Level Distillation Loss (PDL) and Patch-Level Feature Reconstruction (PFR). Specifically, on the one hand, the PDL preserves feature-level consistency between successive models by implementing a distillation loss based on the contributions of patch tokens to new class learning. On the other hand, the PFR promotes classifier alignment by reconstructing old class features from previous tasks that adapt to new task knowledge, thereby preserving well-calibrated decision boundaries. Comprehensive experiments validate the effectiveness of our DIA, revealing significant improvements on NECIL benchmark datasets while maintaining an optimal balance between computational complexity and accuracy. Jiashuo Li, Shaokun Wang, Yuhang He 0001, Xing Wei 0001, Yihong Gong |
CVPR | 5 |
| 2025 | Autoregressive Sequential Pretraining for Visual TrackingabstractRecent advancements in visual object tracking have shifted towards a sequential generation paradigm, where object deformation and motion exhibit strong temporal dependencies. Despite the importance of these dependencies, widely adopted image-level pretrained backbones barely capture the dynamics in the consecutive video, which is the essence of tracking. Thus, we propose AutoRegressive Sequential Pretraining (ARP), an unsupervised spatio-temporal learner, via generating the evolution of object appearance and motion in video sequences. Our method leverages a diffusion model to autoregressively generate the future frame appearance, conditioned on historical embeddings extracted by a general encoder. Furthermore, to ensure trajectory coherence, the same encoder is employed to learn trajectory consistency by generating coordinate sequences in a reverse autoregressive fashion, a process we term backtracking. Further, we integrate the pretrained ARP into AR-TrackV2, creating ARPTrack, which is further fine-tuned for tracking tasks. ARPTrack achieves state-of-the-art performance across multiple benchmarks, becoming the first tracker to surpass 80% AO on GOT-10k, while maintaining high efficiency. These results demonstrate the effectiveness of our approach in capturing temporal dependencies for continuous video tracking. The code will be released soon. Shiyi Liang, Yifan Bai 0001, Yihong Gong, Xing Wei 0001 |
CVPR | 4 |
| 2025 | SCAT: Shared-Convolution Adaptation Tuning for Foreground SegmentationabstractFine-tuning a minimal subset of parameters in large well-trained models has emerged as a popular paradigm for transforming prior knowledge to address downstream tasks in computer vision. Although it has shown promising performance in certain vision tasks such as classification, parameter-efficient tuning remains in its infancy and suffers a significant accuracy drop compared to tuning the entire model particularly in field of segmentation. In this paper, we propose a novel tuning method named SCAT (Shared-Convolution Adaptation Timing), designed to adapt segmentation models to various fine-grained foreground segmentation vision tasks. By injecting strong inductive bias prompts with shared convolutional features into the frozen backbones, SCAT significantly increases the transferability of the pre-trained models with only a few learnable parameters. SCAT delivers superior performance compared to other state-of- the-art fine-tuning methods, domain-specific hand-crafted networks, and even the fully-tuning paradigm across numerous foreground segmentation scenarios. We release our source code at: https://github.com/KevinLi2023/SCAT. Dezheng Gao, Zelin Yang, Xing Wei 0001 |
ICASSP | 4 |
| 2025 | SeqGrowGraph: Learning Lane Topology as a Chain of Graph ExpansionsabstractAccurate lane topology is essential for autonomous driving, yet traditional methods struggle to model the complex, non-linear structures-such as loops and bidirectional lanes-prevalent in real-world road structure. We present SeqGrowGraph, a novel framework that learns lane topology as a chain of graph expansions, inspired by human map-drawing processes. Representing the lane graph as a directed graph $G=(V,E)$, with intersections ($V$) and centerlines ($E$), SeqGrowGraph incrementally constructs this graph by introducing one vertex at a time. At each step, an adjacency matrix ($A$) expands from $n \times n$ to $(n+1) \times (n+1)$ to encode connectivity, while a geometric matrix ($M$) captures centerline shapes as quadratic Bézier curves. The graph is serialized into sequences, enabling a transformer model to autoregressively predict the chain of expansions, guided by a depth-first search ordering. Evaluated on nuScenes and Argoverse 2 datasets, SeqGrowGraph achieves state-of-the-art performance. Mengwei Xie, Shuang Zeng, Xinyuan Chang, Mu Xu, Xing Wei 0001 |
ICCV | 7 |
| 2025 | Positional Prompt Tuning for Efficient 3D Representation Learning
Shaochen Zhang, Zekun Qi, Runpei Dong, Xiuxiu Bai, Xing Wei 0001 |
ACM Multimedia | 5 |
| 2025 | FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous DrivingabstractVision–Language–Action (VLA) models are increasingly used for end-to-end driving due to their world knowledge and reasoning ability. Most prior work, however, inserts textual chains-of-thought (CoT) as intermediate steps tailored to the current scene. Such symbolic compressions can blur spatio-temporal relations and discard fine visual cues, creating a cross-modal gap between perception and planning.
We propose FSDrive, a visual spatio-temporal CoT framework that enables VLAs to think in images. The model first acts as a world model to generate a unified future frame that overlays coarse but physically-plausible priors—future lane dividers and 3D boxes—on the predicted future image. This unified frame serves as the visual CoT, capturing both spatial structure and temporal evolution. The same VLA then functions as an inverse-dynamics model, planning trajectories from current observations and the visual CoT.
To equip VLAs with image generation while preserving understanding, we introduce a unified pre-training paradigm that expands the vocabulary to include visual tokens and jointly optimizes VQA (for semantics) and future-frame prediction (for dynamics). A progressive easy-to-hard scheme first predicts lane/box priors to enforce physical constraints, then completes full future frames for fine details.
On nuScenes and NAVSIM, FSDrive improves trajectory accuracy and reduces collisions under both ST-P3 and UniAD metrics, and attains competitive FID for future-frame generation despite using lightweight autoregression. It also advances scene understanding on DriveLM. Together, these results indicate that visual CoT narrows the cross-modal gap and yields safer, more anticipatory planning.
Code is available at https://github.com/MIV-XJTU/FSDrive. Shuang Zeng, Xinyuan Chang, Mengwei Xie, Yifan Bai 0001, Mu Xu, Xing Wei 0001 |
NeurIPS | 8 |
| 2025 | Adaptive knowledge transfer for data-free low-bit quantization via tiered collaborative learning
Zelin Yang, Xing Wei 0001 |
Neurocomputing | 6 |
| 2025 | Curriculum Dataset DistillationabstractMost dataset distillation methods struggle to accommodate large-scale datasets due to their substantial computational and memory requirements. Recent research has begun to explore scalable disentanglement methods. However, there are still performance bottlenecks and room for optimization in this direction. In this paper, we present a curriculum-based dataset distillation framework aiming to harmonize performance and scalability. This framework strategically distills synthetic images, adhering to a curriculum that transitions from simple to complex. By incorporating curriculum evaluation, we address the issue of previous methods generating images that tend to be homogeneous and simplistic, doing so at a manageable computational cost. Furthermore, we introduce adversarial optimization towards synthetic images to further improve their representativeness and safeguard against their overfitting to the neural network involved in distilling. This enhances the generalization capability of the distilled images across various neural network architectures and also increases their robustness to noise. Extensive experiments demonstrate that our framework sets new benchmarks in large-scale dataset distillation, achieving substantial improvements of 11.1% on Tiny-ImageNet, 9.0% on ImageNet-1K, and 7.3% on ImageNet-21K. Our distilled datasets and code are available at https://github.com/MIV-XJTU/CUDD. Zhiheng Ma, Anjia Cao, Funing Yang, Yihong Gong, Xing Wei 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Evolving Parameterized Prompt Memory for Continual LearningabstractRecent studies have demonstrated the potency of leveraging prompts in Transformers for continual learning (CL). Nevertheless, employing a discrete key-prompt bottleneck can lead to selection mismatches and inappropriate prompt associations during testing. Furthermore, this approach hinders adaptive prompting due to the lack of shareability among nearly identical instances at more granular level. To address these challenges, we introduce the Evolving Parameterized Prompt Memory (EvoPrompt), a novel method involving adaptive and continuous prompting attached to pre-trained Vision Transformer (ViT), conditioned on specific instance. We formulate a continuous prompt function as a neural bottleneck and encode the collection of prompts on network weights. We establish a paired prompt memory system consisting of a stable reference and a flexible working prompt memory. Inspired by linear mode connectivity, we progressively fuse the working prompt memory and reference prompt memory during inter-task periods, resulting in continually evolved prompt memory. This fusion involves aligning functionally equivalent prompts using optimal transport and aggregating them in parameter space with an adjustable bias based on prompt node attribution. Additionally, to enhance backward compatibility, we propose compositional classifier initialization, which leverages prior prototypes from pre-trained models to guide the initialization of new classifiers in a subspace-aware manner. Comprehensive experiments validate that our approach achieves state-of-the-art performance in both class and domain incremental learning scenarios. Muhammad Rifki Kurniawan, Xiang Song 0005, Zhiheng Ma, Yuhang He 0001, Yihong Gong, Xing Wei 0001 |
AAAI | 7 |
| 2024 | ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to DescribeabstractWe present ARTrackV2, which integrates two pivotal aspects of tracking: determining where to look (localization) and how to describe (appearance analysis) the target object across video frames. Building on the foundation of its predecessor, ARTrackV2 extends the concept by introducing a unified generative framework to “read out” object's trajectory and “retell” its appearance in an autoregressive manner. This approach fosters a time-continuous methodology that models the joint evolution of motion and visual features, guided by previous estimates. Furthermore, ARTrackV2 stands out for its efficiency and simplicity, obviating the less efficient intra-frame autoregression and hand-tuned parameters for appearance updates. Despite its simplicity, ARTrackV2 achieves state-of-the-art performance on prevailing benchmark datasets while demonstrating a remarkable efficiency improvement. In particular, ARTrackV2 achieves an AO score of 79. 5% on GOT-10k and an AUC of 86. 1% on TrackingNet while being 3.6× faster than ARTrack. Yifan Bai 0001, Zeyang Zhao, Yihong Gong, Xing Wei 0001 |
CVPR | 4 |
| 2024 | DYSON: Dynamic Feature Space Self-Organization for Online Task-Free Class Incremental LearningabstractIn this paper, we focus on a challenging Online Task-Free Class Incremental Learning (OTFCIL) problem. Dif-ferent from the existing methods that continuously learn the feature space from data streams, we propose a novel compute-and-align paradigm for the OTFCIL. It first com-putes an optimal geometry, i.e., the class prototype distri-bution, for classifying existing classes and updates it when new classes emerge, and then trains a DNN model by aligning its feature space to the optimal geometry. To this end, we develop a novel Dynamic Neural Collapse (DNC) algorithm to compute and update the optimal geometry. The DNC ex-pands the geometry when new classes emerge without loss of the geometry optimality and guarantees the drift distance of old class prototypes with an explicit upper bound. On this basis, we propose a novel DYnamic feature space Self-OrganizatioN (DYSON) method containing three ma-jor components, including 1) a feature extractor, 2) a Dy-namic Feature-Geometry Alignment (DFGA) module aligning the feature space to the optimal geometry computed by DNC and 3) a training-free class-incremental classifier de-rived from the DNC geometry. Experimental comparison results on four benchmark datasets, including CIFAR10, CI-FAR100, CUB200, and CoRe50, demonstrate the efficiency and superiority of the DYSON method. The source code is released at https://github.com/isCDX2IDYSON. Yuhang He 0001, Yuhan Jin, Songlin Dong, Xing Wei 0001, Yihong Gong |
CVPR | 5 |
| 2024 | Region-Aware Sequence-to-Sequence Learning for Hyperspectral Denoising
Jiahua Xiao, Yang Liu 0385, Xing Wei 0001 |
ECCV (60) | 3 |
| 2024 | Projecting Points to Axes: Oriented Object Detection via Point-Axis Representation
Zeyang Zhao, Qilong Xue, Yuhang He 0001, Yifan Bai 0001, Xing Wei 0001, Yihong Gong |
ECCV (28) | 5 |
| 2024 | MultiQ: Multi-model Joint Learning via Synthetic Data for Data-Free QuantizationabstractData-Free quantization is a technique that eliminates reliance on the original data during the quantization process. However, existing methods suffer from low data-fit accuracy and inefficient quantization. To tackle this problem, we propose a method called MultiQ: Multi-model Joint Learning via Synthetic Data for Data-Free Quantization. This method is based on a generative approach and offers several advantages such as high precision, low training time cost, and strong portability. In the quantization process, multiple models with different bitwidths can be trained simultaneously. Our method has been extensively evaluated, particularly on the ImageNet1K dataset using MobileNetV1, and has shown a significant improvement of 7.15%. These results demonstrate the effectiveness of our proposed method in achieving data-free quantization efficiently. Xing Wei 0001, Huazheng Zhao |
ICME | 2 |
| 2024 | Bridging Fourier and Spatial-Spectral Domains for Hyperspectral Image DenoisingabstractRemarkable progresses have been made in hyperspectral image (HSI) denoising. However, the majority of existing methods are predominantly confined to the spatial-spectral domain, overlooking the untapped potential inherent in the Fourier domain. This paper presents a novel approach to address HSI denoising by bridging the information from the Fourier and spatial-spectral domains. Our method highlights key insights into the Fourier properties within spatial and spectral domains through the Fourier transform. Specifically, we note that the amplitude inherently embody noise and photon reflection characteristics, while the phase holds structural information. These insights unveil new perspectives on the physical properties of HSIs, motivating us to leverage complementary information exchange between Fourier and spatial-spectral domains. To this end, we introduce the Fourier-prior Integration Denoising Network (FIDNet), a potent yet straightforward approach that utilizes Fourier insights to synergistically interact with spatial-spectral domains for superior HSI denoising. In FIDNet, we independently extract spatial and Fourier features through dual branches and merge these representations to enhance spectral evolution modeling through the inherent structure consistency constraints and continuing reflection variation revealed in Fourier prior. Our proposed method demonstrates robust generalization across synthetic and real-world benchmark datasets, achieves comparable results with state-of-the-art methods in both quantitative quality and visual results. The code is available at https://github.com/MIV-XJTU/FIDNet. Jiahua Xiao, Yang Liu 0385, Shizhou Zhang, Xing Wei 0001 |
ACM Multimedia | 4 |
| 2024 | Overcoming Catastrophic Forgetting for Multi-Label Class-Incremental LearningabstractDespite the recent progress of class-incremental learning (CIL) methods, their capabilities in real-world scenarios such as multi-label settings remain unexplored. This paper focuses on a more practical CIL problem named multi-label class-incremental learning (MLCIL). MLCIL requires the vision models to overcome catastrophic forgetting of old knowledge while learning new classes from multi-label samples. Direct application of existing CIL methods to MLCIL leads to label absence, representative sample selection, and feature dilution problems. To address these problems, we present a novel AdaPtive Pseudo-Label-drivEn (APPLE) framework consisting of three components. First, the adaptive pseudo-label strategy is proposed to solve the label absence problem, which leverages the old model to annotate old classes for new samples. Second, a cluster sampling strategy is proposed to obtain more diverse samples to alleviate catastrophic forgetting under the MLCIL setting better. Finally, a class attention decoder is designed to mitigate the object feature dilution problem in multi-label samples. The extensive experiments on PASCAL VOC 2007 and MS-COCO demonstrate that our proposed method significantly outperforms other representative state-of-the-art CIL methods. Xiang Song 0005, Kuang Shu, Songlin Dong, Xing Wei 0001, Yihong Gong |
WACV | 5 |
| 2024 | Few-shot online anomaly detection and segmentation
Shenxing Wei, Xing Wei 0001, Zhiheng Ma, Songlin Dong, Shaochen Zhang, Yihong Gong |
Knowl. Based Syst. | 2 |
| 2024 | CONet: Crowd and occlusion-aware network for occluded human pose estimation
Xiuxiu Bai, Xing Wei 0001, Zengying Wang, Miao Zhang 0013 |
Neural Networks | 2 |
| 2024 | Global self-sustaining and local inheritance for source-free unsupervised domain adaptation
Lin Peng 0003, Yuhang He 0001, Shaokun Wang, Xiang Song 0005, Songlin Dong, Xing Wei 0001, Yihong Gong |
Pattern Recognit. | 6 |
| 2024 | Domain Incremental Object Detection Based on Feature Space Topology Preserving StrategyabstractObject detection with the capacity to incrementally adapt to new domains is a crucial yet relatively under-explored research topic. The catastrophic forgetting problem presents a significant challenge to achieve this goal, where the model’s performance improves quickly in new conditions but deteriorates sharply in old ones after several incremental learning sessions. Drawing on recent discoveries in visual memories of the human brain, we introduce the Topology-Preserving Domain Incremental Object Detection (TP-DIOD) approach, which aims to address the catastrophic forgetting problem by extracting the topological structure of the feature space learned by the Convolutional Neural Network (CNN) model and preserving this topology during the subsequent incremental learning sessions. Specifically, we model the feature space topology using the self-organizing map (SOM) and construct an anchor image set based on the centroid vectors of the SOM nodes to memorize the feature space topology. We then develop the anchor loss function to penalize the topological changes of the feature space during the subsequent incremental learning sessions. Experimental evaluations on two sets of datasets demonstrate the effectiveness of the proposed TP-DIOD method in mitigating the catastrophic forgetting problem and achieving high accuracy on both old and new domain datasets. Xiang Song 0005, Yuhang He 0001, Changxin Wang, Songlin Dong, Xing Wei 0001, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Knowledge Synergy Learning for Multi-Modal TrackingabstractBenefiting from the rich information provided by different modalities, multi-modal tracking has shown significant improvements compared to single-modal tracking. However, in practical applications, multi-modal tracking still faces two major challenges. Firstly, it is crucial to effectively integrate the complementary information from different modalities in order to improve tracking performance. Secondly, as trackers are often deployed in dynamic environments, it is difficult to ensure complete multi-modal data. Thus, handling modal-missing issues is essential to achieve robust and reliable tracking. To address these challenges, this paper proposes a Knowledge Synergy Network (KSNet) that integrates multi-modal features into a comprehensive representation and incorporates a modal compensation mechanism to handle modal-missing issues. With this framework, a multi-modal tracker (KSTrack) is built and trained using multi-modal data. KSTrack is capable of handling both complete and incomplete multi-modal data during inference. Comprehensive experiments on four large-scale RGB-Thermal (RGB-T) and RGB-Depth (RGB-D) benchmarks show that KSTrack surpasses state-of-the-art multi-modal trackers when using multi-modal data and outperforms single-modal trackers by a large margin when using single-modal data. Yuhang He 0001, Zhiheng Ma, Xing Wei 0001, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Analogical Learning-Based Few-Shot Class-Incremental LearningabstractFSCIL (Few-shot class-incremental learning) is a prominent research topic in the ML community. It faces two significant challenges: forgetting old class knowledge and overfitting to limited new class training examples. In this paper, we present a novel FSCIL approach inspired by the human brain’s analogical learning mechanism, which enables human beings to form knowledge about a target domain from the knowledge of the source domains that are analogical to the target in some aspects. The proposed analogical learning-based FSCIL (ALFSCIL) method consists of two major components: new class classifier constructor (NCCC) and Meta-Analogical training (MAT). The NCCC module utilizes a multi-head cross-attention transformer to compute analogies between new and old classes, generating new class classifiers by blending old class classifiers based on the computed analogies. The MAT module updates the parameters of the CNN feature extractor, the NCCC module, and the knowledge for each encountered class after each round of the FSCIL session. We turn the optimization process into a bi-level optimization problem(BOP) whose theoretical analysis proves the stability and plasticity of our proposed model. Experimental evaluations reveal that this proposed ALFSCIL method achieves the SOTA performance accuracies on three benchmark datasets: CIFAR100, miniImageNet, and CUB200. Jiashuo Li, Songlin Dong, Yihong Gong, Yuhang He 0001, Xing Wei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Brain Cognition-Inspired Dual-Pathway CNN Architecture for Image ClassificationabstractInspired by the global-local information processing mechanism in the human visual system, we propose a novel convolutional neural network (CNN) architecture named cognition-inspired network (CogNet) that consists of a global pathway, a local pathway, and a top-down modulator. We first use a common CNN block to form the local pathway that aims to extract fine local features of the input image. Then, we use a transformer encoder to form the global pathway to capture global structural and contextual information among local parts in the input image. Finally, we construct the learnable top-down modulator where fine local features of the local pathway are modulated by global representations of the global pathway. For ease of use, we encapsulate the dual-pathway computation and modulation process into a building block, called the global-local block (GL block), and a CogNet of any depth can be constructed by stacking a necessary number of GL blocks one after another. Extensive experimental evaluations have revealed that the proposed CogNets have achieved the state-of-the-art performance accuracies on all the six benchmark datasets and are very effective for overcoming the "texture bias" and the "semantic confusion" problems faced by many CNN models. Songlin Dong, Yihong Gong, Jingang Shi, Miao Shang, Xing Wei 0001, Xiaopeng Hong, Tiangang Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | DKT: Diverse Knowledge Transfer Transformer for Class Incremental LearningabstractIn the context of incremental class learning, deep neural networks are prone to catastrophic forgetting, where the accuracy of old classes declines substantially as new knowledge is learned. While recent studies have sought to address this issue, most approaches suffer from either the stability-plasticity dilemma or excessive computational and parameter requirements. To tackle these challenges, we propose a novel framework, the Diverse Knowledge Transfer Transformer (DKT), which incorporates two knowledge transfer mechanisms that use attention mechanisms to transfer both task-specific and task-general knowledge to the current task, along with a duplex classifier to address the stability-plasticity dilemma. Additionally, we design a loss function that clusters similar categories and discriminates between old and new tasks in the feature space. The proposed method requires only a small number of extra parameters, which are negligible in comparison to the increasing number of tasks. We perform extensive experiments on CIFAR100, ImageNet100, and ImageNet1000 datasets, which demonstrate that our method outperforms other competitive methods and achieves state-of-the-art performance. Our source code is available at https://github.com/MIVXJTU/DKT. Xinyuan Gao, Yuhang He 0001, Songlin Dong, Xing Wei 0001, Yihong Gong |
CVPR | 5 |
| 2023 | Autoregressive Visual TrackingabstractWe present ARTrack, an autoregressive framework for visual object tracking. ARTrack tackles tracking as a coordinate sequence interpretation task that estimates object trajectories progressively, where the current estimate is induced by previous states and in turn affects subsequences. This time-autoregressive approach models the sequential evolution of trajectories to keep tracing the object across frames, making it superior to existing template matching based trackers that only consider the per-frame localization accuracy. ARTrack is simple and direct, eliminating customized localization heads and post-processings. Despite its simplicity, ARTrack achieves state-of-the-art performance on prevailing benchmark datasets. Source code is available at https://github.com/MIV-XJTU/ARTrack. Xing Wei 0001, Yifan Bai 0001, Yongchao Zheng, Dahu Shi, Yihong Gong |
CVPR | 1 |
| 2023 | Knowledge Restore and Transfer for Multi-Label Class-Incremental LearningabstractCurrent class-incremental learning research mainly focuses on single-label classification tasks while multi-label class-incremental learning (MLCIL) with more practical application scenarios is rarely studied. Although there have been many anti-forgetting methods to solve the problem of catastrophic forgetting in single-label class-incremental learning, these methods have difficulty in solving the MLCIL problem due to label absence and information dilution problems. To solve these problems, we propose a Knowledge Restore and Transfer (KRT) framework containing two key components. First, a dynamic pseudo-label (DPL) module is proposed to solve the label absence problem by restoring the knowledge of old classes to the new data. Second, an incremental cross-attention (ICA) module is designed to maintain and transfer the old knowledge to solve the information dilution problem. Comprehensive experimental results on MS-COCO and PASCAL VOC datasets demonstrate the effectiveness of our method for improving recognition performance and mitigating forgetting on multi-label class-incremental learning tasks. The source code is available at https://gith.ub.com/witdsl/KRT-MLCIL. Songlin Dong, Haoyu Luo, Yuhang He 0001, Xing Wei 0001, Yihong Gong |
ICCV | 4 |
| 2023 | Learning Symmetry-Aware Geometry Correspondences for 6D Object Pose EstimationabstractCurrent 6D pose estimation methods focus on handling objects that are previously trained, which limits their applications in real dynamic world. To this end, we propose a geometry correspondence-based framework, termed GCPose, to estimate 6D pose of arbitrary unseen objects without any re-training. Specifically, the proposed method draws the idea from point cloud registration and resorts to object-agnostic geometry features to establish the 3D-3D correspondences between the object-scene point cloud and object-model point cloud. Then the 6D pose parameters are solved by a least-squares fitting algorithm. Taking the symmetry properties of objects into consideration, we design a symmetry-aware matching loss to facilitate the learning of dense point-wise geometry features and improve the performance considerably. Moreover, we introduce an online training data generation with special data augmentation and normalization to empower the network to learn diverse geometry prior. With training on synthetic objects from ShapeNet, our method outperforms previous approaches for unseen object pose estimation by a large margin on T-LESS, LINEMOD, Occluded-LINEMOD, and TUD-L datasets. Code is available at https://github.com/hikvision-research/GCPose. Shenxing Wei, Dahu Shi, Wenming Tan, Zheyang Li, Ye Ren, Xing Wei 0001, Yi Yang 0001, Shiliang Pu |
ICCV | 7 |
| 2023 | Atten-Adapter: A Unified Attention-Based Adapter for Efficient TuningabstractRecently, more and more large pre-trained models have emerged. Several parameter-efficient tuning methods have been studied to transfer the prior knowledge of the pre-trained models to specific downstream tasks and achieve promising results. This paper proposes a simple yet effective method called Atten-Adapter. To the best of our knowledge, this is the first work that utilizes attention with learnable parameters as the internal structure of the adapter in the field of fine-tuning. The attention-based adapter can provide better information fusion ability and pay more attention to the global features compared to the MLP-based adapter. As a plug-and-play module, Atten-Adapter can be easily adapted to different types of vision models such as ConvNets and Transformer architectures in different tasks like classification and segmentation. Moreover, we demonstrate the generality of our proposed adapters by conducting experiments on language models. With small amounts of tunable parameters, our method achieves significant improvements compared to the previous state-of-the-art methods. Wenzhe Gu, Maixuan Xue, Jiahua Xiao, Dahu Shi, Xing Wei 0001 |
ICIP | 6 |
| 2023 | Hyperspectral Image Denoising Using Uncertainty-Aware AdjustorabstractHyperspectral image (HSI) denoising has achieved promising results with the development of deep learning. A mainstream class of methods exploits the spatial-spectral correlations and recovers each band with the aids of neighboring bands, collectively referred to as spectral auxiliary networks. However, these methods treat entire adjacent spectral bands equally. In theory, clearer and nearer bands tend to contain more reliable spectral information than noisier and farther ones with higher uncertainties. How to achieve spectral enhancement and adaptation of each adjacent band has become an urgent problem in HSI denoising. This work presents the UA-Adjustor, a comprehensive adjustor that enhances denoising performance by considering both the band-to-pixel and enhancement-to-adjustment aspects. Specifically, UA-Adjustor consists of three stages that evaluate the importance of neighboring bands, enhance neighboring bands based on uncertainty perception, and adjust the weight of spatial pixels in adjacent bands through estimated uncertainty. For its simplicity, UA-Adjustor can be flexibly plugged into existing spectral auxiliary networks to improve denoising behavior at low cost. Extensive experimental results validate that the proposed solution can improve over recent state-of-the-art (SOTA) methods on both simulated and real-world benchmarks by a large margin. Jiahua Xiao, Xing Wei 0001 |
IJCAI | 2 |
| 2023 | Hyperspectral Image Denoising with Spectrum AlignmentabstractSpectral modeling plays a critical role in denoising hyperspectral images (HSIs), with recent approaches leveraging well-designed network architectures to extract spectral contexts for noise removal. However, these approaches overlook a striking finding: the presence of spectral differences in noisy contexts can pose challenges for the denoising network during the restoration process of each band in the HSI. We attribute this to the varying levels of spectral difference between different bands and the unknown distribution of various noises. These factors can make it difficult for the network to capture consistent features, ultimately leading to suboptimal solutions. We propose a novel concept termed 'spectral displacement,' which views spectral differences as pixel motion displacement along the spectral domain. To eliminate the effect of spectral displacement, we introduce a potential solution: spectral alignment. This approach can increase the mutual information between different spectral bands and enhance the effectiveness of denoising. We then present the Spectral Alignment Recurrent Network (SARN) for efficient and effective displacement estimation and pixel-level alignment between neighboring bands. SARN can serve as a general plug-in for HSI backbones without requiring any model-specific design. Experimental results on several benchmark datasets confirm the effectiveness and superiority of our concept and network. The source code will be available at https://github.com/MIV-XJTU/SARN. Jiahua Xiao, Yantao Ji, Xing Wei 0001 |
ACM Multimedia | 3 |
| 2023 | Sparse Parameterization for Epitomic Dataset DistillationabstractThe success of deep learning relies heavily on large and diverse datasets, but the storage, preprocessing, and training of such data present significant challenges. To address these challenges, dataset distillation techniques have been proposed to obtain smaller synthetic datasets that capture the essential information of the originals. In this paper, we introduce a Sparse Parameterization for Epitomic datasEt Distillation (SPEED) framework, which leverages the concept of dictionary learning and sparse coding to distill epitomes that represent pivotal information of the dataset. SPEED prioritizes proper parameterization of the synthetic dataset and introduces techniques to capture spatial redundancy within and between synthetic images. We propose Spatial-Agnostic Epitomic Tokens (SAETs) and Sparse Coding Matrices (SCMs) to efficiently represent and select significant features. Additionally, we build a Feature-Recurrent Network (FReeNet) to generate hierarchical features with high compression and storage efficiency. Experimental results demonstrate the superiority of SPEED in handling high-resolution datasets, achieving state-of-the-art performance on multiple benchmarks and downstream applications. Our framework is compatible with a variety of dataset matching approaches, generally enhancing their performance. This work highlights the importance of proper parameterization in epitomic dataset distillation and opens avenues for efficient representation learning. Source code is available at https://github.com/MIV-XJTU/SPEED. Xing Wei 0001, Anjia Cao, Funing Yang, Zhiheng Ma |
NeurIPS | 1 |
| 2023 | Topology-preserving transfer learning for weakly-supervised anomaly detection and segmentation
Shenxing Wei, Xing Wei 0001, Muhammad Rifki Kurniawan, Zhiheng Ma, Yihong Gong |
Pattern Recognit. Lett. | 2 |
| 2023 | Semi-Supervised Crowd Counting via Multiple Representation LearningabstractThere has been a growing interest in counting crowds through computer vision and machine learning techniques in recent years. Despite that significant progress has been made, most existing methods heavily rely on fully-supervised learning and require a lot of labeled data. To alleviate the reliance, we focus on the semi-supervised learning paradigm. Usually, crowd counting is converted to a density estimation problem. The model is trained to predict a density map and obtains the total count by accumulating densities over all the locations. In particular, we find that there could be multiple density map representations for a given image in a way that they differ in probability distribution forms but reach a consensus on their total counts. Therefore, we propose multiple representation learning to train several models. Each model focuses on a specific density representation and utilizes the count consistency between models to supervise unlabeled data. To bypass the explicit density regression problem, which makes a strong parametric assumption on the underlying density distribution, we propose an implicit density representation method based on the kernel mean embedding. Extensive experiments demonstrate that our approach outperforms state-of-the-art semi-supervised methods significantly. Xing Wei 0001, Yunfeng Qiu, Zhiheng Ma, Xiaopeng Hong, Yihong Gong |
IEEE Trans. Image Process. | 1 |
| 2022 | SOIT: Segmenting Objects with Instance-Aware TransformersabstractThis paper presents an end-to-end instance segmentation framework, termed SOIT, that Segments Objects with Instance-aware Transformers. Inspired by DETR, our method views instance segmentation as a direct set prediction problem and effectively removes the need for many hand-crafted components like RoI cropping, one-to-many label assignment, and non-maximum suppression (NMS). In SOIT, multiple queries are learned to directly reason a set of object embeddings of semantic category, bounding-box location, and pixel-wise mask in parallel under the global image context. The class and bounding-box can be easily embedded by a fixed-length vector. The pixel-wise mask, especially, is embedded by a group of parameters to construct a lightweight instance-aware transformer. Afterward, a full-resolution mask is produced by the instance-aware transformer without involving any RoI-based operation. Overall, SOIT introduces a simple single-stage instance segmentation framework that is both RoI- and NMS-free. Experimental results on the MS COCO dataset demonstrate that SOIT outperforms state-of-the-art instance segmentation approaches significantly. Moreover, the joint learning of multiple tasks in a unified query embedding can also substantially improve the detection performance. Code is available at https://github.com/yuxiaodongHRI/SOIT. Dahu Shi, Xing Wei 0001, Ye Ren, Tingqun Ye, Wenming Tan |
AAAI | 3 |
| 2022 | End-to-End Multi-Person Pose Estimation with TransformersabstractCurrent methods of multi-person pose estimation typically treat the localization and association of body joints separately. In this paper, we propose the first fully end-to-end multi-person Pose Estimation framework with TRansformers, termed PETR. Our method views pose estimation as a hierarchical set prediction problem and effectively removes the need for many hand-crafted modules like RoI cropping, NMS and grouping post-processing. In PETR, multiple pose queries are learned to directly reason a set of full-body poses. Then a joint decoder is utilized to further refine the poses by exploring the kinematic relations between body joints. With the attention mechanism, the proposed method is able to adaptively attend to the features most relevant to target keypoints, which largely overcomes the feature misalignment difficulty in pose estimation and improves the performance considerably. Extensive experiments on the MS COCO and CrowdPose benchmarks show that PETR plays favorably against state-of-the-art approaches in terms of both accuracy and efficiency. The code and models are available at https://github.com/hikvision-research/opera. Dahu Shi, Xing Wei 0001, Liangqi Li, Ye Ren, Wenming Tan |
CVPR | 2 |
| 2022 | Deep Dynamic Scene Deblurring From Optical FlowabstractDeblurring can not only provide visually more pleasant pictures and make photography more convenient, but also can improve the performance of objection detection as well as tracking. However, removing dynamic scene blur from images is a non-trivial task as it is difficult to model the non-uniform blur mathematically. Several methods first use single or multiple images to estimate optical flow (which is treated as an approximation of blur kernels) and then adopt non-blind deblurring algorithms to reconstruct the sharp images. However, these methods cannot be trained in an end-to-end manner and are usually computationally expensive. In this paper, we explore optical flow to remove dynamic scene blur by using the multi-scale spatially variant recurrent neural network (RNN). We utilize FlowNets to estimate optical flow from two consecutive images in different scales. The estimated optical flow provides the RNN weights in different scales so that the weights can better help RNNs to remove blur in the feature spaces. Finally, we develop a convolutional neural network (CNN) to restore the sharp images from the deblurred features. Both quantitatively and qualitatively evaluations on the benchmark datasets demonstrate that the proposed method performs favorably against state-of-the-art algorithms in terms of accuracy, speed, and model size. Jiawei Zhang 0002, Jinshan Pan, Daoye Wang, Shangchen Zhou, Xing Wei 0001, Furong Zhao, Jimmy S. J. Ren |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Identity-Quantity Harmonic Multi-Object TrackingabstractThe data association problem of multi-object tracking (MOT) aims to assign IDentity (ID) labels to detections and infer a complete trajectory for each target. Most existing methods assume that each detection corresponds to a unique target and thus cannot handle situations when multiple targets occur in a single detection due to detection failure in crowded scenes. To relax this strong assumption for practical applications, we formulate the MOT as a Maximizing An Identity-Quantity Posterior (MAIQP) problem on the basis of associating each detection with an identity and a quantity characteristic and then provide solutions to tackle two key problems arising. Firstly, a local target quantification module is introduced to count the number of targets within one detection. Secondly, we propose an identity-quantity harmony mechanism to reconcile the two characteristics. On this basis, we develop a novel Identity-Quantity HArmonic Tracking (IQHAT) framework that allows assigning multiple ID labels to detections containing several targets. Through extensive experimental evaluations on five benchmark datasets, we demonstrate the superiority of the proposed method. Yuhang He 0001, Xing Wei 0001, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
IEEE Trans. Image Process. | 2 |
| 2022 | ECCNAS: Efficient Crowd Counting Neural Architecture SearchabstractRecent solutions to crowd counting problems have already achieved promising performance across various benchmarks. However, applying these approaches to real-world applications is still challenging, because they are computation intensive and lack the flexibility to meet various resource budgets. In this article, we propose an efficient crowd counting neural architecture search (ECCNAS) framework to search efficient crowd counting network structures, which can fill this research gap. A novel search from pre-trained strategy enables our cross-task NAS to explore the significantly large and flexible search space with less search time and get more proper network structures. Moreover, our well-designed search space can intrinsically provide candidate neural network structures with high performance and efficiency. In order to search network structures according to hardwares with different computational performance, we develop a novel latency cost estimation algorithm in our ECCNAS. Experiments show our searched models get an excellent trade-off between computational complexity and accuracy and have the potential to deploy in practical scenarios with various resource budgets. We reduce the computational cost, in terms of multiply-and-accumulate (MACs), by up to 96% with comparable accuracy. And we further designed experiments to validate the efficiency and the stability improvement of our proposed search from pre-trained strategy. Yabin Wang 0001, Zhiheng Ma, Xing Wei 0001, Yaowei Wang 0001, Xiaopeng Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Few-Shot Class-Incremental Learning via Relation Knowledge DistillationabstractIn this paper, we focus on the challenging few-shot class incremental learning (FSCIL) problem, which requires to transfer knowledge from old tasks to new ones and solves catastrophic forgetting. We propose the exemplar relation distillation incremental learning framework to balance the tasks of old-knowledge preserving and new-knowledge adaptation. First, we construct an exemplar relation graph to represent the knowledge learned by the original network and update gradually for new tasks learning. Then an exemplar relation loss function for discovering the relation knowledge between different classes is introduced to learn and transfer the structural information in relation graph. A large number of experiments demonstrate that relation knowledge does exist in the exemplars and our approach outperforms other state-of-the-art class-incremental learning methods on the CIFAR100, miniImageNet, and CUB200 datasets. Songlin Dong, Xiaopeng Hong, Xinyuan Chang, Xing Wei 0001, Yihong Gong |
AAAI | 5 |
| 2021 | Error-Aware Density Isomorphism Reconstruction for Unsupervised Cross-Domain Crowd CountingabstractThis paper focuses on the unsupervised domain adaptation problem for video-based crowd counting, in which we use labeled data as source domain and unlabelled video data as target domain. It is challenging as there is a huge gap between the source and the target domain and no annotations of samples are available in the target domain. The key issue is how to utilize unlabelled videos in the target domain for knowledge learning and transferring from the source domain. To tackle this problem, we propose a novel Error-aware Density Isomorphism REConstruction Network (EDIREC-Net) for cross-domain crowd counting. EDIREC-Net jointly transfers a pre-trained counting model to target domains using a density isomorphism reconstruction objective and models the reconstruction erroneousness by error reasoning. Specifically, as crowd flows in videos are consecutive, the density maps in adjacent frames turn out to be isomorphic. On this basis, we regard the density isomorphism reconstruction error as a self-supervised signal to transfer the pre-trained counting models to different target domains. Moreover, we leverage an estimation-reconstruction consistency to monitor the density reconstruction erroneousness and suppress unreliable density reconstructions during training. Experimental results on four benchmark datasets demonstrate the superiority of the proposed method and ablation studies investigate the efficiency and robustness. The source code is available at https://github.com/GehenHe/EDIREC-Net. Yuhang He 0001, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
AAAI | 3 |
| 2021 | Learning to Count via Unbalanced Optimal TransportabstractCounting dense crowds through computer vision technology has attracted widespread attention. Most crowd counting datasets use point annotations. In this paper, we formulate crowd counting as a measure regression problem to minimize the distance between two measures with different supports and unequal total mass. Specifically, we adopt the unbalanced optimal transport distance, which remains stable under spatial perturbations, to quantify the discrepancy between predicted density maps and point annotations. An efficient optimization algorithm based on the regularized semi-dual formulation of UOT is introduced, which alternatively learns the optimal transportation and optimizes the density regressor. The quantitative and qualitative results illustrate that our method achieves state-of-the-art counting and localization performance. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yunfeng Qiu, Yihong Gong |
AAAI | 2 |
| 2021 | Efficient Deep Image Denoising via Class Specific ConvolutionabstractDeep neural networks have been widely used in image denoising during the past few years. Even though they achieve great success on this problem, they are computationally inefficient which makes them inappropriate to be implemented in mobile devices. In this paper, we propose an efficient deep neural network for image denoising based on pixel-wise classification. Despite using a computationally efficient network cannot effectively remove the noises from any content, it is still capable to denoise from a specific type of pattern or texture. The proposed method follows such a divide and conquer scheme. We first use an efficient U-net to pixel-wisely classify pixels in the noisy image based on the local gradient statistics.Then we replace part of the convolution layers in existing denoising networks by the proposed Class Specific Convolution layers (CSConv) which use different weights for different classes of pixels. Quantitative and qualitative evaluations on public datasets demonstrate that the proposed method can reduce the computational costs without sacrificing the performance compared to state-of-the-art algorithms. Jiawei Zhang 0002, Xuanye Cheng, Feng Zhang 0047, Xing Wei 0001, Jimmy S. J. Ren |
AAAI | 5 |
| 2021 | Towards A Universal Model for Cross-Dataset Crowd CountingabstractThis paper proposes to handle the practical problem of learning a universal model for crowd counting across scenes and datasets. We dissect that the crux of this problem is the catastrophic sensitivity of crowd counters to scale shift, which is very common in the real world and caused by factors such as different scene layouts and image resolutions. Therefore it is difficult to train a universal model that can be applied to various scenes. To address this problem, we propose scale alignment as a prime module for establishing a novel crowd counting framework. We derive a closed-form solution to get the optimal image rescaling factors for alignment by minimizing the distances between their scale distributions. A novel neural network together with a loss function based on an efficient sliced Wasserstein distance is also proposed for scale distribution estimation. Benefiting from the proposed method, we have learned a universal model that generally works well on several datasets where can even outperform state-of-the-art models that are particularly fine-tuned for each dataset significantly. Experiments also demonstrate the much better generalizability of our model to unseen scenes. Zhiheng Ma, Xiaopeng Hong, Xing Wei 0001, Yunfeng Qiu, Yihong Gong |
ICCV | 3 |
| 2021 | Anomaly Detection Via Self-Organizing MapabstractAnomaly detection plays a key role in industrial manufacturing for product quality control. Traditional methods for anomaly detection are rule-based with limited generalization ability. Recent methods based on supervised deep learning are more powerful but require large-scale annotated datasets for training. In practice, abnormal products are rare thus it is very difficult to train a deep model in a fully supervised way. In this paper, we propose a novel unsupervised anomaly detection approach based on Self-organizing Map (SOM). Our method, Self-organizing Map for Anomaly Detection (SOMAD) maintains normal characteristics by using topological memory based on multi-scale features. SOMAD achieves state-of-the-art performance on unsupervised anomaly detection and localization on the MVTec dataset. Kaitao Jiang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ICIP | 4 |
| 2021 | Direct Measure Matching for Crowd CountingabstractTraditional crowd counting approaches usually use Gaussian assumption to generate pseudo density ground truth, which suffers from problems like inaccurate estimation of the Gaussian kernel sizes. In this paper, we propose a new measure-based counting approach to regress the predicted density maps to the scattered point-annotated ground truth directly. First, crowd counting is formulated as a measure matching problem. Second, we derive a semi-balanced form of Sinkhorn divergence, based on which a Sinkhorn counting loss is designed for measure matching. Third, we propose a self-supervised mechanism by devising a Sinkhorn scale consistency loss to resist scale changes. Finally, an efficient optimization method is provided to minimize the overall loss function. Extensive experiments on four challenging crowd counting datasets namely ShanghaiTech, UCF-QNRF, JHU++ and NWPU have validated the proposed method. Xiaopeng Hong, Zhiheng Ma, Xing Wei 0001, Yunfeng Qiu, Yaowei Wang 0001, Yihong Gong |
IJCAI | 4 |
| 2021 | InsPose: Instance-Aware Networks for Single-Stage Multi-Person Pose EstimationabstractMulti-person pose estimation is an attractive and challenging task. Existing methods are mostly based on two-stage frameworks, which include top-down and bottom-up methods. Two-stage methods either suffer from high computational redundancy for additional person detectors or they need to group keypoints heuristically after predicting all the instance-agnostic keypoints. The single-stage paradigm aims to simplify the multi-person pose estimation pipeline and receives a lot of attention. However, recent single-stage methods have the limitation of low performance due to the difficulty of regressing various full-body poses from a single feature vector. Different from previous solutions that involve complex heuristic designs, we present a simple yet effective solution by employing instance-aware dynamic networks. Specifically, we propose an instance-aware module to adaptively adjust (part of) the network parameters for each instance. Our solution can significantly increase the capacity and adaptive-ability of the network for recognizing various poses, while maintaining a compact end-to-end trainable pipeline. Extensive experiments on the MS-COCO dataset demonstrate that our method achieves significant improvement over existing single-stage methods, and makes a better balance of accuracy and efficiency compared to the state-of-the-art two-stage approaches. Dahu Shi, Xing Wei 0001, Wenming Tan, Ye Ren, Shiliang Pu |
ACM Multimedia | 2 |
| 2021 | Beyond Universal Person Re-Identification AttackabstractDeep learning-based person re-identification (Re-ID) has made great progress and achieved high performance recently. In this paper, we make the first attempt to examine the vulnerability of current person Re-ID models against a dangerous attack method, i.e., the universal adversarial perturbation (UAP) attack, which has been shown to fool classification models with a little overhead. We propose a more universal adversarial perturbation (MUAP) method for both image-agnostic and model-insensitive person Re-ID attack. Firstly, we adopt a list-wise attack objective function to disrupt the similarity ranking list directly. Secondly, we propose a model-insensitive mechanism for cross-model attack. Extensive experiments show that the proposed attack approach achieves high attack performance and outperforms other state of the arts by large margin in cross-model scenario. The results also demonstrate the vulnerability of current Re-ID models to MUAP and further suggest the need of designing more robust Re-ID models. Xing Wei 0001, Rongrong Ji, Xiaopeng Hong, Qi Tian 0001, Yihong Gong |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Person Re-Identification in Aerial ImageryabstractNowadays, with the rapid development of consumer Unmanned Aerial Vehicles (UAVs), visual surveillance by utilizing the UAV platform has been very attractive. Most of the research works for UAV captured visual data are mainly focused on the tasks of object detection and tracking. However, limited attention has been paid to the task of person Re-identification (ReID) which has been widely studied in ordinary surveillance cameras with fixed emplacements. In this paper, to facilitate the research of person ReID in aerial imagery, we collect a large scale airborne person ReID dataset named as Person ReID in Aerial Imagery (PRAI-1581), which consists of 39,461 images of 1581 person identities. The images of the dataset are shot by two DJI consumer UAVs flying at an altitude ranging from 20 to 60 meters above the ground, which covers most of the real UAV surveillance scenarios. In addition, we propose to utilize subspace pooling of convolution feature maps to represent the input person images. Our method can learn a discriminative and compact feature representation for ReID in aerial imagery and can be trained in an end-to-end fashion efficiently. We conduct extensive experiments on the proposed dataset and the experimental results demonstrate that re-identifying persons in aerial imagery is a challenging problem, where our method performs favorably against state of the arts. Shizhou Zhang, Xing Wei 0001, Peng Wang 0015, Bingliang Jiao, Yanning Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Infrared-Visible Cross-Modal Person Re-Identification with an X ModalityabstractThis paper focuses on the emerging Infrared-Visible cross-modal person re-identification task (IV-ReID), which takes infrared images as input and matches with visible color images. IV-ReID is important yet challenging, as there is a significant gap between the visible and infrared images. To reduce this ‘gap’, we introduce an auxiliary X modality as an assistant and reformulate infrared-visible dual-mode cross-modal learning as an X-Infrared-Visible three-mode learning problem. The X modality restates from RGB channels to a format with which cross-modal learning can be easily performed. With this idea, we propose an X-Infrared-Visible (XIV) ReID cross-modal learning framework. Firstly, the X modality is generated by a lightweight network, which is learnt in a self-supervised manner with the labels inherited from visible images. Secondly, under the XIV framework, cross-modal learning is guided by a carefully designed modality gap constraint, with information exchanged cross the visible, X, and infrared modalities. Extensive experiments are performed on two challenging datasets SYSU-MM01 and RegDB to evaluate the proposed XIV-ReID approach. Experimental results show that our method considerably achieves an absolute gain of over 7% in terms of rank 1 and mAP even compared with the latest state-of-the-art methods. Diangang Li, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
AAAI | 2 |
| 2020 | Superpixel Masking and Inpainting for Self-Supervised Anomaly Detection
Kaitao Jiang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
BMVC | 5 |
| 2020 | Few-Shot Class-Incremental LearningabstractThe ability to incrementally learn new classes is crucial to the development of real-world artificial intelligence systems. In this paper, we focus on a challenging but practical few-shot class-incremental learning (FSCIL) problem. FSCIL requires CNN models to incrementally learn new classes from very few labelled samples, without forgetting the previously learned ones. To address this problem, we represent the knowledge using a neural gas (NG) network, which can learn and preserve the topology of the feature manifold formed by different classes. On this basis, we propose the TOpology-Preserving knowledge InCrementer (TOPIC) framework. TOPIC mitigates the forgetting of the old classes by stabilizing NG's topology and improves the representation learning for few-shot new classes by growing and adapting NG to new training samples. Comprehensive experimental results demonstrate that our proposed method significantly outperforms other state-of-the-art class-incremental learning methods on CIFAR100, miniImageNet, and CUB200 datasets. Xiaopeng Hong, Xinyuan Chang, Songlin Dong, Xing Wei 0001, Yihong Gong |
CVPR | 5 |
| 2020 | Topology-Preserving Class-Incremental Learning
Xinyuan Chang, Xiaopeng Hong, Xing Wei 0001, Yihong Gong |
ECCV (19) | 4 |
| 2020 | Complex Spatial-Temporal Attention Aggregation For Video Person Re-IdentificationabstractVideo-based person re-identification (Re-ID) aims to match pedestrian tracklets of the same identity captured by different cameras. Existing works usually compute the video-level feature representation via simple frame-level feature aggregation, such as average pooling and max pooling. However, the performance of such methods degenerates severely under low signal-noise ratio and partial occlusions. In this paper, we propose a novel Complex Spatial-Temporal Attention Aggregation (CAA), which fully exploits the discriminative information in spatial-temporal dimension via the combination of two aggregation method, namely region-aware aggregation and region-regardless aggregation. We evaluate the proposed method in three widely used video Re-ID datasets, including MARS, iLIDS-VID, and PRID-2011. The experimental results demonstrate that the proposed method outperforms the state of the arts. Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ICIP | 2 |
| 2020 | Class-Incremental Learning with Topological Schemas of Memory SpacesabstractClass-incremental learning (CIL) aims to incrementally learn a unified classifier for new classes emerging, which suffers from the catastrophic forgetting problem. To alleviate forgetting and improve the recognition performance, we propose a novel CIL framework, named the topological schemas model (TSM). TSM consists of a Gaussian mixture model arranged on 2D grids (2D-GMM) as the memory of the learned knowledge. To train the 2D-GMM model, we develop a novel competitive expectation-maximization (CEM) method, which contains a global topology embedding step and a local expectation-maximization fine-tuning step. Meanwhile, we choose the image samples of old classes that have the maximum posterior probability with respect to each Gaussian distribution as the episodic points. When finetuning for new classes, we propose the memory preservation loss (MPL) term to ensure episodic points still have maximum probabilities with respect to the corresponding Gaussian distribution. MPL preserves the distribution of 2D-GMM for old knowledge during incremental learning and alleviates catastrophic forgetting. Comprehensive experimental evaluations on two popular CIL benchmarks CIFAR100 and subImageNet demonstrate the superiority of our TSM. Xinyuan Chang, Xiaopeng Hong, Xing Wei 0001, Wei Ke 0003, Yihong Gong |
ICPR | 4 |
| 2020 | Polynomial Universal Adversarial Perturbations for Person Re-IdentificationabstractIn this paper, we focus on Universal Adversarial Perturbations (UAP) attack on state-of-the-art person re-identification (Re-ID) methods. Existing UAP methods usually compute a perturbation image and add it to the images of interest. Such a simple constant form greatly limits the attack power. To address this problem, we extend the formulation of UAP to a polynomial form and propose the Polynomial Universal Adversarial Perturbation (PUAP). Unlike traditional UAP methods which only rely on the additive perturbation signal, the proposed PUAP consists of both an additive perturbation and a multiplicative modulation factor. The additive perturbation produces the fundamental component of the signal, while the multiplicative factor modulates the perturbation signal in line with the unit impulse pattern of the input image. Moreover, we introduce a Pearson correlation coefficient loss to generate universal perturbations, for disrupting the outputs of person Re-ID models. Extensive experiments on DukeMTMC-reID, Market-1501, and MARS show that the proposed method can efficiently improve the attack performance, especially when the magnitude of UAP is constrained to a relatively small value. Xing Wei 0001, Rongrong Ji, Xiaopeng Hong, Yihong Gong |
ICPR | 2 |
| 2020 | Learning Scales from Points: A Scale-aware Probabilistic Model for Crowd CountingabstractCounting people automatically through computer vision technology is a challenging task. Recently, convolution neural network (CNN) based methods have made significant progress. Nonetheless, large scale variations of instances caused by, for example, perspective effects remain unsolved. Moreover, it is problematic to estimate scales with only point annotations. In this paper, we propose a scale-aware probabilistic model to handle this problem. Unlike previous methods that generate a single density map where instances of various scales are processed indiscriminately, we propose a density pyramid network (DPN), where each pyramid level handles instances within a particular scale range. Furthermore, we propose a scale distribution estimator (SDE) to learn scales of people from input data, under the weak supervision of point annotations. Finally, we adopt an instance-level probabilistic scale-aware model (IPSM) to guide the multi-scale training of DPN explicitly. Qualitative and quantitative experimental results demonstrate the effectiveness of the proposed method, which achieves competitive results on four widely used benchmarks. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ACM Multimedia | 2 |
| 2020 | Co-Attentive Lifting for Infrared-Visible Person Re-IdentificationabstractInfrared-visible cross-modality person re-identification (IV-ReID) has attracted much attention with the popularity of dual-mode video surveillance systems, where the RGB mode works in the daytime and automatically switches to the infrared mode at night. Despite its significant application value, IV-ReID remains a difficult problem mainly due to two great challenges. First, it is difficult to identify persons in the infrared image, which lacks color and texture clues. Second, there is a significant gap between the infrared and visible modalities where appearances of the same person vary considerably. This paper proposes a novel attention-based approach to handle the two difficulties in a unified framework. 1) We propose an attention lifting mechanism to learn discriminative features in each modality. 2) We propose a co-attentive learning mechanism to bridge the gap between the two modalities. Our method only makes slight modifications of a given backbone network and requires small computation overhead while improving the performance significantly. We conduct extensive experiments to demonstrate the superiority of our proposed method. Xing Wei 0001, Diangang Li, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
ACM Multimedia | 1 |
| 2020 | Transductive semi-supervised metric learning for person re-identification
Xinyuan Chang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
Pattern Recognit. | 3 |
| 2020 | Multi-Target Multi-Camera Tracking by Tracklet-to-Target AssignmentabstractThis paper focuses on the Multi-Target Multi-Camera Tracking task (MTMCT), which aims at tracking multiple targets within a multi-camera network. As the trajectory of each target is inherently split into multiple sub-trajectories (namely local tracklets) in a multi-camera network, a major challenge of MTMCT is how to accurately match the local tracklets generated within each camera across different cameras and generate a complete global trajectory for each target, i.e., the cross-camera tracklet matching problem. We solve the cross-camera tracklet matching problem by TRACklet-to-Target Assignment (TRACTA), and propose the Restricted Non-negative Matrix Factorization (RNMF) algorithm to compute the optimal assignment solution that meets a set of constraints, which should be in force in practice. TRACTA can correct the tracking errors caused by occlusions and missed detections in local tracklets, and produce a complete global trajectory for each target across all the cameras. Moreover, we also develop an analytical way of estimating the total number of targets in the camera network, which plays an important role to compute the tracklet-to-target assignment. Experimental evaluations and ablation studies on four MTMCT benchmark datasets show the superiority of the proposed TRACTA method. Yuhang He 0001, Xing Wei 0001, Xiaopeng Hong, Weiwei Shi 0003, Yihong Gong |
IEEE Trans. Image Process. | 2 |
| 2019 | Bayesian Loss for Crowd Count Estimation With Point SupervisionabstractIn crowd counting datasets, each person is annotated by a point, which is usually the center of the head. And the task is to estimate the total count in a crowd scene. Most of the state-of-the-art methods are based on density map estimation, which convert the sparse point annotations into a “ground truth” density map through a Gaussian kernel, and then use it as the learning target to train a density map estimator. However, such a "ground-truth" density map is imperfect due to occlusions, perspective effects, variations in object shapes, etc. On the contrary, we propose Bayesian loss, a novel loss function which constructs a density contribution probability model from the point annotations. Instead of constraining the value at every pixel in the density map, the proposed training loss adopts a more reliable supervision on the count expectation at each annotated point. Without bells and whistles, the loss function makes substantial improvements over the baseline loss on all tested datasets. Moreover, our proposed loss function equipped with a standard backbone network, without using any external detectors or multi-scale architectures, plays favourably against the state of the arts. Our method outperforms previous best approaches by a large margin on the latest and largest UCF-QNRF dataset. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ICCV | 2 |
| 2019 | Person Re-identification with Neural Architecture Search
Shizhou Zhang, Xing Wei 0001, Peng Wang 0015, Yanning Zhang 0001 |
PRCV (1) | 3 |
| 2018 | Kernelized Subspace Pooling for Deep Local DescriptorsabstractRepresenting local image patches in an invariant and discriminative manner is an active research topic in computer vision. It has recently been demonstrated that local feature learning based on deep Convolutional Neural Network (CNN) can significantly improve the matching performance. Previous works on learning such descriptors have focused on developing various loss functions, regularizations and data mining strategies to learn discriminative CNN representations. Such methods, however, have little analysis on how to increase geometric invariance of their generated descriptors. In this paper, we propose a descriptor that has both highly invariant and discriminative power. The abilities come from a novel pooling method, dubbed Subspace Pooling (SP) which is invariant to a range of geometric deformations. To further increase the discriminative power of our descriptor, we propose a simple distance kernel integrated to the marginal triplet loss that helps to focus on hard examples in CNN training. Finally, we show that by combining SP with the projection distance metric [13], the generated feature descriptor is equivalent to that of the Bilinear CNN model [22], but outperforms the latter with much lower memory and computation consumptions. The proposed method is simple, easy to understand and achieves good performance. Experimental results on several patch matching benchmarks show that our method outperforms the state-of-the-arts significantly. Xing Wei 0001, Yihong Gong, Nanning Zheng 0001 |
CVPR | 1 |
| 2018 | Grassmann Pooling as Compact Homogeneous Bilinear Pooling for Fine-Grained Visual Classification
Xing Wei 0001, Yihong Gong, Jiawei Zhang 0002, Nanning Zheng 0001 |
ECCV (3) | 1 |
| 2018 | Specular highlight reduction with known surface geometry
Xing Wei 0001, Xiaobin Xu 0001, Jiawei Zhang 0002, Yihong Gong |
Comput. Vis. Image Underst. | 1 |
| 2018 | Joint Contour Filtering
Xing Wei 0001, Qingxiong Yang, Yihong Gong |
Int. J. Comput. Vis. | 1 |
| 2018 | Spatiotemporal GMM for Background Subtraction with Superpixel HierarchyabstractWe propose a background subtraction algorithm using hierarchical superpixel segmentation, spanning trees and optical flow. First, we generate superpixel segmentation trees using a number of Gaussian Mixture Models (GMMs) by treating each GMM as one vertex to construct spanning trees. Next, we use the -smoother to enhance the spatial consistency on the spanning trees and estimate optical flow to extend the -smoother to the temporal domain. Experimental results on synthetic and real-world benchmark datasets show that the proposed algorithm performs favorably for background subtraction in videos against the state-of-the-art methods in spite of frequent and sudden changes of pixel values. Xing Wei 0001, Qingxiong Yang, Qing Li 0001, Gang Wang 0012, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Superpixel HierarchyabstractSuperpixel segmentation has been one of the most important tasks in computer vision. In practice, an object can be represented by a number of segments at finer levels with consistent details or included in a surrounding region at coarser levels. Thus, a superpixel segmentation hierarchy is of great importance for applications that require different levels of image details. However, there is no method that can generate all scales of superpixels accurately in real time. In this paper, we propose the superhierarchy algorithm which is able to generate multi-scale superpixels as accurately as the state-of-the-art methods but with one to two orders of magnitude speed-up. The proposed algorithm can be directly integrated with recent efficient edge detectors to significantly outperform the state-of-the-art methods in terms of segmentation accuracy. Quantitative and qualitative evaluations on a number of applications demonstrate that the proposed algorithm is accurate and efficient in generating a hierarchy of superpixels. Xing Wei 0001, Qingxiong Yang, Yihong Gong, Narendra Ahuja, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Training mixture of weighted SVM for object detection using EM algorithm
De Cheng, Jinjun Wang, Xing Wei 0001, Yihong Gong |
Neurocomputing | 3 |