Qilong Wang 0001

dblp:119/1488 · DBLP profile ↗
← Back
79ranked-venue papers
15as first author
46since 2021 · last 2026
0000-0002-3765-9787ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 57 · 12 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 9 first-author · 27 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 GDT-VLM: Global Distribution Modeling for Visual Token Compression in Efficient Multimodal Large Language Models
abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable success, yet the massive number of visual tokens per image imposes a heavy inference burden. Existing methods attempt extreme compression with a single visual token via spatial reduction or cross-modal attention, but often overlook the statistical information inherent in visual tokens, leading to a suboptimal efficiency-effectiveness trade-off. In this paper, we show that effective statistical characterization of visual features benefits extreme visual token compression in efficient MLLMs. To this end, we propose GDT-VLM, a novel architecture that exploits global distribution modeling of visual tokens for efficient compression. Specifically, GDT-VLM encodes visual features by jointly modeling their global first-order (GAP) and second-order (Brownian Distance Covariance) statistics, enabling a more expressive yet compact representation. By capturing the holistic characteristics of vision tokens, our GDT-VLM yields compact vision information in a single token while effectively preserving statistical content. Extensive experiments on 7 benchmarks show that our approach achieves competitive accuracy while offering a favorable efficiency-effectiveness trade-off.
Jiangtao Xie, Qilong Wang 0001, Peihua Li
ICMR4
2026 DeST: A Decoupled Spatio-Temporal Framework for Action Segmentation
Zhongyu Li 0006, Shanghua Gao, Qilong Wang 0001, Qibin Hou, Ming-Ming Cheng
Int. J. Comput. Vis.4
2026 Global and Local Visual-Textual Alignment for Open Vocabulary Object Detection
abstract
Recently, with the development of the Vision-Language Model (VLM), adopting such VLM (e.g., CLIP) into object detection framework has gradually become a promising and attractive research direction, and the resulted open vocabulary object detection methods can effectively alleviate the limitations in those close-set ones, making the detectors perceive the unseen world. The core issue in open vocabulary object detection is to design an effective and efficient alignment between the visual (e.g., image) and textual (e.g., caption) features in the semantic space, so that the detectors can capture more information around the open-set scene. Current approaches deploy extra uncurated image-text pairs to pre-train a detector for obtaining a better visual-textual alignment in the feature space. Besides, knowledge distillation technology is also adopted to design an appropriate information transferring flow for aligning the visual-textual knowledge. However, large-scale image-text pairs are not always available to obtain, and the pre-training process will inevitable introduce much more computation overhead. While knowledge distillation methods focus on aligning between the local region visual feature in RoI and the textual features of VLM, neglecting the global information alignment between the image and text. For addressing the dilemmas in these alignment manners, we propose a Global and Local Visual-Textual Alignment for Open Vocabulary Object Detection in this paper. Specifically, our proposed method integrates global image-caption and local region-prompt alignments into a unified learning paradigm. The global alignment takes the whole image and caption as the visual and textual inputs, respectively, and matches the image and caption representations from the detector and the text encoder in CLIP by contrastive learning from the overall perspective. Different from global alignment, the local one concentrates on the accordance between regions and prompts from the aspect of portion description. It extracts and aligns the embeddings for the visual patch RoIs from the image encoder in CLIP and discriminating textual token prompts from the text encoder. Moreover, we also design a prompt tuning strategy, which contains global and local components corresponding to the alignment procedure, for better adapting CLIP to downstream task object detection in a parameter-efficient learning manner. By implementation on Faster R-CNN, we conduct experiments on open vocabulary benchmarks OV-COCO and OV-LVIS, respectively. The results verify that our proposed method can achieve clear improvement over counterparts on novel categories, while performing favorably against state-of-the-arts.
Hao Wang 0073, Tong Jia 0001, Shizhuo Deng, Dongyue Chen 0001, Qilong Wang 0001, Wangmeng Zuo
IEEE Trans. Image Process.5
2026 Heterogeneous Federated Dynamic Graph HyperNetwork for Image Classification
abstract
Federated learning (FL) enables privacy-preserving collaboration among distributed clients, but practical deployments often face heterogeneous models and non-IID data, leading to degraded communication and personalization. In addition, real-world FL systems frequently encounter newly joined clients that require rapid adaptation and abnormal clients that may upload corrupted updates, further exacerbating instability and hindering global convergence. To address these challenges in image classification, we propose HFedDGHN, a Heterogeneous Federated Dynamic Graph HyperNetwork that jointly models inter-client relations and personalized parameter generation. Specifically, a graph structure learner adaptively captures client correlations to construct a dynamic collaboration graph, while a graph-convolutional hypernetwork generates model parameters for heterogeneous architectures, enabling implicit knowledge transfer without sharing local data or weights. Moreover, the framework naturally supports meta-learning-based generalization, allowing efficient adaptation to newly joined clients. Furthermore, the dynamic graph enhances robustness by isolating abnormal clients, as they tend to be excluded from most neighborhoods during adaptive graph construction. Extensive experiments across multiple benchmarks demonstrate that HFedDGHN achieves superior accuracy compared to state-of-the-art personalized and heterogeneous FL methods, while naturally improving robustness and scalability in real-world deployments.
Liu Yang 0010, Kegen Chen, Qilong Wang 0001, Zhengyi Xu, Shiqiao Gu, Qinghua Hu
IEEE Trans. Image Process.3
2026 Orientation-Aware Task-Decoupled Learning for Oriented Object Detection
abstract
Recent studies have shown that disentanglement of classification and localization tasks has great potential to improve the performance of general object detection. However, such kind of disentanglement strategies remain not well explored in oriented object detection. Particularly, there exist two challenges lying in task disentanglement for oriented object detection: (1) existing task-decoupled methods ignore the orientation of objects, hardly coping with arbitrarily oriented objects; (2) the targets in oriented object detection (e.g., high-resolution remote sensing images) are generally small-size and fine-grained, making classification more difficult. To handle the above issues, we rethink task-decoupled policy in oriented object detection and propose an effective Orientation-aware Task-Decoupled Learning (OTDL) method. Specifically, our OTDL first presents a light-weight Task-specific Proposal Offset Learning (TPOL) module to generate the eligible proposals for arbitrarily oriented objects, where TPOL module equips classification and localization tasks with individual proposals by learning task-specific and orientation-aware offsets in a local coordinate. Furthermore, we empirically study the effect of various double-head strategies on performance of oriented object detection, while proposing a novel Pyramid Covariance Attention (PCA)-based classification head to cope with small-size and fine-grained targets. Based on the proposed TPOL module and PCA-based classification head, our OTDL explores the potential of task disentanglement for improving the performance of oriented object detection. The experiments are conducted on five oriented object detection benchmarks (i.e., DOTA-v1.0, DOTA-v1.5, HRSC2016, DIOR-R and SODA-A), and the results show our OTDL method significantly outperforms its counterparts, while achieving state-of-the-art performance.
Qilong Wang 0001, Qinghua Hu
IEEE Trans. Multim.2
2025 TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition
abstract
Going beyond few-shot action recognition (FSAR), cross-domain FSAR (CDFSAR) has attracted recent research interests by solving the domain gap lying in source-to-target transfer learning. Existing CDFSAR methods mainly focus on joint training of source and target data to mitigate the side effect of domain gap. However, such kind of methods suffer from two limitations: First, pair-wise joint training requires retraining deep models in case of one source data and multiple target ones, which incurs heavy computation cost, especially for large source and small target data. Second, pre-trained models after joint training are adopted to target domain in a straightforward manner, hardly taking full potential of pre-trained models and then limiting recognition performance. To overcome above limitations, this paper proposes a simple yet effective baseline, namely Temporal-Aware Model Tuning (TAMT) for CDFSAR. Specifically, our TAMT involves a decoupled paradigm by performing pre-training on source data and fine-tuning target data, which avoids retraining for multiple target data with single source. To effectively and efficiently explore the potential of pre-trained models in transferring to target domain, our TAMT proposes a Hierarchical Temporal Tuning Network (HTTN), whose core involves local temporal-aware adapters (TAA) and a global temporal-aware moment tuning (GTMT). Particularly, TAA learns few parameters to recalibrate the intermediate features of frozen pre-trained models, enabling efficient adaptation to target domains. Furthermore, GTMT helps to generate powerful video representations, improving match performance on the target domain. Experiments on several widely used video benchmarks show our TAMT outperforms the recently proposed counterparts by 13% ∼31%, achieving new state-of-the-art CDFSAR results.
Zilin Gao, Qilong Wang 0001, Zhaofeng Chen, Peihua Li, Qinghua Hu
CVPR3
2025 Generative Inbetweening through Frame-wise Conditions-Driven Video Generation
abstract
Generative inbetweening aims to generate intermediate frame sequences by utilizing two key frames as input. Although remarkable progress has been made in video generation models, generative inbetweening still faces challenges in maintaining temporal stability due to the ambiguous interpolation path between two key frames. This issue becomes particularly severe when there is a large motion gap between input frames. In this paper, we propose a straight-forward yet highly effective Frame-wise Conditions-driven Video Generation (FCVG) method that significantly enhances the temporal stability of interpolated video frames. Specifically, our FCVG provides an explicit condition for each frame, making it much easier to identify the interpolation path between two input frames and thus ensuring temporally stable production of visually plausible video frames. To achieve this, we suggest extracting matched lines from two input frames that can then be easily interpolated frame by frame, serving as frame-wise conditions seamlessly integrated into existing video generation models. In extensive evaluations covering diverse scenarios such as natural landscapes, complex human poses, camera movements and animations, existing methods often exhibit incoherent transitions across frames. In contrast, our FCVG demonstrates the capability to generate temporally stable videos using both linear and non-linear interpolation curves. Our project page and code are available at https://fcvg-inbetween.github.io/.
Dongwei Ren, Qilong Wang 0001, Xiaohe Wu, Wangmeng Zuo
CVPR3
2025 Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem Solving
Zixian Guo, Ming Liu 0018, Qilong Wang 0001, Zhilong Ji, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo
ICCV3
2025 Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning
abstract
Recently, remarkable progress has been made in large-scale pre-trained model tuning, and inference efficiency is becoming more crucial for practical deployment. Early exiting in conjunction with multi-stage predictors, when cooperated with a parameter-efficient fine-tuning strategy, offers a straightforward way to achieve an inference-efficient model. However, a key challenge remains unresolved: How can early stages provide low-level fundamental features to deep stages while simultaneously supplying high-level discriminative features to early-stage predictors? To address this problem, we propose a Decoupled Multi-Predictor Optimization (DMPO) method to effectively decouple the low-level representative ability and high-level discriminative ability in early stages. First, in terms of architecture, we introduce a lightweight bypass module into multi-stage predictors for functional decomposition of shallow features from early stages, while a high-order statistics-based predictor is developed for early stages to effectively enhance their discriminative ability. To reasonably train our multi-predictor architecture, a decoupled optimization is proposed to allocate two-phase loss weights for multi-stage predictors during model tuning, where the initial training phase enables the model to prioritize the acquisition of discriminative ability of deep stages via emphasizing representative ability of early stages, and the latter training phase drives discriminative ability towards earlier stages as much as possible. As such, our DMPO can effectively decouple representative and discriminative abilities in early stages in terms of architecture design and model optimization. Experiments across various datasets and pre-trained backbones demonstrate that DMPO clearly outperforms its counterparts when reducing computational cost.
Liwei Luo, Shuaitengyuan Li, Dongwei Ren, Qilong Wang 0001, Pengfei Zhu 0001, Qinghua Hu
ICCV4
2025 Unknown Text Learning for Clip-Based Few-Shot Open-Set Recognition
Qilong Wang 0001, Bing Cao 0002, Qinghua Hu, Yahong Han
ICCV2
2025 DALIP: Distribution Alignment-Based Language-Image Pre-Training for Domain-Specific Data
Jiangtao Xie, Qilong Wang 0001, Qinghua Hu, Peihua Li
ICCV4
2025 Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark
abstract
The dynamic imbalance of the fore-background is a major challenge in video object counting, which is usually caused by the sparsity of target objects. This remains understudied in existing works and often leads to severe under-/over-prediction errors. To tackle this issue in video object counting, we propose a density-embedded Efficient Masked Autoencoder Counting (E-MAC) framework in this paper. To empower the model’s representation ability on density regression, we develop a new Density-Embedded Masked mOdeling (DEMO) method, which first takes the density map as an auxiliary modality to perform multimodal self-representation learning for image and density map. Although DEMO contributes to effective cross-modal regression guidance, it also brings in redundant background information, making it difficult to focus on the foreground regions. To handle this dilemma, we propose an efficient spatial adaptive masking derived from density maps to boost efficiency. Meanwhile, we employ an optical flow-based temporal collaborative fusion strategy to effectively capture the dynamic variations across frames, aligning features to derive multi-frame density residuals. The counting accuracy of the current frame is boosted by harnessing the information from adjacent frames. In addition, considering that most existing datasets are limited to human-centric scenarios, we propose a large video bird counting dataset, DroneBird, in natural scenarios for migratory bird protection. Extensive experiments on three crowd datasets and our DroneBird validate our superiority against the counterparts. The code and dataset are available.
Bing Cao 0002, Quanhao Lu, Jiekang Feng, Qilong Wang 0001, Pengfei Zhu 0001, Qinghua Hu
ICLR4
2025 Asymmetric Factorized Bilinear Operation for Vision Transformer
abstract
As a core component of Transformer-like deep architectures, a feed-forward network (FFN) for channel mixing is responsible for learning features of each token. Recent works show channel mixing can be enhanced by increasing computational burden or can be slimmed at the sacrifice of performance. Although some efforts have been made, existing works are still struggling to solve the paradox of performance and complexity trade-offs. In this paper, we propose an Asymmetric Factorized Bilinear Operation (AFBO) to replace FFN of vision transformer (ViT), which attempts to efficiently explore rich statistics of token features for achieving better performance and complexity trade-off. Specifically, our AFBO computes second-order statistics via a spatial-channel factorized bilinear operation for feature learning, which replaces a simple linear projection in FFN and enhances the feature learning ability of ViT by modeling second-order correlation among token features. Furthermore, our AFBO presents two structured-sparsity channel mapping strategies, namely Grouped Cross Channel Mapping (GCCM) and Overlapped Cycle Channel Mapping (OCCM). They decompose bilinear operation into grouped channel features by considering information interaction between groups, significantly reducing computational complexity while guaranteeing model performance. Finally, our AFBO is built with GCCM and OCCM in an asymmetric way, aiming to achieve a better trade-off. Note that our AFBO is model-agnostic, which can be flexibly integrated with existing ViTs. Experiments are conducted with twenty ViTs on various tasks, and the results show our AFBO is superior to its counterparts while improving existing ViTs in terms of generalization and robustness.
Qilong Wang 0001, Jiangtao Xie, Pengfei Zhu 0001, Qinghua Hu
ICLR2
2025 RoomEditor: High-Fidelity Furniture Synthesis with Parameter-Sharing U-Net
abstract
Virtual furniture synthesis, a critical task in image composition, aims to seamlessly integrate reference objects into indoor scenes while preserving geometric coherence and visual realism. Despite its significant potential in home design applications, this field remains underexplored due to two major challenges: the absence of publicly available and ready-to-use benchmarks hinders reproducible research, and existing image composition methods fail to meet the stringent fidelity requirements for realistic furniture placement. To address these issues, we introduce RoomBench, a ready-to-use benchmark dataset for virtual furniture synthesis, comprising 7,298 training pairs and 895 testing samples across 27 furniture categories. Then, we propose RoomEditor, a simple yet effective image composition method that employs a parameter-sharing dual U-Net architecture, ensuring better feature consistency by sharing weights between dual branches. Technical analysis reveals that conventional dual-branch architectures generally suffer from inconsistent intermediate features due to independent processing of reference and background images. In contrast, RoomEditor enforces unified feature learning through shared parameters, thereby facilitating model optimization for robust geometric alignment and maintaining visual consistency. Experiments show our RoomEditor is superior to state-of-the-arts, while generalizing directly to diverse objects synthesis in unseen scenes without task-specific fine-tuning. Our dataset and code are available at https://github.com/stonecutter-21/roomeditor.
Zhenyi Lin, Xiaofan Ming, Qilong Wang 0001, Dongwei Ren, Wangmeng Zuo, Qinghua Hu
NeurIPS3
2025 A2 M2-Net: Adaptively Aligned Multi-scale Moment for Few-Shot Action Recognition
Zilin Gao, Qilong Wang 0001, Bingbing Zhang 0001, Qinghua Hu, Peihua Li
Int. J. Comput. Vis.2
2025 BackMix: Regularizing Open Set Recognition by Removing Underlying Fore-Background Priors
abstract
Open set recognition (OSR) requires models to classify known samples while detecting unknown samples for real-world applications. Existing studies show impressive progress using unknown samples from auxiliary datasets to regularize OSR models, but they have proved to be sensitive to selecting such known outliers. In this paper, we discuss the aforementioned problem from a new perspective: Can we regularize OSR models without elaborately selecting auxiliary known outliers? We first empirically and theoretically explore the role of foregrounds and backgrounds in open set recognition and disclose that: 1) backgrounds that correlate with foregrounds would mislead the model and cause failures when encounters 'partially' known images; 2) Backgrounds unrelated to foregrounds can serve as auxiliary known outliers and provide regularization via global average pooling. Based on the above insights, we propose a new method, Background Mix (BackMix), that mixes the foreground of an image with different backgrounds to remove the underlying fore-background priors. Specifically, BackMix first estimates the foreground with class activation maps (CAMs), then randomly replaces image patches with backgrounds from other images to obtain mixed images for training. With backgrounds de-correlated from foregrounds, the open set recognition performance is significantly improved. The proposed method is quite simple to implement, requires no extra operation for inferences, and can be seamlessly integrated into almost all of the existing frameworks.
Yu Wang 0106, Junxian Mu, Hongzhi Huang, Qilong Wang 0001, Pengfei Zhu 0001, Qinghua Hu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Automatic Label Assignment for Object Detection
abstract
Label assignment, which aims to classify region proposals as positive or negative samples depending on the correlations between their classification and localization predictions with the corresponding ground truth, is recognized as an essential ingredient in object detection and strongly affects the detection performance. Recently, some dynamic label assignment methods have been proposed to overcome the limitations of the static methods and achieve promising performance improvement. Despite eliminating the restrictions of the human prior sampling knowledge in static methods, existing dynamic principles usually suffer from two weaknesses. First, most of them deploy mixture models or implicit branch in prediction head to coarsely estimate the spatial distribution of the positive samples for objects. They give little attention to the effect of appearance information of the objects. Furthermore, these methods still cannot perceive the quality distribution of the positive samples, and these low-quality samples lead to adverse effects on the detection performance. To address issues, this paper presents a novel automatic label assignment for object detection. Specifically, our method first introduces an instance property branch into object detection pipeline to distinguish the foreground from the background. Then, an objectness prediction module which is composed by the confidence and weight mechanisms is developed to generate the positive and negative weight maps for the objects. The instance property branch and objectness prediction module can provide a coarse-to-fine optimization framework to make our method realize the appearance of the objects. Finally, a positive sample selection strategy is proposed to explore the quality statistical distribution of the positive samples, which are trained by different designed label targets. We evaluate our method on the MS COCO dataset and we achieve 48.4%, 47.9%, 48.0% and 49.3% on ResNet-101, ResNeXt-101, DCN-ResNet-101 and DCN-ResNeXt-101 in terms of AP0.5:0.95, respectively. We evaluate the timing complexity of ALA by calculating the inference speed and the frame per second (FPS) for these four backbones are 11.9, 10.4, 9.9 and 8.0, respectively. The experiment results demonstrate that we can obtain clear improvement over the competing methods with favorable performance compared to the state-of-the-arts.
Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo
IEEE Trans. Circuits Syst. Video Technol.3
2025 WS-SAM: Generalizing SAM to Weakly Supervised Object Detection With Category Label
abstract
Building an effective object detector usually depends on large well-annotated training samples. While annotating such dataset is extremely laborious and costly, where box-level supervision which contains both accurate classification category and localization coordinate is required. Compared to above box-level supervised annotation, those weakly supervised learning manners (e.g,, category, point and scribble) need relatively less laborious annotation cost, and provide a feasible way to mitigate the reliance on the dataset. Because of the lack of sufficient supervised information, current weakly supervised methods cannot achieve satisfactory detection performance. Recently, Segment Anything Model (SAM) has appeared as a task-agnostic foundation model and shown promising performance improvement in many related works due to its powerful generalization and data processing abilities. The properties of the SAM inspire us to adopt such basic benchmark to weakly supervised object detection field to compensate the deficiencies in supervised information. However, directly deploying SAM on weakly supervised object detection task meets with two issues. Firstly, SAM needs meticulously-designed prompts, and such expert-level prompts restrict their applicability and practicality. Besides, SAM is a category unawareness model, and it cannot assign the category labels to the generated predictions. To solve above issues, we propose WS-SAM, which generalizes Segment Anything Model (SAM) to weakly supervised object detection with category label. Specifically, we design an adaptive prompt generator to take full advantages of the spatial and semantic information from the prompt. It employs in a self-prompting manner by taking the output of SAM from the previous iteration as the prompt input to guide the next iteration, where the prompts can be adaptively generated based on the classification activation map. We also develop a segmentation mask refinement module and formulate the label assignment process as a shortest path optimization problem by considering the similarity between each location and prompts. Furthermore, a bidirectional adapter is also implemented to resolve the domain discrepancy by incorporating domain-specific information. We evaluate the effectiveness of our method on several detection datasets (e.g., PASCAL VOC and MS COCO), and the experiment results show that our proposed method can achieve clear improvement over state-of-the-art methods, while performing favorably against state-of-the-arts.
Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo
IEEE Trans. Image Process.3
2025 Fine-Grained Domain Generalization With Feature Structuralization
abstract
Fine-grained domain generalization (FGDG) is a more challenging task than traditional DG tasks due to its small inter-class variations and relatively large intra-class disparities. When domain distribution changes, the vulnerability of subtle features leads to a severe deterioration in model performance. Nevertheless, humans inherently demonstrate the capacity for generalizing to out-of-distribution data, leveraging structured multi-granularity knowledge that emerges from discerning the commonality and specificity within categories. Likewise, we propose a Feature Structuralized Domain Generalization (FSDG) model, wherein features experience structuralization into common, specific, and confounding segments, harmoniously aligned with their relevant semantic concepts, to elevate performance in FGDG. Specifically, feature structuralization (FS) is accomplished through joint optimization of five constraints: a decorrelation function applied to disentangled segments, three constraints ensuring common feature consistency and specific feature distinctiveness, and a prediction calibration term. By imposing these stipulations, FSDG is prompted to disentangle and align features based on multi-granularity knowledge, facilitating robust subtle distinctions among categories. Extensive experimentation on three benchmarks consistently validates the superiority of FSDG over state-of-the-art counterparts, with an average improvement of 6.2% in FGDG performance. Beyond that, the explainability analysis on explicit concept matching intensity between the shared concepts among categories and the model channels, along with experiments on various mainstream model architectures, substantiates the validity of FS.
Wenlong Yu, Dongyue Chen 0001, Qilong Wang 0001, Qinghua Hu
IEEE Trans. Multim.3
2025 Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multitask Learning Perspective
abstract
Human beings can leverage knowledge from relative tasks to improve learning on a primary task. Similarly, multitask learning (MTL) methods suggest using auxiliary tasks to enhance a neural network's performance on a specific primary task. However, previous methods often select auxiliary tasks carefully but treat them as secondary during training. The weights assigned to auxiliary losses are typically smaller than the primary loss weight, leading to insufficient training on auxiliary tasks and ultimately failing to support the main task effectively. To address this issue, we propose an uncertainty-based impartial learning method that ensures balanced training across all tasks. In addition, we consider both gradients and uncertainty information during backpropagation to further improve performance on the primary task. Extensive experiments show that our method achieves performance comparable to or better than state-of-the-art approaches. Moreover, our weighting strategy is effective and robust in enhancing the performance of the primary task regardless of the noise auxiliary tasks' pseudolabels.
Yuanze Li, Chun-Mei Feng 0001, Qilong Wang 0001, Guanglei Yang, Wangmeng Zuo
IEEE Trans. Neural Networks Learn. Syst.3
2025 RD-OpenMax: Rethinking OpenMax for Robust Realistic Open-Set Recognition
abstract
Open-set recognition (OSR) toward a practical open-world setting has attracted increasing research attention in recent years. However, existing OSR settings are either too idealized or focus on specific scenes such as long-tailed distribution and few-shot samples, which fail to capture the complexity of real-world scenarios. In this article, we propose a realistic OSR (ROSR) setting that covers a diverse range of challenging and real-world scenarios, including fine-grained cases with strong semantic correlation and a large number of species, few-shot samples, long-tailed sample distribution, dynamic inputs (e.g., images, spatio-temporal, and multimodal signals) and cross-domain adaptation. In particular, we rethink the simple and basic OpenMax for the ROSR setting and introduce a novel method, regularized discriminative OpenMax (RD-OpenMax), to handle the challenges in the ROSR setting. RD-OpenMax improves upon the basic OpenMax approach by introducing a covariance attention-based covariance pooling (CACP) module as a global aggregation step before the deep architecture's classifier. This module explores rich statistical information on features and provides discriminative distance scores for OpenMax. To address the instability of extreme value theory (EVT) estimation due to insufficient training samples under few-shot and long-tailed scenarios, we propose a regularized EVT (REVT) method based on Monte Carlo sampling to recalibrate the distribution of distance scores. As such, our RD-OpenMax performs a REVT model of distance scores generated by discriminative CACP representations to distinguish known classes and recognize unknown ones effectively and robustly. Extensive experiments are conducted on more than ten visual benchmarks across several scenarios, and the empirical comparisons show that the ROSR setting challenges existing state-of-the-art OSR approaches. Moreover, our RD-OpenMax clearly outperforms its counterparts under the ROSR setting while performing favorably against state-of-the-arts under the traditional OSR setting.
Xiaojie Yin, Bing Cao 0002, Qinghua Hu, Qilong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 CCP-GNN: Competitive Covariance Pooling for Improving Graph Neural Networks
abstract
Graph neural networks (GNNs) have advanced graph classification tasks, where a global pooling to generate graph representations by summarizing node features plays a critical role in the final performance. Most of the existing GNNs are built with a global average pooling (GAP) or its variants, which however, take no full consideration of node specificity while neglecting rich statistics inherent in node features, limiting classification performance of GNNs. Therefore, this article proposes a novel competitive covariance pooling (CCP) based on observation of graph structures, i.e., graphs generally can be identified by a (small) key part of nodes. To this end, our CCP generates node-level second-order representations to explore rich statistics inherent in node features, which are fed to a competitive-based attention module for effectively discovering key nodes through learning node weights. Subsequently, our CCP aggregates node-level second-order representations in conjunction with node weights by summation to produce a covariance representation for each graph, while an iterative matrix normalization is introduced to consider geometry of covariances. Note that our CCP can be flexibly integrated with various GNNs (namely CCP-GNN) to improve the performance of graph classification with little computational cost. The experimental results on seven graph-level benchmarks show that our CCP-GNN is superior or competitive to state-of-the-arts. Our code is available at https://github.com/Jillian555/CCP-GNN.
Pengfei Zhu 0001, Qinghua Hu, Xiao Wang 0017, Qilong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning
abstract
Recently, pre-trained vision-language models (e.g., CLIP) have shown great potential in few-shot learning and attracted a lot of research interest. Although efforts have been made to improve few-shot ability of CLIP, key factors on the effectiveness of existing methods have not been well studied, limiting further exploration of CLIP's potential in few-shot learning. In this paper, we first introduce a uni-fied formulation to analyze CLIP-based few-shot learning methods from a perspective of logit bias, which encourages us to learn an effective logit bias for further improving per-formance of CLIP-based few-shot learning methods. To this end, we disassemble three key components involved in computation of logit bias (i.e., logit features, logit predictor, and logit fusion) and empirically analyze the effect on per-formance of few-shot classification. Based on analysis of key components, this paper proposes a novel AMU-Tuning method to learn effective logit bias for CLIP-based few-shot classification. Specifically, our AMU-Tuning predicts logit bias by exploiting the appropriate Auxiliary features, which are fed into an efficient feature-initialized linear clas-sifier with Multi-branch training. Finally, an Uncertainty-based fusion is developed to incorporate logit bias into CLIP for few-shot classification. The experiments are con-ducted on several widely used benchmarks, and the re-sults show AMU-Tuning clearly outperforms its counter-parts while achieving state-of-the-art performance of CLIP-based few-shot learning without bells and whistles.
Yuwei Tang, Zhenyi Lin, Qilong Wang 0001, Pengfei Zhu 0001, Qinghua Hu
CVPR3
2024 Conditional Controllable Image Fusion
abstract
Image fusion aims to integrate complementary information from multiple input images acquired through various sources to synthesize a new fused image. Existing methods usually employ distinct constraint designs tailored to specific scenes, forming fixed fusion paradigms. However, this data-driven fusion approach is challenging to deploy in varying scenarios, especially in rapidly changing environments. To address this issue, we propose a conditional controllable fusion (CCF) framework for general image fusion tasks without specific training. Due to the dynamic differences of different samples, our CCF employs specific fusion constraints for each individual in practice. Given the powerful generative capabilities of the denoising diffusion model, we first inject the specific constraints into the pre-trained DDPM as adaptive fusion conditions. The appropriate conditions are dynamically selected to ensure the fusion process remains responsive to the specific requirements in each reverse diffusion stage. Thus, CCF enables conditionally calibrating the fused images step by step. Extensive experiments validate our effectiveness in general fusion tasks across diverse scenarios against the competing methods without additional training. The code is publicly available.
Bing Cao 0002, Xingxin Xu, Pengfei Zhu 0001, Qilong Wang 0001, Qinghua Hu
NeurIPS4
2024 Relation Knowledge Distillation by Auxiliary Learning for Object Detection
abstract
Balancing the trade-off between accuracy and speed for obtaining higher performance without sacrificing the inference time is a challenging topic for object detection task. Knowledge distillation, which serves as a kind of model compression techniques, provides a potential and feasible way to handle above efficiency and effectiveness issue through transferring the dark knowledge from the sophisticated teacher detector to the simple student one. Despite demonstrating promising solutions to make harmonies between accuracy and speed, current knowledge distillation for object detection methods still suffer from two limitations. Firstly, most of the methods are inherited or refereed from the frameworks in image classification task, and deploy an implicit manner by imitating or constraining the features from the intermediate layers or the output predictions between the teacher and student models. While little consideration has been raised to the intrinsic relevance of the classification and localization predictions in object detection task. Besides, these methods fail to investigate the relationship between detection and distillation tasks in knowledge distillation pipeline, and they train the whole network by simply integrating losses from these two different tasks through hand-crafted designation parameters. For addressing the aforementioned issues, we propose a novel Relation Knowledge Distillation by Auxiliary Learning for Object Detection (ReAL) method in this paper. Specifically, we first design a prediction relation distillation module which makes the student model directly mimic the output predictions from the teacher one, and conduct self and mutual relation distillation losses to excavate the relation information between teacher and student models. Moreover, for better devolving into the relationship between different tasks in distillation pipeline, we introduce the auxiliary learning into knowledge distillation for object detection and develop a dynamic weight adaptation strategy. Through regarding detection task as primary task and treating distillation task as auxiliary task in auxiliary learning framework, we dynamically adjust and regularize the corresponding weights of the losses for these tasks during the training process. Experiments on MS COCO dataset are conducted using various detector combinations of teacher and student models and the results show that our proposed ReAL can achieve obvious improvement on different distillation model configurations, while performing favorably against state-of-the-arts.
Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo
IEEE Trans. Image Process.3
2024 Layer-Specific Knowledge Distillation for Class Incremental Semantic Segmentation
abstract
Recently, class incremental semantic segmentation (CISS) towards the practical open-world setting has attracted increasing research interest, which is mainly challenged by the well-known issue of catastrophic forgetting. Particularly, knowledge distillation (KD) techniques have been widely studied to alleviate catastrophic forgetting. Despite the promising performance, existing KD-based methods generally use the same distillation schemes for different intermediate layers to transfer old knowledge, while employing manually tuned and fixed trade-off weights to control the effect of KD. These KD-based methods take no consideration of feature characteristics from different intermediate layers, limiting the effectiveness of KD for CISS. In this paper, we propose a layer-specific knowledge distillation (LSKD) method to assign appropriate knowledge schemes and weights for various intermediate layers by considering feature characteristics, aiming to further explore the potential of KD in improving the performance of CISS. Specifically, we present a mask-guided distillation (MD) to alleviate the background shift on semantic features, which performs distillation by masking the features affected by the background. Furthermore, a mask-guided context distillation (MCD) is presented to explore global context information lying in high-level semantic features. Based on them, our LSKD assigns different distillation schemes according to feature characteristics. To adjust the effect of layer-specific distillation adaptively, LSKD introduces a regularized gradient equilibrium method to learn dynamic trade-off weights. Additionally, our LSKD makes an attempt to simultaneously learn distillation schemes and trade-off weights of different layers by developing a bi-level optimization method. Extensive experiments on widely used Pascal VOC 12 and ADE20K show our LSKD clearly outperforms its counterparts while achieving state-of-the-art results.
Qilong Wang 0001, Liu Yang 0010, Wangmeng Zuo, Qinghua Hu
IEEE Trans. Image Process.1
2024 AMS-Net: Modeling Adaptive Multi-Granularity Spatio-Temporal Cues for Video Action Recognition
abstract
Effective spatio-temporal modeling as a core of video representation learning is challenged by complex scale variations in spatio-temporal cues in videos, especially different visual tempos of actions and varying spatial sizes of moving objects. Most of the existing works handle complex spatio-temporal scale variations based on input-level or feature-level pyramid mechanisms, which, however, rely on expensive multistream architectures or explore multiscale spatio-temporal features in a fixed manner. To effectively capture complex scale dynamics of spatio-temporal cues in an efficient way, this article proposes a single-stream architecture (SS-Arch.) with single-input [namely, adaptive multi-granularity spatio-temporal network (AMS-Net)] to model adaptive multi-granularity (Multi-Gran.) Spatio-temporal cues for video action recognition. To this end, our AMS-Net proposes two core components, namely, competitive progressive temporal modeling (CPTM) block and collaborative spatio-temporal pyramid (CSTP) module. They, respectively, capture fine-grained temporal cues and fuse coarse-level spatio-temporal features in an adaptive manner. It admits that AMS-Net can handle subtle variations in visual tempos and fair-sized spatio-temporal dynamics in a unified architecture. Note that our AMS-Net can be flexibly instantiated based on existing deep convolutional neural networks (CNNs) with the proposed CPTM block and CSTP module. The experiments are conducted on eight video benchmarks, and the results show our AMS-Net establishes state-of-the-art (SOTA) performance on fine-grained action recognition (i.e., Diving48 and FineGym), while performing very competitively on widely used Something-Something and Kinetics.
Qilong Wang 0001, Qiyao Hu, Zilin Gao, Peihua Li, Qinghua Hu
IEEE Trans. Neural Networks Learn. Syst.1
2023 Reliable and Interpretable Personalized Federated Learning
abstract
Federated learning can coordinate multiple users to participate in data training while ensuring data privacy. The collaboration of multiple agents allows for a natural connection between federated learning and collective intelligence. When there are large differences in data distribution among clients, it is crucial for federated learning to design a reliable client selection strategy and an interpretable client communication framework to better utilize group knowledge. Herein, a reliable personalized federated learning approach, termed RIPFL, is proposed and fully interpreted from the perspective of social learning. RIPFL reliably selects and divides the clients involved in training such that each client can use different amounts of social information and more effectively communicate with other clients. Simultaneously, the method effectively integrates personal information with the social information generated by the global model from the perspective of Bayesian decision rules and evidence theory, enabling individuals to grow better with the help of collective wisdom. An interpretable federated learning mind is well scalable, and the experimental results indicate that the proposed method has superior robustness and accuracy than other state-of-the-art federated learning algorithms.
Zixuan Qin, Liu Yang 0010, Qilong Wang 0001, Yahong Han, Qinghua Hu
CVPR3
2023 Tuning Pre-trained Model via Moment Probing
abstract
Recently, efficient fine-tuning of large-scale pre-trained models has attracted increasing research interests, where linear probing (LP) as a fundamental module is involved in exploiting the final representations for task-dependent classification. However, most of the existing methods focus on how to effectively introduce a few of learnable parameters, and little work pays attention to the commonly used LP module. In this paper, we propose a novel Moment Probing (MP) method to further explore the potential of LP. Distinguished from LP which builds a linear classification head based on the mean of final features (e.g., word tokens for ViT) or classification tokens, our MP performs a linear classifier on feature distribution, which provides the stronger representation ability by exploiting richer statistical information inherent in features. Specifically, we represent feature distribution by its characteristic function, which is efficiently approximated by using first- and second-order moments of features. Furthermore, we propose a multi-head convolutional cross-covariance (MHC3) to compute second-order moments in an efficient and effective manner. By considering that MP could affect feature learning, we introduce a partially shared module to learn two recalibrating parameters (PSRP) for backbones based on MP, namely MP+. Extensive experiments on ten benchmarks using various models show that our MP significantly outperforms LP and is competitive with counterparts at lower training cost, while our MP+achieves state-of-the-art performance.
Qilong Wang 0001, Zhenyi Lin, Pengfei Zhu 0001, Qinghua Hu
ICCV2
2023 Coarse Helps Fine: A Multi-Granularity Discriminative Adversarial Network for Fine-Grained Open-Set Domain Adaptation
abstract
Open-set domain adaptation (OSDA) aims to align shared classes between the source and the target domain and recognize the private classes of the target domain as unknown. Although unknown classes represent semantic novelty, current OSDA benchmarks lack clear definitions of semantic categories. We propose to use fine-grained visual categorization (FGVC) datasets for the issue because of their specific descriptions of semantic classes. This introduces the new setting named fine-grained OSDA. The entanglement among FGVC, unknown class recognition, and domain adaptation makes fine-grained OSDA a challenging problem. In this paper, we propose a multi-granularity discriminative adversarial network. It utilizes multi-grained labels of the source domain and curriculum learning to improve FGVC performance, exploits discriminative information to recognize unknown classes, and adapts domains through a conditional domain discriminator. Extensive experiments demonstrate our approach outperforms the state-of-the-art methods.
Jing Li 0132, Liu Yang 0010, Qilong Wang 0001, Qinghua Hu
ICME3
2023 Towards a Deeper Understanding of Global Covariance Pooling in Deep Learning: An Optimization Perspective
abstract
Global covariance pooling (GCP) as an effective alternative to global average pooling has shown good capacity to improve deep convolutional neural networks (CNNs) in a variety of vision tasks. Although promising performance, it is still an open problem on how GCP (especially its post-normalization) works in deep learning. In this paper, we make the effort towards understanding the effect of GCP on deep learning from an optimization perspective. Specifically, we first analyze behavior of GCP with matrix power normalization on optimization loss and gradient computation of deep architectures. Our findings show that GCP can improve Lipschitzness of optimization loss and achieve flatter local minima, while improving gradient predictiveness and functioning as a special pre-conditioner on gradients. Then, we explore the effect of post-normalization on GCP from the model optimization perspective, which encourages us to propose a simple yet effective normalization, namely DropCov. Based on above findings, we point out several merits of deep GCP that have not been recognized previously or fully explored, including faster convergence, stronger model robustness and better generalization across tasks. Extensive experimental results using both CNNs and vision transformers on diversified vision tasks provide strong support to our findings while verifying the effectiveness of our method.
Qilong Wang 0001, Jiangtao Xie, Pengfei Zhu 0001, Peihua Li, Wangmeng Zuo, Qinghua Hu
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Multi-head second-order pooling for graph transformer networks
Qilong Wang 0001, Pengfei Zhu 0001
Pattern Recognit. Lett.2
2023 WDAN: A Weighted Discriminative Adversarial Network With Dual Classifiers for Fine-Grained Open-Set Domain Adaptation
abstract
Deep neural networks usually depend on substantial labeled data and suffer from poor generalization to new domains. Domain adaptation can be used to resolve these issues, using a classifier trained with a label-rich source and transferred to a label-scarce target domain. Traditional domain adaptation adopts the close-set assumption that both domains share the same classes. However, real-world applications operate in an open-set scenario where target domains have private categories. This aspect is considered by open-set domain adaptation (OSDA). Nevertheless, current OSDA benchmarks lack clear definitions of semantic classes that are at the core of the open-set concept. In this study, we propose fine-grained visual categorization (FGVC) datasets containing specific descriptions of semantic classes as a solution, introducing the new setting named fine-grained OSDA. Owing to the entanglement among FGVC, unknown class recognition, and domain adaptation, fine-grained OSDA is a challenging task. For this reason, we designed a weighted discriminative adversarial network with dual classifiers (WDAN). It utilizes a selective transformer encoder with overlapping patches and supervised contrastive learning to extract features suitable for FGVC, adversarial training with domain-specific discriminative information to recognize target-private classes, and a weighted conditional domain discriminator to learn domain-invariant features for domain adaptation. Extensive experiments on five benchmarks, including one newly built, demonstrated that WDAN outperforms state-of-the-art methods. This work fills the existing gap in benchmarks for fine-grained OSDA, promoting future developments of real-world applications.
Jing Li 0132, Liu Yang 0010, Qilong Wang 0001, Qinghua Hu
IEEE Trans. Circuits Syst. Video Technol.3
2023 Fully Cascade Consistency Learning for One-Stage Object Detection
abstract
Object detection is usually solved by deploying one single prediction head including classification and localization branches to obtain the final results. Recently proposed works utilize several prediction heads in a cascade learning manner to improve the detection performance. Despite achieving promising performance, existing cascade learning manner methods still meet with two inconsistency issues. Firstly, most of them refine the bounding boxes in different prediction heads only by depending on the localization accuracy (i.e., IoU), while ignoring the inconsistency between classification confidence and localization accuracy. Moreover, simply increasing the IoU threshold by experience to select positive samples makes the inconsistency issue even worse. Secondly, little consideration has been paid on the feature inconsistency between detection-specific features from different prediction heads and detection-generalized ones from backbone model. The extracted feature from backbone model contains the general representation for the whole images. While prediction heads need to be carefully designed to have specific ability which contains more discriminative expressions for the two sub-tasks classification and regression. The different contexture representations of the output features from these two parts lead to the feature inconsistency between backbone model and prediction head in cascade learning architecture. To solve these two inconsistency issues, this paper proposes a novel cascade consistency learning method for one-stage detector. Specifically, a feature adaptation module is firstly developed to calibrate features from different prediction heads and backbone model for solving the feature inconsistency. Then, we design an automatic positive sample threshold selection strategy for further solve the inconsistency between the classification and localization predictions. Moreover, the quality of bounding boxes in cascade learning manner are evaluated by taking both the classification confidence and localization accuracy into consideration. Experiments on MS COCO show that our proposed cascade consistency learning manner (dubbed$\text{C}^{2}\text{L}$) can achieve clear improvement over counterparts based on several different one-stage detectors, while performing favorably against state-of-the-arts.
Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo
IEEE Trans. Circuits Syst. Video Technol.4
2022 Joint Distribution Matters: Deep Brownian Distance Covariance for Few-Shot Classification
abstract
Few-shot classification is a challenging problem as only very few training examples are given for each new task. One of the effective research lines to address this challenge focuses on learning deep representations driven by a similarity measure between a query image and few support images of some class. Statistically, this amounts to measure the dependency of image features, viewed as random vectors in a high-dimensional embedding space. Previous methods either only use marginal distributions without considering joint distributions, suffering from limited representation capability, or are computationally expensive though harnessing joint distributions. In this paper, we propose a deep Brownian Distance Covariance (DeepBDC) method for few-shot classification. The central idea of DeepBDC is to learn image representations by measuring the discrepancy between joint characteristic functions of embedded features and product of the marginals. As the BDC metric is decoupled, we formulate it as a highly modular and efficient layer. Furthermore, we instantiate DeepBDC in two different few-shot classification frameworks. We make experiments on six standard few-shot image benchmarks, covering general object recognition, fine-grained categorization and cross-domain classification. Extensive evaluations show our DeepBDC significantly outperforms the counterparts, while establishing new state-of-the-art results. The source code is available at http://www.peihuali.org/DeepBDC.
Jiangtao Xie, Fei Long 0001, Jiaming Lv, Qilong Wang 0001, Peihua Li
CVPR4
2022 DropCov: A Simple yet Effective Method for Improving Deep Architectures
abstract
Previous works show global covariance pooling (GCP) has great potential to improve deep architectures especially on visual recognition tasks, where post-normalization of GCP plays a very important role in final performance. Although several post-normalization strategies have been studied, these methods pay more close attention to effect of normalization on covariance representations rather than the whole GCP networks, and their effectiveness requires further understanding. Meanwhile, existing effective post-normalization strategies (e.g., matrix power normalization) usually suffer from high computational complexity (e.g., $O(d^{3})$ for $d$-dimensional inputs). To handle above issues, this work first analyzes the effect of post-normalization from the perspective of training GCP networks. Particularly, we for the first time show that \textit{effective post-normalization can make a good trade-off between representation decorrelation and information preservation for GCP, which are crucial to alleviate over-fitting and increase representation ability of deep GCP networks, respectively}. Based on this finding, we can improve existing post-normalization methods with some small modifications, providing further support to our observation. Furthermore, this finding encourages us to propose a novel pre-normalization method for GCP (namely DropCov), which develops an adaptive channel dropout on features right before GCP, aiming to reach trade-off between representation decorrelation and information preservation in a more efficient way. Our DropCov only has a linear complexity of $O(d)$, while being free for inference. Extensive experiments on various benchmarks (i.e., ImageNet-1K, ImageNet-C, ImageNet-A, Stylized-ImageNet, and iNat2017) show our DropCov is superior to the counterparts in terms of efficiency and effectiveness, and provides a simple yet effective method to improve performance of deep architectures involving both deep convolutional neural networks (CNNs) and vision transformers (ViTs).
Qilong Wang 0001, Jiangtao Xie, Peihua Li, Qinghua Hu
NeurIPS1
2022 Temporal grafter network: Rethinking LSTM for effective video recognition
Bingbing Zhang 0001, Qilong Wang 0001, Zilin Gao, Ruiren Zeng, Peihua Li
Neurocomputing2
2022 Delving Deeper Into Pixel Prior for Box-Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) based on bounding box annotations has attracted considerable recent attention and has achieved promising performance. However, most of existing methods focus on generation of high-quality pseudo labels for segmented objects using box indicators, but they fail to fully explore and exploit prior from bounding box annotations, which limits performance of WSSS methods, especially for fine parts and boundaries. To overcome above issues, this paper proposes a novel Pixel-as-Instance Prior (PIP) for WSSS methods by delving deeper into pixel prior from bounding box annotations. Specifically, the proposed PIP is built on two important observations on pixels around bounding boxes. First, since objects are usually irregularity and tightly close to bounding boxes (dubbed irregular-filling prior), so each row or column of bounding boxes basically have at least one pixel belonging to foreground objects and background, respectively. Second, pixels near the bounding boxes tend to be highly ambiguous and more difficult to classify (dubbed label-ambiguity prior). To implement our PIP, a constrained loss alike multiple instance learning (MIL) and a labeling-balance loss are developed to jointly train WSSS models, which regards each pixel as a weighted positive or negative instance while considering more effective prior (i.e., irregular-filling and label-ambiguity priors) from bounding box annotations in an efficient way. Note that our PIP can be flexibly integrated with various WSSS methods, while clearly improving their performance with negligible computational overload in training stage. The experiments are conducted on most widely used PASCAL VOC 2012 and Cityscapes benchmarks, and the results show that our PIP has a good ability to improve performance of various WSSS methods, while achieving very competitive results.
Tianqi Ma, Qilong Wang 0001, Wangmeng Zuo
IEEE Trans. Image Process.2
2022 CrabNet: Fully Task-Specific Feature Learning for One-Stage Object Detection
abstract
Object detection is usually solved by learning a deep architecture involving classification and localization tasks, where feature learning for these two tasks is shared using the same backbone model. Recent works have shown that suitable disentanglement of classification and localization tasks has the great potential to improve performance of object detection. Despite the promising performance, existing feature disentanglement methods usually suffer from two limitations. First, most of them only focus on the disentangled proposals or predication heads for classification and localization tasks after RPN. While little consideration has been given to that the features for these two different tasks actually are obtained by a shared backbone model before RPN. Second, they are suggested for two-stage objectors and are not applicable to one-stage methods. To overcome these limitations, this paper presents a novel fully task-specific feature learning method for one-stage object detection. Specifically, our method first learns disentangled features for classification and localization tasks using two separated backbone models, where auxiliary classification and localization heads are inserted at the end of the two backbone models for providing a fully task-specific features for classification and localization. Then, a feature interaction module is developed for aligning and fusing task-specific features, which are further used to produce the final detection result. Experiments on MS COCO show that our proposed method (dubbed CrabNet) can achieve clear improvement over counterparts with increasing limited inference time, while performing favorably against state-of-the-arts.
Hao Wang 0073, Qilong Wang 0001, Qinghua Hu, Wangmeng Zuo
IEEE Trans. Image Process.2
2021 Detection, Tracking, and Counting Meets Drones in Crowds: A Benchmark
abstract
To promote the developments of object detection, tracking and counting algorithms in drone-captured videos, we construct a benchmark with a new drone-captured large-scale dataset, named as DroneCrowd, formed by 112 video clips with 33, 600 HD frames in various scenarios. Notably, we annotate 20, 800 people trajectories with 4.8 million heads and several video-level attributes. Meanwhile, we design the Space-Time Neighbor-Aware Network (STNNet) as a strong baseline to solve object detection, tracking and counting jointly in dense crowds. STNNet is formed by the feature extraction module, followed by the density map estimation heads, and localization and association subnets. To exploit the context information of neighboring objects, we design the neighboring context loss to guide the association subnet training, which enforces consistent relative position of nearby objects in temporal domain. Extensive experiments on our DroneCrowd dataset demonstrate that STNNet performs favorably against the state-of-the-arts.
Longyin Wen, Dawei Du, Pengfei Zhu 0001, Qinghua Hu, Qilong Wang 0001, Liefeng Bo, Siwei Lyu
CVPR5
2021 Boosting Weakly Supervised Object Detection via Learning Bounding Box Adjusters
abstract
Weakly-supervised object detection (WSOD) has emerged as an inspiring recent topic to avoid expensive instance-level object annotations. However, the bounding boxes of most existing WSOD methods are mainly determined by precomputed proposals, thereby being limited in precise object localization. In this paper, we defend the problem setting for improving localization performance by leveraging the bounding box regression knowledge from a well-annotated auxiliary dataset. First, we use the well-annotated auxiliary dataset to explore a series of learnable bounding box adjusters (LBBAs) in a multi-stage training manner, which is class-agnostic. Then, only LBBAs and a weakly-annotated dataset with non-overlapped classes are used for training LBBA-boosted WSOD. As such, our LBBAs are practically more convenient and economical to implement while avoiding the leakage of the auxiliary well-annotated dataset. In particular, we formulate learning bounding box adjusters as a bi-level optimization problem and suggest an EM-like multi-stage training algorithm. Then, a multi-stage scheme is further presented for LBBA-boosted WSOD. Additionally, a masking strategy is adopted to improve proposal classification. Experimental results verify the effectiveness of our method. Our method performs favorably against state-of-the-art WSOD methods and knowledge transfer model with similar problem setting. Code is publicly available at https://github.com/DongSky/lbba_boosted_wsod.
Bowen Dong 0001, Zitong Huang, Yuelin Guo, Qilong Wang 0001, Zhenxing Niu, Wangmeng Zuo
ICCV4
2021 Temporal-attentive Covariance Pooling Networks for Video Recognition
abstract
For video recognition task, a global representation summarizing the whole contents of the video snippets plays an important role for the final performance. However, existing video architectures usually generate it by using a simple, global average pooling (GAP) method, which has limited ability to capture complex dynamics of videos. For image recognition task, there exist evidences showing that covariance pooling has stronger representation ability than GAP. Unfortunately, such plain covariance pooling used in image recognition is an orderless representative, which cannot model spatio-temporal structure inherent in videos. Therefore, this paper proposes a Temporal-attentive Covariance Pooling (TCP), inserted at the end of deep architectures, to produce powerful video representations. Specifically, our TCP first develops a temporal attention module to adaptively calibrate spatio-temporal features for the succeeding covariance pooling, approximatively producing attentive covariance representations. Then, a temporal covariance pooling performs temporal pooling of the attentive covariance representations to characterize both intra-frame correlations and inter-frame cross-correlations of the calibrated features. As such, the proposed TCP can capture complex temporal dynamics. Finally, a fast matrix power normalization is introduced to exploit geometry of covariance representations. Note that our TCP is model-agnostic and can be flexibly integrated into any video architectures, resulting in TCPNet for effective video recognition. The extensive experiments on six benchmarks (e.g., Kinetics, Something-Something V1 and Charades) using various video architectures show our TCPNet is clearly superior to its counterparts, while having strong generalization ability. The source code is publicly available.
Zilin Gao, Qilong Wang 0001, Bingbing Zhang 0001, Qinghua Hu, Peihua Li
NeurIPS2
2021 ALGeNet: Adaptive Log-Euclidean Gaussian embedding network for time series forecasting
Zongxia Xie, Qilong Wang 0001, Renhui Li
Neurocomputing3
2021 Deep CNNs Meet Global Covariance Pooling: Better Representation and Generalization
abstract
Compared with global average pooling in existing deep convolutional neural networks (CNNs), global covariance pooling can capture richer statistics of deep features, having potential for improving representation and generalization abilities of deep CNNs. However, integration of global covariance pooling into deep CNNs brings two challenges: (1) robust covariance estimation given deep features of high dimension and small sample size; (2) appropriate usage of geometry of covariances. To address these challenges, we propose a global Matrix Power Normalized COVariance (MPN-COV) Pooling. Our MPN-COV conforms to a robust covariance estimator, very suitable for scenario of high dimension and small sample size. It can also be regarded as Power-Euclidean metric between covariances, effectively exploiting their geometry. Furthermore, a global Gaussian embedding network is proposed to incorporate first-order statistics into MPN-COV. For fast training of MPN-COV networks, we implement an iterative matrix square root normalization, avoiding GPU unfriendly eigen-decomposition inherent in MPN-COV. Additionally, progressive 1×1 convolutions and group convolution are introduced to compress covariance representations. The proposed methods are highly modular, readily plugged into existing deep CNNs. Extensive experiments are conducted on large-scale object classification, scene categorization, fine-grained visual recognition and texture classification, showing our methods outperform the counterparts and obtain state-of-the-art performance.
Qilong Wang 0001, Jiangtao Xie, Wangmeng Zuo, Lei Zhang 0006, Peihua Li
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Multi-scale structural kernel representation for object detection
Hao Wang 0073, Qilong Wang 0001, Peihua Li, Wangmeng Zuo
Pattern Recognit.2
2021 Constrained Online Cut-Paste for Object Detection
abstract
Well-annotated training samples show necessity in achieving high performance of object detection, but collection of massive samples is extremely laborious and costly. Recently, cut-paste based methods show the potential to augment the training samples by cutting the foreground instances and pasting them on some background regions. However, existing cut-paste based methods hardly guarantee the quality of synthetic images due to lack of mechanism to ensure rationality of the pasted instances (e.g., context, geometry and diversity), limiting the effectiveness of data augmentation. To overcome above issues, this paper proposes a novel Constrained Online Cut-Paste (COCP) method, making an attempt to effectively and efficiently augment training data for improving performance of object detection. Specifically, our COCP generates synthetic images by switching instances of same class from various image pairs in each training mini-batch, ensuring context coherence between the cut instances and the pasted backgrounds. Furthermore, two constraints based on geometric consistency and sample diversity are developed to eliminate counterproductive and meaningless switched instances those suffer from significant geometric discrepancy or lack variations, further improving quality of the synthetic images. The experiments are conducted on both MS COCO and PASCAL VOC datasets using various state-of-the-art detectors (e.g., Faster R-CNN, RetinaNet, FCOS and Mask R-CNN). The results show that our proposed COCP can be well generalized to various datasets and detectors with clear performance gains, while performing favorably against its counterparts.
Hao Wang 0073, Qilong Wang 0001, Jian Yang 0003, Wangmeng Zuo
IEEE Trans. Circuits Syst. Video Technol.2
2020 Neural Blind Deconvolution Using Deep Priors
abstract
Blind deconvolution is a classical yet challenging low-level vision problem with many real-world applications. Traditional maximum a posterior (MAP) based methods rely heavily on fixed and handcrafted priors that certainly are insufficient in characterizing clean images and blur kernels, and usually adopt specially designed alternating minimization to avoid trivial solution. In contrast, existing deep motion deblurring networks learn from massive training images the mapping to clean image or blur kernel, but are limited in handling various complex and large size blur kernels. To connect MAP and deep models, we in this paper present two generative networks for respectively modeling the deep priors of clean image and blur kernel, and propose an unconstrained neural optimization solution to blind deconvolution. In particular, we adopt an asymmetric Autoencoder with skip connections for generating latent clean image, and a fully-connected network (FCN) for generating blur kernel. Moreover, the SoftMax nonlinearity is applied to the output layer of FCN to meet the non-negative and equality constraints. The process of neural optimization can be explained as a kind of ''zero-shot" self-supervised learning of the generative networks, and thus our proposed method is dubbed SelfDeblur. Experimental results show that our SelfDeblur can achieve notable quantitative gains as well as more visually plausible deblurring results in comparison to state-of-the-art blind deconvolution methods on benchmark datasets and real-world blurry images. The source code is publicly available at https://github.com/csdwren/SelfDeblur.
Dongwei Ren, Kai Zhang 0008, Qilong Wang 0001, Qinghua Hu, Wangmeng Zuo
CVPR3
2020 ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks
abstract
Recently, channel attention mechanism has demonstrated to offer great potential in improving the performance of deep convolutional neural networks (CNNs). However, most existing methods dedicate to developing more sophisticated attention modules for achieving better performance, which inevitably increase model complexity. To overcome the paradox of performance and complexity trade-off, this paper proposes an Efficient Channel Attention (ECA) module, which only involves a handful of parameters while bringing clear performance gain. By dissecting the channel attention module in SENet, we empirically show avoiding dimensionality reduction is important for learning channel attention, and appropriate cross-channel interaction can preserve performance while significantly decreasing model complexity. Therefore, we propose a local cross-channel interaction strategy without dimensionality reduction, which can be efficiently implemented via 1D convolution. Furthermore, we develop a method to adaptively select kernel size of 1D convolution, determining coverage of local cross-channel interaction. The proposed ECA module is both efficient and effective, e.g., the parameters and computations of our modules against backbone of ResNet50 are 80 vs. 24.37M and 4.7e-4 GFlops vs. 3.86 GFlops, respectively, and the performance boost is more than 2% in terms of Top-1 accuracy. We extensively evaluate our ECA module on image classification, object detection and instance segmentation with backbones of ResNets and MobileNetV2. The experimental results show our module is more efficient while performing favorably against its counterparts.
Qilong Wang 0001, Banggu Wu, Pengfei Zhu 0001, Peihua Li, Wangmeng Zuo, Qinghua Hu
CVPR1
2020 What Deep CNNs Benefit From Global Covariance Pooling: An Optimization Perspective
abstract
Recent works have demonstrated that global covariance pooling (GCP) has the ability to improve performance of deep convolutional neural networks (CNNs) on visual classification task. Despite considerable advance, the reasons on effectiveness of GCP on deep CNNs have not been well studied. In this paper, we make an attempt to understand what deep CNNs benefit from GCP in a viewpoint of optimization. Specifically, we explore the effect of GCP on deep CNNs in terms of the Lipschitzness of optimization loss and the predictiveness of gradients, and show that GCP can make the optimization landscape more smooth and the gradients more predictive. Furthermore, we discuss the connection between GCP and second-order optimization for deep CNNs. More importantly, above findings can account for several merits of covariance pooling for training deep CNNs that have not been recognized previously or fully explored, including significant acceleration of network convergence (i.e., the networks trained with GCP can support rapid decay of learning rates, achieving favorable performance while significantly reducing number of training epochs), stronger robustness to distorted examples generated by image corruptions and perturbations, and good generalization ability to different vision tasks, e.g., object detection and instance segmentation. We conduct extensive experiments using various deep CNN architectures on diversified tasks, and the results provide strong support to our findings.
Qilong Wang 0001, Banggu Wu, Dongwei Ren, Peihua Li, Wangmeng Zuo, Qinghua Hu
CVPR1
2020 Learning second-order statistics for place recognition based on robust covariance estimation of CNN features
Zifei Yan, Qilong Wang 0001, Xiaohe Wu, Wangmeng Zuo
Neurocomputing3
2020 Locality-constrained affine subspace coding for image classification and retrieval
Bingbing Zhang 0001, Qilong Wang 0001, Xiaoxiao Lu, Fasheng Wang, Peihua Li
Pattern Recognit.2
2020 Weighted and Class-Specific Maximum Mean Discrepancy for Unsupervised Domain Adaptation
abstract
Although maximum mean discrepancy (MMD) has achieved great success in unsupervised domain adaptation (UDA), most of existing UDA methods ignore the issue of class weight bias across domains, which is ubiquitous and evidently gives rise to the degradation of UDA performance. In this work, we propose two improved MMD metrics, i.e., weighted MMD (WMMD) and class-specific MMD (CMMD), to alleviate the adverse effect caused by the changes of class prior distributions between source and target domains. In WMMD, class-specific auxiliary weights are deployed to reweigh the source samples. In CMMD, we calculate the MMD for each class of source and target samples. Since the class labels of target samples are unknown for UDA problem, we present a classification expectation-maximization algorithm to estimate the pseudo-labels of target samples on the fly and update the model parameters using estimated labels. The proposed methods can be flexibly incorporated into deep convolutional neural networks to form WMMD and CMMD based domain adaptation networks, which we called WDAN and CDAN, respectively. By combining WMMD with CMMD, we present a CWMMD based domain adaptation network (CWDAN) to further improve classification performance. Experiments show that, both WMMD and CMMD benefit the classification accuracy, and our CWDAN can achieve compelling UDA performance in comparison with MMD and the state-of-the-art UDA methods.
Hongliang Yan, Zhetao Li, Qilong Wang 0001, Peihua Li, Yong Xu 0001, Wangmeng Zuo
IEEE Trans. Multim.3
2019 Global Second-Order Pooling Convolutional Networks
abstract
Deep Convolutional Networks (ConvNets) are fundamental to, besides large-scale visual recognition, a lot of vision tasks. As the primary goal of the ConvNets is to characterize complex boundaries of thousands of classes in a high-dimensional space, it is critical to learn higher-order representations for enhancing non-linear modeling capability. Recently, Global Second-order Pooling (GSoP), plugged at the end of networks, has attracted increasing attentions, achieving much better performance than classical, first-order networks in a variety of vision tasks. However, how to effectively introduce higher-order representation in earlier layers for improving non-linear capability of ConvNets is still an open problem. In this paper, we propose a novel network model introducing GSoP across from lower to higher layers for exploiting holistic image information throughout a network. Given an input 3D tensor outputted by some previous convolutional layer, we perform GSoP to obtain a covariance matrix which, after nonlinear transformation, is used for tensor scaling along channel dimension. Similarly, we can perform GSoP along spatial dimension for tensor scaling as well. In this way, we can make full use of the second-order statistics of the holistic image throughout a network. The proposed networks are thoroughly evaluated on large-scale ImageNet-1K, and experiments have shown that they outperform non-trivially the counterparts while achieving state-of-the-art results.
Zilin Gao, Jiangtao Xie, Qilong Wang 0001, Peihua Li
CVPR3
2019 Deep Global Generalized Gaussian Networks
abstract
Recently, global covariance pooling (GCP) has shown great advance in improving classification performance of deep convolutional neural networks (CNNs). However, existing deep GCP networks compute covariance pooling of convolutional activations with assumption that activations are sampled from Gaussian distributions, which may not hold in practice and fails to fully characterize the statistics of activations. To handle this issue, this paper proposes a novel deep global generalized Gaussian network (3G-Net), whose core is to estimate a global covariance of generalized Gaussian for modeling the last convolutional activations. Compared with GCP in Gaussian setting, our 3G-Net assumes the distribution of activations follows a generalized Gaussian, which can capture more precise characteristics of activations. However, there exists no analytic solution for parameter estimation of generalized Gaussian, making our 3G-Net challenging. To this end, we first present a novel regularized maximum likelihood estimator for robust estimating covariance of generalized Gaussian, which can be optimized by a modified iterative re-weighted method. Then, to efficiently estimate the covariance of generaized Gaussian under deep CNN architectures, we approximate this re-weighted method by developing an unrolling re-weighted module and a square root covariance layer. In this way, 3GNet can be flexibly trained in an end-to-end manner. The experiments are conducted on large-scale ImageNet-1K and Places365 datasets, and the results demonstrate our 3G-Net outperforms its counterparts while achieving very competitive performance to state-of-the-arts.
Qilong Wang 0001, Peihua Li, Qinghua Hu, Pengfei Zhu 0001, Wangmeng Zuo
CVPR1
2018 Towards Faster Training of Global Covariance Pooling Networks by Iterative Matrix Square Root Normalization
abstract
Global covariance pooling in convolutional neural networks has achieved impressive improvement over the classical first-order pooling. Recent works have shown matrix square root normalization plays a central role in achieving state-of-the-art performance. However, existing methods depend heavily on eigendecomposition (EIG) or singular value decomposition (SVD), suffering from inefficient training due to limited support of EIG and SVD on GPU. Towards addressing this problem, we propose an iterative matrix square root normalization method for fast end-to-end training of global covariance pooling networks. At the core of our method is a meta-layer designed with loop-embedded directed graph structure. The meta-layer consists of three consecutive nonlinear structured layers, which perform pre-normalization, coupled matrix iteration and post-compensation, respectively. Our method is much faster than EIG or SVD based ones, since it involves only matrix multiplications, suitable for parallel implementation on GPU. Moreover, the proposed network with ResNet architecture can converge in much less epochs, further accelerating network training. On large-scale ImageNet, we achieve competitive performance superior to existing counterparts. By fine-tuning our models pre-trained on ImageNet, we establish state-of-the-art results on three challenging fine-grained benchmarks. The source code and network models will be available at http://www.peihuali.org/iSQRT-COV.
Peihua Li, Jiangtao Xie, Qilong Wang 0001, Zilin Gao
CVPR3
2018 Multi-Scale Location-Aware Kernel Representation for Object Detection
abstract
Although Faster R-CNN and its variants have shown promising performance in object detection, they only exploit simple first-order representation of object proposals for final classification and regression. Recent classification methods demonstrate that the integration of high-order statistics into deep convolutional neural networks can achieve impressive improvement, but their goal is to model whole images by discarding location information so that they cannot be directly adopted to object detection. In this paper, we make an attempt to exploit high-order statistics in object detection, aiming at generating more discriminative representations for proposals to enhance the performance of detectors. To this end, we propose a novel Multi-scale Location-aware Kernel Representation (MLKP) to capture high-order statistics of deep features in proposals. Our MLKP can be efficiently computed on a modified multi-scale feature map using a low-dimensional polynomial kernel approximation. Moreover, different from existing orderless global representations based on high-order statistics, our proposed MLKP is location retentive and sensitive so that it can be flexibly adopted to object detection. Through integrating into Faster R-CNN schema, the proposed MLKP achieves very competitive performance with state-of-the-art methods, and improves Faster R-CNN by 4.9% (mAP), 4.7% (mAP) and 5.0% (AP at IOU=[0.5:0.05:0.95]) on PASCAL VOC 2007, VOC 2012 and MS COCO benchmarks, respectively. Code is available at: https://github.com/Hwang64/MLKP.
Hao Wang 0073, Qilong Wang 0001, Mingqi Gao 0006, Peihua Li, Wangmeng Zuo
CVPR2
2018 Support Vector Metric Learning on Symmetric Positive Definite Manifold
abstract
The manifold of symmetric positive definite (SPD) matrices has drawn significant attention because of its widespread applications. SPD matrices provide compact nonlinear representations of data and form a special type of Riemannian manifold. The direct application of support vector machines on SPD manifold maybe fails due to lack of samples per class. In this paper, we propose a support vector metric learning (SVML) model on SPD manifold. We define a positive definite kernel for point pairs on SPD manifold and transform metric learning on SPD manifold to a point pair classification problem. The metric learning problem can be efficiently solved by standard support vector machines. Compared with classifying points on SPD manifold by support vector machines directly, SVML effectively learns a distance metric for SPD matrices by training a binary support vector machine model. Experiments on video based face recognition, image set classification, and material classification show that SVML outperforms the state-of-the-art metric learning algorithms on SPD manifold.
Hao Cheng 0010, Pengfei Zhu 0001, Qilong Wang 0001, Changqing Zhang 0002, Qinghua Hu
ICME3
2018 Towards Generalized and Efficient Metric Learning on Riemannian Manifold
abstract
Modeling data as points on non-linear Riemannian manifold has attracted increasing attentions in many computer vision tasks, especially visual recognition. Learning an appropriate metric on Riemannian manifold plays a key role in achieving promising performance. For widely used symmetric positive definite (SPD) manifold and Grassmann manifold, most of existing metric learning methods are designed for one manifold, and are not straightforward for the other one. Furthermore, optimizations in previous methods usually rely on computationally expensive iterations. To address above limitations, this paper makes an attempt to propose a generalized and efficient Riemannian manifold metric learning (RMML) method, which can be flexibly adopted to both SPD and Grassmann manifolds. By minimizing the geodesic distance of similar pairs and the interpoint geodesic distance of dissimilar ones on nonlinear manifolds, the proposed RMML is optimized by computing the geodesic mean between inverse of similarity matrix and dissimilarity matrix, benefiting a global closed-form solution and high efficiency. The experiments are conducted on various visual recognition tasks, and the results demonstrate our RMML performs favorably against its counterparts in terms of both accuracy and efficiency.
Pengfei Zhu 0001, Hao Cheng 0010, Qinghua Hu, Qilong Wang 0001, Changqing Zhang 0002
IJCAI4
2018 Beyond Similar and Dissimilar Relations : A Kernel Regression Formulation for Metric Learning
abstract
Most existing metric learning methods focus on learning a similarity or distance measure relying on similar and dissimilar relations between sample pairs. However, pairs of samples cannot be simply identified as similar or dissimilar in many real-world applications, e.g., multi-label learning, label distribution learning or tasks with continuous decision values. To this end, in this paper we propose a novel relation alignment metric learning (RAML) formulation to handle the metric learning problem in those scenarios. Since the relation of two samples can be measured by the difference degree of the decision values, motivated by the consistency of the sample relations in the feature space and decision space, our proposed RAML utilizes the sample relations in the decision space to guide the metric learning in the feature space. Specifically, our RAML method formulates metric learning as a kernel regression problem, which can be efficiently optimized by the standard regression solvers. We carry out several experiments on the single-label classification, multi-label classification, and label distribution learning tasks, to demonstrate that our method achieves favorable performance against the state-of-the-art methods.
Pengfei Zhu 0001, Ren Qi, Qinghua Hu, Qilong Wang 0001, Changqing Zhang 0002, Liu Yang 0010
IJCAI4
2018 Global Gated Mixture of Second-order Pooling for Improving Deep Convolutional Neural Networks
abstract
In most of existing deep convolutional neural networks (CNNs) for classification, global average (first-order) pooling (GAP) has become a standard module to summarize activations of the last convolution layer as final representation for prediction. Recent researches show integration of higher-order pooling (HOP) methods clearly improves performance of deep CNNs. However, both GAP and existing HOP methods assume unimodal distributions, which cannot fully capture statistics of convolutional activations, limiting representation ability of deep CNNs, especially for samples with complex contents. To overcome the above limitation, this paper proposes a global Gated Mixture of Second-order Pooling (GM-SOP) method to further improve representation ability of deep CNNs. To this end, we introduce a sparsity-constrained gating mechanism and propose a novel parametric SOP as component of mixture model. Given a bank of SOP candidates, our method can adaptively choose Top-K (K > 1) candidates for each input sample through the sparsity-constrained gating module, and performs weighted sum of outputs of K selected candidates as representation of the sample. The proposed GM-SOP can flexibly accommodate a large number of personalized SOP candidates in an efficient way, leading to richer representations. The deep networks with our GM-SOP can be end-to-end trained, having potential to characterize complex, multi-modal distributions. The proposed method is evaluated on two large scale image benchmarks (i.e., downsampled ImageNet-1K and Places365), and experimental results show our GM-SOP is superior to its counterparts and achieves very competitive performance. The source code will be available at http://www.peihuali.org/GM-SOP.
Qilong Wang 0001, Zilin Gao, Jiangtao Xie, Wangmeng Zuo, Peihua Li
NeurIPS1
2018 Hyperlayer Bilinear Pooling with application to fine-grained categorization and image retrieval
Qiule Sun, Qilong Wang 0001, Jianxin Zhang 0001, Peihua Li
Neurocomputing2
2018 An Information Geometry-Based Distance Between High-Dimensional Covariances for Scalable Classification
abstract
Modeling images/videos with covariance matrices has attracted increasing attentions in various vision tasks, especially in visual classification. For covariances-based visual classification, measuring the distances between covariances is one of the key issues and has been studied for decades. Since the space of covariances is a Riemannian manifold, the geometrical structure of covariances should be favorably considered when designing distance metrics. Although this problem has been widely studied, designing an effective and efficient metric between high-dimensional covariances (HDCOV) for scalable classification is still an open problem. In this paper, we present an information geometry-based distance (IGBD) to tackle this challenge from the perspective of information geometry. Our idea is based on the fact that each covariance can be viewed as a zero-mean Gaussian distribution, and thus the distances between covariances are measured by those between the corresponding Gaussian distributions. The core of our method is to project each distribution, in the form of a set of random samples, to a vector on the tangent space of a common, known distribution on the statistical manifold, based on Fisher information metric and maximum likelihood method. On the tangent space, the Euclidean norm can be used to measure the distances between those sets of projection vectors (or equivalently distributions). The proposed IGBD for HDCOV is computationally efficient and easily combined with a linear support vector machine, suitable for scalable visual classification. The experiments are conducted on various kinds and sizes of benchmarks, and results show the proposed method is efficient and the combination of HDCOV can achieve very competitive performance.
Qilong Wang 0001, Xiaoxiao Lu, Peihua Li, Zhenguo Gao, Yongri Piao
IEEE Trans. Circuits Syst. Video Technol.1
2017 G2DeNet: Global Gaussian Distribution Embedding Network and Its Application to Visual Recognition
abstract
Recently, plugging trainable structural layers into deep convolutional neural networks (CNNs) as image representations has made promising progress. However, there has been little work on inserting parametric probability distributions, which can effectively model feature statistics, into deep CNNs in an end-to-end manner. This paper proposes a Global Gaussian Distribution embedding Network (G2DeNet) to take a step towards addressing this problem. The core of G2DeNet is a novel trainable layer of a global Gaussian as an image representation plugged into deep CNNs for end-to-end learning. The challenge is that the proposed layer involves Gaussian distributions whose space is not a linear space, which makes its forward and backward propagations be non-intuitive and non-trivial. To tackle this issue, we employ a Gaussian embedding strategy which respects the structures of both Riemannian manifold and smooth group of Gaussians. Based on this strategy, we construct the proposed global Gaussian embedding layer and decompose it into two sub-layers: the matrix partition sub-layer decoupling the mean vector and covariance matrix entangled in the embedding matrix, and the square-rooted, symmetric positive definite matrix sub-layer. In this way, we can derive the partial derivatives associated with the proposed structural layer and thus allow backpropagation of gradients. Experimental results on large scale region classification and fine-grained recognition tasks show that G2DeNet is superior to its counterparts, capable of achieving state-of-the-art performance.
Qilong Wang 0001, Peihua Li, Lei Zhang 0006
CVPR1
2017 Mind the Class Weight Bias: Weighted Maximum Mean Discrepancy for Unsupervised Domain Adaptation
abstract
In domain adaptation, maximum mean discrepancy (MMD) has been widely adopted as a discrepancy metric between the distributions of source and target domains. However, existing MMD-based domain adaptation methods generally ignore the changes of class prior distributions, i.e., class weight bias across domains. This remains an open problem but ubiquitous for domain adaptation, which can be caused by changes in sample selection criteria and application scenarios. We show that MMD cannot account for class weight bias and results in degraded domain adaptation performance. To address this issue, a weighted MMD model is proposed in this paper. Specifically, we introduce class-specific auxiliary weights into the original MMD for exploiting the class prior probability on source and target domains, whose challenge lies in the fact that the class label in target domain is unavailable. To account for it, our proposed weighted MMD model is defined by introducing an auxiliary weight for each class in the source domain, and a classification EM algorithm is suggested by alternating between assigning the pseudo-labels, estimating auxiliary weights and updating model parameters. Extensive experiments demonstrate the superiority of our weighted MMD over conventional MMD for domain adaptation.
Hongliang Yan, Yukang Ding, Peihua Li, Qilong Wang 0001, Yong Xu 0001, Wangmeng Zuo
CVPR4
2017 Is Second-Order Information Helpful for Large-Scale Visual Recognition?
abstract
By stacking layers of convolution and nonlinearity, convolutional networks (ConvNets) effectively learn from lowlevel to high-level features and discriminative representations. Since the end goal of large-scale recognition is to delineate complex boundaries of thousands of classes, adequate exploration of feature distributions is important for realizing full potentials of ConvNets. However, state-of-the-art works concentrate only on deeper or wider architecture design, while rarely exploring feature statistics higher than first-order. We take a step towards addressing this problem. Our method consists in covariance pooling, instead of the most commonly used first-order pooling, of highlevel convolutional features. The main challenges involved are robust covariance estimation given a small sample of large-dimensional features and usage of the manifold structure of covariance matrices. To address these challenges, we present a Matrix Power Normalized Covariance (MPNCOV) method. We develop forward and backward propagation formulas regarding the nonlinear matrix functions such that MPN-COV can be trained end-to-end. In addition, we analyze both qualitatively and quantitatively its advantage over the well-known Log-Euclidean metric. On the ImageNet 2012 validation set, by combining MPN-COV we achieve over 4%, 3% and 2.5% gains for AlexNet, VGG-M and VGG-16, respectively; integration of MPN-COV into 50-layer ResNet outperforms ResNet-101 and is comparable to ResNet-152. The source code will be available on the project page: http://www.peihuali.org/MPN-COV.
Peihua Li, Jiangtao Xie, Qilong Wang 0001, Wangmeng Zuo
ICCV3
2017 Joint distance and similarity measure learning based on triplet-based constraints
Mu Li 0005, Qilong Wang 0001, David Zhang 0001, Peihua Li, Wangmeng Zuo
Inf. Sci.2
2017 Local Log-Euclidean Multivariate Gaussian Descriptor and Its Application to Image Classification
abstract
This paper presents a novel image descriptor to effectively characterize the local, high-order image statistics. Our work is inspired by the Diffusion Tensor Imaging and the structure tensor method (or covariance descriptor), and motivated by popular distribution-based descriptors such as SIFT and HoG. Our idea is to associate one pixel with a multivariate Gaussian distribution estimated in the neighborhood. The challenge lies in that the space of Gaussians is not a linear space but a Riemannian manifold. We show, for the first time to our knowledge, that the space of Gaussians can be equipped with a Lie group structure by defining a multiplication operation on this manifold, and that it is isomorphic to a subgroup of the upper triangular matrix group. Furthermore, we propose methods to embed this matrix group in the linear space, which enables us to handle Gaussians with Euclidean operations rather than complicated Riemannian operations. The resulting descriptor, called Local Log-Euclidean Multivariate Gaussian (L2EMG) descriptor, works well with low-dimensional and high-dimensional raw features. Moreover, our descriptor is a continuous function of features without quantization, which can model the first- and second-order statistics. Extensive experiments were conducted to evaluate thoroughly L2EMG, and the results showed that L2EMG is very competitive with state-of-the-art descriptors in image classification.
Peihua Li, Qilong Wang 0001, Hui Zeng 0001, Lei Zhang 0006
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 High-Order Local Pooling and Encoding Gaussians Over a Dictionary of Gaussians
abstract
Local pooling (LP) in configuration (feature) space proposed by Boureau et al. explicitly restricts similar features to be aggregated, which can preserve as much discriminative information as possible. At the time it appeared, this method combined with sparse coding achieved competitive classification results with only a small dictionary. However, its performance lags far behind the state-of-the-art results as only the zero-order information is exploited. Inspired by the success of high-order statistical information in existing advanced feature coding or pooling methods, we make an attempt to address the limitation of LP. To this end, we present a novel method called high-order LP (HO-LP) to leverage the information higher than the zero-order one. Our idea is intuitively simple: we compute the first- and second-order statistics per configuration bin and model them as a Gaussian. Accordingly, we employ a collection of Gaussians as visual words to represent the universal probability distribution of features from all classes. Our problem is naturally formulated as encoding Gaussians over a dictionary of Gaussians as visual words. This problem, however, is challenging since the space of Gaussians is not a Euclidean space but forms a Riemannian manifold. We address this challenge by mapping Gaussians into the Euclidean space, which enables us to perform coding with common Euclidean operations rather than complex and often expensive Riemannian operations. Our HO-LP preserves the advantages of the original LP: pooling only similar features and using a small dictionary. Meanwhile, it achieves very promising performance on standard benchmarks, with either conventional, hand-engineered features or deep learning-based features.
Peihua Li, Hui Zeng 0001, Qilong Wang 0001, Simon C. K. Shiu, Lei Zhang 0006
IEEE Trans. Image Process.3
2016 RAID-G: Robust Estimation of Approximate Infinite Dimensional Gaussian with Application to Material Recognition
abstract
Infinite dimensional covariance descriptors can provide richer and more discriminative information than their low dimensional counterparts. In this paper, we propose a novel image descriptor, namely, robust approximate infinite dimensional Gaussian (RAID-G). The challenges of RAID-G mainly lie on two aspects: (1) description of infinite dimensional Gaussian is difficult due to its non-linear Riemannian geometric structure and the infinite dimensional setting, hence effective approximation is necessary, (2) traditional maximum likelihood estimation (MLE) is not robust to high (even infinite) dimensional covariance matrix in Gaussian setting. To address these challenges, explicit feature mapping (EFM) is first introduced for effective approximation of infinite dimensional Gaussian induced by additive kernel function, and then a new regularized MLE method based on von Neumann divergence is proposed for robust estimation of covariance matrix. The EFM and proposed regularized MLE allow a closed-form of RAID-G, which is very efficient and effective for high dimensional features. We extend RAID-G by using the outputs of deep convolutional neural networks as original features, and apply it to material recognition. Our approach is evaluated on five material benchmarks and one fine-grained benchmark. It achieves 84.9% accuracy on FMD and 86.3% accuracy on UIUC material database, which are much higher than state-of-the-arts.
Qilong Wang 0001, Peihua Li, Wangmeng Zuo, Lei Zhang 0006
CVPR1
2016 Evaluation of ground distances and features in EMD-based GMM matching for texture classification
Hua Hao, Qilong Wang 0001, Peihua Li, Lei Zhang 0006
Pattern Recognit.2
2016 Towards effective codebookless model for image classification
Qilong Wang 0001, Peihua Li, Lei Zhang 0006, Wangmeng Zuo
Pattern Recognit.1
2015 From dictionary of visual words to subspaces: Locality-constrained affine subspace coding
abstract
The locality-constrained linear coding (LLC) is a very successful feature coding method in image classification. It makes known the importance of locality constraint which brings high efficiency and local smoothness of the codes. However, in the LLC method the geometry of feature space is described by an ensemble of representative points (visual words) while discarding the geometric structure immediately surrounding them. Such a dictionary only provides a crude, piecewise constant approximation of the data manifold. To approach this problem, we propose a novel feature coding method called locality-constrained affine subspace coding (LASC). The data manifold in LASC is characterized by an ensemble of subspaces attached to the representative points (or affine subspaces), which can provide a piecewise linear approximation of the manifold. Given an input descriptor, we find its top-k neighboring subspaces, in which the descriptor is linearly decomposed and weighted to form the first-order LASC vector. Inspired by the success of usage of higher-order information in image classification, we propose the second-order LASC vector based on the Fisher information metric for further performance improvement. We make experiments on challenging benchmarks and experiments have shown the LASC method is very competitive.
Peihua Li, Xiaoxiao Lu, Qilong Wang 0001
CVPR3
2015 Ask the dictionary: Soft-assignment location-orientation pooling for image classification
abstract
The pooling step is one of the key components of the well-known Bag-of-visual words (BoW) model widely used in image classification. In this paper, we propose a novel pooling method, which is called Soft-Assignment Location-Orientation Pooling (SALOP). Inspired by the bag of statistical sampling analysis (Bossa), SALOP also explores the effect of dictionary for pooling method, but leverages both location and orientation information between the local descriptors and the atoms of dictionary to aggregate feature codes. Moreover, different from existing pooling methods, SALOP employs a soft-assignment pooling scheme to handle ambiguity and uncertainty existing in the pooling process. The evaluation is conducted on two image benchmarks: Scene15 and PASCAL VOC 2007. The experimental results show our SALOP can achieve promising performances.
Qilong Wang 0001, Xiaona Deng, Peihua Li, Lei Zhang 0006
ICIP1
2014 Shrinkage Expansion Adaptive Metric Learning
Qilong Wang 0001, Wangmeng Zuo, Lei Zhang 0006, Peihua Li
ECCV (7)1
2013 A Novel Earth Mover's Distance Methodology for Image Matching with Gaussian Mixture Models
abstract
The similarity or distance measure between Gaussian mixture models (GMMs) plays a crucial role in content-based image matching. Though the Earth Mover's Distance (EMD) has shown its advantages in matching histogram features, its potentials in matching GMMs remain unclear and are not fully explored. To address this problem, we propose a novel EMD methodology for GMM matching. We first present a sparse representation based EMD called SR-EMD by exploiting the sparse property of the underlying problem. SR-EMD is more efficient and robust than the conventional EMD. Second, we present two novel ground distances between component Gaussians based on the information geometry. The perspective from the Riemannian geometry distinguishes the proposed ground distances from the classical entropy-or divergence-based ones. Furthermore, motivated by the success of distance metric learning of vector data, we make the first attempt to learn the EMD distance metrics between GMMs by using a simple yet effective supervised pair-wise based method. It can adapt the distance metrics between GMMs to specific classification tasks. The proposed method is evaluated on both simulated data and benchmark real databases and achieves very promising performance.
Peihua Li, Qilong Wang 0001, Lei Zhang 0006
ICCV2
2013 Log-Euclidean Kernels for Sparse Representation and Dictionary Learning
abstract
The symmetric positive definite (SPD) matrices have been widely used in image and vision problems. Recently there are growing interests in studying sparse representation (SR) of SPD matrices, motivated by the great success of SR for vector data. Though the space of SPD matrices is well-known to form a Lie group that is a Riemannian manifold, existing work fails to take full advantage of its geometric structure. This paper attempts to tackle this problem by proposing a kernel based method for SR and dictionary learning (DL) of SPD matrices. We disclose that the space of SPD matrices, with the operations of logarithmic multiplication and scalar logarithmic multiplication defined in the Log-Euclidean framework, is a complete inner product space. We can thus develop a broad family of kernels that satisfies Mercer's condition. These kernels characterize the geodesic distance and can be computed efficiently. We also consider the geometric structure in the DL process by updating atom matrices in the Riemannian space instead of in the Euclidean space. The proposed method is evaluated with various vision problems and shows notable performance gains over state-of-the-arts.
Peihua Li, Qilong Wang 0001, Wangmeng Zuo, Lei Zhang 0006
ICCV2
2013 Variational Earth Mover's Distance for Image Segmentation
abstract
Image segmentation using similarity or dissimilarity measures between probability distributions has been of great research interest in recent years. It is shown that the cross-bin metrics such as EMD is superior to the bin-wise metrics. However, existing segmentation approaches involving EMD are limited to univariate distributions, or one-dimensional marginal distributions of multidimensional features. This paper presents a novel segmentation method based on the variational EMD (VEMD) model, which can exploit joint distributions of multidimensional features. This method formulates the segmentation problem as the minimization of the EMD-based functional, which measures the distance between the foreground (resp. background) distribution and the reference foreground (resp. background) distribution. Using the simplex method and theory of shape derivative, we minimize the functional and obtain the gradient descent flow. We use a Gaussian filtering level-set method to obtain the numerical solution, in which the level-set re-initialization and smoothness constraint commonly imposed by the contour length are not necessary. Experiments show that the proposed method outperforms the state-of-the-art segmentation methods in the presence of illumination changes and noise.
Peihua Li, Qilong Wang 0001
ICIG2
2012 Robust Registration-Based Tracking by Sparse Representation with Model Update
Peihua Li, Qilong Wang 0001
ACCV (3)2
2012 Local Log-Euclidean Covariance Matrix (L2ECM) for Image Representation and Its Applications
Peihua Li, Qilong Wang 0001
ECCV (3)2