Pengfei Zhu 0001

dblp:40/6172-1 · DBLP profile ↗
← Back
160ranked-venue papers
36as first author
87since 2021 · last 2026
0000-0002-4310-9140ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 113 · 25 first-author · 60 since 2021Graphics, computer vision, multimedia, augmented reality and games · 79 · 17 first-author · 46 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 4 since 2021Security and privacy · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Reconcile Gradient Modulation for Harmony Multimodal Learning
abstract
Multimodal learning frequently faces two coupled challenges: modality imbalance, where dominant modalities suppress others during training, and modality conflict, where opposing gradient directions hinder optimization. Existing methods typically address these issues in isolation, yet they are intrinsically correlated and most fundamentally reflected in the gradient space—severe imbalance may obscure conflicts, while suppressing conflict may homogenize features and worsen imbalance, affecting fusion performance. To jointly address this coupled challenge, we propose Reconcile Gradient Modulation (RGM), a unified framework that adaptively adjusts gradient magnitude and direction for harmony multimodal learning. The core of RGM is SynOrth Grad, which minimizes Dirichlet energy to perform minimal-gradient surgery. It enhances cooperation synergy when modalities are aligned and enforces orthogonality to preserve uniqueness in conflict situations, thus promoting stable and balanced learning. To guide this modulation, we propose Cumulative Gradient Energy (CGE) as a convergence-guaranteed measure of modality-wise progress, and construct a Balance-nonConflict Plane (BCP) for real-time diagnosis and control of training dynamics. Experiments on diverse benchmarks validate our effectiveness and generalizability, consistently outperforming counterparts that are designed to handle multimodal imbalance or conflict independently.
Xiyuan Gao, Bing Cao 0002, Baoquan Gong, Pengfei Zhu 0001
AAAI4
2026 Point Cloud Quantization Through Multimodal Prompting for 3D Understanding
abstract
Vector quantization has emerged as a powerful tool in large-scale multimodal models, unifying heterogeneous representations through discrete token encoding. However, its effectiveness hinges on robust codebook design. Current prototype-based approaches relying on trainable vectors or clustered centroids fall short in representativeness and interpretability, even as multimodal alignment demonstrates its promise in vision-language models. To address these limitations, we propose a simple multimodal prompting-driven quantization framework for point cloud analysis. Our methodology is built upon two core insights: 1) Text embeddings from pre-trained models inherently encode visual semantics through many-to-one contrastive alignment, naturally serving as robust prototype priors; and 2) Multimodal prompts enable adaptive refinement of these prototypes, effectively mitigating vision-language semantic gaps. The framework introduces a dual-constrained quantization space, enforced by compactness and separation regularization, which seamlessly integrates visual and prototype features, resulting in hybrid representations that jointly encode geometric and semantic information. Furthermore, we employ Gumbel-Softmax relaxation to achieve differentiable discretization while maintaining quantization sparsity. Extensive experiments on the ModelNet40 and ScanObjectNN datasets clearly demonstrate the superior effectiveness of the proposed method.
Wencheng Zhu, Xinzhong Zhu, Pengfei Zhu 0001
AAAI5
2026 CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion generates all-weather perception-capable images by combining complementary modalities, enhancing environmental awareness for intelligent unmanned systems. Existing methods either focus on pixel-level fusion while overlooking downstream task adaptability or implicitly learn rigid semantics through cascaded detection/segmentation models, unable to interactively address diverse semantic target perception needs. We propose CtrlFuse, a controllable image fusion framework that enables interactive dynamic fusion guided by mask prompts. The model integrates a multi-modal feature extractor, a reference prompt encoder (RPE), and a prompt-semantic fusion module (PSFM). The RPE dynamically encodes task-specific semantic prompts by fine-tuning pre-trained segmentation models with input mask guidance, while the PSFM explicitly injects these semantics into fusion features. Through synergistic optimization of parallel segmentation and fusion branches, our method achieves mutual enhancement between task performance and fusion quality. Experiments demonstrate state-of-the-art results in both fusion controllability and segmentation accuracy, with the adapted task branch even outperforming the original segmentation model.
Yiming Sun 0003, Yuan Ruan, Qinghua Hu, Pengfei Zhu 0001
AAAI4
2026 Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion
abstract
Image fusion aims to integrate comprehensive information from images acquired through multiple sources. However, images captured by diverse sensors often encounter various degradations that can negatively affect fusion quality. Traditional fusion methods generally treat image enhancement and fusion as separate processes, overlooking the inherent correlation between them; notably, the dominant regions in one modality of a fused image often indicate areas where the other modality might benefit from enhancement. Inspired by this observation, we introduce the concept of dominant regions for image enhancement and present a Dynamic Relative EnhAnceMent framework for Image Fusion (Dream-IF). This framework quantifies the relative dominance of each modality across different layers and leverages this information to facilitate reciprocal cross-modal enhancement. By integrating the relative dominance derived from image fusion, our approach supports not only image restoration but also a broader range of image enhancement applications. Furthermore, we employ prompt-based encoding to capture degradation-specific details, which dynamically steer the restoration process and promote coordinated enhancement in both multi-modal image fusion and image enhancement scenarios. Extensive experimental results demonstrate that Dream-IF consistently outperforms its counterparts.
Xingxin Xu, Bing Cao 0002, Dongdong Li 0004, Qinghua Hu, Pengfei Zhu 0001
AAAI5
2026 VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
abstract
Vision-language models bridge visual and linguistic understanding and have proven to be powerful for video recognition tasks. Existing methods primarily rely on parameter-efficient fine-tuning of pre-trained image-text models, suffering from limited interpretability and poor generalization due to inadequate temporal modeling. To address these, we propose a simple yet effective video-to-text discretization framework. Our approach leverages the frozen text encoder to build a visual codebook derived from video class labels, exploiting the many-to-one contrastive alignment between visual and textual embeddings in multimodal pretraining. This enables the transformation of temporal visual features into discrete textual tokens via feature lookups, yielding interpretable video representations through explicit video modeling. Then, to improve robustness against noisy or irrelevant frames, we introduce a confidence-aware fusion module that dynamically weights keyframes based on their semantic relevance, as measured by the codebook. Furthermore, we incorporate learnable text prompts to conduct adaptive codebook updates during training. Experiments on four datasets, including HMDB-51, UCF-101, Something-Something-v2, and Kinetics-400, validate the superiority of our approach, achieving competitive improvements over state-of-the-art approaches.
Wencheng Zhu, Pengfei Zhu 0001
AAAI4
2026 KSCNet: Exploring KAN and state space model collaboration network for small object detection from UAV imagery
Yiming Sun 0003, Pengfei Zhu 0001, Xinzhong Zhu
Expert Syst. Appl.4
2026 DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency
abstract
Given a single labeled examples, in-context segmentation aims to segment corresponding objects. This setting, known as one-shot segmentation in few-shot learning, explores the segmentation model's generalization ability and has been applied to various vision tasks, including scene understanding and image/video editing. While recent Segment Anything Models (SAMs) have achieved state-of-the-art results in interactive segmentation, these approaches are not directly applicable to in-context segmentation. In this work, we propose the Dual Consistency SAM (DC-SAM) method based on prompt-tuning to adapt SAM and SAM2 for in-context segmentation of both images and videos. Our key insights are to enhance the features of the SAM's prompt encoder in segmentation by providing high-quality visual prompts. When generating a mask prior from support images, we fuse the SAM features to better align the prompt encoder rather than relying solely on a pre-trained backbone. Then, we design a cycle-consistent cross-attention on fused features and initial visual prompts. This design leverages coarse masks from the SAM mask decoder to ensure consistency between features and visual prompts. Next, a dual-branch design is provided by using the discriminative positive and negative prompts in the prompt encoder. Furthermore, we design a simple mask-tube training strategy to adopt our proposed dual consistency method into the mask tube. Although the proposed DC-SAM is primarily designed for images, it can be seamlessly extended to the video domain with the support of SAM2. Given the absence of in-context segmentation in the video domain, we manually curate and construct the first benchmark from existing video segmentation datasets, namedIn-Context Video Object Segmentation (IC-VOS), to better assess the in-context capability of the model. Extensive experiments demonstrate that our method achieves 55.5 (+1.4) mIoU on COCO-20$^{i}$, 73.0 (+1.1) mIoU on PASCAL-5$^{i}$, and a$\mathcal {J\&F}$score of 71.52 on the proposed IC-VOS benchmark.
Mengshi Qi, Pengfei Zhu 0001, Xiangtai Li, Xiaoyang Bi, Lu Qi 0001, Huadong Ma, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Bi-directional Self-Registration for misaligned infrared-visible image fusion
Timing Li, Bing Cao 0002, Qinghua Hu, Pengfei Zhu 0001
Pattern Recognit.4
2026 A large-scale drone based thermal infrared benchmark and inception transformer network for crowd counting
abstract
Crowd counting plays a crucial role in applications like public safety and smart cities, where estimating the number of people in images or videos is essential. However, since most of the current crowd counting datasets are visible light images, the counting performance is limited by the light intensity. To tackle this problem, we collect a drone based thermal infrared crowd counting dataset named LYU-DroneInfrared, which includes 64,210 images and 2,997,352 head annotation points, covering different scenes such as schools, streets, squares, sports grounds, etc. In addition, we propose the IncepTNet, an inception transformer network based on the transformer architecture. It consists of two parts: low frequency feature extraction and high frequency feature extraction. In low-frequency feature extraction, average pooling is first applied, followed by multi-head self-attention to capture contextual information in the image. On the other hand, to be able to pay more attention to the fine-grained details of the images, the parallel approach of convolution and maximum pooling are used to extract the high frequency features. The proposed method is validated on two benchmark datasets JHU-Crowd++ and NWPU-Crowd, and a newly collected LYU-DroneInfrared dataset. Extensive experimental results have shown that the IncepTNet method exhibits excellent crowd counting performance on different types of datasets.
Xing Wang 0002, Timing Li, Shuanglong Yao, Pengfei Zhu 0001
Pattern Recognit.7
2026 Incomplete cross-modality class-incremental learning in visible-thermal recognition
abstract
Visible-thermal cross-modality learning enhances downstream task performance by integrating information from multiple sources. In real-world scenarios such as autonomous driving, new classes continually emerge, and data is often incomplete due to sensor occlusions. This raises a key question about how to incrementally update a model with incomplete cross-modality data. To address this problem, we propose a practical task termed incomplete cross-modality class-incremental learning (ICMCIL), which aims to effectively leverage incomplete cross-modality information to learn new knowledge without forgetting the old. We construct a benchmark for ICMCIL and thoroughly analyze its challenges, revealing that (1) different modalities experience varying degrees of forgetting, (2) conventional cross-modality fusion only partially alleviates forgetting, and (3) missing data exacerbates the forgetting of previous classes. To address these issues, we propose Hybrid Fusion via Completion (HFC), a unified framework that integrates completion, fusion, and forgetting prevention. Additionally, we enhance information fusion by introducing a feature interchange mechanism, wherein features are shuffled and channels are reordered to improve information flow. Extensive experiments demonstrate that HFC effectively addresses ICMCIL, significantly mitigating modality forgetting.
Xinjie Yao, Yanxian Bi, Yu Wang 0106, Pengfei Zhu 0001, Ruipu Zhao, Wanyu Lin, Qinghua Hu
Pattern Recognit.4
2026 GC3VG: Generalized Multi-Task Visual Grounding With Coarse-to-Fine Consistency Constraints
abstract
In this work, we propose an efficient and streamlined paradigm to address the challenge of consistency prediction in generalized multi-task visual grounding. While most existing approaches primarily focus on integrating multi-modal information and employing multi-task learning to enhance both visual and linguistic understanding, they often rely on joint supervision at the region and pixel levels to exploit task complementarities. In contrast, C3VG explores the relatively under-addressed problem ofconsistency across multi-task predictions. To this end, a multi-task visual grounding framework based on a coarse-to-fine architecture is introduced. Empirical studies demonstrate that the incorporation of both implicit and explicit consistency constraints substantially enhances the coherence between detection and segmentation outputs. However, C3VG is restricted to single-referent visual grounding scenarios and exhibits limited generalizability to real-world applications, which often involve multi-referents or even absent referent. To overcome these limitations, we proposeGC3VG, which incorporates three key advancements: (1) extension to generalized scenarios, including both multi-referent and non-referent cases; (2) aUnified Coherent Refinement Modulethat implicitly encodes region- and instance-level features while explicitly modeling their relational alignment through an IoUbased constraint; and (3) aGranularity-aware Hard-mining Alignmentstrategy that enforces prediction consistency in the feature space and simultaneously enhances the discriminative power of visual and linguistic representations. Extensive experiments on RefCOCO/+/g and gRefCOCO demonstrate the effectiveness and generalizability of the proposed framework.
Kai Chen 0037, Wenxuan Cheng, Jiedong Zhuang, Zhenhua Feng 0001, Pengfei Zhu 0001, Wankou Yang
IEEE Trans. Circuits Syst. Video Technol.6
2026 Hyperbolic Cycle Alignment for Infrared-Visible Image Fusion
abstract
Accurate alignment is a fundamental prerequisite for multi-modal image fusion, yet aligning multi-modal remains challenging due to nonlinear geometric distortions and substantial appearance discrepancies. This work presents the Hyperbolic Cycle Alignment Network (Hy-CycleAlign), a geometry-aware cyclic alignment framework formulated in hyperbolic space. By embedding multi-level representations into a negatively curved manifold, Hy-CycleAlign departs from conventional Euclidean-space paradigms and provides enhanced sensitivity to spatial perturbations, enabling more reliable modeling of cross-modal correspondences. The framework integrates a dual-path cyclic structure that enforces bidirectional deformation consistency and prevents accumulated alignment drift. In addition, a hyperbolic hierarchy contrastive alignment module jointly constrains semantic and structural representations within a unified hyperbolic embedding domain, promoting coherent alignment across global and local geometric scales. From a theoretical standpoint, we derive the sensitivity properties of the Poincaré model and show that its metric inherently amplifies positional variations, thereby strengthening the discriminability of subtle cross-modal misalignments compared with Euclidean geometry. Extensive experiments on diverse misaligned multi-modal datasets show that Hy-CycleAlign achieves strong overall results in alignment accuracy, structural fidelity, and downstream fusion quality. These results validate the effectiveness of hyperbolic geometric modeling for robust multi-modal image alignment.
Timing Li, Bing Cao 0002, Jiahe Feng, Haifang Cao, Qinghua Hu, Pengfei Zhu 0001
IEEE Trans. Image Process.6
2026 Multi-Granularity Superpoint Graph Learning for Weakly Supervised 3D Semantic Segmentation
abstract
Weakly supervised 3D semantic segmentation has proven effective in alleviating the heavy dependence on dense annotations by generating high-quality pseudo-labels. However, due to the scene complexity and disorder of the point cloud, merely applying the model semantic prediction or hand-crafted feature similarity for pseudo labeling is inefficient and biased. This limitation inevitably results in incorrect pseudo labels. To tackle this challenge, we propose a new method called Multi-granularity Superpoint Graph Learning (MSGL) that leverages the multi-scale local features of point clouds to improve the quality of pseudo labels. We first design a multi-granularity local representation learning module on the superpoint graph to capture the neighboring structure information of each superpoint within complex scenes. Subsequently, the generated structural embedding is utilized to enhance the affinity matrix of label propagation, thereby yielding high-quality pseudo labels. To further enforce the generalization of the structural representation module under scenario changes or data fluctuations, we present a multi-granularity consistency loss in MSGL. This loss is applied across different views of the superpoint graph within each scene to ensure a robust and consistent learning process. Our experiments conducted on three benchmarks show that the proposed method outperforms existing weakly supervised methods under several sparse label settings, and improves the baseline by an average of 7.7% with only 1% extra computation cost. Moreover, our approach even compares favorably to some fully supervised methods with only one point labeled for each thing.
Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Le Hui, Jin Xie 0001, Bin Xiao 0002, Qinghua Hu
IEEE Trans. Multim.3
2026 TEDFuse: Task-Driven Equivariant Consistency Decomposition Network for Multi-Modal Image Fusion
abstract
Multimodal image fusion integrates infrared and visible images by leveraging their complementary strengths. However, most existing fusion techniques primarily focus on pixel level integration, often neglecting the preservation of semantic consistency between the source and fused images. To address this limitation, we propose TEDFuse, a Task-Driven Equivariant Consistency Decomposition Network that ensures semantic con sistency within the image space and across high-level semantic tasks. TEDFuse incorporates two key components: first, a robust decomposition framework with equivariant consistency, ensuring that the fused image retains consistent transformation properties under shifts, rotations, and reflections, thereby enhancing local detail preservation and global semantic alignment; In addition, a task-driven fusion framework that integrates a segmentation module, reinforcing semantic feature preservation through a semantic loss function and ensuring consistency in downstream tasks such as segmentation and detection. The proposed method not only preserves the semantic coherence of the fused image but also improves performance in high-level tasks, demonstrating superior capability in multimodal fusion for complex visual applications. Extensive experiments validate the effectiveness of TEDFuse by analyzing feature evolution, examining the relationship between fusion quality and task performance, and discussing calibration strategies for infrared-visible image fusion. The code is available at https://github.com/Claire-cxy/TEDFuse.
Yiming Sun 0003, Zhen Wang 0033, Hao Cheng 0010, Yongfeng Dong, Pengfei Zhu 0001
IEEE Trans. Multim.6
2025 Asymmetric Reinforcing Against Multi-Modal Representation Bias
abstract
The strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the challenge of dynamic modality contributions, the dominance of different modalities may change with the environments, leading to suboptimal performance in multimodal learning. Current methods mainly enhance weak modalities to balance multimodal representation bias, which inevitably optimizes from a partialmodality perspective, easily leading to performance descending for dominant modalities. To address this problem, we propose an Asymmetric Reinforcing method against Multimodal representation bias (ARM). Our ARM dynamically reinforces the weak modalities while maintaining the ability to represent dominant modalities through conditional mutual information. Moreover, we provide an in-depth analysis that optimizing certain modalities could cause information loss and prevent leveraging the full advantages of multimodal data. By exploring the dominance and narrowing the contribution gaps between modalities, we have significantly improved the performance of multimodal learning, making notable progress in mitigating imbalanced multimodal learning.
Xiyuan Gao, Bing Cao 0002, Pengfei Zhu 0001, Nannan Wang 0001, Qinghua Hu
AAAI3
2025 Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning
abstract
Recently, remarkable progress has been made in large-scale pre-trained model tuning, and inference efficiency is becoming more crucial for practical deployment. Early exiting in conjunction with multi-stage predictors, when cooperated with a parameter-efficient fine-tuning strategy, offers a straightforward way to achieve an inference-efficient model. However, a key challenge remains unresolved: How can early stages provide low-level fundamental features to deep stages while simultaneously supplying high-level discriminative features to early-stage predictors? To address this problem, we propose a Decoupled Multi-Predictor Optimization (DMPO) method to effectively decouple the low-level representative ability and high-level discriminative ability in early stages. First, in terms of architecture, we introduce a lightweight bypass module into multi-stage predictors for functional decomposition of shallow features from early stages, while a high-order statistics-based predictor is developed for early stages to effectively enhance their discriminative ability. To reasonably train our multi-predictor architecture, a decoupled optimization is proposed to allocate two-phase loss weights for multi-stage predictors during model tuning, where the initial training phase enables the model to prioritize the acquisition of discriminative ability of deep stages via emphasizing representative ability of early stages, and the latter training phase drives discriminative ability towards earlier stages as much as possible. As such, our DMPO can effectively decouple representative and discriminative abilities in early stages in terms of architecture design and model optimization. Experiments across various datasets and pre-trained backbones demonstrate that DMPO clearly outperforms its counterparts when reducing computational cost.
Liwei Luo, Shuaitengyuan Li, Dongwei Ren, Qilong Wang 0001, Pengfei Zhu 0001, Qinghua Hu
ICCV5
2025 Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark
abstract
The dynamic imbalance of the fore-background is a major challenge in video object counting, which is usually caused by the sparsity of target objects. This remains understudied in existing works and often leads to severe under-/over-prediction errors. To tackle this issue in video object counting, we propose a density-embedded Efficient Masked Autoencoder Counting (E-MAC) framework in this paper. To empower the model’s representation ability on density regression, we develop a new Density-Embedded Masked mOdeling (DEMO) method, which first takes the density map as an auxiliary modality to perform multimodal self-representation learning for image and density map. Although DEMO contributes to effective cross-modal regression guidance, it also brings in redundant background information, making it difficult to focus on the foreground regions. To handle this dilemma, we propose an efficient spatial adaptive masking derived from density maps to boost efficiency. Meanwhile, we employ an optical flow-based temporal collaborative fusion strategy to effectively capture the dynamic variations across frames, aligning features to derive multi-frame density residuals. The counting accuracy of the current frame is boosted by harnessing the information from adjacent frames. In addition, considering that most existing datasets are limited to human-centric scenarios, we propose a large video bird counting dataset, DroneBird, in natural scenarios for migratory bird protection. Extensive experiments on three crowd datasets and our DroneBird validate our superiority against the counterparts. The code and dataset are available.
Bing Cao 0002, Quanhao Lu, Jiekang Feng, Qilong Wang 0001, Pengfei Zhu 0001, Qinghua Hu
ICLR5
2025 Asymmetric Factorized Bilinear Operation for Vision Transformer
abstract
As a core component of Transformer-like deep architectures, a feed-forward network (FFN) for channel mixing is responsible for learning features of each token. Recent works show channel mixing can be enhanced by increasing computational burden or can be slimmed at the sacrifice of performance. Although some efforts have been made, existing works are still struggling to solve the paradox of performance and complexity trade-offs. In this paper, we propose an Asymmetric Factorized Bilinear Operation (AFBO) to replace FFN of vision transformer (ViT), which attempts to efficiently explore rich statistics of token features for achieving better performance and complexity trade-off. Specifically, our AFBO computes second-order statistics via a spatial-channel factorized bilinear operation for feature learning, which replaces a simple linear projection in FFN and enhances the feature learning ability of ViT by modeling second-order correlation among token features. Furthermore, our AFBO presents two structured-sparsity channel mapping strategies, namely Grouped Cross Channel Mapping (GCCM) and Overlapped Cycle Channel Mapping (OCCM). They decompose bilinear operation into grouped channel features by considering information interaction between groups, significantly reducing computational complexity while guaranteeing model performance. Finally, our AFBO is built with GCCM and OCCM in an asymmetric way, aiming to achieve a better trade-off. Note that our AFBO is model-agnostic, which can be flexibly integrated with existing ViTs. Experiments are conducted with twenty ViTs on various tasks, and the results show our AFBO is superior to its counterparts while improving existing ViTs in terms of generalization and robustness.
Qilong Wang 0001, Jiangtao Xie, Pengfei Zhu 0001, Qinghua Hu
ICLR4
2025 Task-Gated Multi-Expert Collaboration Network for Degraded Multi-Modal Image Fusion
abstract
Multi-modal image fusion aims to integrate complementary information from different modalities to enhance perceptual capabilities in applications such as rescue and security. However, real-world imaging often suffers from degradation issues, such as noise, blur, and haze in visible imaging, as well as stripe noise in infrared imaging, which significantly degrades model performance. To address these challenges, we propose a task-gated multi-expert collaboration network (TG-ECNet) for degraded multi-modal image fusion. The core of our model lies in the task-aware gating and multi-expert collaborative framework, where the task-aware gating operates in two stages: degradation-aware gating dynamically allocates expert groups for restoration based on degradation types, and fusion-aware gating guides feature integration across modalities to balance information retention between fusion and restoration tasks. To achieve this, we design a two-stage training strategy that unifies the learning of restoration and fusion tasks. This strategy resolves the inherent conflict in information processing between the two tasks, enabling all-in-one multi-modal image restoration and fusion. Experimental results demonstrate that TG-ECNet significantly enhances fusion performance under diverse complex degradation conditions and improves robustness in downstream applications. The code is available at https://github.com/LeeX54946/TG-ECNet.
Yiming Sun 0003, Pengfei Zhu 0001, Qinghua Hu, Dongwei Ren, Xinzhong Zhu
ICML3
2025 Socialized Coevolution: Advancing a Better World through Cross-Task Collaboration
abstract
Traditional machine societies rely on data-driven learning, overlooking interactions and limiting knowledge acquisition from model interplay. To address these issues, we revisit the development of machine societies by drawing inspiration from the evolutionary processes of human societies. Motivated by Social Learning (SL), this paper introduces a practical paradigm of Socialized Coevolution (SC). Compared to most existing methods focused on knowledge distillation and multi-task learning, our work addresses a more challenging problem: not only enhancing the capacity to solve new downstream tasks but also improving the performance of existing tasks through inter-model interactions. Inspired by cognitive science, we propose Dynamic Information Socialized Collaboration (DISC), which achieves SC through interactions between models specialized in different downstream tasks. Specifically, we introduce the dynamic hierarchical collaboration and dynamic selective collaboration modules to enable dynamic and effective interactions among models, allowing them to acquire knowledge from these interactions. Finally, we explore potential future applications of combining SL and SC, discuss open questions, and propose directions for future research, aiming to spark interest in this emerging and exciting interdisciplinary field. Our code will be publicly available at https://github.com/yxjdarren/SC.
Xinjie Yao, Yu Wang 0106, Pengfei Zhu 0001, Wanyu Lin, Ruipu Zhao, Zhoupeng Guo, Qinghua Hu
ICML3
2025 Motion-Aware Adaptive Pixel Pruning for Efficient Local Motion Deblurring
Wei Shang 0001, Dongwei Ren, Pengfei Zhu 0001, Qinghua Hu, Wangmeng Zuo
ACM Multimedia4
2025 Multimodal Negative Learning
abstract
Multimodal learning systems often encounter challenges related to modality imbalance, where a dominant modality may overshadow others, thereby hindering the learning of weak modalities. Conventional approaches often force weak modalities to align with dominant ones in "Learning to be (the same)" (Positive Learning), which risks suppressing the unique information inherent in the weak modalities. To address this challenge, we offer a new learning paradigm: "Learning Not to be" (Negative Learning). Instead of enhancing weak modalities’ target-class predictions, the dominant modalities dynamically guide the weak modality to suppress non-target classes. This stabilizes the decision space and preserves modality-specific information, allowing weak modalities to preserve unique information without being over-aligned. We proceed to reveal the multimodal learning from a robustness perspective and theoretically derive the Multimodal Negative Learning (MNL) framework, which introduces a dynamic guidance mechanism tailored for negative learning. Our method provably tightens the robustness lower bound of multimodal learning by increasing the Unimodal Confidence Margin (UCoM) and reduces the empirical error of weak modalities, particularly under noisy and imbalanced scenarios. Extensive experiments across multiple benchmarks demonstrate the effectiveness and generalizability of our approach against the competing methods. The code will be available at: https://github.com/BaoquanGong/Multimodal-Negative-Learning.git
Baoquan Gong, Xiyuan Gao, Pengfei Zhu 0001, Qinghua Hu, Bing Cao 0002
NeurIPS3
2025 Graphs Help Graphs: Multi-Agent Graph Socialized Learning
abstract
Graphs in the real world are fragmented and dynamic, lacking collaboration akin to that observed in human societies. Existing paradigms present collaborative information collapse and forgetting, making collaborative relationships poorly autonomous and interactive information insufficient. Moreover, collaborative information is prone to loss when the graph grows. Effective collaboration in heterogeneous dynamic graph environments becomes challenging. Inspired by social learning, this paper presents a Graph Socialized Learning (GSL) paradigm. We provide insights into graph socialization in GSL and boost the performance of agents through effective collaboration. It is crucial to determine with whom, what, and when to share and accumulate information for effective GSL. Thus, we propose the ''Graphs Help Graphs'' (GHG) method to solve these issues. Specifically, it uses a graph-driven organizational structure to select interacting agents and manage interaction strength autonomously. We produce customized synthetic graphs as an interactive medium based on the demand of agents, then apply the synthetic graphs to build prototypes in the life cycle to help select optimal parameters. We demonstrate the effectiveness of GHG in heterogeneous dynamic graphs by an extensive empirical study. The code is available through https://github.com/Jillian555/GHG.
Yu Wang 0106, Pengfei Zhu 0001, Wanyu Lin, Xinjie Yao, Qinghua Hu
NeurIPS3
2025 Hyperbolic-Euclidean Deep Mutual Learning
abstract
Graph neural networks (GNNs) exhibit powerful performance in handling graph data, with Euclidean and hyperbolic variants excelling in processing grid-based and hierarchical structures, respectively. However, existing methods focus on learning specific structures linked to the inherent properties of the underlying space, failing to fully exploit their complementary properties in distinct geometric spaces, thus limiting their ability to efficiently model complex graph structures. In this paper, we propose a Hyperbolic-Euclidean Deep Mutual Learning (H-EDML) framework, which leverages the unique properties of hyperbolic space to effectively capture the hierarchical relationships present in graph data, while also utilizes the familiar Euclidean space to handle local interactions. Specifically, We design a topology mutual learning module to bolster the capacity of each single model to perceive the holistic topological structure of the graph. Then, we integrate a decision mutual learning module to further advance the models' comprehensive judgment capabilities towards graph data, thereby strengthening the robustness and generalization. Furthermore, we employ an attention-based probabilistic integration strategy for the final prediction to alleviate potential disparities in decision-making among different models. Extensive experiments on node classification are conducted on five real-world graph datasets and the results show that our proposed H-EDML achieves competitive performances compared to the state-of-the-art methods. The source code will be available at: https://github.com/caohaifang123/H-EDML.
Haifang Cao, Yu Wang 0106, Pengfei Zhu 0001, Qinghua Hu
WWW4
2025 Visible-thermal cross-modality class-incremental learning
Xinjie Yao, Yu Wang 0106, Pengfei Zhu 0001, Ruipu Zhao, Shenglei Pei, Wanyu Lin
Expert Syst. Appl.4
2025 Unknown Support Prototype Set for Open Set Recognition
Guosong Jiang, Pengfei Zhu 0001, Bing Cao 0002, Qinghua Hu
Int. J. Comput. Vis.2
2025 BackMix: Regularizing Open Set Recognition by Removing Underlying Fore-Background Priors
abstract
Open set recognition (OSR) requires models to classify known samples while detecting unknown samples for real-world applications. Existing studies show impressive progress using unknown samples from auxiliary datasets to regularize OSR models, but they have proved to be sensitive to selecting such known outliers. In this paper, we discuss the aforementioned problem from a new perspective: Can we regularize OSR models without elaborately selecting auxiliary known outliers? We first empirically and theoretically explore the role of foregrounds and backgrounds in open set recognition and disclose that: 1) backgrounds that correlate with foregrounds would mislead the model and cause failures when encounters 'partially' known images; 2) Backgrounds unrelated to foregrounds can serve as auxiliary known outliers and provide regularization via global average pooling. Based on the above insights, we propose a new method, Background Mix (BackMix), that mixes the foreground of an image with different backgrounds to remove the underlying fore-background priors. Specifically, BackMix first estimates the foreground with class activation maps (CAMs), then randomly replaces image patches with backgrounds from other images to obtain mixed images for training. With backgrounds de-correlated from foregrounds, the open set recognition performance is significantly improved. The proposed method is quite simple to implement, requires no extra operation for inferences, and can be seamlessly integrated into almost all of the existing frameworks.
Yu Wang 0106, Junxian Mu, Hongzhi Huang, Qilong Wang 0001, Pengfei Zhu 0001, Qinghua Hu
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Uncertainty-Aware Superpoint Graph Transformer for Weakly Supervised 3-D Semantic Segmentation
abstract
Weakly supervised 3-D semantic segmentation has successfully mitigated the labor-intensive and time-consuming task of annotating 3-D point clouds. However, reliably utilizing the minimal point-wise annotations for unlabeled data in complex and large-scale scenes is still challenging, such as only 20 points labeled in 2 million points. To tackle this challenge, we propose a new Uncertainty-aware Superpoint Graph Transformer (UaSGT) framework that utilizes minimal annotations for unlabeled data learning through reliable long-range supervision propagation from labeled superpoints to unlabeled superpoints. First, we propose a superpoint graph transformer to achieve long-range supervision propagation along the attention-based fuzzy subsets defined on superpoints. The attention-based fuzzy subset measures the membership of unlabeled superpoints to clusters centered on labeled superpoints. Second, we employ an uncertainty-aware membership rectification technique on the fuzzy subset to ensure reliable propagation among superpoints within the same category. This technique integrates an uncertainty prediction module to mask the influence of unreliable membership and a spatial prior refinement module to reduce uncertainty in intraclass membership degrees. Finally, experimental results on two large-scale benchmarks S3DIS and ScanNet-V2 demonstrate the superiority of our approach compared to the state-of-the-art with at least 90% annotation reduction, and our method also achieves comparable performance to fully supervised methods with less than 0.1% labeled points.
Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Le Hui, Jin Xie 0001, Qinghua Hu
IEEE Trans. Fuzzy Syst.3
2025 RTF: Recursive TransFusion for Multi-Modal Image Synthesis
abstract
Multi-modal image synthesis is crucial for obtaining complete modalities due to the imaging restrictions in reality. Current methods, primarily CNN-based models, find it challenging to extract global representations because of local inductive bias, leading to synthetic structure deformation or color distortion. Despite the significant global representation ability of transformer in capturing long-range dependencies, its huge parameter size requires considerable training data. Multi-modal synthesis solely based on one of the two structures makes it hard to extract comprehensive information from each modality with limited data. To tackle this dilemma, we propose a simple yet effective Recursive TransFusion (RTF) framework for multi-modal image synthesis. Specifically, we develop a TransFusion unit to integrate local knowledge extracted from the individual modality by connecting a CNN-based local representation block (LRB) and a transformer-based global fusion block (GFB) via a feature translating gate (FTG). Considering the numerous parameters introduced by the transformer, we further unfold a TransFusion unit with recursive constraint repeatedly, forming recursive TransFusion (RTF), which progressively extracts multi-modal information at different depths. Our RTF remarkably reduces network parameters while maintaining superior performance. Extensive experiments validate our superiority against the competing methods on multiple benchmarks. The source code will be available at https://github.com/guoliangq/RTF.
Bing Cao 0002, Guoliang Qi, Pengfei Zhu 0001, Qinghua Hu, Xinbo Gao 0001
IEEE Trans. Image Process.4
2025 CKD: Contrastive Knowledge Distillation From a Sample-Wise Perspective
abstract
In this paper, we propose a simple yet effective contrastive knowledge distillation framework that achieves sample-wise logit alignment while preserving semantic consistency. Conventional knowledge distillation approaches exhibit over-reliance on feature similarity per sample, which risks overfitting, and contrastive approaches focus on inter-class discrimination at the expense of intra-sample semantic relationships. Our approach transfers "dark knowledge" through teacher-student contrastive alignment at the sample level. Specifically, our method first enforces intra-sample alignment by directly minimizing teacher-student logit discrepancies within individual samples. Then, we utilize inter-sample contrasts to preserve semantic dissimilarities across samples. By redefining positive pairs as aligned teacher-student logits from identical samples and negative pairs as cross-sample logit combinations, we reformulate these dual constraints into an InfoNCE loss framework, reducing computational complexity lower than sample squares while eliminating dependencies on temperature parameters and large batch sizes. We conduct comprehensive experiments across three benchmark datasets, including the CIFAR-100, ImageNet-1K, and MS COCO datasets, and experimental results clearly confirm the effectiveness of the proposed method on image classification, object detection, and instance segmentation tasks.
Wencheng Zhu, Pengfei Zhu 0001, Yu Wang 0106, Qinghua Hu
IEEE Trans. Image Process.3
2025 CCP-GNN: Competitive Covariance Pooling for Improving Graph Neural Networks
abstract
Graph neural networks (GNNs) have advanced graph classification tasks, where a global pooling to generate graph representations by summarizing node features plays a critical role in the final performance. Most of the existing GNNs are built with a global average pooling (GAP) or its variants, which however, take no full consideration of node specificity while neglecting rich statistics inherent in node features, limiting classification performance of GNNs. Therefore, this article proposes a novel competitive covariance pooling (CCP) based on observation of graph structures, i.e., graphs generally can be identified by a (small) key part of nodes. To this end, our CCP generates node-level second-order representations to explore rich statistics inherent in node features, which are fed to a competitive-based attention module for effectively discovering key nodes through learning node weights. Subsequently, our CCP aggregates node-level second-order representations in conjunction with node weights by summation to produce a covariance representation for each graph, while an iterative matrix normalization is introduced to consider geometry of covariances. Note that our CCP can be flexibly integrated with various GNNs (namely CCP-GNN) to improve the performance of graph classification with little computational cost. The experimental results on seven graph-level benchmarks show that our CCP-GNN is superior or competitive to state-of-the-arts. Our code is available at https://github.com/Jillian555/CCP-GNN.
Pengfei Zhu 0001, Qinghua Hu, Xiao Wang 0017, Qilong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 Boosting Pseudo-Labeling With Curriculum Self-Reflection for Attributed Graph Clustering
abstract
Attributed graph clustering is an unsupervised learning task that aims to partition various nodes of a graph into distinct groups. Existing approaches focus on devising diverse pretext tasks to obtain suitable supervised information for representation learning, among which the predictive methods show great potential. However, these methods 1) generate auxiliary task bias toward the clustering target and 2) introduce label noise due to static thresholds. To address this issue, we propose a new self-supervised learning method, namely, pseudo-labeling with curriculum self-reflection (PLCSR), that learns reliable pseudo-labels by mining its information to achieve progressive processing of nodes in a self-reflection manner. First, a self-auxiliary encoder is constructed using the exponential moving average (EMA) of the original encoder's parameters to replace the auxiliary tasks, which provides an additional perspective of finding highly confident pseudo-labels. Second, a curriculum selection strategy using dynamic thresholds is designed to take full advantage of graph nodes more accurately. Besides simple nodes with high confidence at the initial stage, nodes that yield consistent predictions from both encoders are then assigned pseudo-labels to avoid the under-learning problem. For the rest difficult nodes that are highly uncertain, we abstain from making judgments to minimize their adverse impact on the model. Extensive experiments have shown that PLCSR significantly outperforms the state-of-the-art predictive method CDRS, achieving more than 6% improvements in terms of clustering accuracy. The code is available at: https://github.com/Jillian555/PLCSR.
Pengfei Zhu 0001, Yu Wang 0106, Bin Xiao 0002, Jinglin Zhang 0001, Wanyu Lin, Qinghua Hu
IEEE Trans. Neural Networks Learn. Syst.1
2024 Every Node Is Different: Dynamically Fusing Self-Supervised Tasks for Attributed Graph Clustering
abstract
Attributed graph clustering is an unsupervised task that partitions nodes into different groups. Self-supervised learning (SSL) shows great potential in handling this task, and some recent studies simultaneously learn multiple SSL tasks to further boost performance. Currently, different SSL tasks are assigned the same set of weights for all graph nodes. However, we observe that some graph nodes whose neighbors are in different groups require significantly different emphases on SSL tasks. In this paper, we propose to dynamically learn the weights of SSL tasks for different nodes and fuse the embeddings learned from different SSL tasks to boost performance. We design an innovative graph clustering approach, namely Dynamically Fusing Self-Supervised Learning (DyFSS). Specifically, DyFSS fuses features extracted from diverse SSL tasks using distinct weights derived from a gating network. To effectively learn the gating network, we design a dual-level self-supervised strategy that incorporates pseudo labels and the graph structure. Extensive experiments on five datasets show that DyFSS outperforms the state-of-the-art multi-task SSL methods by up to 8.66% on the accuracy metric. The code of DyFSS is available at: https://github.com/q086/DyFSS.
Pengfei Zhu 0001, Yu Wang 0106, Qinghua Hu
AAAI1
2024 Bi-directional Adapter for Multimodal Tracking
abstract
Due to the rapid development of computer vision, single-modal (RGB) object tracking has made significant progress in recent years. Considering the limitation of single imaging sensor, multi-modal images (RGB, infrared, etc.) are introduced to compensate for this deficiency for all-weather object tracking in complex environments. However, as acquiring sufficient multi-modal tracking data is hard while the dominant modality changes with the open environment, most existing techniques fail to extract multi-modal complementary information dynamically, yielding unsatisfactory tracking performance. To handle this problem, we propose a novel multi-modal visual prompt tracking model based on a universal bi-directional adapter, cross-prompting multiple modalities mutually. Our model consists of a universal bi-directional adapter and multiple modality-specific transformer encoder branches with sharing parameters. The encoders extract features of each modality separately by using a frozen, pre-trained foundation model. We develop a simple but effective light feature adapter to transfer modality-specific information from one modality to another, performing visual feature prompt fusion in an adaptive manner. With adding fewer (0.32M) trainable parameters, our model achieves superior tracking performance in comparison with both the full fine-tuning methods and the prompt learning-based methods. Our code is available: https://github.com/SparkTempest/BAT.
Bing Cao 0002, Junliang Guo, Pengfei Zhu 0001, Qinghua Hu
AAAI3
2024 Dynamic Sub-graph Distillation for Robust Semi-supervised Continual Learning
abstract
Continual learning (CL) has shown promising results and comparable performance to learning at once in a fully supervised manner. However, CL strategies typically require a large number of labeled samples, making their real-life deployment challenging. In this work, we focus on semi-supervised continual learning (SSCL), where the model progressively learns from partially labeled data with unknown categories. We provide a comprehensive analysis of SSCL and demonstrate that unreliable distributions of unlabeled data lead to unstable training and refinement of the progressing stages. This problem severely impacts the performance of SSCL. To address the limitations, we propose a novel approach called Dynamic Sub-Graph Distillation (DSGD) for semi-supervised continual learning, which leverages both semantic and structural information to achieve more stable knowledge distillation on unlabeled data and exhibit robustness against distribution bias. Firstly, we formalize a general model of structural distillation and design a dynamic graph construction for the continual learning progress. Next, we define a structure distillation vector and design a dynamic sub-graph distillation algorithm, which enables end-to-end training and adaptability to scale up tasks. The entire proposed method is adaptable to various CL methods and supervision settings. Finally, experiments conducted on three datasets CIFAR10, CIFAR100, and ImageNet-100, with varying supervision ratios, demonstrate the effectiveness of our proposed approach in mitigating the catastrophic forgetting problem in semi-supervised continual learning scenarios. Our code is available: https://github.com/fanyan0411/DSGD.
Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Qinghua Hu
AAAI3
2024 Exploring Diverse Representations for Open Set Recognition
abstract
Open set recognition (OSR) requires the model to classify samples that belong to closed sets while rejecting unknown samples during test. Currently, generative models often perform better than discriminative models in OSR, but recent studies show that generative models may be computationally infeasible or unstable on complex tasks. In this paper, we provide insights into OSR and find that learning supplementary representations can theoretically reduce the open space risk. Based on the analysis, we propose a new model, namely Multi-Expert Diverse Attention Fusion (MEDAF), that learns diverse representations in a discriminative way. MEDAF consists of multiple experts that are learned with an attention diversity regularization term to ensure the attention maps are mutually different. The logits learned by each expert are adaptively fused and used to identify the unknowns through the score function. We show that the differences in attention maps can lead to diverse representations so that the fused representations can well handle the open space. Extensive experiments are conducted on standard and OSR large-scale benchmarks. Results show that the proposed discriminative method can outperform existing generative models by up to 9.5% on AUROC and achieve new state-of-the-art performance with little computational cost. Our method can also seamlessly integrate existing classification models. Code is available at https://github.com/Vanixxz/MEDAF.
Yu Wang 0106, Junxian Mu, Pengfei Zhu 0001, Qinghua Hu
AAAI3
2024 AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning
abstract
Recently, pre-trained vision-language models (e.g., CLIP) have shown great potential in few-shot learning and attracted a lot of research interest. Although efforts have been made to improve few-shot ability of CLIP, key factors on the effectiveness of existing methods have not been well studied, limiting further exploration of CLIP's potential in few-shot learning. In this paper, we first introduce a uni-fied formulation to analyze CLIP-based few-shot learning methods from a perspective of logit bias, which encourages us to learn an effective logit bias for further improving per-formance of CLIP-based few-shot learning methods. To this end, we disassemble three key components involved in computation of logit bias (i.e., logit features, logit predictor, and logit fusion) and empirically analyze the effect on per-formance of few-shot classification. Based on analysis of key components, this paper proposes a novel AMU-Tuning method to learn effective logit bias for CLIP-based few-shot classification. Specifically, our AMU-Tuning predicts logit bias by exploiting the appropriate Auxiliary features, which are fed into an efficient feature-initialized linear clas-sifier with Multi-branch training. Finally, an Uncertainty-based fusion is developed to incorporate logit bias into CLIP for few-shot classification. The experiments are con-ducted on several widely used benchmarks, and the re-sults show AMU-Tuning clearly outperforms its counter-parts while achieving state-of-the-art performance of CLIP-based few-shot learning without bells and whistles.
Yuwei Tang, Zhenyi Lin, Qilong Wang 0001, Pengfei Zhu 0001, Qinghua Hu
CVPR4
2024 Task-Customized Mixture of Adapters for General Image Fusion
abstract
General image fusion aims at integrating important in-formation from multi-source images. However, due to the significant cross-task gap, the respective fusion mechanism varies considerably in practice, resulting in limited performance across subtasks. To handle this problem, we pro-pose a novel task-customized mixture of adapters (TC-MoA) for general image fusion, adaptively prompting various fusion tasks in a unified model. We borrow the insight from the mixture of experts (MoE), taking the experts as effi-cient tuning adapters to prompt a pre-trained foundation model. These adapters are shared across different tasks and constrained by mutual information regularization, ensuring compatibility with different tasks while complementarity for multi-source images. The task-specific routing networks customize these adapters to extract task-specific information from different sources with dynamic dominant inten-sity, performing adaptive visual feature prompt fusion. No-tably, our TC-MoA controls the dominant intensity bias for different fusion tasks, successfully unifying multiple fusion tasks in a single model. Extensive experiments show that TC-MoA outperforms the competing approaches in learning commonalities while retaining compatibility for gen-eral image fusion (multi-modal, multi-exposure, and multi-focus), and also demonstrating striking controllability on more generalization experiments. The code is available at https://github.com/YangSun22/TC-MoA.
Pengfei Zhu 0001, Bing Cao 0002, Qinghua Hu
CVPR1
2024 Visible and Clear: Finding Tiny Objects in Difference Map
Bing Cao 0002, Haiyu Yao, Pengfei Zhu 0001, Qinghua Hu
ECCV (17)3
2024 Socialized Learning: Making Each Other Better Through Multi-Agent Collaboration
abstract
Learning new knowledge frequently occurs in our dynamically changing world, e.g., humans culturally evolve by continuously acquiring new abilities to sustain their survival, leveraging collective intelligence rather than a large number of individual attempts. The effective learning paradigm during cultural evolution is termed socialized learning (SL). Consequently, a straightforward question arises: Can multi-agent systems acquire more new abilities like humans? In contrast to most existing methods that address continual learning and multi-agent collaboration, our emphasis lies in a more challenging problem: we prioritize the knowledge in the original expert classes, and as we adeptly learn new ones, the accuracy in the original expert classes stays superior among all in a directional manner. Inspired by population genetics and cognitive science, leading to unique and complete development, we propose Multi-Agent Socialized Collaboration (MASC), which achieves SL through interactions among multiple agents. Specifically, we introduce collective collaboration and reciprocal altruism modules, organizing collaborative behaviors, promoting information sharing, and facilitating learning and knowledge interaction among individuals. We demonstrate the effectiveness of multi-agent collaboration in an extensive empirical study. Our code will be publicly available at https://github.com/yxjdarren/SL.
Xinjie Yao, Yu Wang 0106, Pengfei Zhu 0001, Wanyu Lin, Qinghua Hu
ICML3
2024 Dynamic Brightness Adaptation for Robust Multi-modal Image Fusion
Yiming Sun 0003, Bing Cao 0002, Pengfei Zhu 0001, Qinghua Hu
IJCAI3
2024 Following in the Footsteps: Predicting Human Trajectories Using Motion Pattern Memory
Yuxin Yang 0008, Pengfei Zhu 0001, Mengshi Qi, Huadong Ma
MMAsia2
2024 Conditional Controllable Image Fusion
abstract
Image fusion aims to integrate complementary information from multiple input images acquired through various sources to synthesize a new fused image. Existing methods usually employ distinct constraint designs tailored to specific scenes, forming fixed fusion paradigms. However, this data-driven fusion approach is challenging to deploy in varying scenarios, especially in rapidly changing environments. To address this issue, we propose a conditional controllable fusion (CCF) framework for general image fusion tasks without specific training. Due to the dynamic differences of different samples, our CCF employs specific fusion constraints for each individual in practice. Given the powerful generative capabilities of the denoising diffusion model, we first inject the specific constraints into the pre-trained DDPM as adaptive fusion conditions. The appropriate conditions are dynamically selected to ensure the fusion process remains responsive to the specific requirements in each reverse diffusion stage. Thus, CCF enables conditionally calibrating the fused images step by step. Extensive experiments validate our effectiveness in general fusion tasks across diverse scenarios against the competing methods without additional training. The code is publicly available.
Bing Cao 0002, Xingxin Xu, Pengfei Zhu 0001, Qilong Wang 0001, Qinghua Hu
NeurIPS3
2024 Persistence Homology Distillation for Semi-supervised Continual Learning
abstract
Semi-supervised continual learning (SSCL) has attracted significant attention for addressing catastrophic forgetting in semi-supervised data. Knowledge distillation, which leverages data representation and pair-wise similarity, has shown significant potential in preserving information in SSCL. However, traditional distillation strategies often fail in unlabeled data with inaccurate or noisy information, limiting their efficiency in feature spaces undergoing substantial changes during continual learning. To address these limitations, we propose Persistence Homology Distillation (PsHD) to preserve intrinsic structural information that is insensitive to noise in semi-supervised continual learning. First, we capture the structural features using persistence homology by homological evolution across different scales in vision data, where the multi-scale characteristic established its stability under noise interference. Next, we propose a persistence homology distillation loss in SSCL and design an acceleration algorithm to reduce the computational cost of persistence homology in our module. Furthermore, we demonstrate the superior stability of PsHD compared to sample representation and pair-wise similarity distillation methods theoretically and experimentally. Finally, experimental results on three widely used datasets validate that the new PsHD outperforms state-of-the-art with 3.9% improvements on average, and also achieves 1.5% improvements while reducing 60% memory buffer size, highlighting the potential of utilizing unlabeled data in SSCL. Our code is available: https://github.com/fanyan0411/PsHD.
Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Qinghua Hu
NeurIPS3
2024 What Matters in Graph Class Incremental Learning? An Information Preservation Perspective
abstract
Graph class incremental learning (GCIL) requires the model to classify emerging nodes of new classes while remembering old classes. Existing methods are designed to preserve effective information of old models or graph data to alleviate forgetting, but there is no clear theoretical understanding of what matters in information preservation. In this paper, we consider that present practice suffers from high semantic and structural shifts assessed by two devised shift metrics. We provide insights into information preservation in GCIL and find that maintaining graph information can preserve information of old models in theory to calibrate node semantic and graph structure shifts. We correspond graph information into low-frequency local-global information and high-frequency information in spatial domain. Based on the analysis, we propose a framework, Graph Spatial Information Preservation (GSIP). Specifically, for low-frequency information preservation, the old node representations obtained by inputting replayed nodes into the old model are aligned with the outputs of the node and its neighbors in the new model, and then old and new outputs are globally matched after pooling. For high-frequency information preservation, the new node representations are encouraged to imitate the near-neighbor pair similarity of old node representations. GSIP achieves a 10\% increase in terms of the forgetting metric compared to prior methods on large-scale datasets. Our framework can also seamlessly integrate existing replay designs. The code is available through https://github.com/Jillian555/GSIP.
Yu Wang 0106, Pengfei Zhu 0001, Wanyu Lin, Qinghua Hu
NeurIPS3
2024 M-RRFS: A Memory-Based Robust Region Feature Synthesizer for Zero-Shot Object Detection
Peiliang Huang, Dingwen Zhang, De Cheng, Longfei Han, Pengfei Zhu 0001, Junwei Han 0001
Int. J. Comput. Vis.5
2024 Integrated Heterogeneous Graph Attention Network for Incomplete Multi-modal Clustering
Yu Wang 0106, Xinjie Yao, Pengfei Zhu 0001, Qinghua Hu
Int. J. Comput. Vis.3
2024 Improved generative adversarial network with deep metric learning for missing data imputation
Mohammed Al-taezi, Yu Wang 0106, Pengfei Zhu 0001, Qinghua Hu, Abdulrahman Al-Badwi
Neurocomputing3
2024 Stabilizing Multispectral Pedestrian Detection With Evidential Hybrid Fusion
abstract
Multispectral pedestrian detection is an important task due to its critical role in a wide spectrum of applications. Basically, the complementary information from color and thermal images could provide a more accurate and reliable pedestrian detection result. However, multimodal data usually suffer from the issue of dynamic change or corruption for some modalities. At the same time, as a safety-critical task, how to produce a stable and reliable detection result is also a key challenge. To address these challenges, we propose a stable multispectral pedestrian detection (SMPD) algorithm, providing a new paradigm for multispectral detection by dynamically integrating different modalities at an evidence level. Specifically, we introduce the Dirichlet distribution to characterize the distribution of the class probabilities, parameterized with evidence from different modalities. Then, multi-branch fusion, based on Dempster-Shafer theory, can integrate these pieces of evidence to obtain the detection result. In addition, a Plug-and-Play module, termed modal enhancement module, is introduced to enhance cross-modality interaction. This is an end-to-end framework, which can induce accurate detection and uncertainty estimation, and then endows the model with both reliability and robustness against noise or corruption. Extensive experimental results demonstrate the efficiency of our algorithm compared with state-of-the-art methods.
Qing Li 0018, Changqing Zhang 0002, Qinghua Hu, Pengfei Zhu 0001, Huazhu Fu, Lei Chen 0011
IEEE Trans. Circuits Syst. Video Technol.4
2024 Multiview Deep Subspace Clustering Networks
abstract
Multiview subspace clustering aims to discover the inherent structure of data by fusing multiple views of complementary information. Most existing methods first extract multiple types of handcrafted features and then learn a joint affinity matrix for clustering. The disadvantage of this approach lies in two aspects: 1) multiview relations are not embedded into feature learning and 2) the end-to-end learning manner of deep learning is not suitable for multiview clustering. Even when deep features have been extracted, it is a nontrivial problem to choose a proper backbone for clustering on different datasets. To address these issues, we propose the multiview deep subspace clustering networks (MvDSCNs), which learns a multiview self-representation matrix in an end-to-end manner. The MvDSCN consists of two subnetworks, i.e., a diversity network (Dnet) and a universality network (Unet). A latent space is built using deep convolutional autoencoders, and a self-representation matrix is learned in the latent space using a fully connected layer. Dnet learns view-specific self-representation matrices, whereas Unet learns a common self-representation matrix for all views. To exploit the complementarity of multiview representations, the Hilbert-Schmidt independence criterion (HSIC) is introduced as a diversity regularizer that captures the nonlinear, high-order interview relations. Because different views share the same label space, the self-representation matrices of each view are aligned to the common one by universality regularization. The MvDSCN also unifies multiple backbones to boost clustering performance and avoid the need for model selection. Experiments demonstrate the superiority of the MvDSCN.
Pengfei Zhu 0001, Xinjie Yao, Yu Wang 0106, Binyuan Hui, Dawei Du, Qinghua Hu
IEEE Trans. Cybern.1
2024 Autoencoder-Based Collaborative Attention GAN for Multi-Modal Image Synthesis
abstract
Multi-modal images are required in a wide range of practical scenarios, from clinical diagnosis to public security. However, certain modalities may be incomplete or unavailable because of the restricted imaging conditions, which commonly leads to decision bias in many real-world applications. Despite the significant advancement of existing image synthesis techniques, learning complementary information from multi-modal inputs remains challenging. To address this problem, we propose an autoencoder-based collaborative attention generative adversarial network (ACA-GAN) that uses available multi-modal images to generate the missing ones. The collaborative attention mechanism deploys a single-modal attention module and a multi-modal attention module to effectively extract complementary information from multiple available modalities. Considering the significant modal gap, we further developed an autoencoder network to extract the self-representation of target modality, guiding the generative model to fuse target-specific information from multiple modalities. This considerably improves cross-modal consistency with the desired modality, thereby greatly enhancing the image synthesis performance. Quantitative and qualitative comparisons for various multi-modal image synthesis tasks highlight the superiority of our approach over several prior methods by demonstrating more precise and realistic results.
Bing Cao 0002, Haifang Cao, Pengfei Zhu 0001, Changqing Zhang 0002, Qinghua Hu
IEEE Trans. Multim.4
2024 Multi-View Knowledge Ensemble With Frequency Consistency for Cross-Domain Face Translation
abstract
Cross-domain face translation aims to transfer face images from one domain to another. It can be widely used in practical applications, such as photos/sketches in law enforcement, photos/drawings in digital entertainment, and near-infrared (NIR)/visible (VIS) images in security access control. Restricted by limited cross-domain face image pairs, the existing methods usually yield structural deformation or identity ambiguity, which leads to poor perceptual appearance. To address this challenge, we propose a multi-view knowledge (structural knowledge and identity knowledge) ensemble framework with frequency consistency (MvKE-FC) for cross-domain face translation. Due to the structural consistency of facial components, the multi-view knowledge learned from large-scale data can be appropriately transferred to limited cross-domain image pairs and significantly improve the generative performance. To better fuse multi-view knowledge, we further design an attention-based knowledge aggregation module that integrates useful information, and we also develop a frequency-consistent (FC) loss that constrains the generated images in the frequency domain. The designed FC loss consists of a multidirection Prewitt (mPrewitt) loss for high-frequency consistency and a Gaussian blur loss for low-frequency consistency. Furthermore, our FC loss can be flexibly applied to other generative models to enhance their overall performance. Extensive experiments on multiple cross-domain face datasets demonstrate the superiority of our method over state-of-the-art methods both qualitatively and quantitatively.
Bing Cao 0002, Pengfei Zhu 0001, Qinghua Hu, Dongwei Ren, Wangmeng Zuo, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Learning Dynamic Compact Memory Embedding for Deformable Visual Object Tracking
abstract
Recently, template-based trackers have become the leading tracking algorithms with promising performance in terms of efficiency and accuracy. However, the correlation operation between query feature and the given template only achieves accurate target localization, but is prone to state estimation error, especially when the target suffers from severe deformation. To address this issue, segmentation-based trackers are proposed that use per-pixel matching to improve the tracking performance of deformable objects effectively. However, most of the existing trackers only match with the target features of the initial frame, thereby lacking the discrimination for handling a variety of challenging factors, e.g., similar distractors, background clutter, and appearance change. To this end, we propose a dynamic compact memory embedding technique to enhance the discrimination of the segmentation-based visual tracking method that can well tell the target from the background. Specifically, we initialize a memory embedding with the target features in the first frame. During the tracking process, the current target features that have certain correlation with the existing memory are updated to the memory embedding online. To further improve the tracking accuracy for deformable objects, we use a weighted point-to-global matching strategy to measure the correlation between the pixelwise query feature and the whole template, so as to capture more detailed deformation information. Extensive evaluations on six challenging tracking benchmarks including VOT2016, VOT2018, VOT2019, GOT-10K, TrackingNet, and LaSOT demonstrate the superiority of our method over recent remarkable trackers. Besides, our tracker outperforms the excellent segmentation-based trackers, i.e., D3S and SiamMask on the DAVIS2017 benchmark. The code is available at https://github.com/peace-love243/CMEDFL.
Pengfei Zhu 0001, Kaihua Zhang 0001, Yu Wang 0106, Tianzhu Zhang 0001, Qinghua Hu
IEEE Trans. Neural Networks Learn. Syst.2
2024 Multi-Task Credible Pseudo-Label Learning for Semi-Supervised Crowd Counting
abstract
As a widely used semi-supervised learning strategy, self-training generates pseudo-labels to alleviate the labor-intensive and time-consuming annotation problems in crowd counting while boosting the model performance with limited labeled data and massive unlabeled data. However, the noise in the pseudo-labels of the density maps greatly hinders the performance of semi-supervised crowd counting. Although auxiliary tasks, e.g., binary segmentation, are utilized to help improve the feature representation learning ability, they are isolated from the main task, i.e., density map regression and the multi-task relationships are totally ignored. To address the above issues, we develop a multi-task credible pseudo-label learning (MTCP) framework for crowd counting, consisting of three multi-task branches, i.e., density regression as the main task, and binary segmentation and confidence prediction as the auxiliary tasks. Multi-task learning is conducted on the labeled data by sharing the same feature extractor for all three tasks and taking multi-task relations into account. To reduce epistemic uncertainty, the labeled data are further expanded, by trimming the labeled data according to the predicted confidence map for low-confidence regions, which can be regarded as an effective data augmentation strategy. For unlabeled data, compared with the existing works that only use the pseudo-labels of binary segmentation, we generate credible pseudo-labels of density maps directly, which can reduce the noise in pseudo-labels and therefore decrease aleatoric uncertainty. Extensive comparisons on four crowd-counting datasets demonstrate the superiority of our proposed model over the competing methods. The code is available at: https://github.com/ljq2000/MTCP.
Pengfei Zhu 0001, Jingqing Li, Bing Cao 0002, Qinghua Hu
IEEE Trans. Neural Networks Learn. Syst.1
2023 Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image Fusion
abstract
Infrared and visible image fusion aims to integrate comprehensive information from multiple sources to achieve superior performances on various practical tasks, such as detection, over that of a single modality. However, most existing methods directly combined the texture details and object contrast of different modalities, ignoring the dynamic changes in reality, which diminishes the visible texture in good lighting conditions and the infrared contrast in low lighting conditions. To fill this gap, we propose a dynamic image fusion framework with a multi-modal gated mixture of local-to-global experts, termed MoE-Fusion, to dynamically extract effective and comprehensive information from the respective modalities. Our model consists of a Mixture of Local Experts (MoLE) and a Mixture of Global Experts (MoGE) guided by a multi-modal gate. The MoLE performs specialized learning of multi-modal local features, prompting the fused images to retain the local information in a sample-adaptive manner, while the MoGE focuses on the global information that complements the fused image with overall texture detail and contrast. Extensive experiments show that our MoE-Fusion outperforms state-of-the-art methods in preserving multi-modal image texture and contrast through the local-to-global dynamic learning paradigm, and also achieves superior performance on detection tasks. Our code is available: https://github.com/SunYM2020/MoE-Fusion.
Bing Cao 0002, Yiming Sun 0003, Pengfei Zhu 0001, Qinghua Hu
ICCV3
2023 Tuning Pre-trained Model via Moment Probing
abstract
Recently, efficient fine-tuning of large-scale pre-trained models has attracted increasing research interests, where linear probing (LP) as a fundamental module is involved in exploiting the final representations for task-dependent classification. However, most of the existing methods focus on how to effectively introduce a few of learnable parameters, and little work pays attention to the commonly used LP module. In this paper, we propose a novel Moment Probing (MP) method to further explore the potential of LP. Distinguished from LP which builds a linear classification head based on the mean of final features (e.g., word tokens for ViT) or classification tokens, our MP performs a linear classifier on feature distribution, which provides the stronger representation ability by exploiting richer statistical information inherent in features. Specifically, we represent feature distribution by its characteristic function, which is efficiently approximated by using first- and second-order moments of features. Furthermore, we propose a multi-head convolutional cross-covariance (MHC3) to compute second-order moments in an efficient and effective manner. By considering that MP could affect feature learning, we introduce a partially shared module to learn two recalibrating parameters (PSRP) for backbones based on MP, namely MP+. Extensive experiments on ten benchmarks using various models show that our MP significantly outperforms LP and is competitive with counterparts at lower training cost, while our MP+achieves state-of-the-art performance.
Qilong Wang 0001, Zhenyi Lin, Pengfei Zhu 0001, Qinghua Hu
ICCV4
2023 Unsupervised Self-Driving Attention Prediction via Uncertainty Mining and Knowledge Embedding
abstract
Predicting attention regions of interest is an important yet challenging task for self-driving systems. Existing methodologies rely on large-scale labeled traffic datasets that are labor-intensive to obtain. Besides, the huge domain gap between natural scenes and traffic scenes in current datasets also limits the potential for model training. To address these challenges, we are the first to introduce an unsupervised way to predict self-driving attention by uncertainty modeling and driving knowledge integration. Our approach’s Uncertainty Mining Branch (UMB) discovers commonalities and differences from multiple generated pseudo-labels achieved from models pre-trained on natural scenes by actively measuring the uncertainty. Meanwhile, our Knowledge Embedding Block (KEB) bridges the domain gap by incorporating driving knowledge to adaptively refine the generated pseudo-labels. Quantitative and qualitative results with equivalent or even more impressive performance compared to fully-supervised state-of-the-art approaches across all three public datasets demonstrate the effectiveness of the proposed method and the potential of this direction. The code is available at https://github.com/zaplm/DriverAttention.
Pengfei Zhu 0001, Mengshi Qi, Weijian Li 0001, Huadong Ma
ICCV1
2023 SVBRDF Reconstruction by Transferring Lighting Knowledge
abstract
Abstract The problem of reconstructing spatially‐varying BRDFs from RGB images has been studied for decades. Researchers found themselves in a dilemma: opting for either higher quality with the inconvenience of camera and light calibration, or greater convenience at the expense of compromised quality without complex setups. We address this challenge by introducing a two‐branch network to learn the lighting effects in images. The two branches, referred to as Light‐known and Light‐aware, diverge in their need for light information. The Light‐aware branch is guided by the Light‐known branch to acquire the knowledge of discerning light effects and surface reflectance properties, but without the reliance of light positions. Both branches are trained using the synthetic dataset, but during testing on real‐world cases without calibration, only the Light‐aware branch is activated. To facilitate a more effective utilization of various light conditions, we employ gated recurrent units (GRUs) to fuse the features extracted from different images. The two modules mutually benefit when multiple inputs are provided. We present our reconstructed results on both synthetic and real‐world examples, demonstrating high quality while maintaining a lightweight characteristic in comparison to previous methods.
Pengfei Zhu 0001, Shuichang Lai, Mufan Chen, Jie Guo 0001, Yanwen Guo 0001
Comput. Graph. Forum1
2023 Towards a Deeper Understanding of Global Covariance Pooling in Deep Learning: An Optimization Perspective
abstract
Global covariance pooling (GCP) as an effective alternative to global average pooling has shown good capacity to improve deep convolutional neural networks (CNNs) in a variety of vision tasks. Although promising performance, it is still an open problem on how GCP (especially its post-normalization) works in deep learning. In this paper, we make the effort towards understanding the effect of GCP on deep learning from an optimization perspective. Specifically, we first analyze behavior of GCP with matrix power normalization on optimization loss and gradient computation of deep architectures. Our findings show that GCP can improve Lipschitzness of optimization loss and achieve flatter local minima, while improving gradient predictiveness and functioning as a special pre-conditioner on gradients. Then, we explore the effect of post-normalization on GCP from the model optimization perspective, which encourages us to propose a simple yet effective normalization, namely DropCov. Based on above findings, we point out several merits of deep GCP that have not been recognized previously or fully explored, including faster convergence, stronger model robustness and better generalization across tasks. Extensive experimental results using both CNNs and vision transformers on diversified vision tasks provide strong support to our findings while verifying the effectiveness of our method.
Qilong Wang 0001, Jiangtao Xie, Pengfei Zhu 0001, Peihua Li, Wangmeng Zuo, Qinghua Hu
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Multi-head second-order pooling for graph transformer networks
Qilong Wang 0001, Pengfei Zhu 0001
Pattern Recognit. Lett.3
2023 Cross-Drone Transformer Network for Robust Single Object Tracking
abstract
Drones have been widely used in a variety of applications, e.g., aerial photography and military security, because of their high maneuverability and broad views compared with fixed cameras. Multi-drone tracking systems can provide rich information about targets by collecting complementary video clips from different views, especially when targets are occluded or disappear in some views. However, it is challenging to handle cross-drone information interaction and multi-drone information fusion in multi-drone visual tracking. Recently, Transformer has shown significant advantages in automatically modeling the correlation between templates and search regions for visual tracking. To leverage its potential in multi-drone tracking, we propose a novel cross-drone Transformer network (TransMDOT) for visual object tracking tasks. The self-attention mechanism is used to automatically capture the correlation between multiple templates and the corresponding search region to achieve multi-drone feature fusion. During the tracking process, a cross-drone mapping mechanism is proposed by using the surrounding information of the drone with promising tracking status as reference, assisting drones that lost targets to re-calibrate, which implements real-time cross-drone information interaction. As the existing multi-drone evaluation metrics only consider spatial information while ignore temporal information, we further present a system perception index (SPFI) that combines both temporal and spatial information to evaluate the tracking status of multiple drones. Experiments on the MDOT dataset prove that TransMDOT greatly surpasses the state-of-the-art methods in both single-drone performance and multi-drone system fusion performance. Our code will be available onhttps://github.com/cgjacklin/transmdot.
Pengfei Zhu 0001, Bing Cao 0002, Xing Wang 0002, Qinghua Hu
IEEE Trans. Circuits Syst. Video Technol.2
2023 OpenMix+: Revisiting Data Augmentation for Open Set Recognition
abstract
Open set recognition requires models to recognize samples of known classes learned in the training set while reject unknowns not learned. Compared with the structural risk minimization theory for closed-set problems, structural risk in open set tasks remains rarely explored. In this paper, we point out that balancing between structural risk and open space risk is crucial for open set recognition, and re-formalize it as open set structural risk. This brings a new view towards the general relationship between closed set recognition and open set recognition against the common intuition, which argues that a good closed set classifier always benefits for open set recognition. Specifically, we theoretically and experimentally show that recent mix-based data augmentation methods are aggressive closed set regularization methods, which reduce structural risk at cost of sacrificing open space risk. Besides, we show that existing negative data augmentation designed for open space risk reduction also ignore the trade-off problem between structural risk and open space risk, which limits their performance. We propose an efficient negative data augmentation strategy named self-mix and a corresponding method named OpenMix. OpenMix generates high-quality negative samples by mixing samples themselves, which can take care of both risks simultaneously. When combining OpenMix with conservative closed set regularization methods to form OpenMix+, models can achieve lower open set structural risk. Extensive experiments validate the superiority of OpenMix and OpenMix+ in terms of both effectiveness and universality.
Guosong Jiang, Pengfei Zhu 0001, Yu Wang 0106, Qinghua Hu
IEEE Trans. Circuits Syst. Video Technol.2
2023 Confidence-Aware Fusion Using Dempster-Shafer Theory for Multispectral Pedestrian Detection
abstract
Multispectral pedestrian detection is an important and valuable task in many applications, which could provide a more accurate and reliable pedestrian detection result by using the complementary visual information from color and thermal images. However, it faces two open and difficult challenges: 1) how to effectively and dynamically integrate multispectral information according to the confidence of different modalities, and 2) how to produce a reliable prediction result. In this paper, we propose a novel confidence-aware multispectral pedestrian detection (CMPD) method, which flexibly learns the multispectral representation while simultaneously producing a reliable result with confidence estimation. Specifically, a dense fusion strategy is first proposed to extract the multilevel multispectral representation at the feature level. Then, an additional confidence subnetwork is utilized to dynamically estimate the detection confidence for each modality. Finally, Dempster's combination rule is introduced to fuse the results of different branches according to the rectified confidence. Our proposed CMPD method not only effectively integrates multimodal information but also provides a reliable prediction. Extensive experimental results demonstrate the efficiency of our algorithm compared with state-of-the-art methods.
Qing Li 0018, Changqing Zhang 0002, Qinghua Hu, Huazhu Fu, Pengfei Zhu 0001
IEEE Trans. Multim.5
2023 Robust Multi-Drone Multi-Target Tracking to Resolve Target Occlusion: A Benchmark
abstract
Multi-drone multi-target tracking aims at collabo- ratively detecting and tracking targets across multiple drones and associating the identities of objects from different drones, which can overcome the shortcomings of single-drone object tracking. To address the critical challenges of identity association and target occlusion in multi-drone multi-target tracking tasks, we collect an occlusion-aware multi-drone multi-target tracking dataset named MDMT. It contains 88 video sequences with 39,678 frames, including 11,454 different IDs of persons, bicycles, and cars. The MDMT dataset comprises 2,204,620 bounding boxes, of which 543,444 bounding boxes contain target occlusions. We also design a multi-device target association score (MDA) as the evaluation criteria for the ability of cross-view target association in multi-device tracking. Furthermore, we propose a Multi-matching Identity Authentication network (MIA-Net) for the multi-drone multi-target tracking task. The local-global matching algorithm in MIA-Net discovers the topological relationship of targets across drones, efficiently solves the problem of cross-drone association, and effectively complements occluded targets with the advantage of multiple drone view mapping. Extensive experiments on the MDMT dataset validate the effectiveness of our proposed MIA-Net for the task of identity association and multi-object tracking with occlusions.
Timing Li, Yu Wang 0106, Qinghua Hu, Pengfei Zhu 0001
IEEE Trans. Multim.7
2023 Latent Heterogeneous Graph Network for Incomplete Multi-View Learning
abstract
Multi-view learning has progressed rapidly in recent years. Although many previous studies assume that each instance appears in all views, it is common in real-world applications for instances to be missing from some views, resulting in incomplete multi-view data. To tackle this problem, we propose a novel Latent Heterogeneous Graph Network (LHGN) for incomplete multi-view learning, which aims to use multiple incomplete views as fully as possible in a flexible manner. By learning a unified latent representation, a trade-off between consistency and complementarity among different views is implicitly realized. To explore the complex relationship between samples and latent representations, a neighborhood constraint and a view-existence constraint are proposed, for the first time, to construct a heterogeneous graph. Finally, to avoid any inconsistencies between training and test phase, a transductive learning technique is applied based on graph learning for classification tasks. Extensive experimental results on real-world datasets demonstrate the effectiveness of our model over existing state-of-the-art approaches. Our code is available at:https://github.com/yxjdarren/LHGN_TMM_2022.
Pengfei Zhu 0001, Xinjie Yao, Yu Wang 0106, Binyuan Hui, Qinghua Hu
IEEE Trans. Multim.1
2023 Collaborative Decision-Reinforced Self-Supervision for Attributed Graph Clustering
abstract
Attributed graph clustering aims to partition nodes of a graph structure into different groups. Recent works usually use variational graph autoencoder (VGAE) to make the node representations obey a specific distribution. Although they have shown promising results, how to introduce supervised information to guide the representation learning of graph nodes and improve clustering performance is still an open problem. In this article, we propose a Collaborative Decision-Reinforced Self-Supervision (CDRS) method to solve the problem, in which a pseudo node classification task collaborates with the clustering task to enhance the representation learning of graph nodes. First, a transformation module is used to enable end-to-end training of existing methods based on VGAE. Second, the pseudo node classification task is introduced into the network through multitask learning to make classification decisions for graph nodes. The graph nodes that have consistent decisions on clustering and pseudo node classification are added to a pseudo-label set, which can provide fruitful self-supervision for subsequent training. This pseudo-label set is gradually augmented during training, thus reinforcing the generalization capability of the network. Finally, we investigate different sorting strategies to further improve the quality of the pseudo-label set. Extensive experiments on multiple datasets show that the proposed method achieves outstanding performance compared with state-of-the-art methods. Our code is available at https://github.com/Jillian555/TNNLS_CDRS.
Pengfei Zhu 0001, Yu Wang 0106, Bin Xiao 0002, Qinghua Hu
IEEE Trans. Neural Networks Learn. Syst.1
2022 Label-Efficient Hybrid-Supervised Learning for Medical Image Segmentation
abstract
Due to the lack of expertise for medical image annotation, the investigation of label-efficient methodology for medical image segmentation becomes a heated topic. Recent progresses focus on the efficient utilization of weak annotations together with few strongly-annotated labels so as to achieve comparable segmentation performance in many unprofessional scenarios. However, these approaches only concentrate on the supervision inconsistency between strongly- and weakly-annotated instances but ignore the instance inconsistency inside the weakly-annotated instances, which inevitably leads to performance degradation. To address this problem, we propose a novel label-efficient hybrid-supervised framework, which considers each weakly-annotated instance individually and learns its weight guided by the gradient direction of the strongly-annotated instances, so that the high-quality prior in the strongly-annotated instances is better exploited and the weakly-annotated instances are depicted more precisely. Specially, our designed dynamic instance indicator (DII) realizes the above objectives, and is adapted to our dynamic co-regularization (DCR) framework further to alleviate the erroneous accumulation from distortions of weak annotations. Extensive experiments on two hybrid-supervised medical segmentation datasets demonstrate that with only 10% strong labels, the proposed framework can leverage the weak labels efficiently and achieve competitive performance against the 100% strong-label supervised scenario.
Junwen Pan, Qi Bi, Yanzhan Yang, Pengfei Zhu 0001, Cheng Bian
AAAI4
2022 DetFusion: A Detection-driven Infrared and Visible Image Fusion Network
abstract
Infrared and visible image fusion aims to utilize the complementary information between the two modalities to synthesize a new image containing richer information. Most existing works have focused on how to better fuse the pixel-level details from both modalities in terms of contrast and texture, yet ignoring the fact that the significance of image fusion is to better serve downstream tasks. For object detection tasks, object-related information in images is often more valuable than focusing on the pixel-level details of images alone. To fill this gap, we propose a detection-driven infrared and visible image fusion network, termed DetFusion, which utilizes object-related information learned in the object detection networks to guide multimodal image fusion. We cascade the image fusion network with the detection networks of both modalities and use the detection loss of the fused images to provide guidance on task-related information for the optimization of the image fusion network. Considering that the object locations provide a priori information for image fusion, we propose an object-aware content loss that motivates the fusion model to better learn the pixel-level information in infrared and visible images. Moreover, we design a shared attention module to motivate the fusion network to learn object-specific information from the object detection networks. Extensive experiments show that our DetFusion outperforms state-of-the-art methods in maintaining pixel intensity distribution and preserving texture details. More notably, the performance comparison with state-of-the-art image fusion methods in task-driven evaluation also demonstrates the superiority of the proposed method. Our code will be available: https://github.com/SunYM2020/DetFusion.
Yiming Sun 0003, Bing Cao 0002, Pengfei Zhu 0001, Qinghua Hu
ACM Multimedia3
2022 Learning Self-supervised Low-Rank Network for Single-Stage Weakly and Semi-supervised Semantic Segmentation
Junwen Pan, Pengfei Zhu 0001, Kaihua Zhang 0001, Bing Cao 0002, Yu Wang 0106, Dingwen Zhang, Junwei Han 0001, Qinghua Hu
Int. J. Comput. Vis.2
2022 Detection and Tracking Meet Drones Challenge
abstract
Drones, or general UAVs, equipped with cameras have been fast deployed with a wide range of applications, including agriculture, aerial photography, and surveillance. Consequently, automatic understanding of visual data collected from drones becomes highly demanding, bringing computer vision and drones more and more closely. To promote and track the developments of object detection and tracking algorithms, we have organized three challenge workshops in conjunction with ECCV 2018, ICCV 2019 and ECCV 2020, attracting more than 100 teams around the world. We provide a large-scale drone captured dataset, VisDrone, which includes four tracks, i.e., (1) image object detection, (2) video object detection, (3) single object tracking, and (4) multi-object tracking. In this paper, we first present a thorough review of object detection and tracking datasets and benchmarks, and discuss the challenges of collecting large-scale drone-based object detection and tracking datasets with fully manual annotations. After that, we describe our VisDrone dataset, which is captured over various urban/suburban areas of 14 different cities across China from North to South. Being the largest such dataset ever published, VisDrone enables extensive evaluation and investigation of visual analysis algorithms for the drone platform. We provide a detailed analysis of the current state of the field of large-scale object detection and tracking on drones, and conclude the challenge as well as propose future directions. We expect the benchmark largely boost the research and development in video analysis on drone platforms. All the datasets and experimental results can be downloaded from https://github.com/VisDrone/VisDrone-Dataset.
Pengfei Zhu 0001, Longyin Wen, Dawei Du, Xiao Bian, Heng Fan 0001, Qinghua Hu, Haibin Ling
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Multi-granularity episodic contrastive learning for few-shot learning
Pengfei Zhu 0001, Yu Wang 0106, Jinglin Zhang 0001
Pattern Recognit.1
2022 Semi-Supervised Image Deraining Using Knowledge Distillation
abstract
Image deraining has achieved considerable progress based on supervised learning with synthetic training pairs, but is usually limited in handling real-world rainy images. Although semi-supervised methods are suggested to exploit real-world rainy images when training deep deraining models, their performances are still notably inferior. To address this crucial issue, this work proposes a semi-supervised image deraining network with knowledge distillation (SSID-KD) for better exploiting real-world rainy images. In particular, the consistency of feature distribution of rain streaks extracted from synthetic and real-world rainy images is enforced by adopting knowledge distillation. Moreover, as for the backbone in SSID-KD, we propose the multi-scale feature fusion module and the pyramid fusion module to better extract deep features of rainy images. SSID-KD can relieve the problem of over-deraining or under-deraining for real-world rainy images, while it can keep comparable performance with supervised deraining methods on several benchmark datasets. Extensive experiments on both synthetic and real-world rainy images have validated that our SSID-KD not only can achieve better deraining results than existing semi-supervised deraining methods but also are quantitatively comparable with state-of-the-art supervised deraining methods. Benefiting from the well exploration of real-world rainy images, our SSID-KD can obtain more visually plausible deraining results. The source code and trained models are publicly available athttps://github.com/cuiyixin555/SSID-KD.
Cong Wang 0018, Dongwei Ren, Yunjin Chen, Pengfei Zhu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Drone-Based RGB-Infrared Cross-Modality Vehicle Detection Via Uncertainty-Aware Learning
abstract
Drone-based vehicle detection aims at detecting vehicle locations and categories in aerial images. It empowers smart city traffic management and disaster relief. Researchers have made a great deal of effort in this area and achieved considerable progress. However, because of the paucity of data under extreme conditions, drone-based vehicle detection remains a challenge when objects are difficult to distinguish, particularly in low-light conditions. To fill this gap, we constructed a large-scale drone-based RGB-infrared vehicle detection dataset called DroneVehicle, which contains 28, 439 RGB-infrared image pairs covering urban roads, residential areas, parking lots, and other scenarios from day to night. Cross-modal images provide complementary information for vehicle detection, but also introduce redundant information. To handle this dilemma, we further propose an uncertainty-aware cross-modality vehicle detection (UA-CMDet) framework to improve detection performance in complex environments. Specifically, we design an uncertainty-aware module using cross-modal intersection over union and illumination estimation to quantify the uncertainty of each object. Our method takes uncertainty as a weight to boost model learning more effectively while reducing bias caused by high-uncertainty objects. For more robust cross-modal integration, we further perform illumination-aware non-maximum suppression during inference. Extensive experiments on our DroneVehicle and two challenging RGB-infrared object detection datasets demonstrated the advanced flexibility and superior performance of UA-CMDet over competing methods. Our code and DroneVehicle will be available:https://github.com/VisDrone/DroneVehicle.
Yiming Sun 0003, Bing Cao 0002, Pengfei Zhu 0001, Qinghua Hu
IEEE Trans. Circuits Syst. Video Technol.3
2022 Semisupervised Laplace-Regularized Multimodality Metric Learning
abstract
Distance metric learning, which aims at learning an appropriate metric from data automatically, plays a crucial role in the fields of pattern recognition and information retrieval. A tremendous amount of work has been devoted to metric learning in recent years, but much of the work is basically designed for training a linear and global metric with labeled samples. When data are represented with multimodal and high-dimensional features and only limited supervision information is available, these approaches are inevitably confronted with a series of critical problems: 1) naive concatenation of feature vectors can cause the curse of dimensionality in learning metrics and 2) ignorance of utilizing massive unlabeled data may lead to overfitting. To mitigate this deficiency, we develop a semisupervised Laplace-regularized multimodal metric-learning method in this work, which explores a joint formulation of multiple metrics as well as weights for learning appropriate distances: 1) it learns a global optimal distance metric on each feature space and 2) it searches the optimal combination weights of multiple features. Experimental results demonstrate both the effectiveness and efficiency of our method on retrieval and classification tasks.
Jianqing Liang, Pengfei Zhu 0001, Chuangyin Dang, Qinghua Hu
IEEE Trans. Cybern.2
2022 Unsupervised Spectral Feature Selection With Dynamic Hyper-Graph Learning
abstract
Unsupervised spectral feature selection (USFS) methods could output interpretable and discriminative results by embedding a Laplacian regularizer in the framework of sparse feature selection to keep the local similarity of the training samples. To do this, USFS methods usually construct the Laplacian matrix using either a general-graph or a hyper-graph on the original data. Usually, a general-graph could measure the relationship between two samples while a hyper-graph could measure the relationship among no less than two samples. Obviously, the general-graph is a special case of the hyper-graph and the hyper-graph may capture more complex structure of samples than the general graph. However, in previous USFS methods, the construction of the Laplacian matrix is separated from the process of feature selection. Moreover, the original data usually contain noise. Each of them makes difficult to output reliable feature selection models. In this paper, we propose a novel feature selection method by dynamically constructing a hyper-graph based Laplacian matrix in the framework of sparse feature selection. Experimental results on real datasets showed that our proposed method outperformed the state-of-the-art methods in terms of both clustering and segmentation tasks.
Xiaofeng Zhu 0001, Shichao Zhang 0001, Yonghua Zhu, Pengfei Zhu 0001, Yue Gao 0002
IEEE Trans. Knowl. Data Eng.4
2021 Dynamic Hybrid Relation Exploration Network for Cross-Domain Context-Dependent Semantic Parsing
abstract
Semantic parsing has long been a fundamental problem in natural language processing. Recently, cross-domain context-dependent semantic parsing has become a new focus of research. Central to the problem is the challenge of leveraging contextual information of both natural language queries and database schemas in the interaction history. In this paper, we present a dynamic graph framework that is capable of effectively modelling contextual utterances, tokens, database schemas, and their complicated interaction as the conversation proceeds. The framework employs a dynamic memory decay mechanism that incorporates inductive bias to integrate enriched contextual relation representation, which is further enhanced with a powerful reranking model. At the time of writing, we demonstrate that the proposed framework outperforms all existing models by large margins, achieving new state-of-the-art performance on two large-scale benchmarks, the SParC and CoSQL datasets. Specifically, the model attains a 55.8% question-match and 30.8% interaction-match accuracy on SParC, and a 46.8% question-match and 17.0% interaction-match accuracy on CoSQL.
Binyuan Hui, Ruiying Geng, Qiyu Ren, Binhua Li, Yongbin Li 0001, Jian Sun 0021, Fei Huang 0002, Luo Si, Pengfei Zhu 0001, Xiaodan Zhu 0001
AAAI9
2021 Multi-View Information-Bottleneck Representation Learning
abstract
In real-world applications, clustering or classification can usually be improved by fusing information from different views. Therefore, unsupervised representation learning on multi-view data becomes a compelling topic in machine learning. In this paper, we propose a novel and flexible unsupervised multi-view representation learning model termed Collaborative Multi-View Information Bottleneck Networks (CMIB-Nets), which comprehensively explores the common latent structure and the view-specific intrinsic information, and discards the superfluous information in the data significantly improving the generalization capability of the model. Specifically, our proposed model relies on the information bottleneck principle to integrate the shared representation among different views and the view-specific representation of each view, prompting the multi-view complete representation and flexibly balancing the complementarity and consistency among multiple views. We conduct extensive experiments (including clustering analysis, robustness experiment, and ablation study) on real-world datasets, which empirically show promising generalization ability and robustness compared to state-of-the-arts.
Zhibin Wan, Changqing Zhang 0002, Pengfei Zhu 0001, Qinghua Hu
AAAI3
2021 Detection, Tracking, and Counting Meets Drones in Crowds: A Benchmark
abstract
To promote the developments of object detection, tracking and counting algorithms in drone-captured videos, we construct a benchmark with a new drone-captured large-scale dataset, named as DroneCrowd, formed by 112 video clips with 33, 600 HD frames in various scenarios. Notably, we annotate 20, 800 people trajectories with 4.8 million heads and several video-level attributes. Meanwhile, we design the Space-Time Neighbor-Aware Network (STNNet) as a strong baseline to solve object detection, tracking and counting jointly in dense crowds. STNNet is formed by the feature extraction module, followed by the density map estimation heads, and localization and association subnets. To exploit the context information of neighboring objects, we design the neighboring context loss to guide the association subnet training, which enforces consistent relative position of nearby objects in temporal domain. Extensive experiments on our DroneCrowd dataset demonstrate that STNNet performs favorably against the state-of-the-arts.
Longyin Wen, Dawei Du, Pengfei Zhu 0001, Qinghua Hu, Qilong Wang 0001, Liefeng Bo, Siwei Lyu
CVPR3
2021 Adaptive Correlation Filters Feature Fusion Learning for Visual Tracking
Pengfei Zhu 0001
ICANN (5)2
2021 Get to the Point: Content Classification of Animated Graphics Interchange Formats with Key-Frame Attention
abstract
Animated Graphics Interchange Formats (GIFS) are low-bandwidth short image sequences that can continuously display multiple frames without sound. In this paper, we focus on a new content classification task that is important in real-world applications. A key problem for this task is that some frames in an animated GF are irrelevant to the label, which may drastically reduce the classification performance. To this end, we first collect a new dataset of Web animated GIFS (WGF) that includes some typical samples in which only several key-frames are relevant to the ground truth. Then, an attention-based method is designed to learn to produce importance scores of the frames, and subsequently multi-frame predicted scores are merged to obtain the final prediction. Besides, an additional entropy loss is also used to sharpen the attention results to further emphasize the key-frames. Experimental results on WGF show that the proposed approach significantly outperforms various baseline methods.
Yongjuan Ma, Yu Wang 0106, Pengfei Zhu 0001, Junwen Pan
ICIP3
2021 Cross-View Equivariant Auto-Encoder
abstract
Unsupervised representation learning on multi-view data (multiple types of features or modalities) becomes a compelling topic in machine learning. Most existing methods focus on directly projecting different views into a common space to explore the consistency across different views. Al-though simple, the underlying relationships among different views are not guaranteed during the learning process. In this paper, we propose a novel unsupervised multi-view representation learning model termed as Cross-View Equivariant Auto-Encoder (CVE-AE), which jointly conducts data re-construction with view-specific autoencoder for information preservation within each view, and transformation reconstruction with transformation decoder for correlations preservation across different views. Accordingly, the generalization ability of our model is promoted due to the preserved intra-view intrinsic information and underlying inter-view relationships. We conduct extensive experiments on real-world datasets, and the proposed model achieves superior performance over state-of-the-art unsupervised representation learning methods.
Zhibin Wan, Changqing Zhang 0002, Huazhu Fu, Xi Peng 0001, Pengfei Zhu 0001, Qinghua Hu
ICME6
2021 Semi-supervised Single Image Deraining with Discrete Wavelet Transform
Wei Shang 0001, Dongwei Ren, Pengfei Zhu 0001, Yankun Gao
PRICAI (3)4
2021 Evolving Fully Automated Machine Learning via Life-Long Knowledge Anchors
abstract
Automated machine learning (AutoML) has achieved remarkable progress on various tasks, which is attributed to its minimal involvement of manual feature and model designs. However, most of existing AutoML pipelines only touch parts of the full machine learning pipeline, e.g., neural architecture search or optimizer selection. This leaves potentially important components such as data cleaning and model ensemble out of the optimization, and still results in considerable human involvement and suboptimal performance. The main challenges lie in the huge search space assembling all possibilities over all components, as well as the generalization ability over different tasks like image, text, and tabular etc. In this paper, we present a first-of-its-kind fully AutoML pipeline, to comprehensively automate data preprocessing, feature engineering, model generation/selection/training and ensemble for an arbitrary dataset and evaluation metric. Our innovation lies in the comprehensive scope of a learning pipeline, with a novel "life-long" knowledge anchor design to fundamentally accelerate the search over the full search space. Such knowledge anchors record detailed information of pipelines and integrates them with an evolutionary algorithm for joint optimization across components. Experiments demonstrate that the result pipeline achieves state-of-the-art performance on multiple datasets and modalities. Specifically, the proposed framework was extensively evaluated in the NeurIPS 2019 AutoDL challenge, and won the only champion with a significant gap against other approaches, on all the image, video, speech, text and tabular tracks.
Xiawu Zheng, Yang Zhang 0079, Sirui Hong, Huixia Li, Lang Tang, Youcheng Xiong, Yan Wang 0059, Xiaoshuai Sun, Pengfei Zhu 0001, Chenglin Wu 0001, Rongrong Ji
IEEE Trans. Pattern Anal. Mach. Intell.10
2021 Multi-Drone-Based Single Object Tracking With Agent Sharing Network
abstract
Drones equipped with cameras (UAVs) can dynamically track the target in the air from a broader view compared with static cameras or moving sensors over the ground. However, it is still challenging to accurately track the target using a single drone due to several factors such as appearance variations and severe occlusions. To this end, we collect a newMulti-Drone singleObjectTracking (MDOT) dataset that consists of 92 groups of video clips with 113, 918 high resolution frames taken by two drones and 63 groups of video clips with 145, 875 high resolution frames taken by three drones. Besides, two evaluation metrics are specially designed for multi-drone single object tracking,i.e., automatic fusion score (AFS) and ideal fusion score (IFS). Moreover, the agent sharing network (ASNet) is proposed by integrating self-supervised template sharing, target re-detection, and view-aware fusion of the target from multiple drones into a unified framework, which can improve the tracking accuracy significantly compared with single drone tracking. Extensive experiments on MDOT show that our ASNet significantly outperforms recent state-of-the-art trackers. The dataset can be found inhttps://github.com/VisDrone/MultiDrone.
Pengfei Zhu 0001, Jiayu Zheng, Dawei Du, Longyin Wen, Yiming Sun 0003, Qinghua Hu
IEEE Trans. Circuits Syst. Video Technol.1
2021 Graph Regularized Flow Attention Network for Video Animal Counting From Drones
abstract
In this paper, we propose a large-scale video based animal counting dataset collected by drones (AnimalDrone) for agriculture and wildlife protection. The dataset consists of two subsets, i.e., PartA captured on site by drones and PartB collected from the Internet, with rich annotations of more than 4 million objects in 53, 644 frames and corresponding attributes in terms of density, altitude and view. Moreover, we develop a new graph regularized flow attention network (GFAN) to perform density map estimation in dense crowds of video clips with arbitrary crowd density, perspective, and flight altitude. Specifically, our GFAN method leverages optical flow to warp the multi-scale feature maps in sequential frames to exploit the temporal relations, and then combines the enhanced features to predict the density maps. Moreover, we introduce the multi-granularity loss function including pixel-wise density loss and region-wise count loss to enforce the network to concentrate on discriminative features for different scales of objects. Meanwhile, the graph regularizer is imposed on the density maps of multiple consecutive frames to maintain temporal coherency. Extensive experiments are conducted to demonstrate the effectiveness of the proposed method, compared with several state-of-the-art counting algorithms. The AnimalDrone dataset is available at https://github.com/VisDrone/AnimalDrone.
Pengfei Zhu 0001, Dawei Du, Libo Zhang 0001, Qinghua Hu
IEEE Trans. Image Process.1
2021 A Recursive Regularization Based Feature Selection Framework for Hierarchical Classification
abstract
The sizes of datasets in terms of the number of samples, features, and classes have dramatically increased in recent years. In particular, there usually exists a hierarchical structure among class labels as hundreds of classes exist in a classification task. We call these tasks hierarchical classification, and hierarchical structures are helpful for dividing a very large task into a collection of relatively small subtasks. Various algorithms have been developed to select informative features for flat classification. However, these algorithms ignore the semantic hyponymy in the directory of hierarchical classes, and select a uniform subset of the features for all classes. In this paper, we propose a new feature selection framework with recursive regularization for hierarchical classification. This framework takes the hierarchical information of the class structure into account. In contrast to flat feature selection, we select different feature subsets for each node in a hierarchical tree structure with recursive regularization. The proposed framework uses parent-child, sibling, and family relationships for hierarchical regularization. By imposing$\ell _{2,1}$-norm regularization to different parts of the hierarchical classes, we can learn a sparse matrix for the feature ranking at each node. Extensive experiments on public datasets demonstrate the effectiveness and efficiency of the proposed algorithms.
Hong Zhao 0002, Qinghua Hu, Pengfei Zhu 0001, Yu Wang 0106, Ping Wang 0072
IEEE Trans. Knowl. Data Eng.3
2021 Adaptive and Robust Partition Learning for Person Retrieval With Policy Gradient
abstract
Person retrieval aims at effectively matching the pedestrian images over an extensive database given a specified identity. As extracting effective features is crucial in a high-performance retrieval system, recent significant progress was achieved by part-based models that have constructed robust local representations on top of vertically striped part features. However, this kind of models use predefined partitioning strategies, making the number and size of each partition identical even when input images vary a lot. This unchangeable setting usually leads to less flexibility and robustness in capturing visual variance. The primary reason for such a negative effect is that a fixed partitioning strategy is unable to deal with (a) the significant variance from pose, illumination and viewpoint which is common in a pedestrian image dataset, and (b) also the inference error and misalignment of human bodies introduced by the prepositive pedestrian detection module or human pose estimation module. In this paper, we tackle this problem via introducing the novel Adaptive Partition Network (APN). The APN utilizes deep reinforcement learning and applies an agent to generate optimal partitioning strategies dynamically for different input images. The agent inside the APN is optimized with the policy gradient algorithm and maximizes the reward of choosing the best partition setting. By leveraging the supervision cues from the objective partitioning strategies that are generated on a set of held-out training images, the agent is trained jointly with other parts of APN, which ensures the APN's robustness and generalization ability. Extensive experimental results on multiple datasets, including CUHK03, DukeMTMC and Market-1501, demonstrate the superiority of APN over the state-of-the-art models.
Yuxuan Shi 0002, Zhen Wei 0001, Pengfei Zhu 0001, Jialie Shen 0001, Ping Li 0021
IEEE Trans. Multim.5
2020 Collaborative Graph Convolutional Networks: Unsupervised Learning Meets Semi-Supervised Learning
abstract
Graph convolutional networks (GCN) have achieved promising performance in attributed graph clustering and semi-supervised node classification because it is capable of modeling complex graphical structure, and jointly learning both features and relations of nodes. Inspired by the success of unsupervised learning in the training of deep models, we wonder whether graph-based unsupervised learning can collaboratively boost the performance of semi-supervised learning. In this paper, we propose a multi-task graph learning model, called collaborative graph convolutional networks (CGCN). CGCN is composed of an attributed graph clustering network and a semi-supervised node classification network. As Gaussian mixture models can effectively discover the inherent complex data distributions, a new end to end attributed graph clustering network is designed by combining variational graph auto-encoder with Gaussian mixture models (GMM-VGAE) rather than the classic k-means. If the pseudo-label of an unlabeled sample assigned by GMM-VGAE is consistent with the prediction of the semi-supervised GCN, it is selected to further boost the performance of semi-supervised learning with the help of the pseudo-labels. Extensive experiments on benchmark graph datasets validate the superiority of our proposed GMM-VGAE compared with the state-of-the-art attributed graph clustering networks. The performance of node classification is greatly improved by our proposed CGCN, which verifies graph-based unsupervised learning can be well exploited to enhance the performance of semi-supervised learning.
Binyuan Hui, Pengfei Zhu 0001, Qinghua Hu
AAAI2
2020 RGB-T Crowd Counting from Drone: A Benchmark and MMCCN Network
Qing Li 0018, Pengfei Zhu 0001
ACCV (6)3
2020 ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks
abstract
Recently, channel attention mechanism has demonstrated to offer great potential in improving the performance of deep convolutional neural networks (CNNs). However, most existing methods dedicate to developing more sophisticated attention modules for achieving better performance, which inevitably increase model complexity. To overcome the paradox of performance and complexity trade-off, this paper proposes an Efficient Channel Attention (ECA) module, which only involves a handful of parameters while bringing clear performance gain. By dissecting the channel attention module in SENet, we empirically show avoiding dimensionality reduction is important for learning channel attention, and appropriate cross-channel interaction can preserve performance while significantly decreasing model complexity. Therefore, we propose a local cross-channel interaction strategy without dimensionality reduction, which can be efficiently implemented via 1D convolution. Furthermore, we develop a method to adaptively select kernel size of 1D convolution, determining coverage of local cross-channel interaction. The proposed ECA module is both efficient and effective, e.g., the parameters and computations of our modules against backbone of ResNet50 are 80 vs. 24.37M and 4.7e-4 GFlops vs. 3.86 GFlops, respectively, and the performance boost is more than 2% in terms of Top-1 accuracy. We extensively evaluate our ECA module on image classification, object detection and instance segmentation with backbones of ResNets and MobileNetV2. The experimental results show our module is more efficient while performing favorably against its counterparts.
Qilong Wang 0001, Banggu Wu, Pengfei Zhu 0001, Peihua Li, Wangmeng Zuo, Qinghua Hu
CVPR3
2020 Spatial Attention Pyramid Network for Unsupervised Domain Adaptation
Dawei Du, Libo Zhang 0001, Longyin Wen, Tiejian Luo, Pengfei Zhu 0001
ECCV (13)7
2020 SPL-MLL: Selecting Predictable Landmarks for Multi-label Learning
Junbing Li, Changqing Zhang 0002, Pengfei Zhu 0001, Baoyuan Wu, Lei Chen 0011, Qinghua Hu
ECCV (9)3
2020 Multi-label Quadruplet Dictionary Learning
Jiayu Zheng, Wencheng Zhu, Pengfei Zhu 0001
ICANN (2)3
2020 Bilateral Recurrent Network for Single Image Deraining
abstract
Single image deraining has been widely studied in recent years. Motivated by residual learning, most deep learning based deraining approaches devote research attention to extracting rain streaks, usually yielding visual artifacts in final deraining images. To address this issue, we in this paper propose bilateral recurrent network (BRN) to simultaneously exploit rain streak layer and background image layer. Generally, we employ dual residual networks (ResNet) that are recursively unfolded to sequentially extract rain streaks and predict clean background image. Furthermore, we propose bilateral LSTMs into dual ResNets, which not only can respectively propagate deep features across multiple stages, but also bring the interplay between rain streak layer and background image layer. The experimental results demonstrate that our BRN notably outperforms state-of-the-art deep deraining networks on both synthetic datasets and real rainy images. All the source code and pre-trained models are available at https://github.com/shangwei5/BRN.
Wei Shang 0001, Pengfei Zhu 0001, Dongwei Ren
ICASSP2
2020 Progressive Point To Set Metric Learning For Semi-Supervised Few-Shot Classification
abstract
Few-shot learning aims to learn models that can generalize to unseen tasks from very few annotated samples of available tasks. The performance of few-shot learning is greatly affected by the number of samples per class. The massive unlabeled data can help to boost the performance of few shot learning models. In this paper, we propose a novel progressive point to set metric learning (PPSML) model for semisupervised few-shot classification. The distance metric is defined for an image of the query set to a class of the support set by point to set distance. A self-training strategy is designed to select the samples locally or globally with high confidence and use these samples to progressively update the point to set distance. Experiments on benchmark datasets show that our proposed PPSML significantly improves the accuracy of few shot classification and outperforms the state-of-the-art semisupervised few-shot learning methods.
Pengfei Zhu 0001, Mingqi Gu, Changqing Zhang 0002, Qinghua Hu
ICIP1
2020 Learning from Web Data: Improving Crowd Counting via Semi-Supervised Learning
abstract
Deep neural networks have been widely used in crowd counting that aims to give the number of objects in images and videos. The performance of crowd counting models is greatly affected by the size of datasets with high quality annotations. However, collecting and annotating large-scale crowd counting dataset is labor-intensive and time-consuming. In this work, we exploit unlabeled web images to boost the performance of crowd counting models in a semi-supervised manner. Based on the observation that the rotation and splitting operations will not change the number of object in image, we design three auxiliary tasks to improve the feature representation ability of deep models. A semi-supervised multi-task learning framework is proposed by introducing three auxiliary tasks with respect to unlabeled data and our framework can be easily extended to other crowd counting models. An unlabeled dataset (Web-Crowd) with 8679 web images are collected for semi-supervised crowd counting. Experiments shows that our semi-supervised multi-task learning framework can effectively boost the performance of crowd counting models on UCF-QNRF dataset and ShanghaiTech dataset.
Pengfei Zhu 0001
ICPR4
2020 Semi-supervised Learning to Remove Fences from a Single Image
Wei Shang 0001, Pengfei Zhu 0001, Dongwei Ren
PRCV (1)2
2020 Human-in-the-loop image segmentation and annotation
Lianjie Wang, Jin Xie 0001, Pengfei Zhu 0001
Sci. China Inf. Sci.4
2020 Multi-view predictive latent space learning
Jirui Yuan, Pengfei Zhu 0001, Karen Egiazarian
Pattern Recognit. Lett.3
2020 Hybrid Noise-Oriented Multilabel Learning
abstract
For real-world applications, multilabel learning usually suffers from unsatisfactory training data. Typically, features may be corrupted or class labels may be noisy or both. Ignoring noise in the learning process tends to result in an unreasonable model and, thus, inaccurate prediction. Most existing methods only consider either feature noise or label noise in multilabel learning. In this paper, we propose a unified robust multilabel learning framework for data with hybrid noise, that is, both feature noise and label noise. The proposed method, hybrid noise-oriented multilabel learning (HNOML), is simple but rather robust for noisy data. HNOML simultaneously addresses feature and label noise by bi-sparsity regularization bridged with label enrichment. Specifically, the label enrichment matrix explores the underlying correlation among different classes which improves the noisy labeling. Bridged with the enriching label matrix, the structured sparsity is imposed to jointly handle the corrupted features and noisy labeling. We utilize the alternating direction method (ADM) to efficiently solve our problem. Experimental results on several benchmark datasets demonstrate the advantages of our method over the state-of-the-art ones.
Changqing Zhang 0002, Ziwei Yu, Huazhu Fu, Pengfei Zhu 0001, Lei Chen 0011, Qinghua Hu
IEEE Trans. Cybern.4
2020 Deep Fuzzy Tree for Large-Scale Hierarchical Visual Classification
abstract
Deep learning models often use a flat softmax layer to classify samples after feature extraction in visual classification tasks. However, it is hard to make a single decision of finding the true label from massive classes. In this scenario, hierarchical classification is proved to be an effective solution and can be utilized to replace the softmax layer. A key issue of hierarchical classification is to construct a good label structure, which is very significant for classification performance. Several works have been proposed to address the issue, but they have some limitations and are almost designed heuristically. In this article, inspired by fuzzy rough set theory, we propose a deep fuzzy tree model which learns a better tree structure and classifiers for hierarchical classification with theory guarantee. Experimental results show the effectiveness and efficiency of the proposed model in various visual classification datasets.
Yu Wang 0106, Qinghua Hu, Pengfei Zhu 0001, Linhao Li, Bingxu Lu, Jonathan M. Garibaldi, Xianling Li
IEEE Trans. Fuzzy Syst.3
2020 Single Image Deraining Using Bilateral Recurrent Network
abstract
Single image deraining has received considerable progress based on deep convolutional neural network (CNN). In existing deep deraining methods, CNNs are deployed to extract rain streaks while failing in learning direct mapping from rainy image to clean background image, and their architectures become more and more complicated. In this work, we first propose a single recurrent network (SRN) by recursively unfolding a shallow residual network, where a recurrent layer is adopted to propagate deep features across multiple stages. This simple SRN is effective not only in learning residual mapping for extracting rain streaks, but also in learning direct mapping for predicting clean background image. Furthermore, two SRNs are coupled to simultaneously exploit rain streak layer and clean background image layer. Instead of naive combination, we propose bilateral LSTMs, which not only can respectively propagate deep features of rain streak layer and background image layer across stages, but also bring the interplay between these two SRNs, finally forming bilateral recurrent network (BRN). The experimental results demonstrate that our BRN notably outperforms state-of-the-art deep deraining networks on synthetic datasets quantitatively and qualitatively. The proposed methods also perform more favorably in terms of generalization performance on real-world rainy dataset. All the source code and pre-trained models are available at https://github.com/csdwren/RecDerain.
Dongwei Ren, Wei Shang 0001, Pengfei Zhu 0001, Qinghua Hu, Deyu Meng, Wangmeng Zuo
IEEE Trans. Image Process.3
2019 Progressive Image Deraining Networks: A Better and Simpler Baseline
abstract
Along with the deraining performance improvement of deep networks, their structures and learning become more and more complicated and diverse, making it difficult to analyze the contribution of various network modules when developing new deraining networks. To handle this issue, this paper provides a better and simpler baseline deraining network by considering network architecture, input and output, and loss functions. Specifically, by repeatedly unfolding a shallow ResNet, progressive ResNet (PRN) is proposed to take advantage of recursive computation. A recurrent layer is further introduced to exploit the dependencies of deep features across stages, forming our progressive recurrent network (PReNet). Furthermore, intra-stage recursive computation of ResNet can be adopted in PRN and PReNet to notably reduce network parameters with unsubstantial degradation in deraining performance. For network input and output, we take both stage-wise result and original rainy image as input to each ResNet and finally output the prediction of residual image. As for loss functions, single MSE or negative SSIM losses are sufficient to train PRN and PReNet. Experiments show that PRN and PReNet perform favorably on both synthetic and real rainy images. Considering its simplicity, efficiency and effectiveness, our models are expected to serve as a suitable baseline in future deraining research. The source codes are available at https://github.com/csdwren/PReNet.
Dongwei Ren, Wangmeng Zuo, Qinghua Hu, Pengfei Zhu 0001, Deyu Meng
CVPR4
2019 Deep Global Generalized Gaussian Networks
abstract
Recently, global covariance pooling (GCP) has shown great advance in improving classification performance of deep convolutional neural networks (CNNs). However, existing deep GCP networks compute covariance pooling of convolutional activations with assumption that activations are sampled from Gaussian distributions, which may not hold in practice and fails to fully characterize the statistics of activations. To handle this issue, this paper proposes a novel deep global generalized Gaussian network (3G-Net), whose core is to estimate a global covariance of generalized Gaussian for modeling the last convolutional activations. Compared with GCP in Gaussian setting, our 3G-Net assumes the distribution of activations follows a generalized Gaussian, which can capture more precise characteristics of activations. However, there exists no analytic solution for parameter estimation of generalized Gaussian, making our 3G-Net challenging. To this end, we first present a novel regularized maximum likelihood estimator for robust estimating covariance of generalized Gaussian, which can be optimized by a modified iterative re-weighted method. Then, to efficiently estimate the covariance of generaized Gaussian under deep CNN architectures, we approximate this re-weighted method by developing an unrolling re-weighted module and a square root covariance layer. In this way, 3GNet can be flexibly trained in an end-to-end manner. The experiments are conducted on large-scale ImageNet-1K and Places365 datasets, and the results demonstrate our 3G-Net outperforms its counterparts while achieving very competitive performance to state-of-the-arts.
Qilong Wang 0001, Peihua Li, Qinghua Hu, Pengfei Zhu 0001, Wangmeng Zuo
CVPR4
2019 Joint Metric Learning on Riemannian Manifold of Global Gaussian Distributions
Qinqin Nie, Pengfei Zhu 0001, Qinghua Hu, Hao Cheng 0010
ICANN (2)3
2019 Adaptive Graph Fusion for Unsupervised Feature Selection
Sijia Niu, Pengfei Zhu 0001, Qinghua Hu
ICANN (2)2
2019 Multi-task Sparse Regression Metric Learning for Heterogeneous Classification
Haotian Wu 0006, Pengfei Zhu 0001, Qinghua Hu
ICANN (2)3
2019 Flexible Multi-View Representation Learning for Subspace Clustering
abstract
In recent years, numerous multi-view subspace clustering methods have been proposed to exploit the complementary information from multiple views. Most of them perform data reconstruction within each single view, which makes the subspace representation unpromising and thus can not well identify the underlying relationships among data. In this paper, we propose to conduct subspace clustering based on Flexible Multi-view Representation (FMR) learning, which avoids using partial information for data reconstruction. The latent representation is flexibly constructed by enforcing it to be close to different views, which implicitly makes it more comprehensive and well-adapted to subspace clustering. With the introduction of kernel dependence measure, the latent representation can flexibly encode complementary information from different views and explore nonlinear, high-order correlations among these views. We employ the Alternating Direction Minimization (ADM) method to solve our problem. Empirical studies on real-world datasets show that our method achieves superior clustering performance over other state-of-the-art methods.
Ruihuang Li, Changqing Zhang 0002, Qinghua Hu, Pengfei Zhu 0001, Zheng Wang 0008
IJCAI4
2019 Spatio-temporal Active Learning for Visual Tracking
abstract
The success of state-of-the-art deep learning based trackers is fuelled by the large-scale datasets. However, not all training data has a gain on model performance, and some data can even degrade the performance of the model. Therefore, we aim to train state-of-the-art trackers using a part of labeled frame with high information and less data. To this end, we propose a novel framework called STAL(spatio-temporal active learning strategy), which is integrated with an efficient deep tracker and a spatio-temporal active learning strategy. Specifically, we first mine the most informative frames to boost the deep tracker based on the corresponding spatio-temporal response scores of the target. Then, the hard and easy frames are labeled from annotation and machine auto-annotation, respectively. Hard and easy sample pairs are generated from selected frames. To alleviate the impact of sample pairs with large loss, a self-paced fully convolutional Siamese network is proposed by introducing a norm negative regularization. The STAL framework will converge well and output promising tracking performance on several publicly available datasets.
Chenfeng Liu, Pengfei Zhu 0001, Qinghua Hu
IJCNN2
2019 Coupled Dictionary Learning for Multi-label Embedding
abstract
With the booming of social networks, such as Facebook and Flickr, the candidate labels of an instance can be numerous. Hence, traditional multi-label learning algorithms are out of capability to handle a large quantity of labels for the unaffordable time complexity. To alleviate this problem, label space dimension reduction (LSDR) is proposed by transforming the original label space into a lower dimensional one. Inspired by the effectiveness of coupled dictionary learning (CDL) in dealing with cross-modal data, in this paper, we proposed a novel algorithm named Coupled Dictionary Learning for Multi-label Embedding (ML-CDL) to track the problem of LSDR. We novelly treat feature and label as coupled domains. Then CDL is utilized to generate the low-dimensional latent space that leverages the information between feature and label spaces. In particular, the sparse representation coefficients embody the properties of interpretability, discriminability and sparsity. Experimental results on benchmark datasets demonstrate the effectiveness of our algorithm.
Sijia Niu, Pengfei Zhu 0001, Qinghua Hu
IJCNN3
2019 Self-paced Robust Deep Face Recognition with Label Noise
Pengfei Zhu 0001, Wenya Ma, Qinghua Hu
PAKDD (3)1
2019 Fuzzy Rough Set Based Feature Selection for Large-Scale Hierarchical Classification
abstract
The classification of high-dimensional tasks remains a significant challenge for machine learning algorithms. Feature selection is considered to be an indispensable preprocessing step in high-dimensional data classification. In the era of big data, there may be hundreds of class labels, and the hierarchical structure of the classes is often available. This structure is helpful in feature selection and classifier training. However, most current techniques do not consider the hierarchical structure. In this paper, we design a feature selection strategy for hierarchical classification based on fuzzy rough sets. First, a fuzzy rough set model for hierarchical structures is developed to compute the lower and upper approximations of classes organized with a class hierarchy. This model is distinguished from existing techniques by the hierarchical class structure. A hierarchical feature selection problem is then defined based on the model. The new model is more practical than existing feature selection approaches, as many real-world tasks are naturally cast in terms of hierarchical classification. A feature selection algorithm based on sibling nodes is proposed, and this is shown to be more efficient and more versatile than flat feature selection. Compared with the flat feature selection algorithm, the computational load of the proposed algorithm is reduced from 98.0% to 6.5%, while the classification performance is improved on the SAIAPR dataset. The related experiments also demonstrate the effectiveness of the hierarchical algorithm.
Hong Zhao 0002, Ping Wang 0072, Qinghua Hu, Pengfei Zhu 0001
IEEE Trans. Fuzzy Syst.4
2019 One-Step Multi-View Spectral Clustering
abstract
Previous multi-view spectral clustering methods are a two-step strategy, which first learns a fixed common representation (or common affinity matrix) of all the views from original data and then conducts k-means clustering on the resulting common affinity matrix. The two-step strategy is not able to output reasonable clustering performance since the goal of the first step (i.e., the common affinity matrix learning) is not designed for achieving the optimal clustering result. Moreover, the two-step strategy learns the common affinity matrix from original data, which often contain noise and redundancy to influence the quality of the common affinity matrix. To address these issues, in this paper, we design a novel One-step Multi-view Spectral Clustering (OMSC) method to output the common affinity matrix as the final clustering result. In the proposed method, the goal of the common affinity matrix learning is designed to achieving optimal clustering result and the common affinity matrix is learned from low-dimensional data where the noise and redundancy of original high-dimensional data have been removed. We further propose an iterative optimization method to fast solve the proposed objective function. Experimental results on both synthetic datasets and public datasets validated the effectiveness of our proposed method, comparing to the state-of-the-art methods for multi-view clustering.
Xiaofeng Zhu 0001, Shichao Zhang 0001, Wei He 0017, Rongyao Hu, Cong Lei, Pengfei Zhu 0001
IEEE Trans. Knowl. Data Eng.6
2018 Latent Semantic Aware Multi-View Multi-Label Classification
abstract
For real-world applications, data are often associated with multiple labels and represented with multiple views. Most existing multi-label learning methods do not sufficiently consider the complementary information among multiple views, leading to unsatisfying performance. To address this issue, we propose a novel approach for multi-view multi-label learning based on matrix factorization to exploit complementarity among different views. Specifically, under the assumption that there exists a common representation across different views, the uncovered latent patterns are enforced to be aligned across different views in kernel spaces. In this way, the latent semantic patterns underlying in data could be well uncovered and this enhances the reasonability of the common representation of multiple views. As a result, the consensus multi-view representation is obtained which encodes the complementarity and consistence of different views in latent semantic space. We provide theoretical guarantee for the strict convexity for our method by properly setting parameters. Empirical evidence shows the clear advantages of our method over the state-of-the-art ones.
Changqing Zhang 0002, Ziwei Yu, Qinghua Hu, Pengfei Zhu 0001, Xinwang Liu 0002, Xiaobo Wang 0001
AAAI4
2018 Generalized Multi-view Unsupervised Feature Selection
Yue Liu 0008, Changqing Zhang 0002, Pengfei Zhu 0001, Qinghua Hu
ICANN (2)3
2018 Support Vector Metric Learning on Symmetric Positive Definite Manifold
abstract
The manifold of symmetric positive definite (SPD) matrices has drawn significant attention because of its widespread applications. SPD matrices provide compact nonlinear representations of data and form a special type of Riemannian manifold. The direct application of support vector machines on SPD manifold maybe fails due to lack of samples per class. In this paper, we propose a support vector metric learning (SVML) model on SPD manifold. We define a positive definite kernel for point pairs on SPD manifold and transform metric learning on SPD manifold to a point pair classification problem. The metric learning problem can be efficiently solved by standard support vector machines. Compared with classifying points on SPD manifold by support vector machines directly, SVML effectively learns a distance metric for SPD matrices by training a binary support vector machine model. Experiments on video based face recognition, image set classification, and material classification show that SVML outperforms the state-of-the-art metric learning algorithms on SPD manifold.
Hao Cheng 0010, Pengfei Zhu 0001, Qilong Wang 0001, Changqing Zhang 0002, Qinghua Hu
ICME2
2018 FISH-MML: Fisher-HSIC Multi-View Metric Learning
abstract
This work presents a simple yet effective model for multi-view metric learning, which aims to improve the classification of data with multiple views, e.g., multiple modalities or multiple types of features. The intrinsic correlation, different views describing same set of instances, makes it possible and necessary to jointly learn multiple metrics of different views, accordingly, we propose a multi-view metric learning method based on Fisher discriminant analysis (FDA) and Hilbert-Schmidt Independence Criteria (HSIC), termed as Fisher-HSIC Multi-View Metric Learning (FISH-MML). In our approach, the class separability is enforced in the spirit of FDA within each single view, while the consistence among different views is enhanced based on HSIC. Accordingly, both intra-view class separability and inter-view correlation are well addressed in a unified framework. The learned metrics can improve multi-view classification, and experimental results on real-world datasets demonstrate the effectiveness of the proposed method.
Changqing Zhang 0002, Yeqing Liu, Yue Liu 0008, Qinghua Hu, Xinwang Liu 0002, Pengfei Zhu 0001
IJCAI6
2018 Towards Generalized and Efficient Metric Learning on Riemannian Manifold
abstract
Modeling data as points on non-linear Riemannian manifold has attracted increasing attentions in many computer vision tasks, especially visual recognition. Learning an appropriate metric on Riemannian manifold plays a key role in achieving promising performance. For widely used symmetric positive definite (SPD) manifold and Grassmann manifold, most of existing metric learning methods are designed for one manifold, and are not straightforward for the other one. Furthermore, optimizations in previous methods usually rely on computationally expensive iterations. To address above limitations, this paper makes an attempt to propose a generalized and efficient Riemannian manifold metric learning (RMML) method, which can be flexibly adopted to both SPD and Grassmann manifolds. By minimizing the geodesic distance of similar pairs and the interpoint geodesic distance of dissimilar ones on nonlinear manifolds, the proposed RMML is optimized by computing the geodesic mean between inverse of similarity matrix and dissimilarity matrix, benefiting a global closed-form solution and high efficiency. The experiments are conducted on various visual recognition tasks, and the results demonstrate our RMML performs favorably against its counterparts in terms of both accuracy and efficiency.
Pengfei Zhu 0001, Hao Cheng 0010, Qinghua Hu, Qilong Wang 0001, Changqing Zhang 0002
IJCAI1
2018 Beyond Similar and Dissimilar Relations : A Kernel Regression Formulation for Metric Learning
abstract
Most existing metric learning methods focus on learning a similarity or distance measure relying on similar and dissimilar relations between sample pairs. However, pairs of samples cannot be simply identified as similar or dissimilar in many real-world applications, e.g., multi-label learning, label distribution learning or tasks with continuous decision values. To this end, in this paper we propose a novel relation alignment metric learning (RAML) formulation to handle the metric learning problem in those scenarios. Since the relation of two samples can be measured by the difference degree of the decision values, motivated by the consistency of the sample relations in the feature space and decision space, our proposed RAML utilizes the sample relations in the decision space to guide the metric learning in the feature space. Specifically, our RAML method formulates metric learning as a kernel regression problem, which can be efficiently optimized by the standard regression solvers. We carry out several experiments on the single-label classification, multi-label classification, and label distribution learning tasks, to demonstrate that our method achieves favorable performance against the state-of-the-art methods.
Pengfei Zhu 0001, Ren Qi, Qinghua Hu, Qilong Wang 0001, Changqing Zhang 0002, Liu Yang 0010
IJCAI1
2018 Latent Subspace Representation for Multiclass Classification
Changqing Zhang 0002, Xiao Wang 0017, Pengfei Zhu 0001, Zheng Wang 0008, Qinghua Hu
PRICAI (1)4
2018 Co-regularized unsupervised feature selection
Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002
Neurocomputing1
2018 Efficient multi-modal geometric mean metric learning
Jianqing Liang, Qinghua Hu, Pengfei Zhu 0001, Wenwu Wang 0001
Pattern Recognit.3
2018 Multi-view label embedding
Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002, Zhizhao Feng
Pattern Recognit.1
2018 Multi-label feature selection with missing labels
Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002, Hong Zhao 0002
Pattern Recognit.1
2017 Latent Multi-view Subspace Clustering
abstract
In this paper, we propose a novel Latent Multi-view Subspace Clustering (LMSC) method, which clusters data points with latent representation and simultaneously explores underlying complementary information from multiple views. Unlike most existing single view subspace clustering methods that reconstruct data points using original features, our method seeks the underlying latent representation and simultaneously performs data reconstruction based on the learned latent representation. With the complementarity of multiple views, the latent representation could depict data themselves more comprehensively than each single view individually, accordingly makes subspace representation more accurate and robust as well. The proposed method is intuitive and can be optimized efficiently by using the Augmented Lagrangian Multiplier with Alternating Direction Minimization (ALM-ADM) algorithm. Extensive experiments on benchmark datasets have validated the effectiveness of our proposed method.
Changqing Zhang 0002, Qinghua Hu, Huazhu Fu, Pengfei Zhu 0001, Xiaochun Cao
CVPR4
2017 Semi-Supervised Multi-view Multi-label Classification Based on Nonnegative Matrix Factorization
Guangxia Wang, Changqing Zhang 0002, Pengfei Zhu 0001, Qinghua Hu
ICANN (2)3
2017 Unsupervised feature selection by manifold regularized self-representation
abstract
Unsupervised feature selection has been proven to be an efficient technique in mitigating the curse of dimensionality. It helps to understand and analyze the prevalent high-dimensional unlabeled data. Recently, the self-similarity property of objects, which assumes that a feature can be represented by the linear combination of its relevant features, has been successfully used in unsupervised feature selection. However, it does not take the geometry structure of the sample space into consideration. In this paper, we propose a novel algorithm termed manifold regularized self-representation(MRSR). To preserve the local spatial structure, we incorporate an effective manifold regularization into the objective function. An iterative reweighted least square (IRLS) algorithm is developed to solve the optimization problem and the convergence is proved. Extensive experimental results on several benchmark datasets validate the effectiveness of the proposed method.
Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002
ICIP3
2017 Mixed sparsity regularized multi-view unsupervised feature selection
abstract
The traditional learning machines suffer from the curse of dimensionality because of the data explosion in the areas of multi-media, social network, etc. Feature selection is an effective technique to reduce storage burden and time complexity, and improve generalization ability of the learned models. In real-world applications, the data can be collected from different modalities, or described from multi-views as well. Compared with supervised cases, it is more challenging to reduce the feature dimensionality of multi-view data in unsupervised circumstances. The key difficulty with multiview unsupervised feature selection is how to characterize the multi-view relationships. In this paper, we propose a novel method for multi-view unsupervised feature selection by imposing sparsity on both individual features and views. To exploit the complementary information, we also take the view importance into consideration without introducing explicit view weights. Experiments on benchmark datasets show on the proposed algorithm outperforms other unsupervised feature selection methods.
Kennedy W. Wangila, Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002
ICIP3
2017 Independence regularized multi-label ensemble
abstract
In this paper, we focus on promoting multi-label learning task with ensemble learning. Compared to traditional single algorithm methods, it has been recognized that ensemble methods could achieve much better performance than each constituent learned model, especially under the conditional independence of different classifiers. Existing multi-label ensemble algorithms mainly focus on creating diverse component learners by employing different mechanisms, mostly using randomization strategies by smart heuristics. Different from most existing methods, in this paper, we propose an ensemble method to learn the basic classifiers which considers the general independence of the different classifiers. Therefore, each learned multi-label classifier is guaranteed to be diverse and complementary. Furthermore, considering the different qualities of these classifiers, a weight vector is learned to balance these classifiers. Experiments on several benchmark datasets well demonstrate that the proposed method outperforms the state-of-the-art methods.
Ziwei Yu, Changqing Zhang 0002, Qinghua Hu, Pengfei Zhu 0001
ICME4
2017 Multi-view Label Space Dimension Reduction
Pengfei Zhu 0001, Changqing Zhang 0002, Qinghua Hu
ICONIP (1)2
2017 Robust Deep Face Recognition with Label Noise
Jirui Yuan, Wenya Ma, Pengfei Zhu 0001, Karen Egiazarian
ICONIP (2)3
2017 Hierarchical Feature Selection with Recursive Regularization
abstract
In the big data era, the sizes of datasets have increased dramatically in terms of the number of samples, features, and classes. In particular, there exists usually a hierarchical structure among the classes. This kind of task is called hierarchical classification. Various algorithms have been developed to select informative features for flat classification. However, these algorithms ignore the semantic hyponymy in the directory of hierarchical classes, and select a uniform subset of the features for all classes. In this paper, we propose a new technique for hierarchical feature selection based on recursive regularization. This algorithm takes the hierarchical information of the class structure into account. As opposed to flat feature selection, we select different feature subsets for each node in a hierarchical tree structure using the parent-children relationships and the sibling relationships for hierarchical regularization. By imposing $\ell_{2,1}$-norm regularization to different parts of the hierarchical classes, we can learn a sparse matrix for the feature ranking of each node. Extensive experiments on public datasets demonstrate the effectiveness of the proposed algorithm.
Hong Zhao 0002, Pengfei Zhu 0001, Ping Wang 0072, Qinghua Hu
IJCAI2
2017 Non-convex regularized self-representation for unsupervised feature selection
Pengfei Zhu 0001, Wencheng Zhu, Weizhi Wang, Wangmeng Zuo, Qinghua Hu
Image Vis. Comput.1
2017 Subspace clustering guided unsupervised feature selection
Pengfei Zhu 0001, Wencheng Zhu, Qinghua Hu, Changqing Zhang 0002, Wangmeng Zuo
Pattern Recognit.1
2017 Flexible Multi-View Dimensionality Co-Reduction
abstract
Dimensionality reduction aims to map the high-dimensional inputs onto a low-dimensional subspace, in which the similar points are close to each other and vice versa. In this paper, we focus on unsupervised dimensionality reduction for the data with multiple views, and propose a novel method, called Multi-view Dimensionality co-Reduction. Our method flexibly exploits the complementarity of multiple views during the dimensionality reduction and respects the similarity relationships between data points across these different views. The kernel matching constraint based on Hilbert-Schmidt Independence Criterion enhances the correlations and penalizes the disagreement of different views. Specifically, our method explores the correlations within each view independently, and maximizes the dependence among different views with kernel matching jointly. Thus, the locality within each view and the consistence between different views are guaranteed in the subspaces corresponding to different views. More importantly, benefiting from the kernel matching, our method need not depend on a common low-dimensional subspace, which is critical to reduce the influence of the unbalanced dimensionalities of multiple views. Specifically, our method explicitly produces individual low-dimensional projections for individual views, which could be applied for new coming data in the out-of-sample manner. Experiments on both clustering and recognition tasks demonstrate the advantages of the proposed method over the state-of-the-art approaches.
Changqing Zhang 0002, Huazhu Fu, Qinghua Hu, Pengfei Zhu 0001, Xiaochun Cao
IEEE Trans. Image Process.4
2016 Coupled Dictionary Learning for Unsupervised Feature Selection
abstract
Unsupervised feature selection (UFS) aims to reduce the time complexity and storage burden, as well as improve the generalization performance. Most existing methods convert UFS to supervised learning problem by generating labels with specific techniques (e.g., spectral analysis, matrix factorization and linear predictor). Instead, we proposed a novel coupled analysis-synthesis dictionary learning method, which is free of generating labels. The representation coefficients are used to model the cluster structure and data distribution. Specifically, the synthesis dictionary is used to reconstruct samples, while the analysis dictionary analytically codes the samples and assigns probabilities to the samples. Afterwards, the analysis dictionary is used to select features that can well preserve the data distribution. The effective L2p-norm (0 < p <1) regularization is imposed on the analysis dictionary to get much sparse solution and is more effective in feature selection.We proposed an iterative reweighted least squares algorithm to solve the L2p-norm optimization problem and proved it can converge to a fixed point. Experiments on benchmark datasets validated the effectiveness of the proposed method
Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002, Wangmeng Zuo
AAAI1
2016 A Self-Representation Induced Classifier
Pengfei Zhu 0001, Lei Zhang 0006, Wangmeng Zuo, Xiangchu Feng, Qinghua Hu
IJCAI1
2016 Multi-view Representative and Informative Induced Active Learning
Huaxi Huang, Changqing Zhang 0002, Qinghua Hu, Pengfei Zhu 0001
PRICAI4
2016 Set to Set Visual Tracking
Wencheng Zhu, Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002
PRICAI2
2016 Combining neighborhood separable subspaces for classification via sparsity regularized optimization
Pengfei Zhu 0001, Qinghua Hu, Yahong Han, Changqing Zhang 0002
Inf. Sci.1
2016 Data-Distribution-Aware Fuzzy Rough Set Model and its Application to Robust Classification
abstract
Fuzzy rough sets (FRSs) are considered to be a powerful model for analyzing uncertainty in data. This model encapsulates two types of uncertainty: 1) fuzziness coming from the vagueness in human concept formation and 2) roughness rooted in the granulation coming with human cognition. The rough set theory has been widely applied to feature selection, attribute reduction, and classification. However, it is reported that the classical FRS model is sensitive to noisy information. To address this problem, several robust models have been developed in recent years. Nevertheless, these models do not consider a statistical distribution of data, which is an important type of uncertainty. Data distribution serves as crucial information for designing an optimal classification or regression model. Thus, we propose a data-distribution-aware FRS model that considers distribution information and incorporates it in computing lower and upper fuzzy approximations. The proposed model considers not only the similarity between samples, but also the probability density of classes. In order to demonstrate the effectiveness of the proposed model, we design a new sample evaluation index for prototype-based classification based on the model, and a prototype selection algorithm is developed using this index. Furthermore, a robust classification algorithm is constructed with prototype covering and nearest neighbor classification. Experimental results confirm the robustness and effectiveness of the proposed model.
Shuang An, Qinghua Hu, Witold Pedrycz, Pengfei Zhu 0001, Eric C. C. Tsang
IEEE Trans. Cybern.4
2015 Joint representation and pattern learning for robust face recognition
Meng Yang 0001, Pengfei Zhu 0001, Feng Liu 0013, LinLin Shen
Neurocomputing2
2015 Unsupervised feature selection by regularized self-representation
Pengfei Zhu 0001, Wangmeng Zuo, Lei Zhang 0006, Qinghua Hu, Simon C. K. Shiu
Pattern Recognit.1
2014 Local Generic Representation for Face Recognition with Single Sample per Person
Pengfei Zhu 0001, Meng Yang 0001, Lei Zhang 0006, Il-Yong Lee
ACCV (3)1
2014 Towards a scalable resource-driven approach for detecting repackaged Android applications
abstract
Repackaged Android applications (or simply apps) are one of the major sources of mobile malware and also an important cause of severe revenue loss to app developers. Although a number of solutions have been proposed to detect repackaged apps, the majority of them heavily rely on code analysis, thus suffering from two limitations: (1) poor scalability due to the billion opcode problem; (2) unreliability to code obfuscation/app hardening techniques. In this paper, we explore an alternative approach that exploits core resources, which have close relationships with codes, to detect repackaged apps. More precisely, we define new features for characterizing apps, investigate two kinds of algorithms for searching similar apps, and propose a two-stage methodology to speed up the detection. We realize our approach in a system named ResDroid and conduct large scale evaluation on it. The results show that ResDroid can identify repackaged apps efficiently and effectively even if they are protected by obfuscation or hardening systems.
Yuru Shao, Xiapu Luo, Chenxiong Qian, Pengfei Zhu 0001, Lei Zhang 0006
ACSAC4
2014 Multi-granularity distance metric learning via neighborhood granule margin maximization
Pengfei Zhu 0001, Qinghua Hu, Wangmeng Zuo, Meng Yang 0001
Inf. Sci.1
2014 Image Set-Based Collaborative Representation for Face Recognition
abstract
With the rapid development of digital imaging and communication technologies, image set-based face recognition (ISFR) is becoming increasingly important. One key issue of ISFR is how to effectively and efficiently represent the query face image set using the gallery face image sets. The set-to-set distance-based methods ignore the relationship between gallery sets, whereas representing the query set images individually over the gallery sets ignores the correlation between query set images. In this paper, we propose a novel image set-based collaborative representation and classification method for ISFR. By modeling the query set as a convex or regularized hull, we represent this hull collaboratively over all the gallery sets. With the resolved representation coefficients, the distance between the query set and each gallery set can then be calculated for classification. The proposed model naturally and effectively extends the image-based collaborative representation to an image set based one, and our extensive experiments on benchmark ISFR databases show the superiority of the proposed method to state-of-the-art ISFR methods under different set sizes in terms of both recognition rate and efficiency.
Pengfei Zhu 0001, Wangmeng Zuo, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001
IEEE Trans. Inf. Forensics Secur.1
2013 From Point to Set: Extend the Learning of Distance Metrics
abstract
Most of the current metric learning methods are proposed for point-to-point distance (PPD) based classification. In many computer vision tasks, however, we need to measure the point-to-set distance (PSD) and even set-to-set distance (SSD) for classification. In this paper, we extend the PPD based Mahalanobis distance metric learning to PSD and SSD based ones, namely point-to-set distance metric learning (PSDML) and set-to-set distance metric learning (SSDML), and solve them under a unified optimization framework. First, we generate positive and negative sample pairs by computing the PSD and SSD between training samples. Then, we characterize each sample pair by its covariance matrix, and propose a covariance kernel based discriminative function. Finally, we tackle the PSDML and SSDML problems by using standard support vector machine solvers, making the metric learning very efficient for multiclass visual classification tasks. Experiments on gender classification, digit recognition, object categorization and face recognition show that the proposed metric learning methods can effectively enhance the performance of PSD and SSD based classification.
Pengfei Zhu 0001, Lei Zhang 0006, Wangmeng Zuo, David Zhang 0001
ICCV1
2013 Adaptive neighborhood granularity selection and combination based on margin distribution optimization
Pengfei Zhu 0001, Qinghua Hu
Inf. Sci.1
2013 Rule extraction from support vector machines based on consistent region covering reduction
Pengfei Zhu 0001, Qinghua Hu
Knowl. Based Syst.1
2012 Multi-scale Patch Based Collaborative Representation for Face Recognition with Margin Distribution Optimization
Pengfei Zhu 0001, Lei Zhang 0006, Qinghua Hu, Simon C. K. Shiu
ECCV (1)1
2012 Margin distribution based bagging pruning
Zongxia Xie, Yong Xu 0007, Qinghua Hu, Pengfei Zhu 0001
Neurocomputing4
2012 A Novel Algorithm for Finding Reducts With Fuzzy Rough Sets
abstract
Attribute reduction is one of the most meaningful research topics in the existing fuzzy rough sets, and the approach of discernibility matrix is the mathematical foundation of computing reducts. When computing reducts with discernibility matrix, we find that only the minimal elements in a discernibility matrix are sufficient and necessary. This fact motivates our idea in this paper to develop a novel algorithm to find reducts that are based on the minimal elements in the discernibility matrix. Relative discernibility relations of conditional attributes are defined and minimal elements in the fuzzy discernibility matrix are characterized by the relative discernibility relations. Then, the algorithms to compute minimal elements and reducts are developed in the framework of fuzzy rough sets. Experimental comparison shows that the proposed algorithms are effective.
Degang Chen 0002, Lei Zhang 0006, Suyun Zhao, Qinghua Hu, Pengfei Zhu 0001
IEEE Trans. Fuzzy Syst.5
2011 A linear subspace learning approach via sparse coding
abstract
Linear subspace learning (LSL) is a popular approach to image recognition and it aims to reveal the essential features of high dimensional data, e.g., facial images, in a lower dimensional space by linear projection. Most LSL methods compute directly the statistics of original training samples to learn the subspace. However, these methods do not effectively exploit the different contributions of different image components to image recognition. We propose a novel LSL approach by sparse coding and feature grouping. A dictionary is learned from the training dataset, and it is used to sparsely decompose the training samples. The decomposed image components are grouped into a more discriminative part (MDP) and a less discriminative part (LDP). An unsupervised criterion and a supervised criterion are then proposed to learn the desired subspace, where the MDP is preserved and the LDP is suppressed simultaneously. The experimental results on benchmark face image databases validated that the proposed methods outperform many state-of-the-art LSL schemes.
Lei Zhang 0006, Pengfei Zhu 0001, Qinghua Hu, David Zhang 0001
ICCV2
2011 Large-margin nearest neighbor classifiers via sample weight learning
Qinghua Hu, Pengfei Zhu 0001, Yongbin Yang, Daren Yu
Neurocomputing2
2011 Rule learning for classification based on neighborhood covering reduction
Qinghua Hu, Pengfei Zhu 0001, Peijun Ma
Inf. Sci.3
2010 Event Recognition Based on Top-Down Motion Attention
abstract
How to fuse static and dynamic information is a key issue in event analysis. In this paper, a top-down motion guided fusing method is proposed for recognizing events in an unconstrained news video. In the method, the static information is represented as a Bag-of-SIFT-features and motion information is employed to generate event specific attention map to direct the sampling of the interest points. We build class-specific motion histograms for each event so as to give more weight on the interest points that are discriminative to the corresponding event. Experimental results on TRECVID 2005 video corpus demonstrate that the proposed method can improve the mean average accuracy of recognition.
Li Li 0010, Weiming Hu 0004, Bing Li 0001, Chunfeng Yuan, Pengfei Zhu 0001, Wanqing Li 0001
ICPR5
2010 Prototype Learning Using Metric Learning Based Behavior Recognition
abstract
Behavior recognition is an attractive direction in the computer vision domain. In this paper, we propose a novel behavior recognition method based on prototype learning using metric learning. Prototype learning algorithm can improve the classification performance of nearest-neighbor classifier, reduce the storage and computation requirements. And the metric learning algorithm is used to advance the performance of the prototype learning. In this paper, We use a kind of compound feature including local feature and motion feature to recognize human behaviors. The experimental results show the effectiveness of our method.
Pengfei Zhu 0001, Weiming Hu 0004, Chunfeng Yuan, Li Li 0010
ICPR1
2010 Feature Selection via Maximizing Fuzzy Dependency
abstract
Feature selection is an important preprocessing step in pattern analysis and machine learning. The key issue in feature selection is to evaluate quality of candidate features. In this work, we introduce a weighted distance learning algorithm for feature selection via maximizing fuzzy dependency. We maximize fuzzy dependency between features and decision by distance learning and then evaluate the quality of features with the learned weight vector. The features deriving great weights are considered to be useful for classification learning. We test the proposed technique with some classical methods and the experimental results show the proposed algorithm is effective.
Qinghua Hu, Pengfei Zhu 0001, Yongbin Yang, Daren Yu
Fundam. Informaticae2
2009 Segment Model Based Vehicle Motion Analysis
abstract
Motion analysis is a very attractive research direction in computer vision field. In this paper, we propose a framework for analyzing real vehicle motion in visual traffic surveillance by using Segment Model (SM), which is a kind of probabilistic model. SM can grasp the underlying information of observation sequence by using segment distribution. It has been proved to be more precise than that of HMM. In the experiments, we compare our approach with the template matching method based on the Hausdorff distance and the state space method based on the Hidden Markov Model (HMM). The experimental results show the effectiveness of our approach.
Pengfei Zhu 0001, Weiming Hu 0004, Xi Li 0001, Li Li 0010
AVSS1