VLDB 2026 Research / reviewers in the wild / expert
Peihua Li
dblp:80/5257
· DBLP profile ↗
83ranked-venue papers
24as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 16 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 13 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorTheory of computation · 3 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Task-Specific Distance Correlation Matching for Few-Shot Action RecognitionabstractFew-shot action recognition (FSAR) has recently made notable progress through set matching and efficient adaptation of large-scale pre-trained models. However, two key limitations persist. First, existing set matching metrics typically rely on cosine similarity to measure inter-frame linear dependencies and then perform matching with only instance-level information, thus failing to capture more complex patterns such as nonlinear relationships and overlooking task-specific cues. Second, for efficient adaptation of CLIP to FSAR, recent work performing fine-tuning via skip-fusion layers (which we refer to as side layers) has significantly reduced memory cost. However, the newly introduced side layers are often difficult to optimize under limited data conditions. To address these limitations, we propose TS-FSAR, a framework comprising three components: (1) a visual Ladder Side Network (LSN) for efficient CLIP fine-tuning; (2) a metric called Task-Specific Distance Correlation Matching (TS-DCM), which uses alpha-distance correlation to model both linear and nonlinear inter-frame dependencies and leverages a task prototype to enable task-specific matching; and (3) a Guiding LSN with Adapted CLIP (GLAC) module, which regularizes LSN using the adapted frozen CLIP to improve training for better α-distance correlation estimation under limited supervision. Extensive experiments on five widely-used benchmarks demonstrate that our TS-FSAR yields superior performance compared to prior state-of-the-arts. Fei Long 0001, Jiaming Lv, Jiangtao Xie, Peihua Li |
AAAI | 5 |
| 2026 | Parameterized Approximation Algorithms for Dominating Set and Power Dominating Set By Using Leafage
Peihua Li, Jiong Guo |
COCOON | 1 |
| 2026 | GDT-VLM: Global Distribution Modeling for Visual Token Compression in Efficient Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLMs) have achieved remarkable success, yet the massive number of visual tokens per image imposes a heavy inference burden. Existing methods attempt extreme compression with a single visual token via spatial reduction or cross-modal attention, but often overlook the statistical information inherent in visual tokens, leading to a suboptimal efficiency-effectiveness trade-off. In this paper, we show that effective statistical characterization of visual features benefits extreme visual token compression in efficient MLLMs. To this end, we propose GDT-VLM, a novel architecture that exploits global distribution modeling of visual tokens for efficient compression. Specifically, GDT-VLM encodes visual features by jointly modeling their global first-order (GAP) and second-order (Brownian Distance Covariance) statistics, enabling a more expressive yet compact representation. By capturing the holistic characteristics of vision tokens, our GDT-VLM yields compact vision information in a single token while effectively preserving statistical content. Extensive experiments on 7 benchmarks show that our approach achieves competitive accuracy while offering a favorable efficiency-effectiveness trade-off. Jiangtao Xie, Qilong Wang 0001, Peihua Li |
ICMR | 5 |
| 2026 | QHSP-Net: query-aware higher-order statistical pooling network for referring image segmentation
Qiule Sun, Jianxin Zhang 0001, Bingbing Zhang 0001, Peihua Li |
Multim. Syst. | 4 |
| 2026 | IrisMAE: Structure-aware masked image modeling for iris recognition
Shiao Duan, Lingyao Jia, Peihua Li |
Pattern Recognit. | 3 |
| 2025 | TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action RecognitionabstractGoing beyond few-shot action recognition (FSAR), cross-domain FSAR (CDFSAR) has attracted recent research interests by solving the domain gap lying in source-to-target transfer learning. Existing CDFSAR methods mainly focus on joint training of source and target data to mitigate the side effect of domain gap. However, such kind of methods suffer from two limitations: First, pair-wise joint training requires retraining deep models in case of one source data and multiple target ones, which incurs heavy computation cost, especially for large source and small target data. Second, pre-trained models after joint training are adopted to target domain in a straightforward manner, hardly taking full potential of pre-trained models and then limiting recognition performance. To overcome above limitations, this paper proposes a simple yet effective baseline, namely Temporal-Aware Model Tuning (TAMT) for CDFSAR. Specifically, our TAMT involves a decoupled paradigm by performing pre-training on source data and fine-tuning target data, which avoids retraining for multiple target data with single source. To effectively and efficiently explore the potential of pre-trained models in transferring to target domain, our TAMT proposes a Hierarchical Temporal Tuning Network (HTTN), whose core involves local temporal-aware adapters (TAA) and a global temporal-aware moment tuning (GTMT). Particularly, TAA learns few parameters to recalibrate the intermediate features of frozen pre-trained models, enabling efficient adaptation to target domains. Furthermore, GTMT helps to generate powerful video representations, improving match performance on the target domain. Experiments on several widely used video benchmarks show our TAMT outperforms the recently proposed counterparts by 13% ∼31%, achieving new state-of-the-art CDFSAR results. Zilin Gao, Qilong Wang 0001, Zhaofeng Chen, Peihua Li, Qinghua Hu |
CVPR | 5 |
| 2025 | ImagineFSL: Self-Supervised Pretraining Matters on Imagined Base Set for VLM-based Few-shot LearningabstractAdapting CLIP models for few-shot recognition has recently attracted significant attention. Despite considerable progress, these adaptations remain hindered by the pervasive challenge of data scarcity. Text-to-image models, capable of generating abundant photorealistic labeled images, offer a promising solution. However, existing approaches simply treat synthetic images as complements to real images, rather than as standalone knowledge repositories stemming from distinct foundation models. To overcome this limitation, we frame synthetic images as an imagined base set (iBase), i.e., an independent, large-scale synthetic dataset encompassing diverse concepts. Building on this perspective, we introduce ImagineFSL, a novel CLIP adaptation methodology that pretrains on iBase and then fine-tunes for downstream few-shot tasks. We find that, compared to no pretraining, both supervised and self-supervised pretraining are beneficial, with the latter providing better performance. Based on on this finding, we propose an improved self-supervised method tailored for few-shot scenarios, enhancing the transferability of representations from synthetic to real image domains. Additionally, we present a systematic and scalable pipeline that employs chain-of-thought and in-context learning techniques, harnessing foundation models to automatically generate diverse, realistic images. Validated across eleven datasets, our methods consistently outperform state-of-the-art approaches by substantial margins. Haoyuan Yang, Jiaming Lv, Xianjun Cheng, Peihua Li |
CVPR | 6 |
| 2025 | DALIP: Distribution Alignment-Based Language-Image Pre-Training for Domain-Specific Data
Jiangtao Xie, Qilong Wang 0001, Qinghua Hu, Peihua Li |
ICCV | 6 |
| 2025 | BDC-CLIP: Brownian Distance Covariance for Adapting CLIP to Action RecognitionabstractBridging contrastive language-image pre-training (CLIP) to video action recognition has attracted growing interest. Human actions are inherently rich in spatial and temporal contexts, involving dynamic interactions among people, objects, and the environment. Accurately recognizing actions requires effectively capturing these fine-grained elements and modeling their relationships with language. However, most existing methods rely on cosine similarity–practically equivalent to the Pearson correlation coefficient–between global tokens for video-language alignment. As a result, they have limited capacity to model complex dependencies and tend to overlook local tokens that encode critical spatio-temporal cues. To overcome these limitations, we propose BDC-CLIP, a novel framework that leverages Brownian Distance Covariance (BDC) to align visual and textual representations. Our method can capture complex relationships–both linear and nonlinear–between all visual and textual tokens, enabling fine-grained modeling in space, time, and language. BDC-CLIP achieves state-of-the-art performance across zero-shot, few-shot, base-to-novel, and fully supervised action recognition settings, demonstrating its effectiveness and broad applicability. Fei Long 0001, Jiaming Lv, Haoyuan Yang, Xianjun Cheng, Peihua Li |
ICML | 6 |
| 2025 | Instance-aware context with mutually guided vision-language attention for referring image segmentation
Qiule Sun, Jianxin Zhang 0001, Bingbing Zhang 0001, Peihua Li |
Appl. Intell. | 4 |
| 2025 | A2 M2-Net: Adaptively Aligned Multi-scale Moment for Few-Shot Action Recognition
Zilin Gao, Qilong Wang 0001, Bingbing Zhang 0001, Qinghua Hu, Peihua Li |
Int. J. Comput. Vis. | 5 |
| 2025 | Stochastic stylization transformer with self-supervision for iris recognition
Lingyao Jia, Peihua Li |
Multim. Syst. | 3 |
| 2024 | Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge DistillationabstractSince pioneering work of Hinton et al., knowledge distillation based on Kullback-Leibler Divergence (KL-Div) has been predominant, and recently its variants have achieved compelling performance. However, KL-Div only compares probabilities of the corresponding category between the teacher and student while lacking a mechanism for cross-category comparison. Besides, KL-Div is problematic when applied to intermediate layers, as it cannot handle non-overlapping distributions and is unaware of geometry of the underlying manifold. To address these downsides, we propose a methodology of Wasserstein Distance (WD) based knowledge distillation. Specifically, we propose a logit distillation method called WKD-L based on discrete WD, which performs cross-category comparison of probabilities and thus can explicitly leverage rich interrelations among categories. Moreover, we introduce a feature distillation method called WKD-F, which uses a parametric method for modeling feature distributions and adopts continuous WD for transferring knowledge from intermediate layers. Comprehensive evaluations on image classification and object detection have shown (1) for logit distillation WKD-L outperforms very strong KL-Div variants; (2) for feature distillation WKD-F is superior to the KL-Div counterparts and state-of-the-art competitors. Jiaming Lv, Haoyuan Yang, Peihua Li |
NeurIPS | 3 |
| 2024 | AMS-Net: Modeling Adaptive Multi-Granularity Spatio-Temporal Cues for Video Action RecognitionabstractEffective spatio-temporal modeling as a core of video representation learning is challenged by complex scale variations in spatio-temporal cues in videos, especially different visual tempos of actions and varying spatial sizes of moving objects. Most of the existing works handle complex spatio-temporal scale variations based on input-level or feature-level pyramid mechanisms, which, however, rely on expensive multistream architectures or explore multiscale spatio-temporal features in a fixed manner. To effectively capture complex scale dynamics of spatio-temporal cues in an efficient way, this article proposes a single-stream architecture (SS-Arch.) with single-input [namely, adaptive multi-granularity spatio-temporal network (AMS-Net)] to model adaptive multi-granularity (Multi-Gran.) Spatio-temporal cues for video action recognition. To this end, our AMS-Net proposes two core components, namely, competitive progressive temporal modeling (CPTM) block and collaborative spatio-temporal pyramid (CSTP) module. They, respectively, capture fine-grained temporal cues and fuse coarse-level spatio-temporal features in an adaptive manner. It admits that AMS-Net can handle subtle variations in visual tempos and fair-sized spatio-temporal dynamics in a unified architecture. Note that our AMS-Net can be flexibly instantiated based on existing deep convolutional neural networks (CNNs) with the proposed CPTM block and CSTP module. The experiments are conducted on eight video benchmarks, and the results show our AMS-Net establishes state-of-the-art (SOTA) performance on fine-grained action recognition (i.e., Diving48 and FineGym), while performing very competitively on widely used Something-Something and Kinetics. Qilong Wang 0001, Qiyao Hu, Zilin Gao, Peihua Li, Qinghua Hu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Structure correlation-aware attention for Iris recognition
Lingyao Jia, Qiule Sun, Peihua Li |
Neural Comput. Appl. | 3 |
| 2023 | Towards a Deeper Understanding of Global Covariance Pooling in Deep Learning: An Optimization PerspectiveabstractGlobal covariance pooling (GCP) as an effective alternative to global average pooling has shown good capacity to improve deep convolutional neural networks (CNNs) in a variety of vision tasks. Although promising performance, it is still an open problem on how GCP (especially its post-normalization) works in deep learning. In this paper, we make the effort towards understanding the effect of GCP on deep learning from an optimization perspective. Specifically, we first analyze behavior of GCP with matrix power normalization on optimization loss and gradient computation of deep architectures. Our findings show that GCP can improve Lipschitzness of optimization loss and achieve flatter local minima, while improving gradient predictiveness and functioning as a special pre-conditioner on gradients. Then, we explore the effect of post-normalization on GCP from the model optimization perspective, which encourages us to propose a simple yet effective normalization, namely DropCov. Based on above findings, we point out several merits of deep GCP that have not been recognized previously or fully explored, including faster convergence, stronger model robustness and better generalization across tasks. Extensive experimental results using both CNNs and vision transformers on diversified vision tasks provide strong support to our findings while verifying the effectiveness of our method. Qilong Wang 0001, Jiangtao Xie, Pengfei Zhu 0001, Peihua Li, Wangmeng Zuo, Qinghua Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Joint Distribution Matters: Deep Brownian Distance Covariance for Few-Shot ClassificationabstractFew-shot classification is a challenging problem as only very few training examples are given for each new task. One of the effective research lines to address this challenge focuses on learning deep representations driven by a similarity measure between a query image and few support images of some class. Statistically, this amounts to measure the dependency of image features, viewed as random vectors in a high-dimensional embedding space. Previous methods either only use marginal distributions without considering joint distributions, suffering from limited representation capability, or are computationally expensive though harnessing joint distributions. In this paper, we propose a deep Brownian Distance Covariance (DeepBDC) method for few-shot classification. The central idea of DeepBDC is to learn image representations by measuring the discrepancy between joint characteristic functions of embedded features and product of the marginals. As the BDC metric is decoupled, we formulate it as a highly modular and efficient layer. Furthermore, we instantiate DeepBDC in two different few-shot classification frameworks. We make experiments on six standard few-shot image benchmarks, covering general object recognition, fine-grained categorization and cross-domain classification. Extensive evaluations show our DeepBDC significantly outperforms the counterparts, while establishing new state-of-the-art results. The source code is available at http://www.peihuali.org/DeepBDC. Jiangtao Xie, Fei Long 0001, Jiaming Lv, Qilong Wang 0001, Peihua Li |
CVPR | 5 |
| 2022 | Parameterized Approximation Algorithms for TSP
Jianqi Zhou, Peihua Li, Jiong Guo |
ISAAC | 2 |
| 2022 | DropCov: A Simple yet Effective Method for Improving Deep ArchitecturesabstractPrevious works show global covariance pooling (GCP) has great potential to improve deep architectures especially on visual recognition tasks, where post-normalization of GCP plays a very important role in final performance. Although several post-normalization strategies have been studied, these methods pay more close attention to effect of normalization on covariance representations rather than the whole GCP networks, and their effectiveness requires further understanding. Meanwhile, existing effective post-normalization strategies (e.g., matrix power normalization) usually suffer from high computational complexity (e.g., $O(d^{3})$ for $d$-dimensional inputs). To handle above issues, this work first analyzes the effect of post-normalization from the perspective of training GCP networks. Particularly, we for the first time show that \textit{effective post-normalization can make a good trade-off between representation decorrelation and information preservation for GCP, which are crucial to alleviate over-fitting and increase representation ability of deep GCP networks, respectively}. Based on this finding, we can improve existing post-normalization methods with some small modifications, providing further support to our observation. Furthermore, this finding encourages us to propose a novel pre-normalization method for GCP (namely DropCov), which develops an adaptive channel dropout on features right before GCP, aiming to reach trade-off between representation decorrelation and information preservation in a more efficient way. Our DropCov only has a linear complexity of $O(d)$, while being free for inference. Extensive experiments on various benchmarks (i.e., ImageNet-1K, ImageNet-C, ImageNet-A, Stylized-ImageNet, and iNat2017) show our DropCov is superior to the counterparts in terms of efficiency and effectiveness, and provides a simple yet effective method to improve performance of deep architectures involving both deep convolutional neural networks (CNNs) and vision transformers (ViTs). Qilong Wang 0001, Jiangtao Xie, Peihua Li, Qinghua Hu |
NeurIPS | 5 |
| 2022 | Second-order convolutional networks for iris recognition
Lingyao Jia, Xueyu Shi, Qiule Sun, Xingqiang Tang, Peihua Li |
Appl. Intell. | 5 |
| 2022 | Temporal grafter network: Rethinking LSTM for effective video recognition
Bingbing Zhang 0001, Qilong Wang 0001, Zilin Gao, Ruiren Zeng, Peihua Li |
Neurocomputing | 5 |
| 2022 | Constructions of Optimal Uniform Wide-Gap Frequency-Hopping SequencesabstractIn frequency hopping (FH) communication systems, frequency hopping sequences (FHSs) are crucial in determining the system’s anti-jamming performance. If FHSs can ensure a wide-gap between two adjacent frequency points to avoid the frequency points with high interference probability, it will significantly improve the FH communication system’s anti-interference ability. Moreover, if each frequency point appears at the same number of times in a sequence period, the system’s anti-electromagnetic interference will be enhanced. Therefore, it is desirable to employ FHSs with low Hamming autocorrelation, wide frequency-hopping gap, and good uniformity in practical applications. However, to the best of our knowledge, no such infinite classes of FHSs have been reported in the literature to date. This paper aims to present two constructions of uniform wide-gap frequency-hopping sequences (WGFHSs) by concatenating two or three adequately designed sequences. For the first time, we obtain two infinite classes of WGFHSs, which are optimal with respect to the well-known Lempel-Greenberger bound. Peihua Li, Cuiling Fan, Sihem Mesnager, Yang Yang 0005, Zhengchun Zhou |
IEEE Trans. Inf. Theory | 1 |
| 2022 | Detachable Second-Order Pooling: Toward High-Performance First-Order NetworksabstractSecond-order pooling has proved to be more effective than its first-order counterpart in visual classification tasks. However, second-order pooling suffers from the high demand for a computational resource, limiting its use in practical applications. In this work, we present a novel architecture, namely a detachable second-order pooling network, to leverage the advantage of second-order pooling by first-order networks while keeping the model complexity unchanged during inference. Specifically, we introduce second-order pooling at the end of a few auxiliary branches and plug them into different stages of a convolutional neural network. During the training stage, the auxiliary second-order pooling networks assist the backbone first-order network to learn more discriminative feature representations. When training is completed, all auxiliary branches can be removed, and only the backbone first-order network is used for inference. Experiments conducted on CIFAR-10, CIFAR-100, and ImageNet data sets clearly demonstrated the leading performance of our network, which achieves even higher accuracy than second-order networks but keeps the low inference complexity of first-order networks. Lida Li, Jiangtao Xie, Peihua Li, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and LocalizationabstractFor iris recognition in non-cooperative environments, iris segmentation has been regarded as the first most important challenge still open to the biometric community, affecting all downstream tasks from normalization to recognition. In recent years, deep learning technologies have gained significant popularity among various computer vision tasks and also been introduced in iris biometrics, especially iris segmentation. To investigate recent developments and attract more interest of researchers in the iris segmentation method, we organized the 2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization (NIR-ISL 2021) at the 2021 International Joint Conference on Biometrics (IJCB 2021). The challenge was used as a public platform to assess the performance of iris segmentation and localization methods on Asian and African NIR iris images captured in non-cooperative environments. The three best-performing entries achieved solid and satisfactory iris segmentation and localization results in most cases, and their code and models have been made publicly available for reproducibility research. Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Jawad Muhammad, Qi Zhang 0015, Qichuan Tian, Zhaofeng He 0001, Zhenan Sun, Tianbao Liu, Wei Yang 0006, Dongliang Wu, Yingfeng Liu, Ruiye Zhou, Huihai Wu, Junbao Wang, Wantong Xiong, Xueyu Shi, Shao Zeng, Peihua Li, Huijie Wu, Xinhui Zhang, Menghan Zhang, Fadi Boutros, Naser Damer, Arjan Kuijper, Juan E. Tapia, Andres Valenzuela, Christoph Busch 0001, Gourav Gupta, Kiran B. Raja, Xi Wu 0004, Xiaojie Li 0001, Jingfu Yang, Hongyan Jing, Xin Wang 0045, Bin Kong 0001, Youbing Yin, Qi Song 0001, Siwei Lyu, Shu Hu 0001, Leon Premk, Matej Vitek, Vitomir Struc, Peter Peer, Jalil Nourmohammadi-Khiarak, Farhang Jaryani, Samaneh Salehi Nasab, Seyed Naeim Moafinejad, Yasin Amini, Morteza Noshad |
IJCB | 23 |
| 2021 | Temporal-attentive Covariance Pooling Networks for Video RecognitionabstractFor video recognition task, a global representation summarizing the whole contents of the video snippets plays an important role for the final performance. However, existing video architectures usually generate it by using a simple, global average pooling (GAP) method, which has limited ability to capture complex dynamics of videos. For image recognition task, there exist evidences showing that covariance pooling has stronger representation ability than GAP. Unfortunately, such plain covariance pooling used in image recognition is an orderless representative, which cannot model spatio-temporal structure inherent in videos. Therefore, this paper proposes a Temporal-attentive Covariance Pooling (TCP), inserted at the end of deep architectures, to produce powerful video representations. Specifically, our TCP first develops a temporal attention module to adaptively calibrate spatio-temporal features for the succeeding covariance pooling, approximatively producing attentive covariance representations. Then, a temporal covariance pooling performs temporal pooling of the attentive covariance representations to characterize both intra-frame correlations and inter-frame cross-correlations of the calibrated features. As such, the proposed TCP can capture complex temporal dynamics. Finally, a fast matrix power normalization is introduced to exploit geometry of covariance representations. Note that our TCP is model-agnostic and can be flexibly integrated into any video architectures, resulting in TCPNet for effective video recognition. The extensive experiments on six benchmarks (e.g., Kinetics, Something-Something V1 and Charades) using various video architectures show our TCPNet is clearly superior to its counterparts, while having strong generalization ability. The source code is publicly available. Zilin Gao, Qilong Wang 0001, Bingbing Zhang 0001, Qinghua Hu, Peihua Li |
NeurIPS | 5 |
| 2021 | Second-order encoding networks for semantic segmentation
Qiule Sun, Peihua Li |
Neurocomputing | 3 |
| 2021 | Deep CNNs Meet Global Covariance Pooling: Better Representation and GeneralizationabstractCompared with global average pooling in existing deep convolutional neural networks (CNNs), global covariance pooling can capture richer statistics of deep features, having potential for improving representation and generalization abilities of deep CNNs. However, integration of global covariance pooling into deep CNNs brings two challenges: (1) robust covariance estimation given deep features of high dimension and small sample size; (2) appropriate usage of geometry of covariances. To address these challenges, we propose a global Matrix Power Normalized COVariance (MPN-COV) Pooling. Our MPN-COV conforms to a robust covariance estimator, very suitable for scenario of high dimension and small sample size. It can also be regarded as Power-Euclidean metric between covariances, effectively exploiting their geometry. Furthermore, a global Gaussian embedding network is proposed to incorporate first-order statistics into MPN-COV. For fast training of MPN-COV networks, we implement an iterative matrix square root normalization, avoiding GPU unfriendly eigen-decomposition inherent in MPN-COV. Additionally, progressive 1×1 convolutions and group convolution are introduced to compress covariance representations. The proposed methods are highly modular, readily plugged into existing deep CNNs. Extensive experiments are conducted on large-scale object classification, scene categorization, fine-grained visual recognition and texture classification, showing our methods outperform the counterparts and obtain state-of-the-art performance. Qilong Wang 0001, Jiangtao Xie, Wangmeng Zuo, Lei Zhang 0006, Peihua Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Multi-scale structural kernel representation for object detection
Hao Wang 0073, Qilong Wang 0001, Peihua Li, Wangmeng Zuo |
Pattern Recognit. | 3 |
| 2020 | ECA-Net: Efficient Channel Attention for Deep Convolutional Neural NetworksabstractRecently, channel attention mechanism has demonstrated to offer great potential in improving the performance of deep convolutional neural networks (CNNs). However, most existing methods dedicate to developing more sophisticated attention modules for achieving better performance, which inevitably increase model complexity. To overcome the paradox of performance and complexity trade-off, this paper proposes an Efficient Channel Attention (ECA) module, which only involves a handful of parameters while bringing clear performance gain. By dissecting the channel attention module in SENet, we empirically show avoiding dimensionality reduction is important for learning channel attention, and appropriate cross-channel interaction can preserve performance while significantly decreasing model complexity. Therefore, we propose a local cross-channel interaction strategy without dimensionality reduction, which can be efficiently implemented via 1D convolution. Furthermore, we develop a method to adaptively select kernel size of 1D convolution, determining coverage of local cross-channel interaction. The proposed ECA module is both efficient and effective, e.g., the parameters and computations of our modules against backbone of ResNet50 are 80 vs. 24.37M and 4.7e-4 GFlops vs. 3.86 GFlops, respectively, and the performance boost is more than 2% in terms of Top-1 accuracy. We extensively evaluate our ECA module on image classification, object detection and instance segmentation with backbones of ResNets and MobileNetV2. The experimental results show our module is more efficient while performing favorably against its counterparts. Qilong Wang 0001, Banggu Wu, Pengfei Zhu 0001, Peihua Li, Wangmeng Zuo, Qinghua Hu |
CVPR | 4 |
| 2020 | What Deep CNNs Benefit From Global Covariance Pooling: An Optimization PerspectiveabstractRecent works have demonstrated that global covariance pooling (GCP) has the ability to improve performance of deep convolutional neural networks (CNNs) on visual classification task. Despite considerable advance, the reasons on effectiveness of GCP on deep CNNs have not been well studied. In this paper, we make an attempt to understand what deep CNNs benefit from GCP in a viewpoint of optimization. Specifically, we explore the effect of GCP on deep CNNs in terms of the Lipschitzness of optimization loss and the predictiveness of gradients, and show that GCP can make the optimization landscape more smooth and the gradients more predictive. Furthermore, we discuss the connection between GCP and second-order optimization for deep CNNs. More importantly, above findings can account for several merits of covariance pooling for training deep CNNs that have not been recognized previously or fully explored, including significant acceleration of network convergence (i.e., the networks trained with GCP can support rapid decay of learning rates, achieving favorable performance while significantly reducing number of training epochs), stronger robustness to distorted examples generated by image corruptions and perturbations, and good generalization ability to different vision tasks, e.g., object detection and instance segmentation. We conduct extensive experiments using various deep CNN architectures on diversified tasks, and the results provide strong support to our findings. Qilong Wang 0001, Banggu Wu, Dongwei Ren, Peihua Li, Wangmeng Zuo, Qinghua Hu |
CVPR | 5 |
| 2020 | Blind single image super-resolution with a mixture of deep networks
Yifan Wang 0004, Lijun Wang 0001, Hongyu Wang 0001, Peihua Li, Huchuan Lu |
Pattern Recognit. | 4 |
| 2020 | Locality-constrained affine subspace coding for image classification and retrieval
Bingbing Zhang 0001, Qilong Wang 0001, Xiaoxiao Lu, Fasheng Wang, Peihua Li |
Pattern Recognit. | 5 |
| 2020 | Weighted and Class-Specific Maximum Mean Discrepancy for Unsupervised Domain AdaptationabstractAlthough maximum mean discrepancy (MMD) has achieved great success in unsupervised domain adaptation (UDA), most of existing UDA methods ignore the issue of class weight bias across domains, which is ubiquitous and evidently gives rise to the degradation of UDA performance. In this work, we propose two improved MMD metrics, i.e., weighted MMD (WMMD) and class-specific MMD (CMMD), to alleviate the adverse effect caused by the changes of class prior distributions between source and target domains. In WMMD, class-specific auxiliary weights are deployed to reweigh the source samples. In CMMD, we calculate the MMD for each class of source and target samples. Since the class labels of target samples are unknown for UDA problem, we present a classification expectation-maximization algorithm to estimate the pseudo-labels of target samples on the fly and update the model parameters using estimated labels. The proposed methods can be flexibly incorporated into deep convolutional neural networks to form WMMD and CMMD based domain adaptation networks, which we called WDAN and CDAN, respectively. By combining WMMD with CMMD, we present a CWMMD based domain adaptation network (CWDAN) to further improve classification performance. Experiments show that, both WMMD and CMMD benefit the classification accuracy, and our CWDAN can achieve compelling UDA performance in comparison with MMD and the state-of-the-art UDA methods. Hongliang Yan, Zhetao Li, Qilong Wang 0001, Peihua Li, Yong Xu 0001, Wangmeng Zuo |
IEEE Trans. Multim. | 4 |
| 2019 | Global Second-Order Pooling Convolutional NetworksabstractDeep Convolutional Networks (ConvNets) are fundamental to, besides large-scale visual recognition, a lot of vision tasks. As the primary goal of the ConvNets is to characterize complex boundaries of thousands of classes in a high-dimensional space, it is critical to learn higher-order representations for enhancing non-linear modeling capability. Recently, Global Second-order Pooling (GSoP), plugged at the end of networks, has attracted increasing attentions, achieving much better performance than classical, first-order networks in a variety of vision tasks. However, how to effectively introduce higher-order representation in earlier layers for improving non-linear capability of ConvNets is still an open problem. In this paper, we propose a novel network model introducing GSoP across from lower to higher layers for exploiting holistic image information throughout a network. Given an input 3D tensor outputted by some previous convolutional layer, we perform GSoP to obtain a covariance matrix which, after nonlinear transformation, is used for tensor scaling along channel dimension. Similarly, we can perform GSoP along spatial dimension for tensor scaling as well. In this way, we can make full use of the second-order statistics of the holistic image throughout a network. The proposed networks are thoroughly evaluated on large-scale ImageNet-1K, and experiments have shown that they outperform non-trivially the counterparts while achieving state-of-the-art results. Zilin Gao, Jiangtao Xie, Qilong Wang 0001, Peihua Li |
CVPR | 4 |
| 2019 | Deep Global Generalized Gaussian NetworksabstractRecently, global covariance pooling (GCP) has shown great advance in improving classification performance of deep convolutional neural networks (CNNs). However, existing deep GCP networks compute covariance pooling of convolutional activations with assumption that activations are sampled from Gaussian distributions, which may not hold in practice and fails to fully characterize the statistics of activations. To handle this issue, this paper proposes a novel deep global generalized Gaussian network (3G-Net), whose core is to estimate a global covariance of generalized Gaussian for modeling the last convolutional activations. Compared with GCP in Gaussian setting, our 3G-Net assumes the distribution of activations follows a generalized Gaussian, which can capture more precise characteristics of activations. However, there exists no analytic solution for parameter estimation of generalized Gaussian, making our 3G-Net challenging. To this end, we first present a novel regularized maximum likelihood estimator for robust estimating covariance of generalized Gaussian, which can be optimized by a modified iterative re-weighted method. Then, to efficiently estimate the covariance of generaized Gaussian under deep CNN architectures, we approximate this re-weighted method by developing an unrolling re-weighted module and a square root covariance layer. In this way, 3GNet can be flexibly trained in an end-to-end manner. The experiments are conducted on large-scale ImageNet-1K and Places365 datasets, and the results demonstrate our 3G-Net outperforms its counterparts while achieving very competitive performance to state-of-the-arts. Qilong Wang 0001, Peihua Li, Qinghua Hu, Pengfei Zhu 0001, Wangmeng Zuo |
CVPR | 2 |
| 2019 | Resolution-Aware Network for Image Super-ResolutionabstractIn existing deep network-based image super-resolution (SR) methods, each network is only trained for a fixed upscaling factor and can hardly generalize to unseen factors at test time, which is non-scalable in real applications. To mitigate this issue, this paper proposes a resolution-aware network (RAN) for simultaneous SR of multiple factors. The key insight is that SR of multiple factors is essentially different but also shares common operations. To attain stronger generalization across factors, we design an upsampling network (U-Net) consisting of several sub-modules, in which each sub-module implements an intermediate step of the overall image SR and can be shared by SR of different factors. A decision network (D-Net) is further adopted to identify the quality of the input low-resolution image and adaptively select suitable sub-modules to perform SR. U-Net and D-Net together constitute the proposed RAN model, and are jointly trained using a new hierarchical loss function on SR tasks of multiple factors. Experimental evaluations demonstrate that the proposed RAN compares favorably against the state-of-the-art methods and its performance can well generalize across different upscaling factors. Yifan Wang 0004, Lijun Wang 0001, Hongyu Wang 0001, Peihua Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Towards Faster Training of Global Covariance Pooling Networks by Iterative Matrix Square Root NormalizationabstractGlobal covariance pooling in convolutional neural networks has achieved impressive improvement over the classical first-order pooling. Recent works have shown matrix square root normalization plays a central role in achieving state-of-the-art performance. However, existing methods depend heavily on eigendecomposition (EIG) or singular value decomposition (SVD), suffering from inefficient training due to limited support of EIG and SVD on GPU. Towards addressing this problem, we propose an iterative matrix square root normalization method for fast end-to-end training of global covariance pooling networks. At the core of our method is a meta-layer designed with loop-embedded directed graph structure. The meta-layer consists of three consecutive nonlinear structured layers, which perform pre-normalization, coupled matrix iteration and post-compensation, respectively. Our method is much faster than EIG or SVD based ones, since it involves only matrix multiplications, suitable for parallel implementation on GPU. Moreover, the proposed network with ResNet architecture can converge in much less epochs, further accelerating network training. On large-scale ImageNet, we achieve competitive performance superior to existing counterparts. By fine-tuning our models pre-trained on ImageNet, we establish state-of-the-art results on three challenging fine-grained benchmarks. The source code and network models will be available at http://www.peihuali.org/iSQRT-COV. Peihua Li, Jiangtao Xie, Qilong Wang 0001, Zilin Gao |
CVPR | 1 |
| 2018 | Multi-Scale Location-Aware Kernel Representation for Object DetectionabstractAlthough Faster R-CNN and its variants have shown promising performance in object detection, they only exploit simple first-order representation of object proposals for final classification and regression. Recent classification methods demonstrate that the integration of high-order statistics into deep convolutional neural networks can achieve impressive improvement, but their goal is to model whole images by discarding location information so that they cannot be directly adopted to object detection. In this paper, we make an attempt to exploit high-order statistics in object detection, aiming at generating more discriminative representations for proposals to enhance the performance of detectors. To this end, we propose a novel Multi-scale Location-aware Kernel Representation (MLKP) to capture high-order statistics of deep features in proposals. Our MLKP can be efficiently computed on a modified multi-scale feature map using a low-dimensional polynomial kernel approximation. Moreover, different from existing orderless global representations based on high-order statistics, our proposed MLKP is location retentive and sensitive so that it can be flexibly adopted to object detection. Through integrating into Faster R-CNN schema, the proposed MLKP achieves very competitive performance with state-of-the-art methods, and improves Faster R-CNN by 4.9% (mAP), 4.7% (mAP) and 5.0% (AP at IOU=[0.5:0.05:0.95]) on PASCAL VOC 2007, VOC 2012 and MS COCO benchmarks, respectively. Code is available at: https://github.com/Hwang64/MLKP. Hao Wang 0073, Qilong Wang 0001, Mingqi Gao 0006, Peihua Li, Wangmeng Zuo |
CVPR | 4 |
| 2018 | Global Gated Mixture of Second-order Pooling for Improving Deep Convolutional Neural NetworksabstractIn most of existing deep convolutional neural networks (CNNs) for classification, global average (first-order) pooling (GAP) has become a standard module to summarize activations of the last convolution layer as final representation for prediction. Recent researches show integration of higher-order pooling (HOP) methods clearly improves performance of deep CNNs. However, both GAP and existing HOP methods assume unimodal distributions, which cannot fully capture statistics of convolutional activations, limiting representation ability of deep CNNs, especially for samples with complex contents. To overcome the above limitation, this paper proposes a global Gated Mixture of Second-order Pooling (GM-SOP) method to further improve representation ability of deep CNNs. To this end, we introduce a sparsity-constrained gating mechanism and propose a novel parametric SOP as component of mixture model. Given a bank of SOP candidates, our method can adaptively choose Top-K (K > 1) candidates for each input sample through the sparsity-constrained gating module, and performs weighted sum of outputs of K selected candidates as representation of the sample. The proposed GM-SOP can flexibly accommodate a large number of personalized SOP candidates in an efficient way, leading to richer representations. The deep networks with our GM-SOP can be end-to-end trained, having potential to characterize complex, multi-modal distributions. The proposed method is evaluated on two large scale image benchmarks (i.e., downsampled ImageNet-1K and Places365), and experimental results show our GM-SOP is superior to its counterparts and achieves very competitive performance. The source code will be available at http://www.peihuali.org/GM-SOP. Qilong Wang 0001, Zilin Gao, Jiangtao Xie, Wangmeng Zuo, Peihua Li |
NeurIPS | 5 |
| 2018 | Hyperlayer Bilinear Pooling with application to fine-grained categorization and image retrieval
Qiule Sun, Qilong Wang 0001, Jianxin Zhang 0001, Peihua Li |
Neurocomputing | 4 |
| 2018 | Information-Compensated Downsampling for Image Super-ResolutionabstractA large receptive field of deep networks can better incorporate image context and benefits image super-resolution (SR) in many ways. However, common techniques, like strided pooling and convolutional operations, are not directly applicable to SR due to severe image detail losses. In this letter, we circumvent this issue by proposing a new network architecture, namely the information-compensated (IC) downsampling block. It first uses pooling layers to downsample input feature maps and then immediately upsamples the feature maps back to the original size. To further compensate for information loss, skip connections are added to propagate lost features caused by downsampling to the upsampled output. In addition, pixelwise recurrent units are also applied to the downsampled feature maps to model context coherence. Compared with traditional pooling layers, the IC downsampling blocks cannot only enlarge receptive field and better capture image context, but also preserve image details, which are essential to SR. The final network consists of a stack of IC downsampling blocks and can be trained in an end-to-end manner. Experimental results verify that the proposed method performs favorably against the state-of-the-art approaches. Yifan Wang 0004, Lijun Wang 0001, Hongyu Wang 0001, Peihua Li |
IEEE Signal Process. Lett. | 4 |
| 2018 | An Information Geometry-Based Distance Between High-Dimensional Covariances for Scalable ClassificationabstractModeling images/videos with covariance matrices has attracted increasing attentions in various vision tasks, especially in visual classification. For covariances-based visual classification, measuring the distances between covariances is one of the key issues and has been studied for decades. Since the space of covariances is a Riemannian manifold, the geometrical structure of covariances should be favorably considered when designing distance metrics. Although this problem has been widely studied, designing an effective and efficient metric between high-dimensional covariances (HDCOV) for scalable classification is still an open problem. In this paper, we present an information geometry-based distance (IGBD) to tackle this challenge from the perspective of information geometry. Our idea is based on the fact that each covariance can be viewed as a zero-mean Gaussian distribution, and thus the distances between covariances are measured by those between the corresponding Gaussian distributions. The core of our method is to project each distribution, in the form of a set of random samples, to a vector on the tangent space of a common, known distribution on the statistical manifold, based on Fisher information metric and maximum likelihood method. On the tangent space, the Euclidean norm can be used to measure the distances between those sets of projection vectors (or equivalently distributions). The proposed IGBD for HDCOV is computationally efficient and easily combined with a linear support vector machine, suitable for scalable visual classification. The experiments are conducted on various kinds and sizes of benchmarks, and results show the proposed method is efficient and the combination of HDCOV can achieve very competitive performance. Qilong Wang 0001, Xiaoxiao Lu, Peihua Li, Zhenguo Gao, Yongri Piao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | G2DeNet: Global Gaussian Distribution Embedding Network and Its Application to Visual RecognitionabstractRecently, plugging trainable structural layers into deep convolutional neural networks (CNNs) as image representations has made promising progress. However, there has been little work on inserting parametric probability distributions, which can effectively model feature statistics, into deep CNNs in an end-to-end manner. This paper proposes a Global Gaussian Distribution embedding Network (G2DeNet) to take a step towards addressing this problem. The core of G2DeNet is a novel trainable layer of a global Gaussian as an image representation plugged into deep CNNs for end-to-end learning. The challenge is that the proposed layer involves Gaussian distributions whose space is not a linear space, which makes its forward and backward propagations be non-intuitive and non-trivial. To tackle this issue, we employ a Gaussian embedding strategy which respects the structures of both Riemannian manifold and smooth group of Gaussians. Based on this strategy, we construct the proposed global Gaussian embedding layer and decompose it into two sub-layers: the matrix partition sub-layer decoupling the mean vector and covariance matrix entangled in the embedding matrix, and the square-rooted, symmetric positive definite matrix sub-layer. In this way, we can derive the partial derivatives associated with the proposed structural layer and thus allow backpropagation of gradients. Experimental results on large scale region classification and fine-grained recognition tasks show that G2DeNet is superior to its counterparts, capable of achieving state-of-the-art performance. Qilong Wang 0001, Peihua Li, Lei Zhang 0006 |
CVPR | 2 |
| 2017 | Mind the Class Weight Bias: Weighted Maximum Mean Discrepancy for Unsupervised Domain AdaptationabstractIn domain adaptation, maximum mean discrepancy (MMD) has been widely adopted as a discrepancy metric between the distributions of source and target domains. However, existing MMD-based domain adaptation methods generally ignore the changes of class prior distributions, i.e., class weight bias across domains. This remains an open problem but ubiquitous for domain adaptation, which can be caused by changes in sample selection criteria and application scenarios. We show that MMD cannot account for class weight bias and results in degraded domain adaptation performance. To address this issue, a weighted MMD model is proposed in this paper. Specifically, we introduce class-specific auxiliary weights into the original MMD for exploiting the class prior probability on source and target domains, whose challenge lies in the fact that the class label in target domain is unavailable. To account for it, our proposed weighted MMD model is defined by introducing an auxiliary weight for each class in the source domain, and a classification EM algorithm is suggested by alternating between assigning the pseudo-labels, estimating auxiliary weights and updating model parameters. Extensive experiments demonstrate the superiority of our weighted MMD over conventional MMD for domain adaptation. Hongliang Yan, Yukang Ding, Peihua Li, Qilong Wang 0001, Yong Xu 0001, Wangmeng Zuo |
CVPR | 3 |
| 2017 | Is Second-Order Information Helpful for Large-Scale Visual Recognition?abstractBy stacking layers of convolution and nonlinearity, convolutional networks (ConvNets) effectively learn from lowlevel to high-level features and discriminative representations. Since the end goal of large-scale recognition is to delineate complex boundaries of thousands of classes, adequate exploration of feature distributions is important for realizing full potentials of ConvNets. However, state-of-the-art works concentrate only on deeper or wider architecture design, while rarely exploring feature statistics higher than first-order. We take a step towards addressing this problem. Our method consists in covariance pooling, instead of the most commonly used first-order pooling, of highlevel convolutional features. The main challenges involved are robust covariance estimation given a small sample of large-dimensional features and usage of the manifold structure of covariance matrices. To address these challenges, we present a Matrix Power Normalized Covariance (MPNCOV) method. We develop forward and backward propagation formulas regarding the nonlinear matrix functions such that MPN-COV can be trained end-to-end. In addition, we analyze both qualitatively and quantitatively its advantage over the well-known Log-Euclidean metric. On the ImageNet 2012 validation set, by combining MPN-COV we achieve over 4%, 3% and 2.5% gains for AlexNet, VGG-M and VGG-16, respectively; integration of MPN-COV into 50-layer ResNet outperforms ResNet-101 and is comparable to ResNet-152. The source code will be available on the project page: http://www.peihuali.org/MPN-COV. Peihua Li, Jiangtao Xie, Qilong Wang 0001, Wangmeng Zuo |
ICCV | 1 |
| 2017 | Part-based convolutional neural network for visual recognitionabstractMid-level element based representations have been proven to be very effective for visual recognition. We present a method to discover discriminative elements based on deep Convolutional Neural Networks (CNNs), namely Part-based CNN (P-CNN), which acts as the role of encoding module in part-based representation. The P-CNN can be attached at arbitrary layer of a pre-trained CNN and be trained using image-level labels. The training of P-CNN essentially corresponds to the optimization and selection of discriminative mid-level visual elements. For an input image, the output of P-CNN is naturally the part-based coding and can be directly used for image recognition. By applying P-CNN to multiple layers of a pretrained CNN, more diverse visual elements can be obtained for visual recognitions. Experiments are conducted on two recognition tasks and their results demonstrate the effectiveness of the proposed method. Lingxiao Yang, Xiaohua Xie, Peihua Li, David Zhang 0001, Lei Zhang 0006 |
ICIP | 3 |
| 2017 | Ordered over-relaxation based Langevin Monte Carlo sampling for visual tracking
Fasheng Wang, Peihua Li, Xucheng Li, Mingyu Lu |
Neurocomputing | 2 |
| 2017 | Joint distance and similarity measure learning based on triplet-based constraints
Mu Li 0005, Qilong Wang 0001, David Zhang 0001, Peihua Li, Wangmeng Zuo |
Inf. Sci. | 4 |
| 2017 | Local Log-Euclidean Multivariate Gaussian Descriptor and Its Application to Image ClassificationabstractThis paper presents a novel image descriptor to effectively characterize the local, high-order image statistics. Our work is inspired by the Diffusion Tensor Imaging and the structure tensor method (or covariance descriptor), and motivated by popular distribution-based descriptors such as SIFT and HoG. Our idea is to associate one pixel with a multivariate Gaussian distribution estimated in the neighborhood. The challenge lies in that the space of Gaussians is not a linear space but a Riemannian manifold. We show, for the first time to our knowledge, that the space of Gaussians can be equipped with a Lie group structure by defining a multiplication operation on this manifold, and that it is isomorphic to a subgroup of the upper triangular matrix group. Furthermore, we propose methods to embed this matrix group in the linear space, which enables us to handle Gaussians with Euclidean operations rather than complicated Riemannian operations. The resulting descriptor, called Local Log-Euclidean Multivariate Gaussian (L2EMG) descriptor, works well with low-dimensional and high-dimensional raw features. Moreover, our descriptor is a continuous function of features without quantization, which can model the first- and second-order statistics. Extensive experiments were conducted to evaluate thoroughly L2EMG, and the results showed that L2EMG is very competitive with state-of-the-art descriptors in image classification. Peihua Li, Qilong Wang 0001, Hui Zeng 0001, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | High-Order Local Pooling and Encoding Gaussians Over a Dictionary of GaussiansabstractLocal pooling (LP) in configuration (feature) space proposed by Boureau et al. explicitly restricts similar features to be aggregated, which can preserve as much discriminative information as possible. At the time it appeared, this method combined with sparse coding achieved competitive classification results with only a small dictionary. However, its performance lags far behind the state-of-the-art results as only the zero-order information is exploited. Inspired by the success of high-order statistical information in existing advanced feature coding or pooling methods, we make an attempt to address the limitation of LP. To this end, we present a novel method called high-order LP (HO-LP) to leverage the information higher than the zero-order one. Our idea is intuitively simple: we compute the first- and second-order statistics per configuration bin and model them as a Gaussian. Accordingly, we employ a collection of Gaussians as visual words to represent the universal probability distribution of features from all classes. Our problem is naturally formulated as encoding Gaussians over a dictionary of Gaussians as visual words. This problem, however, is challenging since the space of Gaussians is not a Euclidean space but forms a Riemannian manifold. We address this challenge by mapping Gaussians into the Euclidean space, which enables us to perform coding with common Euclidean operations rather than complex and often expensive Riemannian operations. Our HO-LP preserves the advantages of the original LP: pooling only similar features and using a small dictionary. Meanwhile, it achieves very promising performance on standard benchmarks, with either conventional, hand-engineered features or deep learning-based features. Peihua Li, Hui Zeng 0001, Qilong Wang 0001, Simon C. K. Shiu, Lei Zhang 0006 |
IEEE Trans. Image Process. | 1 |
| 2016 | RAID-G: Robust Estimation of Approximate Infinite Dimensional Gaussian with Application to Material RecognitionabstractInfinite dimensional covariance descriptors can provide richer and more discriminative information than their low dimensional counterparts. In this paper, we propose a novel image descriptor, namely, robust approximate infinite dimensional Gaussian (RAID-G). The challenges of RAID-G mainly lie on two aspects: (1) description of infinite dimensional Gaussian is difficult due to its non-linear Riemannian geometric structure and the infinite dimensional setting, hence effective approximation is necessary, (2) traditional maximum likelihood estimation (MLE) is not robust to high (even infinite) dimensional covariance matrix in Gaussian setting. To address these challenges, explicit feature mapping (EFM) is first introduced for effective approximation of infinite dimensional Gaussian induced by additive kernel function, and then a new regularized MLE method based on von Neumann divergence is proposed for robust estimation of covariance matrix. The EFM and proposed regularized MLE allow a closed-form of RAID-G, which is very efficient and effective for high dimensional features. We extend RAID-G by using the outputs of deep convolutional neural networks as original features, and apply it to material recognition. Our approach is evaluated on five material benchmarks and one fine-grained benchmark. It achieves 84.9% accuracy on FMD and 86.3% accuracy on UIUC material database, which are much higher than state-of-the-arts. Qilong Wang 0001, Peihua Li, Wangmeng Zuo, Lei Zhang 0006 |
CVPR | 2 |
| 2016 | Evaluation of ground distances and features in EMD-based GMM matching for texture classification
Hua Hao, Qilong Wang 0001, Peihua Li, Lei Zhang 0006 |
Pattern Recognit. | 3 |
| 2016 | Towards effective codebookless model for image classification
Qilong Wang 0001, Peihua Li, Lei Zhang 0006, Wangmeng Zuo |
Pattern Recognit. | 2 |
| 2015 | From dictionary of visual words to subspaces: Locality-constrained affine subspace codingabstractThe locality-constrained linear coding (LLC) is a very successful feature coding method in image classification. It makes known the importance of locality constraint which brings high efficiency and local smoothness of the codes. However, in the LLC method the geometry of feature space is described by an ensemble of representative points (visual words) while discarding the geometric structure immediately surrounding them. Such a dictionary only provides a crude, piecewise constant approximation of the data manifold. To approach this problem, we propose a novel feature coding method called locality-constrained affine subspace coding (LASC). The data manifold in LASC is characterized by an ensemble of subspaces attached to the representative points (or affine subspaces), which can provide a piecewise linear approximation of the manifold. Given an input descriptor, we find its top-k neighboring subspaces, in which the descriptor is linearly decomposed and weighted to form the first-order LASC vector. Inspired by the success of usage of higher-order information in image classification, we propose the second-order LASC vector based on the Fisher information metric for further performance improvement. We make experiments on challenging benchmarks and experiments have shown the LASC method is very competitive. Peihua Li, Xiaoxiao Lu, Qilong Wang 0001 |
CVPR | 1 |
| 2015 | Ask the dictionary: Soft-assignment location-orientation pooling for image classificationabstractThe pooling step is one of the key components of the well-known Bag-of-visual words (BoW) model widely used in image classification. In this paper, we propose a novel pooling method, which is called Soft-Assignment Location-Orientation Pooling (SALOP). Inspired by the bag of statistical sampling analysis (Bossa), SALOP also explores the effect of dictionary for pooling method, but leverages both location and orientation information between the local descriptors and the atoms of dictionary to aggregate feature codes. Moreover, different from existing pooling methods, SALOP employs a soft-assignment pooling scheme to handle ambiguity and uncertainty existing in the pooling process. The evaluation is conducted on two image benchmarks: Scene15 and PASCAL VOC 2007. The experimental results show our SALOP can achieve promising performances. Qilong Wang 0001, Xiaona Deng, Peihua Li, Lei Zhang 0006 |
ICIP | 3 |
| 2015 | High-order information for robust iris recognition under less controlled conditionsabstractIris recognition has achieved great progress in cooperative environments in the past decades. However, in less controlled conditions it is still an open and challenging problem because of severe noisy factors induced by non-cooperative subjects. For handling this challenging problem, we propose a method called ordinal measure of outer product tensor (O2PT) which leverages the high-order information of image features. O2PT consists of two components. First we compute outer product tensors of raw features (e.g. SIFT) which are vectorized and locally aggregated, characterizing the second-order statistics of raw features. And then we compute the ordinal measure of the aggregated outer product tensors to model the order relation of iris texture, which makes the representation more compact and robust to noise and illumination changes. Furthermore, we combine two modalities to improve the matching performance, namely, O2PT for iris image matching and Fisher Vector (FV), which also exploits the high-order information, for eye image matching. We have achieved competitive matching performance on the challenging UBIRIS.v2 and CASIA-Iris-Thousand databases. Guanglei Yang, Hui Zeng 0001, Peihua Li, Lei Zhang 0006 |
ICIP | 3 |
| 2015 | Manifold Kernel Sparse Representation of Symmetric Positive-Definite Matrices and Its ApplicationsabstractThe symmetric positive-definite (SPD) matrix, as a connected Riemannian manifold, has become increasingly popular for encoding image information. Most existing sparse models are still primarily developed in the Euclidean space. They do not consider the non-linear geometrical structure of the data space, and thus are not directly applicable to the Riemannian manifold. In this paper, we propose a novel sparse representation method of SPD matrices in the data-dependent manifold kernel space. The graph Laplacian is incorporated into the kernel space to better reflect the underlying geometry of SPD matrices. Under the proposed framework, we design two different positive definite kernel functions that can be readily transformed to the corresponding manifold kernels. The sparse representation obtained has more discriminating power. Extensive experimental results demonstrate good performance of manifold kernel sparse codes in image classification, face recognition, and visual tracking. Yuwei Wu 0001, Yunde Jia, Peihua Li, Jian Zhang 0002, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | Shrinkage Expansion Adaptive Metric Learning
Qilong Wang 0001, Wangmeng Zuo, Lei Zhang 0006, Peihua Li |
ECCV (7) | 4 |
| 2014 | The first ICB* competition on iris recognitionabstractIris recognition becomes an important technology in our society. Visual patterns of human iris provide rich texture information for personal identification. However, it is greatly challenging to match intra-class iris images with large variations in unconstrained environments because of noises, illumination variation, heterogeneity and so on. To track current state-of-the-art algorithms in iris recognition, we organized the first ICB* Competition on Iris Recognition in 2013 (or ICIR2013 shortly). In this competition, 8 participants from 6 countries submitted 13 algorithms totally. All the algorithms were trained on a public database (e.g. CASIA-Iris-Thousand [3]) and evaluated on an unpublished database. The testing results in terms of False Non-match Rate (FNMR) when False Match Rate (FMR) is 0.0001 are taken to rank the submitted algorithms. Man Zhang 0005, Jing Liu 0062, Zhenan Sun, Tieniu Tan, Wu Su, Fernando Alonso-Fernandez, Valérian Némesin, Nadia Othman, Koichi Noda, Peihua Li, Edmundo Hoyle, Akanksha Joshi |
IJCB | 10 |
| 2013 | A Novel Earth Mover's Distance Methodology for Image Matching with Gaussian Mixture ModelsabstractThe similarity or distance measure between Gaussian mixture models (GMMs) plays a crucial role in content-based image matching. Though the Earth Mover's Distance (EMD) has shown its advantages in matching histogram features, its potentials in matching GMMs remain unclear and are not fully explored. To address this problem, we propose a novel EMD methodology for GMM matching. We first present a sparse representation based EMD called SR-EMD by exploiting the sparse property of the underlying problem. SR-EMD is more efficient and robust than the conventional EMD. Second, we present two novel ground distances between component Gaussians based on the information geometry. The perspective from the Riemannian geometry distinguishes the proposed ground distances from the classical entropy-or divergence-based ones. Furthermore, motivated by the success of distance metric learning of vector data, we make the first attempt to learn the EMD distance metrics between GMMs by using a simple yet effective supervised pair-wise based method. It can adapt the distance metrics between GMMs to specific classification tasks. The proposed method is evaluated on both simulated data and benchmark real databases and achieves very promising performance. Peihua Li, Qilong Wang 0001, Lei Zhang 0006 |
ICCV | 1 |
| 2013 | Log-Euclidean Kernels for Sparse Representation and Dictionary LearningabstractThe symmetric positive definite (SPD) matrices have been widely used in image and vision problems. Recently there are growing interests in studying sparse representation (SR) of SPD matrices, motivated by the great success of SR for vector data. Though the space of SPD matrices is well-known to form a Lie group that is a Riemannian manifold, existing work fails to take full advantage of its geometric structure. This paper attempts to tackle this problem by proposing a kernel based method for SR and dictionary learning (DL) of SPD matrices. We disclose that the space of SPD matrices, with the operations of logarithmic multiplication and scalar logarithmic multiplication defined in the Log-Euclidean framework, is a complete inner product space. We can thus develop a broad family of kernels that satisfies Mercer's condition. These kernels characterize the geodesic distance and can be computed efficiently. We also consider the geometric structure in the DL process by updating atom matrices in the Riemannian space instead of in the Euclidean space. The proposed method is evaluated with various vision problems and shows notable performance gains over state-of-the-arts. Peihua Li, Qilong Wang 0001, Wangmeng Zuo, Lei Zhang 0006 |
ICCV | 1 |
| 2013 | Variational Earth Mover's Distance for Image SegmentationabstractImage segmentation using similarity or dissimilarity measures between probability distributions has been of great research interest in recent years. It is shown that the cross-bin metrics such as EMD is superior to the bin-wise metrics. However, existing segmentation approaches involving EMD are limited to univariate distributions, or one-dimensional marginal distributions of multidimensional features. This paper presents a novel segmentation method based on the variational EMD (VEMD) model, which can exploit joint distributions of multidimensional features. This method formulates the segmentation problem as the minimization of the EMD-based functional, which measures the distance between the foreground (resp. background) distribution and the reference foreground (resp. background) distribution. Using the simplex method and theory of shape derivative, we minimize the functional and obtain the gradient descent flow. We use a Gaussian filtering level-set method to obtain the numerical solution, in which the level-set re-initialization and smoothness constraint commonly imposed by the contour length are not necessary. Experiments show that the proposed method outperforms the state-of-the-art segmentation methods in the presence of illumination changes and noise. Peihua Li, Qilong Wang 0001 |
ICIG | 1 |
| 2012 | Robust Registration-Based Tracking by Sparse Representation with Model Update
Peihua Li, Qilong Wang 0001 |
ACCV (3) | 1 |
| 2012 | Local Log-Euclidean Covariance Matrix (L2ECM) for Image Representation and Its Applications
Peihua Li, Qilong Wang 0001 |
ECCV (3) | 1 |
| 2012 | Iris recognition using ordinal encoding of Log-Euclidean covariance matrices
Peihua Li, Guolong Wu |
ICPR | 1 |
| 2012 | Weighted co-occurrence phase histogram for iris recognition
Peihua Li |
Pattern Recognit. Lett. | 1 |
| 2012 | Iris recognition in non-ideal imaging conditions
Peihua Li, Hongwei Ma |
Pattern Recognit. Lett. | 1 |
| 2011 | Tracking Objects Using Orientation Covariance Matrices
Peihua Li |
ICIC (1) | 1 |
| 2011 | An Iris Recognition Approach with SIFT Descriptors
Peihua Li |
ICIC (2) | 2 |
| 2010 | Robust and accurate iris segmentation in very noisy iris images
Peihua Li, Lijuan Xiao |
Image Vis. Comput. | 1 |
| 2009 | Robust Acoustic Source Localization with TDOA Based RANSAC Algorithm
Peihua Li, Xianzhe Ma |
ICIC (1) | 1 |
| 2009 | A Parabolic Detection Algorithm Based on Kernel Density Estimation
Peihua Li |
ICIC (1) | 3 |
| 2008 | An incremental method for accurate iris segmentationabstractThe paper presents an incremental method for accurate iris segmentation. Firstly, observing the characteristics of iris images, we search for a square region that contains pupil within or nearby which a specular highlight lies. Means and standard deviations of both pupil and specular highlight are employed for detection of such a square, and Integral Images are used to accelerate the detection procedure. Next Canny edge detection followed by Hough transform are used for accurate localization of pupillary boundary. Secondly, by seeking points with maximum gradient along two line segments radiating from pupil center, the radius of outer, limbic boundary can be coarsely determined. According to the rough radius, two annulus sectors are found out within which limbic boundary is finely localized by Canny edge detection plus Hough transform as well. The incremental technique reduces region for edge detection and parameter space for Hough transform, facilitating accurate and fast iris segmentation. Experiments on publicly available UBIRIS database show that the proposed method has encouraging performance. Peihua Li |
ICPR | 1 |
| 2008 | Improvement on Mean Shift based tracking using second-order informationabstractObject tracking based on Mean Shift (MS) algorithm has been very successful and thus receives significant research interests. Unfortunately, traditional MS based tracking only utilizes the gradient of the similarity function (SF), neglecting completely higher-order information of SF. The paper regards MS based tracking as an optimization problem, and proposes to make use of both the Gradient and Hessian of SF. Specifically, we introduce Newton algorithm with constant, unit step and Newton with varying step lengths, and Trust region algorithm. The advantage of exploiting higher-order information is that higher convergence rate and better performance are achieved. Diverse experiments are made to compare traditional MS based tracking with the proposed algorithms, showing that the proposed algorithms have better performance at comparable computational cost. Lijuan Xiao, Peihua Li |
ICPR | 2 |
| 2008 | Histogram feature-based Fisher linear discriminant for face detection
Haijing Wang, Peihua Li, Tianwen Zhang |
Neural Comput. Appl. | 2 |
| 2008 | An Adaptive Binning Color Model for Mean Shift TrackingabstractThe mean shift (MS) algorithm for object tracking using color has recently received a significant amount of attention thanks to its effectiveness and efficiency. Most current work, unfortunately, failing to notice that object color is usually very compactly distributed, partitions uniformly the whole color space and thus leads to a large number of void bins and limited capability of representing object color distribution. Also, there lacks a systematic way to determine automatically the number of bins. Aiming to address these problems, this paper presents an adaptive binning color model for MS tracking. First, the object color is analyzed based on a clustering algorithm and, according to the clustering result, the color space of the object is partitioned into subspaces by orthonormal transformation. Then, a color model is defined by considering the weighted number of pixels as well as intra-cluster distribution based on independent component analysis (ICA), and a similarity measure is introduced to evaluate likeness between the reference and the candidate models. Finally, the MS algorithm is developed based on the proposed color model and its computational complexity is analyzed. Experiments show that the proposed algorithm has better tracking performance than the conventional MS algorithm at the cost of moderately increasing computational load. Peihua Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Boosted Gaussian Classifier with Integral Histogram for Face DetectionabstractNovel features and weak classifiers are proposed for face detection within the AdaBoost learning framework. Features are histograms computed from a set of spatial templates in filtered images. The filter banks consist of Intensity, Laplacian of Gaussian (Difference of Gaussians), and Gabor filters, aiming to capture spatial and frequency properties of faces at different scales and orientations. Features selected by AdaBoost learning, each of which corresponds to a histogram with a pair of filter and template, can thus be interpreted as boosted marginal distributions of faces. We fit the Gaussian distribution of each histogram feature only for positives (faces) in the sample set as the weak classifier. The results of the experiment demonstrate that classifiers with corresponding features are more powerful in describing the face pattern than haar-like rectangle features introduced by Viola and Jones. Haijing Wang, Peihua Li, Tianwen Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2006 | Histogram Features-Based Fisher Linear Discriminant for Face Detection
Haijing Wang, Peihua Li, Tianwen Zhang |
ACCV (2) | 2 |
| 2006 | A clustering-based color model and integral images for fast object tracking
Peihua Li |
Signal Process. Image Commun. | 1 |
| 2005 | Novel likelihood estimation technique based on boosting detectorabstractThis paper presents novel likelihood estimation to be used for particle filter based object tracking. The likelihood estimation is built upon cascade object detector trained with Gentle AdaBoost (GAB), in order to capture the probability of existence of object. Two strategies are adopted to construct the likelihood functions: probability-intra-stage (PIS) corresponding to real output of each weak classifier in the same stage, and probability-outer-stage (POS) corresponding to the depth reached in the cascade detector. Five kinds of likelihood functions are thus proposed based on the trained GAB detector. Our experiment shows the likelihood functions are able to characterize probabilistically the existence of object accurately, having much higher confidence value in object regions than that in background, and that the integral strategy of PIS and POS is the best choice. Haijing Wang, Peihua Li, Tianwen Zhang |
ICIP (3) | 2 |
| 2005 | A Shape Tracking Algorithm for Visual ServoingabstractThe paper contributes to presenting both an accurate and robust shape tracking algorithm and a novel visual servoing method. Two steps are involved in the tracking algorithm. Firstly the object shape is assumed to vary under an affine model, and the edge detection is performed along the normal lines to the contour. As a result it is possible to use a Kalman filter to perform efficient tracking. The second step concerns image matching based on perspective model, which is achieved iteratively by searching locally along the normal lines also. As to visual servoing we propose to control the translations of the robot with the normalized zeroth and first order image moments, and to control the orientation with rotation axis and angle extracted from a Homography matrix. Two experiments demonstrate that the tracking algorithm is accurate and robust enough to be used in visual servoing, and the novel visual servoing method is superior to traditional ones. Peihua Li, François Chaumette, Omar Tahri |
ICRA | 1 |
| 2004 | Unscented Kalman filter for visual curve tracking
Peihua Li, Tianwen Zhang |
Image Vis. Comput. | 1 |
| 2003 | Visual contour tracking based on particle filters
Peihua Li, Tianwen Zhang, Arthur E. C. Pece |
Image Vis. Comput. | 1 |