EDBT 2026 Demo / reviewers in the wild / expert
Pengfei Fang
dblp:204/7650
· DBLP profile ↗
58ranked-venue papers
14as first author
54since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 9 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 8 first-author · 32 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Intra-Image Mining and Symmetric Maximum Concept Matching for Few Shot Out-of-Distribution DetectionabstractRecent vision-language model (VLM)-based methods have achieved promising results in zero-shot out-of-distribution (OOD) detection by effectively leveraging the local patch features. However, the zero-shot nature inherently comes with two limitations: 1) imperfect local feature prototypes; 2) lack of OOD prototypes. In this paper, we propose Intra-Image Mining (IIM), a lightweight framework designed to overcome these limitations in a few-shot manner. IIM is motivated by the fact that local patches within an image often exhibit diverse semantics, with some patches deviating from the main class concept. Therefore, for each image, we first select the top-k class prototype-related patches as positive samples and leverage them to refine and optimize the local feature prototype. Then, the next top-k among the remaining patches are selected as negatives—serving as OOD signals to construct OOD prototypes. This process yields coherent local positives and challenging negatives, effectively enhancing the model’s local feature discrimination. Besides, we propose a novel inference strategy named Symmetric Maximum Concept Matching (S-MCM). While existing approaches typically adopt an image-to-text scheme—comparing the image features to textual class prototypes—S-MCM further incorporate a text-to-image perspective, leading to more reliable OOD detection. We also propose two benchmarks to analyze the impact of semantic diversity within ID dataset. Built on a frozen VLM, IIM, in conjunction with S-MCM, achieves consistent gains in OOD detection on ImageNet-1k and other benchmarks, outperforming prior methods in FPR95 and AUROC across various few-shot settings. Kaixiang Chen, Pengfei Fang, Hui Xue 0002 |
AAAI | 2 |
| 2026 | HAP: Harmonized Amplitude Perturbation for Cross-Domain Few-Shot LearningabstractCross-Domain Few-Shot Learning (CD-FSL) remains a significant challenge due to substantial distribution shifts between source and target domains. While prior approaches primarily focus on spatial alignment, they often overlook discrepancies in the frequency domain. In this paper, we reveal frequency band discretization as a key phenomenon, characterized by intra-domain low-frequency dominance, inter-domain amplitude divergence, and limited high-frequency variation. This spectral disharmony biases models toward low-frequency components, leading to spectral collapse. We quantify spectral collapse via the effective rank, a principled measure of spectral diversity. To mitigate spectral collapse, we propose Harmonized Amplitude Perturbation (HAP), a frequency-domain augmentation strategy that perturbs the amplitude spectrum via frequency-aware gains sampled from Harmonized Distributions, while fixing the phase spectrum to maintain semantic integrity. Extensive experiments on both Cross-Domain Few-Shot Image Classification and Object Detection benchmarks demonstrate that HAP effectively increases spectral diversity and consistently improves generalization, outperforming state-of-the-art methods without introducing extra model complexity. Wenqian Li 0005, Pengfei Fang, Hui Xue 0002 |
AAAI | 2 |
| 2026 | Adaptive Hyperbolic Kernels: Modulated Embedding in de Branges-Rovnyak SpacesabstractHierarchical data pervades diverse machine learning applications, including natural language processing, computer vision, and social network analysis. Hyperbolic space, characterized by its negative curvature, has demonstrated strong potential in such tasks due to its capacity to embed hierarchical structures with minimal distortion. Previous evidence indicates that the hyperbolic representation capacity can be further enhanced through kernel methods. However, existing hyperbolic kernels still suffer from mild geometric distortion or lack adaptability. This paper addresses these issues by introducing a curvature-aware de Branges–Rovnyak space, a reproducing kernel Hilbert space (RKHS) that is isometric to a Poincaré ball. We design an adjustable multiplier to select the appropriate RKHS corresponding to the hyperbolic space with any curvature adaptively. Building on this foundation, we further construct a family of adaptive hyperbolic kernels, including the novel adaptive hyperbolic radial kernel, whose learnable parameters modulate hyperbolic features in a task-aware manner. Extensive experiments on visual and language benchmarks demonstrate that our proposed kernels outperform existing hyperbolic kernels in modeling hierarchical dependencies. Leping Si, Meimei Yang, Hui Xue 0002, Shipeng Zhu, Pengfei Fang |
AAAI | 5 |
| 2026 | Towards Understanding the Design of Shared Bodily Control via Exoskeleton-based PlayabstractEmerging technologies such as exoskeletons and electrical muscle stimulation can initiate movement within the human body, blurring the boundary between user and machine. While prior research has explored how such systems augment bodily action, most focus on movement execution rather than decision-making. In this work, we investigate what happens when a bodily-integrated system acts with its own logic and initiates bodily movement alongside users. We present three game scenarios where an exoskeleton controls one arm while the user controls the other, designed to evoke different relational framings: proxy, collaboration, and opposition. Through a qualitative study (N = 16), we examine how users interpret such interactions, and how shared bodily control shapes bodily experience and human-machine relationship. We further contribute a set of implications for designing bodily technologies that decide and move together with users, opening up design possibilities for systems that share bodily control, not merely actuate on users’ behalf. Zhuying Li 0001, Rakesh Patibanda, Pengfei Fang, Florian 'Floyd' Mueller |
CHI | 4 |
| 2026 | Task-adaptive Routing Adapters for Multi-label Few-shot Learning
Chenzi Yang, Pengfei Fang |
ICIC (12) | 2 |
| 2025 | SVasP: Self-Versatility Adversarial Style Perturbation for Cross-Domain Few-Shot LearningabstractCross-Domain Few-Shot Learning (CD-FSL) aims to transfer knowledge from seen source domains to unseen target domains, which is crucial for evaluating the generalization and robustness of models. Recent studies focus on utilizing visual styles to bridge the domain gap between different domains. However, the serious dilemma of gradient instability and local optimization problem occurs in those style-based CD-FSL methods. This paper addresses these issues and proposes a novel crop-global style perturbation method, called Self-Versatility Adversarial Style Perturbation (SVasP), which enhances the gradient stability and escapes from poor sharp minima jointly. Specifically, SVasP simulates more diverse potential target domain adversarial styles via diversifying input patterns and aggregating localized crop style gradients, to serve as global style perturbation stabilizers within one image, a concept we refer to as self-versatility. Then a novel objective function is proposed to maximize visual discrepancy while maintaining semantic consistency between global, crop, and adversarial features. Having the stabilized global style perturbation in the training phase, one can obtain a flattened minima in the loss landscape, boosting the transferability of the model to the target domains. Extensive experiments on multiple benchmark datasets demonstrate that our method significantly outperforms existing state-of-the-art methods. Wenqian Li 0005, Pengfei Fang, Hui Xue 0002 |
AAAI | 2 |
| 2025 | PEARL: Input-Agnostic Prompt Enhancement with Negative Feedback Regulation for Class-Incremental LearningabstractClass-incremental learning (CIL) aims to continuously introduce novel categories into a classification system without forgetting previously learned ones, thus adapting to evolving data distributions. Researchers are currently focusing on leveraging the rich semantic information of pre-trained models (PTMs) in CIL tasks. Prompt learning has been adopted in CIL for its ability to adjust data distribution to better align with pre-trained knowledge. This paper critically examines the limitations of existing methods from the perspective of prompt learning, which heavily rely on input information. To address this issue, we propose a novel PTM-based CIL method called Input-Agnostic Prompt Enhancement with NegAtive Feedback ReguLation (PEARL). In PEARL, we implement an input-agnostic global prompt coupled with an adaptive momentum update strategy to reduce the model's dependency on data distribution, thereby effectively mitigating catastrophic forgetting. Guided by negative feedback regulation, this adaptive momentum update addresses the parameter sensitivity inherent in fixed-weight momentum updates. Furthermore, it fosters the continuous enhancement of the prompt for new tasks by harnessing correlations between different tasks in CIL. Experiments on six benchmarks demonstrate that our method achieves state-of-the-art performance. Yongchun Qin, Pengfei Fang, Hui Xue 0002 |
AAAI | 2 |
| 2025 | Multi-Modal Interactive Agent Layer for Few-Shot Universal Cross-Domain Retrieval and BeyondabstractThis paper firstly addresses the challenge of few-shot universal cross-domain retrieval (FS-UCDR), enabling machines trained with limited data to generalize to novel retrieval scenarios, with queries from entirely unknown domains and categories. To achieve this, we first formally define the FS-UCDR task and propose the Multi-Modal Interactive Agent Layer (MAIL), which enhances the cross-modal interaction in vision-language models (VLMs) by aligning the parameter updates of target layer pairs across modalities.
Specifically, MAIL freezes the selected target layer pair and introduces a trainable agent layer pair to approximate localized parameter updates. A bridge function is then introduced to couple the agent layer pair, enabling gradient communication across modalities to facilitate update alignment. The proposed MAIL offers four key advantages: 1) its cross-modal interaction mechanism improves knowledge acquisition from limited data, making it highly effective in low-data scenarios; 2) during inference, MAIL integrates seamlessly into the VLM via reparameterization, preserving inference complexity; 3) extensive experiments validate the superiority of MAIL, which achieves substantial performance gains over data-efficient UCDR methods while requiring significantly fewer training samples; 4) beyond UCDR, MAIL also performs competitively on few-shot classification tasks, underscoring its strong generalization ability. Code. Kaixiang Chen, Pengfei Fang |
NeurIPS | 2 |
| 2025 | DePro: Domain Ensemble using Decoupled Prompts for Universal Cross-Domain RetrievalabstractThis paper investigates the potential of vision-language models (VLMs) in addressing the challenges of universal cross-domain retrieval (UCDR), where queries originate from unseen domains or classes. A common approach to adapting VLMs for downstream tasks involves prompt tuning, which alleviates the computational burden of full fine-tuning. However, this approach often struggles with the domain and semantic shifts inherent in UCDR. To overcome these limitations, we propose a novel prompt decoupling strategy that separates prompts into universal domain prompts (UDPs) and class prompts (CPs). Specifically, UDPs are designed to unify features from both seen and unseen domains into a cohesive universal domain, while CPs are tailored to capture class-specific visual characteristics, enabling robust retrieval across both known and unknown classes. To ensure effective decoupling, we introduce a dedicated decoupling loss that enforces the domain-agnostic nature of CPs. Additionally, we employ a regulation loss to align features from the frozen CLIP domain with those of the universal domain by selectively integrating or excluding UDPs. This mechanism fosters a synergistic domain ensemble effect, enhancing retrieval generalization across diverse domains. Finally, we propose the domain-aware triplet-hard (DaTri) loss to mitigate overfitting by reducing the risk of class collapse. The proposed framework, referred to as Domain Ensemble using Decoupled Prompts (DePro), demonstrates state-of-the-art performance and effectively enhances the model's generalization capacity across unseen domains and classes, as validated through extensive experiments. Code is here. Kaixiang Chen, Pengfei Fang, Hui Xue 0002 |
SIGIR | 2 |
| 2025 | HVQ-VAE: Variational auto-encoder with hyperbolic vector quantization
Shangyu Chen, Pengfei Fang, Mehrtash Harandi, Trung Le 0001, Jianfei Cai 0001, Dinh Q. Phung |
Comput. Vis. Image Underst. | 2 |
| 2025 | OAFN: An efficient open-world audio few-shot learning network for event classification
Fei Chen 0015, Hui Xue 0002, Pengfei Fang |
Knowl. Based Syst. | 3 |
| 2025 | Adaptive indefinite kernels in hyperbolic spaces
Pengfei Fang |
Neural Networks | 1 |
| 2025 | Toward Few-Shot Learning in the Open World: A Review and BeyondabstractHuman intelligence is characterized by our ability to absorb and apply knowledge from the world around us, especially in rapidly acquiring new concepts from minimal examples, underpinned by prior knowledge. Few-shot learning (FSL) aims to mimic this capacity by enabling significant generalizations and transferability. However, traditional FSL frameworks often rely on assumptions of clean, complete, and static data, conditions that are seldom met in real-world environments. Such assumptions falter in the inherently uncertain, incomplete, and dynamic contexts of the open world. This paper presents a comprehensive review of recent advancements designed to adapt FSL to open-world environments. We categorize existing methods into three distinct types of FSL in the open world: those involving varying instances, varying classes, and varying distributions. Each category is discussed in terms of its specific challenges and methods, as well as its strengths and weaknesses. We standardize experimental settings and metric benchmarks across scenarios and provide a comparative analysis of the performance of various methods. In conclusion, we outline potential future research directions for this evolving field. It is our hope that this review will catalyze further development of effective solutions to these complex challenges, thereby advancing the field of artificial intelligence. Hui Xue 0002, Yuexuan An, Yongchun Qin, Wenqian Li 0005, Yixin Wu 0004, Yongjuan Che, Pengfei Fang, Min-Ling Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | On inferring prototypes for multi-label few-shot learning via partial aggregation
Pengfei Fang, Hui Xue 0002 |
Pattern Recognit. | 1 |
| 2025 | Joint state estimation and topology inference for graphical dynamical systems
Pengfei Fang, Wenling Li |
Signal Process. | 1 |
| 2025 | Learning Noisy Few-Shot Classification Without Relying on Pseudo-Noise DataabstractRecently, noisy few-shot learning (NFSL) has been exploring the model robustness to label noise, breaking the limitation of completely accurate labeling in small sample scenarios. Existing NFSL methods directly employ delicately designed pseudo-noise to simulate and adapt to noisy environments. However, determining the optimal combination of pseudo-noise is challenging and improperly configuring pseudo-noise may lead to adverse effects on the training models. To deal with the problems, this letter proposes a novelAdaptiveMultI-viewDenoisingEvaluation (AMIDE) framework, which establishes an adaptive and robust embedding and classifier without relying on pseudo-noise. In the training phase, we design an adaptive label smoothing scheme, where soft labels with learnable smooth coefficients are inferred from data distribution to mitigate overconfident labeling. In the testing stage, we propose a multi-view fused evaluation scheme, where different network layers are treated as distinct views to generate potential clean features and modify prototypes, thereby enhancing the accuracy of evaluation. In this way, the impact of noise is effectively alleviated from two perspectives. Extensive experiments on several few-shot classification benchmarks show the superiority and robustness of our method. Yixin Wu 0004, Hui Xue 0002, Yuexuan An, Pengfei Fang |
IEEE Signal Process. Lett. | 4 |
| 2025 | On Modulating Motion-Aware Visual-Language Representation for Few-Shot Action RecognitionabstractThis paper focuses on few-shot action recognition (FSAR), where the machine is required to understand human actions, with each only seeing a few video samples. Even with only a few explorations, the most cutting-edge methods employ the action textual features, pre-trained by a visual-language model (VLM), as a cue to optimize video prototypes. However, the action textual features used in these methods are generated from a static prompt, causing the network to overlook rich motion cues within videos. To tackle this issue, we propose a novel framework, namely, motion-aware visual-language representation modulation network (MoveNet). The proposed MoveNet utilizes dynamic motion cues within videos to integrate motion-aware textual and visual feature representations, as a way to modulate the video prototypes. In doing so, a long short motion aggregation module (LSMAM) is first proposed to capture diverse motion cues. Having the motion cues at hand, a motion-conditional prompting module (MCPM) utilizes the motion cues as conditions to boost the semantic associations between textual features and action classes. One further develops a motion-guided visual refinement module (MVRM) that adopts motion cues as guidance in enhancing local frame features. The proposed components compensate for each other and contribute to significant performance gains over the FASR task. Thorough experiments on five standard benchmarks demonstrate the effectiveness of the proposed method, considerably outperforming current state-of-the-art methods. Pengfei Fang, Hui Xue 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Leveraging Bilateral Correlations for Multi-Label Few-Shot LearningabstractMulti-label few-shot learning (ML-FSL) refers to the task of tagging previously unseen images with a set of relevant labels, giving a small number of training examples. Modeling the correlations between instances and labels, formulated in the existing methods, allows us to extract more available knowledge from limited examples. However, they simply explore the instance and label correlations with a uniform importance assumption without considering the discrepancy of importance in different instances or labels, making the utilization of instance and label correlations a bottleneck for ML-FSL. To tackle the issue, we propose a unified framework named bilateral correlation reconstruction (BCR) to enable the network to effectively mine underlying instance and label correlations with varying importance information from both instance-to-label and label-to-instance perspectives. Specifically, from the instance-to-label perspective, we refine prototypes per category by reweighting each image with its specific instance-importance degree extracted from the similarity between the instance and the corresponding category. From the label-to-instance perspective, we smooth labels for each image by recovering latent label-importance with considering the integrated topology of all samples in a task. Experimental results on multiple benchmarks validate that BCR could outperform existing ML-FSL methods by large margins. Yuexuan An, Hui Xue 0002, Xingyu Zhao 0002, Ning Xu 0009, Pengfei Fang, Xin Geng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Text Image Inpainting via Global Structure-Guided Diffusion ModelsabstractReal-world text can be damaged by corrosion issues caused by environmental or human factors, which hinder the preservation of the complete styles of texts, e.g., texture and structure. These corrosion issues, such as graffiti signs and incomplete signatures, bring difficulties in understanding the texts, thereby posing significant challenges to downstream applications, e.g., scene text recognition and signature identification. Notably, current inpainting techniques often fail to adequately address this problem and have difficulties restoring accurate text images along with reasonable and consistent styles. Formulating this as an open problem of text image inpainting, this paper aims to build a benchmark to facilitate its study. In doing so, we establish two specific text inpainting datasets which contain scene text images and handwritten text images, respectively. Each of them includes images revamped by real-life and synthetic datasets, featuring pairs of original images, corrupted images, and other assistant information. On top of the datasets, we further develop a novel neural framework, Global Structure-guided Diffusion Model (GSDM), as a potential solution. Leveraging the global structure of the text as a prior, the proposed GSDM develops an efficient diffusion model to recover clean texts. The efficacy of our approach is demonstrated by thorough empirical study, including a substantial boost in both recognition accuracy and image quality. These findings not only highlight the effectiveness of our method but also underscore its potential to enhance the broader field of text image understanding and processing. Code and datasets are available at: https://github.com/blackprotoss/GSDM. Shipeng Zhu, Pengfei Fang, Chenjie Zhu, Zuoyan Zhao, Hui Xue 0002 |
AAAI | 2 |
| 2024 | TSDNet: An Efficient Light Footprint Keyword Spotting Deep Network Base on Tempera Segment NormalizationabstractKeyword spotting plays a pivotal role in smart devices equipped with speech services, serving as the hands-free interface for the activation of pre-specified tasks. It brings the challenge of maintaining high precision with minimal parameters becomes particularly pronounced in devices constrained by limited hardware resources. In this paper, we delve into the research of Temporal Segment Normalization (TSN) and introduce a highly scalable TSN-based ResNet, termed TSDNet, tailored for deployment on target smart devices. Our approach involves meticulous experimental design, wherein we systematically configure experiments to train distinct scaled models. This not only allows us to validate the effectiveness of TSN but also assesses the performance of the proposed network under varying scales. Experimental results, conducted on the Google Speech Commands dataset v1 and v2, demonstrate that the small, medium, and large-scale models consistently attain stateof-the-art results compared to similar methods. Notably, the smallest model achieves a remarkable 97% accuracy on dataset v1 and 97.1% on dataset v2, while operating with a minimum of 7.4K Bytes of parameters. Furthermore, we undertake retraining of ResNet using our proposed training method, achieving new accuracy margins. Fei Chen 0015, Hui Xue 0002, Pengfei Fang |
CSCWD | 3 |
| 2024 | Copula-Nested Spectral Kernel NetworkabstractSpectral Kernel Networks (SKNs) emerge as a promising approach in machine learning, melding solid theoretical foundations of spectral kernels with the representation power of hierarchical architectures. At its core, the spectral density function plays a pivotal role by revealing essential patterns in data distributions, thereby offering deep insights into the underlying framework in real-world tasks. Nevertheless, prevailing designs of spectral density often overlook the intricate interactions within data structures. This phenomenon consequently neglects expanses of the hypothesis space, thus curtailing the performance of SKNs. This paper addresses the issues through a novel approach, the **Co**pula-Nested Spectral **Ke**rnel **Net**work (**CokeNet**). Concretely, we first redefine the spectral density with the form of copulas to enhance the diversity of spectral densities. Next, the specific expression of the copula module is designed to allow the excavation of complex dependence structures. Finally, the unified kernel network is proposed by integrating the corresponding spectral kernel and the copula module. Through rigorous theoretical analysis and experimental verification, CokeNet demonstrates superior performance and significant advancements over SOTA algorithms in the field. Jinyue Tian, Hui Xue 0002, Yanfang Xue, Pengfei Fang |
ICML | 4 |
| 2024 | Stereographic Projection for Embedding Hierarchical Structures in Hyperbolic Space
Shangyu Chen, Xiaohao Yang, Pengfei Fang, Mehrtash Harandi, Dinh Q. Phung, Jianfei Cai 0001 |
ICPR (9) | 3 |
| 2024 | PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-ResolutionabstractScene text image super-resolution (STISR) aims at simultaneously increasing the resolution and readability of low-resolution scene text images, thus boosting the performance of the downstream recognition task. Two factors in scene text images, visual structure and semantic information, affect the recognition performance significantly. To mitigate the effects from these factors, this paper proposes a Prior-Enhanced Attention Network (PEAN). Specifically, an attention-based modulation module is leveraged to understand scene text images by neatly perceiving the local and global dependence of images, despite the shape of the text. Meanwhile, a diffusion-based module is developed to enhance the text prior, hence offering better guidance for the SR network to generate SR images with higher semantic accuracy. Additionally, a multi-task learning paradigm is employed to optimize the network, enabling the model to generate legible SR images. As a result, PEAN establishes new SOTA results on the TextZoom benchmark. Experiments are also conducted to analyze the importance of the enhanced text prior as a means of improving the performance of the SR network. Code is available at https://github.com/jdfxzzy/PEAN. Zuoyan Zhao, Hui Xue 0002, Pengfei Fang, Shipeng Zhu |
ACM Multimedia | 3 |
| 2024 | Reproducing the Past: A Dataset for Benchmarking Inscription RestorationabstractInscriptions on ancient steles, as carriers of culture, encapsulate the humanistic thoughts and aesthetic values of our ancestors. However, these relics often deteriorate due to environmental and human factors, resulting in significant information loss. Since the advent of inscription rubbing technology over a millennium ago, archaeologists and epigraphers have devoted immense effort to manually restoring these cultural imprints, endeavoring to unlock the storied past within each rubbing. This paper approaches this challenge as a multi-modal task, aiming to establish a novel benchmark for the inscription restoration from rubbings. In doing so, we construct the Chinese Inscription Rubbing Image (CIRI) dataset, which includes a wide variety of real inscription rubbing images characterized by diverse calligraphy styles, intricate character structures, and complex degradation forms. Furthermore, we develop a synthesis approach to generate "intact-degraded'' paired data, mirroring real-world degradation faithfully. On top of the datasets, we propose a baseline framework that achieves visual consistency and textual integrity through global and local diffusion-based restoration processes and explicit incorporation of domain knowledge. Comprehensive evaluations confirm the effectiveness of our pipeline, demonstrating significant improvements in visual presentation and textual integrity. The project is available at: https://github.com/blackprotoss/CIRI. Shipeng Zhu, Hui Xue 0002, Na Nie, Chenjie Zhu, Haiyue Liu, Pengfei Fang |
ACM Multimedia | 6 |
| 2024 | Dynamic functional connections analysis with spectral learning for brain disorder detection
Yanfang Xue, Hui Xue 0002, Pengfei Fang, Shipeng Zhu, Lishan Qiao, Yuexuan An |
Artif. Intell. Medicine | 3 |
| 2024 | Multi-Scale Explicit Matching and Mutual Subject Teacher Learning for Generalizable Person Re-IdentificationabstractDomain generalization in person re-identification (DG-ReID) stands out as the most challenging task and practically important branch in the ReID field, which enables the direct deployment of pre-trained models in unseen and real scenarios. Recent works have made significant efforts in this task via the image-matching paradigm, which searches for the local correspondences in the feature maps. A common practice of employing pixel-wise matching is typically used to ensure efficient matching. This, however, makes the matching susceptible to deviations caused by identity-irrelevant pixel features. On the other hand, patch-wise matching also demonstrates that it will disregard the spatial orientation of pedestrians and amplify the impact of noise. To address the mentioned issues, this paper proposes the Multi-Scale Query-Adaptive Convolution (QAConv-MS) framework, which encodes patches in the feature maps to pixels using template kernels of various scales. This enables the matching process to enjoy broader receptive fields and robustness to orientations and noises. To stabilize the matching process and facilitate the independent learning of each sub-kernel within the template kernels to capture diverse local patterns, we propose the OrthoGonal Norm (OGNorm), which consists of two orthogonal normalizations. We also present Mutual Subject Teacher Learning (MSTL) to address the potential issues of overconfidence and overfitting in the model. MSTL allows two models to individually select the most challenging data for training, resulting in more dependable soft labels that can provide mutual supervision. Extensive experiments conducted in both single-source and multi-source setups offer compelling evidence of our framework’s generalization and competitiveness. Kaixiang Chen, Pengfei Fang, Liyan Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Curved Geometric Networks for Visual Anomaly RecognitionabstractLearning a latent embedding to understand the underlying nature of data distribution is often formulated in Euclidean spaces with zero curvature. However, the success of the geometry constraints, posed in the embedding space, indicates that curved spaces might encode more structural information, leading to better discriminative power and hence richer representations. In this work, we investigate the benefits of the curved space for analyzing anomalous, open-set, or out-of-distribution (OOD) objects in data. This is achieved by considering embeddings via three geometry constraints, namely, spherical geometry (with positive curvature), hyperbolic geometry (with negative curvature), or mixed geometry (with both positive and negative curvatures). Three geometric constraints can be chosen interchangeably in a unified design, given the task at hand. Tailored for the embeddings in the curved space, we also formulate functions to compute the anomaly score. Two types of geometric modules (i.e., geometric-in-one (GiO) and geometric-in-two (GiT) models) are proposed to plug in the original Euclidean classifier, and anomaly scores are computed from the curved embeddings. We evaluate the resulting designs under a diverse set of visual recognition scenarios, including image detection (multiclass OOD detection and one-class anomaly detection) and segmentation (multiclass anomaly segmentation and one-class anomaly segmentation). The empirical results show the effectiveness of our proposal through consistent improvement over various scenarios. The code is made available at https://github.com/JHome1/GiO-GiT. Pengfei Fang, Weihao Li 0005, Junlin Han, Lars Petersson, Mehrtash Harandi |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Asymmetric Dual-Decoder U-Net for Joint Rain and Haze RemovalabstractThis work studies the multi-weather restoration problem. In real-life scenarios, rain and haze, two often co-occurring common weather phenomena, can greatly degrade the clarity and quality of the scene images, leading to a performance drop in the visual applications, such as autonomous driving. However, jointly removing the rain and haze in scene images is ill-posed and challenging, where the existence of haze and rain and the change of atmosphere light, can both degrade the scene information. Current methods focus on the contamination removal part, thus ignoring the restoration of the scene information affected by the change of atmospheric light. We propose a novel deep neural network, named Asymmetric Dual-decoder U-Net (ADU-Net), to address the aforementioned challenge. The ADU-Net produces both the contamination residual and the scene residual to efficiently remove the contamination while preserving the fidelity of the scene information. Extensive experiments show our work outperforms the existing state-of-the-art methods by a considerable margin in both synthetic data and real-world data benchmarks, including RainCityscapes, BID Rain, and SPA-Data. For instance, we improve the state-of-the-art PSNR value by 2.26/4.57 on the RainCityscapes/SPA-Data, respectively. Codes will be made available freely to the research community. Yuan Feng 0002, Yaojun Hu, Pengfei Fang, Sheng Liu 0002, Yanhong Yang, Shengyong Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Improving Scene Text Retrieval via Stylized Middle ModalityabstractScene text retrieval addresses the challenge of localizing and searching for all text instances within scene images based on a query text. This cross-modal task has significant applications in various domains, such as intelligent transportation systems and social media analysis. In practice, ensuring consistency of the same content between two modalities is crucial in improving retrieval accuracy. This article addresses the issue by introducing a stylized middle modality, which fuses the graphical query text with the style of the extracted text proposal. To this end, we propose a stylized middle modality learning (SM 2 L) framework. The proposed stylized middle modality enables the network to jointly enforce constraints on visual feature coherence and text semantic feature consistency in the optimization phase, thereby minimizing the modality gap in the retrieval space. This brings in two major advantages: (1) SM 2 L will pave the way to seamlessly benefit the scene text retrieval and (2) the proposed learning paradigm enables the machine to avoid adding redundant computing resources in the inference phase. Substantial experiments demonstrate that the proposed method outperforms the state-of-the-art retrieval performance considerably. Shipeng Zhu, Pengfei Fang, Hui Xue 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | GOSS: towards generalized open-set semantic segmentationabstractAbstract In this paper, we extend Open-set Semantic Segmentation (OSS) into a new image segmentation task called Generalized Open-set Semantic Segmentation (GOSS). Previously, with well-known OSS, the intelligent agents only detect unknown regions without further processing, limiting their perception capacity of the environment. It stands to reason that further analysis of the detected unknown pixels would be beneficial for agents’ decision-making. Therefore, we propose GOSS, which holistically unifies the abilities of two well-defined segmentation tasks, i.e. OSS and generic segmentation. Specifically, GOSS classifies pixels as belonging to known classes, and clusters (or groups) of pixels of unknown class are labelled as such. We propose a metric that balances the pixel classification and clustering aspects to evaluate this newly expanded task. Moreover, we build benchmark tests on existing datasets and propose neural architectures as baselines. Our experiments on multiple benchmarks demonstrate the effectiveness of our baselines. Code is made available at https://github.com/JHome1/GOSS_Segmentor . Weihao Li 0005, Junlin Han, Jiyang Zheng, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
Vis. Comput. | 5 |
| 2024 | Publisher Correction: GOSS: towards generalized open-set semantic segmentation
Weihao Li 0005, Junlin Han, Jiyang Zheng, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
Vis. Comput. | 5 |
| 2023 | Improving Scene Text Image Super-resolution via Dual Prior Modulation NetworkabstractScene text image super-resolution (STISR) aims to simultaneously increase the resolution and legibility of the text images, and the resulting images will significantly affect the performance of downstream tasks. Although numerous progress has been made, existing approaches raise two crucial issues: (1) They neglect the global structure of the text, which bounds the semantic determinism of the scene text. (2) The priors, e.g., text prior or stroke prior, employed in existing works, are extracted from pre-trained text recognizers. That said, such priors suffer from the domain gap including low resolution and blurriness caused by poor imaging conditions, leading to incorrect guidance. Our work addresses these gaps and proposes a plug-and-play module dubbed Dual Prior Modulation Network (DPMN), which leverages dual image-level priors to bring performance gain over existing approaches. Specifically, two types of prior-guided refinement modules, each using the text mask or graphic recognition result of the low-quality SR image from the preceding layer, are designed to improve the structural clarity and semantic accuracy of the text, respectively. The following attention mechanism hence modulates two quality-enhanced images to attain a superior SR result. Extensive experiments validate that our method improves the image quality and boosts the performance of downstream tasks over five typical approaches on the benchmark. Substantial visualizations and ablation studies demonstrate the advantages of the proposed DPMN. Code is available at: https://github.com/jdfxzzy/DPMN. Shipeng Zhu, Zuoyan Zhao, Pengfei Fang, Hui Xue 0002 |
AAAI | 3 |
| 2023 | Hyperbolic Audio-visual Zero-shot LearningabstractAudio-visual zero-shot learning aims to classify samples consisting of a pair of corresponding audio and video sequences from classes that are not present during training. An analysis of the audio-visual data reveals a large degree of hyperbolicity, indicating the potential benefit of using a hyperbolic transformation to achieve curvature-aware geometric learning, with the aim of exploring more complex hierarchical data structures for this task. The proposed approach employs a novel loss function that incorporates cross-modality alignment between video and audio features in the hyperbolic space. Additionally, we explore the use of multiple adaptive curvatures for hyperbolic projections. The experimental results on this very challenging task demonstrate that our proposed hyperbolic approach for zero-shot learning outperforms the SOTA method on three datasets: VGGSound-GZSL, UCF-GZSL, and ActivityNet-GZSL achieving a harmonic mean (HM) improvement of around 3.0%, 7.0%, and 5.3%, respectively. Zeeshan Hayder, Junlin Han, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
ICCV | 4 |
| 2023 | Expanding the Hyperbolic Kernels: A Curvature-aware Isometric Embedding ViewabstractModeling data relation as a hierarchical structure has proven beneficial for many learning scenarios, and the hyperbolic space, with negative curvature, can encode such data hierarchy without distortion. Several recent studies also show that the representation power of the hyperbolic space can be further improved by endowing the kernel methods. Unfortunately, the known kernel methods, developed in hyperbolic space, are limited by the adaptation capacity or distortion issues. This paper addresses the issues through a novel embedding function. To this end, we propose a curvature-aware isometric embedding, which establishes an isometry from the Poincar\'e model to a special reproducing kernel Hilbert space (RKHS). Then we can further define a series of kernels on this RKHS, including several positive definite kernels and an indefinite kernel. Thorough experiments are conducted to demonstrate the superiority of our proposals over existing-known hyperbolic and Euclidean kernels in various learning tasks, e.g., graph learning and zero-shot learning. Meimei Yang, Pengfei Fang, Hui Xue 0002 |
IJCAI | 2 |
| 2023 | CosNet: A Generalized Spectral Kernel NetworkabstractComplex-valued representation exists inherently in the time-sequential data that can be derived from the integration of harmonic waves. The non-stationary spectral kernel, realizing a complex-valued feature mapping, has shown its potential to analyze the time-varying statistical characteristics of the time-sequential data, as a result of the modeling frequency parameters. However, most existing spectral kernel-based methods eliminate the imaginary part, thereby limiting the representation power of the spectral kernel. To tackle this issue, we propose a generalized spectral kernel network, namely, \underline{Co}mplex-valued \underline{s}pectral kernel \underline{Net}work (CosNet), which includes spectral kernel mapping generalization (SKMG) module and complex-valued spectral kernel embedding (CSKE) module. Concretely, the SKMG module is devised to generalize the spectral kernel mapping in the real number domain to the complex number domain, recovering the inherent complex-valued representation for the real-valued data. Then a following CSKE module is further developed to combine the complex-valued spectral kernels and neural networks to effectively capture long-range or periodic relations of the data. Along with the CosNet, we study the effect of the complex-valued spectral kernel mapping via theoretically analyzing the bound of covering number and generalization error. Extensive experiments demonstrate that CosNet performs better than the mainstream kernel methods and complex-valued neural networks. Yanfang Xue, Pengfei Fang, Jinyue Tian, Shipeng Zhu, Hui Xue 0002 |
NeurIPS | 2 |
| 2023 | On learning distribution alignment for video-based visible-infrared person re-identification
Pengfei Fang, Yaojun Hu, Shipeng Zhu, Hui Xue 0002 |
Comput. Vis. Image Underst. | 1 |
| 2023 | Poincaré Kernels for Hyperbolic Representations
Pengfei Fang, Mehrtash Harandi, Zhen-Zhong Lan, Lars Petersson |
Int. J. Comput. Vis. | 1 |
| 2023 | A DNA image encryption based on a new hyperchaotic system
Yuanyuan Hui, Han Liu 0007, Pengfei Fang |
Multim. Tools Appl. | 3 |
| 2023 | Beyond a strong baseline: cross-modality contrastive learning for visible-infrared person re-identification
Pengfei Fang, Zhen-Zhong Lan |
Mach. Vis. Appl. | 1 |
| 2023 | A survey of image encryption algorithms based on chaotic system
Pengfei Fang, Han Liu 0007, Chengmao Wu 0001, Min Liu 0028 |
Vis. Comput. | 1 |
| 2022 | Adaptive Poincaré Point to Set Distance for Few-Shot ClassificationabstractLearning and generalizing from limited examples, i.e., few-shot learning, is of core importance to many real-world vision applications. A principal way of achieving few-shot learning is to realize an embedding where samples from different classes are distinctive. Recent studies suggest that embedding via hyperbolic geometry enjoys low distortion for hierarchical and structured data, making it suitable for few-shot learning. In this paper, we propose to learn a context-aware hyperbolic metric to characterize the distance between a point and a set associated with a learned set to set distance. To this end, we formulate the metric as a weighted sum on the tangent bundle of the hyperbolic space and develop a mechanism to obtain the weights adaptively, based on the constellation of the points. This not only makes the metric local but also dependent on the task in hand, meaning that the metric will adapt depending on the samples that it compares. We empirically show that such metric yields robustness in the presence of outliers and achieves a tangible improvement over baseline models. This includes the state-of-the-art results on five popular few-shot classification benchmarks, namely mini-ImageNet, tiered-ImageNet, Caltech-UCSD Birds-200-2011(CUB), CIFAR-FS, and FC100. Rongkai Ma, Pengfei Fang, Tom Drummond, Mehrtash Harandi |
AAAI | 2 |
| 2022 | Human-in-the-loop Robotic Grasping Using BERT Scene RepresentationabstractCurrent NLP techniques have been greatly applied in different domains. In this paper, we propose a human-in-the-loop framework for robotic grasping in cluttered scenes, investigating a language interface to the grasping process, which allows the user to intervene by natural language commands. This framework is constructed on a state-of-the-art grasping baseline, where we substitute a scene-graph representation with a text representation of the scene using BERT. Experiments on both simulation and physical robot show that the proposed method outperforms conventional object-agnostic and scene-graph based methods in the literature. In addition, we find that with human intervention, performance can be significantly improved. Our dataset and code are available on our project website https://sites.google.com/view/hitl-grasping-bert. Yaoxian Song, Penglei Sun, Pengfei Fang, Linyi Yang, Yanghua Xiao, Yue Zhang 0004 |
COLING | 3 |
| 2022 | Blind Image Decomposition
Junlin Han, Weihao Li 0005, Pengfei Fang, Chunyi Sun, Mohammad Ali Armin, Lars Petersson, Hongdong Li |
ECCV (18) | 3 |
| 2022 | Learning Instance and Task-Aware Dynamic Kernels for Few-Shot Learning
Rongkai Ma, Pengfei Fang, Gil Avraham, Tianyu Zhu 0001, Tom Drummond, Mehrtash Harandi |
ECCV (20) | 2 |
| 2022 | You Only Cut Once: Boosting Data Augmentation with a Single CutabstractWe present You Only Cut Once (YOCO) for performing data augmentations. YOCO cuts one image into two pieces and performs data augmentations individually within each piece. Applying YOCO improves the diversity of the augmentation per sample and encourages neural networks to recognize objects from partial information. YOCO enjoys the properties of parameter-free, easy usage, and boosting almost all augmentations for free. Thorough experiments are conducted to evaluate its effectiveness. We first demonstrate that YOCO can be seamlessly applied to varying data augmentations, neural network architectures, and brings performance gains on CIFAR and ImageNet classification tasks, sometimes surpassing conventional image-level augmentation by large margins. Moreover, we show YOCO benefits contrastive pre-training toward a more powerful representation that can be better transferred to multiple downstream tasks. Finally, we study a number of variants of YOCO and empirically analyze the performance for respective settings. Junlin Han, Pengfei Fang, Weihao Li 0005, Mohammad Ali Armin, Ian D. Reid 0001, Lars Petersson, Hongdong Li |
ICML | 2 |
| 2022 | A block image encryption algorithm based on a hyperchaotic system and generative adversarial networks
Pengfei Fang, Han Liu 0007, Chengmao Wu 0001, Min Liu 0028 |
Multim. Tools Appl. | 1 |
| 2022 | Attention in Attention Networks for Person RetrievalabstractThis paper generalizes the Attention in Attention (AiA) mechanism, in P. Fang et al., 2019 by employing explicit mapping in reproducing kernel Hilbert spaces to generate attention values of the input feature map. The AiA mechanism models the capacity of building inter-dependencies among the local and global features by the interaction of inner and outer attention modules. Besides a vanilla AiA module, termed linear attention with AiA, two non-linear counterparts, namely, second-order polynomial attention and Gaussian attention, are also proposed to utilize the non-linear properties of the input features explicitly, via the second-order polynomial kernel and Gaussian kernel approximation. The deep convolutional neural network, equipped with the proposed AiA blocks, is referred to as Attention in Attention Network (AiA-Net). The AiA-Net learns to extract a discriminative pedestrian representation, which combines complementary person appearance and corresponding part features. Extensive ablation studies verify the effectiveness of the AiA mechanism and the use of non-linear features hidden in the feature map for attention design. Furthermore, our approach outperforms current state-of-the-art by a considerable margin across a number of benchmarks. In addition, state-of-the-art performance is also achieved in the video person retrieval task with the assistance of the proposed AiA blocks. Pengfei Fang, Jieming Zhou, Soumava Kumar Roy, Pan Ji, Lars Petersson, Mehrtash Harandi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | TRL: Transformer based refinement learning for hybrid-supervised semantic segmentation
Pengfei Fang, Yan Yan 0001, Yang Lu 0009, Hanzi Wang |
Pattern Recognit. Lett. | 2 |
| 2022 | TSGB: Target-Selective Gradient Backprop for Probing CNN Visual SaliencyabstractThe explanation for deep neural networks has drawn extensive attention in the deep learning community over the past few years. In this work, we study the visual saliency, a.k.a. visual explanation, to interpret convolutional neural networks. Compared to iteration based saliency methods, single backward pass based saliency methods benefit from faster speed, and they are widely used in downstream visual tasks. Thus, we focus on single backward pass based methods. However, existing methods in this category struggle to successfully produce fine-grained saliency maps concentrating on specific target classes. That said, producing faithful saliency maps satisfying both target-selectiveness and fine-grainedness using a single backward pass is a challenging problem in the field. To mitigate this problem, we revisit the gradient flow inside the network, and find that the entangled semantics and original weights may disturb the propagation of target-relevant saliency. Inspired by those observations, we propose a novel visual saliency method, termed Target-Selective Gradient Backprop (TSGB), which leverages rectification operations to effectively emphasize target classes and further efficiently propagate the saliency to the image space, thereby generating target-selective and fine-grained saliency maps. The proposed TSGB consists of two components, namely, TSGB-Conv and TSGB-FC, which rectify the gradients for convolutional layers and fully-connected layers, respectively. Extensive qualitative and quantitative experiments on the ImageNet and Pascal VOC datasets show that the proposed method achieves more accurate and reliable results than the other competitive methods. Code is available at https://github.com/123fxdx/CNNvisualizationTSGB. Pengfei Fang, Liao Zhang, Chunhua Shen, Hanzi Wang |
IEEE Trans. Image Process. | 2 |
| 2021 | Semantic-Aware Knowledge Distillation for Few-Shot Class-Incremental LearningabstractFew-shot class incremental learning (FSCIL) portrays the problem of learning new concepts gradually, where only a few examples per concept are available to the learner. Due to the limited number of examples for training, the techniques developed for standard incremental learning cannot be applied verbatim to FSCIL. In this work, we introduce a distillation algorithm to address the problem of FSCIL and propose to make use of semantic information during training. To this end, we make use of word embeddings as semantic information which is cheap to obtain and which facilitate the distillation process. Furthermore, we propose a method based on an attention mechanism on multiple parallel embeddings of visual data to align visual and semantic vectors, which reduces issues related to catastrophic forgetting. Via experiments on MiniImageNet, CUB200, and CIFAR100 dataset, we establish new state-of-the-art results by outperforming existing approaches. Ali Cheraghian, Shafin Rahman, Pengfei Fang, Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi |
CVPR | 3 |
| 2021 | Reinforced Attention for Few-Shot Learning and BeyondabstractFew-shot learning aims to correctly recognize query samples from unseen classes given a limited number of support samples, often by relying on global embeddings of images. In this paper, we propose to equip the backbone network with an attention agent, which is trained by reinforcement learning. The policy gradient algorithm is employed to train the agent towards adaptively localizing the representative regions on feature maps over time. We further design a reward function based on the prediction of the held-out data, thus helping the attention mechanism to generalize better across the unseen classes. The extensive experiments show, with the help of the reinforced attention, that our embedding network has the capability to progressively generate a more discriminative representation in few-shot learning. Moreover, experiments on the task of image classification also show the effectiveness of the proposed design. Pengfei Fang, Weihao Li 0005, Tong Zhang 0023, Christian Simon, Mehrtash Harandi, Lars Petersson |
CVPR | 2 |
| 2021 | Synthesized Feature based Few-Shot Class-Incremental Learning on a Mixture of SubspacesabstractFew-shot class incremental learning (FSCIL) aims to incrementally add sets of novel classes to a well-trained base model in multiple training sessions with the restriction that only a few novel instances are available per class. While learning novel classes, FSCIL methods gradually forget base (old) class training and overfit to a few novel class samples. Existing approaches have addressed this problem by computing the class prototypes from the visual or semantic word vector domain. In this paper, we propose addressing this problem using a mixture of subspaces. Subspaces define the cluster structure of the visual domain and help to describe the visual and semantic domain considering the overall distribution of the data. Additionally, we propose to employ a variational autoencoder (VAE) to generate synthesized visual samples for augmenting pseudo-feature while learning novel classes incrementally. The combined effect of the mixture of subspaces and synthesized features reduces the forgetting and overfitting problem of FSCIL. Extensive experiments on three image classification datasets show that our proposed method achieves competitive results compared to state-of-the-art methods. Ali Cheraghian, Shafin Rahman, Sameera Ramasinghe, Pengfei Fang, Christian Simon, Lars Petersson, Mehrtash Harandi |
ICCV | 4 |
| 2021 | Kernel Methods in Hyperbolic SpacesabstractEmbedding data in hyperbolic spaces has proven beneficial for many advanced machine learning applications such as image classification and word embeddings. However, working in hyperbolic spaces is not without difficulties as a result of its curved geometry (e.g., computing the Frechet mean of a set of points requires an iterative algorithm). Furthermore, in Euclidean spaces, one can resort to kernel machines that not only enjoy rich theoretical properties but that can also lead to superior representational power (e.g., infinite-width neural networks). In this paper, we introduce positive definite kernel functions for hyperbolic spaces. This brings in two major advantages, 1. kernelization will pave the way to seamlessly benefit from kernel machines in conjunction with hyperbolic embeddings, and 2. the rich structure of the Hilbert spaces associated with kernel machines enables us to simplify various operations involving hyperbolic data. That said, identifying valid kernel functions on curved spaces is not straightforward and is indeed considered an open problem in the learning community. Our work addresses this gap and develops several valid positive definite kernels in hyperbolic spaces, including the universal ones (e.g., RBF). We comprehensively study the proposed kernels on a variety of challenging tasks including few-shot learning, zero-shot learning, person reidentification and knowledge distillation, showing the superiority of the kernelization for hyperbolic representations. Pengfei Fang, Mehrtash Harandi, Lars Petersson |
ICCV | 1 |
| 2021 | Set Augmented Triplet Loss for Video Person Re-IdentificationabstractModern video person re-identification (re-ID) machines are often trained using a metric learning approach, supervised by a triplet loss. The triplet loss used in video re-ID is usually based on so-called clip features, each aggregated from a few frame features. In this paper, we propose to model the video clip as a set and instead study the distance between sets in the corresponding triplet loss. In contrast to the distance between clip representations, the distance between clip sets considers the pair-wise similarity of each element (i.e., frame representation) between two sets. This allows the network to directly optimize the feature representation at a frame level. Apart from the commonly-used set distance metrics (e.g., ordinary distance and Hausdorff distance), we further propose a hybrid distance metric, tailored for the set-aware triplet loss. Also, we propose a hard positive set construction strategy using the learned class prototypes in a batch. Our proposed method achieves state-of-the-art results across several standard benchmarks, demonstrating the advantages of the proposed method. Pengfei Fang, Pan Ji, Lars Petersson, Mehrtash Harandi |
WACV | 1 |
| 2020 | Channel Recurrent Attention Networks for Video Pedestrian Retrieval
Pengfei Fang, Pan Ji, Jieming Zhou, Lars Petersson, Mehrtash Harandi |
ACCV (6) | 1 |
| 2020 | Cross-Correlated Attention Networks for Person Re-Identification
Jieming Zhou, Soumava Kumar Roy, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
Image Vis. Comput. | 3 |
| 2019 | Bilinear Attention Networks for Person RetrievalabstractThis paper investigates a novel Bilinear attention (Bi-attention) block, which discovers and uses second order statistical information in an input feature map, for the purpose of person retrieval. The Bi-attention block uses bilinear pooling to model the local pairwise feature interactions along each channel, while preserving the spatial structural information. We propose an Attention in Attention (AiA) mechanism to build inter-dependency among the second order local and global features with the intent to make better use of, or pay more attention to, such higher order statistical relationships. The proposed network, equipped with the proposed Bi-attention is referred to as Bilinear ATtention network (BAT-net). Our approach outperforms current state-of-the-art by a considerable margin across the standard benchmark datasets (e.g., CUHK03, Market-1501, DukeMTMC-reID and MSMT17). Pengfei Fang, Jieming Zhou, Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi |
ICCV | 1 |
| 2019 | Recovering 6D object pose from RGB indoor image based on two-stage detection network with multi-task loss
Fuchang Liu, Pengfei Fang, Zhengwei Yao, Ran Fan, Weiguo Sheng 0001, Huansong Yang |
Neurocomputing | 2 |