VLDB 2026 Research / reviewers in the wild / expert
Huafeng Li 0001
dblp:84/11066-1
· DBLP profile ↗
88ranked-venue papers
33as first author
74since 2021 · last 2026
0000-0003-2462-6174ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 13 first-author · 35 since 2021Artificial intelligence and machine learning · 39 · 16 first-author · 34 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 4 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Dynamic Dehazing via Instruction-Driven and Task-Feedback Closed-Loop Optimization for Diverse Downstream Task AdaptationabstractIn real-world vision systems, haze removal is required not only to enhance image visibility but also to meet the specific needs of diverse downstream tasks. To address this challenge, we propose a novel adaptive dynamic dehazing framework that incorporates a closed-loop optimization mechanism. It enables feedback-driven refinement based on downstream task performance and user instruction–guided adjustment during inference, allowing the model to satisfy the specific requirements of multiple downstream tasks without retraining. Technically, our framework integrates two complementary and innovative mechanisms: (1) a task feedback loop that dynamically modulates dehazing outputs based on performance across multiple downstream tasks, and (2) a text instruction interface that allows users to specify high-level task preferences. This dual-guidance strategy enables the model to adapt its dehazing behavior after training, tailoring outputs in real time to the evolving needs of multiple tasks. Extensive experiments across various vision tasks demonstrate the strong effectiveness, robustness, and generalizability of our approach. These results establish a new paradigm for interactive, task-adaptive dehazing that actively collaborates with downstream applications. Shuaitian Song, Huafeng Li 0001, Shujuan Wang, Yu Liu 0023 |
AAAI | 3 |
| 2026 | Hierarchical Prompt Learning for Image- and Text-Based Person Re-IdentificationabstractPerson re-identification (ReID) aims to retrieve target pedestrian images given either visual queries (image-to-image, I2I) or textual descriptions (text-to-image, T2I). Although both tasks share a common retrieval objective, they pose distinct challenges: I2I emphasizes discriminative identity learning, while T2I requires accurate cross-modal semantic alignment. Existing methods often treat these tasks separately, which may lead to representation entanglement and suboptimal performance. To address this, we propose a unified framework named Hierarchical Prompt Learning (HPL), which leverages task-aware prompt modeling to jointly optimize both tasks. Specifically, we first introduce a Task-Routed Transformer, which incorporates dual classification tokens into a shared visual encoder to route features for I2I and T2I branches respectively. On top of this, we develop a hierarchical prompt generation scheme that integrates identity-level learnable tokens with instance-level pseudo-text tokens. These pseudo-tokens are derived from image or text features via modality-specific inversion networks, injecting fine-grained, instance-specific semantics into the prompts. Furthermore, we propose a Cross-Modal Prompt Regularization strategy to enforce semantic alignment in the prompt token space, ensuring that pseudo-prompts preserve source-modality characteristics while enhancing cross-modal transferability. Extensive experiments on multiple ReID benchmarks validate the effectiveness of our method, achieving state-of-the-art performance on both I2I and T2I tasks. Linhan Zhou, Neng Dong, Yonghang Tai, Huafeng Li 0001 |
AAAI | 6 |
| 2026 | Domain-invariant knowledge ensemble distillation for domain generalization intelligent fault diagnosis
Kaixiong Xu, Youqiang Hu, Huafeng Li 0001, Hongying Yan, Yuqiang Liu, Yi Chai 0003, Ke Zhang 0006 |
Adv. Eng. Informatics | 3 |
| 2026 | Complex-order Darwinian particle swarm optimization
Huafeng Li 0001, António M. Lopes 0001, YangQuan Chen, Yi Chai 0003 |
Expert Syst. Appl. | 3 |
| 2026 | FlexiD-Fuse: Flexible number of inputs multi-modal medical image fusion based on diffusion model
Yushen Xu, Xiaosong Li 0004, Xiaoqi Cheng, Huafeng Li 0001, Haishu Tan |
Expert Syst. Appl. | 5 |
| 2026 | Physical Regularization Loss: Integrating Physical Knowledge to Image Segmentation
Huafeng Li 0001, Guanqiu Qi, Baisen Cong, Yunpeng Gong, Zhiqin Zhu |
Int. J. Comput. Vis. | 3 |
| 2026 | Multimodal artificial intelligence for disease diagnosis: Advances, applications, and challenges
Shaozhe Wang, Fan Zhang 0070, Yu Liu 0023, Huafeng Li 0001, Junyu Dong, David Zhang 0001 |
Pattern Recognit. | 6 |
| 2026 | Infrared-assisted single-stage framework for joint restoration and fusion of visible and infrared images under hazy conditions
Huafeng Li 0001, Jiaqi Fang, Yu Liu 0023 |
Pattern Recognit. | 1 |
| 2026 | FlexiSR-Diff: Flexible diffusion for multi-modal medical image fusion & super-resolution
Yushen Xu, Xiaosong Li 0004, Yang Liu 0335, Tao Ye 0002, Huafeng Li 0001 |
Pattern Recognit. | 7 |
| 2026 | Modalities collaboration and granularities interaction for fine-grained sketch-based image retrieval
Junchao Ge, Jiaman Ding, Neng Dong, Shaojie Qiao, Zhengtao Yu 0001, Huafeng Li 0001 |
Pattern Recognit. | 7 |
| 2026 | A Survey on lightweight technology of neural networks for medical image segmentationabstractRecent advances in medical image segmentation have significantly improved segmentation accuracy. Nevertheless, the clinical deployment of large-scale segmentation networks remains constrained by challenges such as excessive parameter counts, complex architectures, and limited adaptability to diverse deployment environments. The absence of lightweight design further restricts their integration into resource-limited edge devices. To address these barriers, lightweight strategies have emerged as an effective solution. Structural optimization simplifies network architectures to reduce computational costs, while model compression techniques shrink model size without sacrificing performance. At the same time, hardware-level acceleration provides additional support for efficient inference in real-world scenarios. This review systematically summarizes recent lightweight methods for medical image segmentation from both software and hardware perspectives. Representative algorithmic approaches are highlighted, including pruning, quantization, knowledge distillation, and efficient network architectures, along with hardware-aware optimization strategies tailored for edge deployment. Moreover, we explored the mainstream approach of integrating large-scale models with lightweight technologies to achieve the optimal balance between segmentation accuracy and computational efficiency. Finally, current limitations and potential research directions are outlined to promote the translation of lightweight segmentation models into routine clinical workflows. By providing a structured reference, this review aims to support researchers and practitioners in advancing the efficient and practical application of medical image segmentation in clinical environments. Zhiqin Zhu, Hanchen Wang 0005, Guanqiu Qi, Neal Mazur, Yu Liu 0023, Huafeng Li 0001, Baisen Cong, Litao Bai |
Pattern Recognit. | 7 |
| 2026 | Spatial-Temporal High-Frequency Learning for Video-Based Visible-Infrared Person Re-IdentificationabstractVideo-based Visible-Infrared Person Re-Identification (VVI-ReID) aims to learn consistent person feature representations across video sequences in different modalities. Existing methods that use an intermediate modality to bridge the gap between visible (RGB) and infrared (IR) sequences tend to be limited by high construction costs, loss of high-frequency details, and lack of temporal cues. Moreover, they typically focus on refining global representation using high-level features, neglecting the enhancement of local details through low-level features. To address these challenges, we propose the novel Spatial-Temporal High-Frequency Learning (STHF) framework, which constructs an appropriate intermediate modality for the VVI-ReID task and alleviates the modality gap via hierarchical feature enhancement. Specifically, we introduce the Spatial-Temporal High-Pass Filter (ST-HPF), which filters out spatial-temporal Low-Frequency Components (LFC), preserving high-frequency details to construct an intermediate modality at the sequence level. We then enhance the local details with low-level features through the Shallow Detail Compensation (SDC) module, which reduces local noise interference. Finally, the Deep Semantic Refinement (DSR) module refines the global representation by modeling spatial-temporal high-frequency semantic associations using high-level features. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches on the publicly available HITSZ-VCM and BUPTCampus datasets. The code is available at https://github.com/TSC95720/STHF. Sichen Tao, Neng Dong, Fan Li 0006, Huafeng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text PerceptionabstractMulti-modality image fusion (MMIF) in adverse weather aims to address the loss of visual information caused by weather-related degradations, providing clearer scene representations. Although a few studies have attempted to incorporate textual information to improve semantic perception, they often lack effective categorization and thorough analysis of textual content. To address these limitations, we propose AWM-Fuse, a unified fusion framework that handles diverse weather degradations via global and local text perception with shared parameters. In particular, a global text perception module leverages BLIP-generated captions to extract overall scene features and identify primary degradation types, thus promoting generalization across various adverse weather conditions. Complementing this, the local module employs detailed scene descriptions produced by ChatGPT to concentrate on specific degradation effects through concrete textual cues, enabling the recovery of subtle details. Furthermore, textual descriptions are used to constrain the generation of fused images, effectively steering the network learning process toward better alignment with semantic labels, thereby promoting the learning of more meaningful visual features. To facilitate text-guided fusion under adverse weather, we construct AWMM-Text, a large-scale benchmark providing paired global and local annotations for multi-modality image pairs. Extensive experiments demonstrate that AWM-Fuse consistently outperforms state-of-the-art methods under complex weather conditions and on multiple downstream tasks. Our code is available at https://github.com/Feecuin/AWM-Fuse. Xilai Li, Huichun Liu, Xiaosong Li 0004, Tao Ye 0002, Zhenyu Kuang, Huafeng Li 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | Soft Supervision-Guided Spatial-Temporal Refinement Network for Video-Based Visible-Infrared Person Re-IdentificationabstractThanks to automatic switch between visible and infrared modes, person re-identification (Re-ID) in 24-hour has been possible through cross-modal retrieval. Instead of exploiting still images, video-based cross-modal person Re-ID is studied in this paper. Specifically, a large-scale dataset 'HITSZ-PVCM' is first collected, consisting of as many as 1,681 identities and 839,632 frames. Generally, videos contain much richer pedestrian appearances. However, most existing works only generate temporal representations by whole frames, inevitably losing fine-grained details. Furthermore, training a network by metric losses (e.g., center loss) is a common strategy, while such point-to-point constraints are too strong and limit model generalization due to existing diversity among intra-class samples. Here, we propose a Soft Supervision guided Spatial-Temporal Refinement (S3TR) network to tackle these problems. Specifically, S3TR refines each frame guided by a coarse temporal feature, so that more discriminative features are extracted and transformed to a sequential representation. Followed by a global-local mutual learning module, the modality gap is then erased without losing fine-grained details. Furthermore, we propose a novel soft-clustering center loss to measure intra-/inter-class similarity/dissimilarity in a group-to-group way, efficiently improving model generalization. To the best of our knowledge, HITSZ-PVCM is the largest dataset and S3TR achieves superior performances compared with state-of-the-arts. Jinxing Li 0003, Chuhao Zhou, Huafeng Li 0001, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Learning to Change by Critique and Correction: A Synergistic Framework for Remote Sensing Change Detection and CaptioningabstractChange Detection (CD) and Change Captioning (CC) are two core tasks for understanding land-cover evolution in remote sensing imagery. Existing approaches have explored CD-CC joint modeling through shared representations, task-specific decoding, feature interaction, and semantic guidance. However, the lack of explicit cross-task feedback mechanisms often leads to mutual interference, making it difficult to achieve both accurate detection and expressive descriptions. To address this issue, we propose the Learning to Change by Critique and Correction (LCCC) framework, which reformulates CD and CC as a critique-correction closed-loop process. In LCCC, CD and CC no longer passively share features but interact through bidirectional critique and correction: the CD task provides explicit spatial constraints for CC, and CC, in turn, supervises CD via a Text-Guided Critique Attention (TGCA) mechanism, establishing a synergistic relationship where both tasks act as critics and correctors. Furthermore, we design a Reciprocal Suppression and Enhancement (RSE) module to purify cross-task representations and propose a Key Complementary Feature Fusion (KCFF) mechanism to bridge the gap between high-level semantics and low-level visual features, ensuring a balance between task specialization and cross-task enhancement. Extensive experiments demonstrate that LCCC significantly outperforms existing methods in both detection accuracy and description quality, validating the effectiveness and generality of the proposed critique-correction paradigm for synergistic multi-task modeling. The code of the proposed method is available at https://github.com/Throb16/Lccc. Huafeng Li 0001, Yamin Zhang, Yunbin Tu, Liang Li 0003 |
IEEE Trans. Image Process. | 1 |
| 2026 | Bidirectional Cross-Modal Collaborative Alignment via Semantic-Guided Visual Embeddings for Partially Relevant Video RetrievalabstractPartially Relevant Video Retrieval (PRVR) aims to retrieve videos that match a given textual query only partially. This task is inherently challenging due to the modality gap between text and video, which is further exacerbated by the partial semantic correspondence between linguistic descriptions and visual content. To address these challenges, we propose a bidirectional cross-modal alignment mechanism that collaboratively optimizes both visual and textual modalities. In the visual modality, a major difficulty lies in the absence of visual cues that directly correspond to textual semantics, limiting the model's ability to align visual representations with textual meanings under unsupervised conditions. To overcome this issue, we construct a semantic-visual association library, which stores paired visual and textual features with semantic annotations. During training, the model dynamically retrieves the most semantically similar visual samples from this library based on the current visual feature vector. These retrieved samples, preliminarily associated with semantics via cross-modal matching, are used to form dynamic anchors that guide visual representation learning. By leveraging these enriched visual features, the model progressively refines the visual representations to achieve better alignment with the corresponding textual inputs, thereby enhancing cross-modal consistency. In the textual modality, we enhance textual representations by integrating semantically aligned visual features selected from the same association library, further narrowing the modality gap. Extensive experiments on benchmark datasets under partial semantic correspondence scenarios demonstrate that our method achieves state-of-the-art performance. The source code of the paper is available at https://github.com/cyanlll/BOA. Huafeng Li 0001, Jialong Zhao, Jie Wen 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | Training-Free Video Corpus Moment Retrieval via Synergistic Collaboration and Adaptive CalibrationabstractVideo Corpus Moment Retrieval (VCMR) is pivotal to multimodal understanding. However, existing methods rely heavily on large-scale annotated data, which limits their generalization and scalability. To address this issue, we propose a training-free VCMR framework, termed Synergistic Collaboration and Adaptive Calibration (SCAC), enabling effective semantic parsing and precise temporal localization without parameter updates. SCAC introduces a Query Event Chain Generation module that leverages large language models to transform complex textual queries into structured event chains, while a Video Event Chain Generation module represents videos as semantically coherent event chains through subtitle segmentation and keyframe aggregation. Built on these structured representations, SCAC performs Event-Chain-Based Cross-Modal Retrieval with mean-variance joint scoring to suppress local mismatches and reinforce global consistency. During localization, a Synergy-Calibration Mechanism dynamically refines temporal boundaries via profit-setback feedback. Extensive experiments show that SCAC achieves comparable or superior results to supervised counterparts under training-free conditions, demonstrating strong cross-modal generalization and adaptive capability. The code of our method is available at https://github.com/cyanlll/SCAC. Jialong Zhao, Huafeng Li 0001, Changchun Hua |
IEEE Trans. Image Process. | 2 |
| 2026 | Seeing Clearly and Detecting Precisely: Perceptual Enhancement and Focus Calibration for Small-Object DetectionabstractSmall-object detection remains challenging due to limited pixel information, blurred boundaries, and weak semantic cues. Although recent advances in multiscale fusion and attention mechanisms have led to improved performance, existing methods still struggle to preserve high-frequency structural details and achieve precise localization-particularly in dense, cluttered, or low-resolution scenarios. These limitations are primarily caused by the loss of fine-grained features during downsampling and the absence of region-aware focus mechanisms. Inspired by the human visual strategy of "see clearly and detect precisely," we propose PEFC-Net, a novel framework that enhances both perceptual clarity and localization accuracy for small-object detection. To mitigate structural degradation, we introduce the hybrid structural perception (HSP) module, which jointly encodes spatial gradients and localized frequency components through wavelet-based decomposition and edge-aware refinement. To further improve region-level focus, we design the axis-aligned focus calibration (AAFC) module, which captures long-range directional context via axis-sensitive pooling and adaptively refines attention with shape-aware calibration. Extensive experiments on four challenging benchmarks-VisDrone-2019, TT100K, NWPU VHR-10, and DIOR-demonstrate that PEFC-Net consistently outperforms state-of-the-art methods, delivering robust performance under occlusion, dense distribution, and scale variation. Zhiqin Zhu, Guanqiu Qi, Huafeng Li 0001, Yu Liu 0023 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2026 | Working Condition-Decoupled and Invariant-Feature Fusion Transformer for Domain Generalization Intelligent Fault DiagnosisabstractFor fault diagnosis under unseen working conditions (WCs), it is crucial to extract general knowledge unrelated to data distribution from available source data and identify transferable discriminative features. However, WC-related information is often tightly coupled with health state (HS)-related information, making it difficult to directly distinguish their contributions, posing challenges to fault diagnosis. To address this issue, a novel approach named WC-decoupled and invariant-feature fusion transformer (WCD-IFFT) is proposed, which aims to minimize the impact of WCs by extracting transferable features closely related to HSs. Specifically, two key components are designed to decouple WC-related features from HS-related features: orthogonality separation and decouple loss. Additionally, to enrich the semantics of HS-related features, time-domain and Fourier phase features are mapped into a unified space and fused, combining the instantaneous changes of time-domain signals with frequency-domain distribution information to enhance the feature representation capability. Extensive experiments on cross-domain fault diagnosis tasks demonstrate the effectiveness of the proposed method. Kaixiong Xu, Huafeng Li 0001, Meichen Lu, Yi Chai 0002, Youqiang Hu, Shenhang Wang, Ke Zhang 0006 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | BSAFusion: A Bidirectional Stepwise Feature Alignment Network for Unaligned Medical Image FusionabstractIf unaligned multimodal medical images can be simultaneously aligned and fused using a single-stage approach within a unified processing framework, it will not only achieve mutual promotion of dual tasks but also help reduce the complexity of the model. However, the design of this model faces the challenge of incompatible requirements for feature fusion and alignment. To address this challenge, this paper proposes an unaligned medical image fusion method called Bidirectional Stepwise Feature Alignment and Fusion (BSFA-F) strategy. To reduce the negative impact of modality differences on cross-modal feature matching, we incorporate the Modal Discrepancy-Free Feature Representation (MDF-FR) method into BSFA-F. MDF-FR utilizes a Modality Feature Representation Head (MFRH) to integrate the global information of the input image. By injecting the information contained in MFRH of the current image into other modality images, it effectively reduces the impact of modality differences on feature alignment while preserving the complementary information carried by different images. In terms of feature alignment, BSFA-F employs a bidirectional stepwise alignment deformation field prediction strategy based on the path independence of vector displacement between two points. This strategy solves the problem of large spans and inaccurate deformation field prediction in single-step alignment. Finally, Multi-Modal Feature Fusion block achieves the fusion of aligned features. The experimental results across multiple datasets demonstrate the effectiveness of our method. Huafeng Li 0001, Dayong Su |
AAAI | 1 |
| 2025 | UniFuse: A Unified All-In-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and Misalignments
Dayong Su, Huafeng Li 0001, Jinxing Li 0003, Yu Liu 0023 |
ICCV | 3 |
| 2025 | Weakly Supervised Visible-Infrared Person Re-Identification via Heterogeneous Expert Collaborative Consistency LearningabstractTo reduce the reliance of visible-infrared person re-identification (ReID) models on labeled cross-modal samples, this paper explores a weakly supervised cross-modal person ReID method that uses only single-modal sample identity labels, addressing scenarios where cross-modal identity labels are unavailable. To mitigate the impact of missing cross-modal labels on model performance, we propose a heterogeneous expert collaborative consistency learning framework, designed to establish robust cross-modal identity correspondences in a weakly supervised manner. This framework leverages labeled data from each modality to independently train dedicated classification experts. To associate cross-modal samples, these classification experts act as heterogeneous predictors, predicting the identities of samples from the other modality. To improve prediction accuracy, we design a cross-modal relationship fusion mechanism that effectively integrates predictions from different experts. Under the implicit supervision provided by cross-modal identity correspondences, collaborative and consistent learning among the experts is encouraged, significantly enhancing the model's ability to extract modality-invariant features and improve cross-modal identity recognition. Experimental results on two challenging datasets validate the effectiveness of the proposed method. Lingqi Kong, Huafeng Li 0001, Jie Wen 0001 |
ICCV | 3 |
| 2025 | DINOv2 Driven Gait Representation Learning for Video-Based Visible-Infrared Person Re-identificationabstractVideo-based Visible-Infrared person re-identification (VVI-ReID) aims to retrieve the same pedestrian across visible and infrared modalities from video sequences. Existing methods tend to exploit modality-invariant visual features but largely overlook gait features, which are not only modality-invariant but also rich in temporal dynamics, thus limiting their ability to model the spatiotemporal consistency essential for cross-modal video matching. To address these challenges, we propose a DINOv2-Driven Gait Representation Learning (DinoGRL) framework that leverages the rich visual priors of DINOv2 to learn gait features complementary to appearance cues, facilitating robust sequence-level representations for cross-modal retrieval. Specifically, we introduce a Semantic-Aware Silhouette and Gait Learning (SASGL) model, which generates and enhances silhouette representations with general-purpose semantic priors from DINOv2 and jointly optimizes them with the ReID objective to achieve semantically enriched and task-adaptive gait feature learning. Furthermore, we develop a Progressive Bidirectional Multi-Granularity Enhancement (PBMGE) module, which progressively refines feature representations by enabling bidirectional interactions between gait and appearance streams across multiple spatial granularities, fully leveraging their complementarity to enhance global representations with rich local details and produce highly discriminative features. Extensive experiments on HITSZ-VCM and BUPT datasets demonstrate the superiority of our approach, significantly outperforming existing state-of-the-art methods. Neng Dong, Fan Li 0006, Huafeng Li 0001 |
ACM Multimedia | 6 |
| 2025 | Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image MatchingabstractWeakly supervised text-to-person image matching, as a crucial approach to reducing models' reliance on large-scale manually labeled samples, holds significant research value. However, existing methods struggle to predict complex one-to-many identity relationships, severely limiting performance improvements. To address this challenge, we propose a local-and-global dual-granularity identity association mechanism. Specifically, at the local level, we explicitly establish cross-modal identity relationships within a batch, reinforcing identity constraints across different modalities and enabling the model to better capture subtle differences and correlations. At the global level, we construct a dynamic cross-modal identity association network with the visual modality as the anchor and introduce a confidence-based dynamic adjustment mechanism, effectively enhancing the model's ability to identify weakly associated samples while improving overall sensitivity. Additionally, we propose an information-asymmetric sample pair construction method combined with consistency learning to tackle hard sample mining and enhance model robustness. Experimental results demonstrate that the proposed method substantially boosts cross-modal matching accuracy, providing an efficient and practical solution for text-to-person image matching. Code is available at https://github.com/syl6312/DGCMIA. Yongle Shang, Huafeng Li 0001 |
ACM Multimedia | 3 |
| 2025 | Domain-adaptive person re-identification without cross-camera paired samples
Huafeng Li 0001, Yanmei Mao, Guanqiu Qi, Zhengtao Yu 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Hybrid V2V/V2I Task Offloading in Vehicular Edge Computing: A Double-Layer Stackelberg Game ApproachabstractTask offloading is a promising way to support computation-intensive applications in a resource limited network such as the Internet of Vehicles (IoVs). However, great challenges exist in applying task offloading efficiently in practical IoV scenarios. For one thing, vehicles and RoadSide Units (RSUs) are reluctant to participate in the cooperation due to the lack of incentive. For another, multiple vehicles generate various computation tasks at the same time, which further intensifies the resource competition among vehicles. Although existing work have explored the potentiality of Stackelberg game theory in stimulating the cooperation between vehicles/servers, etc., from various aspects, the classical single-layer modelling is over simplified and fails to make the best use of all available resources. To solve these issues, this paper focuses on simultaneously incentivizing vehicles and RSUs as computation assistants, to provide their resources in completing the task together. A novel double-layer Stackelberg game-based approach is proposed to perfectly characterize the collaboration and competition among vehicles/RSUs in the hybrid V2V/V2I offloading scenario. The interactions between multiple user vehicles and a RSU are formulated as the first-layer Stackelberg game, and that between a user vehicle and multiple service vehicles are defined as the second-layer Stackelberg game. The existence of Nash Equilibrium in each Stackelberg game is proved and two algorithms are designed to approximate the achievable optimal offloading strategy under practical vehicular environment. Simulation results prove that the proposed approach achieves remarkable performance advantages in terms of completion delay, utilities of vehicles, RSUs and the overall system. Shujuan Wang, Hao Peng 0001, Huafeng Li 0001 |
IEEE Internet Things J. | 3 |
| 2025 | Spatio-temporal information mining and fusion feature-guided modal alignment for video-based visible-infrared person re-identification
Zhigang Zuo, Huafeng Li 0001, Minghong Xie |
Image Vis. Comput. | 2 |
| 2025 | Playing to spot the difference: Enhancing HDR imaging with dual-task synergy and multi-perspective consensus learning
Juncheng Luo, Huafeng Li 0001 |
Knowl. Based Syst. | 3 |
| 2025 | MulFS-CAP: Multimodal Fusion-Supervised Cross-Modality Alignment Perception for Unregistered Infrared-Visible Image FusionabstractIn this study, we propose Multimodal Fusion-supervised Cross-modality Alignment Perception (MulFS-CAP), a novel framework for single-stage fusion of unregistered infrared-visible images. Traditional two-stage methods depend on explicit registration algorithms to align source images spatially, often adding complexity. In contrast, MulFS-CAP seamlessly blends implicit registration with fusion, simplifying the process and enhancing suitability for practical applications. MulFS-CAP utilizes a shared shallow feature encoder to merge unregistered infrared-visible images in a single stage. To address the specific requirements of feature-level alignment and fusion, we develop a consistent feature learning approach via a learnable modality dictionary. This dictionary provides complementary information for unimodal features, thereby maintaining consistency between individual and fused multimodal features. As a result, MulFS-CAP effectively reduces the impact of modality variance on cross-modality feature alignment, allowing for simultaneous registration and fusion. Additionally, in MulFS-CAP, we advance a novel cross-modality alignment approach, creating a correlation matrix to detail pixel relationships between source images. This matrix aids in aligning features across infrared and visible images, further refining the fusion process. The above designs make MulFS-CAP more lightweight, effective and explicit registration-free. Experimental results from different datasets demonstrate the effectiveness of our proposed method and its superiority over the state-of-the-art two-stage methods. Huafeng Li 0001, Zengyi Yang, Wei Jia 0001, Zhengtao Yu 0001, Yu Liu 0023 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Multi-granular inter-frame relation exploration and global residual embedding for video-based person re-identification
Zhiqin Zhu, Sixin Chen, Guanqiu Qi, Huafeng Li 0001, Xinbo Gao 0001 |
Signal Process. Image Commun. | 4 |
| 2025 | TriMatch: Triple Matching for Text-to-Image Person Re-IdentificationabstractText-to-image person re-identification (TIReID) is a cross-modal retrieval task that aims to retrieve target person images based on a given text description. Existing methods primarily focus on mining the semantic associations across modalities, relying on the matching between heterogeneous features for retrieval. However, due to the inherent heterogeneous gaps between modalities, it is challenging to establish precise semantic associations, particularly in fine-grained correspondences, often leading to incorrect retrieval results. To address this issue, this letter proposes an innovative Triple Matching (TriMatch) framework that integrates cross-modal (image-text) matching and unimodal (image-image, text-text) matching for high-precision person retrieval. The framework introduces a generation task that performs cross-modal (image-to-text and text-to-image) feature generation and intra-modal feature alig achieve unimodal matching. By incorporating the generation task, TriMatch considers not only the semantic correlations between modalities but also the semantic consistency within single modalities, thereby effectively enhancing the accuracy of target person retrieval. Extensive experiments on multiple datasets demonstrate the superiority of TriMatch over existing methods. Shuanglin Yan, Neng Dong, Huafeng Li 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Multi-Expert Adaptive Selection: Task-Balancing for All-in-One Image RestorationabstractThe use of a single image restoration framework to achieve multi-task image restoration has garnered significant attention from researchers. However, several practical challenges remain, including meeting the specific and simultaneous demands of different tasks, balancing relationships between tasks, and effectively utilizing task correlations in model design. To address these challenges, this paper explores a multi-expert adaptive selection mechanism. We begin by designing a feature representation method that accounts for both the pixel channel level and the global level, encompassing low-frequency and high-frequency components of the image. Based on this method, we construct a multi-expert selection and ensemble scheme. This scheme adaptively selects the most suitable expert from the expert library according to the content of the input image and the prompts of the current task. It not only meets the individualized needs of different tasks but also achieves balance and optimization across tasks. By sharing experts, our design promotes interconnections between different tasks, thereby enhancing overall performance and resource utilization. Additionally, the multi-expert mechanism effectively eliminates irrelevant experts, reducing interference from them and further improving the effectiveness and accuracy of image restoration. Experimental results demonstrate that our proposed method is both effective and superior to existing approaches, highlighting its potential for practical applications in multi-task image restoration. The source code of the proposed method is available athttps://github.com/zhoushen1/MEASNet. Shen Zhou 0001, Huafeng Li 0001, Liehuang Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Breaking the Paired Sample Barrier in Person Re-Identification: Leveraging Unpaired Samples for Domain GeneralizationabstractDomain generalization (DG) for person re-identification (Re-ID) aims to train models on labeled source domains that generalize well to unseen target domains. However, DG for Re-ID faces a major challenge: existing methods rely solely on labeled paired samples to train DG models and are unable to effectively leverage unpaired samples across cameras. In many cases, cross-camera paired samples are extremely scarce and difficult to annotate. To overcome this limitation, we introduce a novel method specifically tailored for Re-ID. This method leverages cross-camera unpaired samples in model training, thereby reducing the dependence on cross-camera paired samples. We refer to this technique as Unpaired-driven DG (U-DG) person Re-ID. The proposed method leverages a robust image encoder to extract identity-consistent features across various camera views. This capability is further enhanced by integrating a multi-camera person identity classifier, which boosts the encoder’s ability to capture consistent identities, even when viewed from different camera perspectives. To address the scarcity of cross-camera paired samples, we devise a unique model training strategy in our method. Specifically, we use the feature vector from the person identity classifier as a single identity prototype. This prototype serves as a reference for generating identity-related prompts across cameras, effectively compensating for the scarcity of cross-camera paired samples during model training. Additionally, we employ a learnable perturbation prompt to mimic appearance variations exhibited by the same individual across different cameras. Our U-DG offers numerous advantages: it can effectively leverage a large number of unpaired samples for model training, compensating for the scarcity of cross-camera paired samples. Moreover, it does not rely solely on cross-camera paired samples, thereby facilitating the construction of training samples. Experimental results on multiple challenging datasets demonstrate that our approach achieves performance comparable to typical DG person Re-ID, highlighting its feasibility and effectiveness. The source code of our method is available athttps://github.com/lhf12278/DGPS. Huafeng Li 0001, Yaoxin Liu, Jinxing Li 0003, Zhengtao Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | UMCFuse: A Unified Multiple Complex Scenes Infrared and Visible Image Fusion FrameworkabstractInfrared and visible image fusion has emerged as a prominent research area in computer vision. However, little attention has been paid to the fusion task in complex scenes, leading to sub-optimal results under interference. To fill this gap, we propose a unified framework for infrared and visible images fusion in complex scenes, termed UMCFuse. Specifically, we classify the pixels of visible images from the degree of scattering of light transmission, allowing us to separate fine details from overall intensity. Maintaining a balance between interference removal and detail preservation is essential for the generalization capacity of the proposed method. Therefore, we propose an adaptive denoising strategy for the fusion of detail layers. Meanwhile, we fuse the energy features from different modalities by analyzing them from multiple directions. Extensive fusion experiments on real and synthetic complex scenes datasets cover adverse weather conditions, noise, blur, overexposure, fire, as well as downstream tasks including semantic segmentation, object detection, salient object detection, and depth estimation, consistently indicate the superiority of the proposed method compared with the recent representative methods. Our code is available at https://github.com/ixilai/UMCFuse. Xilai Li, Xiaosong Li 0004, Tianshu Tan, Huafeng Li 0001, Tao Ye 0002 |
IEEE Trans. Image Process. | 4 |
| 2025 | Disentangling Inter- and Intra-Video Relations for Multi-Event Video-Text Retrieval and GroundingabstractVideo-text retrieval aims to precisely search for videos most relevant to text queries within a video corpus. However, existing methods are largely limited to single-text (single-event) queries and are not effective at handling multi-text (multi-event) queries. Furthermore, these methods typically focus solely on retrieval and do not attempt to locate multiple events within the retrieved videos. To address these limitations, our paper proposes a novel method named Disentangling Inter- and Intra-Video Relations, which jointly addresses multi-event video-text retrieval and grounding. This method leverages both inter-video and intra-video event relationships to enhance retrieval and grounding performance. At the retrieval level, we devise a Relational Event-Centric Video-Text Retrieval module based on the principle that comprehensive textual information leads to precise correspondence between text and video. It incorporates event relationship features at different hierarchical levels and exploits the hierarchical structure of video relationships to achieve multi-level contrastive learning between events and videos. This approach enhances the richness, accuracy, and comprehensiveness of event descriptions, improving alignment precision between text and video and enabling effective differentiation among videos. For event grounding, we propose Event Contrast-Driven Video Grounding, which accounts for positional differences among events on the 2D temporal score map and achieves precise grounding of multiple events through divergence learning for their locations. Our solution not only provides efficient text-to-video retrieval but also accurately grounds events within the retrieved videos, addressing the shortcomings of existing methods. Extensive experimental results on the ActivityNet Captions and Charades-STA benchmark datasets demonstrate the superior performance of our method, validating its effectiveness. The innovation of this research lies in introducing a new joint framework for video-text retrieval and multi-event grounding while offering new ideas for further research and applications in related fields. The code is available at https://github.com/X7J92/MVT-RG. Mengzhao Wang 0002, Huafeng Li 0001, Jinxing Li 0003, Dapeng Tao, Zhengtao Yu 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Cps-STS: Bridging the Gap Between Content and Position for Coarse-Point-Supervised Scene Text SpotterabstractRecently, weakly supervised methods for scene text spotter are increasingly popular with researchers due to their potential to significantly reduce dataset annotation efforts. The latest progress in this field is text spotter based on single or multi-point annotations. However, this method struggles with the sensitivity of text recognition to the precise annotation location and fails to capture the relative positions and shapes of characters, leading to impaired recognition of texts with extensive rotations and flips. To address these challenges, this paper develops a novel method named Coarse-point-supervised Scene Text Spotter (Cps-STS). Cps-STS first utilizes a few approximate points as text location labels and introduces a learnable position modulation mechanism, easing the accuracy requirements for annotations and enhancing model robustness. Additionally, we incorporate a Spatial Compatibility Attention (SCA) module for text decoding to effectively utilize spatial data such as position and shape. This module fuses compound queries and global feature maps, serving as a bias in the SCA module to express text spatial morphology. In order to accurately locate and decode text content, we introduce features containing spatial morphology information and text content into the input features of the text decoder. By introducing features with spatial morphology information as bias terms into the text decoder, ablation experiments demonstrate that this operation enables the model to effectively identify and utilize the relationship between text content and position to enhance the recognition performance of our model. One significant advantage of Cps-STS is its ability to achieve full supervision-level performance with just a few imprecise coarse points at a low cost. Extensive experiments validate the effectiveness and superiority of Cps-STS over existing approaches. Weida Chen, Jie Jiang 0015, Linfei Wang, Huafeng Li 0001, Yibing Zhan, Dapeng Tao |
IEEE Trans. Multim. | 4 |
| 2025 | Progressive Feature Mining and External Knowledge-Assisted Text-Pedestrian Image RetrievalabstractText-Pedestrian Image Retrieval employs textual description of pedestrian's appearance to identify the corresponding pedestrian image. This task involves modality discrepancy and the challenges posed by textual diversity of pedestrians with the same identity. Although advancements have been made in text-pedestrian image retrieval, current methods do not comprehensively address these challenges. Thus, this paper proposes a progressive feature mining and external knowledge- assisted feature purification method. Specifically, we implement a progressive mining mode, enabling the model to extract discriminative features from overlooked information. This enhances the model's feature representation capabilities and prevents the loss of discriminative information. To further mitigate the challenges posed by modality discrepancy and text diversity in cross-modal matching, we propose to use external knowledge of other samples from the same modality. This approach accentuates identity-consistent features and diminishes identity-inconsistent ones, refining feature representation and reducing interference from textual diversity and negative sample correlation features of the same modality. Extensive experiments on three challenging datasets demonstrate the effectiveness and superiority of the proposed method, with its retrieval performance outstripping that of large-scale model-based methods on large-scale datasets. Huafeng Li 0001, Shedan Yang, Dapeng Tao, Zhengtao Yu 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Dual-Task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and GroundingabstractVideo Paragraph Grounding (VPG) aims to precisely locate the most appropriate moments within a video that are relevant to a given textual paragraph query. However, existing methods typically rely on large-scale annotated temporal labels and assume that the correspondence between videos and paragraphs is known. This is impractical in real-world applications, as constructing temporal labels requires significant labor costs, and the correspondence is often unknown. To address this issue, we propose a Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding method (DMR-JRG). In this method, retrieval and grounding tasks are mutually reinforced rather than being treated as separate issues. DMR-JRG mainly consists of two branches: a retrieval branch and a grounding branch. The retrieval branch uses inter-video contrastive learning to roughly align the global features of paragraphs and videos, reducing modality differences and constructing a coarse-grained feature space to break free from the need for correspondence between paragraphs and videos. Additionally, this coarse-grained feature space further facilitates the grounding branch in extracting fine-grained contextual representations. In the grounding branch, we achieve precise cross-modal matching and grounding by exploring the consistency between local, global, and temporal dimensions of video segments and textual paragraphs. By synergizing these dimensions, we construct a fine-grained feature space for video and textual features, greatly reducing the need for large-scale annotated temporal labels. Meanwhile, we design a grounding reinforcement retrieval module (GRRM) that brings the coarse-grained feature space of the retrieval branch closer to the fine-grained feature space of the grounding branch, thereby reinforcing retrieval branch through grounding branch, and finally achieving mutual reinforcement between tasks. Extensive experiments on three challenging datasets demonstrate the effectiveness of our proposed method. The code is available athttps://github.com/X7J92/DMR-JRG. Mengzhao Wang 0002, Huafeng Li 0001, Jinxing Li 0003, Minghong Xie, Dapeng Tao |
IEEE Trans. Multim. | 2 |
| 2025 | Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction for Visual GroundingabstractVisual grounding has attracted wide attention thanks to its broad application in various visual language tasks. Although visual grounding has made significant research progress, existing methods ignore the promotion effect of the association between text and image features at different hierarchies on cross-modal matching. This paper proposes a Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction Visual Grounding method. It first generates a mask through decoupled sentence phrases, and a text and image hierarchical matching mechanism is constructed, highlighting the role of association between different hierarchies in cross-modal matching. In addition, a corresponding target object position progressive correction strategy is defined based on the hierarchical matching mechanism to achieve accurate positioning for the target object described in the text. This method can continuously optimize and adjust the bounding box position of the target object as the certainty of the text description of the target object improves. This design explores the association between features at different hierarchies and highlights the role of features related to the target object and its position in target positioning. The proposed method is validated on different datasets through experiments, and its superiority is verified by the performance comparison with the state-of-the-art methods. Minghong Xie, Mengzhao Wang 0002, Huafeng Li 0001, Dapeng Tao, Zhengtao Yu 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Focus Affinity Perception and Super-Resolution Embedding for Multifocus Image FusionabstractDespite the fact that there is a remarkable achievement on multifocus image fusion, most of the existing methods only generate a low-resolution image if the given source images suffer from low resolution. Obviously, a naive strategy is to independently conduct image fusion and image super-resolution. However, this two-step approach would inevitably introduce and enlarge artifacts in the final result if the result from the first step meets artifacts. To address this problem, in this article, we propose a novel method to simultaneously achieve image fusion and super-resolution in one framework, avoiding step-by-step processing of fusion and super-resolution. Since a small receptive field can discriminate the focusing characteristics of pixels in detailed regions, while a large receptive field is more robust to pixels in smooth regions, a subnetwork is first proposed to compute the affinity of features under different types of receptive fields, efficiently increasing the discriminability of focused pixels. Simultaneously, in order to prevent from distortion, a gradient embedding-based super-resolution subnetwork is also proposed, in which the features from the shallow layer, the deep layer, and the gradient map are jointly taken into account, allowing us to get an upsampled image with high resolution. Compared with the existing methods, which implemented fusion and super-resolution independently, our proposed method directly achieves these two tasks in a parallel way, avoiding artifacts caused by the inferior output of image fusion or super-resolution. Experiments conducted on the real-world dataset substantiate the superiority of our proposed method compared with state of the arts. Huafeng Li 0001, Jinxing Li 0003, Yu Liu 0023, Guangming Lu 0002, Yong Xu 0001, Zhengtao Yu 0001, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Catalyst for Clustering-Based Unsupervised Object Re-identification: Feature CalibrationabstractClustering-based methods are emerging as a ubiquitous technology in unsupervised object Re-Identification (ReID), which alternate between pseudo-label generation and representation learning. Recent advances in this field mainly fall into two groups: pseudo-label correction and robust representation learning. Differently, in this work, we improve unsupervised object ReID from feature calibration, a completely different but complementary insight from the current approaches. Specifically, we propose to insert a conceptually simple yet empirically powerful Feature Calibration Module (FCM) before pseudo-label generation. In practice, FCM calibrates the features using a nonparametric graph attention network, enforcing similar instances to move together in the feature space while allowing dissimilar instances to separate. As a result, we can generate more reliable pseudo-labels using the calibrated features and further improve subsequent representation learning. FCM is simple, effective, parameter-free, training-free, plug-and-play, and can be considered as a catalyst, increasing the ’chemical reaction’ between pseudo-label generation and representation learning. Moreover, it maintains the efficiency of testing time with negligible impact on training time. In this paper, we insert FCM into a simple baseline. Experiments across different scenarios and benchmarks show that FCM consistently improves the baseline (e.g., 8.2% mAP gain on MSMT17), and achieves the new state-of-the-art results. Code is available at: https://github.com/lhf12278/FCM-ReID. Huafeng Li 0001, Qingsong Hu, Zhanxuan Hu |
AAAI | 1 |
| 2024 | Depth Information Assisted Collaborative Mutual Promotion Network for Single Image DehazingabstractRecovering a clear image from a single hazy image is an open inverse problem. Although significant research progress has been made, most existing methods ignore the effect that downstream tasks play in promoting upstream de-hazing. From the perspective of the haze generation mechanism, there is a potential relationship between the depth information of the scene and the hazy image. Based on this, we propose a dual-task collaborative mutual promotion framework to achieve the dehazing of a single image. This framework integrates depth estimation and de-hazing by a dual-task interaction mechanism and achieves mutual enhancement of their performance. To realize the joint optimization of the two tasks, an alternative imple-mentation mechanism with the difference perception is developed. On the one hand, the difference perception between the depth maps of the dehazing result and the ideal image is proposed to promote the dehazing network to pay attention to the non-ideal areas of the dehazing. On the other hand, by improving the depth estimation performance in the difficult-to-recover areas of the hazy image, the de-hazing network can explicitly use the depth information of the hazy image to assist the clear image recovery. To promote the depth estimation, we propose to use the difference between the dehazed image and the ground truth to guide the depth estimation network to focus on the de-hazed unideal areas. It allows dehazing and depth estimation to leverage their strengths in a mutually reinforcing manner. Experimental results show that the proposed method can achieve better performance than that of the state-of-the-art approaches. The source code is released at https://github.com/zhoushen/IDCMPNet. Shen Zhou 0001, Huafeng Li 0001 |
CVPR | 3 |
| 2024 | Prototype-Guided Dual-Transformer Reasoning for Video Individual CountingabstractVideo Individual Counting (VIC), which focuses on accurately tallying the total number of individuals in a video without duplication, is crucial for urban public space management and densely-populated areas planning. Existing methods suffer from limitations in terms of expensive manual annotation, and the efficiency of location or detection algorithms. In this work, we contribute a novel Prototype-guided Dual-Transformer Reasoning framework, termed PDTR, which takes both similarity and difference of adjacent frames into account to achieve accurate counting in an end-to-end regression manner. Specifically, we first design a multi-receptive field feature fusion module to acquire initial comprehensive representations. Subsequently, the dynamic prototype generation module memorizes consistent representations of similar information to generate prototypes. Additionally, to further dig out the shared and private features from different frames, a prototype cross-guided decoder and a privacy-decoupling module are designed. Extensive experiments conducted on two existing VIC datasets, consistently demonstrate the superiority of PDTR over state-of-the-art baselines. Yishu Liu 0001, Huafeng Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
ACM Multimedia | 3 |
| 2024 | Robust discriminative and modal-consistent feature learning for fine-grained sketch-based image retrieval
Junchao Ge, Huafeng Li 0001 |
MMAsia | 2 |
| 2024 | DSLSM: Dual-kernel-induced statistic level set model for image segmentation
Fan Zhang 0070, Xiaojun Duan, Binglu Wang, Huafeng Li 0001, Junyu Dong, David Zhang 0001 |
Expert Syst. Appl. | 6 |
| 2024 | A Deep Learning Framework for Infrared and Visible Image Fusion Without Strict Registration
Huafeng Li 0001, Junyu Liu, Yu Liu 0023 |
Int. J. Comput. Vis. | 1 |
| 2024 | Supervised adaptive similarity consistent latent representation hashing
Hongbin Wang 0002, Zhenqiu Shu, Huafeng Li 0001 |
Neurocomputing | 5 |
| 2024 | Interactive attack-defense for generalized person re-identification
Huafeng Li 0001, Zhanxuan Hu, Zhengtao Yu 0001 |
Neural Networks | 1 |
| 2024 | Cross co-teaching for semi-supervised medical image segmentation
Fan Zhang 0070, Jinjiang Wang, Huafeng Li 0001, Junyu Dong, David Zhang 0001 |
Pattern Recognit. | 6 |
| 2024 | Generation and Recombination for Multifocus Image Fusion With Free Number of InputsabstractMultifocus image fusion is an effective method to overcome the limitations of optical lenses. The fused results can be obtained from some existing methods by generating decision maps. However, such methods assume that the focused areas of the two source images are complementary, making it impossible to achieve the simultaneous fusion of multiple images. Additionally, existing methods ignore the impact of hard pixels on the fusion performance, limiting the visual quality improvement of fusion images. To address these issues, a combined generation and recombination model called GRFusion is proposed. In GRFusion, the focus property detection of each source image can be independently implemented, enabling the simultaneous fusion of multiple source images and avoiding information loss caused by alternating fusion. It renders the GRFusion free from the limitation of the number of input images. Furthermore, GRFusion investigates the detection of hard pixels with ambiguous focus properties by analyzing the inconsistencies among the detection results of the focus areas in the source images. This allows the hard pixels to be distinguished from the source images. Besides, a multidirectional gradient embedding method is proposed for generating full-focus images. Subsequently, a hard-pixel-guided recombination mechanism for constructing the fused result is devised to integrate the complementary advantages of feature reconstruction-based and focused pixel recombination-based methods. Extensive experimental results demonstrate the effectiveness and superiority of the proposed method. The source code of the proposed method is available at: https://github.com/lhf12278/GRFusion. Huafeng Li 0001, Yuxin Huang 0004, Zhengtao Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | DCPNet: A Dual-Task Collaborative Promotion Network for PansharpeningabstractPansharpening, a type of image fusion, combines a high-resolution panchromatic (PAN) image with a low-resolution multispectral (LRMS) image to produce a high-resolution multispectral output. Recent advancements in deep learning have led to the development of data-driven pansharpening methods, which have shown superior performance over traditional non-data-driven approaches. Nevertheless, current data-driven methods primarily focus on the design of single-task-driven networks. They tend to ignore the positive impact that incorporating auxiliary tasks could have on pansharpening. This oversight restricts further improvements in their performance. To break the constraints of single-task methods, we propose a dual-task collaborative promotion network (DCPNet) for pansharpening. DCPNet incorporates an LRMS image super-resolution (SR) reconstruction network into the pansharpening network, establishing a dual-task parallel collaborative framework that achieves joint collaborative optimization for both pansharpening and SR reconstruction. Additionally, we devise a feature fusion scheme that jointly reinforces spectral and spatial details, establishing a bridge for information interaction between the two task branches. This enables the two branches to extract higher quality features in a collaborative way, thus facilitating the realization of the interaction and fusion of dual-task features, as well as the exploration of potential information within the fused features. Experimental results demonstrate that our approach not only effectively transfers spectral information from the LRMS images to the fusion result but also preserves spatial details from the PAN images. The code is available at https://github.com/lhf12278/DCPNet. Xuji Yang, Huafeng Li 0001, Minghong Xie, Zhengtao Yu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Disentanglement Learning With Adaptive Centroid Alignment for Multiple Target Domains Fault DiagnosisabstractMost current domain adaptation methods for fault diagnosis focus on single target domain. However, test data often comes from multiple target domains, as machines work under different operating conditions, subsequently generating a more complex and extensive distribution of target data. Unfortunately, single target domain adaptation methods are not adaptive for multiple target domains adaptation (MTDA), which results in transfer performance degradation. To this end, a novel disentanglement learning with adaptive centroid alignment is proposed for MTDA. Specifically for disentanglement learning, two encoders and two classifiers are constructed independently for fault-related and domain-related feature extractions and classifications. Followed by the dual-adversarial strategy, only fault-related but domain-irrelevant features are extracted. Furthermore, to achieve the category alignment, we also propose an adaptive centroid alignment strategy, so that the feature centroids of the same fault category in different domains are enforced to be close to each other. Extensive experiments demonstrate the superiority of our proposed method compared with other popular approaches. Yu Gao 0024, Xutao Zheng, Jinxing Li 0003, Lijun Zong, Hongpeng Yin, Huafeng Li 0001, Guangming Lu 0002 |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | Feature Adaptive Modulation and Prototype Learning for Domain Generalization Intelligent Fault DiagnosisabstractMost existing domain generalization fault diagnosis methods concentrate on learning domain-invariant features or global feature distribution alignment. Nevertheless, this could lose vital clues related to the fault categories, and make adapting to unknown working conditions challenging. To this end, a novel approach termed feature adaptive modulation and health state prototype consistency learning (FAMPL) is proposed. Specifically, FAMPL incorporates a feature adaptive modulation module designed to generate modulation parameters, which are utilized to perform affine transformations on the acquired features, yielding modulation features. This approach aims to capture essential clues associated with specific working conditions. To further enhance the ability to distinguish between different fault categories, a specialized health state prototype learning strategy has been developed. This approach significantly refines the model's capacity for feature discrimination, making it more adept at accurately identifying and categorizing various fault types. Numerous cross-domain fault diagnosis experiments have demonstrated the superiority of FAMPL. Kaixiong Xu, Huafeng Li 0001, Yi Chai 0002, Maoyun Guo |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Logical Relation Inference and Multiview Information Interaction for Domain Adaptation Person Re-IdentificationabstractDomain adaptation person re-identification (Re-ID) is a challenging task, which aims to transfer the knowledge learned from the labeled source domain to the unlabeled target domain. Recently, some clustering-based domain adaptation Re-ID methods have achieved great success. However, these methods ignore the inferior influence on pseudo-label prediction due to the different camera styles. The reliability of the pseudo-label plays a key role in domain adaptation Re-ID, while the different camera styles bring great challenges for pseudo-label prediction. To this end, a novel method is proposed, which bridges the gap of different cameras and extracts more discriminative features from an image. Specifically, an intra-to-intermechanism is introduced, in which samples from their own cameras are first grouped and then aligned at the class level across different cameras followed by our logical relation inference (LRI). Thanks to these strategies, the logical relationship between simple classes and hard classes is justified, preventing sample loss caused by discarding the hard samples. Furthermore, we also present a multiview information interaction (MvII) module that takes features of different images from the same pedestrian as patch tokens, obtaining the global consistency of a pedestrian that contributes to the discriminative feature extraction. Unlike the existing clustering-based methods, our method employs a two-stage framework that generates reliable pseudo-labels from the views of the intracamera and intercamera, respectively, to differentiate the camera styles, subsequently increasing its robustness. Extensive experiments on several benchmark datasets show that the proposed method outperforms a wide range of state-of-the-art methods. The source code has been released at https://github.com/lhf12278/LRIMV. Fan Li 0006, Jinxing Li 0003, Huafeng Li 0001, Bob Zhang 0001, Dapeng Tao, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Video-based Visible-Infrared Person Re-Identification via Style Disturbance Defense and Dual InteractionabstractVideo-based visible-infrared person re-identification (VVI-ReID) aims to retrieve video sequences of the same pedestrian from different modalities. The key of VVI-ReID is to learn discriminative sequence-level representations that are invariant to both intra- and inter-modal discrepancies. However, most works only focus on the elimination of modality-gap while ignore the distractors within the modality. Moreover, existing sequence-level representation learning approaches are limited to a single video, failing to mine the correlations among multiple videos of the same pedestrian. In this paper, we propose a Style Augmentation, Attack and Defense network with Graph-based dual interaction (SAADG) to guarantee the semantic consistency against both intra-modal discrepancies and inter-modal gap. Specifically, we first generate diverse styles for video frames by random style variation in image spaces. Followed by the style attack and defense, the intra- and inter-modal discrepancies are modeled as different types of style disturbance (attack), and our model achieves to keep the id-related content invariant under such attack. Besides, a graph-based dual interaction module is further introduced to fully explore the cross-view and cross-modal correlations among various videos of the same identity, which are then transferred to the sequence-level representations. Extensive experiments on the public SYSU-MM01 and HITSZ-VCM datasets show that our approach achieves the remarkable performance compared with state-of-the-arts. The code is available at https://github.com/ChuhaoZhou99/SAADG_VVIReID. Chuhao Zhou, Jinxing Li 0003, Huafeng Li 0001, Guangming Lu 0002, Yong Xu 0001, Min Zhang 0005 |
ACM Multimedia | 3 |
| 2023 | Multi-stage attention network for monaural speech enhancementabstractAbstract Although current attention‐based speech enhancement methods have been proven to be capable of significantly improving the noise reduction performance, a bottleneck has arisen in juggling both detailed features and high‐level features: The more attention paid to such performance indicators as speech intelligibility index and articulation index will lead to the loss of subtle features such as syllable continuity and timbre distortion. To tackle such a problem, we divide the speech enhancement model into multiple stages, with a special attention mechanism introduced in each stage so that the detailed speech information is retained and the features of high levels are captured. In this research, a Multi‐Stage Attention Network (MSANet) is implemented in cascade by using three different attention modules: Self Attention, Channel Attention and Spatial Attention. The attention module in each stage needs only to focus on the features of its corresponding stage, thus giving full play to their respective advantages without affecting or damaging the features of other stages, helping decouple the feature extraction process and obtaining the features capable of complementing one another. The comparisons and ablation experiments show that the performance of MSANet is superior to those of current mainstream time domain or time‐frequency domain state‐of‐the‐art methods. Compared to the baseline, PESQ score of their model (3.11) has increased by 5%, CSIG (4.44) has increased by 2.5%, CBAK (3.63) has increased by 2.8% and COVL (3.81) has increased by 3.8%, which demonstrates the potential of MSANet as speech feature extraction backbones. An implementation of the PyTorch version is available here: http://www.msp‐lab.cn:1436/msp/MSANet . Kunpeng Wang 0002, Wenjing Lu, Juan Yao, Huafeng Li 0001 |
IET Signal Process. | 5 |
| 2023 | Relation-aware attention for video captioning via graph learning
Yunbin Tu, Huafeng Li 0001, Shengxiang Gao, Zhengtao Yu 0001 |
Pattern Recognit. | 4 |
| 2023 | Intermediary-Guided Bidirectional Spatial-Temporal Aggregation Network for Video-Based Visible-Infrared Person Re-IdentificationabstractThis work focuses on the task of Video-based Visible-Infrared Person Re-Identification, a promising technique for achieving 24-hour surveillance systems. Two main issues in this field are modality discrepancy mitigating and spatial–temporal information mining. In this work, we propose a novel method, named Intermediary-guided Bidirectional spatial–temporal Aggregation Network (IBAN), to address both issues at once. Specifically, IBAN is designed to learn modality-irrelevant features by leveraging the anaglyph data of pedestrian images to serve as the intermediary. Furthermore, a bidirectional spatial–temporal aggregation module is introduced to exploit the spatial–temporal information of video data, while mitigating the impact of noisy image frames. Finally, we design an Easy-sample-based loss to guide the final embedding space and further improve the model’s generalization performance. Extensive experiments on Video-based Visible-Infrared benchmarks show that IBAN achieves promising results and outperforms the state-of-the-art ReID methods by a large margin, improving the rank-1/mAP by$1.29\%/3.46\%$at the Infrared to Visible situation, and by$5.04\%/3.27\%$at the Visible to Infrared situation. The source code of the proposed method will be released athttps://github.com/lhf12278/IBAN. Huafeng Li 0001, Minghui Liu 0001, Zhanxuan Hu, Feiping Nie 0001, Zhengtao Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Occluded Person Re-Identification via Defending Against Attacks From ObstaclesabstractDue to incomplete appearance features, the identity matching of occluded pedestrians under multiple cross-camera views is a long-term challenge. Although existing re-identification (re-ID) solutions of occluded pedestrians have made significant progress, most of them achieve accurate identity matching by extracting pedestrian appearance features from unoccluded areas. However, when a pedestrian is partially blocked by the body of another pedestrian, existing methods cannot accurately determine whether the unoccluded body parts belong to the target pedestrian, which brings great difficulties to pedestrian identity matching. To alleviate this problem, this paper introduces the idea of adversarial attack into occluded person re-ID and proposes an adversarial training framework that can defend against attacks from obstacles to resist the interference of obstacles on pedestrian identity matching. Unlike existing solutions, the proposed framework is not limited to extracting features of unoccluded human body areas to achieve occluded person re-ID, but explores how to make the re-ID model more resistant to obstacles. In the proposed framework, the occluded pedestrian images are regarded as adversarial examples and used to attack model training. If the trained model can defend against this kind of attack, its generalization is significantly improved, and the above-mentioned issues are also effectively solved. Specifically, a single-branch dual-stream collaborative network is designed. With the cooperation of the pre-trained verification guidance network, the model realizes the attack and defense of adversarial samples. This work broadens research horizons in robust model design of occluded person re-ID, and expands the scope of adversarial attacks. Compared with existing solutions, a lot of experimental results confirm that the proposed solution achieves better performance on two occluded re-ID datasets and two partial re-ID datasets. Shujuan Wang, Run Liu 0004, Huafeng Li 0001, Guanqiu Qi, Zhengtao Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Joint Image and Feature Levels Disentanglement for Generalizable Vehicle Re-identificationabstractDomain generalization (DG), which doesn’t require any data from target domains during training, is more challenging but practical than unsupervised domain adaptation (UDA). Since different vehicles of the same type have a similar appearance, neural networks always rely on a small amount of useful information to distinguish them, meaning that is more significant to remove ID-unrelated information for vehicle re-Identification (re-ID). Therefore, it is the key to eliminating the interference of a large amount of redundant information for the generalizable vehicle re-ID method. To address this unique challenge, we propose a novel disentanglement learning method that encourages variational autoencoder (VAE) network to reduce ID-unrelated features of vehicles by minimizing image reconstruction errors and providing sufficient representation to vehicle labels. To capture the intrinsic characteristics associated with the DG task, our core idea is to build the identity information streaming framework to separate ID-related and ID-unrelated information at the image and feature levels. In contrast with the general decoupling methods, our method leverages the decoupling of joint image and feature levels to extract more generalizable features. Furthermore, we present a brand-new vehicle dataset of truck types named “Optimus Prime (Opri)”, which includes multiple images of each truck captured by cameras at different high-speed toll gates. Experimental results on public datasets demonstrate that our method can achieve promising results and outperform several state-of-the-art approaches. Our codes and models are available at JIFD. Zhenyu Kuang, Chuchu He, Yue Huang 0001, Xinghao Ding, Huafeng Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Learning Modal-Invariant and Temporal-Memory for Video-based Visible-Infrared Person Re-IdentificationabstractThanks for the cross-modal retrieval techniques, visible-infrared (RGB-IR) person re-identification (Re-ID) is achieved by projecting them into a common space, allowing person Re-ID in 24-hour surveillance systems. However, with respect to the probe-to- gallery, almost all existing RGB-IR based cross-modal person Re-ID methods focus on image-to-image matching, while the video-to-video matching which contains much richer spatial- and temporal-information remains under-explored. In this paper, we primarily study the video-based cross-modal per-son Re-ID method. To achieve this task, a video-based RGB-IR dataset is constructed, in which 927 valid identities with 463,259 frames and 21,863 tracklets captured by 12 RGB/IR cameras are collected. Based on our constructed dataset, we prove that with the increase of frames in a tracklet, the performance does meet more enhancement, demonstrating the significance of video-to-video matching in RGB-IR person Re-ID. Additionally, a novel method is further proposed, which not only projects two modalities to a modal-invariant subspace, but also extracts the temporal-memory for motion-invariant. Thanks to these two strategies, much better results are achieved on our video-based cross-modal person Re-ID. The code and dataset are released at: https://github.com/VCM-project233/MITML. Jinxing Li 0003, Zeyu Ma 0001, Huafeng Li 0001, Kaixiong Xu, Guangming Lu 0002, David Zhang 0001 |
CVPR | 4 |
| 2022 | Cross-Compatible Embedding and Semantic Consistent Feature Construction for Sketch Re-identificationabstractSketch re-identification (Re-ID) refers to using sketches of pedestrians to retrieve their corresponding photos from surveillance videos. It can track pedestrians according to the sketches drawn based on eyewitnesses without querying pedestrian photos. Although the Sketch Re-ID concept has been proposed, the gap between the sketch and the photo still greatly hinders pedestrian identity matching. Based on the idea of transplantation without rejection, we propose a Cross-Compatible Embedding (CCE) approach to narrow the gap. A Semantic Consistent Feature Construction (SCFC) scheme is simultaneously presented to enhance feature discrimination. Under the guidance of identity consistency, the CCE performs cross modal interchange at the local token level in the Transformer framework, enabling the model to extract modal-compatible features. The SCFC improves the representation ability of features by handling the inconsistency of information in the same location of the sketch and the corresponding pedestrian photo. The SCFC scheme divides the local tokens of pedestrian images with different modes into different groups and assigns specific semantic information to each group for constructing a semantic consistent global feature representation. Experiments on the public Sketch Re-ID dataset confirm the effectiveness of the proposed method and its superiority over existing methods. Experiments on Sketch-based image retrieval datasets QMUL-Shoe-v2 and QMUL-Chair-v2 are conducted to assess the method's generalization. The results show that the proposed method outperforms the state-of-the-art works compared. The source code of our method is available at: https://github.com/lhf12278/CCSC. Yongzeng Wang, Huafeng Li 0001 |
ACM Multimedia | 3 |
| 2022 | Hierarchical complementary residual attention learning for defocus blur detection
Huafeng Li 0001 |
Neurocomputing | 2 |
| 2022 | Key point-aware occlusion suppression and semantic alignment for occluded person re-identification
Shujuan Wang, Bochun Huang, Huafeng Li 0001, Guanqiu Qi, Dapeng Tao, Zhengtao Yu 0001 |
Inf. Sci. | 3 |
| 2022 | Haze transfer and feature aggregation network for real-world single image dehazing
Huafeng Li 0001, Jirui Gao, Minghong Xie, Zhengtao Yu 0001 |
Knowl. Based Syst. | 1 |
| 2022 | Dual-stream Reciprocal Disentanglement Learning for domain adaptation person re-identification
Huafeng Li 0001, Kaixiong Xu, Jinxing Li 0003, Zhengtao Yu 0001 |
Knowl. Based Syst. | 1 |
| 2022 | Triple Adversarial Learning and Multi-View Imaginative Reasoning for Unsupervised Domain Adaptation Person Re-IdentificationabstractDue to the importance of practical applications, unsupervised domain adaptation (UDA) person re-identification (re-ID) has attracted increasing attention. However, most of existing methods often lack the multi-view information reasoning and ignore the domain discrepancy of the pedestrian images with the same identity, which constrain the further improvement of recognition performance. So, this paper proposes a triple adversarial learning and multi-view imaginative reasoning network (TAL-MIRN) for UDA person re-ID, which consists of a multi-view imaginative reasoning module (IRM) and a triple adversarial learning module (TALM). IRM makes the classified pedestrian identity features from a single-view image extracted by a feature encoder consistent with the classification results of the aggregated multi-view pedestrian identity features, so the strong multi-view imaginative reasoning ability of the feature encoder is obtained. TALM is composed by the adversarial learning between the camera classifier and feature encoder, adversarial learning of joint distribution alignment, and adversarial learning of the difference between two classifiers used in classification. In particular, the domain-invariant features at camera level are guaranteed by the adversarial learning between the feature extractor and camera classifier. The joint alignment of identity and domain is achieved by the competition between the feature extractor and classifier integrated with identity and domain. The discriminability and robustness of the learned features are enhanced by playing a MinMax game between two different identity classifiers. Furthermore, a simple normalization operation named as cross normalization (CN) is proposed to increase both modeling and generalization capability of the proposed TAL-MIRN across multiple domains. The proposed TAL-MIRN is applied to five benchmark datasets, and the comparative experimental results confirm its superiority over the state-of-the-art methods. The related source codes is available athttps://github.com/lhf12278/TALM-IRM. Huafeng Li 0001, Neng Dong, Zhengtao Yu 0001, Dapeng Tao, Guanqiu Qi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Body Part-Level Domain Alignment for Domain-Adaptive Person Re-Identification With Transformer FrameworkabstractAlthough existing domain-adaptive person re-identification (re-ID) methods have achieved competitive performance, most of them highly rely on the reliability of pseudo-label prediction, which seriously limits their applicability as noisy labels cannot be avoided. This paper designs a Transformer framework based on body part-level domain alignment to solve the above-mentioned issues in domain-adaptive person re-ID. Different parts of the human body (such as head, torso, and legs) have different structures and shapes. Therefore, they usually exhibit different characteristics. The proposed method makes full use of the dissimilarity between different human body parts. Specifically, the local features from the same body part are aggregated by the Transformer to obtain the corresponding class token, which is used as the global representation of this body part. Additionally, a Transformer layer-embedded adversarial learning strategy is designed. This strategy can simultaneously achieve domain alignment and classification of the class token for each human body part in both target and source domains by an integrated discriminator, thereby realizing domain alignment at human body part level. Compared with existing domain-level and identity-level alignment methods, the proposed method has a stronger fine-grained domain alignment capability. Therefore, the information loss or distortion that may occur in the feature alignment process can be effectively alleviated. The proposed method does not need to predict pseudo labels of any target sample, so the negative impact caused by unreliable pseudo labels on re-ID performance can be effectively avoided. Compared with state-of-the-art methods, the proposed method achieves better performance on the datasets that are in line with real-world scene settings. The source codes of this paper will be available at https://github.com/lhf12278/BPDA. Yiming Wang 0005, Guanqiu Qi, Yi Chai 0002, Huafeng Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | TDPN: Texture and Detail-Preserving Network for Single Image Super-ResolutionabstractSingle image super-resolution (SISR) using deep convolutional neural networks (CNNs) achieves the state-of-the-art performance. Most existing SISR models mainly focus on pursuing high peak signal-to-noise ratio (PSNR) and neglect textures and details. As a result, the recovered images are often perceptually unpleasant. To address this issue, in this paper, we propose a texture and detail-preserving network (TDPN), which focuses not only on local region feature recovery but also on preserving textures and details. Specifically, the high-resolution image is recovered from its corresponding low-resolution input in two branches. First, a multi-reception field based branch is designed to let the network fully learn local region features by adaptively selecting local region features in different reception fields. Then, a texture and detail-learning branch supervised by the textures and details decomposed from the ground-truth high resolution image is proposed to provide additional textures and details for the super-resolution process to improve the perceptual quality. Finally, we introduce a gradient loss into the SISR field and define a novel hybrid loss to strengthen boundary information recovery and to avoid overly smooth boundary in the final recovered high-resolution image caused by using only the MAE loss. More importantly, the proposed method is model-agnostic, which can be applied to most off-the-shelf SISR networks. The experimental results on public datasets demonstrate the superiority of our TDPN on most state-of-the-art SISR methods in PSNR, SSIM and perceptual quality. We will share our code on https://github.com/tocaiqing/TDPN. Jinxing Li 0003, Huafeng Li 0001, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Hazy Re-ID: An Interference Suppression Model for Domain Adaptation Person Re-Identification Under Inclement Weather ConditionabstractIn a conventional domain adaptation person Re-identification (Re-ID) task, both the training and test images in target domain are collected under the sunny weather. However, in reality, the pedestrians to be retrieved may be obtained under severe weather conditions such as hazy, dusty and snowing, etc. This paper proposes a novel Interference Suppression Model (ISM) to deal with the interference caused by the hazy weather in domain adaptation person Re-ID. A teacher-student model is used in the ISM to distill the interference information at the feature level by reducing the discrepancy between the clear and the hazy intrinsic similarity matrix. Furthermore, in the distribution level, the extra discriminator is introduced to assist the student model make the interference feature distribution more clear. The experimental results show that the proposed method achieves the superior performance on two synthetic datasets than the state-of-the-art methods. The related code will be released online https://github.com/pangjian123/ISM-ReID. Jian Pang, Dacheng Zhang, Huafeng Li 0001, Weifeng Liu 0001, Zhengtao Yu 0001 |
ICME | 3 |
| 2021 | Joint image fusion and super-resolution for enhanced visualization via semi-coupled discriminative dictionary learning and advantage embedding
Huafeng Li 0001, Moyuan Yang, Zhengtao Yu 0001 |
Neurocomputing | 1 |
| 2021 | Cross adversarial consistency self-prediction learning for unsupervised domain adaptation person re-identification
Huafeng Li 0001, Jian Pang, Dapeng Tao, Zhengtao Yu 0001 |
Inf. Sci. | 1 |
| 2021 | Attribute-Aligned Domain-Invariant Feature Learning for Unsupervised Domain Adaptation Person Re-IdentificationabstractDomain invariance and discrimination of learned features as two crucial factors affect the performance of unsupervised domain adaptation (UDA) person re-identification (Re-ID). Person attributes (such as “backpack”, “boots”, “handbag”, etc) remaining unchanged across multiple domains have been used as mid-level visual-semantic information in UDA person Re-ID. As two main challenges, both misalignment of attribute-related regions across multiple images and domain shift between source and target domains affect the learning of domain-invariant features (DIF). To address the above two challenges, this article proposes to take advantage of the stability of person attributes and the complementarity of person attributes and the corresponding low-level visual features to guide the learning of discriminative DIF. Specifically, the proposed solution contains the generation of latent attribute-correlated visual features (GLAVF), DIF learning under the guidance of person attributes, and the alignment of person attributes corresponding to the local regions of pedestrian images. Due to the gap between person attributes and visual features, person attributes are first converted into latent attribute-correlated visual features (LAVF) without any specific domain information in GLAVF, and then LAVF are used as the substitutions of person attributes to guide the learning of DIF. To enhance the discrimination of learned features, the proposed solution mainly explores the alignment between person attributes and corresponding local regions, and the alignment of the same person attributes across multiple pedestrian images. A fully connected layer is used to achieve the above two types of alignment in the proposed framework, which reduces the adverse impacts of inference information and ensures the semantic consistency between person attributes and corresponding local regions across multiple pedestrian images. The effectiveness of the proposed solution is confirmed on four existing datasets by comparative experiments. Huafeng Li 0001, Dapeng Tao, Zhengtao Yu 0001, Guanqiu Qi |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | Different Input Resolutions and Arbitrary Output Resolution: A Meta Learning-Based Deep Framework for Infrared and Visible Image FusionabstractInfrared and visible image fusion has gained ever-increasing attention in recent years due to its great significance in a variety of vision-based applications. However, existing fusion methods suffer from some limitations in terms of the spatial resolutions of both input source images and output fused image, which prevents their practical usage to a great extent. In this paper, we propose a meta learning-based deep framework for the fusion of infrared and visible images. Unlike most existing methods, the proposed framework can accept the source images of different resolutions and generate the fused image of arbitrary resolution just with a single learned model. In the proposed framework, the features of each source image are first extracted by a convolutional network and upscaled by a meta-upscale module with an arbitrary appropriate factor according to practical requirements. Then, a dual attention mechanism-based feature fusion module is developed to combine features from different source images. Finally, a residual compensation module, which can be iteratively adopted in the proposed framework, is designed to enhance the capability of our method in detail extraction. In addition, the loss function is formulated in a multi-task learning manner via simultaneous fusion and super-resolution, aiming to improve the effect of feature learning. And, a new contrast loss inspired by a perceptual contrast enhancement approach is proposed to further improve the contrast of the fused image. Extensive experiments on widely-used fusion datasets demonstrate the effectiveness and superiority of the proposed method. The code of the proposed method is publicly available at https://github.com/yuliu316316/MetaLearning-Fusion. Huafeng Li 0001, Yueliang Cen, Yu Liu 0023, Xun Chen 0001, Zhengtao Yu 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Person re-identification with dictionary learning regularized by stretching regularization and label consistency constraint
Huafeng Li 0001, Weiyan Zhou, Zhengtao Yu 0001, Huaiping Jin |
Neurocomputing | 1 |
| 2020 | Noise-robust image fusion with low-rank sparse decomposition guided by external patch prior
Huafeng Li 0001, Xiaoge He, Zhengtao Yu 0001, Jiebo Luo 0001 |
Inf. Sci. | 1 |
| 2020 | Structure alignment of attributes and visual features for cross-dataset person re-identification
Huafeng Li 0001, Zhenyu Kuang, Zhengtao Yu 0001, Jiebo Luo 0001 |
Pattern Recognit. | 1 |
| 2020 | Person re-identification by integrating metric learning and support vector machine
Hongbin Wang 0002, Hongpeng Yin, Zhengtao Yu 0001, Huafeng Li 0001 |
Signal Process. | 5 |
| 2020 | Attribute-Identity Embedding and Self-Supervised Learning for Scalable Person Re-IdentificationabstractDue to the domain shift between source dataset and target dataset, most of the existing person re-identification (PRID) algorithms trained by a supervised learning framework often fail to be well generalized to another domain. To address this challenge, we propose a self-supervised learning algorithm based on attribute-identity embedding, which can incrementally optimize the model by selecting unlabeled samples from target domain. Thus the gap between source domain and target domain is bridged. Specifically, we first develop an attribute-identity joint prediction dictionary learning model for simultaneously learning a latent attribute space, a semantic attribute dictionary and an identifier. In our method, the predicted attribute from latent attribute space is used as a bridge to establish a preliminary link between different domains so as to predict the label of the target data sample. Second, to exploit the latent label contained in the predicted samples, we propose a prediction-training cycle self-supervised learning to tune the model variables to make them more adaptive in the target domain. Finally, the similarity measurement of pedestrians is achieved by combining the attribute space with latent identity space. The experiments show that the developed method outperforms some state-of-the-art supervised PRID methods and unsupervised PRID algorithms. Huafeng Li 0001, Shuanglin Yan, Zhengtao Yu 0001, Dapeng Tao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Jointly Learning Commonality and Specificity Dictionaries for Person Re-IdentificationabstractDespite advances in person re-identification (re-ID), it is still far from meeting the needs of real-world applications due to tremendous visual ambiguity in person appearance across cameras. To overcome this problem, we propose a person re-ID method by decomposing a pedestrian's appearance feature into different components. This method assumes that each pedestrian image is composed of person-shared components that reflect the similarities of different pedestrians and person-specific components that reflect unique identity information. Based on this assumption, we propose to reduce the ambiguity in visual features by removing person-shared components from pedestrian visual features. To this end, we develop a framework for learning a pair of commonality and specificity dictionaries, while introducing a distance constraint to force the particularities of the same person over the specificity dictionary to have the same coding coefficients and the coding coefficients of different pedestrians to have weak correlation. Furthermore, considering the similarity of the commonality dictionary and the sparsity of the specificity dictionary, low-rank and sparse regularization terms are introduced into the dictionary learning framework to improve their representation ability and discriminative ability. Extensive experimental results show that the proposed algorithm outperforms or is competitive with the state-of-the-art methods. Huafeng Li 0001, Zhengtao Yu 0001, Jiebo Luo 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Top distance regularized projection and dictionary learning for person re-identification
Huafeng Li 0001, Jinting Zhu, Dapeng Tao, Zhengtao Yu 0001 |
Inf. Sci. | 1 |
| 2019 | Generative image deblurring based on multi-scaled residual adversary network driven by composed prior-posterior loss
Meng Wang 0024, Shengyu Hou, Huafeng Li 0001, Fan Li 0006 |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Joint medical image fusion, denoising and enhancement via discriminative low-rank sparse dictionaries learning
Huafeng Li 0001, Xiaoge He, Dapeng Tao, Yuan Yan Tang, Ruxin Wang 0002 |
Pattern Recognit. | 1 |
| 2017 | Multifocus image fusion via fixed window technique of multiscale images and non-local means filtering
Huafeng Li 0001, Hongmei Qiu, Zhengtao Yu 0001 |
Signal Process. | 1 |
| 2016 | Fractional differential and variational method for image fusion and super-resolution
Huafeng Li 0001, Zhengtao Yu 0001, Cunli Mao |
Neurocomputing | 1 |
| 2016 | Multifocus image fusion by combining with mixed-order structure tensors and multiscale neighborhood
Huafeng Li 0001, Xiaosong Li 0004, Zhengtao Yu 0001, Cunli Mao |
Inf. Sci. | 1 |
| 2016 | Performance improvement scheme of multifocus image fusion derived by difference images
Huafeng Li 0001, Xinkun Liu, Zhengtao Yu 0001 |
Signal Process. | 1 |
| 2013 | A new fusion scheme for multifocus images based on focused pixels detection
Huafeng Li 0001, Yi Chai 0003, Zhaofei Li |
Mach. Vis. Appl. | 1 |