EDBT 2026 Demo / reviewers in the wild / expert
Aihua Zheng
dblp:74/7436
· DBLP profile ↗
84ranked-venue papers
26as first author
59since 2021 · last 2026
0000-0002-9820-4743ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 13 first-author · 39 since 2021Artificial intelligence and machine learning · 39 · 15 first-author · 24 since 2021Security and privacy · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Teacher Interactive Knowledge Distillation Network for Text-to-Visible & Infrared Person RetrievalabstractText-to-visible & infrared person retrieval aims to retrieve the corresponding visible (RGB) and thermal infrared (TIR) images given the text descriptions. Existing methods perform semantic decoupling by aligning RGB and TIR features separately to different attributes, thereby facilitating the alignment between the fused multimodal representation and the text. However, insufficient TIR representation ability and cross-view representation capabilities of RGB and TIR modalities limit the retrieval accuracy and robustness. To address these issues, we propose a novel Dual-teacher Interactive Knowledge Distillation Network called DIKDNet, which performs the interactive knowledge distillation between two modality-specific teachers with rich cross-view representation capabilities to enhance TIR representations and the collaborative knowledge distillation from both teachers to the corresponding students to enhance the cross-modal cross-view representations, for robust text-to-visible & infrared person retrieval. Specifically, to enhance the representation ability of the TIR backbone network while preserving modality-specific characteristics, we design an Interactive Knowledge Distillation Module (IKDM), which introduces a boundary-constrained distillation strategy between RGB and TIR backbones, to transfer the semantic features of RGB backbone to TIR one. To enhance the cross-modal cross-view representation capability, we design a Collaborative Knowledge Distillation Module (CKDM) to transfer the cross-modal similarity relations and the cross-view multimodal representations from teacher networks to student ones. Experimental results demonstrate that our method consistently achieves significant performance gains on both the RGBT-PEDES and RGBNT201-PEDES datasets. The code will be released upon the acceptance. Chenglong Li 0002, Yifei Deng, Aihua Zheng |
AAAI | 4 |
| 2026 | ProxyTTT: Proxy-driven Test-Time Training for Multi-modal Re-identificationabstractMulti-modal object re-identification (ReID) aims to retrieve specific targets by leveraging complementary cues from different sensing modalities. Despite recent progress, two key challenges remain: (1) the limited ability to jointly address both modality and viewpoint discrepancies, and (2) the difficulty of effectively leveraging reliable target-domain data to improve generalization. To address these challenges, we propose Proxy-driven Test-Time Training (ProxyTTT), a unified framework that enhances both multi-modal identity representation learning and model generalization. During training, we propose a Multi-Proxy Learning (MPL) mechanism to address the representation bias across different views and modalities. MPL disentangles fine-grained modality-specific and modality-common identity proxies as semantic anchors to align identity features across diverse perspectives and sensing modalities. This alignment strategy enables the model to learn robust and discriminative global identity representations under heterogeneous modality conditions. At test time, to reliably exploit target domain data, we propose Proxy-guided Entropy-based Selective Adaptation (PESA) for test-time training. Specifically, PESA leverages the semantic structure encoded by identity proxies to estimate prediction uncertainty via entropy, and selectively adapts the model using only high-confidence samples. This selective adaptation effectively mitigates the domain shift between training and deployment environments, improving the model’s generalization in real-world scenarios. Extensive experiments on four public multi-modal ReID benchmarks (RGBNT201, RGBNT100, MSVR310, and WMVeID863) demonstrate the effectiveness of ProxyTTT. Aihua Zheng, Zhaojun Liu, Xixi Wan, Chenglong Li 0002, Jin Tang 0001, Yan Yan 0002 |
AAAI | 1 |
| 2026 | Progressive Multi-modal Knowledge Distillation for Multi-spectral Object Re-identificationabstractIn the field of multi-spectral object re-identification (ReID), multi-modal knowledge and modal-specific knowledge exhibit complementary advantages when handling hard samples, but existing methods rarely integrate this collaborative information. Knowledge distillation is a direct approach for transferring information, however, heterogeneity in model architectures and variations in sample hardness can undermine the stability and controllability of knowledge transfer. To alleviate these limitations, we propose the novel Progressive Multi-modal Knowledge Distillation (PMKD) framework that enables multi-stage knowledge transfer guided by hard sample awareness. In the multi-modal knowledge transfer stage, the source model (pre-trained on multi-modal data) disseminates its learned multi-modal collaborative knowledge to multiple independently modal-specific target models, guiding their adaptation to hard samples within training batches. In the modal-specific knowledge retention stage, the independent models enriched with multi-modal knowledge guide the training phase. The architectural consistency between source-target models ensures more lossless knowledge transfer, effectively mitigating the risk of capability drift, and preserving inherent competence. Moreover, the entire progressive multi-modal knowledge distillation is regulated by the proposed hardness-aware distillation loss, which automatically adapts distillation intensity through hard sample mining, thereby ensuring stable transfer of hard sample handling capabilities. Extensive experiments on benchmark multi-spectral ReID datasets validate the effectiveness and superior performance of the proposed method. Aihua Zheng, Zi Wang 0013, Jin Tang 0001 |
AAAI | 1 |
| 2026 | Semantic-Driven Visual Progressive Refinement for Aerial-Ground Person ReID: A Challenging Large-Scale BenchmarkabstractAerial-Ground Person Re-IDentification (AGPReID) aims to extract identity-discriminative representations from heterogeneous perspectives across different platforms in complex real-world environments. However, existing methods primarily focus on visual appearance modeling and make insufficient use of semantic attribute priors, which limits their ability to bridge the aerial-ground view gap. To address this limitation, we propose a Semantic-driven Visual Progressive Refinement framework for AGPReID (SVPR-ReID), which effectively leverages textual attribute priors to guide the extraction of fine-grained visual cues. Specifically, we design a View-Decoupled Feature Extractor that incorporates view-aware textual prompts to decouple view-invariant identity features. Then, to alleviate inter-class ambiguity, we propose an Attribute-Scattered Mixture-of-Experts module that integrates attribute semantics into the visual space, thereby improving discrimination among visually similar pedestrians. Finally, we design a Context-Vision Progressive Refinement module for progressive refinement of attribute and view-invariant features, obtaining robust cross-view identity representations. In particular, we contribute a comprehensive benchmark for AGPReID, named CP2108, which contains 142,817 images of 2,108 identities annotated with 22 attributes. Notably, it includes 191 identities captured across different times, enabling both short- and long-term ReID evaluation, addressing the limitation of existing datasets that focus only on short-term scenarios. Extensive experimental results validate the effectiveness of our SVPR-ReID on four AGPReID datasets. Aihua Zheng, Xixi Wan, Zi Wang 0013, Jin Tang 0001, Bin Luo 0001 |
AAAI | 1 |
| 2026 | Multi-level alignment network for unsupervised domain adaptive multi-modality object re-identification
Yusong Sheng, Yuhe Ding, Aihua Zheng, Zi Wang 0013, Jin Tang 0001 |
Knowl. Based Syst. | 3 |
| 2026 | Harmonizing class uniformity and separability for transferability estimation
Yuhe Ding, Bo Jiang 0002, Lijun Sheng, Aihua Zheng, Jian Liang 0001 |
Pattern Recognit. | 4 |
| 2026 | SequencePAR: Understanding pedestrian attributes via a sequence generation paradigm
Jiandong Jin, Xiao Wang 0014, Yin Lin, Chenglong Li 0002, Lili Huang 0006, Aihua Zheng, Jin Tang 0001 |
Pattern Recognit. | 6 |
| 2026 | Bidirectional intervention attention network for audio-visual matching
Jiaxiang Wang 0001, Aihua Zheng, Dequan Li, Chenglong Li 0002, Wenjuan Cheng, Ran He 0001 |
Pattern Recognit. | 2 |
| 2026 | Fine-Grained and Granularity-Dynamic Framework for Referring Remote Sensing Image Segmentation
Duzhi Yuan, Guyue Hu 0001, Aihua Zheng, Chenglong Li 0002, Jin Tang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2026 | Reliable Multi-Modal Object Re-Identification via Modality-Aware Graph ReasoningabstractMulti-modal data provides abundant and diverse object information, crucial for effective modal interactions in Re-Identification (ReID) task. However, existing approaches often overlook the quality variations in local features and fail to fully leverage the complementary information across modalities, particularly in cases where features are of low quality. In this paper, we propose to address this issue by leveraging a novel graph reasoning model, termed the Modality-aware Graph Reasoning Network (MGRNet). Specifically, we first construct modality-aware graphs to enhance the extraction of fine-grained local details by effectively capturing and modeling the relationships between patches. Subsequently, the selective graph nodes swap operation is employed to alleviate the adverse effects of low-quality local features by considering both local and global information, enhancing the representation of discriminative information. Finally, the swapped modality-aware graphs are fed into the local-aware graph reasoning module, which propagates multi-modal information to yield a reliable feature representation. Another advantage of the proposed graph reasoning approach is its ability to reconstruct missing modal information by exploiting inherent structural relationships, thereby minimizing disparities between different modalities. Experimental results on four benchmarks (RGBNT201, Market1501-MM, RGBNT100, MSVR310) indicate that the proposed method achieves state-of-the-art performance in multi-modal object ReID. The code for our method will be available upon acceptance. Xixi Wan, Aihua Zheng, Zi Wang 0013, Bo Jiang 0002, Jin Tang 0001, Jixin Ma 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | Text-Visible/Infrared Person Retrieval: Attribute-Guided Feature Decoupling and Collaborative Alignment and a Unified BenchmarkabstractExisting research on text-to-image person retrieval primarily focuses on visible images, which are not suitable under low-light scenarios. Infrared imaging becomes necessary in many visual systems, and matching text with both visible and infrared images is required. However, visible and infrared images are heterogeneous with different visual characteristics, so matching text with them in a unified framework is very challenging. In this work, we design a new task called Text-Visible/Infrared person retrieval and contribute a novel approach and a unified benchmark to promote the research and development of this field. On one hand, we propose a novel Attribute-guided feature decoupling and Collaborative Alignment Network (ACANet) that pursues accurate alignment from the text modality to both visible and infrared modalities in a unified framework according to the texture and color attribute information of text descriptions. In particular, we decouple the color features of visible images supervised by the text labels and integrate them into the infrared features to eliminate the impact of the absence of color information in infrared images during cross-modal collaborative alignment. Moreover, we also decouple the texture information from visible images supervised by the text labels and perform the collaborative alignment of texture and infrared features with a fusion agent. In addition, we extend conventional masked language modeling to a cross-modal paradigm to help ACANet learn uniform fine-grained alignment in multiple image modalities. On the other hand, we contribute a unified high-quality MM01LLCM-Text dataset, which provides person images in both visible and infrared modalities paired with fine-grained text descriptions. Experimental results show that the proposed ACANet outperforms existing state-of-the-art methods on MM01LLCM-Text dataset. Chenglong Li 0002, Yifei Deng, Aihua Zheng, Jin Tang 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | REMIND: Retrieval-Augmented Reconstruction With Dual Memories for Modality-Missing Object Re-IdentificationabstractTo address the modality-missing object Re-Identification (Re-ID) task, a common strategy is to compensate for absent information by exploiting available modalities. However, existing reconstruction-based approaches suffer from two major limitations: 1) they often overlook modality-specific cues inherent in the missing modality; 2) they typically adopt a single-path reconstruction strategy. These issues result in incomplete representations and constrain the capacity to model complex semantic mappings across heterogeneous modalities. To address these challenges, we propose REMIND, a novel framework for modality-missing object Re-Identification, namely REtrieval-AugMented ReconstructIoN With Dual Memories. Specifically, we design a Dual Memory Construction module that, guided by information-theoretic insights, extracts modality-specific and modality-common features through two complementary branches and stores them in dedicated memory banks. These memory banks serve as structured prior knowledge to guide the reconstruction process, ensuring that the features of missing modalities are preserved even under modality-missing conditions. In addition, we have developed a retrieval-augmented missing reconstruction module that enhances the expressiveness and robustness of the reconstruction through multi-path reconstruction and perturbation mechanisms. Adaptive fusion techniques are employed for integration, simultaneously improving the expressiveness and robustness of the reconstructed features. Through the synergy of information-theoretically motivated regularization and retrieval-enhanced reconstruction, REMIND achieves robust feature recovery and delivers highly discriminative representations for reliable modality-missing Re-ID. Extensive experiments on several multi-modal object Re-ID benchmarks demonstrate the effectiveness and superiority of REMIND under various missing modality scenarios. The code is publicly available at: https://github.com/skye-1201/REMIND. Zhendong Xu, Zi Wang 0013, Aihua Zheng, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Ranking Vision-Language Models in Fully Unlabeled TasksabstractVision language models (VLMs) like CLIP show stellar zero-shot capability on classification benchmarks. However, selecting the VLM with the highest performance on the unlabeled downstream task is non-trivial. Existing VLM selection methods focus on the class-name-only setting, relying on supervised auxiliary datasets and large language models, which may not be accessible or feasible during deployment. This paper introduces the problem ofunsupervised vision-language model selection, where only unsupervised downstream datasets are available, with no additional information provided. To solve this problem, we propose a method termed Visual-tExtual Graph Alignment (VEGA), to select VLMs without any annotations by measuring the alignment of the VLM between the two modalities on the downstream task. VEGA is motivated by the pretraining paradigm of VLMs, which aligns features with the same semantics from the visual and textual modalities, thereby mapping both modalities into a shared representation space. Specifically, we first construct two graphs on the vision and textual features, respectively. VEGA is then defined as the overall similarity between the visual and textual graphs at both node and edge levels. Extensive experiments across three different benchmarks, covering a variety of application scenarios and downstream datasets, demonstrate that VEGA consistently provides reliable and accurate estimates of VLMs' performance on unlabeled downstream tasks. Yuhe Ding, Bo Jiang 0002, Aihua Zheng, Jian Liang 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification
Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | DEEP: Decoupled Semantic Prompt Learning, Guiding and Embedding for Multi-Spectral Object Re-IdentificationabstractMulti-spectral object re-identification (ReID) captures diverse object semantics to robustly recognize identity in complex environments. However, without explicit semantic guidance (e.g., attributes, masks, and keypoints), existing modal fusion-based methods struggle to comprehensively capture person or vehicle semantics across spectra. Thanks to the large-scale vision-language pre-training, CLIP effectively aligns visual concepts across different image modalities to a unified semantic prompt. In this paper, we proposeDEEP, aDEcoupled sEmanticPrompt Learning, Guiding and Embedding framework for Multi-Spectral Object ReID. Specifically, to address the challenges posed by low-quality modality noise and spectral style discrepancies, we first propose a Decoupled Semantic Prompt (DSP) strategy, which explicitly decouples the semantic alignment into spectral-style learning with spectral-shared prompts and object content learning with instance-specific inversion token. Second, to lead the model focusing on semantically faithful regions, we propose a Semantic-Guided Spectral Fusion (SGSF) module that builds a semantic interaction bridge between spectra to explore complementary semantics across modalities. Finally, to further empower the spectral representation, we propose a Spectral Semantic Embedding (SSE) module constrained by semantic-aware structural consistency to refine the fine-grained identity semantics in each spectrum. Extensive experiments on five public benchmarks, RGBNT201, Market-MM, MSVR310, WMVEID863, and RGBNT100, demonstrate the proposed method outperforms the state-of-the-art methods. The source code is released at this link:https://github.com/lsh-ahu/DEEP-ReID. Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | MambaPro: Multi-Modal Object Re-identification with Mamba Aggregation and Synergistic PromptabstractMulti-modal object Re-IDentification (ReID) aims to retrieve specific objects by utilizing complementary image information from different modalities. Recently, large-scale pre-trained models like CLIP have demonstrated impressive performance in traditional single-modal ReID tasks. However, they remain unexplored for multi-modal object ReID. Furthermore, current multi-modal aggregation methods have obvious limitations in dealing with long sequences from different modalities. To address above issues, we introduce a novel framework called MambaPro for multi-modal object ReID. To be specific, we first employ a Parallel Feed-Forward Adapter (PFA) for adapting CLIP to multi-modal object ReID. Then, we propose the Synergistic Residual Prompt (SRP) to guide the joint learning of multi-modal features. Finally, leveraging Mamba's superior scalability for long sequences, we introduce Mamba Aggregation (MA) to efficiently model interactions between different modalities. As a result, MambaPro could extract more robust features with lower complexity. Extensive experiments on three multi-modal object ReID benchmarks (i.e., RGBNT201, RGBNT100 and MSVR310) validate the effectiveness of our proposed methods. Xuehu Liu, Tianyu Yan, Aihua Zheng, Huchuan Lu |
AAAI | 5 |
| 2025 | DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-IdentificationabstractMulti-modal object Re-IDentification (ReID) aims to retrieve specific objects by combining complementary information from multiple modalities. Existing multi-modal object ReID methods primarily focus on the fusion of heterogeneous features. However, they often overlook the dynamic quality changes in multi-modal imaging. In addition, the shared information between different modalities can weaken modality-specific information. To address these issues, we propose a novel feature learning framework called DeMo for multi-modal object ReID, which adaptively balances decoupled features using a mixture of experts. To be specific, we first deploy a Patch-Integrated Feature Extractor (PIFE) to extract multi-granularity and multi-modal features. Then, we introduce a Hierarchical Decoupling Module (HDM) to decouple multi-modal features into non-overlapping forms, preserving the modality uniqueness and increasing the feature diversity. Finally, we propose an Attention-Triggered Mixture of Experts (ATMoE), which replaces traditional gating with dynamic attention weights derived from decoupled features. With these modules, our DeMo can generate more robust multi-modal features. Extensive experiments on three object ReID benchmarks verify the effectiveness of our methods. Aihua Zheng |
AAAI | 3 |
| 2025 | Dual-PST: Dual-Branch SpatioTemporal-Planar Network for Video Forgery DetectionabstractWith the advancement of generative AI, distinguishing real and AI-generated faces in videos has become increasingly challenging. However, traditional methods struggle to capture local details and temporal dynamics simultaneously, making it difficult to achieve high detection accuracy while maintaining low computational overhead. To address this problem, we propose a Dual-branch SpatioTemporal-Planar Network (Dual-PST) based on the selective state-space model. It is capable of extracting image features and temporal relations simultaneously, while maintaining linear computational consumption. Specifically, we design a Multi-Selective State-Space module (MS3) that can extract global features from image typography consisting of consecutive video frames by scanning them in multiple sequences. To further enhance temporal modeling capabilities, we propose a Sequential Tri-frame Local module, which captures inter-frame temporal relationships and local features by temporally splicing single-frame features. These features are first extracted using MS3 and then further enhanced through inter-frame masking operations. Experimental results show that Dual-PST significantly improves detection accuracy while maintaining low computational complexity and strong model robustness. Junxian Duan, Jie Cao 0002, Aihua Zheng |
ICASSP | 5 |
| 2025 | Growing to Detect: A Dynamic Prototype Tree with Structured Replay for Incremental Deepfake DetectionabstractThe rapid advancement of deepfake technology poses significant threats to social trust. Recent research has improved detectors by adapting to emerging deepfakes using a limited number of samples through incremental learning. However, these approaches often overlook the scarcity of novel samples, resulting in insufficient learning of forgery patterns. To overcome this challenge, we propose a Replay-based Dynamic Prototype Network that integrates two key modules: the Dynamic Prototype Tree (DPT) module and the Similarity Subtree Replay (SSR) strategy. The DPT module dynamically introduces prototypes through a hierarchical tree structure to effectively adapt to new deepfakes. It expands prototypes based on similarity, thereby retaining the knowledge learned from previous prototypes while learning new forgery patterns. The SSR strategy mitigates catastrophic forgetting by stabilizing learned features through the replay of relevant subtrees. Experimental results demonstrate that our approach outperforms existing methods across five datasets, particularly on high-quality face swap samples generated by diffusion-based methods, achieving an AUC of 85.86% on the cross-dataset task from FaceForensics++ to DiffSwap. Junxian Duan, Jie Cao 0002, Aihua Zheng, Ran He 0001 |
IJCB | 4 |
| 2025 | Feature Decoupling with Modality Modulation for Multimodal Sentiment Analysis
Jiaxiang Wang 0001, Aihua Zheng, Wenjuan Cheng, Xiaofei Sheng |
ICIG (2) | 3 |
| 2025 | Modality Modulation with Adaptive Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis aims at extracting effective information from different modalities such as text, audio, and visual to infer the speaker’s sentiment state. Due to the high heterogeneity across modalities, most existing approaches decouple modalities into specific and invariant features, which can capture effective cross-modal representations to some extent. However, in multimodal tasks, different modalities exhibit varying strengths and weaknesses, with strong modalities dominating the overall optimization direction of the network, leading to under-optimization of weak modalities. To address the modality imbalance issue, we propose a Modality Modulation Adaptive Fusion Network (MMAFNet) to optimize the learning of valuable information from each modality. Specifically, for modality-specific features, we design a specific feature gradient modulation strategy to stimulate the weak modalities learning and adaptively modulate the corresponding gradients by measuring different importance to better optimize each modality. For modality-invariant features, in terms of the distance between different modalities, we propose an invariant feature parameter reset strategy that prevents overfitting irrelevant information while enhancing the feature extraction capability of weaker modalities. Finally, we incorporate an adaptive fusion module to combine modality-specific and invariant features based on their respective weights. Overall, we analyze the characteristics of various features and propose modality modulation strategies that mitigate modality imbalance. Extensive experiments on two multimodal sentiment analysis datasets, demonstrate the superior performance of our method. Aihua Zheng, Jiaxiang Wang 0001, Xiaofei Sheng, Wenjuan Cheng |
IJCNN | 1 |
| 2025 | Modality Modulation with Adaptive Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis aims at extracting effective information from different modalities such as text, audio, and visual to infer the speaker’s sentiment state. Due to the high heterogeneity across modalities, most existing approaches decouple modalities into specific and invariant features, which can capture effective cross-modal representations to some extent. However, in multimodal tasks, different modalities exhibit varying strengths and weaknesses, with strong modalities dominating the overall optimization direction of the network, leading to under-optimization of weak modalities. To address the modality imbalance issue, we propose a Modality Modulation Adaptive Fusion Network (MMAFNet) to optimize the learning of valuable information from each modality. Specifically, for modality-specific features, we design a specific feature gradient modulation strategy to stimulate the weak modalities learning and adaptively modulate the corresponding gradients by measuring different importance to better optimize each modality. For modality-invariant features, in terms of the distance between different modalities, we propose an invariant feature parameter reset strategy that prevents overfitting irrelevant information while enhancing the feature extraction capability of weaker modalities. Finally, we incorporate an adaptive fusion module to combine modality-specific and invariant features based on their respective weights. Overall, we analyze the characteristics of various features and propose modality modulation strategies that mitigate modality imbalance. Extensive experiments on two multimodal sentiment analysis datasets, demonstrate the superior performance of our method. Aihua Zheng, Jiaxiang Wang 0001, Xiaofei Sheng, Wenjuan Cheng |
IJCNN | 1 |
| 2025 | SPromptGL: Semantic Prompt Guided Graph Learning for Multi-modal Brain Disease
Xixi Wan, Bo Jiang 0002, Aihua Zheng |
MICCAI (12) | 4 |
| 2025 | UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-IdentificationabstractMulti-modal object Re-IDentification (ReID) has gained considerable attention with the goal of retrieving specific targets across cameras using heterogeneous visual data sources. At present, multi-modal object ReID faces two core challenges: (1) learning robust features under fine-grained local noise caused by occlusion, frame loss, and other disruptions; and (2) effectively integrating heterogeneous modalities to enhance multi-modal representation. To address the above challenges, we propose a robust approach named Uncertainty-Guided Graph model for multi-modal object ReID (UGG-ReID). UGG-ReID is designed to mitigate noise interference and facilitate effective multi-modal fusion by estimating both local and sample-level aleatoric uncertainty and explicitly modeling their dependencies. Specifically, we first propose the Gaussian patch-graph representation model that leverages uncertainty to quantify fine-grained local cues and capture their structural relationships. This process boosts the expressiveness of modal-specific information, ensuring that the generated embeddings are both more informative and robust. Subsequently, we design an uncertainty-guided mixture of experts strategy that dynamically routes samples to experts exhibiting low uncertainty. This strategy effectively suppresses noise-induced instability, leading to enhanced robustness. Meanwhile, we design an uncertainty-guided routing to strengthen the multi-modal interaction, improving the performance. UGG-ReID is comprehensively evaluated on five representative multi-modal object ReID datasets, encompassing diverse spectral modalities. Experimental results show that the proposed method achieves excellent performance on all datasets and is significantly better than current methods in terms of noise immunity. Our code is available at https://github.com/wanxixi11/UGG-ReID. Xixi Wan, Aihua Zheng, Bo Jiang 0002, Beibei Wang 0006, Chenglong Li 0002, Jin Tang 0001 |
NeurIPS | 2 |
| 2025 | Text-Guided Noise Replacement Visual Prompt Learning for Vision-Language Models
Xiaokang Shao, Mengjin Liu, Zhaojun Liu, Junxian Duan, Aihua Zheng |
PRCV (2) | 6 |
| 2025 | Frustratingly Easy Feature Reconstruction for Out-of-Distribution Detection
Yingsheng Wang, Shuo Lu, Jian Liang 0001, Aihua Zheng, Ran He 0001 |
PRCV (9) | 4 |
| 2025 | THGS: Lifelike Talking Human Avatar Synthesis From Monocular Video Via 3D Gaussian SplattingabstractAbstract Despite the remarkable progress in 3D talking head generation, directly generating 3D talking human avatars still suffers from rigid facial expressions, distorted hand textures and out‐of‐sync lip movements. In this paper, we extend speaker‐specific talking head generation task to talking human avatar synthesis and propose a novel pipeline, THGS, that animates lifelike Talking Human avatars using 3D Gaussian Splatting (3DGS). Given speech audio, expression and body poses as input, THGS effectively overcomes the limitations of 3DGS human re‐construction methods in capturing expressive dynamics, such as mouth movements, facial expressions and hand gestures, from a short monocular video. Firstly, we introduce a simple yet effective Learnable Expression Blendshapes (LEB) for facial dynamics re‐construction, where subtle facial dynamics can be generated by linearly combining the static head model and expression blendshapes. Secondly, a Spatial Audio Attention Module (SAAM) is proposed for lip‐synced mouth movement animation, building connections between speech audio and mouth Gaussian movements. Thirdly, we employ a body pose, expression and skinning weights joint optimization strategy to optimize these parameters on the fly, which aligns hand movements and expressions better with video input. Experimental results demonstrate that THGS can achieve high‐fidelity 3D talking human avatar animation at 150+ fps on a web‐based rendering system, improving the requirements of real‐time applications. Our project page is at https://sora158.github.io/THGS.github.io/ . Lingyun Yu 0002, Quanwei Yang, Aihua Zheng, Hongtao Xie 0001 |
Comput. Graph. Forum | 4 |
| 2025 | Keypoint-guided feature enhancement and alignment for cross-resolution vehicle re-identification
Aihua Zheng, Zi Wang 0013, Chenglong Li 0002, Xiaofei Sheng |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Prompt-Based Cross-Modal Feature Alignment for Weakly Supervised IFERabstractInfrared Facial Expression Recognition (IFER) encounters challenges in data acquisition and annotation under low-light conditions, making fully supervised training difficult. Although pre-trained Vision-Language Models (VLMs) can enhance generalization for downstream tasks, their insufficient attention modeling in cross-domain scenarios leads to ineffective local semantic correlation. To address this, we propose a Prompt-based Cross-modal feature Alignment (PCA) method that improves weakly supervised IFER performance by leveraging RGB facial expression data. The PCA framework comprises two key components: (1) a Cross-modal Prompt Transfer (CPT) strategy that integrates category-specific information to distinguish expressions, and (2) an Image-Guided Alignment (IGA) module that achieves feature alignment using dual-domain feature banks. Experimental results on two benchmark datasets demonstrate that our method significantly outperforms current state-of-the-art approaches, confirming its effectiveness and superiority. Hanqin Shi, Xiaofeng Kang, Jiaxiang Wang 0001, Aihua Zheng, Wenjuan Cheng |
IEEE Signal Process. Lett. | 4 |
| 2025 | Prototype-Based Diversity and Integrity Learning for All-Day Multi-Modal Person Re-IdentificationabstractRecent multi-modal person re-identification methods have improved model performance by leveraging complementary information from multiple spectra. However, existing methods cannot ensure feature stability under varying illumination and rely on inflexible paired data, remaining inadequate against real-world cross-time retrieval and modality-missing challenges. To solve these, we first propose diversity representation that augments illumination-sensitive images to simulate diverse lighting conditions via illumination augmentation and enriches instance features using modality-specific prototypes via multiple interaction modules. Secondly, we propose integrity reconstruction that leverages prototypes and available instance features to recover information, the reconstruction module effectively utilizes identity and modality cues to address unpredictable missing problems. In addition, we build a more comprehensive dataset (AllDay843) to alleviate the inadequate dataset diversity, which comprises 91,371 images of 843 identities captured by multi-modal cameras across various periods throughout the day, while incorporating numerous real-world challenges. By integrating diversity representation and integrity reconstruction, the proposed Prototype-Based Diversity and Integrity learning network (PDINet) establishes excellence on the AllDay843 dataset, surpassing existing state-of-the-art approaches. The data and codes are available in https://github.com/ziwang1121/PDINet. Zi Wang 0013, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Adaptive Interaction and Correction Attention Network for Audio-Visual MatchingabstractAudio-visual matching techniques aim to recognize and match information across different identities by learning a similarity metric across modalities. However, modal differences arise from insufficient cross-modal correlations and noise interference, which substantially hinder the performance of traditional deep metric learning methods in audio-visual matching tasks. To address the modal differences issue, we propose a novel Adaptive Interactive and Correction Attention Network (AICANet). This network efficiently captures deep information connections, generating modality-consistent feature embeddings within a unified metric framework. The core of AICANet is its two-pronged approach to reducing modal differences. First, we propose the Adaptive Interactive Attention (AIA) module, which flexibly establishes associations among cross-modal local features using dynamically generated pseudo-labels. Second, we propose the Adaptive Correction Attention (ACA) mechanism, which employs an adaptive threshold to de-interference effectively and accurately adjust the representation of local feature associations. Notably, the ACA mechanism is suitable for both intra-modal and inter-modal refined attention correction. Additionally, we design a relative distance stretching metric loss (LRDSM), which reinforces the similarity invariance of feature embeddings in a uniform space and enhances matching accuracy. Extensive tests on the VoxCeleb and VoxCeleb2 datasets demonstrate that AICANet outperforms leading existing algorithms across several evaluation metrics, validating its superior performance. The codes can be found at https://github.com/w1018979952/AICANet. Jiaxiang Wang 0001, Aihua Zheng, Lei Liu 0049, Chenglong Li 0002, Ran He 0001, Jin Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Heterogeneous Test-Time Training for Multi-Modal Person Re-identificationabstractMulti-modal person re-identification (ReID) seeks to mitigate challenging lighting conditions by incorporating diverse modalities. Most existing multi-modal ReID methods concentrate on leveraging complementary multi-modal information via fusion or interaction. However, the relationships among heterogeneous modalities and the domain traits of unlabeled test data are rarely explored. In this paper, we propose a Heterogeneous Test-time Training (HTT) framework for multi-modal person ReID. We first propose a Cross-identity Inter-modal Margin (CIM) loss to amplify the differentiation among distinct identity samples. Moreover, we design a Multi-modal Test-time Training (MTT) strategy to enhance the generalization of the model by leveraging the relationships in the heterogeneous modalities and the information existing in the test data. Specifically, in the training stage, we utilize the CIM loss to further enlarge the distance between anchor and negative by forcing the inter-modal distance to maintain the margin, resulting in an enhancement of the discriminative capacity of the ultimate descriptor. Subsequently, since the test data contains characteristics of the target domain, we adapt the MTT strategy to optimize the network before the inference by using self-supervised tasks designed based on relationships among modalities. Experimental results on benchmark multi-modal ReID datasets RGBNT201, Market1501-MM, RGBN300, and RGBNT100 validate the effectiveness of the proposed method. The codes can be found at https://github.com/ziwang1121/HTT. Zi Wang 0013, Huaibo Huang, Aihua Zheng, Ran He 0001 |
AAAI | 3 |
| 2024 | Day-Night Cross-domain Vehicle Re-identificationabstractPrevious advances in vehicle re-identification (ReID) are mostly reported under favorable lighting conditions, while cross-day-and-night performance is neglected, which greatly hinders the development of related traffic intelli-gence applications. This work instead develops a novel Day-Night Dual-domain Modulation (DNDM) vehicle re-identification framework for day-night cross-domain traf-fic scenarios. Specifically, a unique night-domain glare suppression module is provided to attenuate the headlight glare from raw nighttime vehicle images. To enhance ve-hicle features under low-light environments, we propose a dual-domain structure enhancement module in the feature extractor, which enhances geometric structures between ap-pearance features. To alleviate day-night domain discrep-ancies, we develop a cross-domain class awareness mod-ule that facilitates the interaction between appearance and structure features in both domains. In this work, we ad-dress the Day-Night cross-domain ReID (DN-ReID) prob-lem and provide a new cross-domain dataset named DN-Wild, including day and night images of 2,286 identities, giving in total 85,945 daytime images and 54,952 nighttime images. Furthermore, we also take into account the mat-ter of balance between day and night samples, and provide a dataset called DN-348. Exhaustive experiments demon-strate the robustness of the proposed framework in the DN-ReID problem. The code and benchmark are released at https://github.com/chenjingong/DN-ReID. Jingong Chen, Aihua Zheng, Yong Wu 0006, Yonglong Luo |
CVPR | 3 |
| 2024 | Semantic-Aware Detail Enhancement for Blind Face RestorationabstractThe goal of Blind Face Restoration is to recover high-quality images from low-quality images suffering from unknown degradations, posing a significantly challenging problem. In recent years, numerous BFR methods have been proposed, achieving significant success. However, faces possess a unique facial topology, and subtle differences in texture, slight structural imbalances, and minimal asymmetry are easily perceptible in the restored face images. Previous methods often struggle to generate realistically high-quality images from real-world low-quality images and fail to preserve fine features. To more effectively restore image details and textures, providing a more natural and realistic restoration effect, we integrate facial semantic information as prior knowledge into the blind face restoration task. We employ a multi-head cross-attention mechanism to simultaneously consider facial semantic information and context information for modeling. Additionally, we introduce a local detail enhancement module specifically designed to enhance the processing capability of details around the eyes and mouth. Experimental results indicate that our proposed method recovers facial images on synthetic and real datasets more realistically and with higher fidelity. Xiaoqiang Zhou, Jie Cao 0002, Huaibo Huang, Aihua Zheng, Ran He 0001 |
FG | 5 |
| 2024 | Parallel Augmentation and Dual Enhancement for Occluded Person Re-IdentificationabstractOccluded person re-identification (Re-ID), the task of searching for the same person’s images in occluded environments, has attracted lots of attention in the past decades. Recent approaches concentrate on improving performance on occluded data by data/feature augmentation or using extra models to predict occlusions. However, they ignore the imbalance problem in this task and can not fully utilize the information from the training data. To alleviate these two issues, we propose a simple yet effective method with Parallel Augmentation and Dual Enhancement (PADE), which is robust on both occluded and non-occluded data and does not require any auxiliary clues. First, we design a parallel augmentation mechanism (PAM) to generate more suitable occluded data to mitigate the negative effects of unbalanced data. Second, we propose the global and local dual enhancement strategy (DES) to promote the context information and details. Experimental results on three widely used occluded datasets and two non-occluded datasets validate the effectiveness of our method. The code is available at PADE (GitHub). Zi Wang 0013, Huaibo Huang, Aihua Zheng, Chenglong Li 0002, Ran He 0001 |
ICASSP | 3 |
| 2024 | Disentangled generation network for enlarged license plate recognition and a unified dataset
Chenglong Li 0002, Xiaobin Yang, Guohao Wang, Aihua Zheng, Jin Tang 0001 |
Comput. Vis. Image Underst. | 4 |
| 2024 | MAPS: A Noise-Robust Progressive Learning Approach for Source-Free Domain Adaptive Keypoint DetectionabstractExisting cross-domain keypoint detection methods always require accessing the source data during adaptation, which may violate the data privacy law and pose serious security concerns. Instead, this paper considers a realistic problem setting called source-free domain adaptive keypoint detection, where only the well-trained source model is provided to the target domain. For the challenging problem, we first construct a teacher-student learning baseline by stabilizing the predictions under data augmentation and network ensembles. Built on this, we further propose a unified approach, Mixup Augmentation and Progressive Selection (MAPS), to fully exploit the noisy pseudo labels of unlabeled target data during training. On the one hand, MAPS regularizes the model to favor simple linear behavior in-between the target samples via self-mixup augmentation, preventing the model from over-fitting to noisy predictions. On the other hand, MAPS employs the self-paced learning paradigm and progressively selects pseudo-labeled samples from ‘easy’ to ‘hard’ into the training process to reduce noise accumulation. Results on four keypoint detection datasets show that MAPS outperforms the baseline and achieves comparable or even better results in comparison to previous non-source-free counterparts. The code is available athttps://github.com/YuheD/MAPS. Yuhe Ding, Jian Liang 0001, Bo Jiang 0002, Aihua Zheng, Ran He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Public-Private Attributes-Based Variational Adversarial Network for Audio-Visual Cross-Modal MatchingabstractExisting audio-visual cross-modal matching methods focus on mitigating cross-modal heterogeneity but ignore the impact of intra-class discrepancy of the same identity in different scenarios, which might greatly limit the matching performance. To simultaneously handle both problems of intra-class discrepancy and cross-modal heterogeneity, we propose a novel public-private attributes-based variational adversarial network (P2VANet), which captures the consistency within and between classes, for audio-visual cross-modal matching. In particular,P2VANet first uses a variational auto-encoder, which captures the inherent global information in diverse scenarios from the hidden variable through reconstruction, to reduce the intra-class discrepancy. Then it integrates a public attributes guidance module to capture the consistency of audio and visual by supervision of the common high-level semantic information to mitigate cross-modal heterogeneity. In addition,P2VANet designs private attributes embedding module to enhance the discriminative features inherent in each class to decrease inter-class similarity. Extensive experiments on audio-visual cross-modal matching demonstrate the effectiveness of the proposed approach compared with the state-of-the-art methods. Aihua Zheng, Jiaxiang Wang 0001, Chao Tang 0002, Chenglong Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Attribute-Guided Cross-Modal Interaction and Enhancement for Audio-Visual MatchingabstractAudio-visual matching is an essential task that measures the correlation between audio clips and visual images. However, current methods rely solely on the joint embedding of global features from audio clips and face image pairs to learn semantic correlations. This approach overlooks the importance of high-confidence correlations and discrepancies of local subtle features, which are crucial for cross-modal matching. To address this issue, we propose a novel Attribute-guided Cross-modal Interaction and Enhancement Network (ACIENet), which employs multiple attributes to explore the associations of different key local subtle features. The ACIENet contains two novel modules: the Attribute-guided Interaction (AGI) module and the Attribute-guided Enhancement (AGE) module. The AGI module employs global feature alignment similarity to guide cross-modal local feature interactions, which enhances cross-modal association features for the same identity and expands cross-modal distinctive features for different identities. Additionally, the interactive features and original features are fused to ensure intra-class discriminability and inter-class correspondence. The AGE module captures subtle attribute-related features by using an attribute-driven network, thereby enhancing discrimination at the attribute level. Specifically, it strengthens the combined attribute-related features of gender and nationality. To prevent interference between multiple attribute features, we design a multi-attribute learning network as a parallel framework. Experiments conducted on a public benchmark dataset demonstrate the efficacy of the ACIENet method in different scenarios. Code and models are available at https://github.com/w1018979952/ACIENet. Jiaxiang Wang 0001, Aihua Zheng, Yan Yan 0002, Ran He 0001, Jin Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Camera Topology Graph Guided Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) aims to retrieve vehicles across non-overlapping cameras. Most studies consider representation learning from single appearance information of the vehicle images. Some works adopt the spatio-temporal information to remove unreasonable vehicles to refine the results in the testing phase. However, they ignore the potential topological relations among cameras under the Closed Circuit Television (CCTV) camera systems in the training phase, which usually leads to suboptimal results due to the high intra-identity variations. To handle this problem, we propose a novel vehicle re-identification framework, which explicitly models the camera topological relations of all input images to aggregate neighbor images and thus acquires camera-independent representations. Specifically, we first construct a Camera Topology Graph (CTG) to elucidate the topological relations among cameras. It takes different cameras as nodes and constructs edges from four levels of the camera system, position, orientation, and individual. Then, we introduce a Camera Topology-based Graph Convolutional Network (CT-GCN), which suppresses irrelevant neighbor images and learns different camera representation functions. Finally, we propose a topological cross-entropy loss to obtain the more discriminative vehicle representations. The whole network is trained in an end-to-end manner. Extensive experiments on three benchmark datasets demonstrate the effectiveness of the proposed method against state-of-the-art vehicle Re-ID methods. Aihua Zheng, Yonglong Luo |
IEEE Trans. Multim. | 2 |
| 2023 | Modify: Model-Driven Face Stylization Without Style ImagesabstractExisting face stylization methods always acquire the presence of the target (style) domain during the translation process, which violates privacy regulations and limits their applicability in real-world systems. To address this issue, we propose a new method called MODel-drIven Face stYlization (MODIFY), which relies on the generative model to bypass the dependence of the target images. Briefly, MODIFY first trains a generative model in the target domain and then translates a source input to the target domain via the provided style model. To preserve the multimodal style information, MODIFY further introduces an additional remapping network, mapping a known continuous distribution into the encoder’s embedding space. During translation in the source domain, MODIFY fine-tunes the encoder module within the target style-persevering model to capture the content of the source input as precisely as possible. Our method is extremely simple and satisfies versatile training modes for face stylization. Experimental results on several different datasets validate the effectiveness of MODIFY for unsupervised face stylization. Code will be released at https://github.com/YuheD/MODIFY. Yuhe Ding, Jian Liang 0001, Jie Cao 0002, Aihua Zheng, Ran He 0001 |
ICASSP | 4 |
| 2023 | Where to Focus: Central Attention-Based Face Forgery Detection
Jinghui Sun, Yuhe Ding, Jie Cao 0002, Junxian Duan, Aihua Zheng |
PRCV (5) | 5 |
| 2023 | Diverse features discovery transformer for pedestrian attribute recognition
Aihua Zheng, Jiaxiang Wang 0001, Huaibo Huang, Ran He 0001, Amir Hussain 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | ProxyMix: Proxy-based Mixup training with label refinery for source-free domain adaptation
Yuhe Ding, Lijun Sheng, Jian Liang 0001, Aihua Zheng, Ran He 0001 |
Neural Networks | 4 |
| 2023 | Multi-Query Vehicle Re-Identification: Viewpoint-Conditioned Network, Unified Dataset and New MetricabstractExisting vehicle re-identification methods mainly rely on the single query, which has limited information for vehicle representation and thus significantly hinders the performance of vehicle Re-ID in complicated surveillance networks. In this paper, we propose a more realistic and easily accessible task, called multi-query vehicle Re-ID, which leverages multiple queries to overcome viewpoint limitation of single one. Based on this task, we make three major contributions. First, we design a novel viewpoint-conditioned network (VCNet), which adaptively combines the complementary information from different vehicle viewpoints, for multi-query vehicle Re-ID. Moreover, to deal with the problem of missing vehicle viewpoints, we propose a cross-view feature recovery module which recovers the features of the missing viewpoints by learnt the correlation between the features of available and missing viewpoints. Second, we create a unified benchmark dataset, taken by 6142 cameras from a real-life transportation surveillance system, with comprehensive viewpoints and large number of crossed scenes of each vehicle for multi-query vehicle Re-ID evaluation. Finally, we design a new evaluation metric, called mean cross-scene precision (mCSP), which measures the ability of cross-scene recognition by suppressing the positive samples with similar viewpoints from the same camera. Comprehensive experiments validate the superiority of the proposed method against other methods, as well as the effectiveness of the designed metric in the evaluation of multi-query vehicle Re-ID. The codes and dataset are available at: https://github.com/zhangchaobin001/VCNet. Aihua Zheng, Chaobin Zhang, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Looking and Hearing Into Details: Dual-Enhanced Siamese Adversarial Network for Audio-Visual MatchingabstractAudio-visual cross-modal matching aims to explore the intrinsic correspondence between face images and audio clips. Existing methods usually focus on the salient features of identities between visual images and voice clips, while neglecting their subtle differences, which are crucial to distinguishing cross-modal samples. To deal with this problem, we propose a novel Dual-enhanced Siamese Adversarial Network (DSANet), which pursues the adversarial dual enhancement to highlight both salient and subtle features for robust audio-visual cross-modal matching. First, we designed a dual enhancement mechanism to enhance potential subtle features by randomly selecting a region feature for salient feature suppression, while enhancing salient features in the corresponding region to ensure the global discriminative ability. Second, to establish the correlation of subtle features in the process of eliminating cross-modal heterogeneity, we design a siamese adversarial structure to perform modal heterogeneity elimination for both enhanced salient and subtle features in a parallel manner. Moreover, we propose an adaptive masked cross-entropy loss to force the network to focus on the feature differences among hard classes. Experiments on public benchmark datasets validate the effectiveness of the proposed algorithm. Jiaxiang Wang 0001, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Interact, Embed, and EnlargE: Boosting Modality-Specific Representations for Multi-Modal Person Re-identificationabstractMulti-modal person Re-ID introduces more complementary information to assist the traditional Re-ID task. Existing multi-modal methods ignore the importance of modality-specific information in the feature fusion stage. To this end, we propose a novel method to boost modality-specific representations for multi-modal person Re-ID: Interact, Embed, and EnlargE (IEEE). First, we propose a cross-modal interacting module to exchange useful information between different modalities in the feature extraction phase. Second, we propose a relation-based embedding module to enhance the richness of feature descriptors by embedding the global feature into the fine-grained local information. Finally, we propose multi-modal margin loss to force the network to learn modality-specific information for each modality by enlarging the intra-class discrepancy. Superior performance on multi-modal Re-ID dataset RGBNT201 and three constructed Re-ID datasets validate the effectiveness of the proposed method compared with the state-of-the-art approaches. Zi Wang 0013, Chenglong Li 0002, Aihua Zheng, Ran He 0001, Jin Tang 0001 |
AAAI | 3 |
| 2022 | Progressive Attribute Embedding for Accurate Cross-modality Person Re-IDabstractAttributes are important information to bridge the appearance gap across modalities, but have not been well explored in cross-modality person ReID. This paper proposes a progressive attribute embedding module (PAE) to effectively fuse the fine-grained semantic attribute information and the global structural visual information. Through a novel cascade way, we use attribute information to learn the relationship between the person images in different modalities, which significantly relieves the modality heterogeneity. Meanwhile, by embedding attribute information to guide more discriminative image feature generation, it simultaneously reduces the inter-class similarity and the intra-class discrepancy. In addition, we propose an attribute-based auxiliary learning strategy (AAL) to supervise the network to learn modality-invariant and identity-specific local features by joint attribute and identity classification losses. The PAE and AAL are jointly optimized in an end-to-end framework, namely, progressive attribute embedding network (PAENet). One can plug PAE and AAL into current mainstream models, as we implement them in five cross-modality person ReID frameworks to further boost the performance. Extensive experiments on public datasets demonstrate the effectiveness of the proposed method against the state-of-the-art cross-modality person ReID methods. Aihua Zheng, Chenglong Li 0002, Bin Luo 0001, Ruoran Jia |
ACM Multimedia | 1 |
| 2022 | Prior-Guided Multi-scale Fusion Transformer for Face Attribute Recognition
Shaoheng Song, Huaibo Huang, Jiaxiang Wang 0001, Aihua Zheng, Ran He 0001 |
PRCV (1) | 4 |
| 2022 | Attributes Based Visible-Infrared Person Re-identification
Aihua Zheng, Mengya Feng, Bo Jiang 0002, Bin Luo 0001 |
PRCV (1) | 1 |
| 2022 | Pedestrian attribute recognition: A survey
Xiao Wang 0014, Shaofei Zheng, Aihua Zheng, Zhe Chen 0013, Jin Tang 0001, Bin Luo 0001 |
Pattern Recognit. | 4 |
| 2022 | Category-Wise Fusion and Enhancement Learning for Multimodal Remote Sensing Image Semantic SegmentationabstractThis paper presents a simple yet effective method called Category-wise Fusion and Enhancement learning (CaFE), which leverages the category priors to achieve effective feature fusion and imbalance learning, for multi-modal remote sensing image semantic segmentation. In particular, we disentangle the feature fusion process via the categories to achieve the category-wise fusion based on the fact that the feature fusion in the same category regions tends to have similar characteristics. The disentangled fusion would also increase the fusion capacity with a small number of parameters while reducing the dependence on large-scale training data. For the sample imbalance problem, we design a simple yet effective category-wise enhancement learning scheme. In particular, we assign the weight for each category region based on the proportion of samples in this region over the whole image. By this way, the learning algorithm would focus more on the regions with smaller proportion. Note that both category-wise feature fusion and imbalance learning are only performed in the training stage, and the segmentation efficiency is thus not affected. Experimental results on two benchmark datasets demonstrate the effectiveness of our CaFE against other state-of-the-art methods. Aihua Zheng, Jinbo He, Chenglong Li 0002, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Entropy Guided Adversarial Domain Adaptation for Aerial Image Semantic SegmentationabstractRecent advances on aerial image semantic segmentation mainly employ the domain adaption to transfer knowledge from the source domain to the target domain. Despite the remarkable achievement, most methods focus on the global marginal distribution alignment to reduce the domain shift between source and target domains, leading to a wrong mapping of the well-aligned features. In this article, we propose an effective unsupervised domain adaptation approach, which relies on a novel entropy guided adversarial learning algorithm, for aerial image semantic segmentation. In specific, we perform local feature alignment between domains by learning a self-adaptive weight from the target prediction probability map to measure the interdomain discrepancy. To exploit the meaningful structure information among semantic regions, we propose to utilize the graph convolutions for long-range semantic reasoning. Comprehensive experimental results on the benchmark dataset of aerial image semantic segmentation and natural scenes demonstrate the superior performance of the proposed method compared to the state-of-the-art methods. Aihua Zheng, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Attribute and State Guided Structural Embedding Network for Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) is a crucial task in smart city and intelligent transportation, aiming to match vehicle images across non-overlapping surveillance camera scenarios. However, the images of different vehicles may have small visual discrepancies when they have the same/similar attributes, e.g., the same/similar color, type, and manufacturer. Meanwhile, the images from a vehicle may have large visual discrepancies with different states, e.g., different camera views, vehicle viewpoints, and capture time. In this paper, we propose an attribute and state guided structural embedding network (ASSEN) to achieve discriminative feature learning by attribute-based enhancement and state-based weakening for vehicle Re-ID. First, we propose an attribute-based enhancement and expanding module to enhance the discrimination of vehicle features through identity-related attribute information, and we design an attribute-based expanding loss to increase the feature gap between different vehicles. Second, we design a state-based weakening and shrinking module, which not only weakens the state information that interferes with identification but also reduces the intra-class feature gap by a state-based shrinking loss. Third, we propose a global structural embedding module that exploits the attribute information and state information to explore hierarchical relationships between vehicle features, then we use these relationships for feature embedding to learn more robust vehicle features. Extensive experiments on benchmark datasets VeRi-776, VehicleID, and VERI-Wild demonstrate the superior performance and generalization of the proposed method against state-of-the-art vehicle Re-ID methods. The code is available at https://github.com/ttaalle/fast_assen. Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | MsKAT: Multi-Scale Knowledge-Aware Transformer for Vehicle Re-IdentificationabstractExisting vehicle re-identification (Re-ID) methods usually suffer from intra-instance discrepancy and inter-instance similarity. The key to solving this problem lies in filtering out identity-irrelevant interference and collecting identity-relevant vehicle details. In this paper, we aim to design a robust vehicle Re-ID framework that trains a model guided by knowledge vectors yet is able to disentangle the identity-relevant features and identity-irrelevant features. Toward this end, we propose a novel Multi-scale Knowledge-Aware Transformer (MsKAT) to build a knowledge-guided multi-scale feature alignment framework. First, we construct a Knowledge-Aware Transformer (KAT) to interact with semantic knowledge and visual feature. KAT mainly includes State elimination Transformer (SeT) to eliminate state (camera, viewpoint) interference and Attribute aggregation Transformer (AaT) to gather attribute (color, type) information. Second, to learn the knowledge-guided sample differences, we propose to encourage the separation of identity-relevant features and identity-irrelevant features by a Knowledge-Guided Alignment loss ($\mathcal {L}_{KGA}$). Specifically,$\mathcal {L}_{KGA}$suppresses the difference between knowledge-guided positive pairs and the similarity between knowledge-guided negative pairs. Third, with the multi-scale settings of KAT and$\mathcal {L}_{KGA}$, our model can capture knowledge-guided visual consistency features at different scales. Extensive evidence demonstrates our approach achieves new state-of-the-art on three widely-used vehicle re-identification benchmarks. Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Viewpoint-Aware Progressive Clustering for Unsupervised Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) is an active task due to its importance in large-scale intelligent monitoring in smart cities. Despite the rapid progress in recent years, most existing methods handle vehicle Re-ID task in a supervised manner, which is both time and labor-consuming and limits their application to real-life scenarios. Recently, unsupervised person Re-ID methods achieve impressive performance by exploring domain adaption or clustering-based techniques. However, one cannot directly generalize these methods to vehicle Re-ID since vehicle images present huge appearance variations in different viewpoints. To handle this problem, we propose a novel viewpoint-aware clustering algorithm for unsupervised vehicle Re-ID. In particular, we first divide the entire feature space into different subspaces according to the predicted viewpoints and then perform a progressive clustering to mine the accurate relationship among samples. Comprehensive experiments against the state-of-the-art methods on two multi-viewpoint benchmark datasets VeRi-776 and VeRi-Wild validate the promising performance of the proposed method in both with and without domain adaption scenarios while handling unsupervised vehicle Re-ID. Aihua Zheng, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | PH-GCN: Person Retrieval With Part-Based Hierarchical Graph Convolutional NetworkabstractCompact feature representation of person image is important for person re-identification (Re-ID) task. Recently, part-based representation models have been widely studied for extracting the more compact and robust feature representation for person image to improve person Re-ID results. However, existing part-based representation models mostly extract the features of different parts independently which ignore the spatial relationship information among different parts. To address this issue, in this paper we propose a novel deep learning framework, named Part-based Hierarchical Graph Convolutional Network (PH-GCN) for person Re-ID problem. Given a person image, PH-GCN first constructs a hierarchical graph to represent the spatial relationships among different parts. Then, both local and global feature learning is achieved by the feature information passing in PH-GCN, which takes the information of other parts into account for part feature representation. Finally, a perceptron layer is adopted for the final person part label prediction and re-identification. The proposed framework provides a general solution that integrateslocal,globalandstructuralfeature learning simultaneously in a unified end-to-end network representation and learning. Extensive experiments on several widely used benchmark datasets demonstrate the effectiveness and benefits of the proposed PH-GCN approach for person Re-ID task. Bo Jiang 0002, Xixi Wang 0005, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Adversarial-Metric Learning for Audio-Visual Cross-Modal MatchingabstractAudio-visual matching aims to learn the intrinsic correspondence between image and audio clip. Existing works mainly concentrate on learning discriminative features, while ignore the cross-modal heterogeneous issue between audio and visual modalities. To deal with this issue, we propose a novel Adversarial-Metric Learning (AML) model for audio-visual matching. AML aims to generate a modality-independent representation for each person in each modality via adversarial learning, while simultaneously learns a robust similarity measure for cross-modality matching via metric learning. By integrating the discriminative modality-independent representation and robust cross-modality metric learning into an end-to-end trainable deep network, AML can overcome the heterogeneous issue with promising performance for audio-visual matching. Experiments on the various audio-visual learning tasks, including audio-visual matching, audio-visual verification and audio-visual retrieval on benchmark dataset demonstrate the effectiveness of the proposed AML model. The implementation codes are available onhttps://github.com/MLanHu/AML. Aihua Zheng, Menglan Hu, Bo Jiang 0002, Yan Yan 0002, Bin Luo 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Robust Multi-Modality Person Re-identificationabstractTo avoid the illumination limitation in visible person re-identification (Re-ID) and the heterogeneous issue in cross-modality Re-ID, we propose to utilize complementary advantages of multiple modalities including visible (RGB), near infrared (NI) and thermal infrared (TI) ones for robust person Re-ID. A novel progressive fusion network is designed to learn effective multi-modal features from single to multiple modalities and from local to global views. Our method works well in diversely challenging scenarios even in the presence of missing modalities. Moreover, we contribute a comprehensive benchmark dataset, RGBNT201, including 201 identities captured from various challenging conditions, to facilitate the research of RGB-NI-TI multi-modality person Re-ID. Comprehensive experiments on RGBNT201 dataset comparing to the state-of-the-art methods demonstrate the contribution of multi-modality person Re-ID and the effectiveness of the proposed approach, which launch a new benchmark and a new baseline for multi-modality person Re-ID. Aihua Zheng, Zi Wang 0013, Zi-Han Chen, Chenglong Li 0002, Jin Tang 0001 |
AAAI | 1 |
| 2020 | Multi-Spectral Vehicle Re-Identification: A ChallengeabstractVehicle re-identification (Re-ID) is a crucial task in smart city and intelligent transportation, aiming to match vehicle images across non-overlapping surveillance camera views. Currently, most works focus on RGB-based vehicle Re-ID, which limits its capability of real-life applications in adverse environments such as dark environments and bad weathers. IR (Infrared) spectrum imaging offers complementary information to relieve the illumination issue in computer vision tasks. Furthermore, vehicle Re-ID suffers a big challenge of the diverse appearance with different views, such as trucks. In this work, we address the RGB and IR vehicle Re-ID problem and contribute a multi-spectral vehicle Re-ID benchmark named RGBN300, including RGB and NIR (Near Infrared) vehicle images of 300 identities from 8 camera views, giving in total 50125 RGB images and 50125 NIR images respectively. In addition, we have acquired additional TIR (Thermal Infrared) data for 100 vehicles from RGBN300 to form another dataset for three-spectral vehicle Re-ID. Furthermore, we propose a Heterogeneity-collaboration Aware Multi-stream convolutional Network (HAMNet) towards automatically fusing different spectrum features in an end-to-end learning framework. Comprehensive experiments on prevalent networks show that our HAMNet can effectively integrate multi-spectral data for robust vehicle Re-ID in day and night. Our work provides a benchmark dataset for RGB-NIR and RGB-NIR-TIR multi-spectral vehicle Re-ID and a baseline network for both research and industrial communities. The dataset and baseline codes are available at: https://github.com/ttaalle/multi-modal-vehicle-Re-ID. Chenglong Li 0002, Xianpeng Zhu, Aihua Zheng, Bin Luo 0001 |
AAAI | 4 |
| 2020 | Unsupervised Contrastive Photo-to-Caricature Translation based on Auto-distortionabstractPhoto-to-caricature translation aims to synthesize the caricature as a rendered image exaggerating the features through sketching, pencil strokes, or other artistic drawings. Style rendering and geometry deformation are the most important aspects in photo-to-caricature translation task. To take both into consideration, we propose an unsupervised contrastive photo-to-caricature translation architecture. Considering the intuitive artifacts in the existing methods, we propose a contrastive style loss for style rendering to enforce the similarity between the style of rendered photo and the caricature, and simultaneously enhance its discrepancy to the photos. To obtain an exaggerating deformation in an unpaired/unsupervised fashion, we propose a Distortion Prediction Module (DPM) to predict a set of displacements vectors for each input image while fixing some controlling points, followed by the thin plate spline interpolation for warping. The model is trained on unpaired photo and caricature while can offer bidirectional synthesizing via inputting either a photo or a caricature. Extensive experiments demonstrate that the proposed model is effective to generate hand-drawn like caricatures compared with existing competitors. Yuhe Ding, Xin Ma 0031, Mandi Luo, Aihua Zheng, Ran He 0001 |
ICPR | 4 |
| 2020 | Attentional Wavelet Network for Traditional Chinese Painting TransferabstractTraditional Chinese paintings pay more attention to `Gongbi' and `Xieyi' in artworks, which raises a challenging task to generate Chinese paintings from photos. `Xieyi' creates high-level conception for paintings, while `Gongbi' refers to portraying local details in paintings. This paper proposes an attentional wavelet network for photo to Chinese painting transferring. We first introduce wavelets to obtain high-level conception and local details in Chinese paintings via 2-D haar wavelet transform. Moreover, we design high-level transform stream and local enhancement stream to dispose high frequencies and low frequency respectively. Furthermore, we exploit self-attention mechanism to compatibly pick up high-level information which is used to remedy the missing details when reconstructing the Chinese painting. To advance our experiment, we set up a new dataset named P2ADataset, with diverse photos and Chinese paintings on famous mountains around China. Experimental results comparing with the state-of-the-art style transferring algorithms verify the effectiveness of the proposed method. We will release the codes and data to the public. Rui Wang 0124, Huaibo Huang, Aihua Zheng, Ran He 0001 |
ICPR | 3 |
| 2020 | Talking Face Generation via Learning Semantic and Temporal Synchronous LandmarksabstractGiven a speech clip and facial image, the goal of talking face generation is to synthesize a talking face video with accurate mouth synchronization and natural face motion. Recent progress has proven the effectiveness of the landmarks as the intermediate information during talking face generation. However, the large gap between audio and visual modalities makes the prediction of landmarks challenging and limits generation ability. This paper proposes a semantic and temporal synchronous landmark learning method for talking face generation. First, we propose to introduce a word detector to enforce richer semantic information. Then, we propose to preserve the temporal synchronization and consistency between landmarks and audio via the proposed temporal residual loss. Lastly, we employ a U-Net generation network with adaptive reconstruction loss to generate facial images for the predicted landmarks. Experimental results on two benchmark datasets LRW and GRID demonstrate the effectiveness of our model compared to the state-of-the-art methods of talking face generation. Aihua Zheng, Feixia Zhu, Mandi Luo, Ran He 0001 |
ICPR | 1 |
| 2020 | Let's Play Music: Audio-Driven Performance Video GenerationabstractWe propose a new task named Audio-driven Performance Video Generation (APVG), which aims to synthesize the video of a person playing a certain instrument guided by a given music audio clip. It is a challenging task to generate the high-dimensional temporal consistent videos from low-dimensional audio modality. In this paper, we propose a multi-staged framework to generate realistic and synchronized performance video from given music. Firstly, we provide both global appearance and local spatial information by generating the coarse videos and keypoints of body and hands from a given music respectively. Then, we propose to transform the generated keypoints to heatmap via a differentiable space transformer, since the heatmap provides more spatial information but is harder to generate directly from audio. Finally, we propose a Structured Temporal UNet (STU) to extract both intra-frame structured information and interframe temporal consistency. They are obtained via graph-based structure module, and CNN-GRU based high-level temporal module respectively for final video generation. Comprehensive experiments validate the effectiveness of our proposed framework. Yi Li 0018, Feixia Zhu, Aihua Zheng, Ran He 0001 |
ICPR | 4 |
| 2020 | Arbitrary Talking Face Generation via Attentional Audio-Visual Coherence LearningabstractTalking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing methods mainly focus on either disentangling the information in a single image or learning temporal information between frames. However, cross-modality coherence between audio and video information has not been well addressed during synthesis. In this paper, we propose a novel arbitrary talking face generation framework by discovering the audio-visual coherence via the proposed Asymmetric Mutual Information Estimator (AMIE). In addition, we propose a Dynamic Attention (DA) block by selectively focusing the lip area of the input image during the training stage, to further enhance lip synchronization. Experimental results on benchmark LRW dataset and GRID dataset transcend the state-of-the-art methods on prevalent metrics with robust high-resolution synthesizing on gender and pose variations. Huaibo Huang, Yi Li 0018, Aihua Zheng, Ran He 0001 |
IJCAI | 4 |
| 2020 | Multi-modal foreground detection via inter- and intra-modality-consistent low-rank separation
Aihua Zheng, Naipeng Ye, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001 |
Neurocomputing | 1 |
| 2020 | Multi-scale attention vehicle re-identification
Aihua Zheng, Xianmin Lin, Jiacheng Dong, Wenzhong Wang, Jin Tang 0001, Bin Luo 0001 |
Neural Comput. Appl. | 1 |
| 2020 | Joint graph regularized dictionary learning and sparse ranking for multi-modal multi-shot person re-identification
Aihua Zheng, Bo Jiang 0002, Wei-Shi Zheng 0001, Bin Luo 0001 |
Pattern Recognit. | 1 |
| 2020 | A Subspace Learning Approach to Multishot Person ReidentificationabstractThis paper addresses the challenging problem of multishot person reidentification (Re-ID) in real world uncontrolled surveillance systems. A key issue is how to effectively represent and process the multiple data with various appearance information due to the variations of pose, occlusions, and viewpoints. To this end, this paper develops a novel subspace learning approach, which pursues regularized low-rank and sparse representation for multishot person Re-ID. For the images of a person crossing a certain camera, we assume that the appearances of those subset images with similar viewpoints against a camera draw from the same low-rank subspace, and all the images of a person under a camera lie on a union of low-rank subspaces. Based on this assumption, we propose to learn a nonnegative low-rank and sparse graph to represent the person images. Moreover, the recurring pattern prior is integrated into our model to refine the affinities among images. Extensive experiments on four public benchmark datasets yield impressive performance by improving 22.9% on imagery library for intelligent detection systems video re identification (iLIDS-VID), 42.4% on person RE-ID (PRID) dataset 2011, 39.7% and 30.6% on speech, audio, image, and video technology-SoftBio camera 3/8 and camera 5/8, respectively, and 1.6% on motion analysis and re identification set compared to the state-of-the-art methods. Aihua Zheng, Xuehan Zhang, Bo Jiang 0002, Bin Luo 0001, Chenglong Li 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2019 | Person Re-identification with Patch-Based Local Sparse Matching and Metric Learning
Bo Jiang 0002, Yibing Lv, Aihua Zheng, Bin Luo 0001 |
ICIG (2) | 3 |
| 2019 | MMA: Motion Memory Attention Network for Video Object Detection
Huai Hu, Wenzhong Wang, Aihua Zheng, Bin Luo 0001 |
ICIG (2) | 3 |
| 2019 | Learning to Detect License Plates Using Synthesized Data
Yanhui Pang, Wenzhong Wang, Aihua Zheng, Jin Tang 0001 |
ICIG (2) | 3 |
| 2019 | Saliency detection via multi-view graph based saliency optimizationabstractSaliency detection is an important problem in computer vision and pattern recognition area. Many works have been proposed for addressing the saliency detection task. As a popular method, graph based saliency optimization has been widely studied. However, previous works have universally focussed on single graph optimization which fails to consider multi-view feature representation of image content . In this paper, we first provide a general framework for traditional graph based saliency optimization models. Then, we extend the general framework to the multi-view case and propose our general multi-view graph based saliency optimization model. Finally, we present a particular implementation of our general model and derive an effective updating algorithm to solve it. Experimental results using several benchmark datasets demonstrate the effectiveness of our proposed saliency model. Yun Xiao 0003, Bo Jiang 0002, Aihua Zheng, Aiwu Zhou, Amir Hussain 0001, Jin Tang 0001 |
Neurocomputing | 3 |
| 2019 | Background subtraction with multi-scale structured low-rank and sparse factorization
Aihua Zheng, Tian Zou, Yumiao Zhao, Bo Jiang 0002, Jin Tang 0001, Chenglong Li 0002 |
Neurocomputing | 1 |
| 2018 | Exploring Scene Geometry for Scale Adaptive Object Tracking in Surveillance VideosabstractObject tracking is a key technology in video surveillance. Reliable tracker must be adaptive to the constantly changing object sizes. Most of the state-of-the-art methods estimate the object scales using their appearances. Those methods are vulnerable to occlusion, object deformation, illumination change and background clutter. In this paper, we propose to use the geometric context of the surveillance site as a strong clue for scale adaptation. With three reasonable assumptions on the video cameras and the surveillance sites, we deduce a simple geometric model for object scales. The parameters of this model are learned without any human intervention. Then we integrate this model into baseline trackers for robust scale adaptive object tracking. Experimental results on challenging surveillance videos indicate that our approach favorably improves the performance of single-scale baselines, and performs better or comparative to the state-of-the-art multi-scale trackers while significantly improve the speed. Ran Zhong, Wenzhong Wang, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001 |
ICIP | 4 |
| 2018 | Multi-scale Cooperative Ranking for Saliency Detection
Bo Jiang 0002, Xingyue Jiang, Aihua Zheng, Yun Xiao 0003, Jin Tang 0001 |
PRCV (1) | 3 |
| 2018 | Non-negative Dual Graph Regularized Sparse Ranking for Multi-shot Person Re-identification
Aihua Zheng, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
PRCV (1) | 1 |
| 2018 | Spatial-temporal representatives selection and weighted patch descriptor for person re-identification
Aihua Zheng, Foqin Wang, Amir Hussain 0001, Jin Tang 0001, Bo Jiang 0002 |
Neurocomputing | 1 |
| 2017 | Manifold ranking weighted local maximal occurrence descriptor for person re-identificationabstractPerson re-identification is an important task of matching pedestrians across non-overlapping camera views. In this paper, we exploit a weighted feature descriptor for person re-identification. We firstly compute the weights on the superpixel level via graph-based manifold ranking algorithm, then integrate the computed weights into a patch-based feature descriptor, named local maximal occurrence. Finally, the weighted descriptors are fed into a top-push distance learning to mitigate the cross-view gaps. We evaluate the proposed method on three benchmark datasets iLIDS-VID, PRID 450S and VIPeR. The promising experimental results demonstrate the effectiveness of the proposed method comparing with the state-of-the-arts. Foqin Wang, Xuehan Zhang, Jinxin Ma, Jin Tang 0001, Aihua Zheng |
SERA | 5 |
| 2017 | Local-to-global background modeling for moving object detection from non-static cameras
Aihua Zheng, Lei Zhang 0074, Wei Zhang 0012, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
Multim. Tools Appl. | 1 |
| 2017 | Image representation and matching with geometric-edge random structure graph
Bo Jiang 0002, Jin Tang 0001, Aihua Zheng, Bin Luo 0001 |
Pattern Recognit. Lett. | 3 |
| 2015 | Person Re-identification with Density-Distance Unsupervised Salience Learning
Baoliang Zhou, Aihua Zheng, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001 |
ICIG (3) | 2 |
| 2012 | Matching State-Based Sequences with Rich Temporal AspectsabstractA General Similarity Measurement (GSM), which takes into account of both non-temporal and rich temporal aspects including temporal order, temporal duration and temporal gap, is proposed for state-sequence matching. It is believed to be versatile enough to subsume representative existing measurements as its special cases. Aihua Zheng, Jixin Ma 0001, Jin Tang 0001, Bin Luo 0001 |
AAAI | 1 |
| 2012 | Graph matching based on spectral embedding with missing value
Jin Tang 0001, Bo Jiang 0002, Aihua Zheng, Bin Luo 0001 |
Pattern Recognit. | 3 |