VLDB 2026 Research / reviewers in the wild / expert
Zan Gao 0001
dblp:02/2658-1
· DBLP profile ↗
56ranked-venue papers
19as first author
53since 2021 · last 2026
0000-0003-2182-5741ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 10 first-author · 29 since 2021Artificial intelligence and machine learning · 15 · 6 first-author · 15 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 5 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Amplifying Discrepancies: Exploiting Macro and Micro Inconsistencies for Image Manipulation LocalizationabstractThe rapid development of image manipulation technologies poses significant challenges to multimedia forensics, especially in accurate localization of manipulated regions. Existing methods often fail to fully explore the intrinsic discrepancies between manipulated and authentic regions, resulting in sub-optimal performance. To address this limitation, we propose the Focus Region Discrepancy Network (FRD-Net), a novel and efficient framework that significantly enhances manipulation localization by amplifying discrepancies at both macro- and micro-levels. Specifically, our proposed Iterative Clustering Module (ICM) groups features into two discriminative clusters and refines representations via backward propagation from cluster centers, improving the distinction between tampered and authentic regions at the macro level. Thereafter, our Differential Progressive Module (DPM) is constructed to capture fine-grained structural inconsistencies within local neighborhoods and integrate them into a Central Difference Convolution, increasing sensitivity to subtle manipulation details at the micro level. Finally, these complementary modules are seamlessly integrated into a compact architecture that achieves a favorable balance between accuracy and efficiency. Extensive experiments on multiple benchmarks demonstrate that FRD-Net consistently surpasses state-of-the-art methods in terms of manipulation localization performance while maintaining a lower computational cost. Shenghao Chen, Yibo Zhao 0001, Tianyi Wang 0006, Chunjie Ma, Weili Guan, Ming Li 0083, Zan Gao 0001 |
AAAI | 7 |
| 2026 | Unsupervised 2D Image-Based 3D Model Retrieval via Decision Boundary Alignment and Graph Semantic PropagationabstractUnsupervised 2D image-based 3D model retrieval (IBMR) aims to retrieve semantically relevant 3D shapes for a given 2D image query when 3D annotations are unavailable. This setting is challenging due to severe modality gaps, category-imbalanced mini-batches, inconsistent cross-domain decision boundaries, and mismatched semantic neighborhood structures. In this paper, we propose a unified framework that integrates Category-Aligned Sampling (CAS), Decision Boundary Alignment (DBA), and Graph Semantic Propagation (GSP) into a single optimization paradigm. CAS constructs category-consistent mini-batches to stabilize crossmodal learning. Built upon CAS, DBA leverages a masked Margin Disparity Discrepancy to regularize cross-domain class decision boundaries via an adversarial min-max objective, encouraging discriminative separation beyond marginal feature matching. To complement boundary-level regularization, GSP builds a crossdomain affinity graph over 2D and 3D samples and propagates supervision-induced relational structure through semantic message passing, explicitly preserving instance-level neighborhood consistency that is critical for retrieval. Extensive experiments on MI3DOR and MI3DOR-2 demonstrate consistent improvements over representative unsupervised IBMR baselines. Nian Hu, Yibo Zhao 0001, Chen Li 0035, Cong Liu 0012, Zan Gao 0001 |
SIGIR | 6 |
| 2026 | Understanding complex queries: Multi-query comprehension network for temporal sentence grounding
Zan Gao 0001, Shengbo Xiong, Yibo Zhao 0001, Chunjie Ma, Tian Gan 0002, Riwei Wang |
Expert Syst. Appl. | 1 |
| 2026 | Domain adaptive object detection via CLIP-space guidance and LoRA fine-tuning
Enze Qi, Kan Chang, Qingzhi Zhang, Xueyu Zhang, Yehua Ling, Yujian Yuan, Zan Gao 0001 |
Expert Syst. Appl. | 8 |
| 2026 | UECNet: A unified framework for exposure correction utilizing region-level prompts
Shucheng Xia, Kan Chang, Xuxin Tai, Yehua Ling, Yujian Yuan, Zan Gao 0001 |
Knowl. Based Syst. | 8 |
| 2026 | RMPSNet: Occluded person re-identification via regional masking and prompt-distribution synergy
Zan Gao 0001, Shuai Xie, Shengxun Wei, Yibo Zhao 0001, Chunjie Ma, Chen Li 0035 |
Pattern Recognit. | 1 |
| 2026 | Prior-knowledge guidance and dual-domain representation refinement for deepfake detection
Zhiyong Cheng 0001, Tianyi Wang 0006, Chunjie Ma, Yibo Zhao 0001, Zan Gao 0001 |
Pattern Recognit. | 6 |
| 2026 | A Novel Multi-View Perception and Shrinkage Aggregation Network for Inharmonious Region LocalizationabstractWith the popularity of image editing techniques, synthetic images may have inharmonious regions due to color/illumination differences between the manipulated area and the background. The inharmonious region localization task aims to find these regions, which is crucial for blind image harmonization. Existing methods rely on single-view images and do not fully explore multi-scale fusion, which limits their performance. To address these issues, in this paper, we propose a novel multi-view perception and shrinkage aggregation network (MSANet) for the inharmonious region localization task that fully utilizes multi-view images and multi-scale fusion information and can mine subtle cues between candidate objects and the background. Specifically, we first design a multi-view ensemble encoder to fully perceive the inharmonious regions by multi-view interactive learning and then aggregate the feature representations of inharmonious regions. Moreover, we propose a multi-scale shrinkage fusion decoder, where multi-scale features with multi-view prior information are utilized to aggregate adjacent features, adaptively select high-quality information, reduce background interference and gradually locate inharmonious regions. Extensive experimental results on four public datasets (HDobe5K, HCOCO, HFlickr, and Hday2Night) demonstrate that the proposed MSANet can outperform all the SOTA methods in terms of average F1 and average IoU score, while maintaining a lower computational cost1. Shenghao Chen, Chunjie Ma, Yibo Zhao 0001, Meng Liu 0006, Yanbing Xue, Zan Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Learning Generalizable Representations for Deepfake Detection With Realistic Sample Generation and Dual AugmentationabstractDeepfake detection aims to identify manipulated content generated by generative models such as GANs and diffusion models. Although many detection methods have been proposed in recent years, their performance often degrades significantly when the test data includes unknown images or novel forgery types. To address this challenge, we propose a Realistic Sample Generation and Dual Augmentation framework, abbreviated as RSG-DA, to enhance generalization in deepfake detection. The key idea is to explore and expand the forgery feature space in order to learn decision boundaries that can capture diverse forgery patterns. Specifically, we introduce a Dynamic Landmark Diffusion Generator (DLDG) that synthesizes hybrid forgery samples with high visual realism and structural diversity. Additionally, we design a Dual Data Augmentation (DDA) strategy composed of DW-Augmentation and Class-Augmentation, where DW-Augmentation strengthens the representation of authentic image features through multi-scale transformations, while Class-Augmentation enriches the forgery distribution by expanding it with varied manipulations. Finally, we present a Lightweight Generic Forgery Distillation (LGFD) module that integrates the above components into a unified encoder, enabling the learning of robust and transferable forgery representations. Extensive experiments show that our method consistently outperforms state-of-the-art approaches in both intra-dataset and cross-dataset evaluations. Zan Gao 0001, Xinhai Zhu, Yibo Zhao 0001, Chunjie Ma, Chen Li 0035 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | PVLM: Parsing-Aware Vision-Language Model With Dynamic Contrastive Learning for Zero-Shot Deepfake AttributionabstractThe challenge of tracing the source attribution of forged faces has gained significant attention due to the rapid advancement of generative models. However, existing deepfake attribution (DFA) works primarily focus on the interaction among various domains in vision modality, and other modalities such as texts and face parsing are not fully explored. Besides, they tend to fail to assess the generalization performance of deepfake attributors to unseen advanced generators like diffusion in a fine-grained manner. In this paper, we propose a novelparsing-awarevisionlanguagemodel with dynamic contrastive learning (PVLM) method forzero-shotdeepfakeattribution (ZS-DFA), which facilitates effective and fine-grained traceability to unseen advanced generators. Specifically, we conduct a novel and fine-grained ZS-DFA benchmark to evaluate the attribution performance of deepfake attributors to unseen advanced generators like diffusion. Besides, we propose an innovative PVLM attributor based on the vision-language model to capture general and diverse attribution features. We are motivated by the observation that the preservation of source face attributes in facial images generated by GAN and diffusion models varies significantly. We propose to employ the inherent facial attributes preservation differences to capture face parsing-aware forgery representations. Therefore, we devise a novel parsing encoder to focus on global face attribute embeddings, enabling parsing-guided DFA representation learning via dynamic vision-parsing matching. Additionally, we present a novel deepfake attribution contrastive center loss to pull relevant generators closer and push irrelevant ones away, which can be introduced into DFA models to enhance traceability. Experimental results show that our model exceeds the state-of-the-art on the ZS-DFA benchmark via various protocol evaluations. Codes will be available at GitHub. Chunjie Ma, Weili Guan, Tian Gan 0002, Zan Gao 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | LSGNet: A Local-Pattern Separation and Global-Aware Network for Temporal Action DetectionabstractTemporal Action Detection aims to localize and classify action instances within untrimmed videos, yet it remains challenging due to background clutter, high intra-class similarity, and varied temporal scales in real-world scenarios. To address these issues, we propose the Local-Pattern Separation and Global-Aware Network (LSGNet) tailored for temporal action localization. Specifically, the core of LSGNet is the Local Pattern Separation Module (LPSM), which explicitly models both consistency and variation patterns of action segments within local temporal windows. Additionally, to capture comprehensive contextual information, we introduce the Global Context-Aware Representation Module (GCRM), which decouples temporal features across multiple granularities and enables robust modeling of long-range dependencies. Finally, we design the Multi-scale Feature Refinement Module (MFRM) to mitigate the degradation of fine-grained information by performing iterative reconstruction across temporal scales, thereby enriching semantic representations and preserving temporal details. Extensive experiments on THUMOS14, ActivityNet1.3, HACS, and EPIC-Kitchens-100 demonstrate the effectiveness of the proposed LSGNet method. Additional ablation studies on the QVHighlights dataset further confirm the generalization capability of LPSM module in video moment retrieval and highlight detection, achieving consistent improvements in retrieval accuracy and localization precision. Zan Gao 0001, Weilin Yang, Yibo Zhao 0001, Chunjie Ma, Chen Li 0035, Riwei Wang |
IEEE Trans. Image Process. | 1 |
| 2026 | EFIN: A Novel Enhanced Feature Interaction Network for Temporal Sentence Grounding in VideosabstractTemporal sentence grounding in videos (TSGV) is a challenging task that aims to match text queries with semantically relevant segments in untrimmed videos. However, existing methods face limitations in modeling modality features, which constrains the expressive power of candidate moment features. To address this challenge, we propose a novel Enhanced Feature Interaction Network (EFIN) that effectively captures semantic information within each modality and aligns relationships between modalities. Additionally, EFIN enhances the fusion of information between candidate moments and modality features. Specifically, our model begins by extracting modality features to generate candidate moments as priors. Building upon these modality features, we introduce an enhanced feature encoder to extract semantic information within each modality, thereby improving intra-modality feature representation. Simultaneously, the encoder captures alignment relationships between modalities to optimize cross-modality feature representation, enhancing the overall modeling capacity of modality features. Moreover, we design an information fusion module to enrich the comprehension of modality information for candidate moments. Extensive experiments on four benchmark datasets demonstrate the superiority of our proposed EFIN model. Notably, EFIN achieves a maximum performance improvement of approximately 1.67% and 1.91% across different evaluation metrics on TACoS dataset. Chongxu Hu, Xianbin Wen, Yibo Zhao 0001, Chunjie Ma, Weili Guan, Riwei Wang, Zan Gao 0001 |
IEEE Trans. Multim. | 7 |
| 2026 | A Collaborative Hierarchical Aggregation Network for Weakly Supervised Temporal Action LocalizationabstractTemporal action localization is a fundamental task in video understanding that focuses on classifying and temporally localizing action instances in untrimmed videos. Compared to temporal action localization, the Weakly supervised Temporal Action Localization (WTAL) task presents greater challenges, as its training data lacks detailed information about action boundaries. Existing WTAL methods ignore the complementary relationship between modalities and the dependency between snippets, resulting in inaccurate localization results. To solve these issues, we propose a Collaborative Hierarchical Aggregation Network (CHA-Net). Specifically, we first use a modality complementary module to learn the synergies between modalities. Then, a collaborative enhance module is proposed to remove the information irrelevant to actions in RGB modality. Finally, a hierarchical aggregation module is proposed to capture the complete temporal information of action instances to better mine the temporal dependencies between snippets. Extensive experiments on THUMOS14, ActivityNet1.2, and ActivityNet1.3 datasets demonstrate the effectiveness of our method. Compared with F3-Net (TMM2024, Avg{0.1:0.5}) and SPCC-Net (TMM2024, Avg{0.1:0.7}) on the THUMOS14 dataset, the proposed method can achieve improvements of 3.2% and 2.4%, respectively. Zan Gao 0001, Xiaoyi Xu, Yibo Zhao 0001, Chunjie Ma, Yanbing Xue, Riwei Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | FluidGS: Physics Informed Gaussian Splatting for Dynamic Fluid Reconstruction from Sparse Views
Youchen Xie, Chen Li 0035, Sheng Qiu, Zhi-Jun Wang, Chenhui Li 0001, Yibo Zhao 0001, Zan Gao 0001, Changbo Wang |
ACM Multimedia | 7 |
| 2025 | Lightweight Relational Proposal Network with Dual-Branch Distillation for Video Moment Retrieval
Yujia Zhu, Hao Yang 0067, Yibo Zhao 0001, Chunjie Ma, Weili Guan, Zan Gao 0001 |
ACM Multimedia | 6 |
| 2025 | Survey on deep learning-based weakly supervised salient object detection
Lina Du, Chunjie Ma, Huimin Zheng, Xiushan Nie, Zan Gao 0001 |
Expert Syst. Appl. | 6 |
| 2025 | KA-MIN: Knowledge-Aware Multimodal Interaction Network for Emotion Recognition in ConversationabstractEmotion recognition in conversations (ERC) has garnered significant attention for its critical role in human-computer interaction systems. ERC benefits from multimodal data, which offers diverse perspectives on emotional states, and commonsense knowledge (CSK), which enriches the context by incorporating real-world understanding of human behavior. However, existing ERC studies have not fully exploited the potential of multimodal-CSK interactions for complementary information learning from these sources. To address this, we innovatively propose a Knowledge-Aware Multimodal Interaction Network (KA-MIN). KA-MIN is designed to capture complementary emotional information from CSK-multimodal interactions, thereby facilitating the ERC task. To achieve this, KA-MIN begins by combining six relation types of CSK, leveraging their differences between multimodal emotional information. The fused CSK features are then refined to incorporate context and emotional information using multimodal contextual guidance. Subsequently, we construct a novel knowledge-aware multimodal graph structure that allows the CSK information to interact with multimodal information, leading to more comprehensive multimodal and context modeling. During the graph learning process, the CSK-multimodal interactions capture the complementary emotional information between CSK and multimodal features. Finally, we dynamically fuse the multimodal emotional information with the informative CSK and textual guidance to obtain the final utterance representations, which encompass effective emotional information from both multimodal and CSK features. Extensive experiments on two popular multimodal ERC datasets demonstrate the superiority and effectiveness of the proposed KA-MIN framework. Minjie Ren, Xiangdong Huang 0002, Jing Liu 0002, Zan Gao 0001, Yuting Su 0001, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Fine-Grained Modality Relation-Aware Network for Video Moment RetrievalabstractVideo moment retrieval (VMR) involves localizing video segments semantically aligned with given queries within videos. Despite the development of numerous methods for VMR in recent years, there remains a need to better incorporate fine-grained modality relation-aware information both in intra-modality and cross-modality. To address these challenges, we propose a Fine-grained Modality Relation-Aware Network (FMRN) tailored for the video moment retrieval task. FMRN effectively explores fine-grained modality relation-aware information within text queries, videos, and proposals. Our approach begins with a semantic graph encoder to capture deep semantic relations in intra-modality. Besides, we introduce a novel fine-grained cross-modality interaction module comprising a cross-similarity weighting module, an intra-modality weighting module, and an adaptive fusion module. These components comprehensively exploit fine-grained relation information within intra-modality and cross-modality contexts. Specifically, the cross-similarity weighting module leverages similarities between text queries and video snippets, as well as between videos and query words. The intra-modality weighting module determines the importance of words and snippets, while the adaptive fusion module combines cross-similarity weighting and intra-modality weighting. Additionally, we design a proposal relation module to enhance retrieval by capturing fine-grained proposals-relation information in videos. Extensive experiments demonstrate that the proposed method can outperform all state-of-the-art methods on the TACoS dataset and obtain comparable results on the Charades-STA and ActivityNet-Captions datasets. Compared with MCMN (TCSVT2024) and DPHANet (TMM2024), FMRN can achieve average improvements of 3.61 % and 5.44 % on the TACoS dataset, respectively. Yibo Zhao 0001, Zan Gao 0001, Chunjie Ma, Weili Guan, Riwei Wang, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | MFCLIP: Multi-Modal Fine-Grained CLIP for Generalizable Diffusion Face Forgery DetectionabstractThe rapid development of photo-realistic face generation methods has raised significant concerns in society and academia, highlighting the urgent need for robust and generalizable face forgery detection (FFD) techniques. Although existing approaches mainly capture face forgery patterns using image modality, other modalities like fine-grained noises and texts are not fully explored, which limits the generalization capability of the model. In addition, most FFD methods tend to identify facial images generated by GAN, but struggle to detect unseen diffusion-synthesized ones. To address the limitations, we aim to leverage the cutting-edge foundation model, contrastive language-image pre-training (CLIP), to achieve generalizable diffusion face forgery detection (DFFD). In this paper, we propose a novel multi-modal fine-grained CLIP (MFCLIP) model, which mines comprehensive and fine-grained forgery traces across image-noise modalities via language-guided face forgery representation learning, to facilitate the advancement of DFFD. Specifically, we devise a fine-grained language encoder (FLE) that extracts fine global language features from hierarchical text prompts. We design a multi-modal vision encoder (MVE) to capture global image forgery embeddings as well as fine-grained noise forgery patterns extracted from the richest patch, and integrate them to mine general visual forgery traces. Moreover, we build an innovative plug-and-play sample pair attention (SPA) method to emphasize relevant negative pairs and suppress irrelevant ones, allowing cross-modality sample pairs to conduct more flexible alignment. Extensive experiments and visualizations show that our model outperforms the state of the arts on different settings like cross-generator, cross-forgery, and cross-dataset evaluations. Our code will be available at https://github.com/Jenine-321/MFCLIP. Tianyi Wang 0006, Zitong Yu, Zan Gao 0001, LinLin Shen, Shengyong Chen |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Multiple Information Prompt Learning for Cloth-Changing Person Re-IdentificationabstractCloth-changing person re-identification is a subject closer to the real world, which focuses on solving the problem of person re-identification after pedestrians change clothes. The primary challenge in this field is to overcome the complex interplay between intra-class and inter-class variations and to identify features that remain unaffected by changes in appearance. Sufficient data collection for model training would significantly aid in addressing this problem. However, it is challenging to gather diverse datasets in practice. Current methods focus on implicitly learning identity information from the original image or introducing additional auxiliary models, which are largely limited by the quality of the image and the performance of the additional model. To address these issues, inspired by prompt learning, we propose a novel multiple information prompt learning (MIPL) scheme for cloth-changing person ReID, which learns identity robust features through the common prompt guidance of multiple messages. Specifically, the clothing information stripping (CIS) module is designed to decouple the clothing information from the original RGB image features to counteract the influence of clothing appearance. The bio-guided attention (BGA) module is proposed to increase the learning intensity of the model for key information. A dual-length hybrid patch (DHP) module is employed to make the features have diverse coverage to minimize the impact of feature bias. Extensive experiments demonstrate that the proposed method outperforms all state-of-the-art methods on the LTCC, CelebreID, Celeb-reID-light, and CSCC datasets, achieving rank-1 scores of 74.8%, 73.3%, 66.0%, and 88.1%, respectively. When compared to AIM (CVPR23), ACID (TIP23), and SCNet (MM23), MIPL achieves rank-1 improvements of 11.3%, 13.8%, and 7.9%, respectively, on the PRCC dataset. Shengxun Wei, Zan Gao 0001, Chunjie Ma, Yibo Zhao 0001, Weili Guan, Shengyong Chen |
IEEE Trans. Image Process. | 2 |
| 2025 | A Semantic-Aware Attention and Visual Shielding Network for Cloth-Changing Person Re-IdentificationabstractCloth-changing person re-identification (ReID) is a newly emerging research topic that aims to retrieve pedestrians whose clothes are changed. Since the human appearance with different clothes exhibits large variations, it is very difficult for existing approaches to extract discriminative and robust feature representations. Current works mainly focus on body shape or contour sketches, but the human semantic information and the potential consistency of pedestrian features before and after changing clothes are not fully explored or are ignored. To solve these issues, in this work, a novel semantic-aware attention and visual shielding network for cloth-changing person ReID (abbreviated as SAVS) is proposed where the key idea is to shield clues related to the appearance of clothes and only focus on visual semantic information that is not sensitive to view/posture changes. Specifically, a visual semantic encoder is first employed to locate the human body and clothing regions based on human semantic segmentation information. Then, a human semantic attention (HSA) module is proposed to highlight the human semantic information and reweight the visual feature map. In addition, a visual clothes shielding (VCS) module is also designed to extract a more robust feature representation for the cloth-changing task by covering the clothing regions and focusing the model on the visual semantic information unrelated to the clothes. Most importantly, these two modules are jointly explored in an end-to-end unified framework. Extensive experiments demonstrate that the proposed method can significantly outperform state-of-the-art methods, and more robust features can be extracted for cloth-changing persons. Compared with multibiometric unified network (MBUNet) (published in TIP2023), this method can achieve improvements of 17.5% (30.9%) and 8.5% (10.4%) on the LTCC and Celeb-reID datasets in terms of mean average precision (mAP) (rank-1), respectively. When compared with the Swin Transformer (Swin-T), the improvements can reach 28.6% (17.3%), 22.5% (10.0%), 19.5% (10.2%), and 8.6% (10.1%) on the PRCC, LTCC, Celeb, and NKUP datasets in terms of rank-1 (mAP), respectively. Zan Gao 0001, Hongwei Wei, Weili Guan, Jie Nie, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | A Coarse to Fine Detection Method for Prohibited Object in X-ray Images Based on Progressive Transformer Decoder
Chunjie Ma, Lina Du, Zan Gao 0001, Li Zhuo 0001, Meng Wang 0001 |
ACM Multimedia | 3 |
| 2024 | Identity-Guided Collaborative Learning for Cloth-Changing Person ReidentificationabstractCloth-changing person reidentification (ReID) is a newly emerging research topic aimed at addressing the issues of large feature variations due to cloth-changing and pedestrian view/pose changes. Although significant progress has been achieved by introducing extra information (e.g., human contour sketching information, human body keypoints, and 3D human information), cloth-changing person ReID remains challenging because pedestrian appearance representations can change at any time. Moreover, human semantic information and pedestrian identity information are not fully explored. To solve these issues, we propose a novel identity-guided collaborative learning scheme (IGCL) for cloth-changing person ReID, where the human semantic is effectively utilized and the identity is unchangeable to guide collaborative learning. First, we design a novel clothing attention degradation stream to reasonably reduce the interference caused by clothing information where clothing attention and mid-level collaborative learning are employed. Second, we propose a human semantic attention and body jigsaw stream to highlight the human semantic information and simulate different poses of the same identity. In this way, the extraction features not only focus on human semantic information that is unrelated to the background but are also suitable for pedestrian pose variations. Moreover, a pedestrian identity enhancement stream is proposed to enhance the identity importance and extract more favorable identity robust features. Most importantly, all these streams are jointly explored in an end-to-end unified framework, and the identity is utilized to guide the optimization. Extensive experiments on six public clothing person ReID datasets (LaST, LTCC, PRCC, NKUP, Celeb-reID-light, and VC-Clothes) demonstrate the superiority of the IGCL method. It outperforms existing methods on multiple datasets, and the extracted features have stronger representation and discrimination ability and are weakly correlated with clothing. Zan Gao 0001, Shengxun Wei, Weili Guan, Lei Zhu 0002, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Fine-Grained Temporal-Enhanced Transformer for Dynamic Facial Expression RecognitionabstractDynamic facial expression recognition (DFER) plays a vital role in understanding human emotions and behaviors. Existing efforts tend to fall into a single modality self-supervised pretraining learning paradigm, which limits the representation ability of models. Besides, coarse-grained temporal modeling struggles to capture subtle facial expression representations from various inputs. In this letter, we propose a novel method for DFER, termed fine-grained temporal-enhanced transformer (FTET-DFER), which consists of two stages. First, we employ the inherent correlation between visual and auditory modalities in real videos, to capture temporally dense representations such as facial movements and expressions, in a self-supervised audio-visual learning manner. Second, we utilize the learned embeddings as targets, to achieve the DFER. In addition, we design the FTET block to study fine-grained temporal-enhanced facial expression features based on intra-clip locally-enhanced relations as well as inter-clip locally-enhanced global relationships in videos. Extensive experiments show that FTET-DFER outperforms the state-of-the-arts through within-dataset and cross-dataset evaluation. LinLin Shen, Zitong Yu, Zan Gao 0001 |
IEEE Signal Process. Lett. | 5 |
| 2024 | A Semantic Perception and CNN-Transformer Hybrid Network for Occluded Person Re-IdentificationabstractThe objective of the occluded person re-identification (ReID) task is to capture the same person from different camera angles when the pedestrian’s body is partially occluded. In this task, there are two main challenges: 1) pedestrians are often occluded by other persons or objects, and 2) pedestrians change poses. Moreover, these two issues often simultaneously occur. Although many occluded person ReID algorithms have been proposed, many existing methods can often only solve one of these issues well, and the other issue is often ignored. In this work, a novel semantic perception and CNN-transformer hybrid network (abbreviated as SPH) is proposed for occluded person ReID, which consists of a CNN-based human semantic perception stream and a transformer-based pose perception stream. In the former, a human semantic auxiliary module and a human semantic perception module are designed to obtain human semantic information where multi-granularity region features of the human body are extracted to solve the issues of occlusion. In the latter, we propose a token-based pose integration module to obtain the corresponding patch for each pose key-point and the relative position information to solve the change in pedestrian pose. Moreover, these two streams are jointly optimized in a unified framework. In addition, to further solve the issue of occlusion, the human completion strategy is proposed for the query sample where the gallery samples are used to complete the missing parts of the query. Extensive experimental results on three public occluded person ReID datasets, Occluded-DukeMTMC, P-DukeMTMC-reID, and Occluded-REID, demonstrate that the proposed method can outperform all SOTA occluded person ReID methods in terms of the mAP and Rank-1. Compared with PAT (CVPR21) on the Occluded-DukeMTMC and Occluded-REID datasets, the improvements in mAP/Rank-1 reached 10.1%/7.4%, and 10%/1%, respectively. Moreover, when TransReID (ICCV21) was used, SPH achieved improvements of 4.5% (mAP) and 5.5% (Rank-1) on the Occluded-DukeMTMC dataset. Zan Gao 0001, Peng Chen 0047, Tao Zhuo, Meng Liu 0006, Lei Zhu 0002, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Semantic-Aware Contrastive Learning With Proposal Suppression for Video Semantic Role GroundingabstractVideo semantic role grounding has gained substantial interest from both the academic and industrial communities. While existing methods have demonstrated considerable performance improvements, the influence of noisy and intra-object proposals, referring to proposals with the same object label, has yet to be explored in video semantic role grounding. In this study, we propose a semantic-aware contrastive learning network with proposal suppression to enhance the accuracy of grounding referenced objects. To fully exploit the semantic information in each semantic role, we introduce a novel semantic role encoding module that allows for precise representations of each semantic role. We also design a semantic-aware proposal suppression network to reduce the impact of noisy proposals on object representation learning. Additionally, we propose a proposal contrastive loss to improve cross-modal alignment and reduce the effect of irrelevant intra-object proposals. Extensive experiments on four datasets demonstrate that our model achieves significant improvements over state-of-the-art methods. Meng Liu 0006, Jie Guo 0012, Xin Luo 0006, Zan Gao 0001, Liqiang Nie |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | A Snippets Relation and Hard-Snippets Mask Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) is a problem learning an action localization model with only video-level labels available. In recent years, many WTAL methods have developed. However, hard-to-predict snippets near action boundaries are often not considered in these existing approaches, causing action incompleteness and action over-complete issues. To solve these issues, in this work, an end-to-end snippets relation and hard-snippets mask network (SRHN) is proposed. Specifically, a hard-snippets mask module is applied to mask the hard-to-predict snippets adaptively, and in this way, the trained model focuses more on those snippets with low uncertainty. Then, a snippets relation module is designed to capture the relationship among snippets and can make hard-to-predict snippets easy to predict by aggregating the information of multiple temporal receptive fields. Finally, a snippet enhancement loss is further developed to reduce the action probabilities that are not present in videos for hard-to-predict snippets and other snippets, enlarging the action probabilities that exist in videos. Extensive experiments on THUMOS14, ActivityNet1.2, and ActivityNet1.3 datasets demonstrate the effectiveness of the SRHN method. Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Weili Guan, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | GenFace: A Large-Scale Fine-Grained Face Forgery Benchmark and Cross Appearance-Edge LearningabstractThe rapid advancement of photorealistic generators has reached a critical juncture where the discrepancy between authentic and manipulated images is increasingly indistinguishable. Thus, benchmarking and advancing techniques detecting digital manipulation become an urgent issue. Although there have been a number of publicly available face forgery datasets, the forgery faces are mostly generated using GAN-based synthesis technology, which does not involve the most recent technologies like diffusion. The diversity and quality of images generated by diffusion models have been significantly improved and thus a much more challenging face forgery dataset shall be used to evaluate SOTA forgery detection literature. In this paper, we propose a large-scale, diverse, and fine-grained high-fidelity dataset, namely GenFace, to facilitate the advancement of deepfake detection, which contains a large number of forgery faces generated by advanced generators such as the diffusion-based model and more detailed labels about the manipulation approaches and adopted generators. In addition to evaluating SOTA approaches on our benchmark, we design an innovative Cross Appearance-Edge Learning (CAEL) detector to capture multi-grained appearance and edge global representations, and detect discriminative and general forgery traces. Moreover, we devise an Appearance-Edge Cross-Attention (AECA) module to explore the various integrations across two domains. Extensive experiment results and visualizations show that our detection model outperforms the state of the arts on different settings like cross-generator, cross-forgery, and cross-dataset evaluations. Code and datasets will be available athttps://github.com/Jenine-321/GenFace. Zitong Yu, Tianyi Wang 0006, Xiaobin Huang, LinLin Shen, Zan Gao 0001, Jianfeng Ren |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Talking Face Generation With Audio-Deduced Emotional LandmarksabstractThe goal of talking face generation is to synthesize a sequence of face images of the specified identity, ensuring the mouth movements are synchronized with the given audio. Recently, image-based talking face generation has emerged as a popular approach. It could generate talking face images synchronized with the audio merely depending on a facial image of arbitrary identity and an audio clip. Despite the accessible input, it forgoes the exploitation of the audio emotion, inducing the generated faces to suffer from emotion unsynchronization, mouth inaccuracy, and image quality deficiency. In this article, we build a bistage audio emotion-aware talking face generation (AMIGO) framework, to generate high-quality talking face videos with cross-modally synced emotion. Specifically, we propose a sequence-to-sequence (seq2seq) cross-modal emotional landmark generation network to generate vivid landmarks, whose lip and emotion are both synchronized with input audio. Meantime, we utilize a coordinated visual emotion representation to improve the extraction of the audio one. In stage two, a feature-adaptive visual translation network is designed to translate the synthesized landmarks into facial images. Concretely, we proposed a feature-adaptive transformation module to fuse the high-level representations of landmarks and images, resulting in significant improvement in image quality. We perform extensive experiments on the multi-view emotional audio-visual dataset (MEAD) and crowd-sourced emotional multimodal actors dataset (CREMA-D) benchmark datasets, demonstrating that our model outperforms state-of-the-art benchmarks. Shuyan Zhai, Meng Liu 0006, Zan Gao 0001, Lei Zhu 0002, Liqiang Nie |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Universal Relocalizer for Weakly Supervised Referring Expression GroundingabstractThis article introduces the Universal Relocalizer, a novel approach designed for weakly supervised referring expression grounding. Our method strives to pinpoint a target proposal that corresponds to a specific query, eliminating the need for region-level annotations during training. To bolster the localization precision and enrich the semantic understanding of the target proposal, we devise three key modules: the category module, the color module, and the spatial relationship module. The category and color modules assign respective category and color labels to region proposals, enabling the computation of category and color scores. Simultaneously, the spatial relationship module integrates spatial cues, yielding a spatial score for each proposal to enhance localization accuracy further. By adeptly amalgamating the category, color, and spatial scores, we derive a refined grounding score for every proposal. Comprehensive evaluations on the RefCOCO, RefCOCO+, and RefCOCOg datasets manifest the prowess of the Universal Relocalizer, showcasing its formidable performance across the board. Meng Liu 0006, Xuemeng Song, Da Cao, Zan Gao 0001, Liqiang Nie |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | A Novel Temporal Channel Enhancement and Contextual Excavation Network for Temporal Action LocalizationabstractThe temporal action localization (TAL) task aims to locate and classify action instances in untrimmed videos. Most previous methods use classifiers and locators to act on the same feature; thus, the classification and localization processes are relatively independent. Therefore, if the classification results and localization results are fused, there will be a problem that the classification results are correct while the localization results are wrong, resulting in inaccurate final results, and vice versa. To solve this problem, we propose a novel temporal channel enhancement and contextual excavation network (TCN) for the TAL task, which generates robust classification and localization features and refines the final localization results. Specifically, a temporal channel enhancement module is designed to enhance the temporal and channel information of the feature sequence. Then, the temporal semantic contextual excavation module is developed to establish relationships between similar frames. Finally, the features with enhanced contextual information are transferred to a classifier. While executing the classification process, we obtain powerful classification features. Most importantly, with the robust classification features, the final localization features are produced by the refine localization module, which is applied to obtain the final localization results. Extensive experiments show that TCN can outperform all the SOTA methods on the THUMOS14 dataset, and achieves a comparable performance on the ActivityNet1.3 dataset. Compared with ActionFormer (ECCV 2022) and BREM (MM 2022) on the THUMOS14 dataset, the proposed TCN can achieve improvements of 1.8% and 5.0%, respectively. Zan Gao 0001, Xinglei Cui, Yibo Zhao 0001, Tao Zhuo, Weili Guan, Meng Wang 0001 |
ACM Multimedia | 1 |
| 2023 | Attribute-Guided Collaborative Learning for Partial Person Re-IdentificationabstractPartial person re-identification (ReID) aims to solve the problem of image spatial misalignment due to occlusions or out-of-views. Despite significant progress through the introduction of additional information, such as human pose landmarks, mask maps, and spatial information, partial person ReID remains challenging due to noisy keypoints and impressionable pedestrian representations. To address these issues, we propose a unified attribute-guided collaborative learning scheme for partial person ReID. Specifically, we introduce an adaptive threshold-guided masked graph convolutional network that can dynamically remove untrustworthy edges to suppress the diffusion of noisy keypoints. Furthermore, we incorporate human attributes and devise a cyclic heterogeneous graph convolutional network to effectively fuse cross-modal pedestrian information through intra- and inter-graph interaction, resulting in robust pedestrian representations. Finally, to enhance keypoint representation learning, we design a novel part-based similarity constraint based on the axisymmetric characteristic of the human body. Extensive experiments on multiple public datasets have shown that our model achieves superior performance compared to other state-of-the-art baselines. Meng Liu 0006, Ming Yan 0008, Zan Gao 0001, Xiaojun Chang, Liqiang Nie |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | A Multitemporal Scale and Spatial-Temporal Transformer Network for Temporal Action LocalizationabstractTemporal action localization plays an important role in video analysis, which aims to localize and classify actions in untrimmed videos. Previous methods often predict actions on a feature space of a single temporal scale. However, the temporal features of a low-level scale lack sufficient semantics for action classification, while a high-level scale cannot provide the rich details of the action boundaries. In addition, the long-range dependencies of video frames are often ignored. To address these issues, a novel multitemporal-scale spatial–temporal transformer (MSST) network is proposed for temporal action localization, which predicts actions on a feature space of multiple temporal scales. Specifically, we first use refined feature pyramids of different scales to pass semantics from high-level scales to low-level scales. Second, to establish the long temporal scale of the entire video, we use a spatial–temporal transformer encoder to capture the long-range dependencies of video frames. Then, the refined features with long-range dependencies are fed into a classifier for coarse action prediction. Finally, to further improve the prediction accuracy, we propose a frame-level self-attention module to refine the classification and boundaries of each action instance. Most importantly, these three modules are jointly explored in a unified framework, and MSST has an anchor-free and end-to-end architecture. Extensive experiments show that the proposed method can outperform state-of-the-art approaches on the THUMOS14 dataset and achieve comparable performance on the ActivityNet1.3 dataset. Compared with A2Net (TIP20, Avg{0.3:0.7}), Sub-Action (CSVT2022, Avg{0.1:0.5}), and AFSD (CVPR21, Avg{0.3:0.7}) on the THUMOS14 dataset, the proposed method can achieve improvements of 12.6%, 17.4%, and 2.2%, respectively. Zan Gao 0001, Xinglei Cui, Tao Zhuo, Zhiyong Cheng 0001, Anan Liu, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2023 | Toward Fine-Grained Talking Face GenerationabstractTalking face generation is the process of synthesizing a lip-synchronized video when given a reference portrait and an audio clip. However, generating a fine-grained talking video is nontrivial due to several challenges: 1) capturing vivid facial expressions, such as muscle movements; 2) ensuring smooth transitions between consecutive frames; and 3) preserving the details of the reference portrait. Existing efforts have only focused on modeling rigid lip movements, resulting in low-fidelity videos with jerky facial muscle deformations. To address these challenges, we propose a novel Fine-gRained mOtioN moDel (FROND), consisting of three components. In the first component, we adopt a two-stream encoder to capture local facial movement keypoints and embed their overall motion context as the global code. In the second component, we design a motion estimation module to predict audio-driven movements. This enables the learning of local key point motion in the continuous trajectory space to achieve smooth temporal facial movements. Additionally, the local and global motions are fused to estimate a continuous dense motion field, resulting in spatially smooth movements. In the third component, we devise a novel implicit image decoder based on an implicit neural network. This decoder recovers high-frequency information from the input image, resulting in a high-fidelity talking face. In summary, the FROND refines the motion trajectories of facial keypoints into a continuous dense motion field, which is followed by a decoder that fully exploits the inherent smoothness of the motion. We conduct quantitative and qualitative model evaluations on benchmark datasets. The experimental results show that our proposed FROND significantly outperforms several state-of-the-art baselines. Zhicheng Sheng, Liqiang Nie, Meng Liu 0006, Yinwei Wei, Zan Gao 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | TBNet: A Two-Stream Boundary-Aware Network for Generic Image Manipulation LocalizationabstractAbstract - Finding tampered regions in images is a common research topic in machine learning and computer vision. Although many image manipulation location algorithms have been proposed, most of them only focus on RGB images with different color spaces, and the frequency information that contains the potential tampering clues is often ignored. Moreover, among the manipulation operations, splicing and copy-move are two frequently used methods, but as their characteristics are quite different, specific methods have been individually designed for detecting the operations of either splicing or copy-move, and it is very difficult to widely apply these methods in practice. To solve these issues, in this work, a novel end-to-end two-stream boundary-aware network (abbreviated as TBNet) is proposed for generic image manipulation localization where the RGB stream, the frequency stream, and the boundary artifact location are explored in a unified framework. Specifically, we first design an adaptive frequency selection module (AFS) to adaptively select the appropriate frequency to mine inconsistent statistics and eliminate the interference of redundant statistics. Then, an adaptive cross-attention fusion module (ACF) is proposed to adaptively fuse the RGB feature and the frequency feature. Finally, the boundary artifact location network (BAL) is designed to locate the boundary artifacts for which the parameters are jointly updated by the outputs of the ACF, and its results are further fed into the decoder. Thus, the parameters of the RGB stream, the frequency stream, and the boundary artifact location network are jointly optimized, and their latent complementary relationships are fully mined. The results of the extensive experiments performed on six public benchmarks of the image manipulation localization task, namely, CASIA1.0, COVER, Carvalho, In-The-Wild, NIST-16, and IMD-2020, demonstrate that the proposed TBNet can substantially outperform state-of-the-art generic image manipulation localization methods in terms of MCC, F1, and AUC while maintaining robustness with respect to various attacks. Compared with DeepLabV3+ on the CASIA1.0, COVER, Carvalho, In-The-Wild, and NIST-16 datasets, the improvements in MCC/F1 reach 11%/11.1%, 8.2%/10.3%, 10.2%/11.6%, 8.9%/6.2%, and 13.3%/16.0%, respectively. Moreover, when IMD2020 is utilized, its AUC improvement can achieve 14.7%. Zan Gao 0001, Zhiyong Cheng 0001, Weili Guan, Anan Liu, Meng Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Disentangled Graph Neural Networks for Session-Based RecommendationabstractSession-based recommendation (SBR) has drawn increasingly research attention in recent years, due to its great practical value by only exploiting the limited user behavior history in the current session. The key of SBR is to accurately infer the anonymous user purpose in a session which is typically represented as session embedding, and then match it with the item embeddings for the next item prediction. Existing methods typically learn the session embedding at the item level, namely, aggregating the embeddings of items with or without assigned attention weights to items. However, they ignore the fact that a user's intent on adopting an item is driven by certain factors of the item (e.g., theleading actorsof an movie). In other words, they have not explored finer-granularity interests of users at the factor level to generate the session embedding, leading to sub-optimal performance. To address the problem, we propose a novel method called Disentangled Graph Neural Network (Disen-GNN) to capture the session purpose with the consideration of factor-level attention on each item. Specifically, we first employ the disentangled learning technique to cast item embeddings into the embeddings of multiple factors, and then use the gated graph neural network (GGNN) to learn the embedding factor-wisely based on the item adjacent similarity matrix computed for each factor. Moreover, the distance correlation is adopted to enhance the independence between each pair of factors. After representing each item with independent factors, an attention mechanism is designed to learn user intent to different factors of each item in the session. The session embedding is then generated by aggregating the item embeddings with attention weights of each item's factors. To this end, our model takes user intents at the factor level into account to infer the user purpose in a session. Extensive experiments on three benchmark datasets demonstrate the superiority of our method over existing methods. Ansong Li, Zhiyong Cheng 0001, Fan Liu 0008, Zan Gao 0001, Weili Guan, Yuxin Peng 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | A Novel Action Saliency and Context-Aware Network for Weakly-Supervised Temporal Action LocalizationabstractTemporal action localization is a challenging task in computer vision, and it tries to find the start time and the end time of the actions and predict their categories. However, compared to temporal action localization, weakly supervised temporal action localization (WTAL) is a more challenging task due to its poor annotations. With only video-level annotation, some background frames, similar to actions, would be classified as actions and produce inaccurate results. In addition, the two-stream fusion problem, ignored previously, also needs to be further considered. To resolve these issues, we propose a novel action saliency and context-aware network (ASCN) for weakly supervised temporal action localization tasks. Specifically, the temporal saliency and context module is designed to enhance the global saliency and context information of the RGB and the flow features to suppress the backgrounds and enhance the actions. In addition, a hybrid attention mechanism using frame differences and two-stream attention is designed to model the local action context information and further enlarge the scores of the potential action regions and suppress the background regions. Finally, to obtain two-stream consistency and solve the fusion problem, we use the similarity loss and a channel self-attention module to adaptively fuse the enhanced RGB and flow features. Extensive experiments demonstrate that ASCN can outperform all of the SOTA WTAL methods on the THUMOS14 dataset and the ActivityNet1.3 dataset with an average mAP that can reach 37.2% on the THUMOS14 dataset and attains an average mAP of 26.3% on the ActivityNet1.3 dataset. On the ActivityNet1.2 dataset, ASCN can also obtain comparable results. Compared with AdapNet (TNNLS20), MMSD (TIP22), and FTCL (CVPR22) on the THUMOS14 dataset, ASCN can outperform them by 13.5%, 2.9%, and 2.8%, respectively. Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Wen Gao 0001, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Multim. | 3 |
| 2023 | Review Polarity-Wise RecommenderabstractThe de facto review-involved recommender systems, using review information to enhance recommendation, have received increasing interest over the past years. Thereinto, one advanced branch is to extract salient aspects from textual reviews (i.e., the item attributes that users express) and combine them with the matrix factorization (MF) technique. However, the existing approaches all ignore the fact that semantically different reviews often include opposite aspect information. In particular, positive reviews usually express aspects that users prefer, while the negative ones describe aspects that users dislike. As a result, it may mislead the recommender systems into making incorrect decisions pertaining to user preference modeling. Toward this end, in this article, we present a review polarity-wise recommender model, dubbed as RPR, to discriminately treat reviews with different polarities. To be specific, in this model, positive and negative reviews are separately gathered and used to model the user-preferred and user-rejected aspects, respectively. Besides, to overcome the imbalance of semantically different reviews, we further develop an aspect-aware importance weighting strategy to align the aspect importance for these two kinds of reviews. Extensive experiments conducted on eight benchmark datasets have demonstrated the superiority of our model when compared with several state-of-the-art review-involved baselines. Moreover, our method can provide certain explanations to real-world rating prediction scenarios. Jianhua Yin 0001, Zan Gao 0001, Liqiang Nie |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Generic Image Manipulation Localization through the Lens of Multi-scale Spatial InconsistenceabstractImage manipulation localization is of vital importance to public order protection. One dominant approach is to detect the anomalies in images, i.e., visual artifacts, as the tampered edge clue for aiding manipulation prediction. Nevertheless, we argue that these methods struggle with the modeling of spatial inconsistency within multi-scale, resulting in sub-optimal model performance. To overcome this problem, in this paper, we propose a novel end-to-end method to identify the multi-scale spatial inconsistency for image manipulation localization (abbreviated as MSI) where the multi-scale edge-guided attention stream (MEA) and multi-scale context-aware search stream (MCS) are jointly explored in a unified framework, moreover, multi-scale information is efficiently used. In the former, the edge-attention module is designed to precisely locate the tampered regions based upon multi-scale edge boundary features. In the latter, the context-aware search module is designed to model spatial contextual information within multiple scales. To validate the effectiveness of the proposed method, we conduct extensive experiments on six image manipulation localization datasets including NIST-2016, Columbia, CASIA1.0, COVER, DEF-12K, and IMD2020. The experimental results demonstrate that our proposed method can outperform state-of-the-art methods by a significant margin in terms of average F1 score while maintaining robustness with respect to various attacks. Compared with MVSS-Net (Published in ICCV 2021) on the NIST-2016, CASIA1.0, DEF-12K, and IMD2020 datasets, the improvements in F1 score can reach 6.7%, 9.5%, 5.4%, and 8.4%, respectively. Zan Gao 0001, Shenghao Chen, Weili Guan, Jie Nie, Anan Liu |
ACM Multimedia | 1 |
| 2022 | Multiview clustering via consistent and specific nonnegative matrix factorization with graph regularization
Haixia Xu 0003, Limin Gong, Hai-Zhen Xuan, Xusheng Zheng, Zan Gao 0001, Xianbin Wen |
Multim. Syst. | 5 |
| 2022 | Editorial paper for Pattern Recognition Letters VSI on cross model understanding for visual question answering
Shaohua Wan 0001, Zan Gao 0001, Hanwang Zhang, Xiaojun Chang, Chen Chen 0001, Anastasios Tefas |
Pattern Recognit. Lett. | 2 |
| 2022 | Multi-Level View Associative Convolution Network for View-Based 3D Model RetrievalabstractWith the continuous improvement of image processing capabilities, a three-dimensional (3D) model that can contain rich information is becoming the fourth type of multimedia data (in addition to sound, image, and video). Moreover, since there is a wide range of applications of 3D models, how to quickly and effectively obtain the correct target model from the massive data has become a key issue. To date, 3D model retrieval approaches have been proposed, and in these approaches, view-based 3D model retrieval methods can achieve satisfactory performance. In the 3D model retrieval task, the latent relationship mining of all images in a 3D model, the adaptive fusion of different images, and the discriminative feature extraction are the main challenges, but in most existing solutions, these issues are separately performed and they are not explored in an end-to-end network architecture. To solve these issues, in this work, we propose a novel and effective multi-level view associative convolution network (MLVACN) to realize view-based 3D model retrieval, where the relationship exploration of multiple-view images, the fusion of different images, and the feature discrimination learning are realized in a unified end-to-end framework. Specifically, we design the group association layer and the block association layer to study the latent relationships among different views from the view-level and the block-level, respectively. Moreover, the weight fusion layer is further designed to adaptively fuse different views in a 3D model. In addition, these three layers are embedded into theMLVACN. Finally, the pairwise discrimination loss function is proposed to learn the discriminative features of the 3D model. Extensive experimental results on three 3D model retrieval datasets including ModelNet40, ModelNet10, and ShapeNetCore55 demonstrate thatMLVACNcan outperform state-of-the-art methods in term of mAP. When the ModelNet40 dataset is used, the mAP ofMLVACNis improved by 13.25%, 7.75%, 3.95%, and 0.61% as compared to those of the MVCNN, GVCNN, PVNet, and MLVCNN methods, respectively. Zan Gao 0001, Yan Zhang 0154, Hua Zhang 0003, Weili Guan, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | A Novel Multiple-View Adversarial Learning Network for Unsupervised Domain Adaptation Action RecognitionabstractAbstract-domain adaptation action recognition is a hot research topic in machine learning and some effective approaches have been proposed. However, samples in the target domain with label information are often required by these approaches. Moreover, domain-invariant discriminative feature learning, feature fusion, and classifier module learning have not been explored in an end-to-end framework. Thus, in this study, we propose a novel end-to-end multiple-view adversarial learning network (MAN) for unsupervised domain adaptation action recognition in which the fusion of RGB and optical-flow features, domain-invariant discrimination feature learning, and action recognition is conducted in a unified framework. Specifically, a robust spatiotemporal feature extraction network, including a spatial transform network and an adaptive intrachannel weight network, is proposed to improve the scale invariance and robustness of the method. Then, a self-attention mechanism fusion module is designed to adaptively fuse the RGB and optical-flow features. Moreover, a multiview adversarial learning loss is developed to obtain domain-invariant discriminative features. In addition, three benchmark datasets are constructed for unsupervised domain adaptation action recognition, for which all actions and samples are carefully collected from public action datasets, and their action categories are hierarchically augmented, which can guide how to extend existing action datasets. We conduct extensive experiments on four benchmark datasets, and the experimental results demonstrate that our proposed MAN can outperform several state-of-the-art unsupervised domain adaptation action recognition approaches. When the SDAI Action II-6 and SDAI Action II-11 datasets are used, MAN can achieve 3.7% ( H → U ) and 6.1% ( H → U ) improvements over the temporal attentive adversarial adaptation network (published in ICCV 2019) module, respectively. As an added contribution, the SDAI Action II-6, SDAI Action II-11, and SDAI Action II-16 datasets will be released to facilitate future research on domain adaptation action recognition. Zan Gao 0001, Yibo Zhao 0001, Hua Zhang 0003, Da Chen 0002, Anan Liu, Shengyong Chen |
IEEE Trans. Cybern. | 1 |
| 2022 | A Temporal-Aware Relation and Attention Network for Temporal Action LocalizationabstractTemporal action localization is currently an active research topic in computer vision and machine learning due to its usage in smart surveillance. It is a challenging problem since the categories of the actions must be classified in untrimmed videos and the start and end of the actions need to be accurately found. Although many temporal action localization methods have been proposed, they require substantial amounts of computational resources for the training and inference processes. To solve these issues, in this work, a novel temporal-aware relation and attention network (abbreviated as TRA) is proposed for the temporal action localization task. TRA has an anchor-free and end-to-end architecture that fully uses temporal-aware information. Specifically, a temporal self-attention module is first designed to determine the relationship between different temporal positions, and more weight is given to features within the actions. Then, a multiple temporal aggregation module is constructed to aggregate the temporal domain information. Finally, a graph relation module is designed to obtain the aggregated graph features, which are used to refine the boundaries and classification results. Most importantly, these three modules are jointly explored in a unified framework, and temporal awareness is always fully used. Extensive experiments demonstrate that the proposed method can outperform all state-of-the-art methods on the THUMOS14 dataset with an average mAP that reaches 67.6% and obtain a comparable result on the ActivityNet1.3 dataset with an average mAP that reaches 34.4%. Compared with A2Net (TIP20), PCG-TAL (TIP21), and AFSD (CVPR21) TRA can achieve improvements of 11.7%, 4.4%, and 1.8%, respectively on the THUMOS14 dataset. Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Weili Guan, Jie Nie, Anan Liu, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Image Process. | 3 |
| 2022 | Frame-Wise Cross-Modal Matching for Video Moment RetrievalabstractVideo moment retrieval targets at retrieving a golden moment in a video for a given natural language query. The main challenges of this task include 1) the requirement of accurately localizing (i.e., the start time and the end time of) the relevant moment in an untrimmed video stream, and 2) bridging the semantic gap between textual query and video contents. To tackle those problems, early approaches adopt the sliding window or uniform sampling to collect video clips first and then match each clip with the query to identify relevant clips. Obviously, these strategies are time-consuming and often lead to unsatisfied accuracy in localization due to the unpredictable length of the golden moment. To avoid the limitations, researchers recently attempt to directly predict the relevant moment boundaries without the requirement to generate video clips first. One mainstream approach is to generate a multimodal feature vector for the target query and video frames (e.g., concatenation) and then use a regression approach upon the multimodal feature vector for boundary detection. Although some progress has been achieved by this approach, we argue that those methods have not well captured the cross-modal interactions between the query and video frames. In this paper, we propose an Attentive Cross-modal Relevance Matching (ACRM) model which predicts the temporal boundaries based on an interaction modeling between two modalities. In addition, an attention module is introduced to automatically assign higher weights to query words with richer semantic cues, which are considered to be more important for finding relevant video contents. Another contribution is that we propose an additional predictor to utilize the internal frames in the model training to improve the localization accuracy. Extensive experiments on two public datasets TACoS and Charades-STA demonstrate the superiority of our method over several state-of-the-art methods. Ablation studies have been also conducted to examine the effectiveness of different modules in our ACRM model. Haoyu Tang 0002, Jihua Zhu, Meng Liu 0006, Zan Gao 0001, Zhiyong Cheng 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Pairwise Two-Stream ConvNets for Cross-Domain Action Recognition With Small DataabstractIn this work, we target cross-domain action recognition (CDAR) in the video domain and propose a novel end-to-end pairwise two-stream ConvNets (PTC) algorithm for real-life conditions, in which only a few labeled samples are available. To cope with the limited training sample problem, we employ pairwise network architecture that can leverage training samples from a source domain and, thus, requires only a few labeled samples per category from the target domain. In particular, a frame self-attention mechanism and an adaptive weight scheme are embedded into the PTC network to adaptively combine the RGB and flow features. This design can effectively learn domain-invariant features for both the source and target domains. In addition, we propose a sphere boundary sample-selecting scheme that selects the training samples at the boundary of a class (in the feature space) to train the PTC model. In this way, a well-enhanced generalization capability can be achieved. To validate the effectiveness of our PTC model, we construct two CDAR data sets (SDAI Action I and SDAI Action II) that include indoor and outdoor environments; all actions and samples in these data sets were carefully collected from public action data sets. To the best of our knowledge, these are the first data sets specifically designed for the CDAR task. Extensive experiments were conducted on these two data sets. The results show that PTC outperforms state-of-the-art video action recognition methods in terms of both accuracy and training efficiency. It is noteworthy that when only two labeled training samples per category are used in the SDAI Action I data set, PTC achieves 21.9% and 6.8% improvement in accuracy over two-stream and temporal segment networks models, respectively. As an added contribution, the SDAI Action I and SDAI Action II data sets will be released to facilitate future research on the CDAR task. Zan Gao 0001, Leming Guo, Tongwei Ren, Anan Liu, Zhiyong Cheng 0001, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | A Novel Patch Convolutional Neural Network for View-based 3D Model RetrievalabstractIn industrial enterprises, effective retrieval of three-dimensional (3-D) computer-aided design (CAD) models can greatly save time and cost in new product development and manufacturing, thus, many researchers have focused on it. Recently, many view-based 3D model retrieval methods have been proposed and have achieved state-of-the-art performance. However, most of these methods focus on extracting more discriminative view-level features and effectively aggregating the multi-view images of a 3D model, and the latent relationship among these multi-view images is not fully explored. Thus, we tackle this problem from the perspective of exploiting the relationships between patch features to capture long-range associations among multi-view images. To capture associations among views, in this work, we propose a novel patch convolutional neural network (PCNN ) for view-based 3D model retrieval. Specifically, we first employ a CNN to extract patch features of each view image separately. Second, a novel neural network module named PatchConv is designed to exploit intrinsic relationships between neighboring patches in the feature space to capture long-range associations among multi-view images. Then, an adaptive weighted view layer is further embedded into PCNN to automatically assign a weight to each view according to the similarity between each view feature and the view-pooling feature. Finally, a discrimination loss function is employed to extract the discriminative 3D model feature, which consists of softmax loss values generated by the fusion classifier and the specific classifier. Extensive experimental results on two public 3D model retrieval benchmarks, namely, the ModelNet40, and ModelNet10, demonstrate that our proposed PCNN can outperform state-of-the-art approaches, with mAP values of 93.67%, and 96.23%, respectively. Zan Gao 0001, Yuxiang Shao, Weili Guan, Meng Liu 0006, Zhiyong Cheng 0001, Shengyong Chen |
ACM Multimedia | 1 |
| 2021 | Dynamic Modality Interaction Modeling for Image-Text RetrievalabstractImage-text retrieval is a fundamental and crucial branch in information retrieval. Although much progress has been made in bridging vision and language, it remains challenging because of the difficult intra-modal reasoning and cross-modal alignment. Existing modality interaction methods have achieved impressive results on public datasets. However, they heavily rely on expert experience and empirical feedback towards the design of interaction patterns, therefore, lacking flexibility. To address these issues, we develop a novel modality interaction modeling network based upon the routing mechanism, which is the first unified and dynamic multimodal interaction framework towards image-text retrieval. In particular, we first design four types of cells as basic units to explore different levels of modality interactions, and then connect them in a dense strategy to construct a routing space. To endow the model with the capability of path decision, we integrate a dynamic router in each cell for pattern exploration. As the routers are conditioned on inputs, our model can dynamically learn different activated paths for different data. Extensive experiments on two benchmark datasets, i.e., Flickr30K and MS-COCO, verify the superiority of our model compared with several state-of-the-art baselines. Leigang Qu, Meng Liu 0006, Jianlong Wu, Zan Gao 0001, Liqiang Nie |
SIGIR | 4 |
| 2021 | Interest-aware Message-Passing GCN for RecommendationabstractGraph Convolution Networks (GCNs) manifest great potential in recommendation. This is attributed to their capability on learning good user and item embeddings by exploiting the collaborative signals from the high-order neighbors. Like other GCN models, the GCN based recommendation models also suffer from the notorious over-smoothing problem – when stacking more layers, node embeddings become more similar and eventually indistinguishable, resulted in performance degradation. The recently proposed LightGCN and LR-GCN alleviate this problem to some extent, however, we argue that they overlook an important factor for the over-smoothing problem in recommendation, that is, high-order neighboring users with no common interests of a user can be also involved in the user’s embedding learning in the graph convolution operation. As a result, the multi-layer graph convolution will make users with dissimilar interests have similar embeddings. In this paper, we propose a novel Interest-aware Message-Passing GCN (IMP-GCN) recommendation model, which performs high-order graph convolution inside subgraphs. The subgraph consists of users with similar interests and their interacted items. To form the subgraphs, we design an unsupervised subgraph generation module, which can effectively identify users with common interests by exploiting both user feature and graph structure. To this end, our model can avoid propagating negative information from high-order neighbors into embedding learning. Experimental results on three large-scale benchmark datasets show that our model can gain performance improvement by stacking more layers and outperform the state-of-the-art GCN-based recommendation models significantly. Fan Liu 0008, Zhiyong Cheng 0001, Lei Zhu 0002, Zan Gao 0001, Liqiang Nie |
WWW | 4 |
| 2021 | A Pairwise Attentive Adversarial Spatiotemporal Network for Cross-Domain Few-Shot Action Recognition-R2abstractAction recognition is a popular research topic in the computer vision and machine learning domains. Although many action recognition methods have been proposed, only a few researchers have focused on cross-domain few-shot action recognition, which must often be performed in real security surveillance. Since the problems of action recognition, domain adaptation, and few-shot learning need to be simultaneously solved, the cross-domain few-shot action recognition task is a challenging problem. To solve these issues, in this work, we develop a novel end-to-end pairwise attentive adversarial spatiotemporal network (PASTN) to perform the cross-domain few-shot action recognition task, in which spatiotemporal information acquisition, few-shot learning, and video domain adaptation are realised in a unified framework. Specifically, the Resnet-50 network is selected as the backbone of the PASTN, and a 3D convolution block is embedded in the top layer of the 2D CNN (ResNet-50) to capture the spatiotemporal representations. Moreover, a novel attentive adversarial network architecture is designed to align the spatiotemporal dynamics actions with higher domain discrepancies. In addition, the pairwise margin discrimination loss is designed for the pairwise network architecture to improve the discrimination of the learned domain-invariant spatiotemporal feature. The results of extensive experiments performed on three public benchmarks of the cross-domain action recognition datasets, including SDAI Action I, SDAI Action II and UCF50-OlympicSport, demonstrate that the proposed PASTN can significantly outperform the state-of-the-art cross-domain action recognition methods in terms of both the accuracy and computational time. Even when only two labelled training samples per category are considered in the office1 scenario of the SDAI Action I dataset, the accuracy of the PASTN is improved by 6.1%, 10.9%, 16.8%, and 14% compared to that of the $TA^{3}N$ , TemporalPooling, I3D, and P3D methods, respectively. Zan Gao 0001, Leming Guo, Weili Guan, Anan Liu, Tongwei Ren, Shengyong Chen |
IEEE Trans. Image Process. | 1 |
| 2021 | Video Moment Localization via Deep Cross-Modal HashingabstractDue to the continuous booming of surveillance and Web videos, video moment localization, as an important branch of video content analysis, has attracted wide attention from both industry and academia in recent years. It is, however, a non-trivial task due to the following challenges: temporal context modeling, intelligent moment candidate generation, as well as the necessary efficiency and scalability in practice. To address these impediments, we present a deep end-to-end cross-modal hashing network. To be specific, we first design a video encoder relying on a bidirectional temporal convolutional network to simultaneously generate moment candidates and learn their representations. Considering that the video encoder characterizes temporal contextual structures at multiple scales of time windows, we can thus obtain enhanced moment representations. As a counterpart, we design an independent query encoder towards user intention understanding. Thereafter, a cross-model hashing module is developed to project these two heterogeneous representations into a shared isomorphic Hamming space for compact hash code learning. After that, we can effectively estimate the relevance score of each "moment-query" pair via the Hamming distance. Besides effectiveness, our model is far more efficient and scalable since the hash codes of videos can be learned offline. Experimental results on real-world datasets have justified the superiority of our model over several state-of-the-art competitors. Yupeng Hu 0003, Meng Liu 0006, Xiaobin Su, Zan Gao 0001, Liqiang Nie |
IEEE Trans. Image Process. | 4 |
| 2021 | DCR: A Unified Framework for Holistic/Partial Person ReIDabstractPerson reidentification (ReID) is a very popular research topic in machine learning and computer vision. According to the occlusions, it can be divided into holistic person ReID and partial person ReID tasks. Occlusions commonly exist in the partial person ReID task but pose or observation perspective changes often occur in the holistic person ReID task; thus, many different algorithms or different network architectures have been designed for each task. However, this approach increases the cost in practice and hinders the development of ReID techniques. To solve this problem, in this work, a unified framework is proposed for holistic/partial person ReID, which can effectively and efficiently address changes in pose or observation perspective and the occlusions in both tasks. In detail, we first employ a fully convolutional network (FCN) to generate feature maps for an arbitrarily sized image and then use spatial pyramid pooling (SPP) to obtain its spatial pyramid feature. Thereafter, to efficiently solve the matching problem between the query image and gallery images, we build a deep spatial pyramid feature collaborative reconstruction model (DCR). In DCR, the reconstruction errors mainly come from similar blocks (uncovered parts), and the influence of the reconstruction errors of dissimilar blocks (covered parts or changed parts) is minimized. In addition, we also use the deep mutual learning approach to jointly learn the features in the training process and promote model training. Experimental results on two partial person ReID datasets and three holistic person ReID datasets demonstrate that the DCR outperforms the state-of-the-art approaches on both tasks and all datasets. Specifically, it outperforms all competitors with a large margin and achieves an improvement of 9.07% and 5.95% over the DSR method (published in CVPR18) on the Partial ReID and Partial-iLIDS datasets with Rank-1, respectively. Similarly, it also achieves an improvement of 5.08% over the VPM method (published in CVPR19) on the DukeMTMC-ReID dataset with Rank-1. Additionally, the running time of our method for each query is more than 7 faster than that of the DSR or DuATM methods. Zan Gao 0001, Li-Shuai Gao, Hua Zhang 0003, Zhiyong Cheng 0001, Richang Hong, Shengyong Chen |
IEEE Trans. Multim. | 1 |
| 2021 | Introduction to the Special Issue on Fine-grained Visual Computingabstractintroduction Share on Introduction to the Special Issue on Fine-grained Visual Computing Editors: Shaohua Wan Zhongnan University of Economics and Law Zhongnan University of Economics and LawView Profile , Zan Gao Qilu University of Technology Qilu University of TechnologyView Profile , Hanwang Zhang Nanyang Technological University Nanyang Technological UniversityView Profile , Xiaojun Chang Monash University Monash UniversityView Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 17Issue 1sJanuary 2021 Article No.: 11pp 1–3https://doi.org/10.1145/3447532Online:31 March 2021Publication History 0citation70DownloadsMetricsTotal Citations0Total Downloads70Last 12 Months24Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Shaohua Wan 0001, Zan Gao 0001, Hanwang Zhang, Xiaojun Chang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Texture Semantically Aligned with Visibility-aware for Partial Person Re-identificationabstractIn real person re-identification (ReID) tasks, pedestrians are often obscured by other pedestrians or objects; moreover, changes in poses or observation perspectives also commonly exist in partial-person ReID. To the best of our knowledge, few works simultaneously focus on these two issues. In this work, we propose a novel texture semantic alignment (TSA) approach with the visibility-aware for partial person ReID task where the occlusion issue and changes in poses are simultaneously explored in an end-to-end unified framework. Specifically, we first employ a texture alignment scheme with the semantic visibility of a person's image to solve the issue of changes in poses that can enhance the alignment and generalization capability of the models. Second, we design a human pose-based partial region alignment scheme to solve the occlusion problem that makes TSA method emphasize the shared body parts. Finally, these two networks jointly learn these aspects. Extensive experimental results demonstrate that our proposed TSA method is very effective and robust for simultaneously handling occlusion and changes in pose, and it can outperform state-of-the-art approaches by a large margin and achieves an improvement of 5% and 6.4% on the rank-1 accuracy over the visibility-aware part model (VPM) method (published in CVPR 2019) on the Partial ReID and Partial-iLIDS datasets, respectively. Li-Shuai Gao, Hua Zhang 0003, Zan Gao 0001, Weili Guan, Zhiyong Cheng 0001, Meng Wang 0001 |
ACM Multimedia | 3 |
| 2020 | Attention feature matching for weakly-supervised video relocalizationabstractLocalizing the desired video clip for a given query in an untrimmed video has been a hot research topic for multimedia understanding. Recently, a new task named video relocalization, in which the query is a video clip, has been raised. Some methods have been developed for this task, however, these methods often require dense annotations of the temporal boundaries inside long videos for training. A more practical solution is the weakly-supervised approach, which only needs the matching information between the query and video. Haoyu Tang 0002, Jihua Zhu, Zan Gao 0001, Tao Zhuo, Zhiyong Cheng 0001 |
MMAsia | 3 |
| 2019 | Deep Spatial Pyramid Features Collaborative Reconstruction for Partial Person ReIDabstractPartial person re-identification (ReID) is a hot research problem in computer vision. Accurate partial ReID is very challenging due to the common occlusion problem. To address this problem, in this paper, we propose a novel D eep spatial pyramid feature C ollaborative R econstruction approach (DCR ) for partial person ReID, which can effectively and efficiently tackle the occlusion in arbitrary sizes. Specifically, a fully convolutional network (FCN) is first leveraged to extract feature maps of an arbitrary-size image, and then the spatial pyramid pooling (SPP) is adopted to obtain spatial pyramid features. Thereafter, our DCR method is designed to efficiently solve the matching problem between the partial person and the holistic person in the partial person ReID task where the occlusion problem often occurs. Experiments on two partial person ReID datasets demonstrate the efficiency and efficacy of the proposed method by comparing to several state-of-the-art partial person ReID approaches. Our method outperforms all the competitors with a large margin and can achieve an improvement of 9.07% and 5.95% over the DSR method on the Partial REID and Partial-iLIDS Person ReID datasets in terms of the Rank-1 accuracy, respectively. Zan Gao 0001, Li-Shuai Gao, Hua Zhang 0003, Zhiyong Cheng 0001, Richang Hong |
ACM Multimedia | 1 |