EDBT 2026 Demo / reviewers in the wild / expert
Yibo Zhao 0001
dblp:54/10290-1
· DBLP profile ↗
19ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0003-4187-1980ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Amplifying Discrepancies: Exploiting Macro and Micro Inconsistencies for Image Manipulation LocalizationabstractThe rapid development of image manipulation technologies poses significant challenges to multimedia forensics, especially in accurate localization of manipulated regions. Existing methods often fail to fully explore the intrinsic discrepancies between manipulated and authentic regions, resulting in sub-optimal performance. To address this limitation, we propose the Focus Region Discrepancy Network (FRD-Net), a novel and efficient framework that significantly enhances manipulation localization by amplifying discrepancies at both macro- and micro-levels. Specifically, our proposed Iterative Clustering Module (ICM) groups features into two discriminative clusters and refines representations via backward propagation from cluster centers, improving the distinction between tampered and authentic regions at the macro level. Thereafter, our Differential Progressive Module (DPM) is constructed to capture fine-grained structural inconsistencies within local neighborhoods and integrate them into a Central Difference Convolution, increasing sensitivity to subtle manipulation details at the micro level. Finally, these complementary modules are seamlessly integrated into a compact architecture that achieves a favorable balance between accuracy and efficiency. Extensive experiments on multiple benchmarks demonstrate that FRD-Net consistently surpasses state-of-the-art methods in terms of manipulation localization performance while maintaining a lower computational cost. Shenghao Chen, Yibo Zhao 0001, Tianyi Wang 0006, Chunjie Ma, Weili Guan, Ming Li 0083, Zan Gao 0001 |
AAAI | 2 |
| 2026 | Unsupervised 2D Image-Based 3D Model Retrieval via Decision Boundary Alignment and Graph Semantic PropagationabstractUnsupervised 2D image-based 3D model retrieval (IBMR) aims to retrieve semantically relevant 3D shapes for a given 2D image query when 3D annotations are unavailable. This setting is challenging due to severe modality gaps, category-imbalanced mini-batches, inconsistent cross-domain decision boundaries, and mismatched semantic neighborhood structures. In this paper, we propose a unified framework that integrates Category-Aligned Sampling (CAS), Decision Boundary Alignment (DBA), and Graph Semantic Propagation (GSP) into a single optimization paradigm. CAS constructs category-consistent mini-batches to stabilize crossmodal learning. Built upon CAS, DBA leverages a masked Margin Disparity Discrepancy to regularize cross-domain class decision boundaries via an adversarial min-max objective, encouraging discriminative separation beyond marginal feature matching. To complement boundary-level regularization, GSP builds a crossdomain affinity graph over 2D and 3D samples and propagates supervision-induced relational structure through semantic message passing, explicitly preserving instance-level neighborhood consistency that is critical for retrieval. Extensive experiments on MI3DOR and MI3DOR-2 demonstrate consistent improvements over representative unsupervised IBMR baselines. Nian Hu, Yibo Zhao 0001, Chen Li 0035, Cong Liu 0012, Zan Gao 0001 |
SIGIR | 2 |
| 2026 | Understanding complex queries: Multi-query comprehension network for temporal sentence grounding
Zan Gao 0001, Shengbo Xiong, Yibo Zhao 0001, Chunjie Ma, Tian Gan 0002, Riwei Wang |
Expert Syst. Appl. | 3 |
| 2026 | RMPSNet: Occluded person re-identification via regional masking and prompt-distribution synergy
Zan Gao 0001, Shuai Xie, Shengxun Wei, Yibo Zhao 0001, Chunjie Ma, Chen Li 0035 |
Pattern Recognit. | 4 |
| 2026 | Prior-knowledge guidance and dual-domain representation refinement for deepfake detection
Zhiyong Cheng 0001, Tianyi Wang 0006, Chunjie Ma, Yibo Zhao 0001, Zan Gao 0001 |
Pattern Recognit. | 5 |
| 2026 | A Novel Multi-View Perception and Shrinkage Aggregation Network for Inharmonious Region LocalizationabstractWith the popularity of image editing techniques, synthetic images may have inharmonious regions due to color/illumination differences between the manipulated area and the background. The inharmonious region localization task aims to find these regions, which is crucial for blind image harmonization. Existing methods rely on single-view images and do not fully explore multi-scale fusion, which limits their performance. To address these issues, in this paper, we propose a novel multi-view perception and shrinkage aggregation network (MSANet) for the inharmonious region localization task that fully utilizes multi-view images and multi-scale fusion information and can mine subtle cues between candidate objects and the background. Specifically, we first design a multi-view ensemble encoder to fully perceive the inharmonious regions by multi-view interactive learning and then aggregate the feature representations of inharmonious regions. Moreover, we propose a multi-scale shrinkage fusion decoder, where multi-scale features with multi-view prior information are utilized to aggregate adjacent features, adaptively select high-quality information, reduce background interference and gradually locate inharmonious regions. Extensive experimental results on four public datasets (HDobe5K, HCOCO, HFlickr, and Hday2Night) demonstrate that the proposed MSANet can outperform all the SOTA methods in terms of average F1 and average IoU score, while maintaining a lower computational cost1. Shenghao Chen, Chunjie Ma, Yibo Zhao 0001, Meng Liu 0006, Yanbing Xue, Zan Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Learning Generalizable Representations for Deepfake Detection With Realistic Sample Generation and Dual AugmentationabstractDeepfake detection aims to identify manipulated content generated by generative models such as GANs and diffusion models. Although many detection methods have been proposed in recent years, their performance often degrades significantly when the test data includes unknown images or novel forgery types. To address this challenge, we propose a Realistic Sample Generation and Dual Augmentation framework, abbreviated as RSG-DA, to enhance generalization in deepfake detection. The key idea is to explore and expand the forgery feature space in order to learn decision boundaries that can capture diverse forgery patterns. Specifically, we introduce a Dynamic Landmark Diffusion Generator (DLDG) that synthesizes hybrid forgery samples with high visual realism and structural diversity. Additionally, we design a Dual Data Augmentation (DDA) strategy composed of DW-Augmentation and Class-Augmentation, where DW-Augmentation strengthens the representation of authentic image features through multi-scale transformations, while Class-Augmentation enriches the forgery distribution by expanding it with varied manipulations. Finally, we present a Lightweight Generic Forgery Distillation (LGFD) module that integrates the above components into a unified encoder, enabling the learning of robust and transferable forgery representations. Extensive experiments show that our method consistently outperforms state-of-the-art approaches in both intra-dataset and cross-dataset evaluations. Zan Gao 0001, Xinhai Zhu, Yibo Zhao 0001, Chunjie Ma, Chen Li 0035 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | LSGNet: A Local-Pattern Separation and Global-Aware Network for Temporal Action DetectionabstractTemporal Action Detection aims to localize and classify action instances within untrimmed videos, yet it remains challenging due to background clutter, high intra-class similarity, and varied temporal scales in real-world scenarios. To address these issues, we propose the Local-Pattern Separation and Global-Aware Network (LSGNet) tailored for temporal action localization. Specifically, the core of LSGNet is the Local Pattern Separation Module (LPSM), which explicitly models both consistency and variation patterns of action segments within local temporal windows. Additionally, to capture comprehensive contextual information, we introduce the Global Context-Aware Representation Module (GCRM), which decouples temporal features across multiple granularities and enables robust modeling of long-range dependencies. Finally, we design the Multi-scale Feature Refinement Module (MFRM) to mitigate the degradation of fine-grained information by performing iterative reconstruction across temporal scales, thereby enriching semantic representations and preserving temporal details. Extensive experiments on THUMOS14, ActivityNet1.3, HACS, and EPIC-Kitchens-100 demonstrate the effectiveness of the proposed LSGNet method. Additional ablation studies on the QVHighlights dataset further confirm the generalization capability of LPSM module in video moment retrieval and highlight detection, achieving consistent improvements in retrieval accuracy and localization precision. Zan Gao 0001, Weilin Yang, Yibo Zhao 0001, Chunjie Ma, Chen Li 0035, Riwei Wang |
IEEE Trans. Image Process. | 3 |
| 2026 | EFIN: A Novel Enhanced Feature Interaction Network for Temporal Sentence Grounding in VideosabstractTemporal sentence grounding in videos (TSGV) is a challenging task that aims to match text queries with semantically relevant segments in untrimmed videos. However, existing methods face limitations in modeling modality features, which constrains the expressive power of candidate moment features. To address this challenge, we propose a novel Enhanced Feature Interaction Network (EFIN) that effectively captures semantic information within each modality and aligns relationships between modalities. Additionally, EFIN enhances the fusion of information between candidate moments and modality features. Specifically, our model begins by extracting modality features to generate candidate moments as priors. Building upon these modality features, we introduce an enhanced feature encoder to extract semantic information within each modality, thereby improving intra-modality feature representation. Simultaneously, the encoder captures alignment relationships between modalities to optimize cross-modality feature representation, enhancing the overall modeling capacity of modality features. Moreover, we design an information fusion module to enrich the comprehension of modality information for candidate moments. Extensive experiments on four benchmark datasets demonstrate the superiority of our proposed EFIN model. Notably, EFIN achieves a maximum performance improvement of approximately 1.67% and 1.91% across different evaluation metrics on TACoS dataset. Chongxu Hu, Xianbin Wen, Yibo Zhao 0001, Chunjie Ma, Weili Guan, Riwei Wang, Zan Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | A Collaborative Hierarchical Aggregation Network for Weakly Supervised Temporal Action LocalizationabstractTemporal action localization is a fundamental task in video understanding that focuses on classifying and temporally localizing action instances in untrimmed videos. Compared to temporal action localization, the Weakly supervised Temporal Action Localization (WTAL) task presents greater challenges, as its training data lacks detailed information about action boundaries. Existing WTAL methods ignore the complementary relationship between modalities and the dependency between snippets, resulting in inaccurate localization results. To solve these issues, we propose a Collaborative Hierarchical Aggregation Network (CHA-Net). Specifically, we first use a modality complementary module to learn the synergies between modalities. Then, a collaborative enhance module is proposed to remove the information irrelevant to actions in RGB modality. Finally, a hierarchical aggregation module is proposed to capture the complete temporal information of action instances to better mine the temporal dependencies between snippets. Extensive experiments on THUMOS14, ActivityNet1.2, and ActivityNet1.3 datasets demonstrate the effectiveness of our method. Compared with F3-Net (TMM2024, Avg{0.1:0.5}) and SPCC-Net (TMM2024, Avg{0.1:0.7}) on the THUMOS14 dataset, the proposed method can achieve improvements of 3.2% and 2.4%, respectively. Zan Gao 0001, Xiaoyi Xu, Yibo Zhao 0001, Chunjie Ma, Yanbing Xue, Riwei Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | FluidGS: Physics Informed Gaussian Splatting for Dynamic Fluid Reconstruction from Sparse Views
Youchen Xie, Chen Li 0035, Sheng Qiu, Zhi-Jun Wang, Chenhui Li 0001, Yibo Zhao 0001, Zan Gao 0001, Changbo Wang |
ACM Multimedia | 6 |
| 2025 | Lightweight Relational Proposal Network with Dual-Branch Distillation for Video Moment Retrieval
Yujia Zhu, Hao Yang 0067, Yibo Zhao 0001, Chunjie Ma, Weili Guan, Zan Gao 0001 |
ACM Multimedia | 3 |
| 2025 | Fine-Grained Modality Relation-Aware Network for Video Moment RetrievalabstractVideo moment retrieval (VMR) involves localizing video segments semantically aligned with given queries within videos. Despite the development of numerous methods for VMR in recent years, there remains a need to better incorporate fine-grained modality relation-aware information both in intra-modality and cross-modality. To address these challenges, we propose a Fine-grained Modality Relation-Aware Network (FMRN) tailored for the video moment retrieval task. FMRN effectively explores fine-grained modality relation-aware information within text queries, videos, and proposals. Our approach begins with a semantic graph encoder to capture deep semantic relations in intra-modality. Besides, we introduce a novel fine-grained cross-modality interaction module comprising a cross-similarity weighting module, an intra-modality weighting module, and an adaptive fusion module. These components comprehensively exploit fine-grained relation information within intra-modality and cross-modality contexts. Specifically, the cross-similarity weighting module leverages similarities between text queries and video snippets, as well as between videos and query words. The intra-modality weighting module determines the importance of words and snippets, while the adaptive fusion module combines cross-similarity weighting and intra-modality weighting. Additionally, we design a proposal relation module to enhance retrieval by capturing fine-grained proposals-relation information in videos. Extensive experiments demonstrate that the proposed method can outperform all state-of-the-art methods on the TACoS dataset and obtain comparable results on the Charades-STA and ActivityNet-Captions datasets. Compared with MCMN (TCSVT2024) and DPHANet (TMM2024), FMRN can achieve average improvements of 3.61 % and 5.44 % on the TACoS dataset, respectively. Yibo Zhao 0001, Zan Gao 0001, Chunjie Ma, Weili Guan, Riwei Wang, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Multiple Information Prompt Learning for Cloth-Changing Person Re-IdentificationabstractCloth-changing person re-identification is a subject closer to the real world, which focuses on solving the problem of person re-identification after pedestrians change clothes. The primary challenge in this field is to overcome the complex interplay between intra-class and inter-class variations and to identify features that remain unaffected by changes in appearance. Sufficient data collection for model training would significantly aid in addressing this problem. However, it is challenging to gather diverse datasets in practice. Current methods focus on implicitly learning identity information from the original image or introducing additional auxiliary models, which are largely limited by the quality of the image and the performance of the additional model. To address these issues, inspired by prompt learning, we propose a novel multiple information prompt learning (MIPL) scheme for cloth-changing person ReID, which learns identity robust features through the common prompt guidance of multiple messages. Specifically, the clothing information stripping (CIS) module is designed to decouple the clothing information from the original RGB image features to counteract the influence of clothing appearance. The bio-guided attention (BGA) module is proposed to increase the learning intensity of the model for key information. A dual-length hybrid patch (DHP) module is employed to make the features have diverse coverage to minimize the impact of feature bias. Extensive experiments demonstrate that the proposed method outperforms all state-of-the-art methods on the LTCC, CelebreID, Celeb-reID-light, and CSCC datasets, achieving rank-1 scores of 74.8%, 73.3%, 66.0%, and 88.1%, respectively. When compared to AIM (CVPR23), ACID (TIP23), and SCNet (MM23), MIPL achieves rank-1 improvements of 11.3%, 13.8%, and 7.9%, respectively, on the PRCC dataset. Shengxun Wei, Zan Gao 0001, Chunjie Ma, Yibo Zhao 0001, Weili Guan, Shengyong Chen |
IEEE Trans. Image Process. | 4 |
| 2024 | A Snippets Relation and Hard-Snippets Mask Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) is a problem learning an action localization model with only video-level labels available. In recent years, many WTAL methods have developed. However, hard-to-predict snippets near action boundaries are often not considered in these existing approaches, causing action incompleteness and action over-complete issues. To solve these issues, in this work, an end-to-end snippets relation and hard-snippets mask network (SRHN) is proposed. Specifically, a hard-snippets mask module is applied to mask the hard-to-predict snippets adaptively, and in this way, the trained model focuses more on those snippets with low uncertainty. Then, a snippets relation module is designed to capture the relationship among snippets and can make hard-to-predict snippets easy to predict by aggregating the information of multiple temporal receptive fields. Finally, a snippet enhancement loss is further developed to reduce the action probabilities that are not present in videos for hard-to-predict snippets and other snippets, enlarging the action probabilities that exist in videos. Extensive experiments on THUMOS14, ActivityNet1.2, and ActivityNet1.3 datasets demonstrate the effectiveness of the SRHN method. Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Weili Guan, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | A Novel Temporal Channel Enhancement and Contextual Excavation Network for Temporal Action LocalizationabstractThe temporal action localization (TAL) task aims to locate and classify action instances in untrimmed videos. Most previous methods use classifiers and locators to act on the same feature; thus, the classification and localization processes are relatively independent. Therefore, if the classification results and localization results are fused, there will be a problem that the classification results are correct while the localization results are wrong, resulting in inaccurate final results, and vice versa. To solve this problem, we propose a novel temporal channel enhancement and contextual excavation network (TCN) for the TAL task, which generates robust classification and localization features and refines the final localization results. Specifically, a temporal channel enhancement module is designed to enhance the temporal and channel information of the feature sequence. Then, the temporal semantic contextual excavation module is developed to establish relationships between similar frames. Finally, the features with enhanced contextual information are transferred to a classifier. While executing the classification process, we obtain powerful classification features. Most importantly, with the robust classification features, the final localization features are produced by the refine localization module, which is applied to obtain the final localization results. Extensive experiments show that TCN can outperform all the SOTA methods on the THUMOS14 dataset, and achieves a comparable performance on the ActivityNet1.3 dataset. Compared with ActionFormer (ECCV 2022) and BREM (MM 2022) on the THUMOS14 dataset, the proposed TCN can achieve improvements of 1.8% and 5.0%, respectively. Zan Gao 0001, Xinglei Cui, Yibo Zhao 0001, Tao Zhuo, Weili Guan, Meng Wang 0001 |
ACM Multimedia | 3 |
| 2023 | A Novel Action Saliency and Context-Aware Network for Weakly-Supervised Temporal Action LocalizationabstractTemporal action localization is a challenging task in computer vision, and it tries to find the start time and the end time of the actions and predict their categories. However, compared to temporal action localization, weakly supervised temporal action localization (WTAL) is a more challenging task due to its poor annotations. With only video-level annotation, some background frames, similar to actions, would be classified as actions and produce inaccurate results. In addition, the two-stream fusion problem, ignored previously, also needs to be further considered. To resolve these issues, we propose a novel action saliency and context-aware network (ASCN) for weakly supervised temporal action localization tasks. Specifically, the temporal saliency and context module is designed to enhance the global saliency and context information of the RGB and the flow features to suppress the backgrounds and enhance the actions. In addition, a hybrid attention mechanism using frame differences and two-stream attention is designed to model the local action context information and further enlarge the scores of the potential action regions and suppress the background regions. Finally, to obtain two-stream consistency and solve the fusion problem, we use the similarity loss and a channel self-attention module to adaptively fuse the enhanced RGB and flow features. Extensive experiments demonstrate that ASCN can outperform all of the SOTA WTAL methods on the THUMOS14 dataset and the ActivityNet1.3 dataset with an average mAP that can reach 37.2% on the THUMOS14 dataset and attains an average mAP of 26.3% on the ActivityNet1.3 dataset. On the ActivityNet1.2 dataset, ASCN can also obtain comparable results. Compared with AdapNet (TNNLS20), MMSD (TIP22), and FTCL (CVPR22) on the THUMOS14 dataset, ASCN can outperform them by 13.5%, 2.9%, and 2.8%, respectively. Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Wen Gao 0001, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Multim. | 1 |
| 2022 | A Novel Multiple-View Adversarial Learning Network for Unsupervised Domain Adaptation Action RecognitionabstractAbstract-domain adaptation action recognition is a hot research topic in machine learning and some effective approaches have been proposed. However, samples in the target domain with label information are often required by these approaches. Moreover, domain-invariant discriminative feature learning, feature fusion, and classifier module learning have not been explored in an end-to-end framework. Thus, in this study, we propose a novel end-to-end multiple-view adversarial learning network (MAN) for unsupervised domain adaptation action recognition in which the fusion of RGB and optical-flow features, domain-invariant discrimination feature learning, and action recognition is conducted in a unified framework. Specifically, a robust spatiotemporal feature extraction network, including a spatial transform network and an adaptive intrachannel weight network, is proposed to improve the scale invariance and robustness of the method. Then, a self-attention mechanism fusion module is designed to adaptively fuse the RGB and optical-flow features. Moreover, a multiview adversarial learning loss is developed to obtain domain-invariant discriminative features. In addition, three benchmark datasets are constructed for unsupervised domain adaptation action recognition, for which all actions and samples are carefully collected from public action datasets, and their action categories are hierarchically augmented, which can guide how to extend existing action datasets. We conduct extensive experiments on four benchmark datasets, and the experimental results demonstrate that our proposed MAN can outperform several state-of-the-art unsupervised domain adaptation action recognition approaches. When the SDAI Action II-6 and SDAI Action II-11 datasets are used, MAN can achieve 3.7% ( H → U ) and 6.1% ( H → U ) improvements over the temporal attentive adversarial adaptation network (published in ICCV 2019) module, respectively. As an added contribution, the SDAI Action II-6, SDAI Action II-11, and SDAI Action II-16 datasets will be released to facilitate future research on domain adaptation action recognition. Zan Gao 0001, Yibo Zhao 0001, Hua Zhang 0003, Da Chen 0002, Anan Liu, Shengyong Chen |
IEEE Trans. Cybern. | 2 |
| 2022 | A Temporal-Aware Relation and Attention Network for Temporal Action LocalizationabstractTemporal action localization is currently an active research topic in computer vision and machine learning due to its usage in smart surveillance. It is a challenging problem since the categories of the actions must be classified in untrimmed videos and the start and end of the actions need to be accurately found. Although many temporal action localization methods have been proposed, they require substantial amounts of computational resources for the training and inference processes. To solve these issues, in this work, a novel temporal-aware relation and attention network (abbreviated as TRA) is proposed for the temporal action localization task. TRA has an anchor-free and end-to-end architecture that fully uses temporal-aware information. Specifically, a temporal self-attention module is first designed to determine the relationship between different temporal positions, and more weight is given to features within the actions. Then, a multiple temporal aggregation module is constructed to aggregate the temporal domain information. Finally, a graph relation module is designed to obtain the aggregated graph features, which are used to refine the boundaries and classification results. Most importantly, these three modules are jointly explored in a unified framework, and temporal awareness is always fully used. Extensive experiments demonstrate that the proposed method can outperform all state-of-the-art methods on the THUMOS14 dataset with an average mAP that reaches 67.6% and obtain a comparable result on the ActivityNet1.3 dataset with an average mAP that reaches 34.4%. Compared with A2Net (TIP20), PCG-TAL (TIP21), and AFSD (CVPR21) TRA can achieve improvements of 11.7%, 4.4%, and 1.8%, respectively on the THUMOS14 dataset. Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Weili Guan, Jie Nie, Anan Liu, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Image Process. | 1 |