VLDB 2026 Research / reviewers in the wild / expert
Zhaofeng Shi
dblp:298/7613
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-6313-8670ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causality-inspired Federated Learning for Dynamic Spatio-Temporal GraphsabstractFederated Graph Learning (FGL) has emerged as a powerful paradigm for decentralized training of graph neural networks while preserving data privacy. However, existing FGL methods are predominantly designed for static graphs and rely on parameter averaging or distribution alignment, which implicitly assume that all features are equally transferable across clients, overlooking both the spatial and temporal heterogeneity and the presence of client-specific knowledge in real-world graphs. In this work, we identify that such assumptions create a vicious cycle of spurious representation entanglement, client-specific interference, and negative transfer, degrading generalization performance in Federated Learning over Dynamic Spatio-Temporal Graphs (FSTG). To address this issue, we propose a novel causality-inspired framework named SC-FSGL, which explicitly decouples transferable causal knowledge from client-specific noise through representation-level interventions. Specifically, we introduce a Conditional Separation Module that simulates soft interventions through client conditioned masks, enabling the disentanglement of invariant spatio-temporal causal factors from spurious signals and mitigating representation entanglement caused by client heterogeneity. In addition, we propose a Causal Codebook that clusters causal prototypes and aligns local representations via contrastive learning, promoting cross-client consistency and facilitating knowledge sharing across diverse spatio-temporal patterns. Experiments on five diverse heterogeneity Spatio-Temporal Graph (STG) datasets show that SC-FSGL outperforms state-of-the-art methods. Yuxuan Liu 0017, Wenchao Xu 0001, Haozhao Wang, Zhiming He, Zhaofeng Shi, Chongyang Xu, Peichao Wang |
AAAI | 5 |
| 2026 | Ego-PMOVE: Prompt-aware Mixture of View Experts Network for Egocentric Gaze PredictionabstractEgocentric gaze prediction serves as a critical indicator for decoding human visual attention and cognitive processes, but its inherently limited field of view creates prediction challenges. Although exo-view data provides supplementary contextual information, it exhibits significant spatial and semantic gaps. Existing methods focus solely on isolated feature encoding in single-view paradigms, neglecting cross-view gaze correlations. To make up for this gap, we make the first exploration of cross-view gaze relationship for egocentric gaze prediction, and propose Ego-PMOVE, a novel Prompt-aware Mixture of View Experts network. Unlike prior cross-view studies that forcibly align cross-view features thereby introducing inference noise, we leverage the popular Mixture-of-Experts (MoE) and a set of flexible prompts to disentangle features from different views into three parallel experts: a view-shared expert directly modeling common semantic relationships, a view-discrepancy expert adaptively adjusting the spatial position, scale and shifts based on different view-specific features, and an egocentric expert extracting independent features to compensate for the case of missing exocentric data. To balance these experts, we further design a soft router to dynamically weight them for mining useful information while suppressing noise. A view-query gaze decoder then generates view-specific gaze attention maps, jointly optimized by gaze-heamap and cross-view contrastive loss that regularize both shared and divergent features for accurate gaze prediction. Extensive experiments across the multi-view EgoMe dataset and single-view Ego4D and EGTEA Gaze++ datasets demonstrate the effectiveness and generalizability of our approach. Heqian Qiu, Lanxiao Wang, Taijin Zhao, Zhaofeng Shi, Linfeng Xu 0001, Hongliang Li 0001 |
AAAI | 4 |
| 2025 | Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus AdaptationabstractEven from an early age, humans naturally adapt between exocentric (Exo) and egocentric (Ego) perspectives to understand daily procedural activities. Inspired by this cognitive ability, we propose a novel Unsupervised Ego-Exo Dense Procedural Activity Captioning (UE^2 DPAC) task, which aims to transfer knowledge from the labeled source view to predict the time segments and descriptions of action sequences for the target view without annotations. Despite previous works endeavoring to address the fully-supervised single-view or cross-view dense video captioning, they lapse in the proposed task due to the significant inter-view gap caused by temporal misalignment and irrelevant object interference. Hence, we propose a Gaze Consensus-guided Ego-Exo Adaptation Network (GCEAN) that injects the gaze information into the learned representations for the fine-grained Ego-Exo alignment. Specifically, we propose a Score-based Adversarial Learning Module (SALM) that incorporates a discriminative scoring network and compares the scores of distinct views to learn unified view-invariant representations from a global level. Then, the Gaze Consensus Construction Module (GCCM) utilizes the gaze to progressively calibrate the learned representations to highlight the regions of interest and extract the corresponding temporal contexts. Moreover, we adopt hierarchical gaze-guided consistency losses to construct gaze consensus for the explicit temporal and spatial adaptation between the source and target views. To support our research, we propose a new EgoMe-UE^2 DPAC benchmark, and extensive experiments demonstrate the effectiveness of our method, which outperforms many related methods by a large margin. Code is available at https://github.com/ZhaofengSHI/GCEAN. Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
ACM Multimedia | 1 |
| 2025 | D3Net: Dual-Path Decoupling-Distillation for Adaptive Fusion in Continual Egocentric LearningabstractEgocentric continual action recognition faces severe challenges such as sudden viewpoint changes, occlusions, and complex backgrounds. In such scenarios, relying solely on visual modalities is susceptible to interference and lacks sufficient recognition robustness. To overcome the limitations of unimodal approaches, multimodal fusion methods are widely adopted, significantly enhancing recognition performance. However, existing multimodal schemes generally suffer from insufficient exploration of cross-modal complementarity and the vulnerability of modal independence. To address this, this paper proposes a Dual-path Decoupling-Distillation NetWork (D3Net), aiming to achieve more effective dynamic fusion of modal information and knowledge transfer.D3Net first explicitly separates the shared and private features of modalities through a dual-path decoupling module, combined with a dynamic gating mechanism to adaptively adjust the modal fusion weights. Secondly, it designs a complementary distillation module, leveraging cross-modal contrastive learning to effectively mitigate the issues of poor unimodal robustness and vulnerability to interference. Finally, through a cross-task distillation mechanism, it efficiently extracts knowledge from old tasks, alleviating the catastrophic forgetting problem during learning. Experimental results demonstrate that D3Net achieves an average accuracy of 83.97% under the 8×4 task configuration on the UESTC MMEA CL dataset, surpassing baseline method by 5.17%. Chenghao Qi, Heqian Qiu, Zhaofeng Shi, Lanxiao Wang, Hongliang Li 0001 |
MMSP | 3 |
| 2025 | Cognition Transferring and Decoupling for Text-Supervised Egocentric Semantic SegmentationabstractIn this paper, we explore a novel Text-supervised Egocentic Semantic Segmentation (TESS) task that aims to assign pixel-level categories to egocentric images weakly supervised by texts from image-level labels. In this task with prospective potential, the egocentric scenes contain dense wearer-object relations and inter-object interference. However, most recent third-view methods leverage the frozen Contrastive Language-Image Pre-training (CLIP) model, which is pre-trained on the semantic-oriented third-view data and lapses in the egocentric view due to the “relation insensitive” problem. Hence, we propose a Cognition Transferring and Decoupling Network (CTDN) that first learns the egocentric wearer-object relations via correlating the image and text. Besides, a Cognition Transferring Module (CTM) is developed to distill the cognitive knowledge from the large-scale pre-trained model to our model for recognizing egocentric objects with various semantics. Based on the transferred cognition, the Foreground-background Decoupling Module (FDM) disentangles the visual representations to explicitly discriminate the foreground and background regions to mitigate false activation areas caused by foreground-background interferential objects during egocentric relation learning. Extensive experiments on four TESS benchmarks demonstrate the effectiveness of our approach, which outperforms many recent related methods by a large margin. Code will be available athttps://github.com/ZhaofengSHI/CTDN. Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Cross-Modal Cognitive Consensus Guided Audio-Visual SegmentationabstractAudio-Visual Segmentation (AVS) aims to extract the sounding object from a video frame, which is represented by a pixel-wise segmentation mask for application scenarios such as multi-modal video editing, augmented reality, and intelligent robot systems. The pioneering work conducts this task through dense feature-level audio-visual interaction, which ignores the dimension gap between different modalities. More specifically, the audio clip could only provide aGlobalsemantic label in each sequence, but the video frame covers multiple semantic objects across differentLocalregions, which leads to mislocalization of the representationally similar but semantically different object. In this paper, we propose a Cross-modal Cognitive Consensus guided Network (C3N) to align the audio-visual semantics from the global dimension and progressively inject them into the local regions via an attention mechanism. Firstly, a Cross-modal Cognitive Consensus Inference Module (C3IM) is developed to extract a unified-modal label by integrating audio/visual classification confidence and similarities of modality-agnostic label embeddings. Then, we feed the unified-modal label back to the visual backbone as the explicit semantic-level guidance via a Cognitive Consensus guided Attention Module (CCAM), which highlights the local features corresponding to the interested object. Extensive experiments on the Single Sound Source Segmentation (S4) setting and Multiple Sound Source Segmentation (MS3) setting of the AVSBench dataset demonstrate the effectiveness of the proposed method, which achieves state-of-the-art performance. Zhaofeng Shi, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, Hongliang Li 0001 |
IEEE Trans. Multim. | 1 |