Zijie Song

dblp:310/3748 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Fine-grained Text-Video Retrieval with Patch-level Temporal Difference and Aggregation
abstract
Existing Text-Video Retrieval (TVR) methods predominantly rely on global frame representations, often disregarding the fine-grained temporal variations required for precise patch-level alignment. This is critical as video motion is inherently spatially localized; consequently, coarse frame-level modeling tends to be dominated by static backgrounds, overshadowing salient action cues. To address this limitation, we propose TRFG, a novel framework for text-video retrieval that addresses the challenges of modeling Temporal Reasoning and Fine-Grained cross-modal alignment. First, our Temporal Difference module captures frame-to-frame variations at the patch level, effectively suppressing static background noise to highlight "active" motion regions. Second, these differential signals are synthesized via a Temporal Aggregation module to form a coherent representation of the event’s trajectory. Finally, to ensure precise semantic matching, a fine-grained interaction module aligns these dynamic video tokens with textual details. Extensive experiments on MSRVTT, ActivityNet, and DiDeMo demonstrate that TRFG achieves state-of-the-art performance across multiple backbones and retrieval tasks. Ablation studies confirm the complementarity and generalizability of both modules, underscoring the importance of explicit temporal modeling and fine-grained interaction in bridging the modality gap.
Jialong Hu, Zijie Song, Yang Wang 0023, Zhenzhen Hu 0004, Jia Li 0013, Richang Hong
ICMR2
2025 Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
abstract
Video Question Answering (VideoQA) is a complex video-language task that demands a sophisticated understanding of both visual content and temporal dynamics. Traditional Transformer-style architectures, while effective in integrating multimodal data, often simplify temporal dynamics through positional encoding and fail to capture non-linear interactions within video sequences. In this paper, we introduce the Temporal Trio Transformer (T3T), a novel architecture that models time consistency and time variability. The T3T integrates three key components: Temporal Smoothing (TS), Temporal Difference (TD), and Temporal Fusion (TF). The TS module employs Brownian Bridge for capturing smooth, continuous temporal transitions, while the TD module identifies and encodes significant temporal variations and abrupt changes within the video content. Subsequently, the TF module synthesizes these temporal features with textual cues, facilitating a deeper contextual understanding and response accuracy. The efficacy of the T3T is demonstrated through extensive testing on multiple VideoQA benchmark datasets. Our results underscore the importance of a nuanced approach to temporal modeling in improving the accuracy and depth of video-based question answering.
Zijie Song, Zhenzhen Hu 0004, Jia Li 0013, Richang Hong
ICME1
2025 Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
abstract
Metaphorical imagination, the ability to connect seemingly unrelated concepts, is fundamental to human cognition and communication. While understanding linguistic metaphors has advanced significantly, grasping multimodal metaphors, such as those found in internet memes, presents unique challenges due to their unconventional expressions and implied meanings. Existing methods for multimodal metaphor identification often struggle to bridge the gap between literal and figurative interpretations. Additionally, generative approaches that utilize large language models or text-to-image models, while promising, suffer from high computational costs. This paper introduces Concept Drift Guided LayerNorm Tuning (CDGLT), a novel and training-efficient framework for multimodal metaphor identification. CDGLT incorporates two key innovations: (1) Concept Drift, a mechanism that leverages Spherical Linear Interpolation (SLERP) of cross-modal embeddings from a CLIP encoder to generate a new, divergent concept embedding. This drifted concept helps to alleviate the gap between literal features and the figurative task. (2) A prompt construction strategy, that adapts the method of feature extraction and fusion using pre-trained language models for the multimodal metaphor identification task. CDGLT achieves state-of-the-art performance on the MET-Meme benchmark while significantly reducing training costs compared to existing generative methods. Ablation studies demonstrate the effectiveness of both Concept Drift and our adapted LN Tuning approach. Our method represents a significant step towards efficient and accurate multimodal metaphor understanding. The code is available: https://github.com/Qianvenh/CDGLT.
Wenhao Qian, Zhenzhen Hu 0004, Zijie Song, Jia Li 0013
ICMR3
2025 Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
abstract
Recent progress in text–video retrieval has been largely driven by contrastive learning. However, existing methods often overlook the effect of the modality gap, which causes anchor representations to undergo in-place optimization (i.e., optimization tension) that limits their alignment capacity. Moreover, noisy hard negatives further distort the semantics of anchors. To address these issues, we propose GARE, a Gap-Aware Retrieval framework that introduces a learnable, pair-specific increment $\Delta_{ij}$ between text $t_i$ and video $v_j$, redistributing gradients to relieve optimization tension and absorb noise. We derive $\Delta_{ij}$ via a multivariate first-order Taylor expansion of the InfoNCE loss under a trust-region constraint, showing that it guides updates along locally consistent descent directions. A lightweight neural module conditioned on the semantic gap couples increments across batches for structure-aware correction. Furthermore, we regularize $\Delta$ through a variational information bottleneck with relaxed compression, enhancing stability and semantic consistency. Experiments on four benchmarks demonstrate that GARE consistently improves alignment accuracy and robustness, validating the effectiveness of gap-aware tension mitigation.
Zijie Song, Jialong Hu, Zhenzhen Hu 0004, Jia Li 0013, Richang Hong
NeurIPS2
2025 Grid Jigsaw Representation with CLIP: a new perspective on image clustering
Zijie Song, Zhenzhen Hu 0004, Richang Hong
Multim. Syst.1
2025 Adaptive Dual Video Summarization: From Dynamic Keyframes to Captions
abstract
Video summarization and captioning condense content by selecting keyframes and generating language descriptions, integrating both visual and textual perspectives. Existing video-and-language learning models typically select multiple frames as proxies rather than analyzing all frames, which improves computational efficiency but may not adequately represent the original content without redundancy. In this paper, we propose an adaptive dual video summarization framework and demonstrate its effectiveness within the context of video captioning. Given the video frames, we extract visual representations using a video-domain fine-tuned ViT model to narrow the domain shift. The keyframes are summarized based on the frame-level scores. To minimize the number of keyframes while ensuring captioning quality, we introduce a cross-modal video summarizer that selects the most semantically consistent frames according to pseudo score labels. Furthermore, we incorporate an adaptive keyframe selector that determines the optimal number of keyframes based on the video's complexity and content, enhancing the framework's adaptability and generalization. The proposed adaptive keyframe selector enables the framework to handle diverse video content, making it more generalizable and applicable to real-world scenarios.We designed a ranking scheme to assess the video's static appearance and temporal dynamics from score-based and time-based perspectives. To conclude, we use a lightweight LSTM decoder to generate descriptions. Experimental results on the MSR-VTT, MSVD and VATEX benchmarks demonstrate that our adaptive dual video summarization framework can effectively convey the same semantic information as the original video while using a significantly reduced number of keyframes, leading to improved video captioning performance.
Zhenzhen Hu 0004, Zhenshan Wang, Jia Li 0013, Zijie Song, Richang Hong, Meng Wang 0001
IEEE Trans. Multim.5
2024 Dual-Stream Keyframe Enhancement for Video Question Answering
abstract
The redundancy in videos and the quadratic scaling with input length of Transformer models lead to the need for sampling and selection from input videos.During the selection process, differentiable Top-K algorithms are employed to ensure an end-to-end training process.However, these methods not only restrict the level at which temporal information is captured but also introduce sorting noise and inaccuracies.In this paper, we revisit the keyframe selection strategy for VideoQA and propose a novel framework named Dual-Stream Keyframe Enhancement (DSKE) incorporating the enhancement of temporal granularity.To balance end-to-end sorting and hard ranking, we employ a dual-stream keyframe selection strategy by fusing the differentiable and non-differentiable results together to achieve a unified approach.One stream is based on the approximate ranking obtained from the differentiable Top-K algorithm, while the other stream utilizes the results obtained from hard ranking.We separately train decoders on the outputs of each stream and then combine the decoder results to predict the final answer.By integrating both stream results, DSKE effectively balances the inclusion of relevant information while filtering out noise.Additionally, we capture temporal variation information by incorporating a series of overlapping sliding time windows to enrich the temporal granularity.To evaluate the effectiveness of DSKE, we conduct experiments on the NExT-QA and AGQA benchmarks.The results demonstrate that our framework significantly improves the performance of VideoQA by effectively incorporating temporal components and enhancing the keyframe ranking process.
Zhenzhen Hu 0004, Jia Li 0013, Zijie Song, Richang Hong
MMAsia4
2024 Exploring coherence from heterogeneous representations for OCR image captioning
Zijie Song
Multim. Syst.2
2024 Embedded Heterogeneous Attention Transformer for Cross-Lingual Image Captioning
abstract
Cross-lingual image captioning is a challenging task that requires addressing both cross-lingual and cross-modal obstacles in multimedia analysis. The crucial issue in this task is to model the global and the local matching between the image and different languages. Existing cross-modal embedding methods based on the transformer architecture oversee the local matching between the image region and monolingual words, especially when dealing with diverse languages. To overcome these limitations, we propose an Embedded Heterogeneous Attention Transformer (EHAT) to establish cross-domain relationships and local correspondences between images and different languages by using a heterogeneous network. EHAT comprises Masked Heterogeneous Cross-attention (MHCA), Heterogeneous Attention Reasoning Network (HARN), and Heterogeneous Co-attention (HCA). The HARN serves as the core network and it captures cross-domain relationships by leveraging visual bounding box representation features to connect word features from two languages and to learn heterogeneous maps. MHCA and HCA facilitate cross-domain integration in the encoder through specialized heterogeneous attention mechanisms, enabling a single model to generate captions in two languages. We evaluate our approach on the MSCOCO dataset to generate captions in English and Chinese, two languages that exhibit significant differences in their language families. The experimental results demonstrate the superior performance of our method compared to existing advanced monolingual methods. Our proposed EHAT framework effectively addresses the challenges of cross-lingual image captioning, paving the way for improved multilingual image analysis and understanding.
Zijie Song, Zhenzhen Hu 0004, Yuanen Zhou, Ye Zhao 0001, Richang Hong, Meng Wang 0001
IEEE Trans. Multim.1
2023 CDR: Conservative Doubly Robust Learning for Debiased Recommendation
abstract
In recommendation systems (RS), user behavior data is observational rather than experimental, resulting in widespread bias in the data. Consequently, tackling bias has emerged as a major challenge in the field of recommendation systems. Recently, Doubly Robust Learning (DR) has gained significant attention due to its remarkable performance and robust properties. However, our experimental findings indicate that existing DR methods are severely impacted by the presence of so-called Poisonous Imputation, where the imputation significantly deviates from the truth and becomes counterproductive.
Zijie Song, Jiawei Chen 0007, Sheng Zhou 0004, Qihao Shi, Chun Chen 0001, Can Wang 0001
CIKM1
2023 Dual Video Summarization: From Frames to Captions
abstract
Video summarization and video captioning both condense the video content from the perspective of visual and text modes, i.e. the keyframe selection and language description generation. Existing video-and-language learning models commonly sample multiple frames for training instead of observing all. These sampled deputies greatly improve computational efficiency, but do they represent the original video content enough with no more redundancy? In this work, we propose a dual video summarization framework and verify it in the context of video captioning. Given the video frames, we firstly extract the visual representation based on the ViT model fine-tuned on the video-text domain. Then we summarize the keyframes according to the frame-lever score. To compress the number of keyframes as much as possible while ensuring the quality of captioning, we learn a cross-modal video summarizer to select the most semantically consistent frames according to the pseudo score label. Top K frames ( K is no more than 3% of the entire video.) are chosen to form the video representation. Moreover, to evaluate the static appearance and temporal information of video, we design the ranking scheme of video representation from two aspects: feature-oriented and sequence-oriented. Finally, we generate the descriptions with a lightweight LSTM decoder. The experiment results on the MSR-VTT and MSVD dataset reveal that, for the generative task as video captioning, a small number of keyframes can convey the same semantic information to perform well on captioning, or even better than the original sampling.
Zhenzhen Hu 0004, Zhenshan Wang, Zijie Song, Richang Hong
IJCAI3
2023 Grid Feature Jigsaw for Self-supervised Image Clustering
abstract
Image clustering is an essential unsupervised learning task in computer vision. The key issue for image clustering is to learn the representative visual features without annotations to extend class spacing. Jigsaw puzzles, as one pretext task of self-supervised visual representation learning to learn the relative spatial position of image tiles, has attracted the attention of many researchers. Most of existing jigsaw puzzle solving strategies are based on the raw image patches, i.e. original pixels, which makes them only concentrate on the low-level statistics. In this paper, we propose a novel learning strategy named Grid Feature Jigsaw (GFJ) for self-supervised image clustering to increase class spacing by deep mining of single sample feature. We train the model to learn the intra-grid representation via the self-supervised paradigm. By dividing the feature map into grids and arranging adjacent grids in each block, we implement a linear regression from the surrounding grids to represent the reference grid. The experiments of unsupervised computer vision benchmark show the effectiveness on the clustering task with respect to the ACC, NMI and ARI three metrics and we verify GFJ universal performance via various deep convolutional neural networks.
Zijie Song, Zhenzhen Hu 0004, Richang Hong
IJCNN1
2023 Efficient and self-adaptive rationale knowledge base for visual commonsense reasoning
Zijie Song, Zhenzhen Hu 0004, Richang Hong
Multim. Syst.1
2022 OCR-oriented Master Object for Text Image Captioning
abstract
Text image captioning aims to understand the scene text in images for image caption generation. The key issue of this challenging task is to understand the relationship between the text OCR tokens and images. In this paper, we propose a novel text image captioning method by purifying the OCR-oriented scene graph with themaster object. The master object is the object to which the OCR is attached, which is the semantic relationship bridge between the OCR token and the image. We consider the master object as a proxy to connect OCR tokens and other regions in the image. By exploring the master object for each OCR token, we build the purified scene graph based on the master objects and then enrich the visual embedding by the Graph Convolution Network (GCN). Furthermore, we cluster the OCR tokens and feed the hierarchical information to provide a richer representation. Experiments on the TextCaps validation and test dataset demonstrate the effectiveness of the proposed method.
Wenliang Tang, Zhenzhen Hu 0004, Zijie Song, Richang Hong
ICMR3