Leqi Shen

dblp:323/9555 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0002-7742-9142ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Follow-Your-Click: Open-domain Regional Image Animation via Motion Prompts
abstract
Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and tend to move the entire scene. However, human artists may need to control the movement of different objects or regions. Additionally, current I2V methods require users not only to describe the target motion but also to provide redundant detailed descriptions of frame contents.These two issues hinder the practical utilization of current I2V tools. In this paper, we propose a practical framework, named Follow-Your-Click, to achieve image animation with a simple user click (for specifying what to move) and a motion prompt (for specifying how to move). Technically, we propose the first-frame masking strategy, which significantly improves the video generation quality, and a motion-augmented module equipped with a motion prompt dataset to improve the motion prompt following abilities of our model. To further control the motion speed, we propose flow-based motion magnitude control to control the speed of target movement more precisely. Extensive experiments compared with 7 baselines, including both commercial tools and research methods on 8 metrics, suggest the superiority of our approach.
Yue Ma 0016, Yingqing He, Hongfa Wang, Andong Wang, Leqi Shen, Jixuan Ying, Chengfei Cai, Zhifeng Li 0001, Harry Shum, Wei Liu 0005, Qifeng Chen 0001
AAAI5
2025 Beyond Logits: Aligning Feature Dynamics for Effective Knowledge Distillation
abstract
Guoqiang Gong, Jiaxing Wang, Jin Xu, Deping Xiang, Zicheng Zhang, Leqi Shen, Yifeng Zhang, JunhuaShu JunhuaShu, ZhaolongXing ZhaolongXing, Zhen Chen, Pengzhang Liu, Ke Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Guoqiang Gong, Deping Xiang, Leqi Shen, JunhuaShu JunhuaShu, ZhaolongXing ZhaolongXing, Zhen Chen 0046, Pengzhang Liu
ACL (1)6
2025 DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval
abstract
The parameter-efficient adaptation of the image-text pre-training model CLIP for video-text retrieval is a prominent area of research. While CLIP is focused on image-level vision-language matching, video-text retrieval demands comprehensive understanding at the video level. Three key discrepancies emerge in the transfer from image-level to video-level: vision, language, and alignment. However, existing methods mainly focus on vision while neglecting language and alignment. In this paper, we propose Discrepancy Reduction in Vision, Language, and Alignment (DiscoVLA), which simultaneously mitigates all three discrepancies. Specifically, we introduce Image-Video Features Fusion to integrate image-level and video-level features, effectively tackling both vision and language discrepancies. Additionally, we generate pseudo image captions to learn fine-grained image-level alignment. To mitigate alignment discrepancies, we propose Image-To-Video Alignment Distillation, which leverages image-level alignment knowledge to enhance video-level alignment. Extensive experiments demonstrate the superiority of our DiscoVLA. In particular, on MSRVTT with CLIP (ViT-B/16), DiscoVLA outperforms previous methods by 2.2% R@1 and 7.5% R@sum. The code is available at https://github.com/LunarShen/DsicoVLA.
Leqi Shen, Guoqiang Gong, Tianxiang Hao 0001, Pengzhang Liu, Sicheng Zhao, Jungong Han, Guiguang Ding
CVPR1
2025 TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
abstract
Most text-video retrieval methods utilize the text-image pre-trained models like CLIP as a backbone. These methods process each sampled frame independently by the image encoder, resulting in high computational overhead and limiting practical deployment. Addressing this, we focus on efficient text-video retrieval by tackling two key challenges: 1. From the perspective of trainable parameters, current parameter-efficient fine-tuning methods incur high inference costs; 2. From the perspective of model complexity, current token compression methods are mainly designed for images to reduce spatial redundancy but overlook temporal redundancy in consecutive frames of a video. To tackle these challenges, we propose Temporal Token Merging (TempMe), a parameter-efficient and training-inference efficient text-video retrieval architecture that minimizes trainable parameters and model complexity. Specifically, we introduce a progressive multi-granularity framework. By gradually combining neighboring clips, we reduce spatio-temporal redundancy and enhance temporal modeling across different frames, leading to improved efficiency and performance. Extensive experiments validate the superiority of our TempMe. Compared to previous parameter-efficient text-video retrieval methods, TempMe achieves superior performance with just 0.50M trainable parameters. It significantly reduces output tokens by 95% and GFLOPs by 51%, while achieving a 1.8X speedup and a 4.4% R-Sum improvement. With full fine-tuning, TempMe achieves a significant 7.9% R-Sum improvement, trains 1.57X faster, and utilizes 75.2% GPU memory usage. The code is available at https://github.com/LunarShen/TempMe.
Leqi Shen, Tianxiang Hao 0001, Sicheng Zhao, Pengzhang Liu, Yongjun Bao, Guiguang Ding
ICLR1
2025 FastVID: Dynamic Density Pruning for Fast Video Large Language Models
abstract
Video Large Language Models have demonstrated strong video understanding capabilities, yet their practical deployment is hindered by substantial inference costs caused by redundant video tokens. Existing pruning techniques fail to effectively exploit the spatiotemporal redundancy present in video data. To bridge this gap, we perform a systematic analysis of video redundancy from two perspectives: temporal context and visual context. Leveraging these insights, we propose Dynamic Density Pruning for Fast Video LLMs termed FastVID. Specifically, FastVID dynamically partitions videos into temporally ordered segments to preserve temporal structure and applies a density-based token pruning strategy to maintain essential spatial and temporal information. Our method significantly reduces computational overhead while maintaining temporal and visual integrity. Extensive evaluations show that FastVID achieves state-of-the-art performance across various short- and long-video benchmarks on leading Video LLMs, including LLaVA-OneVision, LLaVA-Video, Qwen2-VL, and Qwen2.5-VL. Notably, on LLaVA-OneVision-7B, FastVID effectively prunes $\textbf{90.3\%}$ of video tokens, reduces FLOPs to $\textbf{8.3\%}$, and accelerates the LLM prefill stage by $\textbf{7.1}\times$, while maintaining $\textbf{98.0\%}$ of the original accuracy. The code is available at https://github.com/LunarShen/FastVID.
Leqi Shen, Guoqiang Gong, Pengzhang Liu, Sicheng Zhao, Guiguang Ding
NeurIPS1
2025 Temporal Modeling With Frozen Vision-Language Foundation Models for Parameter-Efficient Text-Video Retrieval
abstract
Temporal modeling plays an important role in the effective adaption of the powerful pretrained text-image foundation model into text-video retrieval. However, existing methods often rely on additional heavy trainable modules, such as transformer or BiLSTM, which are inefficient. In contrast, we avoid introducing such heavy components by leveraging frozen foundation models. To this end, we propose temporal modeling with frozen vision-language foundation models (TFVL) to model the temporal dynamics with fixed encoders. Specifically, text encoder temporal modeling (TextTemp) and image encoder temporal modeling (ImageTemp) apply frozen text and image encoders within the video head and video backbone, respectively. TextTemp uses a frozen text encoder to interpret frame representations as "visual words" within a temporal "sentence," capturing temporal dependencies. On the other hand, ImageTemp uses a frozen image encoder to treat all frame tokens as a unified visual entity, learning spatiotemporal information. The total trainable parameters of our method, comprising a lightweight projection and several prompt tokens, are significantly fewer than those in other existing methods. We evaluate the effectiveness of our method on MSR-VTT, DiDeMo, ActivityNet, and LSMDC. Compared with full fine-tuning on MSR-VTT, our TFVL achieves an average 3.25% gain in R@1 with merely 0.35% of the parameters. Extensive experiments demonstrate that the proposed TFVL outperforms state-of-the-art methods with significantly fewer parameters.
Leqi Shen, Tianxiang Hao 0001, Pengzhang Liu, Sicheng Zhao, Jungong Han, Guiguang Ding
IEEE Trans. Neural Networks Learn. Syst.1
2025 Spatio-Temporal Attention for Text-Video Retrieval
abstract
Text-video retrieval, a fundamental task for associating textual descriptions with video content, has become increasingly important in the video domain. Most existing methods focus on the single-modality features only considering the knowledge within individual video or text modalities, often neglecting cross-modal interactions. However, a text description corresponds to a specific spatio-temporal content within a video, involving a certain segment of a frame sequence and distinct sub-regions within these frames. Therefore, we focus on the text-conditioned video features to bridge the modality gap. In this article, we propose Spatio-Temporal Attention for video-text retrieval, termed STAttn, which utilizes textual information to focus on the spatio-temporal video content. Our final text-conditioned video features are generated from the text-related video frames and the text-related regions within these frames. First, we propose the Spatial Text-Attention Module (STAM) to learn the spatial information within video frames. STAM introduces the text-related salient patches to capture more fine-grained details. Second, we propose the Temporal Text-Attention Module (TTAM) to learn the temporal relationships between video frames. Temporal Triplet loss is proposed in TTAM to enhance the attention toward the text-related frames. Thus, the two modules learn the text-related spatio-temporal content from both intra-frame and inter-frame aspects. Extensive experiments on three benchmark datasets, MSRVTT, ActivityNet, and DiDeMo, demonstrate that our STAttn outperforms the state-of-the-art methods.
Leqi Shen, Sicheng Zhao, Pengzhang Liu, Yongjun Bao, Guiguang Ding
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Balanced Active Sampling for Person Re-identification
abstract
Active learning is attracting more and more attention in person re-identification (Re-ID), as it is promising in the scalability of Re-ID models to satisfy performance with reduced labeling cost. Active sampling of pair-wise images in Re-ID is a highly imbalanced problem, where negative pairs are the vast majority. To avoid sampled pairs being dominated by the negative relationship, previous works tend to sample pairs with confident positive relationships in various ways. However, it is a waste of the labeling budget as most sampled pairs will be positive and already have a very close distance. Thus, there is no significant improvement in the model performance. In this paper, we first argue that balanced sampling is the key to active learning for Re-ID. Along this line, we propose a naïve balanced sampling method based on the global estimation of the most confusing distance. It is further improved by the label-wise estimation and diversity measurement. We also formulate the training of Re-ID models as a constrained clustering problem, where labeled positive and negative pairs are as must-link and cannot-link. Then the model training is based on the pseudo labels. Extensive experiments on benchmarks evaluate the effectiveness and superiority of the proposed methods. Specifically, it achieves comparable performance with supervised counterparts with less than 0.1% pair-wise annotation, which significantly surpasses the state-of-the-art.
Leqi Shen, Guiguang Ding, Zhiheng Zhou 0001, Tianshi Xu, Xiaofeng Jin, Yuheng Huang 0005
ICME2
2024 Camera Bias Regularization for Person Re-identification
abstract
Person re-identification (Re-ID) is to match persons captured by non-overlapping cameras. Due to the discrepancies between cameras caused by illumination, background, or viewpoint, the underlying difficulty for Re-ID is the camera bias problem, which leads to the large gap of within-identity features from different cameras. With limited cross-camera annotation, Re-ID models tend to learn camera-related features, instead of identity-related features. Consequently, Re-ID models suffer from poor transfer ability from seen to unseen domains. In this paper, we investigate the camera bias problem in both supervised and unsupervised learning. In particular, we propose a novel Camera Bias Regularization (CBR) term to reduce the feature distribution gap between cameras. The CBR works by simultaneously enlarging the distance of intra-camera distributions between positive and negative pairs, and reducing the distance of positive pairs’ distributions between intra-camera and cross-camera. In addition, a Cross-Camera (CC) clustering method is also designed for unsupervised learning, which puts more emphasis on cross-camera pairs than intra-camera ones during the clustering process. Extensive experiments are conducted to validate the effectiveness of the proposed CBR and CC. Specifically, with only a plain ResNet-50, it achieves 56.7% mAP and 40.7% mAP on the challenging MSMT17 dataset in supervised and unsupervised settings respectively, which surpasses most state-of-the-arts.
Leqi Shen, Guiguang Ding, Zhiheng Zhou 0001, Tianshi Xu, Xiaofeng Jin, Yuheng Huang 0005
ICME2
2024 X-ReID: Cross-Instance Transformer for Identity-Level Person Re-Identification
abstract
Currently, most existing person re-identification methods use instance-level features, which are extracted only from a single image. However, these instance-level features can easily ignore the discriminative information because the appearance of each identity varies greatly in different images. Thus, it is necessary to exploit identity-level features, which can be shared across different images of each identity. In this paper, we propose a novel training framework, named X-ReID, to promote instance-level features to identity-level features by employing cross-attention to incorporate information from one image to another of the same identity, thus more unified and discriminative pedestrian information can be obtained. Extensive experiments on benchmark datasets show the superiority of our method over existing works. Particularly, on the challenging MSMT17, our proposed method gains 1.1% mAP improvements when compared to the second place.
Leqi Shen, Sicheng Zhao, Zhelun Shen, Tianshi Xu, Guiguang Ding
ICME1
2024 Multi-Label Learning with Block Diagonal Labels
abstract
Collecting large-scale multi-label data with full labels is difficult for real-world scenarios. Many existing studies have tried to address the issue of missing labels caused by annotation but ignored the difficulties encountered during the annotation process. We find that the high annotation workload can be attributed to two reasons: (1) Annotators are required to identify labels on widely varying visual concepts. (2) Exhaustively annotating the entire dataset with all the labels becomes notably difficult and time-consuming. In this paper, we propose a new setting, i.e. block diagonal labels, to reduce the workload on both sides. The numerous categories can be divided into different subsets based on semantics and relevance. Each annotator can only focus on its own subset of labels so that only a small set of highly relevant labels are required to be annotated per image. To deal with the issue of such missing labels, we introduce a simple yet effective method that does not require any prior knowledge of the dataset. In practice, we propose an Adaptive Pseudo-Labeling method to predict the unknown labels with less noise. Formal analysis is conducted to evaluate the superiority of our setting. Extensive experiments are conducted to verify the effectiveness of our method on multiple widely used benchmarks.
Leqi Shen, Sicheng Zhao, Hui Chen 0013, Jundong Zhou, Pengzhang Liu, Yongjun Bao, Guiguang Ding
ACM Multimedia1
2022 SECRET: Self-Consistent Pseudo Label Refinement for Unsupervised Domain Adaptive Person Re-identification
abstract
Unsupervised domain adaptive person re-identification aims at learning on an unlabeled target domain with only labeled data in source domain. Currently, the state-of-the-arts usually solve this problem by pseudo-label-based clustering and fine-tuning in target domain. However, the reason behind the noises of pseudo labels is not sufficiently explored, especially for the popular multi-branch models. We argue that the consistency between different feature spaces is the key to the pseudo labels’ quality. Then a SElf-Consistent pseudo label RefinEmenT method, termed as SECRET, is proposed to improve consistency by mutually refining the pseudo labels generated from different feature spaces. The proposed SECRET gradually encourages the improvement of pseudo labels’ quality during training process, which further leads to better cross-domain Re-ID performance. Extensive experiments on benchmark datasets show the superiority of our method. Specifically, our method outperforms the state-of-the-arts by 6.3% in terms of mAP on the challenging dataset MSMT17. In the purely unsupervised setting, our method also surpasses existing works by a large margin. Code is available at https://github.com/LunarShen/SECRET.
Leqi Shen, Guiguang Ding, Zhenhua Guo 0005
AAAI2