Yepeng Tang

dblp:334/6253 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-4314-7550ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Toward Oriented Multi-Object Tracking for Fisheye Images: Dataset and Framework
abstract
Multi-object tracking (MOT) in fisheye images becomes particularly challenging due to significant radial distortion. In MOT, Camera Motion Compensation (CMC) is crucial for mitigating inter-frame errors caused by camera movement. However, conventional CMC methods degrade in fisheye images, as their rigid motion models cannot characterize the non-uniform motion fields introduced by severe distortion. In this paper, we first establish this limitation through theoretical derivation and then propose a plug-and-play ROI-Centric Motion Compensation (RCMC) mechanism, which serves as a new CMC solution that leverages instance-aware ROI for motion estimation. Moreover, fisheye distortion introduces significant target tilt, distorting object geometry and thereby hindering accurate detection and separation of adjacent instances. To address this issue, we integrate RCMC with tailored adjustments for oriented bounding boxes (OBB) to form FO-MOT, a MOT framework for fisheye images. In addition, to overcome the lack of dedicated evaluation benchmarks, we present FisheyeMOT, a comprehensively annotated dataset with OBB-based labels, designed to support the development and validation of fisheye MOT systems. Finally, we demonstrate the superiority of our framework by qualitative and quantitative experiments. The dataset and code will be released at https://github.com/lukanightfever/FO-MOT.
Chunyu Lin, Lang Nie, Yepeng Tang, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 IA2GNN: Imbalance-Aware Adaptive Graph Construction for Multi-Modal Image Fusion
abstract
Image fusion aims to integrate multi-modal data, enhancing information completeness while reducing redundancy. However, image fusion faces the challenge of information imbalance. This imbalance is reflected in both intra-image variations, where different regions exhibit significant differences in information density and importance, and cross-modal inconsistencies, where corresponding regions in different modalities contribute unequally to the fused result. Most existing image fusion methods, such as models based on CNN and window attention mechanism, adopt a fixed approach that collects the same amount of neighboring information for each pixel. However, this fixed approach lacks adaptability to local variations in information density, ultimately limiting fusion performance. To this end, we propose IA2GNN, an imbalance-aware adaptive graph neural network designed for image fusion. To tackle intra- and cross-modal information imbalance, we leverage the flexibility of graph structures to dynamically adjust node connections during feature extraction and fusion. Specifically, we adjust the connectivity of nodes in the graph to allow regions with rich details to establish more connections, thereby enhancing the extraction and preservation of key information in these regions. For less informative regions, we reduce connections to prevent over-modeling. In this way, we optimize information distribution, achieving a balance between detail retention and information integration in the fusion results. Experimental results demonstrate that the proposed model outperforms baseline methods across multiple image fusion tasks, achieving up to +0.058 improvement in SSIM and +0.351 gain in MI.
Yepeng Tang, Chunjie Zhang 0001, Wei Wang 0108, Xiaolong Zheng 0001, Yao Zhao 0001
IEEE Trans. Multim.2
2025 VRoPE: Rotary Position Embedding for Video Large Language Models
abstract
Rotary Position Embedding (RoPE) has shown strong performance in text-based Large Language Models (LLMs), but extending it to video remains a challenge due to the intricate spatiotemporal structure of video frames.Existing adaptations, such as RoPE-3D, attempt to encode spatial and temporal dimensions separately but suffer from two major limitations: positional bias in attention distribution and disruptions in video-text transitions.To overcome these issues, we propose Video Rotary Position Embedding (VRoPE), a novel positional encoding method tailored for Video-LLMs.Specifically, we introduce a more balanced encoding strategy that mitigates attention biases, ensuring a more uniform distribution of spatial focus.Additionally, our approach restructures positional indices to ensure a smooth transition between video and text tokens.Extensive experiments on different models demonstrate that VRoPE consistently outperforms previous RoPE variants, achieving significant improvements in video understanding, temporal reasoning, and retrieval tasks.Code is available at https://github.com/johncaged/VRoPE.
Longteng Guo, Yepeng Tang, Tongtian Yue, Junxian Cai, Qingbin Liu, Jing Liu 0001
EMNLP3
2025 Visual Relation Diffusion for Human-Object Interaction Detection
Yepeng Tang, Chunjie Zhang 0001, Xiaolong Zheng 0001, Chao Liang 0001, Yunchao Wei, Yao Zhao 0001
ICCV2
2025 Diffusion Feedback Helps CLIP See Better
abstract
Contrastive Language-Image Pre-training (CLIP), which excels at abstracting open-world representations across domains and modalities, has become a foundation for a variety of vision and multimodal tasks. However, recent studies reveal that CLIP has severe visual shortcomings, such as which can hardly distinguish orientation, quantity, color, structure, etc. These visual shortcomings also limit the perception capabilities of multimodal large language models (MLLMs) built on CLIP. The main reason could be that the image-text pairs used to train CLIP are inherently biased, due to the lack of the distinctiveness of the text and the diversity of images. In this work, we present a simple post-training approach for CLIP models, which largely overcomes its visual shortcomings via a self-supervised diffusion process. We introduce DIVA, which uses the DIffusion model as a Visual Assistant for CLIP. Specifically, DIVA leverages generative feedback from text-to-image diffusion models to optimize CLIP representations, with only images (without corresponding text). We demonstrate that DIVA improves CLIP's performance on the challenging MMVP-VLM benchmark which assesses fine-grained visual abilities to a large extent (e.g., 3-7%), and enhances the performance of MLLMs and vision models on multimodal understanding and segmentation tasks. Extensive evaluation on 29 image classification and retrieval benchmarks confirms that our framework preserves CLIP's strong zero-shot capabilities. The code is publicly available at https://github.com/baaivision/DIVA.
Wenxuan Wang 0002, Yepeng Tang, Jing Liu 0001
ICLR4
2024 Entity Dependency Learning Network With Relation Prediction for Video Visual Relation Detection
abstract
Video Visual Relation Detection (VidVRD) is a pivotal task in the field of video analysis. It involves detecting object trajectories in videos, predicting potential dynamic relation between these trajectories, and ultimately representing these relationships in the form oftriplets. Correct prediction of relation is vital for VidVRD. Existing methods mostly adopt the simple fusion of visual and language features of entity trajectories as the feature representation for relation predicates. However, these methods do not take into account the dependency information between the relation predication and the subject and object within the triplet. To address this issue, we propose the entity dependency learning network(EDLN), which can capture the dependency information between relation predicates and subjects, objects, and subject-object pairs. It adaptively integrates these dependency information into the feature representation of relation predicates. Additionally, to effectively model the features of the relation existing between various object entities pairs, in the context encoding phase for relation predicate features, we introduce a fully convolutional encoding approach as a substitute for the self-attention mechanism in the Transformer. Extensive experiments on two public datasets demonstrate the effectiveness of the proposed EDLN.
Guoguang Zhang, Yepeng Tang, Chunjie Zhang 0001, Xiaolong Zheng 0001, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Learnable Feature Augmentation Framework for Temporal Action Localization
abstract
Temporal action localization (TAL) has drawn much attention in recent years, however, the performance of previous methods is still far from satisfactory due to the lack of annotated untrimmed video data. To deal with this issue, we propose to improve the utilization of current data through feature augmentation. Given an input video, we first extract video features with pre-trained video encoders, and then randomly mask various semantic contents of video features to consider different views of video features. To avoid damaging important action-related semantic information, we further develop a learnable feature augmentation framework to generate better views of videos. In particular, a Mask-based Feature Augmentation Module (MFAM) is proposed. The MFAM has three advantages: 1) it captures the temporal and semantic relationships of original video features, 2) it generates masked features with indispensable action-related information, and 3) it randomly recycles some masked information to ensure diversity. Finally, we input the masked features and the original features into shared action detectors respectively, and perform action classification and localization jointly for model learning. The proposed framework can improve the robustness and generalization of action detectors by learning more and better views of videos. In the testing stage, the MFAM can be removed, which does not bring extra computational costs. Extensive experiments are conducted on four TAL benchmark datasets. Our proposed framework significantly improves different TAL models and achieves the state-of-the-art performances.
Yepeng Tang, Weining Wang 0001, Chunjie Zhang 0001, Jing Liu 0001, Yao Zhao 0001
IEEE Trans. Image Process.1
2024 Temporal Action Proposal Generation With Action Frequency Adaptive Network
abstract
As the cornerstone of human-behavior analysis in video understanding, temporal action proposal generation aims to predict the starting and ending time of human action instances in untrimmed videos. Although large achievements in temporal action proposal generation have been achieved, most previous studies ignore the variability of action frequency in raw videos, leading to unsatisfying performances on high-action-frequency videos. In fact, there exists two main issues which should be well addressed: data imbalance between high and low action-frequency videos, and inferior detection of short actions in high-action-frequency videos. To address the above issues, we propose an effective framework by adapting to the variability of action frequency, namely Action Frequency Adaptive Network (AFAN), which can be flexibly built upon any temporal action proposal generation method. AFAN consists of two modules: Learning From Experts (LFE) and Fine-Grained Processing (FGP). The LFE first trains a series of action proposal generators on different subsets of imbalanced data as experts and then teaches a unified student model via knowledge distillation. To better detect short actions, FGP first finds out high-action-frequency videos and then performs fine-grained detection. Extensive experimental results on four benchmark datasets (ActivityNet-1.3, HACS, THUMOS14 and FineAction) demonstrate the effectiveness and generalizability of the proposed AFAN, especially for high-action-frequency videos.
Yepeng Tang, Weining Wang 0001, Chunjie Zhang 0001, Jing Liu 0001, Yao Zhao 0001
IEEE Trans. Multim.1
2023 Anchor-free temporal action localization via Progressive Boundary-aware Boosting
Yepeng Tang, Weining Wang 0001, Chunjie Zhang 0001, Jing Liu 0001
Inf. Process. Manag.1