Yifan Yang 0003

dblp:83/89-3 · DBLP profile ↗
← Back
29ranked-venue papers
0as first author
27since 2021 · last 2026
0000-0003-4237-5874ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Explainable Video Camouflaged Object Detection: SAM2 with Eventstream-Inspired Data
abstract
Video Camouflaged Object Detection (VCOD) poses significant challenges due to the subtle appearance of camouflaged objects, especially under dynamic motion and occlusion. Existing methods predominantly rely on optical flow or black-box features for motion modeling, which often entail substantial computational costs and suffer from limited interpretability. Inspired by the human strategy of identifying abnormal movements between frames and the principle of event camera image formation, we propose an eventstream-inspired dual-branch framework for VCOD. Specifically, we design an eventstream-like data extraction module to capture pixel-level motion variations, effectively distinguishing object motion from background dynamics. This event-based representation is integrated into SAM2 through a dual-branch memory-augmented framework, consisting of Time Bridge Attention and Visual Bridge Attention, enabling joint modeling of motion and appearance cues. In addition, we introduce a Prompt Embedding Generator to eliminate the need for human-provided interactive prompts, facilitating fully automatic VCOD. Extensive experiments on MoCA-Mask and CAD2016 demonstrate that our approach significantly outperforms state-of-the-art methods, achieving both superior segmentation accuracy and interpretable motion modeling. To the best of our knowledge, this is the first work to incorporate eventstream-inspired representations into the VCOD task.
Hong Zhang 0018, Yixuan Lyu, Jianbo Song, Ding Yuan 0001, Yifan Yang 0003
AAAI6
2026 MCTrack: Multi-cue spatio-temporal object tracking
Jianbo Song, Hong Zhang 0018, Yachun Feng, Yifan Yang 0003
Expert Syst. Appl.6
2026 MGFNet: Meta Global Filter Network for multi-size image feature extraction
Hong Zhang 0018, Jiaxu Wan, Jianbo Song, Ding Yuan 0001, Yifan Yang 0003
Int. J. Comput. Vis.6
2026 Efficient early exit single object tracking via general distribution
Yachun Feng, Ding Yuan 0001, Jianbo Song, Yifan Yang 0003, Tianxiao Zhang
Neurocomputing5
2026 SDP-GS: Sparse-view Gaussian splatting via segmentation-aware depth priors
Qi Zhao 0037, Yangyan Deng, Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001
Neurocomputing5
2026 STCTracker: Enhancing sequential temporal consistency in referring multi-object tracking
Hong Zhang 0018, Jiabi Zhao, Ding Yuan 0001, Yifan Yang 0003
Image Vis. Comput.6
2026 LKTrack: a novel tracking framework with large kernel network
Hong Zhang 0018, Huakao Lin, Ding Yuan 0001, Jianbo Song, Yifan Yang 0003
Multim. Syst.6
2026 An end-to-end shadow removal framework with an intuitive interaction scheme
Ding Yuan 0001, Yuqian Meng, Yachun Feng, Hong Zhang 0018, Yifan Yang 0003
Pattern Recognit.6
2026 E2IGB: Enhanced effective-information-guided class-balanced loss for long-tailed object recognition
Hong Zhang 0018, Zhigang Li 0005, Yangyan Deng, Yachun Feng, Ding Yuan 0001, Yifan Yang 0003
Pattern Recognit.7
2026 TransSTC: transformer tracker meets efficient spatial-temporal cues
Hong Zhang 0018, Wanli Xing 0004, Yifan Yang 0003, Ding Yuan 0001
Pattern Recognit.3
2025 SP2T: Sparse Proxy Attention for Dual-Stream Point Transformer
Jiaxu Wan, Hong Zhang 0018, Ziqi He, Yangyan Deng, Qishu Wang, Ding Yuan 0001, Yifan Yang 0003
ICCV7
2025 EMA-GS: Improving sparse point cloud rendering with EMA gradient and anchor upsampling
Ding Yuan 0001, Sizhe Zhang, Hong Zhang 0018, Yangyan Deng, Yifan Yang 0003
Image Vis. Comput.5
2025 AwareTrack: Object awareness for visual tracking via templates interaction
Hong Zhang 0018, Jianbo Song, Yifan Yang 0003, Huimin Ma 0001
Image Vis. Comput.5
2025 CODdiff: Prior leading diffusion model for Camouflage Object Detection
Hong Zhang 0018, Yixuan Lyu, Xuliang Li 0005, Yawei Li 0003, Ding Yuan 0001, Yifan Yang 0003
Knowl. Based Syst.7
2025 Language-guided Visual Tracking: Comprehensive and Effective Multimodal Information Fusion
abstract
Current vision-language trackers often struggle to fuse multimodal information comprehensively and effectively, leading to suboptimal performance in multimodal tasks. This study introduces LGTrack, a novel language-guided visual tracking framework designed to achieve a more comprehensive and efficient fusion of vision and language information. In the encoding stage, an Enhanced Multimodal Interaction Module is proposed to achieve full multimodal fusion, and it is used to construct Early Language Multilevel-guided Multimodal Encoding, which leverages deep semantic information for early and multilevel guidance of vision encoding. In the decoding stage, a multimodal decoding based on Joint Query is proposed, utilizing global features from both vision and language modalities, guiding the efficient operation of the decoding layers. These innovations achieve a more comprehensive fusion of multimodal information. Additionally, a contrastive learning strategy is introduced to align vision-language features in the semantic space, further enhancing the fusion effectiveness. Extensive experiments on multiple benchmarks such as LaSOT, \(\rm{LaSOT_{ext}}\) , TNL2K, and OTB99-Lang demonstrate that our approach outperforms existing state-of-the-art trackers.
Jianbo Song, Hong Zhang 0018, Yachun Feng, Yifan Yang 0003
ACM Trans. Multim. Comput. Commun. Appl.5
2025 P2FTrack: Multi-Object Tracking with Motion Prior and Feature Posterior
abstract
Multiple object tracking (MOT) has emerged as a crucial component of the rapidly developing computer vision. However, existing multi-object tracking methods often overlook the relationship between features and motion, hindering the ability to strike a performance balance between coupled motion and complex scenes. In this work, we propose a novel end-to-end multi-object tracking method that integrates motion and feature information. To achieve this, we introduce a motion prior generator that transforms motion information into attention masks. Additionally, we leverage prior-posterior fusion multi-head attention to combine the motion-derived priors and attention-based posteriors. Our proposed method is extensively evaluated on MOT17 and DanceTrack datasets through comprehensive experiments and ablation studies, demonstrating state-of-the-art performance in the feature-based method with reasonable speed.
Hong Zhang 0018, Jiaxu Wan, Jing Zhang 0017, Ding Yuan 0001, Xuliang Li 0005, Yifan Yang 0003
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Unlocking Attributes' Contribution to Successful Camouflage: A Combined Textual and Visual Analysis Strategy
Hong Zhang 0018, Yixuan Lyu, Qian Yu 0008, Huimin Ma 0001, Ding Yuan 0001, Yifan Yang 0003
ECCV (55)7
2024 Feature Block-Aware Correlation Filters for Real-Time UAV Tracking
abstract
Recently, by virtue of the high computational efficiency and accuracy, discriminative correlation filter (DCF)- based tracking methods have gained attraction in the field of unmanned aerial vehicle (UAV). However, conventional DCF-based methods merely rely on cyclic shift to produce training samples. As a result, the filter trained by these samples owns limited discriminative ability, ineffectively addressing various challenges in the tracking stage. Here, to promote the filter's discriminative ability, we develop a feature block-aware correlation filter (CF) method. Specifically, the extracted feature is divided into two blocks, i.e., target and background feature blocks. These blocks only contain target and background features, respectively, by using different mask matrixes. Then, two regularization terms are proposed to combine both feature blocks into the DCF framework. In addition, we employ effective channel reliability weights to generate target response for precise positioning. Furthermore, substantial experiments have been accomplished on multiple public UAV benchmarks, proving that our tracker possesses superior tracking capabilities and operates at ∼40 frames per second (FPS) on the CPU platform.
Hong Zhang 0018, Yan Li 0094, Ding Yuan 0001, Yifan Yang 0003
IEEE Signal Process. Lett.5
2024 Efficient Template Distinction Modeling Tracker With Temporal Contexts for Aerial Tracking
abstract
In recent years, trackers based on neural networks have demonstrated excellent tracking performance. Compared to general tracking tasks, aerial tracking tasks require more stringent operational efficiency of the tracker. Furthermore, since aerial devices such as unmanned aerial vehicles usually shoot from a high altitude, the tracked target occupies fewer pixels, resulting in scarce target discrimination information and more susceptibility to interference from cluttered backgrounds. Unfortunately, existing trackers usually model the entire template region nondifferently, which tends to confuse target and background information in the template and reduce the robustness of the tracker in complex scenes. Additionally, most tracking networks refuse or only refer to a single feature layer to learn the risky temporal context information, which makes it difficult to accurately supplement the scarce target discrimination information. To this end, we propose an efficient template distinction modeling tracker with temporal contexts, called ETDMT, designed to improve the complex scene robustness of aerial tracking through a combination of template distinction modeling and temporal context analysis. The tracker employs a template distinction modeling transformer network, which distinguishes between target and background elements and adopts different modeling approaches for different elements to alleviate the background interference problem prevalent in aerial tracking. Then, the temporal contexts of the tracker are complemented by a global-local spatial awareness update module, which enhances the tracker’s understanding of the latest target state by comprehensively evaluating and adjusting dynamic templates. Extensive experiments demonstrate that the proposed ETDMT achieves advanced aerial tracking performance and efficient running speed.
Hong Zhang 0018, Wanli Xing 0004, Huakao Lin, Yifan Yang 0003
IEEE Trans. Geosci. Remote. Sens.5
2024 UEDG:Uncertainty-Edge Dual Guided Camouflage Object Detection
abstract
According to Darwinian evolutionary theory, numerous species in the wild have developed remarkable adaptive mechanisms, involving pattern rearrangement and environmental assimilation, to evade predators. These obfuscation strategies pose significant challenges for both individuals and algorithms when performing the Camouflage Object Detection (COD) task in complex and intricate scenarios. Inspired by human strategies in the COD task, which involve assigning uncertainties to the entire input and then focusing on highly uncertain areas with the aid of prior knowledge such as boundary information, we propose the Uncertainty-Edge Dual Guide (UEDG) architecture. UEDG effectively combines probabilistic-derived uncertainty and deterministic-derived edge information to accurately detect concealed objects. The architecture consists of two independent branches dedicated to uncertainty reasoning and edge inference, which are subsequently integrated into a feature fusion module utilizing recursion feedback and feature-reuse techniques. This novel COD framework leverages the benefits of Bayesian learning and convolution-based learning, resulting in a powerful multi-task guided approach. Extensive experiments conducted on four widely employed datasets demonstrate the superior performance of UEDG compared to 12 state-of-the-art approaches, while maintaining an acceptable level of computational complexity. Overall, UEDG presents a promising solution for addressing the challenges of COD in complex environments by combining evolutionary-inspired strategies with advanced computer vision techniques.
Yixuan Lyu, Hong Zhang 0018, Yan Li 0094, Yifan Yang 0003, Ding Yuan 0001
IEEE Trans. Multim.5
2023 Enhancing Spatial Consistency and Class-Level Diversity for Segmenting Fine-Grained Objects
Qi Zhao 0037, Binghao Liu, Shuchang Lyu, Yifan Yang 0003
ICONIP (11)5
2023 SiamST: Siamese network with spatio-temporal awareness for object tracking
Hong Zhang 0018, Wanli Xing 0004, Yifan Yang 0003, Yan Li 0094, Ding Yuan 0001
Inf. Sci.3
2023 RISTrack: Learning Response Interference Suppression Correlation Filters for UAV Tracking
abstract
With the high computation efficiency and tracking accuracy, discriminative correlation filters (DCF) have been applied to UAV tracking. However, in the scenarios (i.e., complex background and temporary occlusion), DCF-based trackers usually generate low credibility response under the influence of background distractors, which contains multiple side peaks and declines the tracking performance. Motivated by the response consistency in adjacent frames and background information penalization, we propose learning a response interference suppression (RIS) correlation filter to tackle this problem. Specifically, we introduce a RIS regularization into the DCF-based framework, which aims to keep the target area response consistent in adjacent frames and repress distractors’ response in the background. Besides, we adopt a response auxiliary strategy (RAS) to smooth the target response, which intends to obtain the precise location and avoid target drift. Furthermore, extensive experiments on three UAV benchmarks demonstrate the excellent performance of the proposed method against other 19 state-of-the-art trackers. Moreover, the tracking speed of the proposed method can reach 42 FPS on a single CPU.
Yan Li 0094, Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001
IEEE Geosci. Remote. Sens. Lett.3
2023 Cross-modality complementary information fusion for multispectral pedestrian detection
Chaoqi Yan, Hong Zhang 0018, Xuliang Li 0005, Yifan Yang 0003, Ding Yuan 0001
Neural Comput. Appl.4
2022 Feature adaptation-based multipeak-redetection spatial-aware correlation filter for object tracking
Wanli Xing 0004, Hong Zhang 0018, Hao Chen 0052, Yifan Yang 0003, Ding Yuan 0001
Neurocomputing4
2022 A feature consistency driven attention erasing network for fine-grained image retrieval
Qi Zhao 0037, Shuchang Lyu, Binghao Liu, Yifan Yang 0003
Pattern Recognit.5
2022 MSAGNet: Multi-Stream Attribute-Guided Network for Occluded Pedestrian Detection
abstract
Pedestrian detection plays an indispensable role in human-centric applications. Although having enjoyed the merits of generic object detectors based on deep learning frameworks, pedestrian detection is still a persistent crucial task since the pedestrians often gather together and occlude each other. In this study, we propose a simple yet effective Multi-Stream Attribute-Guided Network (MSAGNet) to regard occluded pedestrian detection as a standard central point and height estimation problem. Specifically, we focus on searching for the central points of the pedestrians and predicting the scales and offsets of the corresponding pedestrians. Meanwhile, an adaptive weighting parameter, i.e., Intersection over the Visible part region of ground truth (IoV), is utilized to conduct accurate bounding box regression. Furthermore, a novel nonlinear Non-Maximum Suppression (NMS) is proposed to flexibly prune false positives and decrease the miss rate of adjacent overlapping pedestrians. Experimental results on Caltech-USA, CityPersons, CrowdHuman and WiderPerson pedestrian datasets show that the proposed MSAGNet can obtain significant performance boosts, while maintaining a reasonable run-time speed.
Hong Zhang 0018, Chaoqi Yan, Xuliang Li 0005, Yifan Yang 0003, Ding Yuan 0001
IEEE Signal Process. Lett.4
2019 Video denoising for security and privacy in fog computing
abstract
Summary To reduce heavy noise from degraded video in low or predictable latency and preserve privacy, a powerful and efficient video denoising algorithm is proposed based on fog computing for Visual Internet of Things. The conventional method is to remove noise in the cloud; however, this may overload computation and communication and raise security and privacy issues. The proposed denoising algorithm is distributed to heterogeneous devices at network edges to preserve privacy and avoid security risks as noise can be reduced in the fog rather than the cloud. To address the problems of latency, communication rate, and extremely heavy noise, structure registration, inter‐frame and inner‐frame filters, and distribution compensation are applied in the proposed algorithm. A scheme for encrypting the denoised data at network edges is provided so that security and privacy issues may be avoided during transmission and storage. Compared with other denoising approaches under extremely heavy noise conditions, the experimental results demonstrate that the proposed approach achieves superior denoising performance in terms of peak signal‐noise ratio and visual quality at low computational cost, high bandwidth efficiency, and low‐latency response in a fog computing manner.
Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001, Daniel Sun 0004, Jun Zhang 0010, Guoqiang Li 0001, Mingui Sun
Concurr. Comput. Pract. Exp.2
2018 End-to-end temporal attention extraction and human action recognition
Hong Zhang 0018, Miao Xin, Shuhang Wang, Yifan Yang 0003, Lei Zhang 0098, Helong Wang
Mach. Vis. Appl.4