Ding Yuan 0001

dblp:230/9043-1 · DBLP profile ↗
← Back
42ranked-venue papers
3as first author
25since 2021 · last 2026
0000-0001-8107-7218ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Explainable Video Camouflaged Object Detection: SAM2 with Eventstream-Inspired Data
abstract
Video Camouflaged Object Detection (VCOD) poses significant challenges due to the subtle appearance of camouflaged objects, especially under dynamic motion and occlusion. Existing methods predominantly rely on optical flow or black-box features for motion modeling, which often entail substantial computational costs and suffer from limited interpretability. Inspired by the human strategy of identifying abnormal movements between frames and the principle of event camera image formation, we propose an eventstream-inspired dual-branch framework for VCOD. Specifically, we design an eventstream-like data extraction module to capture pixel-level motion variations, effectively distinguishing object motion from background dynamics. This event-based representation is integrated into SAM2 through a dual-branch memory-augmented framework, consisting of Time Bridge Attention and Visual Bridge Attention, enabling joint modeling of motion and appearance cues. In addition, we introduce a Prompt Embedding Generator to eliminate the need for human-provided interactive prompts, facilitating fully automatic VCOD. Extensive experiments on MoCA-Mask and CAD2016 demonstrate that our approach significantly outperforms state-of-the-art methods, achieving both superior segmentation accuracy and interpretable motion modeling. To the best of our knowledge, this is the first work to incorporate eventstream-inspired representations into the VCOD task.
Hong Zhang 0018, Yixuan Lyu, Jianbo Song, Ding Yuan 0001, Yifan Yang 0003
AAAI5
2026 MGFNet: Meta Global Filter Network for multi-size image feature extraction
Hong Zhang 0018, Jiaxu Wan, Jianbo Song, Ding Yuan 0001, Yifan Yang 0003
Int. J. Comput. Vis.5
2026 Efficient early exit single object tracking via general distribution
Yachun Feng, Ding Yuan 0001, Jianbo Song, Yifan Yang 0003, Tianxiao Zhang
Neurocomputing2
2026 SDP-GS: Sparse-view Gaussian splatting via segmentation-aware depth priors
Qi Zhao 0037, Yangyan Deng, Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001
Neurocomputing6
2026 STCTracker: Enhancing sequential temporal consistency in referring multi-object tracking
Hong Zhang 0018, Jiabi Zhao, Ding Yuan 0001, Yifan Yang 0003
Image Vis. Comput.3
2026 LKTrack: a novel tracking framework with large kernel network
Hong Zhang 0018, Huakao Lin, Ding Yuan 0001, Jianbo Song, Yifan Yang 0003
Multim. Syst.3
2026 An end-to-end shadow removal framework with an intuitive interaction scheme
Ding Yuan 0001, Yuqian Meng, Yachun Feng, Hong Zhang 0018, Yifan Yang 0003
Pattern Recognit.1
2026 E2IGB: Enhanced effective-information-guided class-balanced loss for long-tailed object recognition
Hong Zhang 0018, Zhigang Li 0005, Yangyan Deng, Yachun Feng, Ding Yuan 0001, Yifan Yang 0003
Pattern Recognit.6
2026 TransSTC: transformer tracker meets efficient spatial-temporal cues
Hong Zhang 0018, Wanli Xing 0004, Yifan Yang 0003, Ding Yuan 0001
Pattern Recognit.5
2025 SP2T: Sparse Proxy Attention for Dual-Stream Point Transformer
Jiaxu Wan, Hong Zhang 0018, Ziqi He, Yangyan Deng, Qishu Wang, Ding Yuan 0001, Yifan Yang 0003
ICCV6
2025 EMA-GS: Improving sparse point cloud rendering with EMA gradient and anchor upsampling
Ding Yuan 0001, Sizhe Zhang, Hong Zhang 0018, Yangyan Deng, Yifan Yang 0003
Image Vis. Comput.1
2025 CODdiff: Prior leading diffusion model for Camouflage Object Detection
Hong Zhang 0018, Yixuan Lyu, Xuliang Li 0005, Yawei Li 0003, Ding Yuan 0001, Yifan Yang 0003
Knowl. Based Syst.6
2025 P2FTrack: Multi-Object Tracking with Motion Prior and Feature Posterior
abstract
Multiple object tracking (MOT) has emerged as a crucial component of the rapidly developing computer vision. However, existing multi-object tracking methods often overlook the relationship between features and motion, hindering the ability to strike a performance balance between coupled motion and complex scenes. In this work, we propose a novel end-to-end multi-object tracking method that integrates motion and feature information. To achieve this, we introduce a motion prior generator that transforms motion information into attention masks. Additionally, we leverage prior-posterior fusion multi-head attention to combine the motion-derived priors and attention-based posteriors. Our proposed method is extensively evaluated on MOT17 and DanceTrack datasets through comprehensive experiments and ablation studies, demonstrating state-of-the-art performance in the feature-based method with reasonable speed.
Hong Zhang 0018, Jiaxu Wan, Jing Zhang 0017, Ding Yuan 0001, Xuliang Li 0005, Yifan Yang 0003
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Unlocking Attributes' Contribution to Successful Camouflage: A Combined Textual and Visual Analysis Strategy
Hong Zhang 0018, Yixuan Lyu, Qian Yu 0008, Huimin Ma 0001, Ding Yuan 0001, Yifan Yang 0003
ECCV (55)6
2024 Feature Block-Aware Correlation Filters for Real-Time UAV Tracking
abstract
Recently, by virtue of the high computational efficiency and accuracy, discriminative correlation filter (DCF)- based tracking methods have gained attraction in the field of unmanned aerial vehicle (UAV). However, conventional DCF-based methods merely rely on cyclic shift to produce training samples. As a result, the filter trained by these samples owns limited discriminative ability, ineffectively addressing various challenges in the tracking stage. Here, to promote the filter's discriminative ability, we develop a feature block-aware correlation filter (CF) method. Specifically, the extracted feature is divided into two blocks, i.e., target and background feature blocks. These blocks only contain target and background features, respectively, by using different mask matrixes. Then, two regularization terms are proposed to combine both feature blocks into the DCF framework. In addition, we employ effective channel reliability weights to generate target response for precise positioning. Furthermore, substantial experiments have been accomplished on multiple public UAV benchmarks, proving that our tracker possesses superior tracking capabilities and operates at ∼40 frames per second (FPS) on the CPU platform.
Hong Zhang 0018, Yan Li 0094, Ding Yuan 0001, Yifan Yang 0003
IEEE Signal Process. Lett.4
2024 UEDG:Uncertainty-Edge Dual Guided Camouflage Object Detection
abstract
According to Darwinian evolutionary theory, numerous species in the wild have developed remarkable adaptive mechanisms, involving pattern rearrangement and environmental assimilation, to evade predators. These obfuscation strategies pose significant challenges for both individuals and algorithms when performing the Camouflage Object Detection (COD) task in complex and intricate scenarios. Inspired by human strategies in the COD task, which involve assigning uncertainties to the entire input and then focusing on highly uncertain areas with the aid of prior knowledge such as boundary information, we propose the Uncertainty-Edge Dual Guide (UEDG) architecture. UEDG effectively combines probabilistic-derived uncertainty and deterministic-derived edge information to accurately detect concealed objects. The architecture consists of two independent branches dedicated to uncertainty reasoning and edge inference, which are subsequently integrated into a feature fusion module utilizing recursion feedback and feature-reuse techniques. This novel COD framework leverages the benefits of Bayesian learning and convolution-based learning, resulting in a powerful multi-task guided approach. Extensive experiments conducted on four widely employed datasets demonstrate the superior performance of UEDG compared to 12 state-of-the-art approaches, while maintaining an acceptable level of computational complexity. Overall, UEDG presents a promising solution for addressing the challenges of COD in complex environments by combining evolutionary-inspired strategies with advanced computer vision techniques.
Yixuan Lyu, Hong Zhang 0018, Yan Li 0094, Yifan Yang 0003, Ding Yuan 0001
IEEE Trans. Multim.6
2023 Redefined target sample-based background-aware correlation filters for object tracking
Wanli Xing 0004, Hong Zhang 0018, Yujie Wu 0001, Yawei Li 0003, Ding Yuan 0001
Appl. Intell.5
2023 SiamST: Siamese network with spatio-temporal awareness for object tracking
Hong Zhang 0018, Wanli Xing 0004, Yifan Yang 0003, Yan Li 0094, Ding Yuan 0001
Inf. Sci.5
2023 RISTrack: Learning Response Interference Suppression Correlation Filters for UAV Tracking
abstract
With the high computation efficiency and tracking accuracy, discriminative correlation filters (DCF) have been applied to UAV tracking. However, in the scenarios (i.e., complex background and temporary occlusion), DCF-based trackers usually generate low credibility response under the influence of background distractors, which contains multiple side peaks and declines the tracking performance. Motivated by the response consistency in adjacent frames and background information penalization, we propose learning a response interference suppression (RIS) correlation filter to tackle this problem. Specifically, we introduce a RIS regularization into the DCF-based framework, which aims to keep the target area response consistent in adjacent frames and repress distractors’ response in the background. Besides, we adopt a response auxiliary strategy (RAS) to smooth the target response, which intends to obtain the precise location and avoid target drift. Furthermore, extensive experiments on three UAV benchmarks demonstrate the excellent performance of the proposed method against other 19 state-of-the-art trackers. Moreover, the tracking speed of the proposed method can reach 42 FPS on a single CPU.
Yan Li 0094, Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001
IEEE Geosci. Remote. Sens. Lett.5
2023 Cross-modality complementary information fusion for multispectral pedestrian detection
Chaoqi Yan, Hong Zhang 0018, Xuliang Li 0005, Yifan Yang 0003, Ding Yuan 0001
Neural Comput. Appl.5
2022 R-SSD: refined single shot multibox detector for pedestrian detection
Chaoqi Yan, Hong Zhang 0018, Xuliang Li 0005, Ding Yuan 0001
Appl. Intell.4
2022 Feature adaptation-based multipeak-redetection spatial-aware correlation filter for object tracking
Wanli Xing 0004, Hong Zhang 0018, Hao Chen 0052, Yifan Yang 0003, Ding Yuan 0001
Neurocomputing5
2022 MSAGNet: Multi-Stream Attribute-Guided Network for Occluded Pedestrian Detection
abstract
Pedestrian detection plays an indispensable role in human-centric applications. Although having enjoyed the merits of generic object detectors based on deep learning frameworks, pedestrian detection is still a persistent crucial task since the pedestrians often gather together and occlude each other. In this study, we propose a simple yet effective Multi-Stream Attribute-Guided Network (MSAGNet) to regard occluded pedestrian detection as a standard central point and height estimation problem. Specifically, we focus on searching for the central points of the pedestrians and predicting the scales and offsets of the corresponding pedestrians. Meanwhile, an adaptive weighting parameter, i.e., Intersection over the Visible part region of ground truth (IoV), is utilized to conduct accurate bounding box regression. Furthermore, a novel nonlinear Non-Maximum Suppression (NMS) is proposed to flexibly prune false positives and decrease the miss rate of adjacent overlapping pedestrians. Experimental results on Caltech-USA, CityPersons, CrowdHuman and WiderPerson pedestrian datasets show that the proposed MSAGNet can obtain significant performance boosts, while maintaining a reasonable run-time speed.
Hong Zhang 0018, Chaoqi Yan, Xuliang Li 0005, Yifan Yang 0003, Ding Yuan 0001
IEEE Signal Process. Lett.5
2021 Source Data-free Unsupervised Domain Adaptation for Semantic Segmentation
abstract
Deep\footnote learning-based semantic segmentation methods require a huge amount of training images with pixel-level annotations. Unsupervised domain adaptation (UDA) for semantic segmentation enables transferring knowledge learned from the synthetic data (source domain) with low-cost annotations to the real images (target domain). However, current UDA methods mostly require full access to the source domain data for feasible adaptation, which limits their applications in real-world scenarios with privacy, storage, or transmission issues. To this end, this paper identifies and addresses a more practical but challenging problem of UDA for semantic segmentation, where access to the original source domain data is forbidden. In other words, only the pre-trained source model and unlabelled target domain data are available for adaptation. To tackle the problem, we propose to construct a set of source domain virtual data to mimic the source domain distribution by identifying the target domain high-confidence samples predicted by the pre-trained source model. Then by analyzing the data properties in the cross-domain semantic segmentation tasks, we propose an uncertainty and prior distribution-aware domain adaptation method to align the virtual source domain and the target domain with both adversarial learning and self-training strategies. Extensive experiments on three cross-domain semantic segmentation datasets with in-depth analyses verify the effectiveness of the proposed method.
Mucong Ye, Jing Zhang 0017, Jinpeng Ouyang, Ding Yuan 0001
ACM Multimedia4
2021 Dynamic Scene Video Deblurring Using Robust Incremental Weighted Fourier Aggregation
abstract
Motion blur is an inevitable problem when shooting a video in motion, and it will cause serious degradation of images. In this letter, we propose a new robust incremental weighted Fourier aggregation algorithm for dynamic scene video deblurring. This is motivated by the fact that existing multiframe aggregation methods usually require a tedious iterative process to obtain sharp information in distant frames. We propose to perform alignment and aggregation once for each frame using the current frame and the previous frame result, which includes the sharp information of all the previous frames. Moreover, a criterion for evaluating the alignment of blurred images is proposed. Based on this, a robust aggregation approach is also proposed to help reduce artifacts caused by image patch alignment failure. Specifically, we decompose the current frame into a series of image patches, find consistent image patches in the deblurring results of the previous frame by a coarse-to-fine alignment algorithm, and aggregate them in the Fourier domain. Experiments show that our method can obtain improved deblurring results and reduce the computational complexity compared to the traditional aggregation method.
Yawei Li 0003, Hong Zhang 0018, Yujie Wu 0001, Ding Yuan 0001
IEEE Signal Process. Lett.4
2019 Video denoising for security and privacy in fog computing
abstract
Summary To reduce heavy noise from degraded video in low or predictable latency and preserve privacy, a powerful and efficient video denoising algorithm is proposed based on fog computing for Visual Internet of Things. The conventional method is to remove noise in the cloud; however, this may overload computation and communication and raise security and privacy issues. The proposed denoising algorithm is distributed to heterogeneous devices at network edges to preserve privacy and avoid security risks as noise can be reduced in the fog rather than the cloud. To address the problems of latency, communication rate, and extremely heavy noise, structure registration, inter‐frame and inner‐frame filters, and distribution compensation are applied in the proposed algorithm. A scheme for encrypting the denoised data at network edges is provided so that security and privacy issues may be avoided during transmission and storage. Compared with other denoising approaches under extremely heavy noise conditions, the experimental results demonstrate that the proposed approach achieves superior denoising performance in terms of peak signal‐noise ratio and visual quality at low computational cost, high bandwidth efficiency, and low‐latency response in a fog computing manner.
Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001, Daniel Sun 0004, Jun Zhang 0010, Guoqiang Li 0001, Mingui Sun
Concurr. Comput. Pract. Exp.3
2019 Leveraging semantic segmentation with learning-based confidence measure
Feiyang Cheng, Hong Zhang 0018, Ding Yuan 0001, Mingui Sun
Neurocomputing3
2017 Learning discriminative action and context representations for action recognition in still images
abstract
Action recognition in still images is a challenging task in computer vision. Recent successes in deep feature-learning advance this research, employing robust and rich-semantic feature representation. However, the issue that recognition fails when two action images share similar contexts is long-standing. In this paper, we employ metric learning method to address within-class and between-class confusions in action recognition. We propose a novel loss function, named composite-triplet loss. Supervised by this loss function, our method directly learns a similarity function from data. Employing a customized human action and context detection network, we obtain highly discriminative action image embeddings, which can be used in action image recognition and other tasks. Our approach is evaluated on the still-image action recognition task and the image caption generation task. On the PASCAL VOC dataset, our approach outperforms state-of-the-art methods, achieving 90.6% mean AP.
Miao Xin, Hong Zhang 0018, Ding Yuan 0001, Mingui Sun
ICME3
2017 3D lunar craters detection based on stereo matching
abstract
In this paper, we focus on the 3D crater detection problem on lunar surface, which helps high-precision spacecraft landing and rover navigation in moon exploration projects. A random structured forests method is firstly applied to detect the 2D edges of craters, and then dense correspondence between CCD stereo images estimates the elevations of craters. Finally, we propose a 3D crater detection model, which is solved by an iterative optimization algorithm with the initial 2D edge. Our method is evaluated on three pairs of stereo images captured by Chang'E-I satellite. Experimental results show that the 3D craters in each pair are quickly and accurately located, which demonstrate the effectiveness and efficiency.
Hongmei Zhu, Jihao Yin, Ding Yuan 0001
IGARSS3
2017 SVCV: segmentation volume combined with cost volume for stereo matching
abstract
Stereo matching between binocular stereo images is fundamental to many computer vision tasks, such as three‐dimensional (3D) reconstruction and robot navigation. Various structures of real 3D scenes lead stereo matching to be an old yet still challenging problem. In this study, the authors proposed a novel adaptive support weights technique which exploits the hierarchical information provided by multilevel segmentation to preserve the robustness to imaging conditions and spatial proximity in cost aggregation. Besides, a generalisable cost refinement strategy is designed to remove the matching ambiguity in large weakly textured regions. The proposed strategy utilises both the fluctuation of the filtered cost volume and the colour information to further improve the matching accuracy. Experimental results of 50 stereo images demonstrate the effectiveness and efficiency of the proposed method. Furthermore, a systematic evaluation is developed to assess the conventional steps in local stereo methods and then reliable suggestions are given to the beginners and researchers outside the stereo matching field.
Hongmei Zhu, Jihao Yin, Ding Yuan 0001
IET Comput. Vis.3
2017 Sparse representation over discriminative dictionary for stereo matching
Jihao Yin, Hongmei Zhu, Ding Yuan 0001, Tianfan Xue
Pattern Recognit.3
2016 3-D point cloud normal estimation based on fitting algebraic spheres
abstract
In this paper, we proposed a novel method to estimate the normal information of the unorganized point cloud, which plays an essential part in 3D reconstruction. The original point cloud is firstly divided into cubes with different sizes by the octree method. Then, we fit algebraic sphere in each cube instead of planar surface to improve the accuracy of normal estimation. Finally, the raw normals are refined by a weighting function which increases along with the depth of octree. For evaluation, we compute the intersection angles between the estimated normals and the corresponding groundtruth. Besides, the estimated normals are also plugged into the Poisson surface reconstruction algorithm for intuitive comparison. Experimental results demonstrate the effectiveness of our normal estimating methods. Moreover, the strategy that normal estimation after division saves much more computing time, which promises the efficiency of our method.
Ding Yuan 0001, Hongmei Zhu, Jihao Yin
ICIP2
2016 Dem-based shadow detection and removal for lunar craters
abstract
In this paper, we focus on the shadow problem of the lunar surface, which hinders the implementation of visual tasks and visual processing for moon exploration projects. A random walker model is firstly applied to detect the shadowed pixels in lunar craters. Then, the detected shadows are removed by rectifying the illumination coefficients and detail coefficients, which are obtained by using the multi-scale decomposition technique. The proposed algorithm is evaluated on three groups of CCD images and the corresponding DEM data. Satisfactory results, which achieve comparative illumination yet preserve enough details, demonstrate the effectiveness of the proposed algorithm.
Hongmei Zhu, Jihao Yin, Ding Yuan 0001, Guangyun Zhang
IGARSS3
2016 Recurrent Temporal Sparse Autoencoder for attention-based action recognition
abstract
Visual context is fundamental to understand human actions in videos. However, to efficiently employ temporal context information presents an enormous challenge to this area. Two main problems are long-standing: (1) video frames are redundant while discriminative information is sparse; (2) large amount of interference information is mixed in frame sequences. These factors results in redundant computation and recognition failures. In this paper, we propose a learnable temporal attention mechanism to automatically select important time points from action sequences. We design an unsupervised Recurrent Temporal Sparse Autoencoder (RTSAE) network, which learns to extract sparse key-frames to sharpen discriminative yet to retain descriptive capability, as well to shield interfere information. By applying this technique to a recent proposed action recognition model Adaptive Recurrent-convolutional Hybrid network (ARCH), we significantly improve its performance in both speed and accuracy. Experiments demonstrate that, with the help of the RTSAE, ARCH outperforms most state-of-the-art methods on UCF101 and HMDB51 datasets.
Miao Xin, Hong Zhang 0018, Mingui Sun, Ding Yuan 0001
IJCNN4
2016 ARCH: Adaptive recurrent-convolutional hybrid networks for long-term action recognition
Miao Xin, Hong Zhang 0018, Helong Wang, Mingui Sun, Ding Yuan 0001
Neurocomputing5
2015 Fluctuations of disparity space image for stereo matching in untextured regions
abstract
Allocating the disparities to the untextured regions in stereo image still remains an intractable and challenging problem. In this paper, we present a novel local stereo matching algorithm for large untextured regions. The core ideas behind our method are from two aspects: 1) the fluctuating characteristics of cost volume are first exploited to distinguish ambiguous and unambiguous image regions; 2) the matching costs of pixels in ambiguous regions are regularized with an adaptive cost aggregation. The WTA strategy is performed on the regularized cost volume followed by postprocessing to obtain accurate disparity map. Comparative experiments are conducted on different data sets and the results demonstrate the effectiveness and efficiency of our method.
Hongmei Zhu, Jihao Yin, Ding Yuan 0001, Wei Sui
ICIP3
2015 Single image dehazing using the change of detail prior
Jiafeng Li 0001, Hong Zhang 0018, Ding Yuan 0001, Mingui Sun
Neurocomputing3
2015 Cross-trees, edge and superpixel priors-based cost aggregation for stereo matching
Feiyang Cheng, Hong Zhang 0018, Mingui Sun, Ding Yuan 0001
Pattern Recognit.4
2015 Camera motion estimation through monocular normal flow vectors
Ding Yuan 0001, Jihao Yin, Jiankun Hu
Pattern Recognit. Lett.1
2014 Cross-Trees for Stereo Matching with Priors
abstract
We propose a cross-trees structure to perform the non-local cost aggregation for dense stereo matching. The cross-trees structure consists of a horizontal-tree and a vertical-tree. Compared to other spanning trees, the significant superiority of the cross-trees is that the trees' constructions are efficient and independent on any local or global property. Moreover, the trees are exactly unique. By traversing the two crossed trees successively, a fast non-local cost aggregation algorithm is performed to filter the matching cost volume and then the disparity maps are established with the Winner-Take-All (WTA) strategy. Additionally, two different priors: edge prior and super pixel prior, are proposed to tackle the false smoothing at the depth boundaries. Hence, our method contains two different algorithms in terms of the cross-trees prior in this paper. Performance evaluation on the 27 Middlebury data sets shows that both our algorithms outperform the other two tree-based methods, namely minimum spanning tree (MST) and segment-tree (ST). By performing the non-local cost aggregation on different trees, MST, ST and our method all have competitive rankings on the Middlebury website compared to the local cost aggregation methods.
Feiyang Cheng, Hong Zhang 0018, Mingui Sun, Helong Wang, Ding Yuan 0001
ICPR5
2014 Stereo matching by using the global edge constraint
Feiyang Cheng, Hong Zhang 0018, Ding Yuan 0001, Mingui Sun
Neurocomputing3
2012 Stereo matching with Global Edge Constraint and Graph Cuts
Hong Zhang 0018, Feiyang Cheng, Ding Yuan 0001, Yuecheng Li, Mingui Sun
ICPR3