VLDB 2026 Research / reviewers in the wild / expert
Yanyang Yan
dblp:208/4206
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UMDATrack: Unified Multi-Domain Adaptive Tracking under Adverse Weather ConditionsabstractVisual object tracking has gained promising progress in past decades. Most of the existing approaches focus on learning target representation in well-conditioned daytime data, while for the unconstrained real-world scenarios with adverse weather conditions, e.g. nighttime or foggy environment, the tremendous domain shift leads to significant performance degradation. In this paper, we propose UMDATrack, which is capable of maintaining high-quality target state prediction under various adverse weather conditions within a unified domain adaptation framework. Specifically, we first use a controllable scenario generator to synthesize a small amount of unlabeled videos (less than 2% frames in source daytime datasets) in multiple weather conditions under the guidance of different text prompts. Afterwards, we design a simple yet effective domain-customized adapter (DCA), allowing the target objects' representation to rapidly adapt to various weather conditions without redundant model updating. Furthermore, to enhance the localization consistency between source and target domains, we propose a target-aware confidence alignment module (TCA) following optimal transport theorem. Extensive experiments demonstrate that UMDATrack can surpass existing advanced visual trackers and lead new state-of-the-art performance by a significant margin. Our code is available at https://github.com/Z-Z188/UMDATrack. Siyuan Yao, Wenqi Ren, Yanyang Yan, Xiaochun Cao |
ICCV | 5 |
| 2025 | LLDNet: Joint Low-light Enhancement and Local Motion Deblurring in the DarkabstractLocal motion blur in the dark often occurs in the real world due to long exposure, leading to serious challenges in real-world activities, such as night photography and autonomous driving. Although existing local motion deblurring methods and low-light enhancement methods can solve each problem separately. The simple cascade of these methods cannot handle the joint degradation of the low-light and local motion blur. Therefore, in this paper, we jointly address both low-light conditions and local motion blur, aiming to achieve efficient restoration of low-light local motion blur images. Specifically, to solve the attention noise issue, we introduce the channel differential Transformer that guides the model to focus on essential regions and channels during image enhancement. Besides, we present the window differential Transformer that interpolates the window differential self-attention and the multi-scale feed-forward module to focus on local blurry regions. In addition, the phase content-aware fusion module is employed to enhance the transmission of phase information from encoder to decoder. Extensive experimental results demonstrate the effectiveness of our method for low-light local motion deblurring on both synthetic and real-world datasets. Haigen Liu, Yanyang Yan, Wenqi Ren |
ICME | 2 |
| 2025 | CAN: Cascade Augmentations Against Noise for Image RestorationabstractImage restoration aims to recover the latent clean image from a degraded counterpart. In general, the prevailing state-of-the-art image restoration methods concentrate on solving only a specific degradation type according to the task, e.g., deblurring or deraining. However, if the corresponding well-trained frameworks confront other real-world image corruptions, i.e., the corruptions are not covered in the training phase, and state-of-the-art restoration models will suffer from a lack of generalization ability. We have observed that an image restoration model can be easily confused by noise corruption. Towards improving the robustness of image restoration networks, in this paper, we focus on alleviating the corruption of noise in various image restoration tasks, which is almost inevitable in real-world scenes. To this end, we devise a novel Cascade Augmentation strategy against Noise (CAN) to enhance the robustness of specific image restoration. Specifically, the given degraded images are sequentially augmented from different perspectives, i.e., noise-aware augmentation and model-aware augmentation. The noise-aware augmentation is proposed to enrich the samples by introducing various noise operations. Moreover, to adapt to more unknown corruptions, we propose a novel model-aware augmentation mechanism, which enhances the scalability by exploring useful both spatial and frequency clues with the help of model randomness. It is worth noting that the proposed augmentation scheme is model-agnostic, and it can plug and play into arbitrary state-of-the-art image restoration architectures. In addition, we construct noise corruption benchmark datasets, derived from the validation set of standard image restoration datasets, to assist us in evaluating the robustness of restoration networks. Extensive quantitative and qualitative evaluations demonstrate that the proposed method has strong generalization capability, which can enhance the robustness of various image restoration frameworks when facing diverse noises. Yanyang Yan, Siyuan Yao, Wenqi Ren, Rui Zhang 0040, Qi Guo 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2025 | UncTrack: Reliable Visual Object Tracking With Uncertainty-Aware Prototype Memory NetworkabstractTransformer-based trackers have achieved promising success and become the dominant tracking paradigm because of their accuracy and efficiency. Despite the substantial progress, most of the existing approaches handle object tracking as a deterministic coordinate regression problem, while the target localization uncertainty has been largely overlooked, which hampers trackers' ability to maintain reliable target state prediction in challenging scenarios. To address this issue, we propose UncTrack, a novel uncertainty-aware transformer-based tracker that predicts the target localization uncertainty and incorporates this uncertainty information for accurate target state inference. Specifically, UncTrack uses a transformer encoder to perform feature interactions between the template and search images. The output features are passed into an uncertainty-aware localization decoder (ULD) to coarsely predict the corner-based localization and the corresponding localization uncertainty. Then, the localization uncertainty is sent into a prototype memory network (PMN) to excavate valuable historical information to identify whether the target state prediction is reliable. To enhance the template representation, the samples with high confidence are fed back into the prototype memory bank for memory updating, which makes the tracker more robust to challenging appearance variations. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. Our code is available at https://github.com/ManOfStory/UncTrack. Siyuan Yao, Yanyang Yan, Wenqi Ren, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2024 | INformer: Inertial-Based Fusion Transformer for Camera Shake DeblurringabstractInertial measurement units (IMU) in the capturing device can record the motion information of the device, with gyroscopes measuring angular velocity and accelerometers measuring acceleration. However, conventional deblurring methods seldom incorporate IMU data, and existing approaches that utilize IMU information often face challenges in fully leveraging this valuable data, resulting in noise issues from the sensors. To address these issues, in this paper, we propose a multi-stage deblurring network named INformer, which combines inertial information with the Transformer architecture. Specifically, we design an IMU-image Attention Fusion (IAF) block to merge motion information derived from inertial measurements with blurry image features at the attention level. Furthermore, we introduce an Inertial-Guided Deformable Attention (IGDA) block for utilizing the motion information features as guidance to adaptively adjust the receptive field, which can further refine the corresponding blur kernel for pixels. Extensive experiments on comprehensive benchmarks demonstrate that our proposed method performs favorably against state-of-the-art deblurring approaches. Wenqi Ren, Linrui Wu, Yanyang Yan, Shengyao Xu, Feng Huang 0007, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2021 | Multi-Scale Separable Network for Ultra-High-Definition Video DeblurringabstractAlthough recent research has witnessed a significant progress on the video deblurring task, these methods struggle to reconcile inference efficiency and visual quality simultaneously, especially on ultra-high-definition (UHD) videos (e.g., 4K resolution). To address the problem, we propose a novel deep model for fast and accurate UHD Video Deblurring (UHDVD). The proposed UHDVD is achieved by a separable-patch architecture, which collaborates with a multi-scale integration scheme to achieve a large receptive field without adding the number of generic convolutional layers and kernels. Additionally, we design a residual channel-spatial attention (RCSA) module to improve accuracy and reduce the depth of the network appropriately. The proposed UHDVD is the first real-time deblurring model for 4K videos at 35 fps. To train the proposed model, we build a new dataset comprised of 4K blurry videos and corresponding sharp frames using three different smartphones. Comprehensive experimental results show that our network performs favorably against the state-of-the-art methods on both the 4K dataset and public benchmarks in terms of accuracy, speed, and model size. Senyou Deng, Wenqi Ren, Yanyang Yan, Fenglong Song, Xiaochun Cao |
ICCV | 3 |
| 2021 | SRGAT: Single Image Super-Resolution With Graph Attention NetworkabstractDeep neural networks have demonstrated remarkable reconstruction for single-image super-resolution (SISR). However, most existing CNN-based SISR methods directly learn the relation between low-resolution (LR) and high-resolution (HR) images, neglecting to explore the recurrence of internal patches, hence hindering the representational power of CNNs. In this paper, we propose a novel single image Super-Resolution network based on Graph ATtention network (SRGAT) to make full use of the internal patch-recurrence in a natural image. The proposed model employs a feature mapping block with a recurrent structure to refine low-level representations with high-level information. Especifically, the feature mapping block contains a parallel graph similarity branch and a content branch, where the graph similarity branch aims at exploiting the similarity and symmetry across different image patches in low-resolution feature space and provides additional priors for the content branch to enhance texture details. Specifically, we consider the internal patch-recurrence of an image by constructing a graph network on image feature patches. In this way, the information from neighboring patches can be interacted using graph attention network (GAT) to help it recover additional textures, which complements the textures learned from the content branch. Extensive quantitative and qualitative evaluations on five benchmark datasets demonstrate that the proposed algorithm performs favorably against the state-of-the-art super-resolution methods. Yanyang Yan, Wenqi Ren, Xiaobin Hu, Kun Li 0029, Haifeng Shen, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2019 | Recolored Image Detection via a Deep Discriminative ModelabstractImage recoloring is a technique that can transfer image color or theme and result in an imperceptible change in human eyes. Although image recoloring is one of the most important image manipulation techniques, there is no special method designed for detecting this kind of forgery. In this paper, we propose a trainable end-to-end system for distinguishing recolored images from natural images. The proposed network takes the original image and two derived inputs based on illumination consistency and inter-channel correlation of the original input into consideration and outputs the probability that it is recolored. Our algorithm adopts a convolutional neural network (CNN)- based deep architecture, which consists of three feature extraction blocks and a feature fusion module. To train the deep neural network, we synthesize a data set comprised of recolored images and corresponding ground truth using different recoloring methods. Extensive experimental results on the recolored images generated by various methods show that our proposed network is well generalized and very robust. Yanyang Yan, Wenqi Ren, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | Image Deblurring via Extreme Channels PriorabstractCamera motion introduces motion blur, affecting many computer vision tasks. Dark Channel Prior (DCP) helps the blind deblurring on scenes including natural, face, text, and low-illumination images. However, it has limitations and is less likely to support the kernel estimation while bright pixels dominate the input image. We observe that the bright pixels in the clear images are not likely to be bright after the blur process. Based on this observation, we first illustrate this phenomenon mathematically and define it as the Bright Channel Prior (BCP). Then, we propose a technique for deblurring such images which elevates the performance of existing motion deblurring algorithms. The proposed method takes advantage of both Bright and Dark Channel Prior. This joint prior is named as extreme channels prior and is crucial for achieving efficient restorations by leveraging both the bright and dark information. Extensive experimental results demonstrate that the proposed method is more robust and performs favorably against the state-of-the-art image deblurring methods on both synthesized and natural images. Yanyang Yan, Wenqi Ren, Yuanfang Guo, Rui Wang 0032, Xiaochun Cao |
CVPR | 1 |