VLDB 2026 Research / reviewers in the wild / expert
Ye Deng 0005
dblp:75/7246-5
· DBLP profile ↗
17ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0003-4616-3318ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical frequency adaptation for all-in-one image restoration
Yang Wu 0001, Ye Deng 0005, Siqi Hui, Yuhan Liu 0006, Kangyi Wu, Wenli Huang 0004, Jinjun Wang |
Knowl. Based Syst. | 2 |
| 2026 | DictCR-former: Content-aware dictionary transformer for cloud removal
Wenli Huang 0004, Yang Wu 0001, Sanping Zhou, Xiaomeng Xin, Xiaobo Jia, Ye Deng 0005 |
Pattern Recognit. | 7 |
| 2026 | Frequency-guided generalizable representation learning for cross-domain few-shot learning
Siqi Hui, Sanping Zhou, Ye Deng 0005, Wenli Huang 0004, Yang Wu 0001, Jinjun Wang |
Pattern Recognit. | 3 |
| 2025 | Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense PredictionabstractSufficient cross-task interaction is crucial for success in multi-task dense prediction. However, sufficient interaction often results in high computational complexity, forcing existing methods to face the trade-off between interaction completeness and computational efficiency. To address this limitation, this work proposes a Bidirectional Interaction Mamba (BIM), which incorporates novel scanning mechanisms to adapt the Mamba modeling approach for multi-task dense prediction. On the one hand, we introduce a novel Bidirectional Interaction Scan (BI-Scan) mechanism, which constructs task-specific representations as bidirectional sequences during interaction. By integrating task-first and position-first scanning modes within a unified linear complexity architecture, BI-Scan efficiently preserves critical cross-task information. On the other hand, we employ a Multi-Scale Scan~(MS-Scan) mechanism to achieve multi-granularity scene modeling. This design not only meets the diverse granularity requirements of various tasks but also enhances nuanced cross-task feature interactions. Extensive experiments on two challenging benchmarks, \emph{i.e.}, NYUD-V2 and PASCAL-Context, show the superiority of our BIM vs its state-of-the-art competitors. Mang Cao, Sanping Zhou, Ye Deng 0005, Wenli Huang 0004, Le Wang 0003 |
ICCV | 4 |
| 2025 | Auxiliary Loss Reweighting for Image InpaintingabstractImage inpainting aims to reconstruct missing regions in corrupted images with semantically consistent content. While modern methods employ perceptual and style losses to enhance inpainting quality by supervising deep feature representations, two key challenges persist: (i) existing approaches necessitate time-consuming grid searches to determine optimal loss weights, and (ii) heterogeneous auxiliary loss terms are assigned fixed weights, limiting their adaptive contributions. To address these limitations, we propose a framework featuring dynamically weighted auxiliary losses and an automated weight adaptation mechanism. Specifically, we introduce Tunable Perceptual Loss (TPL) and Tunable Style Loss (TSL), which generalize traditional perceptual and style losses by incorporating tunable weights that independently scale distinct loss components according to their auxiliary potential. These are optimized via our Adaptive Weight Adjustment (AWA) algorithm, which dynamically reweights TPL and TSL during training by prioritizing loss terms that maximally improve inpainting performance. Empirical evaluations on public datasets demonstrate that our framework enhances state-of-the-art inpainting performance while eliminating manual weight tuning. Wenli Huang 0004, Siqi Hui, Ye Deng 0005, Xiaomeng Xin, Yang Wu 0001, Jinjun Wang |
IECON | 3 |
| 2025 | Meta channel masking for cross-domain few-shot image classificationabstractCross-domain Few-shot Learning (CD-FSL) aims to address the challenges of FSL where significant domain gaps exist between source and target image datasets. Unlike many existing CD-FSL methods that utilize an auxiliary target dataset with a few labeled target images to enhance model generalization, our approach directly tackles the limitations imposed by the reliance on source-specific knowledge. We observe that models trained on unbalanced datasets tend to overfit to source-specific features, which, while effective in the source domain, generalize poorly to the target image domain. To address this, we introduce a novel dropout-based framework named Meta Channel Masking (MCM). This framework attenuates the learning of model channels on the source domain by dynamically masking source feature channels during training. In contrast to traditional dropout techniques that manually set masking probabilities based on statistical assumptions about the source data, our MCM framework employs a meta-learning process that automatically adjusts channel mask probabilities. This adjustment is informed by auxiliary target data, effectively minimizing few-shot loss on the auxiliary target dataset and thereby enhancing the model’s generalization capabilities in the target domain. Our extensive experiments across various image classification benchmark datasets demonstrate that our framework outperforms state-of-the-art methods. Siqi Hui, Sanping Zhou, Ye Deng 0005, Pengna Li, Jinjun Wang |
Neurocomputing | 3 |
| 2025 | Dual-View Prompting for Cloud RemovalabstractCloud cover significantly impedes the utilization of remote sensing data, limiting the effectiveness of satellite imagery in critical applications such as environmental monitoring and disaster response. While deep learning methods have advanced cloud removal, existing models predominantly focus on spatial-domain feature discrepancies, often overlooking distinctive spectral difference introduced by clouds. To address this gap, we propose a Dual-view Prompting Network (DVPNet) that integrates spatial and frequency information via prompt learning to generate robust guidance features. The core innovation, the Dual-view Prompting Block (DVPB), operates cascadedly: first, a spatial gating module refines features to capture contextual cues; these features are then transformed into the Fourier domain, where a frequency-gating structure and a learnable spectral prompt further calibrate and enhance representations. The holistically refined dual-view prompt is integrated into the decoder through an efficient windowed cross-attention mechanism, enabling precise cloud removal. Extensive experiments on benchmark datasets demonstrate that DVPNet achieves state-of-the-art performance. This work validates the critical role of frequency-domain modeling in cloud removal and establishes a new spatial-frequency collaborative paradigm for remote sensing image restoration. The code will be made available at https://github.com/huangwenwenlili/DVPNet. Ye Deng 0005, Wenli Huang 0004, Jiang Duan |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Semi-independent Convolution for Image InpaintingabstractIn typical image inpainting tasks, the locations and shapes of damaged or masked areas are often random and irregular. Vanilla convolutions, commonly employed in learning-based inpainting models, treat all spatial features as valid and share parameters across different regions. This approach can struggle with irregular damage patterns, leading to inpainted results that may suffer from color discrepancies and blurriness. In this paper, we introduce a novel operator known as Semi-Independent Convolution (SIConv) to tackle this challenge. The proposed SIConv, on top of the regular convolution with shared weights, also introduces dynamic terms that assign their own independent weights to each part of the image, and the overall computation is formulated as a shared convolution parameter with an additional term to describe the local structure. Qualitative and quantitative experiments demonstrate that our method outperforms the state-of-the-art, yielding clearer, more coherent, and visually convincing inpainting results. Wenli Huang 0004, Ye Deng 0005, Xiaomeng Xin, Jinbao He, Jinjun Wang |
IECON | 2 |
| 2024 | Gradient-guided channel masking for cross-domain few-shot learning
Siqi Hui, Sanping Zhou, Ye Deng 0005, Yang Wu 0001, Jinjun Wang |
Knowl. Based Syst. | 3 |
| 2024 | Sparse self-attention transformer for image inpainting
Wenli Huang 0004, Ye Deng 0005, Siqi Hui, Yang Wu 0001, Sanping Zhou, Jinjun Wang |
Pattern Recognit. | 2 |
| 2024 | Attentive Contextual Attention for Cloud RemovalabstractCloud cover can significantly hinder the use of remote sensing images for Earth observation, prompting urgent advancements in cloud removal technology. Recently, deep learning strategies, especially convolutional neural networks (CNNs) with attention mechanisms, have shown strong potential in restoring cloud-obscured areas. These methods utilize convolution to extract intricate local features and attention mechanisms to gather long-range information, improving the overall comprehension of the scene. However, a common drawback of these approaches is that the resulting images often suffer from blurriness, artifacts, and inconsistencies. This is partly because attention mechanisms apply weights to all features based on generalized similarity scores, which can inadvertently introduce noise and irrelevant details from cloud-covered areas. To overcome this limitation and better capture relevant distant context, we introduce a novel approach named attentive contextual attention (AC-Attention). This method enhances conventional attention mechanisms by dynamically learning data-driven attentive selection scores, enabling it to filter out noise and irrelevant features effectively. By integrating the AC-Attention module into the DSen2-CR cloud removal framework, we significantly improve the model’s ability to capture essential distant information, leading to more effective cloud removal. Our extensive evaluation of various datasets shows that our method outperforms existing ones regarding image reconstruction quality. In addition, we conducted ablation studies by integrating AC-Attention into multiple existing methods and widely used network architectures. These studies demonstrate the effectiveness and adaptability of AC-Attention and reveal its ability to focus on relevant features, thereby improving the overall performance of the networks. The code is available athttps://github.com/huangwenwenlili/ACA-CRNet. Wenli Huang 0004, Ye Deng 0005, Yang Wu 0001, Jinjun Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | CR-former: Single-Image Cloud Removal With Focused Taylor AttentionabstractCloud removal aims to restore high-quality images from cloud-contaminated captures, which is essential in remote sensing applications. Effectively modeling the long-range relationships between image features is key to achieving high-quality cloud-free images. While self-attention mechanisms excel at modeling long-distance relationships, their computational complexity scales quadratically with image resolution, limiting their applicability to high-resolution remote sensing images. Current cloud removal methods have mitigated this issue by restricting the global receptive field to smaller regions or adopting channel attention to model long-range relationships. However, these methods either compromise pixel-level long-range dependencies or lose spatial information, potentially leading to structural inconsistencies in restored images. In this work, we propose the focused Taylor attention (FT-Attention), which captures pixel-level long-range relationships without limiting the spatial extent of attention and achieves the$\mathcal {O}(N)$computational complexity, where N represents the image resolution. Specifically, we utilize Taylor series expansions to reduce the computational complexity of the attention mechanism from$\mathcal {O}(N^{2})$to$\mathcal {O}(N)$, enabling efficient capture of pixel relationships directly in high-resolution images. Additionally, to fully leverage the informative pixel, we develop a new normalization function for the query and key, which produces more distinguishable attention weights, enhancing focus on important features. Building on FT-Attention, we design a U-net style network, termed the CR-former, specifically for cloud removal. Extensive experimental results on representative cloud removal datasets demonstrate the superior performance of our CR-former. The code is available athttps://github.com/wuyang2691/CR-former. Yang Wu 0001, Ye Deng 0005, Sanping Zhou, Yuhan Liu 0006, Wenli Huang 0004, Jinjun Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Context Adaptive Network for Image InpaintingabstractIn a typical image inpainting task, the location and shape of the damaged or masked area is often random and irregular. The vanilla convolutions widely used in learning-based inpainting models treat all spatial features as valid and share parameters across regions, making it difficult for them to cope with those irregular damages, and models tend to produce inpainting results with color discrepancy and blurriness. In this paper, we propose a novel Context Adaptive Network (CANet) to address this issue. The main idea of the proposed CANet is able to generate different weights depending on the miscellaneous input, which may help to complement images with multiple broken forms in a flexible way. Specifically, the proposed CANet has two novel context adaptive modules, namely, the context adaptive block (CAB) and the cross-scale contextual attention (CSCA), which utilize attention mechanisms to cope with diverse content breakdowns. The proposed CAB, during the forward propagation, uses an adaptive term to determine the importance between adaptive term and convolution kernel, so as to dynamically balance features based on the degree of breakage (confidence level or soft mask), and the overall calculation is formulated as a classic convolution implementation with an additional attention term to describe local structure. Besides, the proposed CSCA, not only takes advantage of the contextual attention module, but also considers cross-scale information transfer to generate reasonable features for damaged areas, thus alleviating the inefficiency of the long-range modeling capability of convolutional neural networks. Qualitative and quantitative experiments show that our method performs better than state-of-the-arts, producing clearer, more coherent and visually plausible inpainting results. The code can be found at github.com/dengyecode/CANet_image_inpainting. Ye Deng 0005, Siqi Hui, Sanping Zhou, Wenli Huang 0004, Jinjun Wang |
IEEE Trans. Image Process. | 1 |
| 2022 | Hourglass Attention Network for Image Inpainting
Ye Deng 0005, Siqi Hui, Rongye Meng, Sanping Zhou, Jinjun Wang |
ECCV (18) | 1 |
| 2022 | T-former: An Efficient Transformer for Image InpaintingabstractBenefiting from powerful convolutional neural networks (CNNs), learning-based image inpainting methods have made significant breakthroughs over the years. However, some nature of CNNs (e.g. local prior, spatially shared parameters) limit the performance in the face of broken images with diverse and complex forms. Recently, a class of attention-based network architectures, called transformer, has shown significant performance on natural language processing fields and high-level vision tasks. Compared with CNNs, attention operators are better at long-range modeling and have dynamic weights, but their computational complexity is quadratic in spatial resolution, and thus less suitable for applications involving higher resolution images, such as image inpainting. In this paper, we design a novel attention linearly related to the resolution according to Taylor expansion. And based on this attention, a network called T-former is designed for image inpainting. Experiments on several benchmark datasets demonstrate that our proposed method achieves state-of-the-art accuracy while maintaining a relatively low number of parameters and computational complexity. Ye Deng 0005, Siqi Hui, Sanping Zhou, Deyu Meng, Jinjun Wang |
ACM Multimedia | 1 |
| 2021 | Learning Contextual Transformer Network for Image InpaintingabstractFully Convolutional Networks with attention modules have been proven effective for learning-based image inpainting. While many existing approaches could produce visually reasonable results, the generated images often show blurry textures or distorted structures around corrupted areas. The main reason is due to the fact that convolutional neural networks have limited capacity for modeling contextual information with long range dependencies. Although the attention mechanism can alleviate this problem to some extent, existing attention modules tend to emphasize similarities between the corrupted and the uncorrupted regions while ignoring the dependencies from within each of them. Hence, this paper proposes the Contextual Transformer Network (CTN) which not only learns relationships between the corrupted and the uncorrupted regions but also exploits their respective internal closeness. Besides, instead of a fully convolutional network, in our CTN, we stack several transformer blocks to replace convolution layers to better model the long range dependencies. Finally, by dividing the image into patches of different sizes, we propose a multi-scale multi-head attention module to better model the affinity among various image regions. Experiments on several benchmark datasets demonstrate superior performance by our proposed approach. Ye Deng 0005, Siqi Hui, Sanping Zhou, Deyu Meng, Jinjun Wang |
ACM Multimedia | 1 |
| 2020 | Image Inpainting Using Parallel NetworkabstractDue to the lack of contextual information and the difficulty to directly learn the distribution of a complete image, existing image inpainting methods always use a two-stages approach to make plausible prediction for missing pixels in a coarse-to-fine manner. In this paper, we propose a novel inpainting method with two parallel pipelines. The first pipeline is a standard image completion path that takes the corrupted image as input and outputs the predicted complete image. The second pipeline exists only during the training phase that inputs a complementary image of the corrupted one and still outputs the same complete image. The two pipelines operate simultaneously, and they share identical encoder and most parameters in the decoder. Furthermore, inspired by VAE, random Gaussian noise are added to the features not only to improve the robustness of the model but also to enable generating diverse and plausible results. We evaluated our model on several public datasets and demonstrated that the proposed method outperforms several state-of-the-arts approaches. Ye Deng 0005, Jinjun Wang |
ICIP | 1 |