EDBT 2026 Demo / reviewers in the wild / expert
Yuning Cui 0001
dblp:275/7800-1
· DBLP profile ↗
40ranked-venue papers
27as first author
38since 2021 · last 2027
0000-0002-1279-5539ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 20 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 9 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Training-free multi-scale super-resolution with diffusion models
Aiping Zhang, Yuning Cui 0001, Jianhou Gan, Wenqi Ren |
Expert Syst. Appl. | 2 |
| 2026 | A frequency-guided denoising framework based on convolutional transformer for electrocardiogram signals
Mingyue Cui, Yewei Gan, Jiepeng Chen, Yanchong Xie, Daosong Hu, Yuning Cui 0001, Kai Huang 0001 |
Eng. Appl. Artif. Intell. | 7 |
| 2026 | DRFIR: A dimensionality reduction framework for all-in-one image restoration in spatial and frequency domains
Yuning Cui 0001, Leah Strand, Huilin Yin, Alois C. Knoll |
Expert Syst. Appl. | 2 |
| 2026 | Focal Modulation for Image RestorationabstractAbstract Image restoration aims to recover a sharp image from its degraded counterpart by removing degradations ( e.g., noise, haze, and blur) and restoring missing details. It plays an important role in many fields, such as remote sensing and medical imaging. How to effectively capture critical information parsimoniously for high-quality reconstruction has long been a pivotal problem in this domain. This study aims to develop an efficient and effective focal modulation scheme for image restoration. Inspired by the fact that different regions in a corrupted image always undergo degradations in various degrees, we introduce a dual-domain selection mechanism to emphasize crucial information for restoration, such as edge signals and hard regions. Moreover, a channel modulation module is developed to facilitate channel interactions by exploring the utility of the Fourier transform in channel dimensions. In addition, we split high-resolution features to insert multi-scale receptive fields into the network, improving efficiency and performance. Incorporating these designs into a U-shaped convolutional backbone, the network achieves state-of-the-art performance on 13 different datasets for five general image restoration tasks, including dehazing, desnowing, deraining, motion/defocus deblurring, and low-light enhancement. To further demonstrate the effectiveness of our focal modulation strategy, we apply it to the all-in-one image restoration setting, and the obtained model performs favorably against state-of-the-art all-in-one algorithms. Moreover, our module extends effectively to tasks such as composite degradation, medical imaging, and ultra-high-definition image restoration. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
Int. J. Comput. Vis. | 1 |
| 2026 | Visual-in-Visual: A Unified and Efficient Baseline for Image RestorationabstractRecent years have witnessed remarkable progress in image restoration, yet achieving both high performance and efficiency remains a persistent challenge. To address this issue, we present VIVNet, a strong and efficient unified baseline designed to balance accuracy and practicality. Drawing inspiration from the high efficiency of the human visual system, VIVNet embeds a biologically inspired micro visual module into each block of a macro U-shaped vision architecture. This module mimics key perceptual processes such as retinal encoding, lateral inhibition, and high-order processing by combining lightweight depth-wise convolutions for multi-receptive-field feature extraction, a similarity-aware weighting mechanism to emphasize informative signals, and high-order interactions implemented via iterative element-wise multiplication to capture complex dependencies. This design enhances the model's representational capacity while maintaining computational efficiency. Unlike most existing methods that are limited to narrow task settings, we evaluate VIVNet across a wide range of scenarios, including general, all-in-one, and composite degradation tasks, as well as ultra-high-definition (UHD), underwater, medical, and remote sensing datasets. Extensive experiments show that VIVNet delivers competitive performance with high efficiency. Yuning Cui 0001, Wenqi Ren, Boxin Shi, Alois C. Knoll |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | StarIR: Convolutional Image Restoration With Spatial-Frequency FusionabstractVision Transformer (ViT) has shown impressive performance in image restoration due to its ability to capture a large receptive field. However, its complexity grows quadratically with input resolution, limiting its applicability for high-resolution images. In contrast, Convolutional Neural Networks (CNNs) are computationally efficient but are constrained by their inherently local receptive fields, which limit their ability to capture long-range pixel relationships. To address these challenges, we propose StarIR, which possesses the efficiency of CNNs while also capturing a large receptive field, similar to Transformers. StarIR incorporates two key innovations: 1) a dual-domain representation learning framework, with one branch processing spatial details and the other focusing on mesoscale interactions in the frequency domain; and 2) a high-dimensional feature fusion mechanism, the Star operation, which fuses information from both domains through element-wise multiplication, thereby enhancing representational capacity without increasing network width and depth. Our Star operation is followed by a channel attention unit to facilitate global feature modeling and enhance channel-wise interactions. Building on our straightforward yet powerful design principles, StarIR achieves state-of-the-art performance across 21 datasets covering six single-degradation image restoration tasks. Furthermore, our model performs favorably against leading algorithms in two all-in-one settings and demonstrates robustness on two composite-degradation datasets. In addition, StarIR extends well to several domain-specific applications, including ultra-high-definition (UHD) imaging, remote sensing, medical imaging, and underwater image enhancement. Yuning Cui 0001, Syed Waqas Zamir, Ming-Hsuan Yang 0001, Alois C. Knoll, Fahad Shahbaz Khan, Salman Khan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Ultra-High-Definition Image Restoration via High-Frequency Enhanced TransformerabstractTransformer-based architectures exhibit substantial promise in the realm of ultra-high-definition (UHD) image restoration (IR). Nevertheless, they encounter significant challenges in maintaining high-frequency (HF) details, which are crucial for the reconstruction of texture. Conventional methods tackle computational complexity by significantly reducing the resolution (by a factor of 4 to 8). Moreover, the majority of high-frequency components are eliminated due to the inherent characteristics of self-attention mechanisms, as these mechanisms tend to naturally suppress high-frequency elements during non-local feature integration. This paper proposes a dual-branch transformer architecture that synergistically combines native-resolution HF preservation with efficient contextual modeling, named HiFormer. The high-resolution branch utilizes a directionally-sensitive large-kernel decomposition to effectively address anisotropic degradations with fewer parameters and applies depthwise separable convolutions for localized high-frequency (HF) information extraction. Concurrently, the low-resolution branch assimilates these localized HF elements using adaptive channel modulation to offset spectral losses induced by the inherent smoothing effect of self-attention. Comprehensive experiments across numerous UHD image restoration tasks reveal that our approach surpasses current leading methods in both quantitative metrics and qualitative analysis. The code is available at https://github.com/5chen/HiFormer. Chen Wu 0006, Zhuoran Zheng, Weidong Jiang, Yuning Cui 0001, Jingyuan Xia |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Efficient All-in-One Image Restoration With Adaptive Frequency EnhancementabstractAll-in-one image restoration has recently attracted considerable attention for its ability to address multiple degradation types within a single, unified framework. However, existing methods often incur substantial computational overhead, especially when incorporating explicit degradation priors via complex auxiliary branches, hindering their practical deployment. In this paper, we propose AdaptIR, an efficient all-in-one image restoration network equipped with adaptive frequency enhancement. Recognizing that different degradations impact distinct frequency subbands and exhibit spatially varying restoration demands, we design an Adaptive Frequency Enhancement Module (AFEM) that couples frequency learning with adaptive convolutions to better capture frequency-aware information. Specifically, AFEM learns pixel-wise adaptive attention weights to modulate the spectra of dynamic convolutions, enabling spatially adaptive and content-aware restoration. Furthermore, we introduce a lightweight backbone featuring a Receptive Field Expansion Module (RFEM), which enlarges the receptive field of a convolutional U-shaped architecture by convolving wavelet-transform coefficients. By integrating the plug-and-play AFEM into the bottleneck of the baseline model, AdaptIR achieves state-of-the-art performance on all-in-one image restoration tasks involving multiple degradations, while maintaining high computational efficiency. Moreover, the proposed model can be readily extended to single-degradation tasks (e.g., dehazing, desnowing, and deraining) and domain-specific applications, including ultra-high-definition (UHD), medical, and remote sensing image restoration. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
IEEE Trans. Image Process. | 1 |
| 2026 | CDIR: LoRA-Inspired Attention for Efficient Composite Degradation Image RestorationabstractSpecialized image restoration methods have been extensively explored, each targeting a specific type of degradation. However, real-world images often suffer from composite degradations, prompting growing interest in unified restoration approaches. While recent unified models have shown promising results, many are hindered by high computational complexity, limiting their deployment in resource-constrained settings. Motivated by the parameter-efficient design of Low-Rank Adaptation (LoRA), we propose an efficient attention module specifically designed for composite degradation image restoration. The proposed method adopts a dual-branch architecture, where one branch processes features at full resolution, and the other operates with reduced spatial and channel dimensions to improve efficiency. To better adapt to diverse degradation patterns, the latter branch is further divided into two sub-branches, each incorporating dynamic operations guided by local and contextual priors. These context priors are iteratively updated within each module, drawing inspiration from feedback mechanisms in reinforcement learning, thereby enabling the model to effectively perceive and handle multiple degradation types within a unified structure. Additionally, we introduce a multi-scale feed-forward network to further enhance both performance and computational efficiency. Extensive experiments on two composite degradation benchmarks demonstrate that our proposed network, CDIR, achieves state-of-the-art performance with significantly reduced complexity and fast inference speed. In addition, CDIR shows strong adaptability to various task-specific image restoration scenarios, such as dehazing, desnowing, and deraining. It also performs robustly on domain-specific applications such as ultra-high-definition (UHD), remote sensing, and medical image restoration, highlighting its versatility and practical applicability. Yuning Cui 0001, Wenqi Ren, Boxin Shi, Jianhou Gan, Alois C. Knoll |
IEEE Trans. Image Process. | 1 |
| 2026 | Disentangle to Fuse: Toward Content Preservation and Cross-Modality Consistency for Multi-Modality Image FusionabstractMulti-modal image fusion (MMIF) aims to integrate complementary information from heterogeneous sensor modalities. However, substantial cross-modality discrepancies hinder joint scene representation and lead to semantic degradation in the fused output. To address this limitation, we propose C2MFuse, a novel framework designed to preserve content while ensuring cross-modality consistency. To the best of our knowledge, this is the first MMIF approach to explicitly disentangle style and content representations across modalities for image fusion. C2MFuse introduces a content-preserving style normalization mechanism that suppresses modality-specific variations while maintaining the underlying scene structure. The normalized features are then progressively aggregated to enhance fine-grained details and improve content completeness. In light of the lack of ground truth and the inherent ambiguity of the fused distribution, we further align the fused representation with a well-defined source modality, thereby enhancing semantic consistency and reducing distributional uncertainty. Additionally, we introduce an adaptive consistency loss with learnable transformation, which provides dynamic, modality-aware supervision by enforcing global consistency across heterogeneous inputs. Extensive experiments on five datasets across three representative MMIF tasks demonstrate that C2MFuse achieves efficient and high-quality fusion, surpasses existing methods, and generalizes effectively to downstream visual applications. Xinran Qin, Yuning Cui 0001, Shangquan Sun, Ruoyu Chen 0001, Wenqi Ren, Alois C. Knoll, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2025 | Prior-guided Hierarchical Harmonization Network for Efficient Image DehazingabstractImage dehazing is a crucial task that involves the enhancement of degraded images to recover their sharpness and textures. While vision Transformers have exhibited impressive results in diverse dehazing tasks, their quadratic complexity and lack of dehazing priors pose significant drawbacks for real-world applications. In this paper, guided by triple priors, Bright Channel Prior (BCP), Dark Channel Prior (DCP), and Histogram Equalization (HE), we propose a Prior-guided Hierarchical Harmonization Network (PGHHNet) for image dehazing. PGHNet is built upon the UNet-like architecture with an efficient encoder and decoder, consisting of two module types: (1) Prior aggregation module that injects BCP/DCP and selects diverse contexts with gating attention. (2) Feature harmonization modules that subtract low-frequency components from spatial and channel aspects and learn more informative feature distributions to equalize the feature maps. Inspired by observing the sparsity of BCP/DCP and the histogram equalization, we harmonize the deep features using a histogram equation-guided module and further leverage BCP/DCP to guide spatial attention through a sandwich module as the bottleneck. Comprehensive experiments demonstrate that our model efficiently attains the highest level of performance among existing methods across four different datasets for image dehazing tasks. Xiongfei Su, Yuning Cui 0001, Yulun Zhang 0001, Zheng Chen 0014, Zongliang Wu, Zedong Wang, Yuanlong Zhang, Xin Yuan 0002 |
AAAI | 3 |
| 2025 | AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and ModulationabstractIn the image acquisition process, various forms of degradation, including noise, blur, haze, and rain, are frequently introduced. These degradations typically arise from the inherent limitations of cameras or unfavorable ambient conditions. To recover clean images from their degraded versions, numerous specialized restoration methods have been developed, each targeting a specific type of degradation. Recently, all-in-one algorithms have garnered significant attention by addressing different types of degradations within a single model without requiring the prior information of the input degradation type. However, most methods purely operate in the spatial domain and do not delve into the distinct frequency variations inherent to different degradation types. To address this gap, we propose an adaptive all-in-one image restoration network based on frequency mining and modulation. Our approach is motivated by the observation that different degradation types impact the image content on different frequency subbands, thereby requiring different treatments for each restoration task. Specifically, we first mine low- and high-frequency information from the input features, guided by the adaptively decoupled spectra of the degraded image. The extracted features are then modulated by a bidirectional operator to facilitate interactions between different frequency components. Finally, the modulated features are merged into the original input for a progressively guided restoration. With this approach, the model achieves adaptive reconstruction by accentuating the informative frequency subbands according to different input degradations. Extensive experiments demonstrate that the proposed method, AdaIR, achieves state-of-the-art performance on different image restoration tasks, including image denoising, dehazing, deraining, motion deblurring, and low-light image enhancement. The code is available at https://github.com/c-yn/AdaIR. Yuning Cui 0001, Syed Waqas Zamir, Salman Khan 0001, Alois C. Knoll, Mubarak Shah, Fahad Shahbaz Khan |
ICLR | 1 |
| 2025 | DepthCNE: Contrastive Neighbor Embedding in Self-Supervised Learning for Point CloudsabstractContrastive learning, as one of the self-supervised learning paradigms, has demonstrated substantial effectiveness in boosting various downstream tasks. However, its inherent lack of explicit control over the embedding space structure often results in suboptimal performance and limited feature visualization. To address these challenges, we propose DepthCNE, an innovative self-supervised learning framework for point clouds that integrates contrastive learning with dimensionality reduction. Specifically, DepthCNE utilizes a 3D point-based encoder and a projection head, consistent with conventional contrastive learning designs. In addition, we introduce a dimensionality reduction head, which projects high-dimensional features into a two-dimensional space to improve the model’s representation capability. Extensive experiments demonstrate the competitive performance of DepthCNE on both synthetic and real-world datasets across various downstream tasks. Moreover, our method exhibits promising performance in unsupervised visualization. The codes are available at https://github.com/MingyuLiu1/DepthCNE.git. Ekim Yurtsever, Yuning Cui 0001, Alois C. Knoll |
IJCNN | 4 |
| 2025 | MambaSFLNet: A Mamba-based Model for Low-Light Image Enhancement with Spatial and Frequency FeaturesabstractLow-light image enhancement (LLIE) aims to enhance the illumination of images that are captured under dark conditions, which is critical for various applications in dim environments, such as robotics and autonomous driving. Existing convolutional neural network (CNN)-based methods usually struggle to capture long-range dependencies, while transformer-based methods, despite their effectiveness, are resource-consuming. Besides, the frequency domain includes important lightness degradation information. To this end, we propose a Mamba-based framework called MambaSFLNet to effectively address LLIE by integrating spatial and frequency features. Our approach utilizes the Visual State Space Module to establish relationships across different regions of the input image while maintaining low model complexity. Furthermore, The spatial module not only balances illumination distribution but also suppresses noise and artifacts during enhancement. In addition, the frequency module enhances image contrast and sharpness by leveraging frequency-domain information. Extensive experiments on nine widely used benchmarks demonstrate that our approach achieves superior performance and exhibits strong generalization capabilities compared to existing methods. The codes are available at https://github.com/MingyuLiu1/MambaSFLNet.git Yuning Cui 0001, Leah Strand, Xingcheng Zhou, Alois C. Knoll |
IROS | 2 |
| 2025 | Bio-Inspired Image RestorationabstractImage restoration aims to recover sharp, high-quality images from degraded, low-quality inputs. Existing methods have progressively advanced from task-specific designs to general architectures, all-in-one frameworks, and composite degradation handling. Despite these advances, computational efficiency remains a critical factor for practical deployment. In this work, we present BioIR, an efficient and universal image restoration framework inspired by the human visual system. Specifically, we design two bio-inspired modules, Peripheral-to-Foveal (P2F) and Foveal-to-Peripheral (F2P), to emulate the perceptual processes of human vision, with a particular focus on the functional interplay between foveal and peripheral pathways. P2F delivers large-field contextual signals to foveal regions based on pixel-to-region affinity, while F2P propagates fine-grained spatial details through a static-to-dynamic two-stage integration strategy. Leveraging the biologically motivated design, BioIR achieves state-of-the-art performance across three representative image restoration settings: single-degradation, all-in-one, and composite degradation. Moreover, BioIR maintains high computational efficiency and fast inference speed, making it highly suitable for real-world applications. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
NeurIPS | 1 |
| 2025 | NDFormer: A Mixed-Scale Transformer with Enhanced Nonlinearity for Nighttime Image Deraining
Zhirui Liu, Shangquan Sun, Yuning Cui 0001, Dehong Kong, Wenqi Ren, Kin-Man Lam 0001 |
PRCV (9) | 3 |
| 2025 | EENet: An effective and efficient network for single image dehazing
Yuning Cui 0001, Chaopeng Li, Wenqi Ren, Alois C. Knoll |
Pattern Recognit. | 1 |
| 2025 | LIEDNet: A Lightweight Network for Low-Light Enhancement and DeblurringabstractImages captured at nighttime often face challenges such as low light and blur, primarily caused by dim environments and the frequent use of long exposure. Existing methods either handle the two types of degradations independently or rely on carefully designed priors generated by complex mechanisms, resulting in poor generalization ability and high model complexity. To address these challenges, we propose an end-to-end framework named LIEDNet to efficiently and effectively restore high-quality images on both real-world and synthetic data. Specifically, the introduced LIEDNet consists of three essential components: the Visual State Space Module (VSSM), the Local Feature Module (LFM), and the Dual Gated-Dconv Feedforward Network (DGDFFN). The integration of VSSM and LFM enables the model to capture both global and local features while maintaining low computational overhead. Additionally, the DGDFFN improves image fidelity by extracting multi-scale structural information. Extensive experiments on real-world and synthetic datasets demonstrate the superior performance of LIEDNet in restoring low-light, blurry images. The code is available athttps://github.com/MingyuLiu1/LIEDNethttps://github.com/MingyuLiu1/LIEDNet. Yuning Cui 0001, Wenqi Ren, Juxiang Zhou, Alois C. Knoll |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Exploring the Potential of Pooling Techniques for Universal Image RestorationabstractImage restoration involves recovering a clean image from its degraded counterpart. In recent years, we have witnessed a paradigm shift from convolutional neural networks to Transformers, which have quadratic complexity with respect to the input size. Instead of designing more complex modules based on recent techniques, this paper presents an efficient and effective mechanism for image restoration by exploring the potential of ubiquitous pooling techniques. We leverage different pooling operators as tools for implicit dual-domain representation learning. Specifically, the average and max pooling can be used as extractors for implicit low- and high-frequency signals, respectively. Then, we utilize lightweight learnable parameters to modulate the resulting frequency components. Furthermore, the intermediate high-frequency features can serve as attention maps to highlight the spatial edge information. Our pooling module is built by incorporating the aforementioned dual-domain modulation across multiple scales and various shapes. We demonstrate the effectiveness of our module in single-degradation, composite-degradation, and all-in-one image restoration tasks. Extensive experimental results show that the resulting network achieves state-of-the-art performance on 15 datasets for five single-degradation and two composite-degradation image restoration tasks by deploying our module. Moreover, our method can be extended to all-in-one scenarios and performs favorably against state-of-the-art all-in-one algorithms under two settings. The code is available at https://github.com/c-yn/PoolNet. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
IEEE Trans. Image Process. | 1 |
| 2025 | Frequency-Prompted Image Restoration to Enhance Perception in Intelligent Transportation SystemsabstractHigh perceptual image quality is crucial for intelligent transportation systems (ITS), including autonomous vehicles, digital twins, and surveillance infrastructure. However, images captured in adverse weather conditions or dynamic environments often suffer from various visibility degradations. To address this issue, image restoration aims to recover missing details and remove distortions from degraded observations, thereby enhancing the usability of visual data in intelligent transportation applications. Inspired by the success of prompt learning in natural language processing, recent studies have explored prompt-based approaches for various image restoration tasks. However, most of these methods operate in the spatial domain. Given the importance of frequency learning in image restoration, particularly in reducing the spectral discrepancy between degraded and sharp image pairs, this study investigates the use of frequency prompts through a plug-and-play mechanism consisting of a prompt generation module and a prompt integration module. Specifically, the prompt generation module encodes frequency information by aggregating pre-defined learnable parameters, guided by the implicitly decomposed spectra of the input features. The learned prompts are then integrated into the feature spectra via dual-dimensional attention, dynamically guiding the reconstruction process and enabling more effective frequency-aware learning. To validate the effectiveness of the proposed plug-in module, we integrate it into both CNN-based and Transformer-based backbones. Extensive experiments demonstrate that the CNN-based variant achieves state-of-the-art performance on 15 datasets across five representative image restoration tasks. Furthermore, it generalizes well to composite degradation scenarios. The Transformer-based model performs competitively with state-of-the-art methods under two all-in-one image restoration settings. Finally, the effectiveness of our models in enhancing perception for ITS is empirically verified. Yuning Cui 0001, Xiongfei Su, Alois C. Knoll |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Enhancing Perception for Autonomous Vehicles: A Multi-Scale Feature Modulation Network for Image RestorationabstractAccurate environmental perception is essential for the effective operation of autonomous vehicles. However, visual images captured in dynamic environments or adverse weather conditions often suffer from various degradations. Image restoration focuses on reconstructing clear and sharp images by eliminating undesired degradations from corrupted inputs. These degradations typically vary in size and severity, making it crucial to employ robust multi-scale representation learning techniques. In this paper, we propose Multi-Scale Feature Modulation (MSFM), a novel deep convolutional architecture for image restoration. MSFM modulates multi-scale features in both frequency and spatial domains to make features sharper and closer to that of clean images. Specifically, our multi-scale frequency attention module transforms features into multiple scales and then modulates each scale in the implicit frequency domain using pooling and attention. Moreover, we develop a multi-scale spatial modulation module to refine pixels with the guidance of local features. The proposed frequency and spatial modules enable MSFM to better handle degradations of different sizes. Experimental results demonstrate that MSFM achieves state-of-the-art performance on 12 datasets for a range of image restoration tasks, i.e., image dehazing, image defocus/motion deblurring, and image desnowing. Furthermore, the restored images significantly improve the environmental perception of autonomous vehicles. Yuning Cui 0001, Jianyong Zhu, Alois C. Knoll |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Modumer: Modulating Transformer for Image RestorationabstractImage restoration aims to recover clean images from degraded counterparts. While Transformer-based approaches have achieved significant advancements in this field, they are limited by high complexity and their inability to capture omni-range dependencies, hindering their overall performance. In this work, we develop Modumer for effective and efficient image restoration by revisiting the Transformer block and modulation design, which processes input through a convolutional block and projection layers and fuses features via elementwise multiplication. Specifically, within each unit of Modumer, we integrate the cascaded modulation design with the downsampled Transformer block to build the attention layers, enabling omni-kernel modulation and mapping inputs into high-dimensional feature spaces. Moreover, we introduce a bioinspired parameter-sharing mechanism to attention layers, which not only enhances efficiency but also improves performance. In addition, a dual-domain feed-forward network (DFFN) strengthens the representational power of the model. Extensive experimental evaluations demonstrate that the proposed Modumer achieves state-of-the-art performance across ten datasets in five single-degradation image restoration tasks, including image motion deblurring, deraining, dehazing, desnowing, and low-light enhancement. Moreover, the model exhibits strong generalization capabilities in all-in-one image restoration tasks. Additionally, it demonstrates competitive performance in composite-degradation image restoration. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Omni-Kernel Network for Image RestorationabstractImage restoration aims to reconstruct a high-quality image from a degraded low-quality observation. Recently, Transformer models have achieved promising performance on image restoration tasks due to their powerful ability to model long-range dependencies. However, the quadratically growing complexity with respect to the input size makes them inapplicable to practical applications. In this paper, we develop an efficient convolutional network for image restoration by enhancing multi-scale representation learning. To this end, we propose an omni-kernel module that consists of three branches, i.e., global, large, and local branches, to learn global-to-local feature representations efficiently. Specifically, the global branch achieves a global perceptive field via the dual-domain channel attention and frequency-gated mechanism. Furthermore, to provide multi-grained receptive fields, the large branch is formulated via different shapes of depth-wise convolutions with unusually large kernel sizes. Moreover, we complement local information using a point-wise depth-wise convolution. Finally, the proposed network, dubbed OKNet, is established by inserting the omni-kernel module into the bottleneck position for efficiency. Extensive experiments demonstrate that our network achieves state-of-the-art performance on 11 benchmark datasets for three representative image restoration tasks, including image dehazing, image desnowing, and image defocus deblurring. The code is available at https://github.com/c-yn/OKNet. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
AAAI | 1 |
| 2024 | Omnidirectional Image Super-resolution via Bi-projection FusionabstractWith the rapid development of virtual reality, omnidirectional images (ODIs) have attracted much attention from both the industrial community and academia. However, due to storage and transmission limitations, the resolution of current ODIs is often insufficient to provide an immersive virtual reality experience. Previous approaches address this issue using conventional 2D super-resolution techniques on equirectangular projection without exploiting the unique geometric properties of ODIs. In particular, the equirectangular projection (ERP) provides a complete field-of-view but introduces significant distortion, while the cubemap projection (CMP) can reduce distortion yet has a limited field-of-view. In this paper, we present a novel Bi-Projection Omnidirectional Image Super-Resolution (BPOSR) network to take advantage of the geometric properties of the above two projections. Then, we design two tailored attention methods for these projections: Horizontal Striped Transformer Block (HSTB) for ERP and Perspective Shift Transformer Block (PSTB) for CMP. Furthermore, we propose a fusion module to make these projections complement each other. Extensive experiments demonstrate that BPOSR achieves state-of-the-art performance on omnidirectional image super-resolution. The code is available at https://github.com/W-JG/BPOSR. Yuning Cui 0001, Yawen Li 0001, Wenqi Ren, Xiaochun Cao |
AAAI | 2 |
| 2024 | Hybrid Frequency Modulation Network for Image Restoration
Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
IJCAI | 1 |
| 2024 | RFFNet: Towards Robust and Flexible Fusion for Low-Light Image DenoisingabstractLow-light environments will introduce high-intensity noise into images. Containing fine details with reduced noise, near-infrared/flash images can serve as guidance to facilitate noise removal. However, existing fusion-based methods fail to effectively suppress artifacts caused by inconsistency between guidance/noisy image pairs and do not fully excavate the useful information contained in guidance images. In this paper, we propose a robust and flexible fusion network (RFFNet) for low-light image denoising. Specifically, we present a multi-scale inconsistency calibration module to address inconsistency before fusion by first mapping the guidance features to multi-scale spaces and calibrating them with the aid of pre-denoising features in a coarse-to-fine manner. Furthermore, we develop a dual-domain adaptive fusion module to adaptively extract useful high-/low-frequency signals from the guidance features and then highlight the informative frequencies. Extensive experimental results demonstrate that our method achieves state-of-the-art performance on NIR-guided RGB image denoising and flash-guided no-flash image denoising. Yuning Cui 0001, Yawen Li 0001, Yaping Ruan, Ben Zhu, Wenqi Ren |
ACM Multimedia | 2 |
| 2024 | Dual-domain strip attention for image restoration
Yuning Cui 0001, Alois C. Knoll |
Neural Networks | 1 |
| 2024 | Image Restoration via Frequency SelectionabstractImage restoration aims to reconstruct the latent sharp image from its corrupted counterpart. Besides dealing with this long-standing task in the spatial domain, a few approaches seek solutions in the frequency domain by considering the large discrepancy between spectra of sharp/degraded image pairs. However, these algorithms commonly utilize transformation tools, e.g., wavelet transform, to split features into several frequency parts, which is not flexible enough to select the most informative frequency component to recover. In this paper, we exploit a multi-branch and content-aware module to decompose features into separate frequency subbands dynamically and locally, and then accentuate the useful ones via channel-wise attention weights. In addition, to handle large-scale degradation blurs, we propose an extremely simple decoupling and modulation module to enlarge the receptive field via global and window-based average pooling. Furthermore, we merge the paradigm of multi-stage networks into a single U-shaped network to pursue multi-scale receptive fields and improve efficiency. Finally, integrating the above designs into a convolutional backbone, the proposed Frequency Selection Network (FSNet) performs favorably against state-of-the-art algorithms on 20 different benchmark datasets for 6 representative image restoration tasks, including single-image defocus deblurring, image dehazing, image motion deblurring, image desnowing, image deraining, and image denoising. Yuning Cui 0001, Wenqi Ren, Xiaochun Cao, Alois C. Knoll |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Revitalizing Convolutional Network for Image RestorationabstractImage restoration aims to reconstruct a high-quality image from its corrupted version, playing essential roles in many scenarios. Recent years have witnessed a paradigm shift in image restoration from convolutional neural networks (CNNs) to Transformer-based models due to their powerful ability to model long-range pixel interactions. In this paper, we explore the potential of CNNs for image restoration and show that the proposed simple convolutional network architecture, termed ConvIR, can perform on par with or better than the Transformer counterparts. By re-examing the characteristics of advanced image restoration algorithms, we discover several key factors leading to the performance improvement of restoration models. This motivates us to develop a novel network for image restoration based on cheap convolution operators. Comprehensive experiments demonstrate that our ConvIR delivers state-of-the-art performance with low computation complexity among 20 benchmark datasets on five representative image restoration tasks, including image dehazing, image motion/defocus deblurring, image deraining, and image desnowing. Yuning Cui 0001, Wenqi Ren, Xiaochun Cao, Alois C. Knoll |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Omni-Kernel Modulation for Universal Image RestorationabstractImage restoration is the process of recovering a clean image from a degraded observation. In order to achieve this, it is essential to refine features at multiple scales. This paper develops an effective omni-kernel modulation module to enhance multi-scale representation learning for image restoration. The module consists of three branches, namely global, large, and local branches, which are designed to learn global-to-local feature representations efficiently. Specifically, the global branch achieves a global perceptive field via the dual-domain channel attention and frequency-gated mechanism. Furthermore, to provide multi-grained receptive fields, the large branch is formulated using different shapes of depth-wise convolutions with unusually large kernel sizes. Moreover, we complement local information with a point-wise depth-wise convolution. Finally, we demonstrate the effectiveness of our omni-kernel modulation module in two cases: general image restoration and all-in-one image restoration tasks. Incorporating our method into a convolutional backbone results in a model that achieves state-of-the-art performance on the 15 datasets for three representative image restoration tasks, including image dehazing, desnowing, and defocus deblurring. Moreover, by integrating our module into a pure Transformer-based backbone, the model demonstrates competitive performance against state-of-the-art algorithms in two all-in-one image restoration settings: the three-task and five-task settings. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Enhancing Local-Global Representation Learning for Image RestorationabstractVision systems are the core element in industrial systems, such as intelligent transportation systems and inspection robots. However, undesired degradations caused by bad weather or low-end devices reduce the visibility of images. Image restoration aims to reconstruct a sharp image from a degraded counterpart and plays an important role in industrial systems. Recent transformer-based architectures leverage the self-attention unit and convolutions to model long-range dependencies and local connectivity, respectively, achieving promising performance for image restoration. However, these methods have quadratic complexity with respect to the input size. In addition, convolution operators are ineffective enough to recover the local details. This article presents a joint local and global representation learning framework for image restoration, called LoGoNet. Specifically, to enhance global contexts, we excavate the potential of pooling techniques to refine large-scale feature maps, which help handle large-size degradations. Furthermore, we develop a novel module to emphasize local edges with the implicit Laplace operator. With these designs, the proposed LoGoNet produces powerful feature representations for image restoration, which is helpful for perceiving objects of different sizes in industrial systems. Extensive experiments demonstrate that LoGoNet achieves state-of-the-art performance on nine datasets for four image restoration tasks: image defocus/motion deblurring, image dehazing, and image desnowing. Yuning Cui 0001, Alois C. Knoll |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Dual-Domain Attention for Image DeblurringabstractAs a long-standing and challenging task, image deblurring aims to reconstruct the latent sharp image from its degraded counterpart. In this study, to bridge the gaps between degraded/sharp image pairs in the spatial and frequency domains simultaneously, we develop the dual-domain attention mechanism for image deblurring. Self-attention is widely used in vision tasks, however, due to the quadratic complexity, it is not applicable to image deblurring with high-resolution images. To alleviate this issue, we propose a novel spatial attention module by implementing self-attention in the style of dynamic group convolution for integrating information from the local region, enhancing the representation learning capability and reducing computational burden. Regarding frequency domain learning, many frequency-based deblurring approaches either treat the spectrum as a whole or decompose frequency components in a complicated manner. In this work, we devise a frequency attention module to compactly decouple the spectrum into distinct frequency parts and accentuate the informative part with extremely lightweight learnable parameters. Finally, we incorporate attention modules into a U-shaped network. Extensive comparisons with prior arts on the common benchmarks show that our model, named Dual-domain Attention Network (DDANet), obtains comparable results with a significantly improved inference speed. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
AAAI | 1 |
| 2023 | Focal Network for Image RestorationabstractImage restoration aims to reconstruct a sharp image from its degraded counterpart, which plays an important role in many fields. Recently, Transformer models have achieved promising performance on various image restoration tasks. However, their quadratic complexity remains an intractable issue for practical applications. The aim of this study is to develop an efficient and effective framework for image restoration. Inspired by the fact that different regions in a corrupted image always undergo degradations in various degrees, we propose to focus more on the important areas for reconstruction. To this end, we introduce a dual-domain selection mechanism to emphasize crucial information for restoration, such as edge signals and hard regions. In addition, we split high-resolution features to insert multi-scale receptive fields into the network, which improves both efficiency and performance. Finally, the proposed network, dubbed FocalNet, is built by incorporating these designs into a U-shaped backbone. Extensive experiments demonstrate that our model achieves state-of-the-art performance on ten datasets for three tasks, including single-image defocus deblurring, image dehazing, and image desnowing. Our code is available at https://github.com/c-yn/FocalNet. Yuning Cui 0001, Wenqi Ren, Xiaochun Cao, Alois C. Knoll |
ICCV | 1 |
| 2023 | Selective Frequency Network for Image Restoration
Yuning Cui 0001, Zhenshan Bing, Wenqi Ren, Xinwei Gao, Xiaochun Cao, Kai Huang 0001, Alois C. Knoll |
ICLR | 1 |
| 2023 | IRNeXt: Rethinking Convolutional Network Design for Image RestorationabstractWe present IRNeXt, a simple yet effective convolutional network architecture for image restoration. Recently, Transformer models have dominated the field of image restoration due to the powerful ability of modeling long-range pixels interactions. In this paper, we excavate the potential of the convolutional neural network (CNN) and show that our CNN-based model can receive comparable or better performance than Transformer models with low computation overhead on several image restoration tasks. By re-examining the characteristics possessed by advanced image restoration algorithms, we discover several key factors leading to the performance improvement of restoration models. This motivates us to develop a novel network for image restoration based on cheap convolution operators. Comprehensive experiments demonstrate that IRNeXt delivers state-of-the-art performance among numerous datasets on a range of image restoration tasks with low computational complexity, including image dehazing, single-image defocus/motion deblurring, image deraining, and image desnowing. https://github.com/c-yn/IRNeXt. Yuning Cui 0001, Wenqi Ren, Sining Yang, Xiaochun Cao, Alois C. Knoll |
ICML | 1 |
| 2023 | Strip Attention for Image RestorationabstractAs a long-standing task, image restoration aims to recover the latent sharp image from its degraded counterpart. In recent years, owing to the strong ability of self-attention in capturing long-range dependencies, Transformer based methods have achieved promising performance on multifarious image restoration tasks. However, the canonical self-attention leads to quadratic complexity with respect to input size, hindering its further applications in image restoration. In this paper, we propose a Strip Attention Network (SANet) for image restoration to integrate information in a more efficient and effective manner. Specifically, a strip attention unit is proposed to harvest the contextual information for each pixel from its adjacent pixels in the same row or column. By employing this operation in different directions, each location can perceive information from an expanded region. Furthermore, we apply various receptive fields in different feature groups to enhance representation learning. Incorporating these designs into a U-shaped backbone, our SANet performs favorably against state-of-the-art algorithms on several image restoration tasks. The code is available at https://github.com/c-yn/SANet. Yuning Cui 0001, Luoxi Jing, Alois C. Knoll |
IJCAI | 1 |
| 2023 | Exploring the potential of channel interactions for image restoration
Yuning Cui 0001, Alois C. Knoll |
Knowl. Based Syst. | 1 |
| 2021 | Two-stage Local Spatio-temporal Event Filter based on Adaptive ThresholdsabstractRecently, the emerging event camera shows great potential for robotics and AR/VR application thanks to its advantages including low latency, high dynamic range, low power consumption, etc. However, the output of event cameras usually has a large amount of noise which will affect unlocking their potentials and obtaining wider applications. In this paper, we present a two-stage local spatio-temporal event filter (LSTEF) based on adaptive thresholds. Considering the spatio-temporal constraints of generated event streams, a local sliding window is adopted for the noise candidate selection stage and noise filtering stage. An adaptive thresholding mechanism is also introduced into the filter in order to improve the generalization performance. Corresponding experimental evaluations are performed on the public datasets to prove the efficiency of the proposed filter. Results show that the presented LSTEF can successfully achieve the event denoising, and at the same time, effectively preserve useful scene information. Luoxi Jing, Jun Luo 0011, Dian-xi Shi, Ruihao Li 0001, Huachi Xu, Yuning Cui 0001 |
IJCNN | 7 |
| 2020 | IDNet: A Single-Shot Object Detector Based on Feature FusionabstractThis paper proposes a novel single shot network for object detection. The proposed network, termed IDNet, explores the strategies of the feature fusion to alleviate the scale variation problem in object detection. IDNet mainly consists of two feature fusion modules: an indirect feature fusion module (IF) and a direct feature fusion module (DF). The IF shares long-range dependencies within pyramidal layers and based on these information, IDNet learns to emphasize informative regions and suppress the less useful ones on each layer. The DF is a feature fusion strategy based on modified lateral connection inspired by feature pyramid networks (FPN). It utilizes the averaging operation to reduce the change of feature maps' order of magnitude during fusing features to further improve the performance for detecting small instances. Comprehensive experiments are performed and the results indicate the effectiveness of IDNet, which reaches 80.3 mAP on PASCAL VOC 2007 benchmark. Yuning Cui 0001, Dian-xi Shi, Yongjun Zhang 0006, Qianchong Sun |
ICTAI | 1 |
| 2020 | Selective Feature Network for Object DetectionabstractScale variation is one of the important challenges in object detection. Many state-of-the-art objectors tackle this problem by utilizing the feature pyramids. However, the current methods of producing feature pyramids are still inefficient to integrate the semantic information from other layers. In this work, our motivation is to build a feature pyramid efficiently with the selected contextual feature by integrating the informative features and suppressing the useless ones. To achieve this goal, we propose a novel single-stage detection network termed Selective Feature Network(SFNet) which consists of a semantic-enhanced module and a selective feature module. The semantic-enhanced module improves the semantics of basic pyramids via a lightweight architecture. In conjunction with that, a selective feature module is employed to combine features across different channels and scales by attention mechanism. The resulting contextual feature is then injected into the pyramidal features. Comprehensive experiments are performed on PASCAL VOC and MS COCO datasets. Results demonstrate that, with a VGG16 based SFNet, our approach obtains significant improvements over the competitors without losing real-time processing speed. Yuning Cui 0001, Dian-xi Shi, Yongjun Zhang 0006, Qianchong Sun |
IJCNN | 1 |