Jian-Nan Su

dblp:321/7047 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
14since 2021 · last 2025
0000-0002-8206-3675ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 IniRetinex: Rethinking Retinex-type Low-Light Image Enhancer via Initialization Perspective
abstract
Retinex-based methods have become a general approach for solving low-light image enhancement (LLIE). However, traditional methods require post-processing of illumination (e.g., gamma correction), which lacks adaptability and disrupts the illumination structure. Retinex-based deep networks typically follow a ‘decomposition-adjustment-exposure control’ process, which is redundant and lacks robustness. One major issue is the inaccuracy in estimating and decomposing the initial illumination. Accurate initial illumination can prevent further post-processing instability. We propose IniRetinex, rethinking the Retinex-based LLIE method from the perspective of initialization. By using neural networks to provide reasonable initial illumination and solving for smooth illumination through optimization, higher performance LLIE is achieved. We construct a two-layer convolutional neural network to capture the low-frequency structure of the image, adaptively compensating for classical initial illumination and avoiding additional post-processing. The network requires no pre-training and can be implemented in an unsupervised manner with just a few iterations, making it highly efficient. Additionally, we propose a new illumination optimization strategy by introducing an additional proximal penalty term, improving illumination in areas with varying levels and enhancing image details. Extensive experiments on various low-light image datasets demonstrate that our method achieves state-of-the-art (SOTA) results on multiple benchmarks, offering higher stability and inference efficiency compared to current advanced methods.
Zishu Yao, Guang-Yong Chen, Jian-Nan Su, Min Gan
AAAI4
2025 Clinically Robust Polyp Segmentation: Enhanced Generalization and Perturbation Resistance
abstract
Colonoscopy is vital for detecting colorectal polyps, which are closely linked to colorectal cancer. Accurate segmentation of polyps in colonoscopic images is essential for diagnosis and surgical planning but is challenging due to variability in polyp size, shape, and unclear boundaries. The Segment Anything Model (SAM) has shown promise in polyp segmentation but relies heavily on user-provided prompts and involves a large number of parameters, limiting its practicality in clinical settings. To address the limitations of SAM in clinical practice, we introduced Low Rank and Perturbation Segment Anything Model (LP-SAM) to improve segmentation accuracy and generalization ability while reducing the parameter count and complexity of user input. LP-SAM showed enhanced generalization and a lightweight design, making it more suitable for clinical applications where precise user input may not always be feasible. Comparative evaluations demonstrate that LP-SAM outperforms state-of-the-art methods on datasets such as CVC-ColonDB, CVC-300, and ETIS.
Shanchuan Wang, Jian-Nan Su, Min Gan
ICASSP3
2025 IGSENet: A Clinically Robust Polyp Segmentation Method via Interactive Fusion and Guided-Selective Enhancement
Shanchuan Wang, Tianzong Nie, Jian-Nan Su, Min Gan
PRCV (13)3
2025 Knowledge-prompted intracranial hemorrhage segmentation on brain computed tomography
Tianzong Nie, Feiyan Chen, Jian-Nan Su, Guang-Yong Chen, Min Gan
Expert Syst. Appl.3
2025 LPFSformer: Location Prior Guided Frequency and Spatial Interactive Learning for Nighttime Flare Removal
abstract
When capturing images under strong light sources at night, intense lens flare artifacts often appear, significantly degrading visual quality and impacting downstream computer vision tasks. Although transformer-based methods have achieved remarkable results in nighttime flare removal, they fail to adequately distinguish between flare and non-flare regions. This unified processing overlooks the unique characteristics of these regions, leading to suboptimal performance and unsatisfactory results in real-world scenarios. To address this critical issue, we propose a novel approach incorporating Location Prior Guidance (LPG) and a specialized flare removal model, LPFSformer. LPG is designed to accurately learn the location of flares within an image and effectively capture the associated glow effects. By employing Location Prior Injection (LPI), our method directs the model’s focus towards flare regions through the interaction of frequency and spatial domains. Additionally, to enhance the recovery of high-frequency textures and capture finer local details, we designed a Global Hybrid Feature Compensator (GHFC). GHFC aggregates different expert structures, leveraging the diverse receptive fields and CNN operations of each expert to effectively utilize a broader range of features during the flare removal process. Extensive experiments demonstrate that our LPFSformer achieves state-of-the-art flare removal performance compared to existing methods. Our code and a pre-trained LPFSformer have been uploaded to GitHub for validation.
Guang-Yong Chen, Jian-Nan Su, Min Gan, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.4
2025 Real-World Image Reflection Removal: An Ultra-High-Definition Dataset and an Efficient Baseline
abstract
Reflection removal is a crucial issue in image reconstruction, especially for high-definition images. Removing undesirable reflections can greatly enhance the performance of various visual systems, such as medical imaging, autonomous driving, and security surveillance. However, the resolution of existing reflection removal datasets is not high and the training data heavily relies on synthetic data, which hampers the performance of reflection removal methods and restricts the development of effective techniques tailored for high-definition images. Therefore, this paper introduces a new dataset, Real-world Reflection Removal in 4K (RR4K). This novel dataset, with its large capacity and high resolution of$6000\times 4000$pixels, represents a significant advancement in the field, ensuring a realistic and high quality benchmark. Furthermore, building upon the dataset, we propose an efficient method for single-image reflection removal, optimized for high-definition processing. This method employs the U-Net architecture, enhanced with large kernel distillation and scale-aware features, enabling it to effectively handle complex reflection scenarios while reducing computational demands. Comprehensive testing on the RR4K dataset and existing low-resolution datasets has demonstrated the method’s superior efficiency and effectiveness. We believe that our constructed RR4K dataset can better evaluate and design algorithms for removing undesirable reflection from real-world high-definition images. Our dataset and code are available athttps://github.com/jengchauwei/RR4K.
Guang-Yong Chen, Chao-Wei Zheng, Jian-Nan Su, Min Gan, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.4
2024 Weighted Adaptive Clustering Attention for Efficient Image Super-Resolution
abstract
The non-local attention has attracted widespread attention from researchers in enhancing the ability of deep single-image super-resolution methods to mine self-similarity information. However, non-local attention often faces problems of high computational complexity and inaccurate calculation of cross-correlation of deep features when capturing long-distance information. To solve these problems, we propose an innovative and efficient Weighted Adaptive Clustering Attention (WACA). Specifically, WACA consists of Sparse Adaptive Clustering Attention (SACA) and Weighted Residual Attention (WRA). SACA significantly reduces the noise signal in the process of feature correlation calculation with the help of asymmetric local sensitive hashing, and reduces the computational cost from quadratic to asymptotically linear with the sequence length. In addition, our designed WRA utilizes the high correlation characteristics of the self-similarity matrix between self-attention layers, and further significantly reduces the computational consumption of associated non-local information by co-optimizing the self-similarity matrix. To verify the effectiveness of WACA, we introduce corresponding modules on a residual backbone and construct a framework named Weighted Adaptive Clustering Network (WACN). Experimental results demonstrate that WACN has competitive performance in both quantitative and qualitative evaluations.
Yu-Bin Liu, Jian-Nan Su, Guang-Yong Chen, Yi-Gang Zhao
IJCNN2
2024 Decoupled Non-Local Attention for Single Image Super-Resolution
abstract
Self-similarity-based deep Single Image Super-Resolution (SISR) methods have gained popularity in recent years, especially with the integration of Non-Local Attention (NLA) in deep SISR. However, NLA suffers from the drawback of mixing relevant and irrelevant features, as it computes the response of each query by aggregating information from all non-local features. In this work, we propose a novel approach to exploit self-similarity more effectively using FlyHash, which is inspired by the fruit fly olfactory circuit and has exhibited outstanding performance in approximate similarity search. By limiting the association area of non-local features with FlyHash, we develop the Decoupled Non-Local Attention (DNLA) method, which tackles the difficulties of modeling a large amount of irrelevant non-local features while considerably lowering the computational complexity from quadratic to nearly linear. We verify the effectiveness of our DNLA approach with comprehensive ablation studies to show its capability of capturing nonlocal information for deep SISR. Furthermore, we build a deep Decoupled Non-Local Attention Network (DNLAN) with DNLA, which attains excellent results in both objective evaluation and subjective perception for SISR.
Yi-Gang Zhao, Jian-Nan Su, Guang-Yong Chen, Yu-Bin Liu
IJCNN2
2024 FISTA acceleration inspired network design for underwater image enhancement
Bing-Yuan Chen, Jian-Nan Su, Guang-Yong Chen, Min Gan
J. Vis. Commun. Image Represent.2
2024 Revealing the Dark Side of Non-Local Attention in Single Image Super-Resolution
abstract
Single Image Super-Resolution (SISR) aims to reconstruct a high-resolution image from its corresponding low-resolution input. A common technique to enhance the reconstruction quality is Non-Local Attention (NLA), which leverages self-similar texture patterns in images. However, we have made a novel finding that challenges the prevailing wisdom. Our research reveals that NLA can be detrimental to SISR and even produce severely distorted textures. For example, when dealing with severely degrade textures, NLA may generate unrealistic results due to the inconsistency of non-local texture patterns. This problem is overlooked by existing works, which only measure the average reconstruction quality of the whole image, without considering the potential risks of using NLA. To address this issue, we propose a new perspective for evaluating the reconstruction quality of NLA, by focusing on the sub-pixel level that matches the pixel-wise fusion manner of NLA. From this perspective, we provide the approximate reconstruction performance upper bound of NLA, which guides us to design a concise yet effective Texture-Fidelity Strategy (TFS) to mitigate the degradation caused by NLA. Moreover, the proposed TFS can be conveniently integrated into existing NLA-based SISR models as a general building block. Based on the TFS, we develop a Deep Texture-Fidelity Network (DTFN), which achieves state-of-the-art performance for SISR. Our code and a pre-trained DTFN are available on GitHub†for verification.
Jian-Nan Su, Min Gan, Guang-Yong Chen, Wenzhong Guo, C. L. Philip Chen
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Dynamic Degradation Intensity Estimation for Adaptive Blind Super-Resolution: A Novel Approach and Benchmark Dataset
abstract
Blind Super-Resolution (BlindSR) aims to reconstruct high-resolution (HR) images from low-resolution (LR) images without prior knowledge of the image degradation process. This is a challenging problem in real-world applications, where the degradation can be complex and unknown. Recent unsupervised learning-based BlindSR methods can estimate the image degradation in an unsupervised manner, but they suffer from limited adaptability to different types and intensities of degradation. They tend to capture the average level of degradation across all training samples, resulting in over-smoothing or over-sharpening effects for some images. As a result, the final reconstruction may exhibit the mean effect. Moreover, existing synthetic datasets do not reflect the real-world degradation scenarios, making it difficult to evaluate the performance of BlindSR methods. To address these issues, we propose a novel Degradation Intensity Estimation Module (DIEM) method, which can estimate the pixel-level degradation information of the input image more specifically and use it to guide image reconstruction. Furthermore, we construct a benchmark dataset under real scenarios, which is closer to the real-world BlindSR problem than existing synthetic datasets, and can provide a more reasonable evaluation of BlindSR methods. Extensive experimental results demonstrate that our DIEM-guided BlindSR method can achieve state-of-the-art image reconstruction results. Our code and pre-trained models have been uploaded to GitHub† for validation.
Guang-Yong Chen, Wu-Ding Weng, Jian-Nan Su, Min Gan, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.3
2024 Unsupervised Degradation Aware and Representation for Real-World Remote Sensing Image Super-Resolution
abstract
Blind super-resolution (BlindSR) has recently attracted attention in the field of remote sensing. Due to the lack of paired data, most works assume that the acquired remote sensing images are high-resolution (HR) and use predefined degradation models to synthesize low-resolution (LR) images for training and evaluation. However, these acquired remote sensing images are often degraded by various factors, which still require super-resolution reconstruction to meet practical needs. Using them as ground truth images will limit the model’s ability to restore fine details, resulting in blurry and noisy reconstructions. To overcome these limitations, we propose an unsupervised degradation-aware network which transforms natural images into the degraded domain as real-world remote sensing images. It uses natural images containing rich texture information as a reference for fine-grained restoration of the network, enabling the network to produce clearer reconstructions. Furthermore, we discovered the remarkable capability of patch-wise discriminator to perceive the degradation type of different regions within the acquired remote sensing image. Inspired by this finding, we design a novel degradation representation module (DRM) that can estimate the degradation information from LR images and guide the network to perform adaptive restoration. Comprehensive experimental results demonstrate that our proposed unsupervised blind super-resolution framework (UDASR) achieves state-of-the-art restoration performance. Our code and pre-trained models have been uploaded to GitHub† for validation.
Wenzhong Guo, Wu-Ding Weng, Guang-Yong Chen, Jian-Nan Su, Min Gan, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.4
2024 High-Similarity-Pass Attention for Single Image Super-Resolution
abstract
Recent developments in the field of non-local attention (NLA) have led to a renewed interest in self-similarity-based single image super-resolution (SISR). Researchers usually use the NLA to explore non-local self-similarity (NSS) in SISR and achieve satisfactory reconstruction results. However, a surprising phenomenon that the reconstruction performance of the standard NLA is similar to that of the NLA with randomly selected regions prompted us to revisit NLA. In this paper, we first analyzed the attention map of the standard NLA from different perspectives and discovered that the resulting probability distribution always has full support for every local feature, which implies a statistical waste of assigning values to irrelevant non-local features, especially for SISR which needs to model long-range dependence with a large number of redundant non-local features. Based on these findings, we introduced a concise yet effective soft thresholding operation to obtain high-similarity-pass attention (HSPA), which is beneficial for generating a more compact and interpretable distribution. Furthermore, we derived some key properties of the soft thresholding operation that enable training our HSPA in an end-to-end manner. The HSPA can be integrated into existing deep SISR models as an efficient general building block. In addition, to demonstrate the effectiveness of the HSPA, we constructed a deep high-similarity-pass attention network (HSPAN) by integrating a few HSPAs in a simple backbone. Extensive experimental results demonstrate that HSPAN outperforms state-of-the-art approaches on both quantitative and qualitative evaluations. Our code and a pre-trained model were uploaded to GitHub (https://github.com/laoyangui/HSPAN) for validation.
Jian-Nan Su, Min Gan, Guang-Yong Chen, Wenzhong Guo, C. L. Philip Chen
IEEE Trans. Image Process.1
2023 Global Learnable Attention for Single Image Super-Resolution
abstract
Self-similarity is valuable to the exploration of non-local textures in single image super-resolution (SISR). Researchers usually assume that the importance of non-local textures is positively related to their similarity scores. In this paper, we surprisingly found that when repairing severely damaged query textures, some non-local textures with low-similarity which are closer to the target can provide more accurate and richer details than the high-similarity ones. In these cases, low-similarity does not mean inferior but is usually caused by different scales or orientations. Utilizing this finding, we proposed a Global Learnable Attention (GLA) to adaptively modify similarity scores of non-local textures during training instead of only using a fixed similarity scoring function such as the dot product. The proposed GLA can explore non-local textures with low-similarity but more accurate details to repair severely damaged textures. Furthermore, we propose to adopt Super-Bit Locality-Sensitive Hashing (SB-LSH) as a preprocessing method for our GLA. With the SB-LSH, the computational complexity of our GLA is reduced from quadratic to asymptotic linear with respect to the image size. In addition, the proposed GLA can be integrated into existing deep SISR models as an efficient general building block. Based on the GLA, we constructed a Deep Learnable Similarity Network (DLSN), which achieves state-of-the-art performance for SISR tasks of different degradation types (e.g., blur and noise). Our code and a pre-trained DLSN have been uploaded to GitHub†for validation.
Jian-Nan Su, Min Gan, Guang-Yong Chen, Jia-Li Yin, C. L. Philip Chen
IEEE Trans. Pattern Anal. Mach. Intell.1