Shuigen Wang

dblp:142/0117 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0003-3598-0443ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 Thermal-Physics Guided Infrared Image Super-Resolution with Dynamic High-Frequency Amplification
abstract
The practical deployment of infrared imaging is hindered by its inherent output of low-resolution (LR) images. While the super-resolution (SR) technique is a promising remedy, we discover two major challenges concerning infrared image SR: preserving accurate thermal distributions, which are fundamental to infrared imaging, and addressing the ambiguity of high-frequency elements compared to visible images. To tackle these issues, we propose ThesIS, a tailored framework that utilizes Thermal-Physics guidance and dynamic high-frequency amplification for Infrared image Super-resolution to produce high-resolution (HR) images with accurate physical properties and delicate visual details. Specifically, Thermal Regularization is introduced to reconstruct the accurate thermal radiation distribution via the introduced Infrared Radiation Intensity Alignment Loss, mitigating the adverse effects of complex degradations while conducting initial upscaling. Additionally, we design a guidance mechanism to counter the randomness of the diffusion model, further refining the preservation of physical information. The proposed Dynamic High-Frequency Amplification effectively strengthens the ambiguous high-frequency information present in infrared images, leading to improved texture details and superior visual quality. Extensive experiments demonstrate that ThesIS successfully recovers accurate thermal information while delivering visually satisfying results with state-of-the-art performance. Furthermore, we introduce the InfraredSR dataset, which comprises 39,833 images at a resolution of 512 × 512, hoping to advance research in this field.
Mingxuan Zhou, Yirui Shen, Yutang Zhang, Shuigen Wang
AAAI6
2026 EGHCNet: edge-guided hierarchical cost optimization for stereo matching
Guangfen Wei, Tuo Zhou, Shuigen Wang, Yongwen Zhang
Multim. Syst.4
2026 Removing Multiple Hybrid Adverse Weather in Video via a Unified Model
abstract
Videos captured under real-world adverse weather conditions typically suffer from uncertain hybrid weather efforts. However, existing algorithms can only remove one type of weather degradation at a time and deal with different weather conditions with separate models, thus may fail to handle real-world stochastic hybrid scenarios. Besides, the model training is also infeasible due to the lack of paired video data to characterize the coexistence of multiple weather. To ameliorate the aforementioned issue, we propose a novel unified model, dubbed UniWRV, to remove multiple heterogeneous video weather degradations in an all-in-one fashion. Specifically, to tackle degenerate spatial feature heterogeneity, we propose a tailored weather prior guided module that queries exclusive priors for different instances as prompts to steer spatial feature characterization. To tackle degenerate temporal feature heterogeneity, we propose a dynamic routing aggregation module that can automatically select optimal fusion paths for different instances to dynamically integrate temporal features. Furthermore, we propose a real-world adaptation training scheme that leverages CLIP priors to provide semantic supervision for unlabeled real-world weather-degraded videos, thereby enabling the model to better cope with the diverse and complex real-world weather conditions. Additionally, we managed to construct a new synthetic video dataset, termed HWVideo, for learning and benchmarking multiple hybrid adverse weather removal, which contains 15 hybrid weather conditions with a total of 1500 adverse-weather/clean paired video clips. Real-world hybrid weather videos are also collected to facilitate model generalizability. Comprehensive experiments demonstrate that our UniWRV exhibits robust and superior adaptation capability in multiple heterogeneous degradations learning scenarios, including various generic video restoration tasks beyond weather removal.
Yecong Wan, Ming-Wen Shao, Yuanshuo Cheng, Shuigen Wang
IEEE Trans. Circuits Syst. Video Technol.5
2025 H2D-Net: High-resolution Guided Hierarchical Discriminative Network for Infrared Small Target Detection
abstract
Following detection-by-segmentation paradigm, U-net and its variants have recently achieved competitive performance in infrared small target detection (IRSTD) benchmarks. However, when the targets only occupy few pixels, the U-shape deep network tends to favor global background patterns over local appearance of targets in the feature encoding stage, and indiscriminately amplifies false feature response in the decoder. Such representation bias and error accumulation degrade identification capability when target-similar distractors occur. Here, by introducing high-resolution cues, we advocate our High-resolution Guided Hierarchical Discriminative Network (H2D-Net), where High Resolution Guidance (HRG) module and Holistic Distractor Filter (HDF) module are devised to tackle the aforementioned issues. Specifically, an extra hierarchical network with fixed scale embedding, i.e., high-resolution cues, is parallelly assigned to rectify the representation bias of the U-shape network via a group of the HRG modules, which facilitate bidirectional interaction between the fine-grained spatial details and multi-scale representations. Furthermore, the refining HDF module is embedded into the bottleneck between the encoder and decoder for the purpose of interrupting feedforward propagation of the false feature response. Extensive experiments demonstrate that the H2D-Net significantly enhances the detection performance of infrared small targets, particularly in reducing false alarms, outperforming state-of-the-art methods across multiple real-world infrared datasets.
Xiangpan Fan, Dongshun Cui, Shuigen Wang
IJCNN6
2025 PromptSeg: Prompt for Universal Remote Sensing Semantic Segmentation
Jie Zhang 0133, Ming-Wen Shao, Lingzhuang Meng, Xiangyong Cao, Shuigen Wang
IEEE Trans. Geosci. Remote. Sens.5
2024 Multi-Domain Multi-Scale Diffusion Model for Low-Light Image Enhancement
abstract
Diffusion models have achieved remarkable progress in low-light image enhancement. However, there remain two practical limitations: (1) existing methods mainly focus on the spatial domain for the diffusion process, while neglecting the essential features in the frequency domain; (2) conventional patch-based sampling strategy inevitably leads to severe checkerboard artifacts due to the uneven overlapping. To address these limitations in one go, we propose a Multi-Domain Multi-Scale (MDMS) diffusion model for low-light image enhancement. In particular, we introduce a spatial-frequency fusion module to seamlessly integrates spatial and frequency information. By leveraging the Multi-Domain Learning (MDL) paradigm, our proposed model is endowed with the capability to adaptively facilitate noise distribution learning, thereby enhancing the quality of the generated images. Meanwhile, we propose a Multi-Scale Sampling (MSS) strategy that follows a divide-ensemble manner by merging the restored patches under different resolutions. Such a multi-scale learning paradigm explicitly derives patch information from different granularities, thus leading to smoother boundaries. Furthermore, we empirically adopt the Bright Channel Prior (BCP) which indicates natural statistical regularity as an additional restoration guidance. Experimental results on LOL and LOLv2 datasets demonstrate that our method achieves state-of-the-art performance for the low-light image enhancement task. Codes are available at https://github.com/Oliiveralien/MDMS.
Kai Shang 0001, Ming-Wen Shao, Chao Wang 0102, Yuanshuo Cheng, Shuigen Wang
AAAI5
2024 Subjective Quality Assessment of Thermal Infrared Images
abstract
Thermal infrared images (TIIs) can be distorted by multiple factors, resulting in noise, low contrast, limited dynamic range, and fuzziness, which greatly impede their usefulness. It is crucial to evaluate the quality of TIIs. Unfortunately, there have been very few attempts to study this problem. In this study, we collected 1,000 authentically distorted TIIs using thermal infrared acquisition equipment and conducted strict subjective experiments to obtain a thermal infrared image quality assessment (IQA) database. Each image’s quality score was obtained under strict scoring rules. Finally, we investigated the feasibility of several no-reference (NR) IQA methods in quality assessment of TIIs. We found that existing NR-IQA methods achieve ordinary performance in such a task, and there is an urgent need to develop a specific IQA methods for TIIs. The findings together with the constructed database are expected to pave the way for the development of more advanced IQA methods for further development of this field.
Guanghui Yue 0001, Jinxia Zhang, Zhaofei Xu, Shuigen Wang, Tianwei Zhou, Yuanhao Gong, Wei Zhou 0021
ICIP5
2024 When guided diffusion model meets zero-shot image super-resolution
Huan Liu 0012, Ming-Wen Shao, Kai Shang 0001, Yuanjian Qiao 0001, Shuigen Wang
Eng. Appl. Artif. Intell.5
2024 Colorectal endoscopic image enhancement via unsupervised deep learning
Guanghui Yue 0001, Lvyin Duan, Jingfeng Du, Weiqing Yan, Shuigen Wang, Tianfu Wang 0001
Multim. Tools Appl.6
2024 Boundary-Aware Spatial and Frequency Dual-Domain Transformer for Remote Sensing Urban Images Segmentation
abstract
Semantic segmentation of remote sensing (RS) images refers to labeling each pixel with a class to identify objects or land cover types. Existing mainstream spatial-domain semantic segmentation methods are mainly categorized into convolutional neural network (CNN)-based and vision transformer (ViT)-based approaches. The former excels at capturing local features, while the latter is adept at extracting global features. Several recent approaches consider combining CNN and ViT to efficiently capture local and global features. However, these approaches still struggle to capture complete features of the RS images, resulting in inaccurate segmentation. To address this issue, we introduce the fast Fourier transform (FFT), which transforms images into the frequency domain for feature extraction, acquiring the image-size receptive field that can complement spatial-domain methods. Based on this, we propose a boundary-aware spatial and frequency dual-domain transformer, termed dual-domain transformer. Specifically, our dual-domain transformer incorporates a dual-domain mixer (DualM), where the spatial-domain branch combines depthwise convolution and the attention mechanism to extract local and global features effectively, while the frequency-domain branch uses FFT to extract image-size features. The two branches complement each other, enabling a more comprehensive feature extraction of RS images. Meanwhile, a boundary-guided training strategy utilizing a boundary-aware module (BAM) is devised to constrain the model extract and predict boundary detail texture, which is an auxiliary task. In addition, the decoder incorporates a scale-feature fusion module (SFM) for adaptive information fusion between the encoder and decoder. Comprehensive experiments on the Zeebrugge and ISPRS datasets, including Vaihingen and Potsdam, showcase that the dual-domain transformer significantly outperforms state-of-the-art (SOTA) methods.
Jie Zhang 0133, Ming-Wen Shao, Yecong Wan, Lingzhuang Meng, Xiangyong Cao, Shuigen Wang
IEEE Trans. Geosci. Remote. Sens.6
2024 MHW-GAN: Multidiscriminator Hierarchical Wavelet Generative Adversarial Network for Multimodal Image Fusion
abstract
Image fusion technology aims to obtain a comprehensive image containing a specific target or detailed information by fusing data of different modalities. However, many deep learning-based algorithms consider edge texture information through loss functions instead of specifically constructing network modules. The influence of the middle layer features is ignored, which leads to the loss of detailed information between layers. In this article, we propose a multidiscriminator hierarchical wavelet generative adversarial network (MHW-GAN) for multimodal image fusion. First, we construct a hierarchical wavelet fusion (HWF) module as the generator of MHW-GAN to fuse feature information at different levels and scales, which avoids information loss in the middle layers of different modalities. Second, we design an edge perception module (EPM) to integrate edge information from different modalities to avoid the loss of edge information. Third, we leverage the adversarial learning relationship between the generator and three discriminators for constraining the generation of fusion images. The generator aims to generate a fusion image to fool the three discriminators, while the three discriminators aim to distinguish the fusion image and edge fusion image from two source images and the joint edge image, respectively. The final fusion image contains both intensity information and structure information via adversarial learning. Experiments on public and self-collected four types of multimodal image datasets show that the proposed algorithm is superior to the previous algorithms in terms of both subjective and objective evaluation.
Cheng Zhao 0003, Peng Yang 0011, Feng Zhou 0003, Guanghui Yue 0001, Shuigen Wang, Huisi Wu, Guoliang Chen 0005, Tianfu Wang 0001, Bai Ying Lei
IEEE Trans. Neural Networks Learn. Syst.5
2023 On the Difficulty of Unpaired Infrared-to-Visible Video Translation: Fine-Grained Content-Rich Patches Transfer
abstract
Explicit visible videos can provide sufficient visual information and facilitate vision applications. Unfortunately, the image sensors of visible cameras are sensitive to light conditions like darkness or overexposure. To make up for this, recently, infrared sensors capable of stable imaging have received increasing attention in autonomous driving and monitoring. However, most prosperous vision models are still trained on massive clear visible data, facing huge visual gaps when deploying to infrared imaging scenarios. In such cases, transferring the infrared video to a distinct visible one with fine-grained semantic patterns is a worthwhile endeavor. Previous works improve the outputs by equally optimizing each patch on the translated visible results, which is unfair for enhancing the details on content-rich patches due to the long-tail effect of pixel distribution. Here we propose a novel CPTrans framework to tackle the challenge via balancing gradients of different patches, achieving the fine-grained Content-rich Patches Transferring. Specifically, the content-aware optimization module encourages model optimization along gradients of target patches, ensuring the improvement of visual details. Additionally, the content-aware temporal normalization module enforces the generator to be robust to the motions of target patches. Moreover, we extend the existing dataset InfraredCity to more challenging adverse weather conditions (rain and snow), dubbed as InfraredCity-Adverse11The code and dataset are available at https://github.com/BIT-DA/12V-Processing, Extensive experiments show that the proposed CPTrans achieves state-of-the-art performance under diverse scenes while requiring less training time than competitive methods.
Zhenjie Yu, Shuang Li 0008, Yirui Shen, Chi Harold Liu, Shuigen Wang
CVPR5
2023 Style Transfer Meets Super-Resolution: Advancing Unpaired Infrared-to-Visible Image Translation with Detail Enhancement
abstract
The problem of unpaired infrared-to-visible image translation has gained significant attention due to its ability to generate visible images with color information from low-detail grayscale infrared inputs. However, current methodologies often depend on conventional style transfer techniques, which constrain the spatial resolution of the visible output to be equivalent to that of the input infrared image. The fixed generation pattern results in blurry generated results when translating low-resolution infrared inputs, and utilizing high-resolution infrared inputs as a solution necessitates greater computational resources. This spurs us to investigate the challenging unpaired image translation from low-resolution infrared inputs to high-resolution visible outputs, with the ultimate goal of enhancing image details while reducing computational costs. Therefore, we propose a unified framework that integrates the super-resolution process into our unpaired infrared-to-visible image transfer, yielding realistic and high-resolution results. Specifically, we propose the Detail Consistency Loss to establish a connection between the two aforementioned modules, thereby enhancing the quality of visual detail in style transfer results through the super-resolution module. Furthermore, our Texture Perceptual Loss is designed to ensure that the generator generates high-quality visual details accurately and reliably. Experimental results indicate that our method outperforms other comparative approaches when utilizing low-resolution infrared inputs. Remarkably, our approach even surpasses techniques that use high-resolution infrared inputs to generate visible images. Last but equally important, we propose a new and challenging dataset, dubbed as InfraredCity-HD, which comprises 512X512 resolution images, to advance research on high-resolution infrared-related fields.
Yirui Shen, Jingxuan Kang, Shuang Li 0008, Zhenjie Yu, Shuigen Wang
ACM Multimedia5
2022 ROMA: Cross-Domain Region Similarity Matching for Unpaired Nighttime Infrared to Daytime Visible Video Translation
abstract
Infrared cameras are often utilized to enhance the night vision since the visible light cameras exhibit inferior efficacy without sufficient illumination. However, infrared data possesses inadequate color contrast and representation ability attributed to its intrinsic heat-related imaging principle, which hinders its application. Although, the domain gaps between unpaired nighttime infrared and daytime visible videos are even huger than paired ones that captured at the same time, establishing an effective translation mapping will greatly contribute to various fields. In this case, the structural knowledge within nighttime infrared videos and semantic information contained in the translated daytime visible pairs could be utilized simultaneously. To this end, we propose a tailored framework ROMA that couples with our introduced cRoss-domain regiOn siMilarity mAtching technique for bridging the huge gaps. To be specific, ROMA could efficiently translate the unpaired nighttime infrared videos into fine-grained daytime visible ones, meanwhile maintain the spatiotemporal consistency via matching the cross-domain region similarity. Furthermore, we design a multiscale region-wise discriminator to distinguish the details from synthesized visible results and real references. Moreover, we provide a new and challenging dataset encouraging further research for unpaired nighttime infrared and daytime visible video translation, named InfraredCity, which is $20$ times larger than the recently released infrared-related dataset IRVI. Codes and datasets are available https://github.com/BIT-DA/ROMA here.
Zhenjie Yu, Kai Chen 0030, Shuang Li 0008, Bingfeng Han, Chi Harold Liu, Shuigen Wang
ACM Multimedia6
2022 Sparse LiDAR and Binocular Stereo Fusion Network for 3D Object Detection
Weiqing Yan, Kaiqi Su, Jinlai Ren, Runmin Cong, Shuigen Wang
PRCV (3)6
2021 I2V-GAN: Unpaired Infrared-to-Visible Video Translation
abstract
Human vision is often adversely affected by complex environmental factors, especially in night vision scenarios. Thus, infrared cameras are often leveraged to help enhance the visual effects via detecting infrared radiation in the surrounding environment, but the infrared videos are undesirable due to the lack of detailed semantic information. In such a case, an effective video-to-video translation method from the infrared domain to the visible light counterpart is strongly needed by overcoming the intrinsic huge gap between infrared and visible fields. To address this challenging problem, we propose an infrared-to-visible (I2V) video translation method I2V-GAN to generate fine-grained and spatial-temporal consistent visible light videos by given unpaired infrared videos. Technically, our model capitalizes on three types of constraints: 1) adversarial constraint to generate synthetic frames that are similar to the real ones, 2) cyclic consistency with the introduced perceptual loss for effective content conversion as well as style preservation, and 3) similarity constraints across and within domains to enhance the content and motion consistency in both spatial and temporal spaces at a fine-grained level. Furthermore, the current public available infrared and visible light datasets are mainly used for object detection or tracking, and some are composed of discontinuous images which are not suitable for video tasks. Thus, we provide a new dataset for infrared-to-visible video translation, which is named IRVI. Specifically, it has 12 consecutive video clips of vehicle and monitoring scenes, and both infrared and visible light videos could be apart into 24352 frames. Comprehensive experiments on IRVI validate that I2V-GAN is superior to the compared state-of-the-art methods in the translation of infrared-to-visible videos with higher fluency and finer semantic details. Moreover, additional experimental results on the flower-to-flower dataset indicate I2V-GAN is also applicable to other video translation tasks. The code and IRVI dataset are available at https://github.com/BIT-DA/I2V-GAN.
Shuang Li 0008, Bingfeng Han, Zhenjie Yu, Chi Harold Liu, Kai Chen 0030, Shuigen Wang
ACM Multimedia6
2020 Blind Noisy Image Quality Assessment Using Sub-Band Kurtosis
abstract
Noise that afflicts natural images, regardless of the source, generally disturbs the perception of image quality by introducing a high-frequency random element that, when severe, can mask image content. Except at very low levels, where it may play a purpose, it is annoying. There exist significant statistical differences between distortion-free natural images and noisy images that become evident upon comparing the empirical probability distribution histograms of their discrete wavelet transform (DWT) coefficients. The DWT coefficients of low- or no-noise natural images have leptokurtic, peaky distributions with heavy tails; while noisy images tend to be platykurtic with less peaky distributions and shallower tails. The sample kurtosis is a natural measure of the peakedness and tail weight of the distributions of random variables. Here, we study the efficacy of the sample kurtosis of image wavelet coefficients as a feature driving, an extreme learning machine which learns to map kurtosis values into perceptual quality scores. The model is trained and tested on five types of noisy images, including additive white Gaussian noise, additive Gaussian color noise, impulse noise, masked noise, and high-frequency noise from the LIVE, CSIQ, TID2008, and TID2013 image quality databases. The experimental results show that the trained model has better quality evaluation performance on noisy images than existing blind noise assessment models, while also outperforming general-purpose blind and full-reference image quality assessment methods.
Chenwei Deng, Shuigen Wang, Alan C. Bovik, Guang-Bin Huang, Baojun Zhao
IEEE Trans. Cybern.2
2019 Cloud Detection in Satellite Images Based on Natural Scene Statistics and Gabor Features
abstract
Cloud detection is an important task in remote sensing (RS) image processing. Numerous cloud detection algorithms have been developed. However, most existing methods suffer from the weakness of omitting small and thin clouds, and from an inability to discriminate clouds from photometrically similar regions, such as buildings and snow. Here, we derive a novel cloud detection algorithm for optical RS images, whereby test images are separated into three classes: thick clouds, thin clouds, and noncloudy. First, a simple linear iterative clustering algorithm is adopted that is able to segment potential clouds, including small clouds. Then, a natural scene statistics model is applied to the superpixels to distinguish between clouds and surface buildings. Finally, Gabor features are computed within each superpixel and a support vector machine is used to distinguish clouds from snow regions. The experimental results indicate that the proposed model outperforms state-of-the-art methods for cloud detection.
Chenwei Deng, Zhen Li 0017, Shuigen Wang, Linbo Tang, Alan C. Bovik
IEEE Geosci. Remote. Sens. Lett.4
2019 Content-Insensitive Blind Image Blurriness Assessment Using Weibull Statistics and Sparse Extreme Learning Machine
abstract
Most of the existing image blurriness assessment algorithms are proposed based on measuring image edge width, gradient, high-frequency energy, or pixel intensity variation. However, these methods are content sensitive with little consideration of image content variations, which causes variant estimations for images with different contents but same blurriness degrees. In this paper, a content-insensitive blind image blurriness assessment metric is developed utilizing Weibull statistics. Inspired by the property that the statistics of image gradient magnitude (GM) follows Weibull distribution, we parameterize the GM using$\beta$(scale parameter) and$\gamma$(shape parameter) of Weibull distribution. We also adopt skewness ($\eta$) to measure the asymmetry of the GM distribution. In order to reduce the influence of image content and achieve more robust performance, divisive normalization is then incorporated to moderate the$\beta$,$\gamma$, and$\eta$. The final image quality is predicted using a sparse extreme learning machine. Performances evaluation on the blur image subsets in LIVE, CSIQ, TID2008, and TID2013 databases demonstrate that the proposed method is highly correlated with human perception and robust with image contents. In addition, our method has low computational complexity which is suitable for online applications.
Chenwei Deng, Shuigen Wang, Zhen Li 0017, Guang-Bin Huang, Weisi Lin
IEEE Trans. Syst. Man Cybern. Syst.2
2017 Cloud-cover assessment: From spectral properties to spatial domain natural scene statistic
abstract
Cloud contamination is the most common defect leading to quality degradation in remote sensing images. Numerous cloud-cover assessment (CCA) methods have been developed in the literature. The traditional Landsat 7 CCA algorithm attempted to detect clouds by taking advantages of different spectral properties from five spectral bands. However, it suffers the weakness of omitting thin cirrus clouds and the requirement of thermal bands. In this paper, we derived an automated CCA (ACCA) model that measures statistical deviations in spatial domain between cloud and clear images. Moreover, it only conducts on panchromatic band image, which can successfully address the limitation of unavailable thermal bands for satellite missions without thermal infrared sensors on board. A database with 400 clear/cloud images is then built for performance testing. Experimental results on the database show that our approach is more consistent with ground truths than the latest Landsat 8 ACCA results.
Shuigen Wang, Chenwei Deng, Baojun Zhao
IGARSS1
2017 NMF-Based Image Quality Assessment Using Extreme Learning Machine
abstract
Numerous state-of-the-art perceptual image quality assessment (IQA) algorithms share a common two-stage process: distortion description followed by distortion effects pooling. As for the first stage, the distortion descriptors or measurements are expected to be effective representatives of human visual variations, while the second stage should well express the relationship among quality descriptors and the perceptual visual quality. However, most of the existing quality descriptors (e.g., luminance, contrast, and gradient) do not seem to be consistent with human perception, and the effects pooling is often done in ad-hoc ways. In this paper, we propose a novel full-reference IQA metric. It applies non-negative matrix factorization (NMF) to measure image degradations by making use of the parts-based representation of NMF. On the other hand, a new machine learning technique [extreme learning machine (ELM)] is employed to address the limitations of the existing pooling techniques. Compared with neural networks and support vector regression, ELM can achieve higher learning accuracy with faster learning speed. Extensive experimental results demonstrate that the proposed metric has better performance and lower computational complexity in comparison with the relevant state-of-the-art approaches.
Shuigen Wang, Chenwei Deng, Weisi Lin, Guang-Bin Huang, Baojun Zhao
IEEE Trans. Cybern.1
2016 Gradient-based no-reference image blur assessment using extreme learning machine
Shuigen Wang, Chenwei Deng, Baojun Zhao, Guang-Bin Huang, Baoxian Wang
Neurocomputing1
2016 Fast and Accurate Spatiotemporal Fusion Based Upon Extreme Learning Machine
abstract
Spatiotemporal fusion is important in providing high spatial resolution earth observations with a dense time series, and recently, learning-based fusion methods have been attracting broad interest. These algorithms project image patches onto a feature space with the enforcement of a simple mapping to predict the fine resolution patches from the corresponding coarse ones. However, the sophisticated projection, e.g., sparse representation, is always computationally complex and difficult to be implemented on large patches, which cannot grasp enough local structural information in the coarse patches. To address these issues, a novel spatiotemporal fusion method is proposed in this letter, using a powerful learning technique, i.e., extreme learning machine (ELM). Unlike traditional approaches, we devote to learning a mapping function on difference images directly, rather than the sophisticated feature representation followed by a simple mapping. Characterized by good generalization performance and fast speed, the ELM is employed to achieve accurate and fast fine patches prediction. The proposed algorithm is evaluated by five actual data sets of Landsat enhanced thematic mapper plus-moderate resolution imaging spectroradiometer acquisitions and experimental results show that our method obtains better fusion results while achieving much greater speed.
Chenwei Deng, Shuigen Wang, Guang-Bin Huang, Baojun Zhao, Paula Lauren
IEEE Geosci. Remote. Sens. Lett.3
2015 Kurtosis-Based Blind Noisy Image Quality Assessment in Wavelet Domain
abstract
Noise distortions introduced in natural images generally break the initial probability distributions by dispersing image pixels randomly. We found that there exists a big difference between the distributions of Discrete Wavelet Transform (DWT) coefficients of natural images and noisy images: (1) for natural images, their distributions are sharp with high peaked ness and slight tail, (2) for noisy images, the shapes are much flatter with lower peaked ness and heavier tail. Kurtosis is able to measure and differentiate the probability distributions of noisy images with various noise levels. Moreover, the kurtosis values of DWT coefficients are stable for varying frequency filters. In this paper, we propose a Blind Noisy Image Quality Assessment model using Kurtosis (BNIQAK). Five types of noisy images in the three biggest databases are taken for testing BNIQAK. Experimental results show that BNIQAK has better evaluation performance compared with existing blind noisy models, as well as some general blind and full-reference (FR) methods.
Shuigen Wang, Chenwei Deng, Baojun Zhao
SMC1
2013 A novel SVD-based image quality assessment metric
abstract
Image distortion can be categorized into two aspects: content-dependent degradation and content-independent one. An existing full-reference image quality assessment (IQA) metric cannot deal with these two different impacts well. Singular value decomposition (SVD) as a useful mathematical tool has been used in various image processing applications. In this paper, SVD is employed to separate the structural (content-dependent) and the content-independent components. For each portion, we design a specific assessment model to tailor for its corresponding distortion properties. The proposed models are then fused to obtain the final quality score. Experimental results with the TID database demonstrate that the proposed metric achieves better performance in comparison with the relevant state-of-the-art quality metrics.
Shuigen Wang, Chenwei Deng, Weisi Lin, Baojun Zhao, Jie Chen 0026
ICIP1