Chenwei Deng

dblp:73/8538 · DBLP profile ↗
← Back
61ranked-venue papers
17as first author
12since 2021 · last 2026
0000-0002-3747-5128ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 22 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 10 · 3 first-authorSystems, architecture and hardware · 6Computer networks · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 Active Defocus Blurring for Background Clutter Suppression in LWIR DoFP Polarimeter
Jia Hao 0004, Xinling Yao, Yunyi Bian, Geng Tong, Jia Cao, Xiaochang Yu, Jiancun Zhao, Honglong Chang, Chenwei Deng, Yiting Yu
IEEE Trans. Circuits Syst. Video Technol.11
2025 Toward Effective Knowledge Distillation for Fine-Grained Object Recognition in Remote Sensing
abstract
With advancements in on-board computing devices deployed on remote sensing platforms, the demand for efficiently processing remote sensing imagery has become increasingly prominent. Knowledge distillation, as an effective lightweight method, has been introduced into this domain. Intuitively, distillation from a larger teacher model is expected to yield better performance. However, in our investigation of fine-grained object recognition in remote sensing imagery, we observed a counter-intuitive phenomenon: as the size of the teacher model increases, the performance of the student model initially improves but then degrades. This capacity gap issue hinders the effective utilization of stronger teacher models. To address this issue, we propose a novel distillation framework named BL-KD. It integrates two tailored components: the Class-level Learnable Orthogonal Projection (CLOP) module and the Object Re-Balance (ORB) module, which are jointly optimized to mitigate the negative impact of the capacity gap while effectively adapting to the unique distributional patterns and challenges inherent in remote sensing imagery. Experiments conducted on multiple fine-grained object recognition tasks in remote sensing demonstrate that our method consistently improves student performance, particularly in scenarios involving large teacher-student gaps, and outperforms several widely used distillation baselines.
Yangte Gao, Chenwei Deng, Liang Chen 0004
IEEE Geosci. Remote. Sens. Lett.2
2025 Satellite Attitude Estimation Based on Whole-to-Part Decoupling
abstract
Satellite attitude estimation is a key technology in on-orbit servicing missions. However, occlusions on the satellite’s surface in space introduce local omissions in satellite images, resulting in incomplete extraction of global satellite features. This causes significant biases in the mapping from features to attitude. Besides, there is a coupling between the parameters in rotation representations such as Euler angles and quaternions, making it difficult to improve the accuracy of each attitude parameter simultaneously. Existing methods focus on improving feature extraction performance and lack an evaluation of the relationship between satellite structure and attitude. To address these issues, we proposed a satellite attitude estimation method based on whole-to-part decoupling to separate features and reconstruct rotation representation, balancing attitude estimation biases. Specifically, we constructed a multibranch mapping with feature decoupling to select robust local features that contribute to attitude estimation, mitigating the impact of occlusion on feature extraction. Meanwhile, we designed an explicit rotation representation to decouple attitude parameters and match the relevant mappings for each parameter, reducing the estimation biases. The experimental results on public datasets demonstrate that the proposed method outperforms existing methods.
Chenwei Deng, Yuqi Han, Zhuokai Li
IEEE Geosci. Remote. Sens. Lett.3
2025 Decoupling-Based Cross-Layer Connection Removal for Compact Remote Sensing Image Classification
abstract
Cross-layer connections can couple different layers of neural networks in diverse paths, improving the classification accuracy of remote sensing images. However, some structures are unnecessary for inference and not conducive to be deployed on starboard/airborne platforms. Roughly removing them lengthens the backpropagation path and causes the gradient vanishing problem, degrading the performance of existing end-to-end training based lightweighting methods. To address this issue, we propose a decoupling-based cross-layer connection removal framework, which decomposes the complete deep network into multiple shallower stages based on the density of connections. By optimizing each stage independently, the backpropagation path is shortened and the gradient vanishing problem is expected to be alleviated. Moreover, to approximate the original network in each stage, we perform concret reconstruction error suppression on branches, layers and channels respectively. Combining these strategies, redundant structures are removed while maintaining performance. Extensive experiments on multiple public datasets and hardware platforms demonstrate that this compact architecture achieves substantially improved deployability and inference efficiency at the cost of only marginal accuracy degradation.
Yuqi Han, Chenwei Deng
IEEE Trans. Geosci. Remote. Sens.4
2024 Feature-Based Knowledge Distillation for Infrared Small Target Detection
abstract
Infrared small target detection is an extremely challenging task because of its low resolution, noise interference, and the weak thermal signal of small targets. Nevertheless, despite these difficulties, there is a growing interest in this field due to its significant application value in areas such as military, security, and unmanned aerial vehicles. In light of this, we propose a feature-based knowledge distillation method(IRKD) for infrared small target detection which can efficiently transfer detailed knowledge to students. The key idea behind IRKD is to assign varying importance to the features of teachers and students in different areas during the distillation process. Treating all features equally would negatively impact the distillation results. Therefore, We have developed a Unified Channel-Spatial Attention(UCSA) module that adaptively enhancing the crucial learning areas within the features. The experimental results show that compared to other knowledge distillation methods, our student detector achieved significant improvement in mean average precision(mAP). For example, IRKD improves ResNet50 based RetinaNet from 50.5% to 59.1% mAP, and improver ResNet50 based FCOS from 53.7% to 56.6% on Roboflow.
Jinglei Xue, Jianan Li 0001, Yuqi Han, Chenwei Deng, Tingfa Xu
IEEE Geosci. Remote. Sens. Lett.5
2024 Satellite Component Detection Based Upon Searchable Multiview Features Mapping
abstract
Satellite component detection (SCD) is one of the key technologies for on-orbit service (OOS) in space exploration. Unlike detection tasks on the ground, the single illumination in the space environment often leads to significant shadow occlusion when capturing satellite components from multiple viewing angles, resulting in the loss of component information. Existing component detection methods utilize this information to estimate their high-dimensional feature distributions and lack descriptions of the correlations between multiview features. To address this issue, we proposed an SCD based upon searchable multiview features mapping, which describes the distribution of satellite components by capturing both the discrimination between positive and negative samples and the commonality of intraclass components. Specifically, to enhance the connection of multiview features, orthogonal subspaces for different classes are constructed by the improved cosine distance between multiview features which is performed as a subspace metric to represent the discrimination and is used to cluster the mapped features. Besides, considering inherent errors in the mapping process, a loose constraint of the metric is introduced to search for common features, maintaining the local sparsity of multiview features. In the proposed dataset and the public dataset of satellite components, the proposed method achieved mean average precisions (mAP) of 0.878 and 0.862, respectively, outperforming existing state-of-the-art methods.
Chenwei Deng, Yuqi Han
IEEE Geosci. Remote. Sens. Lett.3
2024 Learning a Robust Topological Relationship for Online Multiobject Tracking in UAV Scenarios
abstract
Many existing multiobject tracking (MOT) methods tend to model each object’s feature individually. However, under acute viewpoint variation and occlusion, there may exist significant differences between the current and historical features of objects, which easily leads to object loss. To alleviate these issues, the topological relationships (i.e., geometric shapes formed by objects) should be modeled as a supplement to individual object features to maintain stability. In this article, we propose a novel MOT framework, which consists of a frame graph and association graph, to leverage the topological relationships both spatially and temporally. Technically, the frame graph models distance and angle among objects to resist viewpoint change, while the association graph utilizes the interframe temporal consistency of topological features to recover occluded objects. Extensive experiments on mainstream datasets demonstrate the effectiveness.
Chenwei Deng, Jiapeng Wu, Yuqi Han, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2023 Cross-AD: Multispectral and Hyperspectral High-Speed Artificial Imitation Object Anomaly Detection
abstract
Multispectral and hyperspectral imaging systems have shown great potential in detecting artificial imitations at natural backgrounds. However, the general spectral anomaly detection(AD) algorithms lack attention to the camouflage characteristics and are not fully suitable for real-time inspection applications. Since the imitative targets have local extent similarity but global differences between real backgrounds, we proposed the Cross-AD method for imitation anomaly detection (IAD), which is based on the horizontal and vertical adaptive background estimation at the pixel level. It combines the scattered distribution features of camouflage and the directional model has a better suppression effect on the common band noise, while the execution time is close to the RX detector. Furthermore, the improved Cross-Box with protection window and Cross-Index with feature factor is proposed to better deal with large imitations and dense vegetation environments. To validate the algorithms, the hyperspectral imitation anomaly detection (HSIAD) dataset is constructed based on the actual camouflage scene. Compared with the real-time and fast spectral AD detection algorithms, the cross-series methods achieve the optimal balance of execution time and detection performance on both multi-and hyperspectral IAD datasets. The implementation code is available at https://github.com/XingshiLuo/Cross-AD.
Xingshi Luo, Chenwei Deng
IGARSS3
2023 GradQuant: Low-Loss Quantization for Remote-Sensing Object Detection
abstract
Convolutional neural network based methods have shown remarkable performance in remote sensing object detection. However, their deployment on resource-limited embedded devices is hindered by their high computational complexity. Neural network quantization methods have been proven effective in compressing and accelerating CNN models by clipping outlier activations and utilizing low-precision values to represent weights and clipped activations. Nonetheless, the clipping of outlier activations leads to distortion of object local features. Furthermore, the lack of enhanced overall feature mining exacerbates the degradation of detection accuracy. To address the limitations above, we propose an innovative clipping-free quantization method called GradQuant, which mitigates model’s quantization accuracy loss caused by clipping outlier activations and the lack of overall feature mining. Specifically, a bounded activation function (sigmoid-weighted tanh, SiTanh) is carefully designed to ensure that object features are represented within a limited range without clipping. On the basis of this, an activation substitute training (AST) method is co-designed to prompt models to focus more on non-outlier object features instead of outlier-like local ones. Extensive experiments on public remote-sensing datasets demonstrate the effectiveness of GradQuant method compared with other state-of-the-art quantization methods.
Chenwei Deng, Yuqi Han, Donglin Jing, Hong Zhang 0018
IEEE Geosci. Remote. Sens. Lett.1
2023 Toward Hierarchical Adaptive Alignment for Aerial Object Detection in Remote Sensing Images
abstract
The aerial objects tend to distribute with a major variation on the scale and arbitrary orientations in remote sensing images. To meet such characteristics of the aerial object, most of the existing anchor-based detectors rely on preset anchors with variable scales, angles, and aspect ratios, which leads to the misalignment of the selection of candidate regions, the extraction of object features, and the label assignment of preset boxes, interfering the performance of the detector. To address this issue, we propose a Hierarchical Adaptive Alignment Network (HAA-Net). Specifically, we first design the Region Refinement Module (RRM), Feature Alignment module (FAM), and Potential Label Assignment Module (PLAM) to alleviate the misalignment of the region, feature, and label levels respectively; furthermore, we use the gradient equalization strategy to jointly optimize these modules at different levels, so that the whole network can be fully trained to significantly improve detection performance. Extensive experiments demonstrate that our approach can achieve superior performance in three common aerial object datasets (e.g., DOTA, HRSC2016, and UCAS-AOD) when compared with state-of-the-art detectors.
Chenwei Deng, Donglin Jing, Yuqi Han, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2022 FAR-Net: Fast Anchor Refining for Arbitrary-Oriented Object Detection
abstract
Compared with natural images, targets in remote-sensing images are often distributed with more flexible orientation, aspect ratio, and scale. Thus, anchor-based algorithms often employ plenty of preset anchors to encode the above-mentioned attributes in object detection tasks. However, they often suffer from the following issues: 1) significant computational burden caused by dense-sampling anchors; 2) serious background interference since many anchors only cover small parts of the actual target; and 3) feature misalignment between the targets with the preset anchors due to the absence of the most discriminant features for target extraction. Therefore, in this letter, a fast anchor refining network (FAR-Net) is advocated to address the remaining issues for arbitrary-oriented object detection in the remote-sensing field. To be specific, a rotation alignment module (RAM) and balanced regression loss function (BR-loss) are carefully designed in the FAR-Net. The RAM is capable of generating high-quality anchors based on a refinement convolution and adaptively aligning the convolutional features by complying with the anchor boxes to reduce redundant calculation. The BR-loss is designed by employing a balanced loss function to prevent misaligned anchors from causing major gradient descents, thereby achieving a more stable network training procedure. Extensive experiments on public remote-sensing datasets (HRSC2016 and UCAS-AOD) demonstrate the excellent detection performance of our algorithm in comparison with numerous existing detectors.
Chenwei Deng, Donglin Jing, Yuqi Han, Shuliang Wang 0001, Hongshuo Wang
IEEE Geosci. Remote. Sens. Lett.1
2021 Learning Dynamic Spatial-Temporal Regularization for UAV Object Tracking
abstract
With the wide vision and high flexibility, unmanned aerial vehicle (UAV) has been widely used into object tracking in recent years. However, its limited computing capability poses a great challenges to tracking algorithms. On the other hand, Discriminative Correlation Filter (DCF) based trackers have attracted great attention due to their computational efficiency and superior accuracy. Many studies introduce spatial and temporal regularization into the DCF framework to achieve a more robust appearance model and further enhance the tracking performance. However, such algorithms generally set fixed spatial or temporal regularization parameters, which lack flexibility and adaptability under cluttered and challenging scenarios. To tackle such issue, in this letter, we propose a novel DCF tracking model by introducing dynamic spatial regularization weight, which encourage the filter focuses on more reliable region during training stage. Furthermore, our method could optimize the spatial and temporal regularization weight simultaneously using Alternative Direction Method of Multiplies (ADMM) technique method, where each sub-problem has closed-form solution. Through the joint optimization, our tracker could not only suppress the potential distractors but also construct robust target appearance on the basis of reliable historical information. Experiments on two UAV benchmarks have demonstrated that our tracker performs favorably against other state-of-the-art algorithms.
Chenwei Deng, Shuangcheng He, Yuqi Han, Boya Zhao
IEEE Signal Process. Lett.1
2020 Few-shot traffic sign recognition with clustering inductive bias and random neural network
Chenwei Deng, Zhengquan Piao, Baojun Zhao
Pattern Recognit.2
2020 High-Performance Visual Tracking With Extreme Learning Machine Framework
abstract
In real-time applications, a fast and robust visual tracker should generally have the following important properties: 1) feature representation of an object that is not only efficient but also has a good discriminative capability and 2) appearance modeling which can quickly adapt to the variations of foreground and backgrounds. However, most of the existing tracking algorithms cannot achieve satisfactory performance in both of the two aspects. To address this issue, in this paper, we advocate a novel and efficient visual tracker by exploiting the excellent feature learning and classification capabilities of an emerging learning technique, that is, extreme learning machine (ELM). The contributions of the proposed work are as follows: 1) motivated by the simplicity and learning ability of the ELM autoencoder (ELM-AE), an ELM-AE-based feature extraction model is presented, and this model can provide a compact and discriminative representation of the inputs efficiently and 2) due to the fast learning speed of an ELM classifier, an ELM-based appearance model is developed for feature classification, and is able to rapidly distinguish the object of interest from its surroundings. In addition, in order to cope with the visual changes of the target and its backgrounds, the online sequential ELM is used to incrementally update the appearance model. Plenty of experiments on challenging image sequences demonstrate the effectiveness and robustness of the proposed tracker.
Chenwei Deng, Yuqi Han, Baojun Zhao
IEEE Trans. Cybern.1
2020 Blind Noisy Image Quality Assessment Using Sub-Band Kurtosis
abstract
Noise that afflicts natural images, regardless of the source, generally disturbs the perception of image quality by introducing a high-frequency random element that, when severe, can mask image content. Except at very low levels, where it may play a purpose, it is annoying. There exist significant statistical differences between distortion-free natural images and noisy images that become evident upon comparing the empirical probability distribution histograms of their discrete wavelet transform (DWT) coefficients. The DWT coefficients of low- or no-noise natural images have leptokurtic, peaky distributions with heavy tails; while noisy images tend to be platykurtic with less peaky distributions and shallower tails. The sample kurtosis is a natural measure of the peakedness and tail weight of the distributions of random variables. Here, we study the efficacy of the sample kurtosis of image wavelet coefficients as a feature driving, an extreme learning machine which learns to map kurtosis values into perceptual quality scores. The model is trained and tested on five types of noisy images, including additive white Gaussian noise, additive Gaussian color noise, impulse noise, masked noise, and high-frequency noise from the LIVE, CSIQ, TID2008, and TID2013 image quality databases. The experimental results show that the trained model has better quality evaluation performance on noisy images than existing blind noise assessment models, while also outperforming general-purpose blind and full-reference image quality assessment methods.
Chenwei Deng, Shuigen Wang, Alan C. Bovik, Guang-Bin Huang, Baojun Zhao
IEEE Trans. Cybern.1
2019 Multimodal-Temporal Fusion: Blending Multimodal Remote Sensing Images to Generate Image Series With High Temporal Resolution
abstract
This paper aims to tackle a general but interesting cross-modality problem in remote sensing community: can multimodal images help to generate synthetic images in time series and improve temporal resolution? To this end, we explore multimodal-temporal fusion, in which we attempt to leverage the availability of additional cross-modality images to simulate the missing images in time series. We propose a multimodal-temporal fusion framework, and mainly focus on two kinds of information for the simulation: intra-modal cross-modality information and inter-modal temporal information. To exploit the cross-modality information, we adopt available paired images and learn a mapping between different modality images using a deep neural network. Considering temporal dependency among time-series images, we formulate a temporal constraint in the learning to encourage temporal consistent results. Experiments are conducted on two cross-modality image simulation applications (SAR to visible and visible to SWIR), and both visual and quantitative results demonstrate that the proposed model can successfully simulate missing images with cross-modality data.
Chenwei Deng, Baojun Zhao, Jocelyn Chanussot
IGARSS2
2019 GenELM: Generative Extreme Learning Machine feature representation
Chenwei Deng, Guang-Bin Huang, Baojun Zhao
Neurocomputing2
2019 Cloud Detection in Satellite Images Based on Natural Scene Statistics and Gabor Features
abstract
Cloud detection is an important task in remote sensing (RS) image processing. Numerous cloud detection algorithms have been developed. However, most existing methods suffer from the weakness of omitting small and thin clouds, and from an inability to discriminate clouds from photometrically similar regions, such as buildings and snow. Here, we derive a novel cloud detection algorithm for optical RS images, whereby test images are separated into three classes: thick clouds, thin clouds, and noncloudy. First, a simple linear iterative clustering algorithm is adopted that is able to segment potential clouds, including small clouds. Then, a natural scene statistics model is applied to the superpixels to distinguish between clouds and surface buildings. Finally, Gabor features are computed within each superpixel and a support vector machine is used to distinguish clouds from snow regions. The experimental results indicate that the proposed model outperforms state-of-the-art methods for cloud detection.
Chenwei Deng, Zhen Li 0017, Shuigen Wang, Linbo Tang, Alan C. Bovik
IEEE Geosci. Remote. Sens. Lett.1
2019 Spatial-Temporal Context-Aware Tracking
abstract
Discriminative correlation filters (DCFs) have recently achieved competitive performance in visual tracking benchmarks. However, most of the existing DCF trackers only consider the spatial features of the target and could hardly benefit from the inter-frame and historical information, which may degrade the tracking performance when occlusion and deformation occurs. To tackle the above-mentioned issues, in this letter, by introducing the temporal constrain into the DCF tracker, we advocate our spatial-temporal context-aware tracker. Through jointly modeling the spatial context and historical target information, our tracker could not only adapt the appearance change but also maintain a relatively stable filter due to the small target variation between inter-frames. Furthermore, we show that the proposed objective formula could be directly solved using the Alternating Direction Method of Multipliers (ADMM) technique with low computational cost. Experiments on the large-scale benchmark demonstrate that the proposed trackers perform favorably against other state-of-the-art methods.
Yuqi Han, Chenwei Deng, Boya Zhao, Baojun Zhao
IEEE Signal Process. Lett.2
2019 StfNet: A Two-Stream Convolutional Neural Network for Spatiotemporal Image Fusion
abstract
Spatiotemporal image fusion is considered as a promising way to provide Earth observations with both high spatial resolution and frequent coverage, and recently, learning-based solutions have been receiving broad attention. However, these algorithms treating spatiotemporal fusion as a single image super-resolution problem, generally suffers from the significant spatial information loss in coarse images, due to the large upscaling factors in real applications. To address this issue, in this paper, we exploit temporal information in fine image sequences and solve the spatiotemporal fusion problem with a two-stream convolutional neural network called StfNet. The novelty of this paper is twofold. First, considering the temporal dependence among image sequences, we incorporate the fine image acquired at the neighboring date to super-resolve the coarse image at the prediction date. In this way, our network predicts a fine image not only from the structural similarity between coarse and fine image pairs but also by exploiting abundant texture information in the available neighboring fine images. Second, instead of estimating each output fine image independently, we consider the temporal relations among time-series images and formulate a temporal constraint. This temporal constraint aiming to guarantee the uniqueness of the fusion result and encourages temporal consistent predictions in learning and thus leads to more realistic final results. We evaluate the performance of the StfNet using two actual data sets of Landsat-Moderate Resolution Imaging Spectroradiometer (MODIS) acquisitions, and both visual and quantitative evaluations demonstrate that our algorithm achieves state-of-the-art performance.
Chenwei Deng, Jocelyn Chanussot, Danfeng Hong, Baojun Zhao
IEEE Trans. Geosci. Remote. Sens.2
2019 State-Aware Anti-Drift Object Tracking
abstract
Correlation filter (CF) based trackers have aroused increasing attentions in visual tracking field due to the superior performance on several datasets while maintaining high running speed. For each frame, an ideal filter is trained in order to discriminate the target from its surrounding background. Considering that the target always undergoes external and internal interference during tracking procedure, the trained tracker should not only have the ability to judge the current state when failure occurs, but also to resist the model drift caused by challenging distractions. To this end, we present a State-aware Anti-drift Tracker (SAT) in this paper, which jointly model the discrimination and reliability information in filter learning. Specifically, global context patches are incorporated into filter training stage to better distinguish the target from backgrounds. Meanwhile, a color-based reliable mask is learned to encourage the filter to focus on more reliable regions suitable for tracking. We show that the proposed optimization problem could be efficiently solved using Alternative Direction Method of Multipliers and fully carried out in Fourier domain. Furthermore, a Kurtosis-based updating scheme is advocated to reveal the tracking condition as well as guarantee a high-confidence template updating. Extensive experiments are conducted on OTB-100 and UAV-20L datasets to compare the SAT tracker with other relevant state-of-the-art methods. Both quantitative and qualitative evaluations further demonstrate the effectiveness and robustness of the proposed work.
Yuqi Han, Chenwei Deng, Baojun Zhao, Dacheng Tao
IEEE Trans. Image Process.2
2019 Content-Insensitive Blind Image Blurriness Assessment Using Weibull Statistics and Sparse Extreme Learning Machine
abstract
Most of the existing image blurriness assessment algorithms are proposed based on measuring image edge width, gradient, high-frequency energy, or pixel intensity variation. However, these methods are content sensitive with little consideration of image content variations, which causes variant estimations for images with different contents but same blurriness degrees. In this paper, a content-insensitive blind image blurriness assessment metric is developed utilizing Weibull statistics. Inspired by the property that the statistics of image gradient magnitude (GM) follows Weibull distribution, we parameterize the GM using$\beta$(scale parameter) and$\gamma$(shape parameter) of Weibull distribution. We also adopt skewness ($\eta$) to measure the asymmetry of the GM distribution. In order to reduce the influence of image content and achieve more robust performance, divisive normalization is then incorporated to moderate the$\beta$,$\gamma$, and$\eta$. The final image quality is predicted using a sparse extreme learning machine. Performances evaluation on the blur image subsets in LIVE, CSIQ, TID2008, and TID2013 databases demonstrate that the proposed method is highly correlated with human perception and robust with image contents. In addition, our method has low computational complexity which is suitable for online applications.
Chenwei Deng, Shuigen Wang, Zhen Li 0017, Guang-Bin Huang, Weisi Lin
IEEE Trans. Syst. Man Cybern. Syst.1
2018 Constrained Nonnegative Matrix Factorization for Robust Hyperspectral Unmixing
abstract
Hyperspectral unmixng (HU) is an essential step for hyperspectral image (HSI) analysis. In real HSI, there often are abnormal fluctuations existing in specific bands, which can be described as sparse noise. This type of corruption will seriously disrupt the hyperspectral image quality, causing extra difficulties during unmixing process. However, the influence of sparse noise is often ignored by existing unmixing methods, which leads to the reduction of robustness and accuracy for HU tasks. Therefore, we propose a new unmixing model which takes noise corruption into consideration. By designing and imposing constraints considering the sparsity of noise, properties of endmember and abundance on nonnegative matrix factorization (NMF), the proposed method can resist the sparse noise and achieve more robust and accurate unmixing results. Adequate experiments have been conducted on both synthetic and real hyperspectral data. And the results confirm the superiority of proposed method compared to state-of-the-art methods.
Chenwei Deng, Jiahui Dai, Baojun Zhao
IGARSS2
2018 Feature-Level Loss for Multispectral Pan-Sharpening with Machine Learning
abstract
Multispectral pan-sharpening plays an important role in providing earth observation with both high-spatial and high-spectral resolutions, and recently pan-sharpening with machine learning has been attracting broad interest. However, these algorithms minimizing the pixel-wise mean squared error, generally suffer from over-smoothed results that lack of high-frequency details in both spatial and spectral dimensions. In this paper, we propose to tackle this problem by shifting the learning loss from pixel-wise error to a higher-level feature loss. The new loss function, formulated by spatial structure similarity and spectral angle mapping, pushes the model to generate results that have similar feature representations with ground truth, rather than match with pixel-wise accuracy. Consequently, more realistic fusion results can be produced. Visual and quantitative analysis both demonstrate that our approach achieves better performance in comparison with state-of-the-art algorithms. Furthermore, experiments on high-level remote sensing task further confirm the superiority of the proposed method in real applications.
Chenwei Deng, Baojun Zhao, Jocelyn Chanussot
IGARSS2
2017 Adaptive feature representation for visual tracking
abstract
Robust feature representation plays significant role in visual tracking. However, it remains a challenging issue, since many factors may affect the experimental performance. The existing method, which combine different features by setting them equally with the fixed weight, could hardly solve the issues, due to the different statistical properties of different features across various of scenarios and attributes. In this paper, by exploiting the internal relationship among these features, we develop a robust method to construct a more stable feature representation. More specifically, we utilize a co-training paradigm to formulate the intrinsic complementary information of multi-feature template into the efficient correlation filter framework. We test our approach on challenging sequences with illumination variation, scale variation, deformation etc. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods favorably.
Yuqi Han, Chenwei Deng, Zengshuo Zhang, Jiatong Li 0004, Baojun Zhao
ICIP2
2017 Low-rank and sparse tensor recovery for hyperspectral anomaly detection
abstract
Anomaly is generally defined as an object that strays away from the background clutter. As for hyperspectral anomaly detection, most of the previous methods fail to fully take advantage of the knowledge in both spatial and spectral domain. In this paper, we propose a novel method based on tensor recovery in which spatial structures and spectral characters are reasonably considered to separate the hypercube into a low-rank background and sparse anomalies. Since background is highly consistent not only in spectral domain, but also in spatial domain, we impose low-rank constraints on three unfolding matrix of the hypercube respectively to capture the global structure together. To better describe the local irregularities with low probability, a general l1norm constraint and an extra sparse regularization are imposed on pixels in the spectral mode alone, for that we consider each spectrum as an entirety. Extensive experiments on two real datasets show outstanding anomaly detection performance of the proposed method in comparison with the state-of-the-art methods.
Jiahui Dai, Chenwei Deng
IGARSS2
2017 Cloud-cover assessment: From spectral properties to spatial domain natural scene statistic
abstract
Cloud contamination is the most common defect leading to quality degradation in remote sensing images. Numerous cloud-cover assessment (CCA) methods have been developed in the literature. The traditional Landsat 7 CCA algorithm attempted to detect clouds by taking advantages of different spectral properties from five spectral bands. However, it suffers the weakness of omitting thin cirrus clouds and the requirement of thermal bands. In this paper, we derived an automated CCA (ACCA) model that measures statistical deviations in spatial domain between cloud and clear images. Moreover, it only conducts on panchromatic band image, which can successfully address the limitation of unavailable thermal bands for satellite missions without thermal infrared sensors on board. A database with 400 clear/cloud images is then built for performance testing. Experimental results on the database show that our approach is more consistent with ground truths than the latest Landsat 8 ACCA results.
Shuigen Wang, Chenwei Deng, Baojun Zhao
IGARSS2
2017 Effective visual tracking by pairwise metric learning
Chenwei Deng, Baoxian Wang, Weisi Lin, Guang-Bin Huang, Baojun Zhao
Neurocomputing1
2017 NMF-Based Image Quality Assessment Using Extreme Learning Machine
abstract
Numerous state-of-the-art perceptual image quality assessment (IQA) algorithms share a common two-stage process: distortion description followed by distortion effects pooling. As for the first stage, the distortion descriptors or measurements are expected to be effective representatives of human visual variations, while the second stage should well express the relationship among quality descriptors and the perceptual visual quality. However, most of the existing quality descriptors (e.g., luminance, contrast, and gradient) do not seem to be consistent with human perception, and the effects pooling is often done in ad-hoc ways. In this paper, we propose a novel full-reference IQA metric. It applies non-negative matrix factorization (NMF) to measure image degradations by making use of the parts-based representation of NMF. On the other hand, a new machine learning technique [extreme learning machine (ELM)] is employed to address the limitations of the existing pooling techniques. Compared with neural networks and support vector regression, ELM can achieve higher learning accuracy with faster learning speed. Extensive experimental results demonstrate that the proposed metric has better performance and lower computational complexity in comparison with the relevant state-of-the-art approaches.
Shuigen Wang, Chenwei Deng, Weisi Lin, Guang-Bin Huang, Baojun Zhao
IEEE Trans. Cybern.2
2017 Robust Object Tracking With Discrete Graph-Based Multiple Experts
abstract
Variations of target appearances due to illumination changes, heavy occlusions, and target deformations are the major factors for tracking drift. In this paper, we show that the tracking drift can be effectively corrected by exploiting the relationship between the current tracker and its historical tracker snapshots. Here, a multi-expert framework is established by the current tracker and its historical trained tracker snapshots. The proposed scheme is formulated into a unified discrete graph optimization framework, whose nodes are modeled by the hypotheses of the multiple experts. Furthermore, an exact solution of the discrete graph exists giving the object state estimation at each time step. With the unary and binary compatibility graph scores defined properly, the proposed framework corrects the tracker drift via selecting the best expert hypothesis, which implicitly analyzes the recent performance of the multi-expert by only evaluating graph scores at the current frame. Three base trackers are integrated into the proposed framework to validate its effectiveness. We first integrate the online SVM on a budget algorithm into the framework with significant improvement. Then, the regression correlation filters with hand-crafted features and deep convolutional neural network features are introduced, respectively, to further boost the tracking performance. The proposed three trackers are extensively evaluated on three data sets: TB-50, TB-100, and VOT2015. The experimental results demonstrate the excellent performance of the proposed approaches against the state-of-the-art methods.
Jiatong Li 0004, Chenwei Deng, Dacheng Tao, Baojun Zhao
IEEE Trans. Image Process.2
2016 Textural and Gradient Feature Extraction from JPEG2000 Codestream for Airfield Detection
abstract
There has been huge growth in the use of optical images in object detection, and numerous images over satellite and aerial vehicle are typically stored and transmitted in the compressed form such as JPEG2000. However, most of the existing detection algorithms are built in the pixel domain, which leads to enlarge data volume and requires compulsory decompression before detecting. Hence, these issues may hinder their practical use in real-time and storage limited applications. In this paper, we extract textural and gradient features from JPEG2000 compressed codestream, and these robust features are utilized for coarse-to-fine airfield detection. The contributions of proposed work are as follows: 1) packet header is effectively exploited for presenting textural information of images, while gradient calculation is achieved in packet body, 2) a novel detection framework taking full advantage of compressed features is proposed. Validated by experiments with large optical images, the proposed method achieves faster computation with higher detection accuracy, in comparison with the existing relevant state-of-the-art approaches.
Chenwei Deng, Baojun Zhao
DCC2
2016 Spatiotemporal reflectance fusion based on location regularized sparse representation
abstract
Spatiotemporal reflectance fusion plays an important role in providing earth observation with both high-spatial and high-temporal resolutions, and sparse representation is one of the popular strategies to implement spatiotemporal fusion. However, the existing methods generally suffers from instability of sparse representation for the fine and coarse image pairs. In this paper, we demonstrate that such instability can be addressed by exploiting spatial correlations among the neighboring fine images, which is mathematically formulated as a location regularized term. A fast iterative shrinkage-thresholding algorithm (FISTA) is then employed to find the optimal solution. Experimental results show that the performance of proposed method outperforms other relevant state-of-the-art fusion approaches.
Chenwei Deng, Baojun Zhao
IGARSS2
2016 Gradient-based no-reference image blur assessment using extreme learning machine
Shuigen Wang, Chenwei Deng, Baojun Zhao, Guang-Bin Huang, Baoxian Wang
Neurocomputing2
2016 Fast and Accurate Spatiotemporal Fusion Based Upon Extreme Learning Machine
abstract
Spatiotemporal fusion is important in providing high spatial resolution earth observations with a dense time series, and recently, learning-based fusion methods have been attracting broad interest. These algorithms project image patches onto a feature space with the enforcement of a simple mapping to predict the fine resolution patches from the corresponding coarse ones. However, the sophisticated projection, e.g., sparse representation, is always computationally complex and difficult to be implemented on large patches, which cannot grasp enough local structural information in the coarse patches. To address these issues, a novel spatiotemporal fusion method is proposed in this letter, using a powerful learning technique, i.e., extreme learning machine (ELM). Unlike traditional approaches, we devote to learning a mapping function on difference images directly, rather than the sophisticated feature representation followed by a simple mapping. Characterized by good generalization performance and fast speed, the ELM is employed to achieve accurate and fast fine patches prediction. The proposed algorithm is evaluated by five actual data sets of Landsat enhanced thematic mapper plus-moderate resolution imaging spectroradiometer acquisitions and experimental results show that our method obtains better fusion results while achieving much greater speed.
Chenwei Deng, Shuigen Wang, Guang-Bin Huang, Baojun Zhao, Paula Lauren
IEEE Geosci. Remote. Sens. Lett.2
2016 Time Varying Metric Learning for visual tracking
Jiatong Li 0004, Baojun Zhao, Chenwei Deng
Pattern Recognit. Lett.3
2016 Extreme Learning Machine for Multilayer Perceptron
abstract
Extreme learning machine (ELM) is an emerging learning algorithm for the generalized single hidden layer feedforward neural networks, of which the hidden node parameters are randomly generated and the output weights are analytically computed. However, due to its shallow architecture, feature learning using ELM may not be effective for natural signals (e.g., images/videos), even with a large number of hidden nodes. To address this issue, in this paper, a new ELM-based hierarchical learning framework is proposed for multilayer perceptron. The proposed architecture is divided into two main components: 1) self-taught feature extraction followed by supervised feature classification and 2) they are bridged by random initialized hidden weights. The novelties of this paper are as follows: 1) unsupervised multilayer encoding is conducted for feature extraction, and an ELM-based sparse autoencoder is developed via l1 constraint. By doing so, it achieves more compact and meaningful feature representations than the original ELM; 2) by exploiting the advantages of ELM random feature mapping, the hierarchically encoded outputs are randomly projected before final decision making, which leads to a better generalization with faster learning speed; and 3) unlike the greedy layerwise training of deep learning (DL), the hidden layers of the proposed framework are trained in a forward manner. Once the previous layer is established, the weights of the current layer are fixed without fine-tuning. Therefore, it has much better learning efficiency than the DL. Extensive experiments on various widely used classification data sets show that the proposed algorithm achieves better and faster convergence than the existing state-of-the-art hierarchical learning methods. Furthermore, multiple applications in computer vision further confirm the generality and capability of the proposed learning scheme.
Jiexiong Tang, Chenwei Deng, Guang-Bin Huang
IEEE Trans. Neural Networks Learn. Syst.2
2015 Kurtosis-Based Blind Noisy Image Quality Assessment in Wavelet Domain
abstract
Noise distortions introduced in natural images generally break the initial probability distributions by dispersing image pixels randomly. We found that there exists a big difference between the distributions of Discrete Wavelet Transform (DWT) coefficients of natural images and noisy images: (1) for natural images, their distributions are sharp with high peaked ness and slight tail, (2) for noisy images, the shapes are much flatter with lower peaked ness and heavier tail. Kurtosis is able to measure and differentiate the probability distributions of noisy images with various noise levels. Moreover, the kurtosis values of DWT coefficients are stable for varying frequency filters. In this paper, we propose a Blind Noisy Image Quality Assessment model using Kurtosis (BNIQAK). Five types of noisy images in the three biggest databases are taken for testing BNIQAK. Experimental results show that BNIQAK has better evaluation performance compared with existing blind noisy models, as well as some general blind and full-reference (FR) methods.
Shuigen Wang, Chenwei Deng, Baojun Zhao
SMC2
2015 Extreme learning machines: new trends and applications
Chenwei Deng, Guang-Bin Huang, Jiexiong Tang
Sci. China Inf. Sci.1
2015 Visual acuity inspired saliency detection by using sparse features
Yuming Fang 0001, Weisi Lin, Zhijun Fang 0001, Zhenzhong Chen 0001, Chia-Wen Lin, Chenwei Deng
Inf. Sci.6
2015 Exploiting entropy masking in perceptual graphic rendering
Lu Dong 0001, Yuming Fang 0001, Weisi Lin, Chenwei Deng, Ce Zhu, Seah Hock Soon
Signal Process. Image Commun.4
2015 Scale and Orientation Invariant Text Segmentation for Born-Digital Compound Images
abstract
Many recent applications require text segmentation for born-digital compound images. To this end, we propose a coarse-to-fine framework for segmenting texts of arbitrary scales and orientations in born-digital compound images. In the coarse stage, the local image activity measure is designed based upon the variation distribution of characters, to highlight the difference between textual and pictorial regions. This stage outputs a coarse textual layer including textual regions as well as a few pictorial regions with high activity. In the fine stage, a textual connected component (TCC) based refinement is proposed to eliminate the survived pictorial regions. In particular, a scale and orientation invariant grouping algorithm is proposed to adaptively generate TCCs with uniform statistical features. The minimum average distance and morphological operations are employed to assist the formation of candidate TCCs. Then, three string-level features (i.e., shapeness, color similarity, and mean activity level) are designed to distinguish the true TCCs from the false positive ones that are formed by connecting the high activity pictorial components. Extensive experiments show that the proposed framework can segment textual regions precisely from born-digital compound images, while preserving the integrity of texts with varied scales and orientations, and avoiding over-connection of textual regions.
Huan Yang 0001, Shiqian Wu, Chenwei Deng, Weisi Lin
IEEE Trans. Cybern.3
2015 Compressed-Domain Ship Detection on Spaceborne Optical Image Using Deep Neural Network and Extreme Learning Machine
abstract
Ship detection on spaceborne images has attracted great interest in the applications of maritime security and traffic control. Optical images stand out from other remote sensing images in object detection due to their higher resolution and more visualized contents. However, most of the popular techniques for ship detection from optical spaceborne images have two shortcomings: 1) Compared with infrared and synthetic aperture radar images, their results are affected by weather conditions, like clouds and ocean waves, and 2) the higher resolution results in larger data volume, which makes processing more difficult. Most of the previous works mainly focus on solving the first problem by improving segmentation or classification with complicated algorithms. These methods face difficulty in efficiently balancing performance and complexity. In this paper, we propose a ship detection approach to solving the aforementioned two issues using wavelet coefficients extracted from JPEG2000 compressed domain combined with deep neural network (DNN) and extreme learning machine (ELM). Compressed domain is adopted for fast ship candidate extraction, DNN is exploited for high-level feature representation and classification, and ELM is used for efficient feature pooling and decision making. Extensive experiments demonstrate that, in comparison with the existing relevant state-of-the-art approaches, the proposed method requires less detection time and achieves higher detection accuracy.
Jiexiong Tang, Chenwei Deng, Guang-Bin Huang, Baojun Zhao
IEEE Trans. Geosci. Remote. Sens.2
2014 A fast learning algorithm for multi-layer extreme learning machine
abstract
Extreme learning machine (ELM) is an efficient training algorithm originally proposed for single-hidden layer feedforward networks (SLFNs), of which the input weights are randomly chosen and need not to be fine-tuned. In this paper, we present a new stack architecture for ELM, to further improve the learning accuracy of ELM while maintaining its advantage of training speed. By exploiting the hidden information of ELM random feature space, a recovery-based training model is developed and incorporated into the proposed ELM stack architecture. Experimental results of the MNIST handwriting dataset demonstrate that the proposed algorithm achieves better and much faster convergence than the state-of-the-art ELM and deep learning methods.
Jiexiong Tang, Chenwei Deng, Guang-Bin Huang, Junhui Hou
ICIP2
2014 Study on subjective quality assessment of Digital Compound Images
abstract
Quality assessment of digital compound images is a less investigated research topic. In this paper, we present a study for subjective quality assessment of Digital Compound Images (DCIs), and investigate whether existing Image Quality Assessment (IQA) methods are effective to evaluate the quality of distorted DCIs. A new Compound Image Quality Assessment Database (CIQAD) is constructed, including 24 reference DCIs and their 576 distorted versions. The Paired Comparison (PC) method is employed for the subjective viewing, and the Hodgerank decomposition is adopted to generate incomplete but balanced comparison pairs, so as to reduce the execution time while guaranteeing the reliability of the results. In our experiment, correlation of 14 existing IQA methods with the obtained Mean Opinion Score (MOS) values on the CIQAD is calculated, which indicates that the 14 IQA methods are not consistent with human visual perception when judging DCIs in different conditions. Therefore, objective quality assessment metrics should be specifically designed for DCIs. Our subjective study has delivered convincing information to guide the construction of objective metrics. Furthermore, we has also published the database online to favor future research on quality assessment of DCIs.
Huan Yang 0001, Weisi Lin, Chenwei Deng, Long Xu 0001
ISCAS3
2013 A novel SVD-based image quality assessment metric
abstract
Image distortion can be categorized into two aspects: content-dependent degradation and content-independent one. An existing full-reference image quality assessment (IQA) metric cannot deal with these two different impacts well. Singular value decomposition (SVD) as a useful mathematical tool has been used in various image processing applications. In this paper, SVD is employed to separate the structural (content-dependent) and the content-independent components. For each portion, we design a specific assessment model to tailor for its corresponding distortion properties. The proposed models are then fused to obtain the final quality score. Experimental results with the TID database demonstrate that the proposed metric achieves better performance in comparison with the relevant state-of-the-art quality metrics.
Shuigen Wang, Chenwei Deng, Weisi Lin, Baojun Zhao, Jie Chen 0026
ICIP2
2013 To exploit uncertainty masking for adaptive image rendering
abstract
For high-quality image rendering using Monte Carlo methods, a large number of samples are required to be computed for each pixel. Adaptive sampling aims to decrease the total number of samples by concentrating samples on difficult regions. However, existing adaptive sampling schemes haven't fully exploited the potential of image regions with complex structures to the reduction of sample numbers. To solve this problem, we propose to exploit uncertainty masking in adaptive sampling. Experimental results show that incorporation of uncertainty information leads to significant sample reduction and therefore time-savings.
Lu Dong 0001, Weisi Lin, Chenwei Deng, Ce Zhu, Seah Hock Soon
ISCAS3
2013 A saliency detection model based on sparse features and visual acuity
abstract
In this paper, we propose a novel computational model of visual attention based on the relevant characteristics of the Human Visual System (HVS). The input image is firstly divided into small image patches. Then the sparse features for each image patch are extracted based on the learned sparse coding basis. The human visual acuity is adopted in the calculation of the center-surround feature differences for saliency detection. In addition, the neighboring image patches for computing the saliency value of each center image patch are selected based on the characteristics of HVS. Experimental results show that the proposed saliency detection algorithm outperforms other existing schemes tested with a large public image database.
Yuming Fang 0001, Weisi Lin, Zhenzhong Chen 0001, Chia-Wen Lin, Zhijun Fang 0001, Chenwei Deng
ISCAS6
2013 Overview of quality assessment for visual signals and newly emerged trends
abstract
Quality assessment is not only essential on its own for testing, optimizing, benchmarking, monitoring and inspecting related systems and services, but also plays an essential role in the design of virtually all visual signal processing and communication algorithms, as well as various related decision making processes. In the paper, we provide an overview of the quality assessment approaches for traditional visual signals, as well as the newly emerged ones, which covers the subjective quality evaluation and objective quality metrics of scalable and mobile videos, high dynamic range (HDR) images, image segmentation results, 3D images/videos, and retargeted images. Also the challenges for designing effective quality metrics and corresponding applications are discussed.
Lin Ma 0002, Chenwei Deng, King Ngi Ngan, Weisi Lin
ISCAS2
2013 Adaptive parameter estimation for total variation image denoising
abstract
In this paper, we propose an adaptive parameter estimation algorithm for total variation image denoising. The de-noising framework consists of two-stage regularization parameter estimation. Firstly, we consider the fidelity of denoised image, and model a convex optimization function of denoised result. Under the results of fast gradient projection (FGP) method with a series of regularization parameters, the convex function converges to an optimal solution, which corresponds to the firststage optimal value of regularization parameter. Second, considering parameter estimation error and noise sensitivity, we build an iterative link between the dual approach function and regularization parameter. At the end of iteration, the regularization parameter reaches a stable value while the corresponding denoised result has a better visual quality. Comparing with several state-of-the-art algorithms, a large number of numerical experiments confirm that the proposed parameter estimation is highly effective, and the final denoised image has a good performance in PSNR and SSIM, especially in low SNR environment.
Baoxian Wang, Baojun Zhao, Chenwei Deng, Linbo Tang
ISCAS3
2012 Study of subjective and objective quality assessment of retargeted images
abstract
This paper presents the result of a recent large-scale subjective study of image retargeting quality on a collection of images generated by several representative image retargeting methods. Owning to many approaches to image retargeting that are developed, there is a need for a diverse independent public database of the retargeted images and the corresponding subjective scores that is freely available. We build an image retargeting quality database, in which 171 retargeted images (obtained from 57 natural source images of different contents) were generated by several representative image retargeting methods. The perceptual quality of each image is evaluated by at least 30 human subjects and the mean opinion scores (MOS) were recorded. Furthermore, several publicly available quality metrics for the retargeted images are evaluated on the built database. The database is made available [1] to the research community in order to further research on the perceptual quality assessment of the retargeted images.
Lin Ma 0002, Weisi Lin, Chenwei Deng, King Ngi Ngan
ISCAS3
2012 Learning based screen image compression
abstract
There are usually two components in computer screen images: textual and pictorial parts. The pictorial part can be compressed efficiently by classical coding approaches (e.g. JPEG, JPEG2000), while the compression of the textual part is still far away from being satisfactory for the reason that the textual content is usually of high-frequency. In this paper, a learning approach is used to construct a tailored dictionary for text representation. Based on the learned dictionary, a novel screen image compression algorithm is proposed through adopting different basis functions for the textual and pictorial components respectively. The screen images are firstly segmented into textual and pictorial parts. Then we employ traditional discrete cosine transformation (DCT) to facilitate the compression of pictorial part, while the learned dictionary is used to represent the textual part in screen images. Experimental results demonstrate the effectiveness of the proposed compression algorithm.
Huan Yang 0001, Weisi Lin, Chenwei Deng
MMSP3
2012 DECA: Recovering fields of physical quantities from incomplete sensory data
abstract
Although wireless sensor networks (WSNs) are powerful in monitoring physical events, the data collected from a WSN are almost always incomplete if the surveyed physical event spreads over a wide area. The reason for this incompleteness is twofold: i) insufficient network coverage and ii) data aggregation for energy saving. Whereas the existing recovery schemes only tackle the second aspect, we develop Dual-lEvel Compressed Aggregation (DECA) as a novel framework to address both aspects. Specifically, DECA allows a high fidelity recovery of a widespread event, under the situations that the WSN only sparsely covers the event area and that an in-network data aggregation is applied for traffic reduction. Exploiting both the low-rank nature of real-world events and the redundancy in sensory data, DECA combines matrix completion with a fine-tuned compressed sensing technique to conduct a dual-level reconstruction process. We demonstrate that DECA can recover a widespread event with less than 5% of the data (with respect to the dimension of the event) being collected. Performance evaluation based on both synthetic and real data sets confirms the recovery fidelity and energy efficiency of our DECA framework.
Liu Xiang, Jun Luo 0001, Chenwei Deng, Athanasios V. Vasilakos, Weisi Lin
SECON3
2012 Content-Based Image Compression for Arbitrary-Resolution Display Devices
abstract
The existing image coding methods cannot support content-based spatial scalability with high compression. In mobile multimedia communications, image retargeting is generally required at the user end. However, content-based image retargeting (e.g., seam carving) is with high computational complexity and is not suitable for mobile devices with limited computing power. The work presented in this paper addresses the increasing demand of visual signal delivery to terminals with arbitrary resolutions, without heavy computational burden to the receiving end. In this paper, the principle of seam carving is incorporated into a wavelet codec (i.e., SPIHT ). For each input image, block-based seam energy map is generated in the pixel domain. In the meantime, multilevel discrete wavelet transform (DWT) is performed. Different from the conventional wavelet-based coding schemes, DWT coefficients here are grouped and encoded according to the resultant seam energy map. The bitstream is then transmitted in energy descending order. At the decoder side, the end user has the ultimate choice for the spatial scalability without the need to examine the visual content; an image with arbitrary aspect ratio can be reconstructed in a content-aware manner based upon the side information of the seam energy map. Experimental results show that, for the end users, the received images with an arbitrary resolution preserve important content while achieving high coding efficiency for transmission.
Chenwei Deng, Weisi Lin, Jianfei Cai 0001
IEEE Trans. Multim.1
2012 Robust Image Coding Based Upon Compressive Sensing
abstract
Multiple description coding (MDC) is one of the widely used mechanisms to combat packet-loss in non-feedback systems. However, the number of descriptions in the existing MDC schemes is very small (typically 2). With the number of descriptions increasing, the coding complexity increases drastically and many decoders would be required. In this paper, the compressive sensing (CS) principles are studied and an alternative coding paradigm with a number of descriptions is proposed based upon CS for high packet loss transmission. Two-dimentional discrete wavelet transform (DWT) is applied for sparse representation. Unlike the typical wavelet coders (e.g., JPEG 2000), DWT coefficients here are not directly encoded, but re-sampled towards equal importance of information instead. At the decoder side, by fully exploiting the intra-scale and inter-scale correlation of multiscale DWT, two different CS recovery algorithms are developed for the low-frequency subband and high-frequency subbands, respectively. The recovery quality only depends on the number of received CS measurements (not on which of the measurements that are received). Experimental results show that the proposed CS-based codec is much more robust against lossy channels, while achieving higher rate-distortion (R-D) performance compared with conventional wavelet-based MDC methods and relevant existing CS-based coding schemes.
Chenwei Deng, Weisi Lin, Bu-Sung Lee, Chiew Tong Lau
IEEE Trans. Multim.1
2011 Content-Based Image Compression for Arbitrary-Resolution Display Devices
abstract
The work presented in this paper addresses the increasing demand of visual signal delivery to terminals with arbitrary resolutions (like in mobile multimedia communication and cloud computing) for universal access and presentation, without extra computational burden to the receiving end. Scalable image compression and transmission are essential, and to be effective and meaningful, it has to be content-based. The existing coding methods cannot support content-based spatial scalability with high compression. In this paper, the principle of seam carving (SC) is incorporated into a wavelet codec. After multi-level discrete wavelet transform (DWT), SC is performed in the low frequency subband. Different from the conventional wavelet-based coding schemes, DWT coefficients here are encoded and transmitted according to the energy map of resultant seams. At the decoder side, the end user has the ultimate choice for the scalability without the need to examine the visual content; an image with arbitrary aspect ratio can be reconstructed in a content-aware manner based upon the encoded information. Simulation results show that the resized images preserve important content while achieving high coding efficiency in transmission.
Chenwei Deng, Weisi Lin, Jianfei Cai 0001
ICC1
2011 Performance analysis, parameter selection and extensions to H.264/AVC FRExt for high resolution video coding
Chenwei Deng, Weisi Lin, Bu-Sung Lee, Chiew Tong Lau, Ming-Ting Sun
J. Vis. Commun. Image Represent.1
2011 Optimal Compression Plane for Efficient Video Coding
abstract
All existing video coding standards developed so far deem video as a sequence of natural frames (formed in the XY plane), and treat spatial redundancy (redundancy along X and Y directions) and temporal redundancy (redundancy along T direction) differently and separately. In this paper, we investigate into a new compression (redundancy reduction) method for video in which the frames are allowed to be formed in a non-XY plane. We are to exploit fuller extent of video redundancy, and propose an adaptive optimal compression plane determination process to be used as a preprocessing step prior to any standard video coding scheme. The essence of the scheme is to form the frames in the plane formed by two axes (among X, Y, and T) corresponding to signal correlation evaluation, which enables better prediction (therefore better compression). In spite of the simplicity of the proposed method, it can be used for both lossless and lossy compression, and with and without interframe prediction. Extensive experimental results show that the new coding method improves the performance of the video coding for a number of coding methods (inclusive of lossless and near-lossless Motion JPEG-LS, Motion JPEG, Motion JPG2K, H.264 intraonly profile, and H.264) and videos with different visual content.
Anmin Liu, Weisi Lin, Manoranjan Paul, Fan Zhang 0093, Chenwei Deng
IEEE Trans. Image Process.5
2010 Comparison between H.264/AVC and Motion jpeg2000 for super-high definition video coding
abstract
H.264/AVC FRExt (Fidelity Range Extensions) and Motion JPEG2000 are the latest inter-frame and intra-frame video coding standards, respectively. It is well known that an inter-frame method achieves higher coding efficiency compared with an intra-frame one, and the Motion JPEG2000 has been selected for digital cinema compression. In this paper, we attempt to compare these two different schemes with theoretical and experimental analysis for super-HD (high definition) visual signals. One additional contribution of the paper is that the impact of block partition, motion search range and skipped block size for inter-frame coding is discussed. Based on the analysis, we extend the standard H.264/AVC FRExt by using larger block size and search range. The experimental results show that this extension leads to higher coding efficiency and makes the H.264/AVC FRExt more suitable for super-HD video coding.
Chenwei Deng, Weisi Lin, Bu-Sung Lee, Chiew Tong Lau, Manoranjan Paul
ICIP1
2010 Robust image compression based on compressive sensing
abstract
The existing image compression methods (e.g., JPEG2000, etc.) are vulnerable to bit-loss, and this is usually tackled by channel coding that follows. However, source coding and channel coding have conflicting requirement. In this paper, we address the problem with an alternative paradigm, and a novel compressive sensing (CS) based compression scheme is therefore proposed. Discrete wavelet transform (DWT) is applied for sparse representation, and based on the property of 2-D DWT, a fast CS measurements taking method is presented. Unlike the unequally important discrete wavelet coefficients, the resultant CS measurements carry nearly the same amount of information and have minimal effects for bit-loss. At the decoder side, one can simply reconstruct the image via l1minimization. Experimental results show that the proposed CS-based image codec without resorting to error protection is more robust compared with existing CS technique and relevant joint source channel coding (JSCC) schemes.
Chenwei Deng, Weisi Lin, Bu-Sung Lee, Chiew Tong Lau
ICME1
2010 Optimal compression plane (OCP) - A new framework for H.264 video coding
abstract
This paper presents a framework for video coding, which is compatible with the existing H.264/AVC standard (for preprocessing). Since the video sequences are of 3D data matrices and the traditional XY image frame plane is not always the best for video coding in terms of rate-distortion (RD) performance, optimal compression plane (OCP) determination models are developed to improve the coding efficiency. We design coding plane level RD cost function (in analog to the one in H.264 macroblock level) to measure the RD performance of each coding plane. We jointly consider the available computational resource and the RD performance in order to determine an adaptive OCP. The performance of the proposed coding framework is assessed with a number video sequences. Extensive experimental results show that the new coding framework can achieve better RD performance at approximately the same computational complexity, for videos with different visual content.
Anmin Liu, Weisi Lin, Manoranjan Paul, Chenwei Deng
ICME4
2010 Just Noticeable Difference for Images With Decomposition Model for Separating Edge and Textured Regions
abstract
In just noticeable difference (JND) models, evaluation of contrast masking (CM) is a crucial step. More specifically, CM due to edge masking (EM) and texture masking (TM) needs to be distinguished due to the entropy masking property of the human visual system. However, TM is not estimated accurately in the existing JND models since they fail to distinguish TM from EM. In this letter, we propose an enhanced pixel domain JND model with a new algorithm for CM estimation. In our model, total-variation based image decomposition is used to decompose an image into structural image (i.e., cartoon like, piecewise smooth regions with sharp edges) and textural image for estimation of EM and TM, respectively. Compared with the existing models, the proposed one shows its advantages brought by the better EM and TM estimation. It has been also applied to noise shaping and visual distortion gauge, and favorable results are demonstrated by experiments on different images.
Anmin Liu, Weisi Lin, Manoranjan Paul, Chenwei Deng, Fan Zhang 0093
IEEE Trans. Circuits Syst. Video Technol.4