VLDB 2026 Research / reviewers in the wild / expert
Baojun Zhao
dblp:84/10182
· DBLP profile ↗
36ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-focus image fusion based on multi-scale feature extraction and edge preservation techniques
Baojun Zhao, Fei Luo 0002, Luis Rojas Pino, Weichao Ding |
Knowl. Based Syst. | 1 |
| 2025 | A review on multi-focus image fusion using deep learning
Fei Luo 0002, Baojun Zhao, Joel Fuentes, Weichao Ding, Chunhua Gu, Luis Rojas Pino |
Neurocomputing | 2 |
| 2025 | MA-MFIF: When misaligned multi-focus Image fusion meets deep homography estimation
Baojun Zhao, Fei Luo 0002, Joel Fuentes, Weichao Ding, Chunhua Gu |
Multim. Tools Appl. | 1 |
| 2025 | Enhanced Aircraft Detection in Compressed Remote Sensing Images Using CMSFF-YOLOv8abstractObject detection is an important part of remote sensing image analysis to accurately identify key features in Earth observation data while minimizing resource usage. However, challenges such as multi-scale object representation and the identification of small-sized objects, particularly in compressed images, persist. To tackle this problem, we propose a state-of-the-art detector called Compressed Multiscale Feature Fusion YOLOv8 (CMSFF-YOLOv8) that employs a four-channel DWT coefficient image representation for preprocessing. This method processes the four coefficients of the 3-DWT result, arranging them into a 4-channel stacked image representation that retains the important spatial-frequency characteristics of the information. The improved preprocessing architecture design employs frequency domain filtering to retain important edge information, adaptive histogram equalization to adjust the intensity distribution locally, and multi-scale feature transfer to enable the extraction of important edge features at different scales. Squeeze-and-Excitation Convolution (SEConv) for channel-wise attention and Multi-Granularity Enhanced Feature Aggregation (MGEFA) for multi-level feature extraction of DWT coefficients are integrated into the backbone of the CMSFF-YOLOv8 architecture, enabling the selective prioritization of structural and edge features in complex settings. Spatial Pyramid Pooling Fast (SPPF) combined with shallow feature fusion networks (SFFNs) enhances small object detection. A custom neck design with upsampling facilitates multi-scale feature fusion. The result is mapped back to the LL3 channel in the compressed domain, reducing computational complexity while retaining critical global context. The proposed CMSFF-YOLOv8 detector, combined with a four-channel image representation and advanced preprocessing, demonstrates an efficient and accurate solution for object identification in compressed-domain remote sensing applications, achieving an F1-score of 0.8473, a recall of 0.82165, a precision of 0.8746, and mAP values of 0.90895 at IoU = 0.5 and 0.60063 at IoU = 0.75 on the HRPlanesV2 dataset. Sharan Thapa, Yuqi Han, Baojun Zhao, Senlin Luo |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Unsupervised Domain-Adaptive Object Detection via Localization Regression AlignmentabstractUnsupervised domain-adaptive object detection uses labeled source domain data and unlabeled target domain data to alleviate the domain shift and reduce the dependence on the target domain data labels. For object detection, the features responsible for classification and localization are different. However, the existing methods basically only consider classification alignment, which is not conducive to cross-domain localization. To address this issue, in this article, we focus on the alignment of localization regression in domain-adaptive object detection and propose a novel localization regression alignment (LRA) method. The idea is that the domain-adaptive localization regression problem can be transformed into a general domain-adaptive classification problem first, and then adversarial learning is applied to the converted classification problem. Specifically, LRA first discretizes the continuous regression space, and the discrete regression intervals are treated as bins. Then, a novel binwise alignment (BA) strategy is proposed through adversarial learning. BA can further contribute to the overall cross-domain feature alignment for object detection. Extensive experiments are conducted on different detectors in various scenarios, and the state-of-the-art performance is achieved; these results demonstrate the effectiveness of our method. The code will be available at: https://github.com/zqpiao/LRA. Zhengquan Piao, Linbo Tang, Baojun Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | AccLoc: Anchor-Free and two-stage detector for accurate object localization
Zhengquan Piao, Linbo Tang, Baojun Zhao |
Pattern Recognit. | 4 |
| 2022 | Modality Translation in Remote Sensing Time SeriesabstractModality translation, which aims to translate images from a source modality to a target one, has attracted a growing interest in the field of remote sensing recently. Compared to translation problems in multimedia applications, modality translation in remote sensing often suffers from inherent ambiguities, i.e., a single input image could correspond to multiple possible outputs, and the results may not be valid in the following image interpretation tasks, such as classification and change detection. To address these issues, we make the attempt to utilizing time-series data to resolve the ambiguities. We propose a novel multimodality image translation framework, which exploits temporal information from two aspects: 1) by introducing a guidance image from given temporally neighboring images in the target modality, we employ a feature mask module and transfer semantic information from temporal images to the output without requiring the use of any semantic labels and 2) while incorporating multiple pairs of images in time series, a temporal constraint is formulated during the learning process in order to guarantee the uniqueness of the prediction result. We also build a multimodal and multitemporal dataset that contains synthetic aperture radar (SAR), visible, and short-wave length infrared band (SWIR) image time series of the same scene to encourage and promote research on modality translation in remote sensing. Experiments are conducted on the dataset for two cross-modality translation tasks (SAR to visible and visible to SWIR). Both qualitative and quantitative results demonstrate the effectiveness and superiority of the proposed model. Danfeng Hong, Jocelyn Chanussot, Baojun Zhao, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Lightweight Fine-Grained Recognition Method Based on Multilevel Feature Weighted FusionabstractFine-grained recognition in remote sensing images has played a critical role in military and civil fields. Recently, with the rapid growth of convolutional neural networks (CNNs), many fine-grained recognition methods have been proposed. However, due to the large amount of parameters and computational complexity, it is difficult to apply these methods in practical applications. To this end, we propose a novel lightweight fine-grained recognition method based on multilevel feature weighted fusion. First, we design a lightweight CNN (LCNN) framework. Second, we propose a multilevel feature weighted fusion method to improve the recognition accuracy. Third, we adopt a feature channel based loss function to train the proposed model end-to-end. Experiments are conducted on the challenging remote sensing dataset MTARSI to evaluate our proposed method. The results show that the proposed method can achieve state-of-the-art performance. Linbo Tang, Baojun Zhao |
IGARSS | 3 |
| 2021 | Robust Shadow Tracking for Video SARabstractVideo synthetic aperture radar (Video SAR) has sparked lots of research attention since it provides the ability of continuously observing and tracking the object of interest. Instead of tracking a target directly, in Video SAR, it is preferred to track its shadow due to the target's instable backscattering characteristic and defocused patterns. However, the current tracking algorithms do not work well since they always suffer from the interference caused by surrounding clutters and cannot locate the target precisely. To address these issues, we propose a robust shadow tracking algorithm for Video SAR in this letter. We equip our tracker with the spatial-temporal (ST) information and saliency-based detection mechanism (SD) against the distractions and background clutters. The proposed optimization formula could be solved efficiently using the alternating direction method of multipliers (ADMM) technique and fully carried in the Fourier domain at a low computational burden. Experiments have demonstrated that the proposed algorithm performs favorably against other state-of-the-art methods. Baojun Zhao, Yuqi Han, Hongshuo Wang, Linbo Tang, Taoyang Wang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Few-shot traffic sign recognition with clustering inductive bias and random neural network
Chenwei Deng, Zhengquan Piao, Baojun Zhao |
Pattern Recognit. | 4 |
| 2020 | High-Performance Visual Tracking With Extreme Learning Machine FrameworkabstractIn real-time applications, a fast and robust visual tracker should generally have the following important properties: 1) feature representation of an object that is not only efficient but also has a good discriminative capability and 2) appearance modeling which can quickly adapt to the variations of foreground and backgrounds. However, most of the existing tracking algorithms cannot achieve satisfactory performance in both of the two aspects. To address this issue, in this paper, we advocate a novel and efficient visual tracker by exploiting the excellent feature learning and classification capabilities of an emerging learning technique, that is, extreme learning machine (ELM). The contributions of the proposed work are as follows: 1) motivated by the simplicity and learning ability of the ELM autoencoder (ELM-AE), an ELM-AE-based feature extraction model is presented, and this model can provide a compact and discriminative representation of the inputs efficiently and 2) due to the fast learning speed of an ELM classifier, an ELM-based appearance model is developed for feature classification, and is able to rapidly distinguish the object of interest from its surroundings. In addition, in order to cope with the visual changes of the target and its backgrounds, the online sequential ELM is used to incrementally update the appearance model. Plenty of experiments on challenging image sequences demonstrate the effectiveness and robustness of the proposed tracker. Chenwei Deng, Yuqi Han, Baojun Zhao |
IEEE Trans. Cybern. | 3 |
| 2020 | Blind Noisy Image Quality Assessment Using Sub-Band KurtosisabstractNoise that afflicts natural images, regardless of the source, generally disturbs the perception of image quality by introducing a high-frequency random element that, when severe, can mask image content. Except at very low levels, where it may play a purpose, it is annoying. There exist significant statistical differences between distortion-free natural images and noisy images that become evident upon comparing the empirical probability distribution histograms of their discrete wavelet transform (DWT) coefficients. The DWT coefficients of low- or no-noise natural images have leptokurtic, peaky distributions with heavy tails; while noisy images tend to be platykurtic with less peaky distributions and shallower tails. The sample kurtosis is a natural measure of the peakedness and tail weight of the distributions of random variables. Here, we study the efficacy of the sample kurtosis of image wavelet coefficients as a feature driving, an extreme learning machine which learns to map kurtosis values into perceptual quality scores. The model is trained and tested on five types of noisy images, including additive white Gaussian noise, additive Gaussian color noise, impulse noise, masked noise, and high-frequency noise from the LIVE, CSIQ, TID2008, and TID2013 image quality databases. The experimental results show that the trained model has better quality evaluation performance on noisy images than existing blind noise assessment models, while also outperforming general-purpose blind and full-reference image quality assessment methods. Chenwei Deng, Shuigen Wang, Alan C. Bovik, Guang-Bin Huang, Baojun Zhao |
IEEE Trans. Cybern. | 5 |
| 2019 | Multimodal-Temporal Fusion: Blending Multimodal Remote Sensing Images to Generate Image Series With High Temporal ResolutionabstractThis paper aims to tackle a general but interesting cross-modality problem in remote sensing community: can multimodal images help to generate synthetic images in time series and improve temporal resolution? To this end, we explore multimodal-temporal fusion, in which we attempt to leverage the availability of additional cross-modality images to simulate the missing images in time series. We propose a multimodal-temporal fusion framework, and mainly focus on two kinds of information for the simulation: intra-modal cross-modality information and inter-modal temporal information. To exploit the cross-modality information, we adopt available paired images and learn a mapping between different modality images using a deep neural network. Considering temporal dependency among time-series images, we formulate a temporal constraint in the learning to encourage temporal consistent results. Experiments are conducted on two cross-modality image simulation applications (SAR to visible and visible to SWIR), and both visual and quantitative results demonstrate that the proposed model can successfully simulate missing images with cross-modality data. Chenwei Deng, Baojun Zhao, Jocelyn Chanussot |
IGARSS | 3 |
| 2019 | GenELM: Generative Extreme Learning Machine feature representation
Chenwei Deng, Guang-Bin Huang, Baojun Zhao |
Neurocomputing | 5 |
| 2019 | Spatial-Temporal Context-Aware TrackingabstractDiscriminative correlation filters (DCFs) have recently achieved competitive performance in visual tracking benchmarks. However, most of the existing DCF trackers only consider the spatial features of the target and could hardly benefit from the inter-frame and historical information, which may degrade the tracking performance when occlusion and deformation occurs. To tackle the above-mentioned issues, in this letter, by introducing the temporal constrain into the DCF tracker, we advocate our spatial-temporal context-aware tracker. Through jointly modeling the spatial context and historical target information, our tracker could not only adapt the appearance change but also maintain a relatively stable filter due to the small target variation between inter-frames. Furthermore, we show that the proposed objective formula could be directly solved using the Alternating Direction Method of Multipliers (ADMM) technique with low computational cost. Experiments on the large-scale benchmark demonstrate that the proposed trackers perform favorably against other state-of-the-art methods. Yuqi Han, Chenwei Deng, Boya Zhao, Baojun Zhao |
IEEE Signal Process. Lett. | 4 |
| 2019 | StfNet: A Two-Stream Convolutional Neural Network for Spatiotemporal Image FusionabstractSpatiotemporal image fusion is considered as a promising way to provide Earth observations with both high spatial resolution and frequent coverage, and recently, learning-based solutions have been receiving broad attention. However, these algorithms treating spatiotemporal fusion as a single image super-resolution problem, generally suffers from the significant spatial information loss in coarse images, due to the large upscaling factors in real applications. To address this issue, in this paper, we exploit temporal information in fine image sequences and solve the spatiotemporal fusion problem with a two-stream convolutional neural network called StfNet. The novelty of this paper is twofold. First, considering the temporal dependence among image sequences, we incorporate the fine image acquired at the neighboring date to super-resolve the coarse image at the prediction date. In this way, our network predicts a fine image not only from the structural similarity between coarse and fine image pairs but also by exploiting abundant texture information in the available neighboring fine images. Second, instead of estimating each output fine image independently, we consider the temporal relations among time-series images and formulate a temporal constraint. This temporal constraint aiming to guarantee the uniqueness of the fusion result and encourages temporal consistent predictions in learning and thus leads to more realistic final results. We evaluate the performance of the StfNet using two actual data sets of Landsat-Moderate Resolution Imaging Spectroradiometer (MODIS) acquisitions, and both visual and quantitative evaluations demonstrate that our algorithm achieves state-of-the-art performance. Chenwei Deng, Jocelyn Chanussot, Danfeng Hong, Baojun Zhao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | State-Aware Anti-Drift Object TrackingabstractCorrelation filter (CF) based trackers have aroused increasing attentions in visual tracking field due to the superior performance on several datasets while maintaining high running speed. For each frame, an ideal filter is trained in order to discriminate the target from its surrounding background. Considering that the target always undergoes external and internal interference during tracking procedure, the trained tracker should not only have the ability to judge the current state when failure occurs, but also to resist the model drift caused by challenging distractions. To this end, we present a State-aware Anti-drift Tracker (SAT) in this paper, which jointly model the discrimination and reliability information in filter learning. Specifically, global context patches are incorporated into filter training stage to better distinguish the target from backgrounds. Meanwhile, a color-based reliable mask is learned to encourage the filter to focus on more reliable regions suitable for tracking. We show that the proposed optimization problem could be efficiently solved using Alternative Direction Method of Multipliers and fully carried out in Fourier domain. Furthermore, a Kurtosis-based updating scheme is advocated to reveal the tracking condition as well as guarantee a high-confidence template updating. Extensive experiments are conducted on OTB-100 and UAV-20L datasets to compare the SAT tracker with other relevant state-of-the-art methods. Both quantitative and qualitative evaluations further demonstrate the effectiveness and robustness of the proposed work. Yuqi Han, Chenwei Deng, Baojun Zhao, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2018 | Constrained Nonnegative Matrix Factorization for Robust Hyperspectral UnmixingabstractHyperspectral unmixng (HU) is an essential step for hyperspectral image (HSI) analysis. In real HSI, there often are abnormal fluctuations existing in specific bands, which can be described as sparse noise. This type of corruption will seriously disrupt the hyperspectral image quality, causing extra difficulties during unmixing process. However, the influence of sparse noise is often ignored by existing unmixing methods, which leads to the reduction of robustness and accuracy for HU tasks. Therefore, we propose a new unmixing model which takes noise corruption into consideration. By designing and imposing constraints considering the sparsity of noise, properties of endmember and abundance on nonnegative matrix factorization (NMF), the proposed method can resist the sparse noise and achieve more robust and accurate unmixing results. Adequate experiments have been conducted on both synthetic and real hyperspectral data. And the results confirm the superiority of proposed method compared to state-of-the-art methods. Chenwei Deng, Jiahui Dai, Baojun Zhao |
IGARSS | 6 |
| 2018 | Feature-Level Loss for Multispectral Pan-Sharpening with Machine LearningabstractMultispectral pan-sharpening plays an important role in providing earth observation with both high-spatial and high-spectral resolutions, and recently pan-sharpening with machine learning has been attracting broad interest. However, these algorithms minimizing the pixel-wise mean squared error, generally suffer from over-smoothed results that lack of high-frequency details in both spatial and spectral dimensions. In this paper, we propose to tackle this problem by shifting the learning loss from pixel-wise error to a higher-level feature loss. The new loss function, formulated by spatial structure similarity and spectral angle mapping, pushes the model to generate results that have similar feature representations with ground truth, rather than match with pixel-wise accuracy. Consequently, more realistic fusion results can be produced. Visual and quantitative analysis both demonstrate that our approach achieves better performance in comparison with state-of-the-art algorithms. Furthermore, experiments on high-level remote sensing task further confirm the superiority of the proposed method in real applications. Chenwei Deng, Baojun Zhao, Jocelyn Chanussot |
IGARSS | 3 |
| 2017 | An Efficient and Robust Visual Tracking via Color-Based Context Prior Model
Chun-lei Yu, Baojun Zhao, Zengshuo Zhang, Maowen Li |
ICIG (1) | 2 |
| 2017 | Adaptive feature representation for visual trackingabstractRobust feature representation plays significant role in visual tracking. However, it remains a challenging issue, since many factors may affect the experimental performance. The existing method, which combine different features by setting them equally with the fixed weight, could hardly solve the issues, due to the different statistical properties of different features across various of scenarios and attributes. In this paper, by exploiting the internal relationship among these features, we develop a robust method to construct a more stable feature representation. More specifically, we utilize a co-training paradigm to formulate the intrinsic complementary information of multi-feature template into the efficient correlation filter framework. We test our approach on challenging sequences with illumination variation, scale variation, deformation etc. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods favorably. Yuqi Han, Chenwei Deng, Zengshuo Zhang, Jiatong Li 0004, Baojun Zhao |
ICIP | 5 |
| 2017 | Cloud-cover assessment: From spectral properties to spatial domain natural scene statisticabstractCloud contamination is the most common defect leading to quality degradation in remote sensing images. Numerous cloud-cover assessment (CCA) methods have been developed in the literature. The traditional Landsat 7 CCA algorithm attempted to detect clouds by taking advantages of different spectral properties from five spectral bands. However, it suffers the weakness of omitting thin cirrus clouds and the requirement of thermal bands. In this paper, we derived an automated CCA (ACCA) model that measures statistical deviations in spatial domain between cloud and clear images. Moreover, it only conducts on panchromatic band image, which can successfully address the limitation of unavailable thermal bands for satellite missions without thermal infrared sensors on board. A database with 400 clear/cloud images is then built for performance testing. Experimental results on the database show that our approach is more consistent with ground truths than the latest Landsat 8 ACCA results. Shuigen Wang, Chenwei Deng, Baojun Zhao |
IGARSS | 6 |
| 2017 | Effective visual tracking by pairwise metric learning
Chenwei Deng, Baoxian Wang, Weisi Lin, Guang-Bin Huang, Baojun Zhao |
Neurocomputing | 5 |
| 2017 | NMF-Based Image Quality Assessment Using Extreme Learning MachineabstractNumerous state-of-the-art perceptual image quality assessment (IQA) algorithms share a common two-stage process: distortion description followed by distortion effects pooling. As for the first stage, the distortion descriptors or measurements are expected to be effective representatives of human visual variations, while the second stage should well express the relationship among quality descriptors and the perceptual visual quality. However, most of the existing quality descriptors (e.g., luminance, contrast, and gradient) do not seem to be consistent with human perception, and the effects pooling is often done in ad-hoc ways. In this paper, we propose a novel full-reference IQA metric. It applies non-negative matrix factorization (NMF) to measure image degradations by making use of the parts-based representation of NMF. On the other hand, a new machine learning technique [extreme learning machine (ELM)] is employed to address the limitations of the existing pooling techniques. Compared with neural networks and support vector regression, ELM can achieve higher learning accuracy with faster learning speed. Extensive experimental results demonstrate that the proposed metric has better performance and lower computational complexity in comparison with the relevant state-of-the-art approaches. Shuigen Wang, Chenwei Deng, Weisi Lin, Guang-Bin Huang, Baojun Zhao |
IEEE Trans. Cybern. | 5 |
| 2017 | Robust Object Tracking With Discrete Graph-Based Multiple ExpertsabstractVariations of target appearances due to illumination changes, heavy occlusions, and target deformations are the major factors for tracking drift. In this paper, we show that the tracking drift can be effectively corrected by exploiting the relationship between the current tracker and its historical tracker snapshots. Here, a multi-expert framework is established by the current tracker and its historical trained tracker snapshots. The proposed scheme is formulated into a unified discrete graph optimization framework, whose nodes are modeled by the hypotheses of the multiple experts. Furthermore, an exact solution of the discrete graph exists giving the object state estimation at each time step. With the unary and binary compatibility graph scores defined properly, the proposed framework corrects the tracker drift via selecting the best expert hypothesis, which implicitly analyzes the recent performance of the multi-expert by only evaluating graph scores at the current frame. Three base trackers are integrated into the proposed framework to validate its effectiveness. We first integrate the online SVM on a budget algorithm into the framework with significant improvement. Then, the regression correlation filters with hand-crafted features and deep convolutional neural network features are introduced, respectively, to further boost the tracking performance. The proposed three trackers are extensively evaluated on three data sets: TB-50, TB-100, and VOT2015. The experimental results demonstrate the excellent performance of the proposed approaches against the state-of-the-art methods. Jiatong Li 0004, Chenwei Deng, Dacheng Tao, Baojun Zhao |
IEEE Trans. Image Process. | 5 |
| 2016 | Textural and Gradient Feature Extraction from JPEG2000 Codestream for Airfield DetectionabstractThere has been huge growth in the use of optical images in object detection, and numerous images over satellite and aerial vehicle are typically stored and transmitted in the compressed form such as JPEG2000. However, most of the existing detection algorithms are built in the pixel domain, which leads to enlarge data volume and requires compulsory decompression before detecting. Hence, these issues may hinder their practical use in real-time and storage limited applications. In this paper, we extract textural and gradient features from JPEG2000 compressed codestream, and these robust features are utilized for coarse-to-fine airfield detection. The contributions of proposed work are as follows: 1) packet header is effectively exploited for presenting textural information of images, while gradient calculation is achieved in packet body, 2) a novel detection framework taking full advantage of compressed features is proposed. Validated by experiments with large optical images, the proposed method achieves faster computation with higher detection accuracy, in comparison with the existing relevant state-of-the-art approaches. Chenwei Deng, Baojun Zhao |
DCC | 3 |
| 2016 | Spatiotemporal reflectance fusion based on location regularized sparse representationabstractSpatiotemporal reflectance fusion plays an important role in providing earth observation with both high-spatial and high-temporal resolutions, and sparse representation is one of the popular strategies to implement spatiotemporal fusion. However, the existing methods generally suffers from instability of sparse representation for the fine and coarse image pairs. In this paper, we demonstrate that such instability can be addressed by exploiting spatial correlations among the neighboring fine images, which is mathematically formulated as a location regularized term. A fast iterative shrinkage-thresholding algorithm (FISTA) is then employed to find the optimal solution. Experimental results show that the performance of proposed method outperforms other relevant state-of-the-art fusion approaches. Chenwei Deng, Baojun Zhao |
IGARSS | 3 |
| 2016 | Gradient-based no-reference image blur assessment using extreme learning machine
Shuigen Wang, Chenwei Deng, Baojun Zhao, Guang-Bin Huang, Baoxian Wang |
Neurocomputing | 3 |
| 2016 | Fast and Accurate Spatiotemporal Fusion Based Upon Extreme Learning MachineabstractSpatiotemporal fusion is important in providing high spatial resolution earth observations with a dense time series, and recently, learning-based fusion methods have been attracting broad interest. These algorithms project image patches onto a feature space with the enforcement of a simple mapping to predict the fine resolution patches from the corresponding coarse ones. However, the sophisticated projection, e.g., sparse representation, is always computationally complex and difficult to be implemented on large patches, which cannot grasp enough local structural information in the coarse patches. To address these issues, a novel spatiotemporal fusion method is proposed in this letter, using a powerful learning technique, i.e., extreme learning machine (ELM). Unlike traditional approaches, we devote to learning a mapping function on difference images directly, rather than the sophisticated feature representation followed by a simple mapping. Characterized by good generalization performance and fast speed, the ELM is employed to achieve accurate and fast fine patches prediction. The proposed algorithm is evaluated by five actual data sets of Landsat enhanced thematic mapper plus-moderate resolution imaging spectroradiometer acquisitions and experimental results show that our method obtains better fusion results while achieving much greater speed. Chenwei Deng, Shuigen Wang, Guang-Bin Huang, Baojun Zhao, Paula Lauren |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2016 | Time Varying Metric Learning for visual tracking
Jiatong Li 0004, Baojun Zhao, Chenwei Deng |
Pattern Recognit. Lett. | 2 |
| 2015 | Kurtosis-Based Blind Noisy Image Quality Assessment in Wavelet DomainabstractNoise distortions introduced in natural images generally break the initial probability distributions by dispersing image pixels randomly. We found that there exists a big difference between the distributions of Discrete Wavelet Transform (DWT) coefficients of natural images and noisy images: (1) for natural images, their distributions are sharp with high peaked ness and slight tail, (2) for noisy images, the shapes are much flatter with lower peaked ness and heavier tail. Kurtosis is able to measure and differentiate the probability distributions of noisy images with various noise levels. Moreover, the kurtosis values of DWT coefficients are stable for varying frequency filters. In this paper, we propose a Blind Noisy Image Quality Assessment model using Kurtosis (BNIQAK). Five types of noisy images in the three biggest databases are taken for testing BNIQAK. Experimental results show that BNIQAK has better evaluation performance compared with existing blind noisy models, as well as some general blind and full-reference (FR) methods. Shuigen Wang, Chenwei Deng, Baojun Zhao |
SMC | 5 |
| 2015 | Compressed-Domain Ship Detection on Spaceborne Optical Image Using Deep Neural Network and Extreme Learning MachineabstractShip detection on spaceborne images has attracted great interest in the applications of maritime security and traffic control. Optical images stand out from other remote sensing images in object detection due to their higher resolution and more visualized contents. However, most of the popular techniques for ship detection from optical spaceborne images have two shortcomings: 1) Compared with infrared and synthetic aperture radar images, their results are affected by weather conditions, like clouds and ocean waves, and 2) the higher resolution results in larger data volume, which makes processing more difficult. Most of the previous works mainly focus on solving the first problem by improving segmentation or classification with complicated algorithms. These methods face difficulty in efficiently balancing performance and complexity. In this paper, we propose a ship detection approach to solving the aforementioned two issues using wavelet coefficients extracted from JPEG2000 compressed domain combined with deep neural network (DNN) and extreme learning machine (ELM). Compressed domain is adopted for fast ship candidate extraction, DNN is exploited for high-level feature representation and classification, and ELM is used for efficient feature pooling and decision making. Extensive experiments demonstrate that, in comparison with the existing relevant state-of-the-art approaches, the proposed method requires less detection time and achieves higher detection accuracy. Jiexiong Tang, Chenwei Deng, Guang-Bin Huang, Baojun Zhao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2014 | Content-Adaptive Rain and Snow Removal Algorithms for Single Image
Shujian Yu, Yixiao Zhao, Yi Mou, Jinghui Wu, Xiaopeng Yang 0002, Baojun Zhao |
ISNN | 7 |
| 2014 | Data-Driven Bridge Detection in Compressed Domain from Panchromatic Satellite Imagery
Yixiao Zhao, Shujian Yu, Jinghui Wu, Zijing Chen, Xiaopeng Yang 0002, Baojun Zhao |
ISNN | 7 |
| 2013 | A novel SVD-based image quality assessment metricabstractImage distortion can be categorized into two aspects: content-dependent degradation and content-independent one. An existing full-reference image quality assessment (IQA) metric cannot deal with these two different impacts well. Singular value decomposition (SVD) as a useful mathematical tool has been used in various image processing applications. In this paper, SVD is employed to separate the structural (content-dependent) and the content-independent components. For each portion, we design a specific assessment model to tailor for its corresponding distortion properties. The proposed models are then fused to obtain the final quality score. Experimental results with the TID database demonstrate that the proposed metric achieves better performance in comparison with the relevant state-of-the-art quality metrics. Shuigen Wang, Chenwei Deng, Weisi Lin, Baojun Zhao, Jie Chen 0026 |
ICIP | 4 |
| 2013 | Adaptive parameter estimation for total variation image denoisingabstractIn this paper, we propose an adaptive parameter estimation algorithm for total variation image denoising. The de-noising framework consists of two-stage regularization parameter estimation. Firstly, we consider the fidelity of denoised image, and model a convex optimization function of denoised result. Under the results of fast gradient projection (FGP) method with a series of regularization parameters, the convex function converges to an optimal solution, which corresponds to the firststage optimal value of regularization parameter. Second, considering parameter estimation error and noise sensitivity, we build an iterative link between the dual approach function and regularization parameter. At the end of iteration, the regularization parameter reaches a stable value while the corresponding denoised result has a better visual quality. Comparing with several state-of-the-art algorithms, a large number of numerical experiments confirm that the proposed parameter estimation is highly effective, and the final denoised image has a good performance in PSNR and SSIM, especially in low SNR environment. Baoxian Wang, Baojun Zhao, Chenwei Deng, Linbo Tang |
ISCAS | 2 |