EDBT 2026 Demo / reviewers in the wild / expert
Chunping Hou
dblp:72/4265
· DBLP profile ↗
102ranked-venue papers
0as first author
20since 2021 · last 2024
0000-0001-8772-0198ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 77 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 8 since 2021Artificial intelligence and machine learning · 10 · 2 since 2021Computer networks · 4Systems, architecture and hardware · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multi-resolution feature perception network for UAV person re-identification
Meiyan Huang, Chunping Hou, Xuebo Zheng |
Multim. Tools Appl. | 2 |
| 2024 | WSPTGAN for Global Ocean Surface Wind Speed Generation With High Temporal Resolution and Spatial CoverageabstractObtaining global ocean surface wind speed data with high temporal resolution and spatial coverage is a challenging task. Due to the lack of widely applicable direct measurement methods and algorithms, current research and data products can only achieve good performance in a small spatial range or at low temporal resolution. In this article, a generative adversarial network (GAN) with a transformer structure called Wind Speed Prediction transformer-GAN (WSPTGAN) is proposed to generate wind speed data with good spatial coverage and high temporal resolution for areas. The WSPTGAN is trained with the proposed image-like wind speed data combined partial missing dataset (CPMD), which is combined with the fifth generation of the European Center for Medium-Range Weather Forecast (ECMWF) reanalysis data and Advanced Scatterometer (ASCAT) data from Meteorological Operational satellites. Thanks to the defective data learning mechanism (DDLM), sequential-wise multihead self-attention mechanism (SMSM), and sequence feature adaptive verification mechanism (SFAVM) in the proposed algorithm, the obtained model has good wind speed prediction accuracy with root mean square error (RMSE) of 0.8984 m/s and can achieve multistep 10-min wind speed data generation within the global ocean. After comparison with five state-of-the-art prediction models, it is confirmed that the algorithm in this article is able to make better use of the defective data for learning and prediction of wind field trends in global ocean regions. Yonghong Hou, Xiaowei Song 0001, Chunping Hou, Zixiang Xiong, Dan Ma 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Self-Attention-Guided Multiindicator Retrieval for Ocean Surface Wind Field With Multimodal Data Augmentation and FusionabstractThe deployment of global navigation satellite system reflectometry (GNSS-R) emerges as a compelling approach for the extraction of ocean surface wind field, primarily due to its exceptional cost-effectiveness, all-weather robustness, and excellent spatiotemporal coverage. Despite these advantages, the insufficient use of various data and the lack of ability to perform multiindicator retrieval limit the performance of existing methods in practical ocean wind field retrieval. To overcome these limitations, this article introduces a novel self-attention-guided ocean surface wind field multiindicator retrieval algorithm based on multimodal observation data augmentation and fusion. Initially, data generation modules are employed to complement high-quality observation data that are not fully provided by GNSS-R system. Subsequently, the multiscale data fusion encoder (MDFE) is implemented to extract and fuse multiple data features of different scales to enhance the utilization ability of data and improve the accuracy of wind field retrieval. Finally, the self-attention multiindicator predictor (SAMIP) is put to use for optimizing the feature attention strategy, achieving accurate retrieval of ocean surface wind speed and direction simultaneously. The proposed method provides a novel solution for the comprehensive utilization of various data products in the GNSS-R system, simultaneously achieves synchronous retrieval of multiple wind field indicators, which are wind speed and direction. In the context of recent advancements in algorithmic development, the proposed algorithm exhibits a notable enhancement in the precision of wind speed and direction retrieval. Benchmarking against ERA5 wind field data, the proposed algorithm achieves a root-mean-square error (RMSE) of 1.23 m/s in wind speed retrieval, demonstrating at least a 9% accuracy improvement compared to five state-of-the-art algorithms from recent years. Furthermore, the RMSE in wind direction retrieval stands at 20.7°, surpassing comparison algorithms by achieving a reduction of over 8% in error. Collectively, these metrics robustly validate the efficacy of the proposed algorithm. Yonghong Hou, Xiaowei Song 0001, Chunping Hou, Zixiang Xiong, Dan Ma 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | A Density and Distance-Based Method for ICESat-2 Photon-Counting Data DenoisingabstractThe Ice, Cloud, and Land Elevation Satellite-2 (ICESat-2) dynamically monitors water depth in shallow waters around islands and reefs. Noise removal is a prerequisite for accurate reconstruction of seafloor topography based on ICESat-2 data products. To this end, we propose a density and distance-based method (DDBM) to extract seafloor signal photons from ICESat-2 data. The DDBM first separates the photons into three parts: above water, water surface, and water column. The water-column photons consist of seafloor signal photons and noise photons. The DDBM adopts a two-step denoising strategy to remove noise in water-column photons to obtain pure and complete signal photons. In the first step, an orientation-variable adaptive ellipse filter is developed, which can adaptively adjust the parameters according to the water depth to remove low-density noise photons. In the second step, a novel distance-based filter (DBF) is designed for stubborn high-density noise clusters. These noise clusters are far from the signal photons, and the DBF removes them by a distance threshold derived from the Otsu threshold method. We select three high-density ICESat-2 datasets to validate the DDBM. Compared with the reference data, the comprehensive evaluation indexes$F$of the DDBM in all datasets are above 0.99, and the highest is 0.998, showing superior performance. Xuebo Zheng, Chunping Hou, Meiyan Huang, Dan Ma 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Locality-Aware Transformer for Video-Based Sign Language TranslationabstractRecently, the application of transformer makes significant progress in sign language translation. However, several characteristics of sign videos are neglected in existing transformer-based methods that hinder translation performance. Firstly, in sign videos, multiple consecutive frames represent a single sign gloss thus the local temporal relations are crucial. Secondly, the inconsistency between video and text demands the non-local and global context modeling ability of the model. To address these issues, a locality-aware transformer is proposed for sign language translation. Concretely, the multi-stride position encoding scheme assigns the same position index to adjacent frames with various strides to enhance the local dependency. Afterward, the adaptive temporal interaction module is utilized to capture non-local and flexible local frame correlation simultaneously. Moreover, a gloss counting task is designed to facilitate the holistic understanding of sign videos. Experimental results on two benchmark datasets demonstrate the effectiveness of the proposed framework. Zihui Guo, Yonghong Hou, Chunping Hou |
IEEE Signal Process. Lett. | 3 |
| 2023 | Reasoning and Tuning: Graph Attention Network for Occluded Person Re-IdentificationabstractOccluded person re-identification (re-id) aims to match occluded person images to holistic ones. Most existing works focus on matching collective-visible body parts by discarding the occluded parts. However, only preserving the collective-visible body parts causes great semantic loss for occluded images, decreasing the confidence of feature matching. On the other hand, we observe that the holistic images can provide the missing semantic information for occluded images of the same identity. Thus, compensating the occluded image with its holistic counterpart has the potential for alleviating the above limitation. In this paper, we propose a novel Reasoning and Tuning Graph Attention Network (RTGAT), which learns complete person representations of occluded images by jointly reasoning the visibility of body parts and compensating the occluded parts for the semantic loss. Specifically, we self-mine the semantic correlation between part features and the global feature to reason the visibility scores of body parts. Then we introduce the visibility scores as the graph attention, which guides Graph Convolutional Network (GCN) to fuzzily suppress the noise of occluded part features and propagate the missing semantic information from the holistic image to the occluded image. We finally learn complete person representations of occluded images for effective feature matching. Experimental results on occluded benchmarks demonstrate the superiority of our method. Meiyan Huang, Chunping Hou |
IEEE Trans. Image Process. | 2 |
| 2022 | Dynamic-boosting attention for self-supervised video representation learning
Chunping Hou, Guanghui Yue 0001 |
Appl. Intell. | 2 |
| 2022 | Instance interactive association graph convolutional network for domain adaptive person re-identification
Chunping Hou, Meiyan Huang |
Appl. Intell. | 2 |
| 2022 | Unsupervised anomaly detection via dual transformation-aware embeddingsabstractAbstract Unsupervised anomaly detection refers to the discovery of unconventional images that are globally or locally different from the training set. Recently, reconstruction‐based anomaly detection methods have made great progress. However, most of the existing methods take reconstructing the original image as the goal of latent feature learning. Due to lack of effective semantic guidance, latent features have intrinsic characteristics which retain redundant details of spatial structure. Such information is too general and cause over‐expression problem. To solve this problem, in this paper, dual transformation‐aware embeddings are coined which aims to achieve a stable model to learn high‐level latent features in a self‐supervised manner. To be more specific, the authors try to extract transformation‐detectable feature embeddings for both structure and content views which explore the regular pattern under different transformations in normal situations. In addition, the relationship between the original feature and the transformed feature is established. Based on such relationship, the latent feature of generated image to predict transformation parameter is extracted. Then, a transformation‐consistency regularization is proposed to constrain decoder to generate high‐quality image with high‐level consistency and achieve a more stable model. Experiments on MVTec‐AD and CIFAR10 datasets prove the effectiveness and robustness of the proposed method. Chunping Hou, Bangbang Ge, Zhicheng Dong 0003, Zhiqiang Wu 0001 |
IET Image Process. | 2 |
| 2022 | Cloud Detection From Remote Sensing Imagery Based on Domain Translation NetworkabstractCloud detection in optical imagery has drawn remarkable attention in the era of big Earth observation data analytic. While multiple supervised learning models have been developed for such purpose, large volumes of paired training samples annotated at the pixel level are essential to ensure the model’s generalization capacity. However, constructing a comprehensive cloud detection training database is a tedious and time-consuming process. To tackle this dilemma, we simply regard cloud-contaminated remote sensing (RS) imagery as the combination of cloud and background domains and propose a cloud detection framework based on image-to-image domain translation network (DTNet) to separate cloud-contaminated RS imagery into two target domains of cloud and background object images without using any paired and pixel-level annotation training data. The framework was evaluated with multispectral images from two types of sensors, Landsat-8 Operational Land Imager (OLI) (30 m) and GaoFen-1 (16 m), and demonstrated superior or comparable performance compared with several state-of-the-art cloud detection models. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Yang Chen 0015, Chunping Hou, Kun Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Body size measurement based on deep learning for image segmentation by binocular stereovision system
Xiaowei Song 0001, Xianli Song, Lei Yang 0050, Chunping Hou, Zixiang Xiong |
Multim. Tools Appl. | 5 |
| 2022 | SISO Radar-Based Human Movement Direction Determination Using Micro-Doppler SignaturesabstractHuman movement direction determination (HMDD) is a significant task in human detection and recognition applications, but it remains a challenge when utilizing single-input and single-output (SISO) radar because angle information cannot be accessed without multiple receiving antennas. Moreover, adopting multiple-output radar systems for this task would limit their applicability in a broader range of scenarios, as these systems require a larger placement area and a more extensive calibration procedure than SISO radar. Tackling this problem, this paper presents an effective method for the SISO-radar-based human movement direction determination task. Our method combines the joint time-frequency analysis (JTFA) technique with a proposed bio-inspired feature extraction process, thereby producing an accurate perception of moving direction based on the analysis of micro-Doppler signatures. The radar simulation and measurements are separately conducted and used to establish the corresponding dataset, so the superior performance of our method could be verified on the HMDD task. Furthermore, this article delves into why existing criteria for evaluating an HMDD model’s performance are insufficient in multi-direction situations, followed by an introduction of “small error concentration" and “omnidirectional error uniformity", as well as their evaluation protocols, to describe and measure the bias of an HMDD model’s results on multi-direction determination problems. By comparing with existing both traditional and deep-learning methods, we confirm our method’s superior in the HMDD task, and we believe that our research will aid in the advancement of human detection and recognition applications using SISO radar. Chunying Song, Yang Yang 0045, Yue Lang, Chunping Hou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Deep edge map guided depth super resolution
Zhongyu Jiang, Huanjing Yue, Yukun Lai, Jing-Yu Yang 0002, Yonghong Hou, Chunping Hou |
Signal Process. Image Commun. | 6 |
| 2021 | Deep noise estimation and removal for real-world noisy images
Huanjing Yue, Zhongyu Jiang, Shengdi Zhou, Jing-Yu Yang 0002, Yonghong Hou, Chunping Hou |
Signal Process. Image Commun. | 6 |
| 2021 | Reference guided image super-resolution via efficient dense warping and adaptive fusion
Huanjing Yue, Zhongyu Jiang, Jing-Yu Yang 0002, Chunping Hou |
Signal Process. Image Commun. | 5 |
| 2021 | A One-Class Classification Method for Human Gait Authentication Using Micro-Doppler SignaturesabstractIn this letter, a radar-based gait authentication method is proposed. We focus on the overfitting problem on the target category caused by limited training data in authentication models and propose a one-class classification model to alleviate this problem. The effectiveness of such model is verified by establishing a radar-based gait dataset, which is composed of gait micro-Doppler spectrograms derived from nine human subjects. The experimental results demonstrate that, under the condition of limited training data, the performances of an authentication model degrade because misclassification of the non-target samples easily occurs. The proposed method effectively avoids this risk, performing the other existing authentication and one-class classification methods on the metric Equal Error Rate. Haoran Ji, Chunping Hou, Yang Yang 0045, Francesco Fioranelli, Yue Lang |
IEEE Signal Process. Lett. | 2 |
| 2021 | Recaptured Screen Image DemoiréingabstractIn many situations, such as transferring data between devices and recording precious moments, we would like to capture the contents on screens using digital cameras for convenience. These recaptured screen images and videos suffer from a special type of degradation called “moiré pattern”, which is caused by the aliasing between the grid of display screen and the array of camera sensor. However, few works are proposed to tackle this problem. Considering the great success of convolutional neural networks (CNNs) in image restoration, we propose a CNN-based moiré removal method for recaptured screen images. There are mainly two contributions in this paper. First, for the generation of training data, we propose an image registration algorithm via global homography transform and local patch matching to compensate the significant viewpoint disparity between the recaptured screen image and the moiré-free image obtained via screenshot. We construct a moiré removal and brightness improvement (MRBI) database with aligned moiré-free and moiré images. Second, we propose a convolutional neural Network with Additive and Multiplicative modules (termed as AMNet) to transfer the low light moiré image to the bright moiré-free image. The proposed network is trained with pixel-wise loss, perceptual loss, and adversarial loss. Extensive experiments on 340 test images demonstrate that the proposed method outperforms state-of-the-art moiré removal methods. Huanjing Yue, Lipu Liang, Hongteng Xu, Chunping Hou, Jing-Yu Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Landsat-8 OLI Multispectral Image Dehazing Based on Optimized Atmospheric Scattering ModelabstractOptical satellite images are often affected by haze atmospheric conditions, which degrades the quality of remote sensing (RS) data and reduces the accuracy of interpretation and classification. Hence, haze removal becomes a necessary preprocessing step for most of the applications of RS image. In this article, we propose a novel haze removal method for Landsat-8 OLI multispectral image based on an optimized atmospheric scattering model. We focus on adaptively estimating the haze transmission map of each band by taking into account the effect of both wavelength and haze atmospheric conditions (haze particle size and haze particle concentration) thus improving dehazing performance. The experimental results on Landsat-8 OLI multispectral images show that the proposed dehazing model is able to remove haze successfully and significantly improve the image visibility as well as correct the spectral bias to some degree. Moreover, this method is simple and feasible, and has good practical value. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | CDnetV2: CNN-Based Cloud Detection for Remote Sensing Imagery With Cloud-Snow CoexistenceabstractCloud detection is a crucial preprocessing step for optical satellite remote sensing (RS) images. This article focuses on the cloud detection for RS imagery with cloud-snow coexistence and the utilization of the satellite thumbnails that lose considerable amount of high resolution and spectrum information of original RS images to extract cloud mask efficiently. To tackle this problem, we propose a novel cloud detection neural network with an encoder-decoder structure, named CDnetV2, as a series work on cloud detection. Compared with our previous CDnetV1, CDnetV2 contains two novel modules, that is, adaptive feature fusing model (AFFM) and high-level semantic information guidance flows (HSIGFs). AFFM is used to fuse multilevel feature maps by three submodules: channel attention fusion model (CAFM), spatial attention fusion model (SAFM), and channel attention refinement model (CARM). HSIGFs are designed to make feature layers at decoder of CDnetV2 be aware of the locations of the cloud objects. The high-level semantic information of HSIGFs is extracted by a proposed high-level feature fusing model (HFFM). By being equipped with these two proposed key modules, AFFM and HSIGFs, CDnetV2 is able to fully utilize features extracted from encoder layers and yield accurate cloud detection results. Experimental results on the ZY-3 satellite thumbnail data set demonstrate that the proposed CDnetV2 achieves accurate detection accuracy and outperforms several state-of-the-art methods. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | RSDehazeNet: Dehazing Network With Channel Refinement for Multispectral Remote Sensing ImagesabstractMultispectral remote sensing (RS) images are often contaminated by the haze that degrades the quality of RS data and reduces the accuracy of interpretation and classification. Recently, the emerging deep convolutional neural networks (CNNs) provide us new approaches for RS image dehazing. Unfortunately, the power of CNNs is limited by the lack of sufficient hazy-clean pairs of RS imagery, which makes supervised learning impractical. To meet the data hunger of supervised CNNs, we propose a novel haze synthesis method to generate realistic hazy multispectral images by modeling the wavelength-dependent and spatial-varying characteristics of haze in RS images. The proposed haze synthesis method not only alleviates the lack of realistic training pairs in multispectral RS image dehazing but also provides a benchmark data set for quantitative evaluation. Furthermore, we propose an end-to-end RSDehazeNet for haze removal. We utilize both local and global residual learning strategies in RSDehazeNet for fast convergence with superior performance. Channel attention modules are incorporated to exploit strong channel correlation in multispectral RS images. Experimental results show that the proposed network outperforms the state-of-the-art methods for synthetic data and real Landsat-8 OLI multispectral RS images. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | 3D Motion Recovery via Low Rank Matrix Restoration with Hankel-Like AugmentationabstractThis paper proposes a 3D skeleton recovery model equipped with a joint augmented low-rank and sparse prior and an articulation-graph-based isometric constraint to exploit temporal and spatial correlation, respectively. The corrupted 3D skeleton sequence is represented as a matrix and constrained by several priors in the proposed model. A Hankel-like augmentation is adopted to strengthen the low-rankness and we integrate a decoupling technique to reduce the internal interferences of the data. We solve our model via an alternating direction method under the augmented Lagrangian multiplier framework with a Gauss-Newton solver for the subproblem of isometric optimization. Experimental results on two skeleton datasets demonstrate the effectiveness and superiority of the proposed model in motion reconstruction and skeleton recovery, compared with state-of-the-art methods. Jing-Yu Yang 0002, Jiabin Shi, Yuyuan Zhu, Kun Li 0001, Chunping Hou |
ICME | 5 |
| 2020 | Feature-segmentation strategy based convolutional neural network for no-reference image quality assessment
Lili Shen, Ning Hang, Chunping Hou |
Multim. Tools Appl. | 3 |
| 2020 | Spatiotemporally scalable matrix recovery for background modeling and moving object detection
Jing-Yu Yang 0002, Huanjing Yue, Kun Li 0001, Chunping Hou |
Signal Process. | 6 |
| 2020 | Adaptive Irregular Graph Construction-Based Salient Object DetectionabstractSaliency detection represents a vital pre-processing stage of computer vision. Most existing propagation-based salient object detection methods construct a k-regular graph for saliency propagation. Applying a regular graph to a vast smooth region is potentially prone to unnecessary or prolonged propagation errors, leading to the excessive highlighting of the background regions. To mitigate such problems, we substitute the conventional k-regular graph with an adaptive irregular graph for saliency value propagation, thereby avoiding unnecessary iterations over a vast smooth region. We first perform a clustering analysis based on the smoothness, color, and other features of regions. The new graph boosts an adaptive link density by considering the clustering result. In addition, we propose a seeding strategy for the propagation. Based on our experimental studies of six major benchmark datasets, our method performed favorably against the other state-of-the-art methods, both quantitatively and qualitatively. Yuan Zhou 0006, Shuwei Huo, Chunping Hou, Sun-Yuan Kung |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Omnidirectional Motion Classification With Monostatic Radar System Using Micro-Doppler SignaturesabstractIn remote sensing, micro-Doppler signatures are widely used in moving target detection and automatic target recognition. However, since Doppler signatures are easily affected by the moving direction of the target, prior information of aspect angle is essential for spectral analysis. Thus, a micro-Doppler-based classifier is considered to be “angle-sensitive.” In this article, we propose an angle-insensitive classifier for the omnidirectional classification problem using the monostatic radar through a proposed new convolutional neural network. We further provide a sensible definition of “angle sensitivity,” and perform experiments on two data sets obtained through simulations and measurements. The results demonstrate that the proposed algorithm outperforms both feature-based and existing deep-learning-based counterparts, and resolve the issue of angle sensitivity in micro-Doppler-based classification. Yang Yang 0045, Chunping Hou, Yue Lang, Takuya Sakamoto, Yuan He 0009, Wei Xiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Stereoscopic Image Stitching via Disparity-Constrained Warping and BlendingabstractAs a significant branch of virtual reality, stereoscopic image stitching aims to generating wide perspectives and natural-looking scenes. Existing 2D image stitching methods cannot be successfully applied to the stereoscopic images without considering the disparity consistency of stereoscopic images. To address this issue, this paper presents a stereoscopic image stitching method based on disparity-constrained warping and blending, which could avoid visual distortion and preserve disparity consistency. First, a point-line-driven homography based disparity minimization method is designed to pre-align the left and right images and reduce vertical disparity. Afterwards, a multi-constraint warping is proposed to further align the left and right images, where the initial disparity map is introduced to control the consistency of disparities. Finally, a disparity consistency seam-cutting and blending method is presented to determine the optimal seam and conduct stereoscopic image stitching. Experimental results demonstrate that the proposed method achieves competitive performance compared with other state-of-the-art methods. Xiaoting Fan, Jianjun Lei 0001, Yuming Fang 0001, Qingming Huang, Nam Ling, Chunping Hou |
IEEE Trans. Multim. | 6 |
| 2019 | Blind Quality Evaluator for Screen Content Images via Analysis of StructureabstractExisting blind evaluators for screen content images (SCIs) are mainly learning-based and require a number of training images with co-registered human opinion scores. However, the size of existing databases is small, and it is labor-, time-consuming and expensive to largely generate human opinion scores. In this study, we propose a novel blind quality evaluator without training. Specifically, the proposed method first calculates the gradient similarity between a distorted image and its translated versions in four directions to estimate the structural distortion, the most obvious distortion in SCIs. Given that the edge region is easier to be distorted, the inter-scale gradient similarity is then calculated as the weighting map. Finally, the proposed method is derived by incorporating the gradient similarity map with the weighting map. Experimental results demonstrate its effectiveness and efficiency on a public available SCI database. Guanghui Yue 0001, Chunping Hou, Weisi Lin |
ICASSP | 2 |
| 2019 | An End-to-End Multi-Scale Residual Reconstruction Network for Image Compressive SensingabstractRecently, deep-learning based reconstruction methods have been proposed to improve recovery performance of compressive sensed image and overcome expensive time complexity drawbacks of iteration-based traditional algorithms. In this paper, we propose an end-to-end multi-scale residual convolutional neural network (CNN), dubbed MSRNet, to simulate image compressive sensing (CS) and inverse reconstruction process in real situation. In the reconstruction stage of MSRNet, we apply three parallel channels with different convolution kernel sizes to exploit different-scale feature information. Besides, residual learning is introduced to accelerate training process and enhance prediction accuracy of network. Moreover, different from generating CS measurements by random measurement matrix in previous methods, we integrate compressive sample process into MSRNet, which means measurement matrix can be adaptively learned by training the network. Experiments on benchmark datasets show our method outperforms other state-of-the-art algorithms by large margins and set a new level for CS reconstruction with competitive time complexity. Renhe Liu, Sumei Li, Chunping Hou |
ICIP | 3 |
| 2019 | No-Reference Stereoscopic Image Quality Assessment Based On Shuffle-Convolutional Neural NetworkabstractWith the development of stereoscopic imaging technology, stereoscopic image quality assessment (SIQA) has been gaining great attention. In this work, to find a better SIQA method conforming to the perceptual characteristics of our brain, we propose a two-channel convolutional neural network (CNN) based on shuffle unit, which is called SCNN, for no-reference SIQA. The shuffle unit is used to mix up the features extracted from the left and right views to complete information communication between the two views. Different from other SIQA methods, the four shuffle units among proposed model achieve the multiple binocular fusions while processing the left and right views. Moreover, the Shuffle v2 block before the global pooling layer further improves the accuracy of SCNN. In addition, it is worth noting that we employ decorrelated batch normalization (DBN) to obtain the better generalization ability. Experimental results demonstrate that the proposed model outperforms the state-of-the-art no-reference SIQA methods. Sumei Li, Chunping Hou |
VCIP | 3 |
| 2019 | Joint Motion Classification and Person Identification via Multitask Learning for Smart HomesabstractIn a smart home environment, assisted living has been a topic of great research over the past decade. Human motion analysis is considered as a key technology for living states recognition in an assisted living system. Recent research has proved that rich information can be obtained from human movements, such as the motion category, moving patterns, and human identity. In this paper, a nonintrusive human movement sensing system is established with a mono-static ultrawide bandwidth radar. Then, we propose a well-designed joint motion classification (MCL) and person identification (PID) convolutional neural network (named as “JMI-CNN”). To recognize human motions and identities simultaneously, the network employs a multitask learning scheme as well as the attention mechanism and the hierarchical feature reuse strategies. We report the experimental result on the data from 15 individuals, each performing six motions. It shows that the model achieves a promising performance of 80.57% on the joint task, while the accuracy for MCL and PID are 98.50% and 80.92%, respectively. Moreover, we carry out ablation studies to evaluate the design principles of the proposed method. Discussions on the impact of signal noise ratio and slow-time window length are also conducted. Yue Lang, Qing Wang 0015, Yang Yang 0045, Chunping Hou, Haiping Liu, Yuan He 0009 |
IEEE Internet Things J. | 4 |
| 2019 | Unsupervised Domain Adaptation for Micro-Doppler Human Motion Classification via Feature FusionabstractMicro-Doppler-based human motion classification has become a topical area of research recently. However, the current research is limited by the lack of labeled training data. Domain adaptation, namely, the ability to take advantage of knowledge from an available source data set and apply it to an unlabeled target data set, is useful in this situation. A typical strategy for this transfer learning technique is to extract domain-invariant feature representations. In this letter, an unsupervised domain adaptation method for micro-Doppler classification is proposed. Given no available measurement training samples, we creatively utilize the motion capture database as an auxiliary and adapt its interior knowledge to the measurement data set. To achieve domain-invariant features, three types of features are extracted and fused including low-level deep features from the convolutional neural network, empirical features, and statistical features. After feature fusion, a k-nearest neighbor classifier is applied to the measurement data to classify seven human activities. Experimental results show that our approach outperforms several state-of-the-art unsupervised domain adaptation methods. The impact of the output from different convolution layers is further investigated, and ablation studies of the efficacy of each feature are also carried out in this letter. Yue Lang, Qing Wang 0015, Yang Yang 0045, Chunping Hou, Danyang Huang, Wei Xiang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2019 | Open-set human activity recognition based on micro-Doppler signatures
Yang Yang 0045, Chunping Hou, Yue Lang, Dai Guan, Danyang Huang, Jinchen Xu |
Pattern Recognit. | 2 |
| 2019 | Transferred deep learning based waveform recognition for cognitive passive radar
Qing Wang 0015, Panfei Du, Jing-Yu Yang 0002, Guohua Wang 0002, Jianjun Lei 0001, Chunping Hou |
Signal Process. | 6 |
| 2019 | Person Re-Identification by Semantic Region Representation and Topology ConstraintabstractPerson re-identification is a popular research topic which aims at matching the specific person in a multi-camera network automatically. Feature representation and metric learning are two important issues for person re-identification. In this paper, we propose a novel person re-identification method, which consists of a reliable representation called semantic region representation (SRR), and an effective metric learning with mapping space topology constraint (MSTC). The SRR integrates semantic representations to achieve effective similarity comparison between the corresponding regions via parsing the body into multiple parts, which focuses on the foreground context against the background interference. To learn a discriminant metric, the MSTC is proposed to consider the topological relationship among all samples in the feature space. It considers two-fold constraints: the distribution of positive pairs should be more compact than the average distribution of negative pairs with regard to the same probe, while the average distance between different classes should be larger than that between same classes. These two aspects cooperate to maintain the compactness of the intra-class as well as the sparsity of the inter-class. Extensive experiments conducted on five challenging person re-identification datasets, VIPeR, SYSU-sReID, QUML GRID, CUHK03, and Market-1501, show that the proposed method achieves competitive performance with the state-of-the-art approaches. Jianjun Lei 0001, Lijie Niu, Huazhu Fu, Bo Peng 0007, Qingming Huang, Chunping Hou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | An Iterative Co-Saliency Framework for RGBD ImagesabstractAs a newly emerging and significant topic in computer vision community, co-saliency detection aims at discovering the common salient objects in multiple related images. The existing methods often generate the co-saliency map through a direct forward pipeline which is based on the designed cues or initialization, but lack the refinement-cycle scheme. Moreover, they mainly focus on RGB image and ignore the depth information for RGBD images. In this paper, we propose an iterative RGBD co-saliency framework, which utilizes the existing single saliency maps as the initialization, and generates the final RGBD co-saliency map by using a refinement-cycle model. Three schemes are employed in the proposed RGBD co-saliency framework, which include the addition scheme, deletion scheme, and iteration scheme. The addition scheme is used to highlight the salient regions based on intra-image depth propagation and saliency propagation, while the deletion scheme filters the saliency regions and removes the non-common salient regions based on interimage constraint. The iteration scheme is proposed to obtain more homogeneous and consistent co-saliency map. Furthermore, a novel descriptor, named depth shape prior, is proposed in the addition scheme to introduce the depth information to enhance identification of co-salient objects. The proposed method can effectively exploit any existing 2-D saliency model to work well in RGBD co-saliency scenarios. The experiments on two RGBD co-saliency datasets demonstrate the effectiveness of our proposed framework. Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Weisi Lin, Qingming Huang, Xiaochun Cao, Chunping Hou |
IEEE Trans. Cybern. | 7 |
| 2019 | Semi-Supervised Salient Object Detection Using a Linear Feedback Control System ModelabstractTo overcome the challenging problems in saliency detection, we propose a novel semi-supervised classifier which makes good use of a linear feedback control system (LFCS) model by establishing a relationship between control states and salient object detection. First, we develop a boundary homogeneity model to estimate the initial saliency and background likelihoods, which are regarded as the labeled samples in our semi-supervised learning procedure. Then in order to allocate an optimized saliency value to each superpixel, we present an iterative semi-supervised learning framework which integrates multiple saliency cues and image features using an LFCS model. Via an innovative iteration method, the system gradually converges an optimized stable state, which is associating with an accurate saliency map. This paper also covers comprehensive simulation study based on public datasets, which demonstrates the superiority of the proposed approach. Yuan Zhou 0006, Shuwei Huo, Wei Xiang 0001, Chunping Hou, Sun-Yuan Kung |
IEEE Trans. Cybern. | 4 |
| 2019 | Video Saliency Detection via Sparsity-Based Reconstruction and PropagationabstractVideo saliency detection aims to continuously discover the motion-related salient objects from the video sequences. Since it needs to consider the spatial and temporal constraints jointly, video saliency detection is more challenging than image saliency detection. In this paper, we propose a new method to detect the salient objects in video based on sparse reconstruction and propagation. With the assistance of novel static and motion priors, a single-frame saliency model is first designed to represent the spatial saliency in each individual frame via the sparsity-based reconstruction. Then, through a progressive sparsity-based propagation, the sequential correspondence in the temporal space is captured to produce the inter-frame saliency map. Finally, these two maps are incorporated into a global optimization model to achieve spatio-temporal smoothness and global consistency of the salient object in the whole video. The experiments on three large-scale video saliency datasets demonstrate that the proposed method outperforms the state-of-the-art algorithms both qualitatively and quantitatively. Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Fatih Porikli, Qingming Huang, Chunping Hou |
IEEE Trans. Image Process. | 6 |
| 2019 | Combining Local and Global Measures for DIBR-Synthesized Image Quality EvaluationabstractDepth-Image-Based-Rendering (DIBR) techniques are significant for three-dimensional (3D) video applications, e.g., 3D television and free viewpoint video (FVV). Unfortunately, the DIBR-synthesized image suffers from various distortions, which induce an annoying viewing experience for the entire FVV. Proposing a quality evaluator for DIBR-synthesized images is fundamental for the design of perceptual friendly FVV systems. Since the associated reference image is usually not accessible, full-reference (FR) methods cannot be directly applied for quality evaluation of the synthesized image. In addition, most traditional no-reference (NR) methods fail to effectively measure the specifically DIBR-related distortions. In this paper, we propose a novel NR quality evaluation method accounting for two categories of DIBR-related distortions, i.e., geometric distortions and sharpness. First, the disoccluded regions, as one of the most obvious geometric distortions, are captured by analyzing local similarity. Then, another typical geometric distortion (i.e., stretching) is detected and measured by calculating the similarity between it and its equal-size adjacent region. Second, considering the property of scale invariance, the global sharpness is measured as the distance between the distorted image and its downsampled version. Finally, the perceptual quality is estimated by linearly pooling the scores of two geometric distortions and sharpness together. Experimental results verify the superiority of the proposed method over the prevailing FR and NR metrics. More specifically, it is superior to all competing methods except APT in terms of effectiveness, but greatly outmatches APT in terms of implementation time. Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Tianwei Zhou, Guangtao Zhai |
IEEE Trans. Image Process. | 2 |
| 2019 | No-Reference Quality Evaluator of Transparently Encrypted ImagesabstractIn past years, various encrypted algorithms have been proposed to fully or partially protect the multimedia content in view of practical applications. In the context of digital TV broadcasting, transparent encryption only protects partial content and fulfills both security and quality requirements. To date, only a few reference-based works have been reported to evaluate the quality of transparently encrypted images. However, these works are incapable of reference-unavailable conditions. In this paper, we conduct the first attempt that proposes a novel quality evaluator in the absence of reference images. The key strategy of the proposed metric lies in extracting features by considering the motivation of transparently encrypted images. Specifically, given that encrypted images prevent content from being easily recognized, several features, including correlation coefficient, information entropy, and intensity statistic, are preliminarily extracted to estimate visual recognizability. Meanwhile, considering that encrypted images are avoided since they are of extremely low quality, we also capture many features to measure the distortions on multiple quality-sensitive image attributes, such as naturalness, structure, and texture. Finally, the quality evaluator is built by bridging all extracted features and corresponding quality scores via a regression module. Experimental results demonstrate that the proposed method is superior to the mainstream no-reference quality evaluation methods designed for synthetically distorted images and possesses a close approximation to state-of-the-art reference-based methods designed for encrypted images. Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Tianwei Zhou, Hantao Liu |
IEEE Trans. Multim. | 2 |
| 2019 | Subtitle Region Selection of S3D Images in Consideration of Visual Discomfort and Viewing HabitabstractSubtitles, serving as a linguistic approximation of the visual content, are an essential element in stereoscopic advertisement and the film industry. Due to the vergence accommodation conflict, the stereoscopic 3D (S3D) subtitle inevitably causes visual discomfort. To meet the viewing experience, the subtitle region should be carefully arranged. Unfortunately, very few works have been dedicated to this area. In this article, we propose a method for S3D subtitle region selection in consideration of visual discomfort and viewing habit. First, we divide the disparity map into multiple depth layers according to the disparity value. The preferential processed depth layer is determined by considering the disparity value of the foremost object. Second, the optimal region and coarse disparity value for S3D subtitle insertion are chosen by convolving the selective depth layer with the mean filter. Specifically, the viewing habit is considered during the region selection. Finally, after region selection, the disparity value of the subtitle is further modified by using the just noticeable depth difference (JNDD) model. Given that there is no public database reported for the evaluation of S3D subtitle insertion, we collect 120 S3D images as the test platform. Both objective and subjective experiments are conducted to evaluate the comfort degree of the inserted subtitle. Experimental results demonstrate that the proposed method can obtain promising performance in improving the viewing experience of the inserted subtitle. Guanghui Yue 0001, Chunping Hou, Tianwei Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Image Alignment via Multi-Model Geometric Fitting and Hierarchical Homography EstimationabstractIt is challenging to achieve accurate alignment for building images containing multiple planes. We propose a multi-model geometric fitting and hierarchical homography estimation method to improve the alignment performance for building images. We first extract scale-invariant feature transform (SIFT) features of the images, and then adopt the multi-homography fitting algorithm to classify the feature points into different deformation models. According to the deduced deformation models, we partition the source image into base and transition regions. For the base regions, we adopt the moving direct linear transformation (Moving DLT) to estimate homographies. For the transition regions, we propose a hierarchical homography estimation method to select appropriate homographies. Experimental results show that our method achieves more accurate alignment results compared with state-of-the-art alignment methods for building images. Jing-Yu Yang 0002, Huanjing Yue, Kun Li 0001, Chunping Hou |
ICASSP | 5 |
| 2018 | Cyclopean Image Based Stereoscopic Image Quality Assessment by Using Sparse Representationabstract3D image quality assessment confronts more difficulties than 2D image quality assessment. In this paper a 3D image quality assessment metric based on sparse representation was proposed. The contributions of the proposed method mainly include the following points: a color cyclopean image is used to better simulate the process of image processing in human brain, which is also very suitable for evaluating the quality of asymmetric distortion image. Meanwhile, for during sparse reconstruction some important information will be lost, we use the corresponding color cyclopean image to do compensation before feature extracting. And the paper creatively extracts spatial and spectral entropy feature of the distortion color cyclopean image and the corresponding reconstruction cyclopean image, respectively. Finally, we uses SVR to evaluate the quality of stereoscopic image. Experimental results show that the proposed method is very much in line with human visual perception. Yongli Chang, Sumei Li, Chunping Hou |
ICIP | 4 |
| 2018 | Deep Joint Noise Estimation and Removal for High ISO JPEG ImagesabstractCapturing images under high ISO mode introduces much noise. The statistics of high ISO noise is quite different from that of Gaussian noise. Therefore, this kind of noise is difficult to be removed by traditional Gaussian noise removal methods. This paper proposes a convolutional neural network (CNN) based method to jointly estimate and remove high ISO noise. There are two contributions in this paper. First, we propose a CNN based noise estimation method to estimate the pixel-wise noise level. Due to the Bayer down-sampling process in imaging, the noise variance map is characterized by Bayer patterns. Therefore, we propose packing 2 × 2 blocks in a noisy image into 4D vectors, which makes the pixels with similar noise levels be neighbors. Second, the noise variance map is correlated with the image content. Thus, we propose concatenating the estimated noise variance map with the noisy image, and feed the fused data to the denoising network. The two networks are trained together in an end-to-end fashion. Experimental results demonstrate that the proposed method outperforms state-of-the-art noise estimation and removal methods. Huanjing Yue, Shengdi Zhou, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Chunping Hou |
ICPR | 5 |
| 2018 | Multiple Residual Learning Network for Single Image Super-ResolutionabstractDeep residual convolutional neural network (CNN) has recently achieved great success in image super-resolution (SR). Because residual learning accelerates convergence rate and eases the difficulty for reconstructing high-resolution (HR) image, these CNN models can achieve higher peak signal to noise ratio (PSNR) values with lower training cost. However, residual image used in present residual network still contains much high frequency information, which increases learning burden and limits learning ability of residual network. Moreover, training a very deep network faces many obstacles and costs too much time. In this paper, we propose a multiple residual learning network (MRLN), which not only further simplifies information complexity of residual image and improves the accuracy of residual network, but also obviously reduces time cost for training a very deep CNN. In MRLN, we use a shallow network formed by 30-layer convolutional layers as basic model and train it for multiple times. The output of previous basic model is used as the HR input of the next one. In this way, an extremely large CNN is converted into a series connection of shallow networks. Fig. 1 shows PSNR of recent state-of-the-art CNN models for scale factor 2 on Set5, our method performs better than other methods and set a new level for SR. Renhe Liu, Sumei Li, Chunping Hou, Guoqing Lei |
VCIP | 3 |
| 2018 | A two-channel convolutional neural network for image super-resolution
Sumei Li, Ru Fan, Guoqing Lei, Guanghui Yue 0001, Chunping Hou |
Neurocomputing | 5 |
| 2018 | Anomaly detection in crowded scenes using motion energy model
Tianyu Chen 0008, Chunping Hou, Hua Chen 0004 |
Multim. Tools Appl. | 2 |
| 2018 | No-reference stereoscopic 3D image quality assessment via combined model
Lili Shen, Jinyi Lei, Chunping Hou |
Multim. Tools Appl. | 3 |
| 2018 | Analysis of maximum tolerant depth distortion in view synthesis
Laihua Wang, Chunping Hou, Sumin Qi, Lanlan Jiang |
Multim. Tools Appl. | 2 |
| 2018 | Author Correction: Analysis of maximum tolerant depth distortion in view synthesis
Laihua Wang, Chunping Hou, Sumin Qi, Lanlan Jiang |
Multim. Tools Appl. | 2 |
| 2018 | Blind stereoscopic 3D image quality assessment via analysis of naturalness, structure, and binocular asymmetry
Guanghui Yue 0001, Chunping Hou, Qiuping Jiang, Yang Yang 0045 |
Signal Process. | 2 |
| 2018 | Fast Mode Decision Based on Grayscale Similarity and Inter-View Correlation for Depth Map Coding in 3D-HEVCabstractThe 3D extension of High Efficiency Video Coding significantly improves the coding efficiency of 3D video at the expense of computational complexity. This paper presents a novel fast mode decision algorithm for depth map coding based on the grayscale similarity and inter-view correlation. First, depth map grayscale similarity is adopted to judge whether the reference frame could assist the coding of the current frame. When the difference in the average grayscale between the co-located coding unit (CU) and the current CU is smaller than the similarity threshold, the depth level of the current CU will be restricted by that of the coded reference CU. Second, the grayscale similarity and inter-view correlation are jointly used for dependent views to achieve early decision on the best prediction unit (PU) mode. The mode decision procedure will be determined early when the co-located CU, which has a grayscale similarity with the current CU, selects Merge or Inter 2N ×2N as the best prediction mode. Moreover, when the corresponding CU in the independent view selects Merge or Inter 2N × 2N as the best prediction mode, the current CU will skip other PU modes checking based on the strong inter-view correlation. Finally, different strategies are proposed for the P-frames and B-frames of dependent views in view of the characteristics of different prediction structures. For B frames, the PU mode information of the coded independent view is utilized as reference to skip the unnecessary mode decision processes. For P frames, the spatial-temporal correlation is considered in the process of early mode decision to determine whether to choose the Merge mode or Inter 2N × 2N as the best mode. Experimental results show that our proposed scheme achieves considerable time saving with negligible degradation of coding performance. Jianjun Lei 0001, Jinhui Duan, Feng Wu 0001, Nam Ling, Chunping Hou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Region Adaptive R-λ Model-Based Rate Control for Depth Maps CodingabstractIn this paper, a novel rate-control algorithm based on the region adaptive R-λ model is proposed for depth maps coding. First, in order to obtain an accurate rate control for depth maps coding, a modified frame level bit allocation method based on coding bits statistical distribution of depth maps is proposed. Second, considering that different areas in a depth map have an imparity effect on virtual view rendering, the blocks of the depth map are divided into two types, namely, interested blocks for virtual view rending (IBV) and noninterested blocks for virtual view rending (NIBV). Then, two different R-λ models are derived for IBV and NIBV, respectively. The optimal bitrates for IBV and NIBV are determined by solving an optimization problem. After that, based on the regional R-λ models, the optimal Lagrange multipliers are calculated for both IBV and NIBV. Finally, the largest coding unit (LCU) level rate control is performed by adaptively adjusting the Lagrange multiplier to avoid blocking artifacts and smooth the quality of coding. Experimental results demonstrate that the proposed method can achieve considerable BD-PSNR gains compared with the unified rate-quantization model and conventional R-λ modelbased algorithms in terms of rendered virtual views quality. Jianjun Lei 0001, Xiaoxu He, Hui Yuan 0001, Feng Wu 0001, Nam Ling, Chunping Hou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2018 | Optimal Region Selection for Stereoscopic Video Subtitle InsertionabstractStereoscopic subtitle insertion is a fundamental and essential element in stereoscopic film and TV industry. However, little work has been dedicated to the optimal region selection for stereoscopic subtitle insertion. In addition, there is no public database reported for the performance evaluation of it. In this paper, we build the first large-scale video database (TJU3D) for stereoscopic video subtitle insertion, which includes 50 video sequences with rich screen scenes. Compared with 2D subtitle region selection, there are several problems we have to consider in stereoscopic subtitle region selection: 1) the subtitle should avoid depth cue collision and occlusion from objects in stereoscopic video sequences; 2) the disparity value of the subtitle must be minimized to reduce visual discomfort; and 3) the temporal coherence constraint must be considered during region selection for subtitles in video sequences. By considering these constraints, we propose an optimal region selection algorithm for stereoscopic subtitle insertion. First, we compute the disparity map of each video frame in video sequences. For each frame, the optimal position and disparity value of the subtitle are determined by a subtitle region selection algorithm, which contains two parts (i.e., the coarse selection and fine selection). After that, by considering the temporal consistency between adjacent frames, the position and disparity value of each frame are further classified and processed in order to avoid the subtitle jitter. We evaluate the proposed method on TJU3D video database through two visual discomfort prediction metrics and one subjective experiment. To further verify the effectiveness of the proposed method, we also validate the performance of the proposed method on video comfort assessment database, i.e., IEEE-SA Stereo Database. Experimental results demonstrate that the visual discomfort is greatly reduced when using the proposed method compared with the basic method. Guanghui Yue 0001, Chunping Hou, Jianjun Lei 0001, Yuming Fang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Co-Saliency Detection for RGBD Images Based on Multi-Constraint Feature Matching and Cross Label PropagationabstractCo-saliency detection aims at extracting the common salient regions from an image group containing two or more relevant images. It is a newly emerging topic in computer vision community. Different from the most existing co-saliency methods focusing on RGB images, this paper proposes a novel co-saliency detection model for RGBD images, which utilizes the depth information to enhance identification of co-saliency. First, the intra saliency map for each image is generated by the single image saliency model, while the inter saliency map is calculated based on the multi-constraint feature matching, which represents the constraint relationship among multiple images. Then, the optimization scheme, namely cross label propagation, is used to refine the intra and inter saliency maps in a cross way. Finally, all the original and optimized saliency maps are integrated to generate the final co-saliency result. The proposed method introduces the depth information and multi-constraint feature matching to improve the performance of co-saliency detection. Moreover, the proposed method can effectively exploit any existing single image saliency model to work well in co-saliency scenarios. Experiments on two RGBD co-saliency datasets demonstrate the effectiveness of our proposed model. Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Qingming Huang, Xiaochun Cao, Chunping Hou |
IEEE Trans. Image Process. | 6 |
| 2018 | Depth Super-Resolution From RGB-D Pairs With Transform and Spatial Domain RegularizationabstractThis paper proposes a depth super-resolution method with both transform and spatial domain regularization. In the transform domain regularization, nonlocal correlations are exploited via an auto-regressive model, where each patch is further sparsified with a locally-trained transform to consider intra-patch correlations. In the spatial domain regularization, we propose a multi-directional total variation (MTV) prior to characterize the geometrical structures spatially orientated at arbitrary directions in depth maps. To achieve adaptive regularization, the MTV is weighted for each directional finite difference considering local characteristics of RGB-D data. We develop an accelerated proximal gradient algorithm to solve the proposed model. Quantitative and qualitative evaluations compared with state-of-the-art methods demonstrate that the proposed method achieves superior depth super-resolution performance for various configurations of magnification factors and datasets. Zhongyu Jiang, Yonghong Hou, Huanjing Yue, Jing-Yu Yang 0002, Chunping Hou |
IEEE Trans. Image Process. | 5 |
| 2018 | Iterative Feedback Control-Based Salient Object SegmentationabstractIn this paper, we establish a mathematical model that relates the control states and the saliency values in salient object detection. We show that a linear feedback control system (LFCS) is amenable to saliency detection tasks owing to its functional properties. This inspired us to employ an LFCS to detect salient objects in static images. Based on the novel iteration method, the system gradually converges to an optimized stable state, which is associated with an accurate saliency map. In addition, to initialize the system, we propose a so-called boundary homogeneity based on a priori knowledge of the boundary in order to estimate the background likelihood and indirectly obtain a foreground (saliency) map. The experimental results indicate that such a feedback control model can offer significant improvement in salient object detection performance. Shuwei Huo, Yuan Zhou 0006, Jianjun Lei 0001, Nam Ling, Chunping Hou |
IEEE Trans. Multim. | 5 |
| 2018 | Analysis of Structural Characteristics for Quality Assessment of Multiply Distorted ImagesabstractPerceptual image quality assessment (IQA) plays an important role in numerous applications, including image restoration, compression, enhancement, and others. Although many works have been conducted on individually distorted IQA problems and have achieved encouraging results, few studies have been conducted on multiple distorted (MD) IQA problems. Thus, limited progress has been made. In this paper, we propose a novel no reference image quality assessment (NR-IQA) method, named improved multiscale local binary pattern (IMLBP), for addressing multiply distorted IQA problems. The image structures are sensitive to image distortions, which motivates us to utilize the structural characteristics for overall image quality prediction. We improved the local binary pattern (LBP) by considering the human visual mechanism to better extract the structural information. The IMLBP contains two parts, the LBP and the radius difference LBP (DLBP). The DLBP reflects the values' changes in the radial direction. Specifically, when the radius value is small, the proposed descriptor is computed to represent microstructural information. Conversely, it represents macrostructural information when the radius becomes large. Moreover, to better mimick the human visual mechanism, the IMLBP is computed with the multiscale strategy and the operation is based on a patch unit whose size is proportional to the radius value. The frequency histogram of feature maps is transformed to feature vectors. Subsequently, a predictable function trained by the support vector regression is used to infer the overall quality score. Experimental results show that the proposed method outperforms most state-of-the-art IQA metrics on publicly available multiply distorted image databases. Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Nam Ling, Beichen Li 0002 |
IEEE Trans. Multim. | 2 |
| 2017 | Depth-weighted correlation method for visual tracking with occlusion detectionabstractDespite the significant progress, it remains a challenging task for a tracker to distinguish a target from the background when the target is occluded. In this paper, we propose a new tracking method, named as depth-weighted correlation method(DWCM), to handle heavy occlusion. The proposed method uses depth cues as the weights of candidate objects and applies the framework of spatio-temporal context (STC). We also propose a scale update scheme for DWCM, so as to obtain an appropriate scale for the target. Encouraging experimental results show that the proposed tracker obtains state-of-the-art results and handles occlusion better than competing tracking methods. Chenghao Li 0004, Yuan Zhou 0006, Bo Cut, Chunping Hou |
ICIP | 4 |
| 2017 | Underwater image enhancement based on structure-texture decompositionabstractUnderwater images generally suffer from low contrast, serious noise and color distortion. The main challenges of underwater image enhancement are to preserve details in dark regions while avoiding oversaturetion in bright regions. This paper proposes a novel underwater image enhancement method based on image decomposition. By decomposing the high-frequency texture and noise into the texture layer, the transmission map is estimated from the noise-free structure layer to avoid the noise amplification problem in underwater image enhancement. Both the structure layer and texture layer are descattered with the estimated transmission map. After denoising by gradient residual minimizition, the texture layer is enhanced and added back into the structure layer to recover the final enhanced image. Experimental results verify that the proposed approach can recover the high-quality images with fine details and edges while improving contrast and color naturalness, especially for images taken in the high turbidity environment. Jing-Yu Yang 0002, Huanjing Yue, Xiaomei Fu, Chunping Hou |
ICIP | 5 |
| 2017 | Image noise estimation and removal considering the bayer pattern of noise varianceabstractTraditional image denoising methods are designed for Gaussian or Poisson noise, which are not suitable for realistic noise introduced in the complicated imaging pipeline. We observe that, due to the demosaicing process in imaging, the noise variance maps of captured JPEG images are characterized by Bayer patterns. In this paper, we propose a novel noise estimation and removal method based on the Bayer pattern of noise variance maps. There are two key contributions in the proposed method. First, to the best of our knowledge, we are the first to consider the Bayer patterns of noise variance maps in noise estimation and denoising. Second, we extend the state-of-the-art denoising method CBM3D to deal with realistic noise by integrating the estimated noise variance map and Bayer-pattern down-sampling into the denoising process. Experimental results show that the proposed method achieves the best noise estimation performance compared with two state-of-the-art methods. In addition, the denoising performance of CBM3D for realistic noise is significantly improved using the proposed approach and outperforms state-of-the-art blind denoising methods. Huanjing Yue, Jing-Yu Yang 0002, Truong Q. Nguyen, Chunping Hou |
ICIP | 5 |
| 2017 | Label propagation based saliency detection via graph designabstractSaliency detection has been widely used as the pre-processing of the computer-vision tasks. Existing propagation based saliency detection methods simply select a k-regular graph for saliency propagation, which usually leads to the mistaken highlighting of the long-range smooth background regions. In this paper, we design a novel graph for label propagation based saliency detection by considering the local consistency and the global symmetry of the image scene and updating the graph model based on smoothness assumption and cluster assumption. Then, we label the reliable seeds and propagate the saliency value through the designed graph. On two widely used large open benchmark data sets, the proposed method significantly outperforms thirteen state-of-the-arts under either quantitative or qualitative evaluation. Yuan Zhou 0006, Shuwei Huo, Chunping Hou |
ICIP | 4 |
| 2017 | Joint nonlocal sparse representation for depth map super-resolutionabstractDepth image super-resolution reconstruction has gained significant popularity due to its practicability. However, conventional depth image super-resolution reconstruction methods access high frequency information either from a high-resolution depth image database or from a high-resolution color image of the same scene, which is limited in specific applications. In this paper, a novel joint nonlocal sparse representation model is proposed, which is able to capture the interdependency of low-resolution depth and intensity information. As a relative new and not well addressed problem, we reconstruct a high-resolution depth image from a single low-resolution depth image with a low-resolution color image as reference. Experiment results demonstrate that the proposed method outperforms many current state-of-the-art depth map super-resolution approaches on both visual effects and objective image quality. Yeda Zhang, Yuan Zhou 0006, Aihua Wang, Chunping Hou |
ICIP | 5 |
| 2017 | Subjective quality assessment of animation imagesabstractIn the past few decades, many attempts have been maken to evaluate the image quality assessment (IQA) of natural scene images. However, the IQA research of animation images (AIs) has been highly overlooked. In this article, we carry out in-depth study on perceptual quality assessment of AIs. As the lack of a public and diverse testing database currently, this paper builds a large-scale Animation Images Quality Assessment Database (AIQAD). This database totally includes 1050 distorted images derived from 30 source images by corrupting seven distortion types with multiple distortion levels. Then, a subjective experiment, which is the basic and accurate quality evaluation measurement, is conducted to obtain the mean opinion score (MOS) for each image. Furthermore, we also investigate the feasibility of utilizing existing mainstream full reference (FR) IQA metrics to solve the IQA problem of AIs. Experimental results demonstrate that existing mainstream FR IQA metrics merely achieve fair performance on the proposed database. Guanghui Yue 0001, Chunping Hou, Ke Gu 0001 |
VCIP | 2 |
| 2017 | No reference image blurriness assessment with local binary patterns
Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Nam Ling |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Region-based bit allocation and rate control for depth video in HEVC
Jianjun Lei 0001, Xiaoxu He, Chunping Hou |
Multim. Tools Appl. | 6 |
| 2017 | A divide-and-conquer hole-filling method for handling disocclusion in single-view renderingabstractLarge holes are unavoidably generated in depth image based rendering (DIBR) using a single color image and its associated depth map. Such holes are mainly caused by disocclusion, which occurs around the sharp depth discontinuities in the depth map. We propose a divide-and-conquer hole-filling method which refines the background depth pixels around the sharp depth discontinuities to address the disocclusion problem. Firstly, the disocclusion region is detected according to the degree of depth discontinuity, and the target area is marked as a binary mask. Then, the depth pixels located in the target area are modified by a linear interpolation process, whose pixel values decrease from the foreground depth value to the background depth value. Finally, in order to remove the isolated depth pixels, median filtering is adopted to refine the depth map. In these ways, disocclusion regions in the synthesized view are divided into several small holes after DIBR, and are easily filled by image inpainting. Experimental results demonstrate that the proposed method can effectively improve the quality of the synthesized view subjectively and objectively. Jianjun Lei 0001, Cuicui Zhang, Kefeng Fan, Chunping Hou |
Multim. Tools Appl. | 6 |
| 2017 | Content-aware disparity adjustment for different stereo displays
Weiqing Yan, Chunping Hou, Baoliang Wang, Laihua Wang |
Multim. Tools Appl. | 2 |
| 2017 | ESPRIT-like two-dimensional direction finding for mixed circular and strictly noncircular sources based on joint diagonalization
Hua Chen 0004, Chunping Hou, Wei-Ping Zhu 0001, Wei Liu 0001, Zongju Peng, Qing Wang 0015 |
Signal Process. | 2 |
| 2017 | Stereoscopic Image Stitching Based on a Hybrid Warping ModelabstractTraditional image editing techniques cannot be directly used to process stereoscopic media, as extra constraints are required to ensure consistent changes between the left and right images. In this paper, we propose a hybrid warping model for stereoscopic image stitching by combining projective and content-preserving warping. First, a uniform homography algorithm is proposed to prewarp the left and right images, and thus ensure consistent changes. Second, a content-preserving warping is introduced to locally refine alignment and reduce vertical disparities. Finally, a seam-cutting-based algorithm is used to find a blending seam, and the multiband blending algorithm is used to produce the final stitched image. Experimental results show that the proposed method can effectively stitch stereoscopic images, which not only avoids local distortions, but also reduces vertical disparities reasonably. Weiqing Yan, Chunping Hou, Jianjun Lei 0001, Yuming Fang 0001, Zhouye Gu, Nam Ling |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Hyperspectral and Multispectral Image Fusion Based on Local Low Rank and Coupled Spectral UnmixingabstractHyperspectral images (HSIs) usually have high spectral and low spatial resolution. Conversely, multispectral images (MSIs) usually have low spectral and high spatial resolution. The fusion of HSI and MSI aims to create spectral images with high spectral and spatial resolution. In this paper, we propose a fusion algorithm by combining linear spectral unmixing with the local low-rank property. By taking advantage of the local low-rank property, we first partition the corresponding spectral image into patches. For each patch pair, we cast the fusion problem as a coupled spectral unmixing problem that extracts the abundance and the endmembers of MSI and HSI, respectively. It then updates the abundance and the endmember through an alternating update algorithm. In fact, the convergence of the alternative update algorithm can be mathematically and empirically supported. We also propose a multiscale postprocessing procedure to combine fusion results obtained under different patch sizes. In experiments on three data sets, the proposed fusion algorithms outperformed state-of-the-art fusion algorithms in both spatial and spectral domains. Yuan Zhou 0006, Liyang Feng, Chunping Hou, Sun-Yuan Kung |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Depth Map Super-Resolution Considering View Synthesis QualityabstractAccurate and high-quality depth maps are required in lots of 3D applications, such as multi-view rendering, 3D reconstruction and 3DTV. However, the resolution of captured depth image is much lower than that of its corresponding color image, which affects its application performance. In this paper, we propose a novel depth map super-resolution (SR) method by taking view synthesis quality into account. The proposed approach mainly includes two technical contributions. First, since the captured low-resolution (LR) depth map may be corrupted by noise and occlusion, we propose a credibility based multi-view depth maps fusion strategy, which considers the view synthesis quality and interview correlation, to refine the LR depth map. Second, we propose a view synthesis quality based trilateral depth-map up-sampling method, which considers depth smoothness, texture similarity and view synthesis quality in the up-sampling filter. Experimental results demonstrate that the proposed method outperforms state-of-the-art depth SR methods for both super-resolved depth maps and synthesized views. Furthermore, the proposed method is robust to noise and achieves promising results under noise-corruption conditions. Jianjun Lei 0001, Lele Li, Huanjing Yue, Feng Wu 0001, Nam Ling, Chunping Hou |
IEEE Trans. Image Process. | 6 |
| 2017 | Textured Image Demoiréing via Signal Decomposition and Guided FilteringabstractMoiré artifacts are generally caused by the interference between the overlap of the sensor's sampling grid and high-frequency (nearly) periodic textures, and heavily affect the image quality. However, it is difficult to effectively remove moiré artifacts from textured images as the structure of moiré patterns is similar to that of textures in some sense. In this paper, we propose a novel textured image demoiréing method by signal decomposition and guided filtering. Given a textured image with moiré artifacts, we first remove moiré artifacts in the green (G) channel using the proposed low-rank and sparse matrix decomposition model. This model regularizes the texture layer by the low-rank prior in spatial domain and the moiré layer by sparse representation in frequency domain. An alternating direction method under the augmented Lagrangian multiplier framework is used to solve the matrix decomposition model. Then, since the red (R) and blue (B) channels are more heavily polluted by moiré artifacts than the G channel, we propose to remove moiré artifacts in its R and B channels via guided filtering by the obtained texture layer of the G channel. Experimental results demonstrate that our method outperforms the state-of-the-art methods for both synthetic and real images. Jing-Yu Yang 0002, Fanglei Liu, Huanjing Yue, Xiaomei Fu, Chunping Hou, Feng Wu 0001 |
IEEE Trans. Image Process. | 5 |
| 2017 | Reconstruction of Structurally-Incomplete Matrices With Reweighted Low-Rank and Sparsity PriorsabstractMost matrix reconstruction methods assume that missing entries randomly distribute in the incomplete matrix, and the low-rank prior or its variants are used to well pose the problem. However, in practical applications, missing entries are structurally rather than randomly distributed, and cannot be handled by the rank minimization prior individually. To remedy this, this paper introduces new matrix reconstruction models using double priors on the latent matrix, named Reweighted Low-rank and Sparsity Priors (ReLaSP). In the proposed ReLaSP models, the matrix is regularized by a low-rank prior to exploit the inter-column and inter-row correlations, and its columns (rows) are regularized by a sparsity prior under a dictionary to exploit intra-column (-row) correlations. Both the low-rank and sparse priors are reweighted on the fly to promote low-rankness and sparsity, respectively. Numerical algorithms to solve our ReLaSP models are derived via the alternating direction method under the augmented Lagrangian multiplier framework. Results on synthetic data, image restoration tasks, and seismic data interpolation show that the proposed ReLaSP models are quite effective in recovering matrices degraded by highly structural missing and various types of noise, complementing the classic matrix reconstruction models that handle random missing only. Jing-Yu Yang 0002, Xuemeng Yang, Xinchen Ye, Chunping Hou |
IEEE Trans. Image Process. | 4 |
| 2017 | Contrast Enhancement Based on Intrinsic Image DecompositionabstractIn this paper, we propose to introduce intrinsic image decomposition priors into decomposition models for contrast enhancement. Since image decomposition is a highly illposed problem, we introduce constraints on both reflectance and illumination layers to yield a highly reliable solution. We regularize the reflectance layer to be piecewise constant by introducing a weighted ℓ1norm constraint on neighboring pixels according to the color similarity, so that the decomposed reflectance would not be affected much by the illumination information. The illumination layer is regularized by a piecewise smoothness constraint. The proposed model is effectively solved by the Split Bregman algorithm. Then, by adjusting the illumination layer, we obtain the enhancement result. To avoid potential color artifacts introduced by illumination adjusting and reduce computing complexity, the proposed decomposition model is performed on the value channel in HSV space. Experiment results demonstrate that the proposed method performs well for a wide variety of images, and achieves better or comparable subjective and objective quality compared with the state-of-the-art methods. Huanjing Yue, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Feng Wu 0001, Chunping Hou |
IEEE Trans. Image Process. | 5 |
| 2017 | Depth-Preserving Stereo Image Retargeting Based on Pixel FusionabstractIn this paper, we propose a pixel fusion-based stereo image retargeting method, which could adaptively retarget stereo images with flexible aspect ratios, simultaneously preserving the depth. Retargeting each image independently by the pixel fusion method ignores the disparity relationship between pixels in the image pair and hence will introduce the distortion of disparity. To address this issue, we advocate to extend the single pixel fusion-based way to be applicable for stereo image pair. First, seams are selected based on the energy function, which simultaneously considers the seam selecting and seam matching. Second, a seam-matching-based matching map is proposed to preserve the disparity relationship between image pair. Then, the scaling factors for the left image are assigned considering both the important object and depth preservation. Subsequently, the scaling factors for the right image are obtained according to the proposed matching map. Based on these scaling factors, the stereo image pair is retargeted with pixel fusion. In contrast to removing pixels to resize image, the way of pixel fusion can obtain more smooth results with less depth distortion. Experimental results demonstrate that our method achieves more preferable qualities in both depth and shape preservation for stereo image retargeting. Jianjun Lei 0001, Changqing Zhang 0002, Feng Wu 0001, Nam Ling, Chunping Hou |
IEEE Trans. Multim. | 6 |
| 2016 | Depth refinement for binocular kinect RGB-D camerasabstractThis paper presents a novel depth refinement framework for binocular Kinect RGB-D cameras for obtaining high quality depth map. Firstly, we build a binocular depth sensing system using two Kinect v2 cameras, and analyze the systematic error of the system from two aspects, i.e., camera interaction and intrinsic characteristics. Then, the captured depth maps from different views are fused to fully exploit the inter-view correlations, and an error compensation method is proposed to remove the systematic errors from the fused depth map. Finally, an edge-guided depth propagation scheme is used to refine the depth map from binocular depth map. Experimental results show that the proposed framework is able to substantially improve the quality of depth image. Jinghui Bai, Jing-Yu Yang 0002, Xinchen Ye, Chunping Hou |
VCIP | 4 |
| 2016 | Depth recovery via decomposition of polynomial and piece-wise constant signalsabstractThis paper proposes a novel decomposition model for high-quality depth recovery (DMDR) from low quality depth measurement accompanied by high-resolution RGB image. We observe that depth patches extracted from the depth map containing smooth regions separated by curves, can be decomposed simultaneously by a low-order polynomial surface and a piece-wise constant signal. In our model, the polynomial surface component is regularized by least-square polynomial smoothing, while the piece-wise constant component is constrained by total variation filtering. The model is effectively solved by the alternating direction method under the augmented Lagrangian multiplier (ALM-ADM) algorithm. Experimental results show that our method is able to handle various types of depth degradation under the designed signal decomposition model, and produces high-quality depth recovery results. Xinchen Ye, Jing-Yu Yang 0002, Chunping Hou, Yao Wang 0001 |
VCIP | 4 |
| 2016 | Salient object detection using color spatial distribution and minimum spanning tree weight
Chang Tang, Chunping Hou, Pichao Wang, Zhanjie Song |
Multim. Tools Appl. | 2 |
| 2016 | A setup for panoramic stereo imaging
Chunping Hou |
Multim. Tools Appl. | 2 |
| 2016 | Saliency Detection for Stereoscopic Images Based on Depth Confidence Analysis and Multiple Cues FusionabstractStereoscopic perception is an important part of human visual system that allows the brain to perceive depth. However, depth information has not been well explored in existing saliency detection models. In this letter, a novel saliency detection method for stereoscopic images is proposed. First, we propose a measure to evaluate the reliability of depth map, and use it to reduce the influence of poor depth map on saliency detection. Then, the input image is represented as a graph, and the depth information is introduced into graph construction. After that, a new definition of compactness using color and depth cues is put forward to compute the compactness saliency map. In order to compensate the detection errors of compactness saliency when the salient regions have similar appearances with background, foreground saliency map is calculated based on depth-refined foreground seeds' selection (DRSS) mechanism and multiple cues contrast. Finally, these two saliency maps are integrated into a final saliency map through weighted-sum method according to their importance. Experiments on two publicly available stereo data sets demonstrate that the proposed method performs better than other ten state-of-the-art approaches. Runmin Cong, Jianjun Lei 0001, Changqing Zhang 0002, Qingming Huang, Xiaochun Cao, Chunping Hou |
IEEE Signal Process. Lett. | 6 |
| 2016 | A Universal Framework for Salient Object DetectionabstractIn this paper, we propose a novel universal framework for salient object detection, which aims to enhance the performance of any existing saliency detection method. First, rough salient regions are extracted from any existing saliency detection model with distance weighting, adaptive binarization, and morphological closing. With the superpixel segmentation, a Bayesian decision model is adopted to refine the rough saliency map to obtain a more accurate saliency map. An iterative optimization method is designed to obtain better saliency results by exploiting the characteristics of the output saliency map each time. Through the iterative optimization process, the rough saliency map is updated step by step with better and better performance until an optimal saliency map is obtained. Experimental results on the public salient object detection datasets with ground truth demonstrate the promising performance of the proposed universal framework subjectively and objectively. Jianjun Lei 0001, Bingren Wang, Yuming Fang 0001, Weisi Lin, Patrick Le Callet, Nam Ling, Chunping Hou |
IEEE Trans. Multim. | 7 |
| 2015 | RIFO: Restoring images with fence occlusionsabstractMany scenes, e.g., zoos, parks, and gardens, are guarded by fences, and people can only take pictures through the fences. It is desirable to remove visually-annoying fence occlusions from images. This paper proposes a novel approach to restore images from fence occlusions (RIFO). The proposed method consists of two steps: fence detection, and disocclusion restoration. In fence detection, the image is first clustered into superpixels, which are fitted into rectangles. We collect fence pixels from superpixels by determining the elongation of their fitted rectangles. A primary shape of the fence is obtained by a color-based classifier learned from sampled pixels. Then, multi-RANSAC and moving least squares (MLS) are used for sketching the fence structure. Complete fence is detected by expanding the fence structure. Disoccluded regions are restorated by a patch-based approach using matrix completion. Experimental results show that our method detects complete fences from images, and the disoccluded regions are faithfully recovered, yielding clean and complete images. Jing-Yu Yang 0002, Leijie Liu, Chunping Hou |
MMSP | 4 |
| 2015 | View generation with DIBR for 3D display system
Laihua Wang, Chunping Hou, Jianjun Lei 0001, Weiqing Yan |
Multim. Tools Appl. | 2 |
| 2015 | A depth estimating method from a single image using FoE CRF
Chunping Hou, Liangzhou Pu, Yonghong Hou |
Multim. Tools Appl. | 2 |
| 2015 | Direction finding and mutual coupling estimation for uniform rectangular arrays
Chunping Hou, Hua Chen 0004, Wei Liu 0001, Qing Wang 0015 |
Signal Process. | 2 |
| 2015 | Depth Coding Based on Depth-Texture Motion and Structure SimilaritiesabstractThis paper addresses high performance depth coding in 3D video by making good use of its coded texture video counterpart. The relationship between the depth and its associated texture video in terms of coding mode and motion vector is carefully examined. Our statistical study suggests that the skip-coding mode and its associated motion vectors in the coded texture can be shared for depth coding by saving bit rate at the cost of little increase of distortion, which subsequently results in a nonsequential coding of the depth map. In this sense, coding/prediction of a block can be performed using the skip-coded blocks below and right, which are not available in the conventional sequential coding, thus producing the so-called omnidirectional blocks predicted in the intra-coding by making the best use of (at most) four neighboring blocks. Moreover, in view of the depth-texture structure similarity, a depth-texture cooperative clustering-based prediction method is proposed for cluster-based depth prediction in the intra-coding, which exploits the structure similarity for the current coding block and its neighboring pixels around the block. On the other hand, some large prediction errors may be present for the depth-texture misaligned pixels, which may greatly compromise the coding performance. To deal with these large residuals induced by the depth-texture misalignment, a simple yet effective detection and rectification approach is incorporated in the proposed depth coding scheme. Experimental results show that our proposed depth coding scheme achieves superior rate-distortion performance compared with other relevant coding methods. Jianjun Lei 0001, Shuai Li 0005, Ce Zhu, Ming-Ting Sun, Chunping Hou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Foreground-Background Separation From Video Clips via Motion-Assisted Matrix RestorationabstractSeparation of video clips into foreground and background components is a useful and important technique, making recognition, classification, and scene analysis more efficient. In this paper, we propose a motion-assisted matrix restoration (MAMR) model for foreground-background separation in video clips. In the proposed MAMR model, the backgrounds across frames are modeled by a low-rank matrix, while the foreground objects are modeled by a sparse matrix. To facilitate efficient foreground-background separation, a dense motion field is estimated for each frame, and mapped into a weighting matrix which indicates the likelihood that each pixel belongs to the background. Anchor frames are selected in the dense motion estimation to overcome the difficulty of detecting slowly moving objects and camouflages. In addition, we extend our model to a robust MAMR model against noise for practical applications. Evaluations on challenging datasets demonstrate that our method outperforms many other state-of-the-art methods, and is versatile for a wide range of surveillance videos. Xinchen Ye, Jing-Yu Yang 0002, Kun Li 0001, Chunping Hou, Yao Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Graph-Based Segmentation for RGB-D Data Using 3-D Geometry Enhanced SuperpixelsabstractWith the advances of depth sensing technologies, color image plus depth information (referred to as RGB-D data hereafter) is more and more popular for comprehensive description of 3-D scenes. This paper proposes a two-stage segmentation method for RGB-D data: 1) oversegmentation by 3-D geometry enhanced superpixels and 2) graph-based merging with label cost from superpixels. In the oversegmentation stage, 3-D geometrical information is reconstructed from the depth map. Then, a K-means-like clustering method is applied to the RGB-D data for oversegmentation using an 8-D distance metric constructed from both color and 3-D geometrical information. In the merging stage, treating each superpixel as a node, a graph-based model is set up to relabel the superpixels into semantically-coherent segments. In the graph-based model, RGB-D proximity, texture similarity, and boundary continuity are incorporated into the smoothness term to exploit the correlations of neighboring superpixels. To obtain a compact labeling, the label term is designed to penalize labels linking to similar superpixels that likely belong to the same object. Both the proposed 3-D geometry enhanced superpixel clustering method and the graph-based merging method from superpixels are evaluated by qualitative and quantitative results. By the fusion of color and depth information, the proposed method achieves superior segmentation performance over several state-of-the-art algorithms. Jing-Yu Yang 0002, Ziqiao Gan, Kun Li 0001, Chunping Hou |
IEEE Trans. Cybern. | 4 |
| 2015 | Fast Mode Decision Using Inter-View and Inter-Component Correlations for Multiview Depth Video CodingabstractWith the development of three-dimensional (3-D) display technologies, 3-D video has attracted more and more interest. Multiview video plus depth (MVD) is one of the most popular representation formats of 3-D video. In MVD coding system, multiview depth video needs to be coded and transmitted in addition to the texture video. This paper presents a novel fast mode decision (FMD) method for odd views in multiview depth video coding. First, the inter-view and inter-component coding correlations are analyzed to provide efficient reference information. Then, with a view to the characteristics of different types of frames, different early termination strategies are proposed. For the nonanchor frame, the early termination criterion is based on the rate-distortion cost information of the even views and the coded block pattern information. For the anchor frame, the criterion is set stricter to maintain the coding accuracy. Experimental results show that the proposed method can reduce 78.07% coding time on average, without significant loss of video quality. Jianjun Lei 0001, Jing Sun 0010, Zhaoqing Pan, Sam Kwong, Jinhui Duan, Chunping Hou |
IEEE Trans. Ind. Informatics | 6 |
| 2015 | Estimation of Signal-Dependent Noise Level Function in Transform Domain via a Sparse Recovery ModelabstractThis paper proposes a novel algorithm to estimate the noise level function (NLF) of signal-dependent noise (SDN) from a single image based on the sparse representation of NLFs. Noise level samples are estimated from the high-frequency discrete cosine transform (DCT) coefficients of nonlocal-grouped low-variation image patches. Then, an NLF recovery model based on the sparse representation of NLFs under a trained basis is constructed to recover NLF from the incomplete noise level samples. Confidence levels of the NLF samples are incorporated into the proposed model to promote reliable samples and weaken unreliable ones. We investigate the behavior of the estimation performance with respect to the block size, sampling rate, and confidence weighting. Simulation results on synthetic noisy images show that our method outperforms existing state-of-the-art schemes. The proposed method is evaluated on real noisy images captured by three types of commodity imaging devices, and shows consistently excellent SDN estimation performance. The estimated NLFs are incorporated into two well-known denoising schemes, nonlocal means and BM3D, and show significant improvements in denoising SDN-polluted images. Jing-Yu Yang 0002, Ziqiao Gan, Zhaoyang Wu, Chunping Hou |
IEEE Trans. Image Process. | 4 |
| 2015 | Depth Sensation Enhancement for Multiple Virtual View RenderingabstractDepth information is an indispensable element in depth image-based rendering (DIBR) for three-dimensional (3-D) display. In this paper, we propose a novel depth sensation enhancement method to address the problems in multiple virtual view rendering. First, as the depth sensation is decreased when rendering intermediate multiple virtual views, the basic principle of depth sensation enhancement is derived according to the number of rendering views. Second, with the increase of the scene complexity, it is difficult to ensure the depth sensation of all neighboring objects. The saliency analysis is adopted to give preferred guarantee to the depth sensation between the salient object and its neighbors. Then, the depth sensation enhancement for multiple virtual view rendering is performed based on a defined energy function built by the number of rendering views and the saliency analysis. Finally, considering the temporal consistency between adjacent frames, the depth sensation enhancement is extended to video applications with a newly designed energy function with energy term of temporal consistency preservation. Experimental results on a public database demonstrate that the proposed method can obtain promising performance in depth sensation. Jianjun Lei 0001, Cuicui Zhang, Yuming Fang 0001, Zhouye Gu, Nam Ling, Chunping Hou |
IEEE Trans. Multim. | 6 |
| 2014 | Rate control of hierarchical B prediction structure for multi-view video coding
Jianjun Lei 0001, Meimin Wu, Shuai Li 0005, Chunping Hou |
Multim. Tools Appl. | 5 |
| 2014 | Pixel-Based Inter Prediction in Coded Texture Assisted Depth CodingabstractThis letter presents a pixel-based motion estimation scheme assisted with the coded texture video for depth inter-prediction, in view of motion similarity between depth and texture video. The proposed scheme can achieve higher inter-prediction gain without transmitting any motion vector in the pixel-based motion estimation. Coupled with depth-texture structure similarity, the inter prediction method is further extended to an integrated prediction approach by making use of both intra and inter information. Experimental results show that our proposed method achieves superior rate-distortion performance. Shuai Li 0005, Jianjun Lei 0001, Ce Zhu, Lu Yu 0003, Chunping Hou |
IEEE Signal Process. Lett. | 5 |
| 2014 | Color-Guided Depth Recovery From RGB-D Data Using an Adaptive Autoregressive ModelabstractThis paper proposes an adaptive color-guided autoregressive (AR) model for high quality depth recovery from low quality measurements captured by depth cameras. We observe and verify that the AR model tightly fits depth maps of generic scenes. The depth recovery task is formulated into a minimization of AR prediction errors subject to measurement consistency. The AR predictor for each pixel is constructed according to both the local correlation in the initial depth map and the nonlocal similarity in the accompanied high quality color image. We analyze the stability of our method from a linear system point of view, and design a parameter adaptation scheme to achieve stable and accurate depth recovery. Quantitative and qualitative evaluation compared with ten state-of-the-art schemes show the effectiveness and superiority of our method. Being able to handle various types of depth degradations, the proposed method is versatile for mainstream depth sensors, time-of-flight camera, and Kinect, as demonstrated by experiments on real systems. Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001, Chunping Hou, Yao Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2013 | Evaluation and modeling of depth feature incorporated visual attention for salient object segmentation
Jianjun Lei 0001, Hailong Zhang 0011, Chunping Hou, Laihua Wang |
Neurocomputing | 4 |
| 2012 | Depth Recovery Using an Adaptive Color-Guided Auto-Regressive Model
Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001, Chunping Hou |
ECCV (5) | 4 |
| 2012 | Estimation of signal-dependent sensor noise via sparse representation of noise level functionsabstractThis paper proposes a noise estimation method for signal-dependent sensor noise based on sparse representation of noise level functions (NLFs). Homogeneous blocks are detected by an image structure analyzer, and grouped to estimate noise levels for various image intensities with confidences. The noise level function is recovered from the incomplete and noisy estimated samples by solving its sparse representation under a trained basis. Experimental results show that our proposed method accurately estimates NLFs for both smooth and highly-textured images over various noise levels. Jing-Yu Yang 0002, Zhaoyang Wu, Chunping Hou |
ICIP | 3 |
| 2012 | A novel UEP scheme based upon rateless codesabstractIn this paper, we propose a novel method suitable for unequal error protection(UEP) and unequal recovery time (URT) properties, namely the Duplicating-Expanding Window Fountain(D-EWF) codes. We implement duplicating windows and expanding windows techniques to improve the performance of both more important bits (MIB) and less important bits (LIB). We analyze the proposed method over binary erasure channels (BEC) by using asymptotic analysis. The D-EWF codes inherit the advantages of both EWF codes and the duplicating windows method. Therefore the UEP property of D-EWF codes is more obvious. Furthermore BER performance of LIB of D-EWF codes converges very fast. Compared with the previous UEP schemes, simulation results show that the D-EWF codes can provide better UEP performance. Chunya Ni, Chunping Hou, Wei Xiang 0001 |
WCNC | 2 |
| 2012 | An Improved Nyquist-Shannon Irregular Sampling Theorem From Local AveragesabstractThe Nyquist–Shannon sampling theorem is on the reconstruction of a band-limited signal from its uniformly sampled samples. The higher the signal bandwidth gets, the more challenging the uniform sampling may become. To deal with this problem, signal reconstruction from local averages has been studied in the literature. In this paper, we obtain an improved Nyquist–Shannon sampling theorem from general local averages. In practice, the measurement apparatus gives a weighted average over an asymmetrical interval. As a special case, for local averages from symmetrical interval, we show that the sampling rate is much lower than that of a result by Gröchenig. Moreover, we obtain two exact dual frames from local averages, one of which improves a result by Sun and Zhou. At the end of this paper, as an example application of local average sampling, we consider a reconstruction algorithm: the piecewise linear approximations. Zhanjie Song, Yanwei Pang, Chunping Hou, Xuelong Li 0001 |
IEEE Trans. Inf. Theory | 4 |
| 2011 | Channel Distortion Modeling for Multi-View Video Transmission Over Packet-Switched NetworksabstractChannel distortion modeling for generic multi-view video transmission remains a unfilled blank, despite that intensive research efforts have been devoted to model traditional 2-D video transmission. This paper aims to fill this blank through developing a recursive distortion model for multi-view video transmission over lossy packet-switched networks. Based on the study on the characteristics of multi-view video coding and the propagating behavior of transmission error due to random frame losses, a recursive mathematical model is derived to estimate the expected channel-induced distortion at both the frame and sequence levels. The model we develop explicitly considers both temporal and inter-view dependencies, induced by motion-compensated and disparity-compensated coding, respectively. The derived model is applicable to all multi-view video encoders using the classical block-based motion-/disparity-compensated prediction framework. Both objective and subjective experimental results are presented to demonstrate that the proposed model is capable of effectively model channel-induced distortion for multi-view video. Yuan Zhou 0006, Chunping Hou, Wei Xiang 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Modeling of Transmission Distortion for Multi-View Video in Packet Lossy NetworksabstractIn this paper, a mathematical model is proposed to estimate the distortion caused by random packet losses for multi-view video transmission. Based on the study of multi-view video coding, the proposed model takes into account the disparity/motion compensation which relates the channel-induced distortion in the current frame with that in the previous frame or the adjacent view, and allows for any motion-compensated and disparity-compensated concealment method at the decoder. Comparative studies between the modeled and simulated distortion results demonstrates that the proposed model is able to estimate the transmission distortion of multi-view video with high accuracy. Yuan Zhou 0006, Chunping Hou, Wei Xiang 0001 |
GLOBECOM | 2 |
| 2010 | Image coding via sparse contourlet representationabstractWhen contourlet coefficients are directly coded, the benefits from directional representation are discounted by the redundancy of the transform. In this paper, we propose a contourlet-based image coding scheme under the sparse representation framework, in which contourlet coefficients are sparsified by the iterative thresholding method before compression. Dependency analysis is performed to reveal correlations among sparsified contourlet coefficients. Considering the characteristics of coefficient correlations, the obtained coefficients are coded with an intra-subband coder. Experimental results show that the coding performances of the contourlet-based schemes are significantly boosted via sparse representations. Jing-Yu Yang 0002, Chunping Hou, Wenli Xu |
ISCAS | 2 |