EDBT 2026 Demo / reviewers in the wild / expert
Fei Zhou 0001
dblp:84/5639-1
· DBLP profile ↗
66ranked-venue papers
17as first author
25since 2021 · last 2026
0000-0003-1216-2181ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 16 first-author · 23 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ultra-High-Definition Video Quality Assessment Based on Adaptively Spatial Fusion and Temporal SelectionabstractUltra high definition video quality assessment (UHD VQA) is challenging due to the high spatio-temporal dimensions of UHD content. In this work, we propose a spatio-temporal framework based on feature fusion and selection to facilitate UHD VOA. First, we introduce a spatially adaptive fusion net (SAFNet) that integrates all patch-wise features into a compact frame-wise representation via local aggregation and global interaction. During the aggregation and the interaction, the contribution of each patch-wise feature is initially estimated via visual saliency-derived weights and then adaptively refined by contextual adjustments. Second, we design a temporally adaptive selection net (TASNet) to reduces temporal redundancy. It retains a small number of temporal positions through thresholding a dynamically estimated importance vector, which captures both representativeness and inter-frame variation. Finally, extensive experiments conducted on 4 datasets with various 4K videos demonstrate the effectiveness of the proposed method in comparison with some traditional and state-of-the-art VQA methods. Zhijie Liang, Wei Sheng, Fei Zhou 0001, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Towards Smart Point-and-Shoot PhotographyabstractHundreds of millions of people routinely take photos using their smartphones as point and shoot (PAS) cameras, yet very few would have the photography skills to compose a good shot of a scene. While traditional PAS cameras have built-in functions to ensure a photo is well focused and has the right brightness, they cannot tell the users how to compose the best shot of a scene. In this paper, we present a first of its kind smart point and shoot (SPAS) system to help users to take good photos. Our SPAS proposes to help users to compose a good shot of a scene by automatically guiding the users to adjust the camera pose live on the scene. We first constructed a large dataset containing 320K images with camera pose information from 4000 scenes. We then developed an innovative CLIP-based Composition Quality Assessment (CCQA) model to assign pseudo labels to these images. The CCQA introduces a unique learnable text embedding technique to learn continuous word embeddings capable of discerning subtle visual quality differences in the range covered by five levels of quality description words {bad, poor, fair, good, perfect}. And finally we have developed a camera pose adjustment model (CPAM) which first determines if the current view can be further improved and if so it outputs the adjust suggestion in the form of two camera pose adjustment angles. The two tasks of CPAM make decisions in a sequential manner and each involves different sets of training samples, we have developed a mixture-of-experts model with a gated loss function to train the CPAM in an end-to-end manner. We will present extensive results to demonstrate the performances of our SPAS system using publicly available image composition datasets. Jiawan Li, Fei Zhou 0001, Zhipeng Zhong, Jiongzhi Lin, Guoping Qiu |
CVPR | 2 |
| 2025 | Geometric Distortion Guided Transformer for Omnidirectional Image Super-ResolutionabstractAs virtual and augmented reality applications gain popularity, omnidirectional image (ODI) super-resolution has become increasingly important. Unlike 2D plain images that are formed on a plane, ODIs are projected onto spherical surfaces. Applying established image super-resolution methods to ODIs, therefore, requires performing equirectangular projection (ERP) to map the ODIs onto a plane. ODI super-resolution needs to take into account geometric distortion resulting from ERP. However, without considering such geometric distortion of ERP images, previous methods only utilize a limited range of pixels and may easily miss self-similar textures for reconstruction. In this paper, we introduce a novel Geometric Distortion Guided Transformer for Omnidirectional image Super-Resolution (GDGT-OSR). Specifically, a distortion modulated rectangle-window selfattention mechanism, integrated with deformable self-attention, is proposed to better perceive the distortion and thus involve more self-similar textures. Distortion modulation is achieved through a newly devised distortion guidance generator that produces guidance for the rectangular windows by exploiting the variability of distortion across latitudes. Furthermore, we propose a dynamic feature aggregation scheme to adaptively fuse the features from different self-attention modules. We present extensive experimental results on public datasets and show that the new GDGT-OSR outperforms methods in existing literature. Cuixin Yang, Rongkang Dong, Jun Xiao 0010, Kin-Man Lam 0001, Fei Zhou 0001, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Vision-Language Knowledge Exploration for Video Saliency Prediction
Fei Zhou 0001, Baitao Huang, Guoping Qiu |
PRCV (9) | 1 |
| 2024 | Blind Image Quality Assessment Based on Separate Representations and Adaptive Interaction of Content and DistortionabstractThe visual quality of an image mainly relies on its content and its distortions. However, the adaptability between their contributions to the image quality has not be well investigated yet. Besides, albeit of many promising efforts, lacking sufficient labeled data still hinders the robust representation of quality-related information. In this work, we first design a self-supervised architecture, named collaborative autoencoder (COAE), to separately represent the content and the distortion information, and then develop a Self-Adaptive Weighting based quAlity predictoR (SAWAR) to balance the individual representations of the content and the distortions in the prediction of image quality. Specifically, the COAE is trained with large-scale unlabeled data, consisting of a content autoencoder (CAE) and a distortion autoencoder (DAE) that work collaboratively and individually. While the CAE is a standard autoencoder for the content representation, the design of the DAE is unique. We introduce the CAE-encoded content representation as an extra input to the decoder of the DAE to learn to reconstruct distorted images, thus effectively forcing it to extract the distortion representation. The SAWAR, whose parameter number is much smaller than that of the COAE, is trained with labeled data in existing IQA datasets. It takes advantage of the interaction between the image content and the distortions to adaptively balance their contributions. Extensive experiments show that the COAE effectively extracts quality-related representations and the SAWAR achieves the state-of-the-art performance. Zehong Zhou, Fei Zhou 0001, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | A Dataset and Model for the Visual Quality Assessment of Inversely Tone-Mapped HDR VideosabstractTo enhance the viewer experience of standard dynamic range (SDR) video content on high dynamic range (HDR) displays, inverse tone mapping (ITM) is employed. Objective visual quality assessment (VQA) models are needed for effective evaluation of ITM algorithms. However, there is a lack of specialized VQA models for assessing the visual quality of inversely tone-mapped HDR videos (ITM-HDR-Videos). This paper addresses both an algorithmic and a dataset gap by introducing a novel SDR referenced HDR (SD-R-HD) VQA model tailored for ITM-HDR-Videos, along with the first public dataset specifically constructed for this purpose. The innovations of the SD-R-HD VQA model include 1) utilizing available SDR video as a reference signal, 2) extracting features that characterize standard ITM operations such as global mapping and local compensation, and 3) directly modeling interframe inconsistencies introduced by ITM operations. The newly created ITM-HDR-VQA dataset comprises 200 ITM-HDR-Videos annotated with mean opinion scores, gathered over 320 man-hours of psychovisual experiments. Experimental results demonstrate that the SD-R-HD VQA model significantly outperforms existing state-of-the-art VQA models. Fei Zhou 0001, Shuhong Yuan, Zhijie Liang, Jiang Duan, Guoping Qiu |
IEEE Trans. Image Process. | 1 |
| 2023 | Aesthetically Relevant Image CaptioningabstractImage aesthetic quality assessment (AQA) aims to assign numerical aesthetic ratings to images whilst image aesthetic captioning (IAC) aims to generate textual descriptions of the aesthetic aspects of images. In this paper, we study image AQA and IAC together and present a new IAC method termed Aesthetically Relevant Image Captioning (ARIC). Based on the observation that most textual comments of an image are about objects and their interactions rather than aspects of aesthetics, we first introduce the concept of Aesthetic Relevance Score (ARS) of a sentence and have developed a model to automatically label a sentence with its ARS. We then use the ARS to design the ARIC model which includes an ARS weighted IAC loss function and an ARS based diverse aesthetic caption selector (DACS). We present extensive experimental results to show the soundness of the ARS concept and the effectiveness of the ARIC model by demonstrating that texts with higher ARS’s can predict the aesthetic ratings more accurately and that the new ARIC model can generate more accurate, aesthetically more relevant and more diverse image captions. Furthermore, a large new research database containing 510K images with over 5 million comments and 350K aesthetic scores, and code for implementing ARIC, are available at https://github.com/PengZai/ARIC Zhipeng Zhong, Fei Zhou 0001, Guoping Qiu |
AAAI | 2 |
| 2023 | Collaborative Auto-encoding for Blind Image Quality AssessmentabstractBlind image quality assessment (BIQA) is a challenging problem with important real-world applications. Recent efforts attempting to exploit powerful representations by deep neural networks (DNN) are hindered by the lack of subjectively annotated data. This paper presents a novel BIQA method which overcomes this fundamental obstacle. Specifically, we design a pair of collaborative autoencoders (COAE) consisting of a content autoencoder (CAE) and a distortion autoencoder (DAE) that work together to extract content and distortion representations, which are shown to be highly descriptive of image quality. While the CAE follows a standard codec procedure, we introduce the CAE-encoded feature as an extra input to the DAE's decoder for reconstructing distorted images, thus effectively forcing DAE's encoder to extract distortion representations. The self-supervised learning framework allows the COAE including two feature extractors to be trained by almost unlimited amount of data, thus leaving limited samples with annotations to finetune a BIQA model. We will show that the proposed BIQA method achieves state-of-the-art performance and has superior generalization capability over other learning based models. All the related codes, trained models, and the supplementary material are available at: https://github.com/Macro-Zhou/NRIQA-VISOR/. Zehong Zhou, Fei Zhou 0001, Guoping Qiu |
ICME | 2 |
| 2023 | Video Inverse Tone Mapping Network with Luma and Chroma Mapping
Peihuan Huang, Gaofeng Cao, Fei Zhou 0001, Guoping Qiu |
ACM Multimedia | 3 |
| 2023 | Restoration of Multiple Image Distortions using a Semi-dynamic Deep Neural NetworkabstractRestoring multiple image distortions with a single model is difficult because different distortions require fundamentally different processing mechanisms, e.g., deblurring requires high-pass filtering, while denoising requires low-pass filtering operations. This paper presents a dynamic universal image restoration (DUIR) system capable of simultaneously processing multiple distortions. The new model features several innovative designs: (i) a distortion embedding module (DEM) to automatically encode the distortion information of an input, (ii) a distortion attention module (DAM) that uses a bi-directional long short-term memory (LSTM) to encode the distortion into a sequence of forward and backward interdependent modulating signals, and (iii) a dynamically adaptive image restoration deep convolutional neural network (DAIR-DCNN) featuring unique semi-dynamic layers (SDLs) in which part of their parameters are dynamically modulated by the distortion signals. DEM, DAM, and SDLs together make DAIR-DCNN adaptive to the distortions of the current input, which in turn equips the DUIR system with the capability of simultaneously processing multiple image distortions with a single trained model. We present extensive experimental results to show that the new technique achieves superior performance to state-of-the-art models on both synthetic and real data. We further demonstrate that a trained DUIR system can simultaneously handle different distortions, including those with conflicting demands, such as denoising, deblurring, and compression artifact removal. Hongming Luo, Fei Zhou 0001, Zehong Zhou, Kin-Man Lam 0001, Guoping Qiu |
ACM Multimedia | 2 |
| 2023 | Super-resolving compressed images via parallel and series integration of artefacts removal and resolution enhancement
Hongming Luo, Fei Zhou 0001, Guangsen Liao, Guoping Qiu |
Signal Process. | 2 |
| 2023 | Super-resolution image visual quality assessment based on structure-texture features
Fei Zhou 0001, Wei Sheng, Zitao Lu, Bo Kang, Mianyi Chen, Guoping Qiu |
Signal Process. Image Commun. | 1 |
| 2023 | Statistical hypothesis testing as a novel perspective of pooling for image quality assessmentabstractImage quality assessment is usually achieved by pooling local quality scores. However, commonly used pooling strategies, based on simple sample statistics, are not always sensitive to distortions. In this short communication, we propose a novel perspective of pooling: reliable pooling through statistical hypothesis testing, which enables effective detection of subtle changes of population parameters when the underlying distribution of local quality scores is affected by distortions. To illustrate the significance of this novel perspective, we design a new pooling strategy utilising simple one-sided one-sample t-test. The experiments on benchmark databases show the reliability of hypothesis testing-based pooling, compared with state-of-the-art pooling strategies. Rui Zhu 0006, Fei Zhou 0001, Wenming Yang, Jing-Hao Xue |
Signal Process. Image Commun. | 2 |
| 2023 | A Decoupled Kernel Prediction Network Guided by Soft Mask for Single Image HDR ReconstructionabstractRecent works on single image high dynamic range (HDR) reconstruction fail to hallucinate plausible textures, resulting in information missing and artifacts in large-scale under/over-exposed regions. In this article, a decoupled kernel prediction network is proposed to infer an HDR image from a low dynamic range (LDR) image. Specifically, we first adopt a simple module to generate a preliminary result, which can precisely estimate well-exposed HDR regions. Meanwhile, an encoder-decoder backbone network with a soft mask guidance module is presented to predict pixel-wise kernels, which is further convolved with the preliminary result to obtain the final HDR output. Instead of traditional kernels, our predicted kernels are decoupled along the spatial and channel dimensions. The advantages of our method are threefold at least. First, our model is guided by the soft mask so that it can focus on the most relevant information for under/over-exposed regions. Second, pixel-wise kernels are able to adaptively solve the different degradations for differently exposed regions. Third, decoupled kernels can avoid information redundancy across channels and reduce the solution space of our model. Thus, our method is able to hallucinate fine details in the under/over-exposed regions and renders visually pleasing results. Extensive experiments demonstrate that our model outperforms state-of-the-art ones. Gaofeng Cao, Fei Zhou 0001, Kanglin Liu, Anjie Wang, Leidong Fan |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | KPN-MFI: A Kernel Prediction Network with Multi-frame Interaction for Video Inverse Tone MappingabstractUp to now, the image-based inverse tone mapping (iTM) models have been widely investigated, while there is little research on video-based iTM methods. It would be interesting to make use of these existing image-based models in the video iTM task. However, directly transferring the imagebased iTM models to video data without modeling spatial-temporal information remains nontrivial and challenging. Considering both the intra-frame quality and the inter-frame consistency of a video, this article presents a new video iTM method based on a kernel prediction network (KPN), which takes advantage of multi-frame interaction (MFI) module to capture temporal-spatial information for video data. Specifically, a basic encoder-decoder KPN, essentially designed for image iTM, is trained to guarantee the mapping quality within each frame. More importantly, the MFI module is incorporated to capture temporal-spatial context information and preserve the inter-frame consistency by exploiting the correction between adjacent frames. Notably, we can readily extend any existing image iTM models to video iTM ones by involving the proposed MFI module. Furthermore, we propose an inter-frame brightness consistency loss function based on the Gaussian pyramid to reduce the video temporal inconsistency. Extensive experiments demonstrate that our model outperforms state-ofthe-art image and video-based methods. The code is available at https://github.com/caogaofeng/KPNMFI. Gaofeng Cao, Fei Zhou 0001, Han Yan 0003, Anjie Wang, Leidong Fan |
IJCAI | 2 |
| 2022 | Restoration of User Videos Shared on Social MediaabstractUser videos shared on social media platforms usually suffer from degradations caused by unknown proprietary processing procedures, which means that their visual quality is poorer than that of the originals. This paper presents a new general video restoration framework for the restoration of user videos shared on social media platforms. In contrast to most deep learning-based video restoration methods that perform end-to-end mapping, where feature extraction is mostly treated as a black box, in the sense that what role a feature plays is often unknown, our new method, termed Video restOration through adapTive dEgradation Sensing (VOTES), introduces the concept of a degradation feature map (DFM) to explicitly guide the video restoration process. Specifically, for each video frame, we first adaptively estimate its DFM to extract features representing the difficulty of restoring its different regions. We then feed the DFM to a convolutional neural network (CNN) to compute hierarchical degradation features to modulate an end-to-end video restoration backbone network, such that more attention is paid explicitly to potentially more difficult to restore areas, which in turn leads to enhanced restoration performance. We will explain the design rationale of the VOTES framework and present extensive experimental results to show that the new VOTES method outperforms various state-of-the-art techniques both quantitatively and qualitatively. In addition, we contribute a large scale real-world database of user videos shared on different social media platforms. Codes and datasets are available at https://github.com/luohongming/VOTES.git Hongming Luo, Fei Zhou 0001, Kin-Man Lam 0001, Guoping Qiu |
ACM Multimedia | 2 |
| 2022 | Metric of choosing the optimal parameter setting for edge aware filteringabstractAbstract Most of the existing edge aware filters have a number of parameters. The optimal settings of these parameters guarantee the best performance but they depend on the input image. It would be difficult for inexperienced users to empirically get the optimal parameters. This paper proposes a new metric for choosing the optimal parameter settings of edge aware filtering, which is called metric of edge aware filtering (MEAF). MEAF evaluates the quality of filtered images from three aspects: the color distance, the distance of the prominent structure, and smoothness. The colour distance is calculated by rooted mean square error. The distance of the prominent structure is calculated by the proposed SSIM map masked by Sobel edges (MASKED‐SSIM). In MASKED‐SSIM, Sobel detector is used to detect the prominent structure from the input image and the calculation of structure distance is constrained on the prominent structures. Number of gradients and relative total variation are further defined to measure the smoothness of the filtered image. MEAF is an objective metric, which is specially designed for choosing the optimal parameters of edge aware filtering and can be used universally for arbitrary input image. Experiments on 12 state‐of‐the‐art edge aware filters show the effectiveness of MEAF. Fei Zhou 0001, Bin Li 0011 |
IET Image Process. | 2 |
| 2022 | ASSP: An adaptive sample statistics-based pooling for full-reference image quality assessment
Yurong Ling, Fei Zhou 0001, Kun Guo 0004, Jing-Hao Xue |
Neurocomputing | 2 |
| 2022 | Tone mapping high dynamic range images based on region-adaptive self-supervised deep learning
Fei Zhou 0001, Guangsen Liao, Jiang Duan, Guoping Qiu |
Signal Process. Image Commun. | 1 |
| 2022 | Towards Disentangling Latent Space for Unsupervised Semantic Face EditingabstractFacial attributes in StyleGAN generated images are entangled in the latent space which makes it very difficult to independently control a specific attribute without affecting the others. Supervised attribute editing requires annotated training data which is difficult to obtain and limits the editable attributes to those with labels. Therefore, unsupervised attribute editing in an disentangled latent space is key to performing neat and versatile semantic face editing. In this paper, we present a new technique termed Structure-Texture Independent Architecture with Weight Decomposition and Orthogonal Regularization (STIA-WO) to disentangle the latent space for unsupervised semantic face editing. By applying STIA-WO to GAN, we have developed a StyleGAN termed STGAN-WO which performs weight decomposition through utilizing the style vector to construct a fully controllable weight matrix to regulate image synthesis, and employs orthogonal regularization to ensure each entry of the style vector only controls one independent feature matrix. To further disentangle the facial attributes, STGAN-WO introduces a structure-texture independent architecture which utilizes two independently and identically distributed (i.i.d.) latent vectors to control the synthesis of the texture and structure components in a disentangled way. Unsupervised semantic editing is achieved by moving the latent code in the coarse layers along its orthogonal directions to change texture related attributes or changing the latent code in the fine layers to manipulate structure related ones. We present experimental results which show that our new STGAN-WO can achieve better attribute editing than state of the art methods. Kanglin Liu, Gaofeng Cao, Fei Zhou 0001, Jiang Duan, Guoping Qiu |
IEEE Trans. Image Process. | 3 |
| 2021 | Noise Robust Video Super-Resolution Without Training on Noisy Data
Fei Zhou 0001, Zitao Lu, Hongming Luo, Cuixin Yang |
ICIG (3) | 1 |
| 2021 | Self-Supervised Video Super-Resolution by Spatial Constraint and Temporal Fusion
Cuixin Yang, Hongming Luo, Guangsen Liao, Zitao Lu, Fei Zhou 0001, Guoping Qiu |
PRCV (3) | 5 |
| 2021 | A brightness-adaptive kernel prediction network for inverse tone mapping
Gaofeng Cao, Fei Zhou 0001, Kanglin Liu |
Neurocomputing | 2 |
| 2021 | Visual Saliency via Selecting and Reweighting Features in Hierarchical Fusion NetworkabstractRecently, computational models based on deep neural networks have made impressive progress in predicting the visual saliency of human beings. Relying on the powerful capability of some pre-trained networks, various models can extract diverse deep features. However, they are unaware of the problem of feature selection and reweighting when predicting saliency. This situation gives rise to features describing scene distractors potentially also contributing to the saliency maps. In this paper, we propose a feature selection and reweighting module (FSRM) for deep saliency prediction models. Through the FSRM, we wish to highlight the saliency-related features in a manner similar to channel attention and simultaneously exclude distractor features by reducing the channel number of deep features. Specifically, in the FSRM, we obtain an importance descriptor of feature channels, where some saliency knowledge including the center prior and rarity is encoded. Furthermore, the number of feature channels is reduced via a transformation matrix derived from the importance descriptor. To predict the saliency, the FSRM is embedded in a hierarchical fusion network that makes use of multi-level features. Experiments and ablation studies show the effectiveness and generalization capability of the FSRM in the saliency prediction. Fei Zhou 0001, Junhua Chen 0004 |
IEEE Signal Process. Lett. | 1 |
| 2021 | Image Defogging Quality Assessment: Real-World Database and MethodabstractFog removal from an image is an active research topic in computer vision. However, current literature is weak in the following two areas which in many ways are hindering progress for developing defogging algorithms. First, there is no true real-world and naturally occurring foggy image datasets suitable for developing defogging models. Second, there is no suitable mathematically simple and easy to use image quality assessment (IQA) methods for evaluating the visual quality of defogged images. We address these two aspects in this paper. We first introduce a new foggy image dataset called multiple real-world foggy image dataset (MRFID). MRFID contains foggy and clear images of 200 outdoor scenes. For each scene, one clear image and 4 foggy images of different densities defined as slightly foggy, moderately foggy, highly foggy, and extremely foggy, are manually selected from images taken from these scenes over the course of one calendar year. We then process the foggy images of MRFID using 16 defogging methods to obtain 12,800 defogged images (DFIs) and perform a comprehensive subjective evaluation of the visual quality of the DFIs. Through collecting the mean opinion score (MOS) of 120 subjects and evaluating a variety of fog-relevant image features, we have developed a new Fog-relevant Feature based SIMilarity index (FRFSIM) for assessing the visual quality of DFIs. We present extensive experimental results to show that our new visual quality assessment measure, the FRFSIM, is more consistent with the MOS than other IQA methods and is therefore more suitable for evaluating defogged images than other state-of-the-art IQA methods. Our dataset and relevant code are available at http://www.vistalab.ac.cn/MRFID-for-defogging/. Wei Liu 0123, Fei Zhou 0001, Tao Lu 0001, Jiang Duan, Guoping Qiu |
IEEE Trans. Image Process. | 2 |
| 2020 | VHS to HDTV Video Translation Using Multi-task Adversarial Learning
Hongming Luo, Guangsen Liao, Xianxu Hou, Fei Zhou 0001, Guoping Qiu |
MMM (1) | 5 |
| 2020 | Deep learning for image super-resolution
Wenming Yang, Fei Zhou 0001, Rui Zhu 0006, Kazuhiro Fukui, Guijin Wang, Jing-Hao Xue |
Neurocomputing | 2 |
| 2020 | Spectral regularization for combating mode collapse in GANs
Kanglin Liu, Guoping Qiu, Wenming Tang, Fei Zhou 0001 |
Image Vis. Comput. | 4 |
| 2020 | Defocus map estimation from a single image using improved likelihood feature and edge-based basis
Qingmin Liao, Jing-Hao Xue, Fei Zhou 0001 |
Pattern Recognit. | 4 |
| 2020 | Special Issue on Advances in Statistical Methods-based Visual Quality Assessment
Fei Zhou 0001, Wenming Yang, Xinbo Gao 0001, Hantao Liu, Rui Zhu 0006, Jing-Hao Xue |
Signal Process. Image Commun. | 1 |
| 2020 | Structure and Texture-Aware Image Decomposition via Training a Neural NetworkabstractStructure-texture image decomposition is a funda-mental but challenging topic in computational graphics and image processing. In this paper, we introduce a structure-aware and a texture-aware measures to facilitate the structure-texture de-composition (STD) of images. Edge strengths and spatial scales that have been widely-used in previous STD researches cannot describe the structures and textures of images well. The proposed two measures differentiate image textures from image structures based on their distinctive characteristics. Specifically, the first one aims to measure the anisotropy of local gradients, and the second one is designed to measure the repeatability degree of signal pat-terns in a neighboring region. Since these two measures describe different properties of image structures and textures, they are complementary to each other. The STD is achieved by optimizing an objective function based on the two new measures. As using traditional optimization methods to solve the optimization prob-lem will require designing different optimizers for different func-tional spaces, we employ an architecture of deep neural network to optimize the STD cost function in a unified manner. The ex-perimental results demonstrate that, as compared with some state-of-the-art methods, our method can better separate image structure and texture and result in shaper edges in the structural component. Furthermore, to demonstrate the usefulness of the proposed STD method, we have successfully applied it to several applications including detail enhancement, edge detection, and visual quality assessment of super-resolved images. Fei Zhou 0001, Guoping Qiu |
IEEE Trans. Image Process. | 1 |
| 2020 | Visual Saliency via Embedding Hierarchical Knowledge in a Deep Neural NetworkabstractDeep neural networks (DNNs) have been extensively applied in image processing, including visual saliency map pre-diction of images. A major difficulty in using a DNN for visual saliency prediction is the lack of labeled ground truth of visual saliency. A powerful DNN usually contains a large number of trainable parameters. This condition can easily lead to model over-fitting. In this study, we develop a novel method that over-comes such difficulty by embedding hierarchical knowledge of existing visual saliency models in a DNN. We achieve the objective of exploiting the knowledge contained in the existing visual sali-ency models by using saliency maps generated by local, global, and semantic models to tune and fix about 92.5% of the parame-ters in our network in a hierarchical manner. As a result, the number of trainable parameters that need to be tuned by the ground truth is considerably reduced. This reduction enables us to fully utilize the power of a large DNN and overcome the issue of over-fitting at the same time. Furthermore, we introduce a simple but very effective center prior in designing the learning cost function of the DNN by attaching high importance to the errors around the image center. We also present extensive experimental results on four commonly used public databases to demonstrate the superiority of the proposed method over classical and state-of-the-art methods on various evaluation metrics. Fei Zhou 0001, Rongguo Yao, Guangsen Liao, Guoping Qiu |
IEEE Trans. Image Process. | 1 |
| 2019 | Spectral Regularization for Combating Mode Collapse in GANsabstractDespite excellent progress in recent years, mode collapse remains a major unsolved problem in generative adversarial networks (GANs). In this paper, we present spectral regularization for GANs (SR-GANs), a new and robust method for combating the mode collapse problem in GANs. Theoretical analysis shows that the optimal solution to the discriminator has a strong relationship to the spectral distributions of the weight matrix. Therefore, we monitor the spectral distribution in the discriminator of spectral normalized GANs (SN-GANs), and discover a phenomenon which we refer to as spectral collapse, where a large number of singular values of the weight matrices drop dramatically when mode collapse occurs. We show that there are strong evidence linking mode collapse to spectral collapse; and based on this link, we set out to tackle spectral collapse as a surrogate of mode collapse. We have developed a spectral regularization method where we compensate the spectral distributions of the weight matrices to prevent them from collapsing, which in turn successfully prevents mode collapse in GANs. We provide theoretical explanations for why SR-GANs are more stable and can provide better performances than SN-GANs. We also present extensive experimental results and analysis to show that SR-GANs not only always outperform SN-GANs but also always succeed in combating mode collapse where SN-GANs fail. Kanglin Liu, Guoping Qiu, Wenming Tang, Fei Zhou 0001 |
ICCV | 4 |
| 2019 | Visual Quality Assessment for Super-Resolved Images: Database and MethodabstractImage super-resolution (SR) has been an active research problem which has recently received renewed interest due to the introduction of new technologies such as deep learning. However, the lack of suitable criteria to evaluate the SR performance has hindered technology development. In this paper, we fill a gap in the literature by providing the first publicly available database as well as a new image quality assessment (IQA) method specifically designed for assessing the visual quality of super-resolved images (SRIs). In constructing the quality assessment database for SRIs (QADS), we carefully selected 20 reference images and created 980 SRIs using 21 image SR methods. Mean opinion score (MOS) for these SRIs is collected through 100 individuals participating in a suitably designed psychovisual experiment. Extensive numerical and statistical analysis is performed to show that the MOS of QADS has excellent suitability and reliability. The psychovisual experiment has led to the discovery that, unlike distortions encountered in other IQA databases, artifacts of the SRIs degenerate the image structure as well as the image texture. Moreover, the structural and textural degenerations have distinctive perceptual properties. Based on these insights, we propose a novel method to assess the visual quality of SRIs by separately considering the structural and textural components of images. Observing that textural degenerations are mainly attributed to dissimilar texture or checkerboard artifacts, we propose to measure the changes of textural distributions. We also observe that structural degenerations appear as blurring and jaggies artifacts in SRIs and develop separate similarity measures for different types of structural degenerations. A new pooling mechanism is then used to fuse the different similarities together to give the final quality score for an SRI. The experiments conducted on the QADS demonstrate that our method significantly outperforms the classical as well as current state-of-the-art IQA methods. Fei Zhou 0001, Rongguo Yao, Guoping Qiu |
IEEE Trans. Image Process. | 1 |
| 2018 | Full-Reference Quality Assessment of Contrast Changed Images Based on Local Linear ModelabstractThis paper presents a new full-reference method to assess the quality of contrast changed images. In this method, we employ a linear model to describe the relationship between local patches of reference images and contrast changed images. With parameters of this model, three quality measures considering contrast comparison, structure variation, and luminance change are defined. Among them, the first measure produces larger quality scores for higher contrast, which is different from traditional forms of quality measures used in most existing full-reference methods. Experiments on four benchmark databases show that the proposed method is superior to state-of-the-art methods in assessing the quality of contrast changed images. Wenming Yang, Fei Zhou 0001, Qingmin Liao |
ICASSP | 3 |
| 2018 | MvSSIM: A quality assessment index for hyperspectral images
Rui Zhu 0006, Fei Zhou 0001, Jing-Hao Xue |
Neurocomputing | 2 |
| 2018 | SPSIM: A Superpixel-Based Similarity Index for Full-Reference Image Quality AssessmentabstractFull-reference image quality assessment algorithms usually perform comparisons of features extracted from square patches. These patches do not have any visual meanings. On the contrary, a superpixel is a set of image pixels that share similar visual characteristics and is thus perceptually meaningful. Features from superpixels may improve the performance of image quality assessment. Inspired by this, we propose a new superpixel-based similarity index by extracting perceptually meaningful features and revising similarity measures. The proposed method evaluates image quality on the basis of three measurements, namely, superpixel luminance similarity, superpixel chrominance similarity, and pixel gradient similarity. The first two measurements assess the overall visual impression on local images. The third measurement quantifies structural variations. The impact of superpixel-based regional gradient consistency on image quality is also analyzed. Distorted images showing high regional gradient consistency with the corresponding reference images are visually appreciated. Therefore, the three measurements are further revised by incorporating the regional gradient consistency into their computations. A weighting function that indicates superpixel-based texture complexity is utilized in the pooling stage to obtain the final quality score. Experiments on several benchmark databases demonstrate that the proposed method is competitive with the state-of-the-art metrics. Qingmin Liao, Jing-Hao Xue, Fei Zhou 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Wavelet-based single image super-resolution with an overall enhancement procedureabstractIn this paper, we address the problem of generating a super-resolution image based on a dictionary of low- and high-resolution exemplars from a single input image in wavelet domain with a overall enhancement procedure. Most methods extract different kinds of features in low-resolution image and high-resolution images to establish the mapping relation. But in this paper, we implement wavelet-transform to extract the same kind of feature to make the mapping more reasonable. Meanwhile we implement local Lipschitz regularity constraint and structure-keeping constraint to preserve the local singularity and edge in our method. Compared with current state-of-art methods on standard images, our method obtains both visual and PSNR improvement. Zongqing Lu 0001, Quan Zou 0001, Fei Zhou 0001, Qingmin Liao |
ICASSP | 3 |
| 2017 | A robust feature descriptor based on multiple gradient-related featuresabstractIn this paper, we propose a robust descriptor named as multiple gradient-related features (MGRF) in virtue of local and overall order encoding. Specifically, three types of features are introduced, including multidirectional gradient, gradient orientation, and first derivative of gradient orientation, each of which represents different aspect of region of interest (ROI). To extract these features, we also propose a novel sampling pattern of tree structure. Furthermore, each gradient-related feature is encoded with both local and overall order information of ROI, and the encoding results are respectively called local and overall gradient order code (GOC). Finally, our descriptor is formed by concatenating the respective feature vector of each type of feature, which is computed as a 2-D joint histogram of GOC and ordinal bin. The experiments conducted on Oxford dataset demonstrate that the proposed descriptor significantly outperforms other state-of-the-art descriptors. Zhaomang Sun, Fei Zhou 0001, Qingmin Liao |
ICASSP | 2 |
| 2017 | Learning adaptive local distance metric for face hallucinationabstractIn this paper, we propose a novel method for face hallucination by learning a new distance metric in the low-resolution (LR) patch space (source space). Local patch-based face hallucination methods usually assume that the two manifolds formed by LR and high-resolution (HR) image patches have similar local geometry. However, this assumption does not hold well in practice. Motivated by metric learning in machine learning, we propose to learn a new distance metric in the source space, under the supervision of the true local geometry in the target space (HR patch space). The learned new metric gives more freedom to the presentation of local geometry in the source space, and thus the local geometries of source and target space turn to be more consistent. Experiments conducted on two datasets demonstrate that the proposed method is superior to the state-of-the-art face hallucination and image super-resolution (SR) methods. Yuanpeng Zou, Fei Zhou 0001, Qingmin Liao |
ICASSP | 2 |
| 2017 | Real-time 3D face reconstruction from one single image by displacement mappingabstractIn this paper, we present a fast and robust method to reconstruct a plausible three-dimension (3D) face from one single frontal face image. In training phase, we classify the faces into several groups based on the facial structures and propose to learn a mapping, known as the displacement mapping (DM) in this paper, for each group. DM relates two displacements: One displacements, denoted as 2D displacements, represent the differences between the positions of feature points on the 2D training faces and those on the reference 2D face that has been pre-defined for the corresponding group; another displacements, denoted as 3D displacements, are the differences between the positions of vertices on the reconstructed 3D face and those on the reference 3D face that is also pre-defined. During the reconstruction phase, we first classify the input face as one of the groups and calculate the 2D displacements. Then we take advantage of the 2D displacements and the learned DM to estimate the 3D displacements. Subsequently, 3D displacements can be used to obtain the precise 3D face by shifting the 3D reference face. Experiments on Basel face model (BFM) database as well as some real-world 2D face images demonstrate the effectiveness and efficiency of the proposed method, in comparison with some state-of-arts methods. Fei Zhou 0001, Qingmin Liao |
ICIP | 2 |
| 2017 | Microstructure analysis of silk samples using mueller matrix determination and sparse representationabstractIn this paper, we propose to use Mueller matrix determination and sparse representation to classify silk samples washed in different detergents. Different detergents have different effects on the same silk samples after washing, and we distinguish their diversities in the Mueller matrix images(MMI) instead of visible light images(VLI). Compared with VLI, Mueller matrix, also known as polarization image, reflects the wavelength-scale microstructure and some optical properties of samples, and focuses on extracting the index to research the polarization property. To achieve a good performance with the microstructure analysis, we utilize the method of sparse representation which uses the reconstruction error for classification. Generally speaking, we introduce to combine Mueller matrix with sparse representation in the classification of the same silk samples washed in different detergents, and the high precision in experimental results indicates that our method works well. Fei Zhou 0001, Hui Ma 0003, Qingmin Liao |
ICIP | 2 |
| 2017 | MDID: A multiply distorted image database for image quality assessment
Fei Zhou 0001, Qingmin Liao |
Pattern Recognit. | 2 |
| 2017 | Single-Image Super-Resolution by Subdictionary Coding and Kernel RegressionabstractIn this paper, we present a new learning-based single-image super-resolution (SR) approach, inspired by existing sparse representation-based methods. As a promising image modeling theory, sparse representation has been effectively applied to solve the image SR problem, usually with the use of pretrained coupled or semi-coupled dictionaries. In our proposed method, we train independent dictionaries for high-resolution (HR) and low-resolution (LR) image patches to endow them more flexibility of expression. We use local subdictionaries to adaptively code image patches, which can characterize image local structures better and ensure the sparsity property of the image. Furthermore, we use kernel regression to relate HR and LR coding coefficients to capture and map the intrinsic nonlinear relationship between them. Such mapping is of central importance in the image SR problem, because high-order statistics play a significant role in the reconstruction of the detail structure of an HR image. The proposed model is generic for image SR in terms of two categories of blurring kernel. Experimental results show that our method can effectively reconstruct image details and outperform state-of-the-art algorithms in both quantitative and visual comparisons. Wenming Yang, Tingrong Yuan, Wei Wang 0194, Fei Zhou 0001, Qingmin Liao |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2016 | Saliency detection based on integration of central bias, reweighting and multi-scale for superpixelsabstractSaliency detection has been a significant problem in computer vision and helpful to object detection. In this paper, we propose a new computational saliency detection model under the Bayesian framework. First, central bias and the reweighting of the salient regions in the convex hull are applied to guide the prior map. Then, multi-scale for superpixels is proposed to detect objects with various scales. At last, the Bayes formula is adopted to obtain the final saliency map. Experimental results on a standard database show that the proposed model outperforms state-of-the-art methods. Xiaoling Hu 0002, Wenming Yang, Fei Zhou 0001, Qingmin Liao |
ICASSP | 3 |
| 2016 | A fast 3D face reconstruction method from a single image using adjustable modelabstractIn this paper, we propose a fast and robust method which uses only a single frontal face image as input to reconstruct a plausible 3D face. Our method mainly consists of three stages: feature point detection, model adaptation in X-Y plane and model adjustment on Z-axis direction. At first stage, we detect some face regions such as face contour and facial components automatically. In these regions, we extract several feature points which can generally describe the structure of face. Subsequently, we apply several deformation processes and optimization procedures on an adjustable 3D face model in the X-Y plane based on these feature points. Finally, we present a method of insertion to obtain a dense and smooth model. Experimental results demonstrate the effectiveness and efficiency of our method as well as the robust adaptation to the complex imaging condition. Fei Zhou 0001, Qingmin Liao |
ICASSP | 2 |
| 2016 | Anchored neighborhood regression based single image super-resolution from self-examplesabstractIn this paper, we present a novel self-learning single image super-resolution (SR) method, which restores a high-resolution (HR) image from self-examples extracted from the low-resolution (LR) input image itself without relying on extra external training images. In the proposed method, we directly use sampled image patches as the anchor points, and then learn multiple linear mapping functions based on anchored neighborhood regression to transform LR space into HR space. Moreover, we utilize the flipped and rotated versions of the self-examples to expand the internal patch space. Experimental comparison on standard benchmarks with state-of-the-art methods validates the effectiveness of the proposed approach. Yapeng Tian, Fei Zhou 0001, Wenming Yang, Xuesen Shang, Qingmin Liao |
ICIP | 2 |
| 2016 | Visual domain adaptation using weighted subspace alignmentabstractDomain Adaptation (DA) has attracted a lot of attention in recent years. DA aims at overcoming the covariate shift in dataset and aligning multiple existing but partially related data collections. In this paper, we propose a new DA algorithm which aligns the weighted subspaces generated from source samples and target samples. The weighted subspaces of source samples are generated using weighted Principal Component Analysis (PCA). Specifically, the source samples closer to the target domain are given higher weights during the construction of subspaces, which is definitely beneficial for building an adaptable classifier. Subsequently, the weighted subspaces of source samples and the subspaces of target samples are aligned to achieve domain adaptation. Experimental results on standard datasets demonstrate the advantages of our approach over state-of-the-art DA approaches. Shuo Chen 0010, Fei Zhou 0001, Qingmin Liao |
VCIP | 2 |
| 2016 | Defocus Map Estimation From a Single Image Based on Two-Parameter Defocus ModelabstractDefocus map estimation (DME) is highly important in many computer vision applications. Nearly, all existing approaches for DME from a single image are based on a one-parameter defocus model, which does not allow for the variation of depth over edges. In this paper, a novel two-parameter model of defocused edges is proposed for DME from a single image. We can estimate the defocus amounts for each side of the edges through this proposed model, and the confidence that the edge is a pattern edge, where the depth remains the same over the edge, can be generated. Then, we modify the TV-L1 algorithm for structure-texture decomposition by taking advantage of this confidence to eliminate pattern edges while preserving structural ones. Finally, the defocus amounts estimated at the edge positions are used as initial values, and the structure component is employed as a guidance in the following Laplacian matting procedure to avoid the influence of pattern edges on the final defocus map. Experiment results show that the proposed method can effectively eliminate the influence of pattern edges compared with the state-of-art method. Furthermore, the estimated defocus map is feasible in applications of depth estimation and foreground/background segmentation. Fei Zhou 0001, Qingmin Liao |
IEEE Trans. Image Process. | 2 |
| 2016 | Consistent Coding Scheme for Single-Image Super-Resolution Via Independent DictionariesabstractIn this paper, we present a unified frame based on collaborative representation (CR) for single-image super-resolution (SR), which learns low-resolution (LR) and high-resolution (HR) dictionaries independently in the training stage and adopts a consistent coding scheme (CCS) to guarantee the prediction accuracy of HR coding coefficients during SR reconstruction. The independent LR and HR dictionaries are learned based on CR with l2-norm regularization, which can well describe the corresponding LR and HR patch space, respectively. Furthermore, a mapping function is learned to map LR coding coefficients onto the corresponding HR coding coefficients. Propagation filtering can achieve smoothing over an image while preserving image context like edges or textural regions. Moreover, to preserve the edge structures of a super-resolved image and suppress artifacts, a propagation filtering-based constraint and image nonlocal self-similarity regularization are introduced into the SR reconstruction framework. Experimental comparison with state-of-the-art single image SR algorithms validates the effectiveness of proposed approach. Wenming Yang, Yapeng Tian, Fei Zhou 0001, Qingmin Liao, Hai Chen, Chenglin Zheng |
IEEE Trans. Multim. | 3 |
| 2015 | Single-frame image super-resolution inspired by perceptual criteriaabstractIn this study, the authors consider the problem of image super‐resolution (SR) in terms of the perceptual criteria. Existing SR methods treat the traditional mean‐squared error (MSE) as an irreplaceable objective function. However, MSE has been widely criticised since it is inconsistent with visual perception of human beings. The perceptual criteria, including the structural similarity (SSIM) index and feature similarity (FSIM) index, have been reported to be more effective in assessing image quality. Therefore SSIM and FSIM are included for the SR task in this study. Specifically, the authors first propose to reform principal component analysis (PCA), which is named as visual perceptual PCA (VP‐PCA), by adopting SSIM as the object function. Subsequently, to accomplish the SR task, the authors cluster the training data and perform VP‐PCA on each cluster to calculate the coefficients. Finally, based on the principle of FSIM, the traditional SR results and the SR results using VP‐PCA are combined to form our fused results. Experimental results are provided to show the superiority of the proposed method over several state‐of‐the‐art methods in both quantitative and visual comparisons. Fei Zhou 0001, Qingmin Liao |
IET Image Process. | 1 |
| 2015 | Single-Image Super-Resolution Based on Compact KPCA Coding and Kernel RegressionabstractIn this letter, we propose a novel approach for single-image super-resolution (SR). Our method is based on the idea of learning a dictionary which can capture the high-order statistics of high-resolution (HR) images. It is of central importance in image SR application, since the high-order statistics play a significant role in the reconstruction of HR image structure. Kernel principal component analysis (KPCA) is adopted to learn such a dictionary. A compact solution is adopted to reduce the time complexity of learning and testing for KPCA. Meanwhile, kernel ridge regression is employed to connect the input low-resolution (LR) image patches with the HR coding coefficients. Experimental results show that the proposed method is effective and efficient in comparison with state-of-art algorithms. Fei Zhou 0001, Tingrong Yuan, Wenming Yang, Qingmin Liao |
IEEE Signal Process. Lett. | 1 |
| 2015 | Coupled Attribute Similarity Learning on Categorical DataabstractAttribute independence has been taken as a major assumption in the limited research that has been conducted on similarity analysis for categorical data, especially unsupervised learning. However, in real-world data sources, attributes are more or less associated with each other in terms of certain coupling relationships. Accordingly, recent works on attribute dependency aggregation have introduced the co-occurrence of attribute values to explore attribute coupling, but they only present a local picture in analyzing categorical data similarity. This is inadequate for deep analysis, and the computational complexity grows exponentially when the data scale increases. This paper proposes an efficient data-driven similarity learning approach that generates a coupled attribute similarity measure for nominal objects with attribute couplings to capture a global picture of attribute similarity. It involves the frequency-based intra-coupled similarity within an attribute and the inter-coupled similarity upon value co-occurrences between attributes, as well as their integration on the object level. In particular, four measures are designed for the inter-coupled similarity to calculate the similarity between two categorical values by considering their relationships with other attributes in terms of power set, universal set, joint set, and intersection set. The theoretical analysis reveals the equivalent accuracy and superior efficiency of the measure based on the intersection set, particularly for large-scale data sets. Intensive experiments of data structure and clustering algorithms incorporating the coupled dissimilarity metric achieve a significant performance improvement on state-of-the-art measures and algorithms on 13 UCI data sets, which is confirmed by the statistical analysis. The experiment results show that the proposed coupled attribute similarity is generic, and can effectively and efficiently capture the intrinsic and global interactions within and between attributes for especially large-scale categorical data sets. In addition, two new coupled categorical clustering algorithms, i.e., CROCK and CLIMBO are proposed, and they both outperform the original ones in terms of clustering quality on UCI data sets and bibliographic data. Can Wang 0004, Xiangjun Dong 0001, Fei Zhou 0001, Longbing Cao, Chihung Chi |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Image super-resolution via Kernel regression of sparse coefficientsabstractIn this paper, we present a sparse coding (SC) inspired method to reconstruct a high-resolution (HR) image from one single low-resolution (LR) image. Instead of restricting the coding coefficients of LR and HR image patches to be equal or linearly mapped, we introduce kernel regression to nonlinearly relate the coding coefficients of LR patches and those of corresponding HR ones in an implicit fashion. Meanwhile, principal component analysis (PCA) is employed to train independent dictionaries which can well express image geometrical structure and ensure image sparse property. Experimental results show that the proposed method can effectively reconstruct image details and outperforms state-of-the-art algorithms in both quantitative and visual comparisons. Tingrong Yuan, Fei Zhou 0001, Wenming Yang, Qingmin Liao |
ICASSP | 2 |
| 2014 | Image amplification based on pixel-splittingabstractIn this paper, we propose a pixel-splitting based image amplification method, which involves two main operations: edge-keeping and mean-keeping. Our main idea is to splitting each parent-pixel into multiple sub-pixels. Based on the edge detection and image gradient, we separate edge into vertical and horizontal short rods. The main orientation of each rod is estimated. Then the intensities of sub-pixels around the rod are calculated according to the intensities and edge orientation of parent-pixels. The rest sub-pixels are estimated based on the principle of mean-keeping. The experimental results demonstrate that our method is effective in zigzagging artifacts reduction, de-blurring, and contrast enhancement. Xiangyi Fu, Fei Zhou 0001, Qingmin Liao |
ICIP | 2 |
| 2014 | Single image super-resolution via sparse KPCA and regressionabstractIn this paper, we present a new approach to single image super-resolution (SR). The basic idea is to learn a dictionary which can capture the high-order statistics of high-resolution (HR) images. This is of central importance in image SR application, since the high-order statistics play a significant role in the reconstruction of HR image structure. Kernel principal component analysis (KPCA) is used to learn such a dictionary. To reduce the time complexity of learning and testing for KPCA, a sparse solution is adopted. Meanwhile, kernel ridge regression is employed to relate the input low-resolution (LR) image patches and the HR coding coefficients. Experimental results show that the proposed method can effectively reconstruct image details and outperform state-of-the-art algorithms in both quantitative and visual comparisons. Tingrong Yuan, Wenming Yang, Fei Zhou 0001, Qingmin Liao |
ICIP | 3 |
| 2014 | Face hallucination via position-based dictionaries coding in kernel feature spaceabstractIn this paper, we present a new method to reconstruct a high-resolution (HR) face image from a low-resolution (LR) observation. Inspired by position-patch based face hallucination approach, we design position-based dictionaries to code image patches, and recovery HR patch using the coding coefficients as reconstruction weights. In order to capture nonlinear similarity of face features, we implicitly map the data into a high dimensional feature space. By applying kernel principal analysis (KPCA) on the mapped data in the high dimensional feature space, we can obtain reconstruction coefficients in a reduced subspace. Experimental results show that the proposed method can effectively reconstruct details of face images and outperform state-of-the-art algorithms in both quantitative and visual comparisons. Wenming Yang, Tingrong Yuan, Fei Zhou 0001, Qingmin Liao |
SMARTCOMP | 3 |
| 2014 | Super-resolution for facial image using multilateral affinity function
Fei Zhou 0001, Qingmin Liao |
Neurocomputing | 1 |
| 2014 | Comparative competitive coding for personal identification by using finger vein and finger dorsal texture fusion
Wenming Yang, Xiaola Huang, Fei Zhou 0001, Qingmin Liao |
Inf. Sci. | 3 |
| 2014 | Nonlocal Pixel Selection for Multisurface Fitting-Based Super-ResolutionabstractIn this paper, we address a super-resolution (SR) problem that constructs a high-resolution (HR) frame/image from a short sequence of low-resolution (LR) frames/images. It is well known that SR is a difficult problem, especially when the number of LR inputs is small. In particular, our previous work involving multisurface fitting-based SR exhibits relatively poor performance in the above case. To cope with this problem, we take advantage of nonlocal pixels to fit local surfaces. The pixels from nonlocal spatial-temporal positions are selected and weighted based on patch similarity and outlier removal. With this method, the fitted surfaces become more elaborate so that more details can be retrieved in SR results. Experiments demonstrate that the proposed method is very effective in producing HR frames through a small number of LR inputs when compared with some state-of-the-art methods. Fei Zhou 0001, Shutao Xia, Qingmin Liao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Joint sparse representation based cepstral-domain dereverberation for distant-talking speech recognitionabstractIn this paper we address reducing the mismatch between training and testing conditions for robust distant-talking speech recognition under realistic reverberant environments. It is well known that the distortions caused by reverberation, background noise, etc., are highly nonlinear in the cepstral domain. In this paper we propose to capture the complex relationships between clean and reverberant speech via joint dictionary learning. Given a test reverberant speech with a sequence of feature vectors we first find their sparse representations, and then estimate the underlying clean feature vectors using the dictionary of clean speech. Based on speech recognition experiments conducted under realistic reverberation conditions, the proposed method is shown to perform very well, resulting in an average relative improvement of 59.1% compared with the baseline front-ends. Weifeng Li 0001, Longbiao Wang, Fei Zhou 0001, Qingmin Liao |
ICASSP | 3 |
| 2013 | Iterative Super-Resolution for Facial Image by Local and Global Regression
Fei Zhou 0001, Wenming Yang, Qingmin Liao |
MMM (1) | 1 |
| 2013 | Feature Denoising Using Joint Sparse Representation for In-Car Speech RecognitionabstractWe address reducing the mismatch between training and testing conditions for hands-free in-car speech recognition. It is well known that the distortions caused by background noise, channel effects, etc., are highly nonlinear in the log-spectral or cepstral domain. This letter introduces a joint sparse representation (JSR) to estimate the underlying clean feature vector from a noisy feature vector. Performing a joint dictionary learning by sharing the same representation coefficients, the proposed method intends to capture the complex relationships (or mapping functions) between clean and noisy speech. Speech recognition experiments on realistic in-car data demonstrate that the proposed method shows excellent recognition performance with a relative improvement of 39.4% compared with the “baseline” frontends. Weifeng Li 0001, Yicong Zhou, Norman Poh, Fei Zhou 0001, Qingmin Liao |
IEEE Signal Process. Lett. | 4 |
| 2012 | A Coarse-to-Fine Subpixel Registration Method to Recover Local Perspective Deformation in the Application of Image Super-ResolutionabstractIn this paper, a coarse-to-fine framework is proposed to register accurately the local regions of interest (ROIs) of images with independent perspective motions by estimating their deformation parameters. A coarse registration approach based on control points (CPs) is presented to obtain the initial perspective parameters. This approach exploits two constraints to solve the problem with a very limited number of CPs. One is named the point-point-line topology constraint, and the other is named the color and intensity distribution of segment constraint. Both of the constraints describe the consistency between the reference and sensed images. To obtain a finer registration, we have converted the perspective deformation into affine deformations in local image patches so that affine refinements can be used readily. Then, the local affine parameters that have been refined are utilized to recover precise perspective parameters of a ROI. Moreover, the location and dimension selections of local image patches are discussed by mathematical demonstrations to avoid the aperture effect. Experiments on simulated data and real-world sequences demonstrate the accuracy and the robustness of the proposed method. The experimental results of image super-resolution are also provided, which show a possible practical application of our method. Fei Zhou 0001, Wenming Yang, Qingmin Liao |
IEEE Trans. Image Process. | 1 |
| 2012 | Interpolation-Based Image Super-Resolution Using Multisurface FittingabstractIn this paper, we propose a new interpolation-based method of image super-resolution reconstruction. The idea is using multisurface fitting to take full advantage of spatial structure information. Each site of low-resolution pixels is fitted with one surface, and the final estimation is made by fusing the multisampling values on these surfaces in the maximum a posteriori fashion. With this method, the reconstructed high-resolution images preserve image details effectively without any hypothesis on image prior. Furthermore, we extend our method to a more general noise model. Experimental results on the simulated and real-world data show the superiority of the proposed method in both quantitative and visual comparisons. Fei Zhou 0001, Wenming Yang, Qingmin Liao |
IEEE Trans. Image Process. | 1 |
| 2010 | Object Tracking and Local Appearance Capturing in a Remote Scene Video Surveillance System with Two Cameras
Wenming Yang, Fei Zhou 0001, Qingmin Liao |
MMM | 2 |