Wen Lu 0004

dblp:31/6140-4 · DBLP profile ↗
← Back
80ranked-venue papers
7as first author
30since 2021 · last 2026
0000-0002-8193-6016ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 26 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HVS-inspired blind image quality index with prominent perception learning and multi-level progressive integration
Taiyang Chen, Bo Hu 0008, Chunyi Li 0001, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001
Neurocomputing6
2026 LRAR: Luminance-ranking autoregressive for low-light image enhancement
Yuntai Liao, Zongfang Ma, Wen Lu 0004, Luze Jia, Qiguang Miao
Inf. Sci.4
2026 Multiscale Window Attention Channel Enhanced for Remote Sensing Image Super-Resolution
abstract
Transformer-based methods for remote sensing image super-resolution face challenges in reconstructing high-frequency textures due to the interference from large flat regions, such as farmlands and water bodies. To address these limitations, we propose a channel-enhanced multiscale window attention mechanism, which is designed to minimize the impact of flat regions on high-frequency area reconstruction while effectively utilizing the intrinsic multiscale features of remote sensing images.To better capture the multi-scale features of remote sensing images, we introduce a series of depthwise separable convolution kernels of varying sizes during the shallow feature extraction stage. Experimental results demonstrate that the proposed method achieves superior PSNR and SSIM scores across multiple remote sensing benchmark datasets and scaling factors, validating its effectiveness.
Jingfan Wang, Wen Lu 0004, Zeming Zhang, Zhaoyang Wang 0003, Zhe Li 0062
IEEE Geosci. Remote. Sens. Lett.2
2026 EyeSim-VQA: A Free-Energy-Guided Eye Simulation Framework for Video Quality Assessment
abstract
Modeling visual perception in a manner consistent with human subjective evaluation has become a central direction in both video quality assessment (VQA) and broader visual understanding tasks. While free-energy-guided self-repair mechanisms—reflecting human observational experience—have proven effective in image quality assessment, extending them to VQA remains non-trivial. In addition, biologically inspired paradigms such as holistic perception, local analysis, and gaze-driven scanning have achieved notable success in high-level vision tasks, yet their potential within the VQA context remains largely underexplored. To address these issues, we propose EyeSimVQA, a novel VQA framework that incorporates free-energy-based self-repair. It adopts a dual-branch architecture, with an aesthetic branch for global perceptual evaluation and a technical branch for fine-grained structural and semantic analysis. Each branch integrates specialized enhancement modules tailored to distinct visual inputs—resized full-frame images and patch-based fragments—to simulate adaptive repair behaviors. We also explore a principled strategy for incorporating high-level visual features without disrupting the original backbone. In addition, we design a biologically inspired prediction head that models sweeping gaze dynamics to better fuse global and local representations for quality prediction. Experiments on five public VQA benchmarks demonstrate that EyeSimVQA achieves competitive or superior performance compared to state-of-the-art methods, while offering improved interpretability through its biologically grounded design. Our code will be publicly available at https://github.com/handsomewzy/EyeSim-VQA.
Zhaoyang Wang 0003, Wen Lu 0004, Jie Li 0001, Lihuo He, Maoguo Gong, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 CognitionCapturer: Decoding Visual Stimuli from Human EEG Signal with Multimodal Information
abstract
Electroencephalogram (EEG) signals have attracted significant attention from researchers due to their non-invasive nature and high temporal sensitivity in decoding visual stimuli. However, most recent studies have focused solely on the relationship between EEG and image data pairs, neglecting the valuable "beyond-image-modality" information embedded in EEG signals. This results in the loss of critical multimodal information in EEG. To address the limitation, this paper proposes a unified framework that fully leverages multimodal data to represent EEG signals, named CognitionCapturer. Specifically, CognitionCapturer trains modality expert encoders for each modality to extract cross-modal information from the EEG modality. Then, it introduces a diffusion prior to map the EEG embedding space to the CLIP embedding space, followed by using a pretrained generative model, the proposed framework can reconstruct visual stimuli with high semantic and structural fidelity. Notably, the framework does not require any fine-tuning of the generative models and can be extended to incorporate more modalities. Through extensive experiments, we demonstrate that CognitionCapturer outperforms state-of-the-art methods both qualitatively and quantitatively.
Kaifan Zhang, Lihuo He, Wen Lu 0004, Di Wang 0011, Xinbo Gao 0001
AAAI4
2025 AGIAA-2K: A Fine-grained Dataset for Aesthetic and Alignment Evaluation of AI-Generated Images
abstract
With the advancement of AI-generated content technologies, AI-generated images (AGIs) have become increasingly influential in artistic creation and visual communication. However, the aesthetic quality of AGIs varies significantly due to technical limitations and the influence of user input, underscoring the urgent need for systematic aesthetic evaluation of AGIs. In addition, it is difficult to ensure the consistency of text-to-image, which compresses the application space of AGIs. To address these issues, a fine-grained dataset for Aesthetic and Alignment evaluation of AGIs (AGIAA-2K) is presented. This dataset contains 2,064 images generated using 172 well-designed prompts across six different AGI models, with each image annotated based on the subjective experiment. Then, the rationality of the dataset is verified by data analysis. Finally, the performances of the existing algorithms are evaluated in terms of image aesthetic assessment and text-to-image alignment of AGIs. The results demonstrate that these algorithms cannot effectively evaluate these two aspects. The AGIAA-2K is available at https://github.com/BoHu90/AGIAA-2K.
Bo Hu 0008, Nanxiang Li, Lihuo He, Wen Lu 0004, Leida Li, Xinbo Gao 0001
ICASSP4
2025 MACA-VQA: Quality Assessment of UGC Videos via Multi-level Distortion Adaptation and Spatiotemporal Cross-Attention Fusion
abstract
User-generated content (UGC) videos often exhibit complex distortions and diverse content, posing significant challenges for traditional video quality assessment (VQA) methods. Approaches that directly merge distortion and semantic information risk feature conflicts and the loss of details. In addition, simple concatenation of spatiotemporal features fails to capture vital interactions, limiting predictive accuracy. Motivated by these challenges, this paper proposes a Multi-level Distortion Adaptation and Spatiotemporal Cross-Attention Fusion framework for VQA, named MACA-VQA. Specifically, a novel multi-level adaptive strategy progressively incorporates distortion information into each Transformer layer of the CLIP model, enabling layer-wise fusion of semantic and distortion features. Furthermore, a newly introduced cross-attention fusion mechanism dynamically integrates spatiotemporal features, capturing complex, multidimensional interactions. Extensive experiments demonstrate that MACA-VQA achieves state-of-the-art performance on multiple public datasets, validating its effectiveness and robustness in both intra-dataset and inter-dataset scenarios. The source code is available at https://github.com/BoHu90/MACA-VQA
Bo Hu 0008, Yimeng Zhao, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001
ICME5
2025 Inferring normality from noised samples: Enhanced deep autoencoder with image denoising for anomaly detection
Xinbo Gao 0001, Wen Lu 0004, Jie Li 0001
Inf. Sci.3
2025 Blind Quality Assessment of Wide-Angle Videos Based on Deformation Representation Learning and Multi-Dimensional Feature Fusion
abstract
Wide-angle videos shot with short-focus lenses often exhibit deformation distortions, which poses significant challenges for video quality assessment (VQA). Although current VQA methods focus primarily on video content and distortion perception, there has been little explicit research on the impact of deformation characteristics on the perception of wide-angle video quality. To this end, this paper makes the first attempt to construct a novel wide-angle video quality assessment method based on deformation representation learning and multi-dimensional feature fusion, termed DRLMF. Specifically, we first analyze the deformation distribution characteristics of wide-angle videos based on the deformation camera model. Based on this, a three-stream video perception and assessment network is proposed. The first branch extracts global semantics using the image encoder of CLIP. The second branch introduces an effective deformation region selection strategy and proposes an interpretable deformation representation learning module. This module leverages the perception advantages of convolutional neural networks (CNNs) in local distortions and considers the correlation between patch size and distortion perception. The third branch extracts motion features using an action recognition network. Finally, an effective multi-dimensional feature fusion module is proposed to integrate more refined and richer semantic, deformation, and motion features. Extensive experiments on wide-angle VQA datasets and standard video datasets show that the DRLMF outperforms the state-of-the-arts in terms of prediction monotonicity and accuracy. The codes will be available at https://github.com/BoHu90/DRLMF.
Bo Hu 0008, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Active Learning for Object Detection With Vectorized Dual Pseudo Loss and Multiple Instance Offset Constraint
abstract
Existing active learning methods for object detection face challenges, such as the lack of ground truth labels for regression loss, insufficient representation of unlabeled instance samples information, and discrepancies in information quality between image-level and multiple anchor-level instances. To address these issues, we propose an active learning method for object detection with vectorized dual pseudo loss and multiple instance offset constraint. This method implements a two-stage framework. The first stage focuses on evaluating the information quality of detection images. We first pioneer a dual pseudo loss formulation that provides theoretically grounded regression loss estimation. The regression loss is calculated as the norm of the offset discrepancy loss vector between the enhanced and original base box vector, further constrained by the cosine value of the angle between the anchor box feature and regressor parameters vector. The distance entropy from the base box feature vector to each category's feature prototype vector is used as a weighting factor for the regression and classification information quality of instance samples. Subsequently, the second stage employs diversity-driven sampling on high-information images, leveraging instance-level cosine similarity to effectively remove redundant images. The proposed method outperforms state-of-the-art active learning approaches for object detection on PASCAL VOC and MS COCO datasets. Additionally, the proposed dual pseudo regression loss robustly captures regression information quality, demonstrating its effectiveness for active learning in object detection.
Jiasai Wu, Shuai Xiao 0001, Jiabao Wen, Qinggang Meng, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Cybern.6
2025 Difficulty-Guided Variant Degradation Learning for Blind Image Super-Resolution
abstract
Recent blind super-resolution (BSR) methods are explored to handle unknown degradations and achieve impressive performance. However, the prevailing assumption in most BSR methods is the spatial invariance of degradation kernels across the entire image, which leads to significant performance declines when faced with spatially variant degradations caused by object motion or defocusing. Additionally, these methods do not account for the human visual system's tendency to focus differently on areas of varying perceptual difficulty, as they uniformly process each pixel during reconstruction. To cope with these issues, we propose a difficulty-guided variant degradation learning network for BSR, named difficulty-guided degradation learning (DDL)-BSR, which explores the relationship between reconstruction difficulty and degradation estimation. Accordingly, the proposed DDL-BSR consists of three customized networks: reconstruction difficulty prediction (RDP), space-variant degradation estimation (SDE), and degradation and difficulty-informed reconstruction (DDR). Specifically, RDP learns the reconstruction difficulty with the proposed reconstruction-distance supervision. Then, SDE is designed to estimate space-variant degradation kernels according to the difficulty map. Finally, both degradation kernels and reconstruction difficulty are fed into DDR, which takes into account such two prior knowledge information to guide super-resolution (SR). Experimental analysis on various synthetic datasets demonstrates that DDL-BSR invariably surpasses state-of-the-art (SOTA) methods, producing SR images with enhanced realism and texture quality. Code is available at https://github.com/JiaWang0704/DDL-BSR.
Jiaxu Leng, Jia Wang 0036, Mengjingcheng Mo, Ji Gan, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 An Information-Assisted Deep Reinforcement Learning Path Planning Scheme for Dynamic and Unknown Underwater Environment
abstract
An autonomous underwater vehicle (AUV) has shown impressive potential and promising exploitation prospects in numerous marine missions. Among its various applications, the most essential prerequisite is path planning. Although considerable endeavors have been made, there are several limitations. A complete and realistic ocean simulation environment is critically needed. As most of the existing methods are based on mathematical models, they suffer from a large gap with reality. At the same time, the dynamic and unknown environment places high demands on robustness and generalization. In order to overcome these limitations, we propose an information-assisted reinforcement learning path planning scheme. First, it performs numerical modeling based on real ocean current observations to establish a complete simulation environment with the grid method, including 3-D terrain, dynamic currents, local information, and so on. Next, we propose an information compression (IC) scheme to trim the mutual information (MI) between reinforcement learning neural network layers to improve generalization. A proof based on information theory provides solid support for this. Moreover, for the dynamic characteristics of the marine environment, we elaborately design a confidence evaluator (CE), which evaluates the correlation between two adjacent frames of ocean currents to provide confidence for the action. The performance of our method has been evaluated and proven by numerical results, which demonstrate a fair sensitivity to ocean currents and high robustness and generalization to cope with the dynamic and unknown underwater environment.
Meng Xi 0001, Jiabao Wen, Zhengjian Li, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 General Deformable RoI Pooling and Semi-Decoupled Head for Object Detection
abstract
Object detection aims to classify interest objects within an image and pinpoint their positions using predicted rectangular bounding boxes. However, classification and localization tasks are heterogeneous, not only spatially misaligned but also differing in properties and feature requirements. Modern detectors commonly share the spatial region and detection head for both tasks, making them challenging to achieve optimal performance altogether, resulting in inconsistent accuracy. Specifically, the predicted bounding box may have higher classification confidence but lower localization quality, or vice versa. To tackle this issue, the spatial decoupling mechanism via general deformable RoI pooling is first proposed. This mechanism separately pursues the favorable regions for classification and localization, and subsequently extracts the corresponding features. Then, the semi-decoupled head is designed. Compared to the decoupled head that utilizes independent classification and localization networks, potentially leading to excessive decoupling and compromised detection performance, the semi-decoupled head enables the networks to mutually enhance each other while concentrating on their respective tasks. In addition, the semi-decoupled head also introduces a redundancy suppression module to filter out redundant task-irrelevant information of features extracted by separate networks and reinforce task-related information. By combining the spatial decoupling mechanism with the semi-decoupled head, the proposed detector achieves an impressive 43.7 AP in Faster R-CNN framework with ResNet-101 as backbone network. Without bells and whistles, extensive experimental results on the popular MS COCO dataset demonstrate that the proposed detector suppresses the baseline by a significant margin and outperforms some state-of-the-art detectors. Code is available athttps://github.com/HB-X/gdpool_semi_dehead.
Bo Han 0004, Lihuo He, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Multim.4
2024 Image Aesthetics Assessment Based on Hypernetwork of Emotion Fusion
abstract
Research in psychology demonstrates that visual features and semantic content can convey various emotions. Furthermore, studies have proved that image emotion and aesthetics are inextricably linked. During the image aesthetic assessment process (IAA), images elicit emotional responses from individuals, leading to emotional resonance and influencing the evaluation of images. This article proposes an image aesthetics assessment method based on hypernetwork of emotion fusion (HNEF). Our method incorporates the emotions depicted in images into the process of IAA. To accomplish this, we extract both aesthetic and emotional features from the images. Additionally, we employed the self-attention mechanism of the transformer to comprehensively investigate the intimate connection between aesthetics and emotion. Additionally, the hypernetwork is designed to establish perception rules governing the high-level semantic information in images. The experimental results validate the strong correlation between emotion and aesthetics. Furthermore, the proposed method exhibits a significantly competitive advantage when compared to existing methods on the Aesthetic Visual Analysis (AVA) dataset.
Guipeng Lan, Shuai Xiao 0001, Yanshuang Zhou, Jiabao Wen, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Multim.6
2023 A deep semi-dense compression network for reinforcement learning based on information theory
Jiabao Wen, Meng Xi 0001, Taiqiu Xiao, Wen Lu 0004, Xinbo Gao 0001
Neurocomputing6
2023 No-reference quality index of tone-mapped images based on authenticity, preservation, and scene expressiveness
Yang Zhao 0027, Shuai Xiao 0001, Wen Lu 0004, Xinbo Gao 0001
Signal Process.4
2023 MetaMP: Metalearning-Based Multipatch Image Aesthetics Assessment
abstract
Image aesthetics assessment (IAA) is a subjective and complex task. The aesthetics of different themes vary greatly in content and aesthetic results, whether they are in the same aesthetic community or not. In aesthetic evaluation tasks, the pretrained network with direct fine-tune may not be able to quickly adapt to tasks on various themes. This article introduces a metalearning-based multipatch (MetaMP) IAA method to adapt to various thematic tasks quickly. The network is trained based on metalearning to obtain content-oriented aesthetic expression. In addition, we design a complete-information patch selection scheme and a multipatch (MP) network to make the fine details fit the overall impression. Experimental results demonstrate the superiority of the proposed method in comparison with the state-of-the-art models based on aesthetic visual analysis (AVA) benchmark datasets. In addition, the evaluation of the dataset shows the effectiveness of our metalearning training model, which not only improves MetaMP assessment accuracy but also provides valuable guidance for network initialization of IAA.
Yanshuang Zhou, Yang Zhao 0027, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Cybern.4
2022 Magic Mirror by Bidirectional GAN: A Video Data Theft-Proof Framework
abstract
Video data is always seen as a kind of high-value data due to its salient features on rich and intuitive presentations, which also makes it an important objective of cyber data theft. Traditionally, people adapt various encryption algorithms to protect sensitive video data. However, such methods yield to intensive computations, complicated key management, and unexpected key leakages. In this paper, we fundamentally change the logic of video protections from stored raw data encryption to machine learning based video frame image learning and retrieval without raw data storage. By building inversions of Generative Adversarial Network (GAN), we firstly transform video’s raw data to a series of parameters in the latent space with raw data discarded. When retrieving, the trained GAN model will reassemble frame images on the basis of latent codes, which is protected by a Honey Encryption (HE) algorithm. Once the retrieving password is correct, the correct latent codes will be selected to generate the right image. Otherwise, wrong latent codes are chosen towards a misunderstanding figure to deceive attackers, therefore achieving a fine-grained video protection with lower storage costs. Our experiment results show that, the framework demonstrates excellent performance regarding security design, indistinguishability of generated fake image, reconstruction accuracy of original images, consistency of the framework, as well as storage cost reductions.
Wen Lu 0004, Kezhang Lin
ICC2
2022 Anomaly Warning: Learning and Memorizing Future Semantic Patterns for Unsupervised Ex-ante Potential Anomaly Prediction
abstract
Existing video anomaly detection methods typically utilize reconstruction or prediction error to detect anomalies in the current frame. However, these methods cannot predict ex-ante potential anomalies in future frames, which is imperative in real scenes. Inspired by the ex-ante prediction ability of humans, we propose an unsupervised Ex-ante Potential Anomaly Prediction Network (EPAP-Net), which learns to build a semantic pool to memorize the normal semantic patterns of future frames for indirect anomaly prediction. At the training time, the memorized patterns are encouraged to be discriminated through our Semantic Pool Building Module (SPBM) with the novel padding and updating strategies. Moreover, we present a novel Semantic Similarity Loss (SSLoss) at the feature level to maximize the semantic consistency of memorized items and corresponding future frames. Specially, to enhance the value of our work, we design a Multiple Frames Prediction module (MFP) to achieve anomaly prediction in future multiple frames. At the test time, we utilize the trained semantic pool instead of ground truth to evaluate the anomalies of future frames. Besides, to obtain better feature representations for our task, we introduce a novel Channel-selected Shift Encoder (CSE), which shifts channels along the temporal dimension between the input frames to capture motion information without generating redundant features. Experimental results demonstrate that the proposed EPAP-Net can effectively predict the potential anomalies in future frames and exhibit superior or competitive performance on video anomaly detection.
Jiaxu Leng, Mingpi Tan, Xinbo Gao 0001, Wen Lu 0004, Zongyi Xu
ACM Multimedia4
2022 QoEVMA'22: 2nd Workshop on Quality of Experience (QoE) in Visual Multimedia Applications
abstract
Nowadays, people spend dramatically more time on watching videos through different devices. The advanced hardware technology and network allow for the increasing demands of users viewing experience. Thus, enhancing the Quality of Experience of end-users in advanced multimedia is the ultimate goal of service providers, as good services would attract more consumers. Quality assessment is thus important. The second workshop on "Quality of Experience (QoE) in visual multimedia applications" (QoEVMA'22) focuses on the QoE assessment of any visual multimedia applications both subjectively and objectively. The topics include 1) QoE assessment on different visual multimedia applications, including VoD for movies, dramas, variety shows, UGC on social networks, live streaming videos for gaming/shopping/social, etc. 2) QoE assessment for different video formats in multimedia services, including 2D, stereoscopic 3D, High Dynamic Range (HDR), Augmented Reality (AR), Virtual Reality (VR), 360, Free-Viewpoint Video(FVV), etc. 3) Key performance indicators (KPI) analysis for QoE. This summary gives a brief overview of the workshop, which took place on October 14, 2022 in Lisbon, Portugal, as a half-day workshop. The complete QOEVMA'22 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3552469
Jing Li 0026, Patrick Le Callet, Xinbo Gao 0001, Zhi Li 0001, Wen Lu 0004, Junle Wang
ACM Multimedia5
2022 Systemic distortion analysis with deep distortion directed image quality assessment models
Xinbo Gao 0001, Wen Lu 0004, Jie Li 0001
Signal Process. Image Commun.3
2022 No Reference Quality Assessment for Screen Content Images Using Stacked Autoencoders in Pictorial and Textual Regions
abstract
Recently, the visual quality evaluation of screen content images (SCIs) has become an important and timely emerging research theme. This article presents an effective and novel blind quality evaluation metric for SCIs by using stacked autoencoders (SAE) based on pictorial and textual regions. Since the SCI consists of not only the pictorial area but also the textual area, the human visual system (HVS) is not equally sensitive to their different distortion types. First, the textual and pictorial regions can be obtained by dividing an input SCI via an SCI segmentation metric. Next, we extract quality-aware features from the textual region and pictorial region, respectively. Then, two different SAEs are trained via an unsupervised approach for quality-aware features that are extracted from these two regions. After the training procedure of the SAEs, the quality-aware features can evolve into more discriminative and meaningful features. Subsequently, the evolved features and their corresponding subjective scores are input into two regressors for training. Each regressor can obtain one output predictive score. Finally, the final perceptual quality score of a test SCI is computed by these two predicted scores via a weighted model. Experimental results on two public SCI-oriented databases have revealed that the proposed scheme can compare favorably with the existing blind image quality assessment metrics.
Yang Zhao 0027, Jiacheng Liu 0003, Bin Jiang 0003, Qinggang Meng, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Cybern.6
2021 Deep blind image quality assessment based on multiple instance regression
Xinbo Gao 0001, Wen Lu 0004, Jie Li 0001
Neurocomputing3
2021 MSCAN: Multimodal Self-and-Collaborative Attention Network for image aesthetic prediction tasks
Xiaodan Zhang 0005, Xinbo Gao 0001, Lihuo He, Wen Lu 0004
Neurocomputing4
2021 Video quality assessment with dense features and ranking pooling
Yu Zhang 0062, Lihuo He, Wen Lu 0004, Jie Li 0001, Xinbo Gao 0001
Neurocomputing3
2021 Staged-Learning: Assessing the Quality of Screen Content Images from Distortion Information
abstract
The small volume of the existing screen content images (SCIs) database with human ratings restricts the training processes of no-reference (NR) image quality assessment models based on traditional machine learning and deep learning. In this letter, we propose an NR model called the multi-task distortion-learning network to jointly analyse the distortion types and distortion degree of SCIs to be the prior knowledge for predicting the SCIs quality. Specifically, we first generate sufficient distorted SCIs labelled with the distortion type and degree, which does not need much effort to conduct subjective scoring experiments. Then, relying on these data, we pre-train a multi-task learning network to obtain strong prior knowledge about assessing the image quality. Finally, we further jointly train a quality assessment network with an attention module that simulates the mechanism of processing visual signals in the human eyes. The experimental results on the public SCIs databases show that the proposed model is competitive against other state-of-art approaches and achieves better consistency with the human vision system.
Zilin Bian, Yang Zhao 0027, Wen Lu 0004, Xinbo Gao 0001
IEEE Signal Process. Lett.4
2021 MTD-Net: Learning to Detect Deepfakes Images by Multi-Scale Texture Difference
abstract
With the rapid development of face manipulation technology, it is difficult for human eyes to distinguish fake face images. On the contrary, Convolutional Neural Network (CNN) discriminators can quickly reach high accuracy in identifying fake/real face images. In this study, we explore the behavior of CNN models in distinguish fake/real faces. We find multi-scale texture difference information plays an important role in face forgery detection. Motivated by the above observation, we propose a new Multi-scale Texture Difference model coined as MTD-Net for robust face forgery detection, which leverages central difference convolution (CDC) and atrous spatial pyramid pooling (ASPP). CDC combines the pixel intensity information and the pixel gradient information to give a stationary description of texture difference information. Simultaneously, based on the ASPP, multi-scale information fusion can keep the texture features from being destroyed. Experimental results on several databases, Faceforensics++, DeeperForensics-1.0, Celeb-DF and DFDC prove that our MTD-Net outperforms existing approaches. The MTD-Net is more robust to image distortion, e.g., JPEG compression and blur, which is urgently needed in the wild world.
Aiyun Li, Shuai Xiao 0001, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2021 Interpretable Detail-Fidelity Attention Network for Single Image Super-Resolution
abstract
Benefiting from the strong capabilities of deep CNNs for feature representation and nonlinear mapping, deep-learning-based methods have achieved excellent performance in single image super-resolution. However, most existing SR methods depend on the high capacity of networks that are initially designed for visual recognition, and rarely consider the initial intention of super-resolution for detail fidelity. To pursue this intention, there are two challenging issues that must be solved: (1) learning appropriate operators which is adaptive to the diverse characteristics of smoothes and details; (2) improving the ability of the model to preserve low-frequency smoothes and reconstruct high-frequency details. To solve these problems, we propose a purposeful and interpretable detail-fidelity attention network to progressively process these smoothes and details in a divide-and-conquer manner, which is a novel and specific prospect of image super-resolution for the purpose of improving detail fidelity. This proposed method updates the concept of blindly designing or using deep CNNs architectures for only feature representation in local receptive fields. In particular, we propose a Hessian filtering for interpretable high-profile feature representation for detail inference, along with a dilated encoder-decoder and a distribution alignment cell to improve the inferred Hessian features in a morphological manner and statistical manner respectively. Extensive experiments demonstrate that the proposed method achieves superior performance compared to the state-of-the-art methods both quantitatively and qualitatively. The code is available at github.com/YuanfeiHuang/DeFiAN.
Yuanfei Huang, Jie Li 0001, Xinbo Gao 0001, Yanting Hu, Wen Lu 0004
IEEE Trans. Image Process.5
2021 No-Reference Quality Assessment for Screen Content Images Using Visual Edge Model and AdaBoosting Neural Network
abstract
In this paper, a competitive no-reference metric is proposed to assess the perceptive quality of screen content images (SCIs), which uses the human visual edge model and AdaBoosting neural network. Inspired by the existing theory that the edge information which reflects the visual quality of SCI is effectively captured by the human visual difference of the Gaussian (DOG) model, we compute two types of multi-scale edge maps via the DOG operator firstly. Specifically, two types of edge maps contain contour and edge information respectively. Then after locally normalizing edge maps, L -moments distribution estimation is utilized to fit their DOG coefficients, and the fitted L -moments parameters can be regarded as edge features. Finally, to obtain the final perceptive quality score, we use an AdaBoosting back-propagation neural network (ABPNN) to map the quality-aware features to the perceptual quality score of SCIs. The reason why the ABPNN is regarded as the appropriate approach for the visual quality assessment of SCIs is that we abandon the regression network with a shallow structure, try a regression network with a deep architecture, and achieve a good generalization ability. The proposed method delivers highly competitive performance and shows high consistency with the human visual system (HVS) on the public SCI-oriented databases.
Zilin Bian, Jiacheng Liu 0003, Bin Jiang 0003, Wen Lu 0004, Xinbo Gao 0001, Houbing Song
IEEE Trans. Image Process.5
2021 Beyond Vision: A Multimodal Recurrent Attention Convolutional Neural Network for Unified Image Aesthetic Prediction Tasks
abstract
Over the past few years, image aesthetic prediction has attracted increasing attention because of its wide applications, such as image retrieval, photo album management and aesthetic-driven image enhancement. However, previous studies in this area only achieve limited success because 1) they primarily depend on visual features and ignore textual information. 2) they tend to focus equally on to each part of images and ignore the selective attention mechanism. This paper overcomes these limitations by proposing a novel multimodal recurrent attention convolutional neural network (MRACNN). More specifically, the MRACNN consists of two streams: the vision stream and the language stream. The former employs the recurrent attention network to tune out irrelevant information and focuses on some key regions to extract visual features. The latter utilizes the Text-CNN to capture the high-level semantics of user comments. Finally, a multimodal factorized bilinear (MFB) pooling approach is used to achieve effective fusion of textual and visual features. Extensive experiments demonstrate that the proposed MRACNN significantly outperforms state-of-the-art methods for unified aesthetic prediction tasks: (i) aesthetic quality classification; (ii) aesthetic score regression; and (iii) aesthetic score distribution prediction.
Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He, Jie Li 0001
IEEE Trans. Multim.3
2020 Learning More Accurate Features for Semantic Segmentation in CycleNet
Linzi Qu, Lihuo He, Junji Ke, Xinbo Gao 0001, Wen Lu 0004
ACCV (1)5
2020 Cross-layer Information Refining Network for Single Image Super-Resolution
Wen Lu 0004, Xiaopeng Sun 0001
ICPR2
2020 PHC-GAN: Physical Constraint Generative Adversarial Network for Single Image Dehazing
abstract
Recently, most existing single image dehazing methods adopt the physical scattering model to generate clear images. The model variables are often estimated by trainable neural networks. However, estimating the variables heavily rely on dataset that usually does not take into account of the physical process induced by the scattering model. In this scheme, error accumulation cannot be avoided when applying the end-to-end methods. In this paper, we propose a physical constraint generative adversarial network (PHC-GAN) for single image dehazing. The PHC-GAN is a physics aware model that leveraging the physical scattering process as an additional constraint. To the best of our knowledge, we are the first introducing physical constraint in learning an end-to-end image dehazing model. Our proposed model not only effectively reduce the error accumulation, but can be well adapted in complex and realistic natural scenes compared to the existing methods. In detail, we realize the physical constraint in terms of a double discriminator architecture. The self-attention module is also utilized to guarantee fast convergence. In experiments, quantitative and qualitative results on both synthetic and natural images demonstrate that PHC-GAN is superior to state-of-the-art dehazing methods.
Gang Long, Wen Lu 0004, Lin Zha
ICTAI2
2020 QoEVMA'20: 1st Workshop on Quality of Experience (QoE) in Visual Multimedia Applications
abstract
Nowadays, people spend dramatically more time on watching videos through different devices. The advanced hardware technology and network allow for the increasing demands of users viewing experience. Thus, enhancing the Quality of Experience of end-users in advanced multimedia is the ultimate goal of service providers, as good services would attract more consumers. Quality assessment is thus important. The first workshop on "Quality of Experience (QoE) in visual multimedia applications" (QoEVMA'20) focuses on the QoE assessment of any visual multimedia applications both subjectively and objectively. The topics include 1)QoE assessment on different visual multimedia applications, including VoD for movies, dramas, variety shows, UGC on social networks, live streaming videos for gaming/shopping/social, etc. 2)QoE assessment for different video formats in multimedia services, including 2D, stereoscopic 3D, High Dynamic Range (HDR), Augmented Reality (AR), Virtual Reality (VR), 360, Free-Viewpoint Video(FVV), etc. 3)Key performance indicators (KPI) analysis for QoE. This summary gives a brief overview of the workshop, which took place at October 16, 2020 in Seattle (U.S.), as a half-day workshop.
Xinbo Gao 0001, Patrick Le Callet, Jing Li 0026, Zhi Li 0001, Wen Lu 0004
ACM Multimedia5
2020 Deep multi-label learning for image distortion identification
Xinbo Gao 0001, Wen Lu 0004, Lihuo He
Signal Process.3
2020 No-reference image quality assessment based on neighborhood co-occurrence matrix
Ziheng Zhou 0004, Wen Lu 0004, Weiquan He
Signal Process. Image Commun.2
2020 No-Reference Quality Assessment of Stereoscopic Videos With Inter-Frame Cross on a Content-Rich Database
abstract
With the wide application of stereoscopic video technology, the quality of stereoscopic video has attracted people's attention. Objective stereoscopic video quality assessment (SVQA) is highly challenging, but essential, particularly the no-reference (NR) SVQA method, where reference information is not needed and a large number of samples are required for training and testing sets. However, as far as we know, there are only a few samples in the established stereo video database, which is unsuitable for NR quality assessment and seriously hampers the development of NR-SVQA method. For these difficulties that we encountered, we carry out a comprehensive subjective evaluation of stereoscopic video quality in our newly established TJU-SVQA databases that contain various contents, mixed resolution coding and symmetrically/asymmetrically distorted stereoscopic videos. Furthermore, we propose a new inter-frame cross map to predict the objective quality scores. We compare and analyze the performance of several state-of-the-art 2D and 3D quality evaluation methods on our new databases. The experimental results on our established databases and a public database demonstrate that the proposed method can robustly predict the quality of stereoscopic videos.
Yang Zhao 0027, Bin Jiang 0003, Qinggang Meng, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2020 A Deep Evaluator for Image Retargeting Quality by Geometrical and Contextual Interaction
abstract
An image is compressed or stretched during the multidevice displaying, which will have a very big impact on perception quality. In order to solve this problem, a variety of image retargeting methods have been proposed for the retargeting process. However, how to evaluate the results of different image retargeting is a very critical issue. In various application systems, the subjective evaluation method cannot be applied on a large scale. So we put this problem in the accurate objective-quality evaluation. Currently, most of the image retargeting quality assessment algorithms use simple regression methods as the last step to obtain the evaluation result, which are not corresponding with the perception simulation in the human vision system (HVS). In this paper, a deep quality evaluator for image retargeting based on the segmented stacked AutoEnCoder (SAE) is proposed. Through the help of regularization, the designed deep learning framework can solve the overfitting problem. The main contributions in this framework are to simulate the perception of retargeted images in HVS. Especially, it trains two separated SAE models based on geometrical shape and content matching. Then, the weighting schemes can be used to combine the obtained scores from two models. Experimental results in three well-known databases show that our method can achieve better performance than traditional methods in evaluating different image retargeting results.
Bin Jiang 0003, Qinggang Meng, Baihua Li, Wen Lu 0004
IEEE Trans. Cybern.5
2020 No-Reference Quality Evaluation of Stereoscopic Video Based on Spatio-Temporal Texture
abstract
Due to the wide application of stereoscopic display technology, stereoscopic video quality assessment (SVQA) is facing great challenges, but worthwhile. Stereoscopic videos contain a great deal of information, which involves not only the spatial domain but also the spatio-temporal domain. Motion in stereoscopic video plays a critical role in quality perception, while the existing SVQA methods rarely refer to motion factors, and the performance of these methods is restrained. In this article, a novel SVQA based on motion perception is introduced and its performance is superior to that of existing excellent methods. Particularly, to appropriately reduce the amount of data processing, we extract the key-frame sequences according to the influence of movement intensity on binocular visual quality perception. The binocular summation and difference operations are implemented on extracted sequences, and then spatial texture and spatio-temporal texture statistic measurement are extracted simultaneously with local binary patterns from three orthogonal planes (LBP-TOP). Experiments are implemented on two publicly available databases and the results demonstrate the effectiveness and robustness of our algorithm for various categories of distortion stereoscopic video pairs.
Yang Zhao 0027, Bin Jiang 0003, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Multim.4
2020 Precise Measurement of Position and Attitude Based on Convolutional Neural Network and Visual Correspondence Relationship
abstract
Accurate measurement of position and attitude information is particularly important. Traditional measurement methods generally require high-precision measurement equipment for analysis, leading to high costs and limited applicability. Vision-based measurement schemes need to solve complex visual relationships. With the extensive development of neural networks in related fields, it has become possible to apply them to the object position and attitude. In this paper, we propose an object pose measurement scheme based on convolutional neural network and we have successfully implemented end-to-end position and attitude detection. Furthermore, to effectively expand the measurement range and reduce the number of training samples, we demonstrated the independence of objects in each dimension and proposed subadded training programs. At the same time, we generated generating image encoder to guarantee the detection performance of the training model in practical applications.
Jiabao Man, Meng Xi 0001, Xinbo Gao 0001, Wen Lu 0004, Qinggang Meng
IEEE Trans. Neural Networks Learn. Syst.5
2020 Objective Video Quality Assessment Combining Transfer Learning With CNN
abstract
Nowadays, video quality assessment (VQA) is essential to video compression technology applied to video transmission and storage. However, small-scale video quality databases with imbalanced samples and low-level feature representations for distorted videos impede the development of VQA methods. In this paper, we propose a full-reference (FR) VQA metric integrating transfer learning with a convolutional neural network (CNN). First, we imitate the feature-based transfer learning framework to transfer the distorted images as the related domain, which enriches the distorted samples. Second, to extract high-level spatiotemporal features of the distorted videos, a six-layer CNN with the acknowledged learning ability is pretrained and finetuned by the common features of the distorted image blocks (IBs) and video blocks (VBs), respectively. Notably, the labels of the distorted IBs and VBs are predicted by the classic FR metrics. Finally, based on saliency maps and the entropy function, we conduct a pooling stage to obtain the quality scores of the distorted videos by weighting the block-level scores predicted by the trained CNN. In particular, we introduce a preprocessing and a postprocessing to reduce the impact of inaccurate labels predicted by the FR-VQA metric. Due to feature learning in the proposed framework, two kinds of experimental schemes including train-test iterative procedures on one database and tests on one database with training other databases are carried out. The experimental results demonstrate that the proposed method has high expansibility and is on a par with some state-of-the-art VQA metrics on two widely used VQA databases with various compression distortions.
Yu Zhang 0062, Xinbo Gao 0001, Lihuo He, Wen Lu 0004, Ran He 0002
IEEE Trans. Neural Networks Learn. Syst.4
2019 Non-Local Hierarchical Residual Network for Single Image Super-Resolution
abstract
Recently, deep convolutional neural networks (CNNs) have been demonstrated excellent performance on single image super-resolution (SISR). However, most deep learning based methods lack the ability to distinguish features in network. For image super-resolution, it is important to design an effective prior to learn the correlation of various feature. To solve this problem, we propose a non-local hierarchical residual network (NHRN) for SISR. Specifically, we introduce a non-local module to measure the self-similarity between each pixels in the feature map, and obtain a weight matrix guiding the deep network to find more precise relationship between L-R and HR images. Thus our method reconstruct images with a sharper edge. In addition, we employ group convolutions to build a hierarchical residual structure, which enable the network extract image features hierarchically. It can reduce the executive time while ensuring reconstruct image quality. Extensive experiments show that our NHRN achieves better accuracy and speed against state-of-the-art methods.
Furui Bai, Wen Lu 0004, Lin Zha, Xiaopeng Sun 0001, Ruoxuan Guan
ICIP2
2019 Improving Image Super-Resolution via Feature Re-Balancing Fusion
abstract
Recently, benefiting from the strong ability of feature representation, deep-learning-based methods have achieved excellent performance in single image super-resolution (SR). Furthermore, skip connection and feature fusion have been demonstrated to be a commendable strategy to deal with various features in different depth for informative reconstruction. Nevertheless, cross-layer features show diverse characteristic in detail representation, blindly fusion then introduces unavoidable interference of features. In this paper, we propose a novel feature fusion unit by utilizing alternative dilated convolutions for re-balancing diverse cross-layer features, named Feature Re-Balancing Fusion Network (RBFNet), which is theoretically and experimentally demonstrated to be robust to the interference in feature fusion for SR. Extensive experiments show that the proposed method achieves excellent performances quantitatively and qualitatively against the state-of-the-art methods.
Yuanfei Huang, Jie Li 0001, Xinbo Gao 0001, Wen Lu 0004, Yanting Hu
ICME4
2019 Distilling with Residual Network for Single Image Super Resolution
abstract
Recently, the deep convolutional neural network (CNN) has made remarkable progress in single image super resolution(SISR). However, blindly using the residual structure and dense structure to extract features from LR images, can cause the network to be bloated and difficult to train. To address these problems, we propose a simple and efficient distilling with residual network(DRN) for SISR. In detail, we propose residual distilling block(RDB) containing two branches, while one branch performs a residual operation and the other branch distills effective information. To further improve efficiency, we design residual distilling group(RDG) by stacking some RDBs and one long skip connection, which can effectively extract local features and fuse them with global features. These efficient features beneficially contribute to image reconstruction. Experiments on benchmark datasets demonstrate that our DRN is superior to the state-of-the-art methods, specifically has a better trade-off between performance and model size.
Xiaopeng Sun 0001, Wen Lu 0004, Rui Wang 0173, Furui Bai
ICME2
2019 Beauty Aware Network: An Unsupervised Method for Makeup Product Retrieval
abstract
Makeup product retrieval has gained more and more attention for its wide application prospects. However, the challenging problem is that the dataset crawled from Internet doesn't have annotated labels. Therefore, existing methods are unable to obtain well-trained networks. To solve this problem, this paper proposes a trainable network named Beauty Aware Network (BAN) for makeup product retrieval. The core of proposed method is using an unsupervised cluster method to train the beauty classification network. And then a covariance pooling layer is introduced to leverage the statistical information. Finally, a multi-layer fusion strategy is used to capture informative clues in images. The proposed method can get simpler but more efficient features for beauty product retrieval with less computation cost. The experiments conduct on Perfect-500k dataset which has more than half-million images. The results demonstrate the effectiveness of beauty aware network by competitive performance.
Linzi Qu, Lihuo He, Wen Lu 0004, Xinbo Gao 0001
ACM Multimedia4
2019 No-Reference Image Quality Assessment via Multi-order Perception Similarity
Ziheng Zhou 0004, Wen Lu 0004, Shishuai Han
PRCV (2)2
2019 Densely convolutional attention network for image super-resolution
Furui Bai, Wen Lu 0004, Yuanfei Huang, Lin Zha
Neurocomputing2
2019 A spatiotemporal model of video quality assessment via 3D gradient differencing
Wen Lu 0004, Ran He 0002, Changcheng Jia, Xinbo Gao 0001
Inf. Sci.1
2019 Stereoscopic image quality assessment combining statistical features and binocular theory
Huifang Xu, Yang Zhao 0027, Hehan Liu, Wen Lu 0004
Pattern Recognit. Lett.5
2019 Fusion global and local deep representations with neural attention for aesthetic quality assessment
abstract
In recent years, deep-learning based aesthetics assessment methods have shown promising results. However, existing methods can only achieve limited success because 1) most of the methods take one fixed-size patch as the training example, which loses the fine grained details and the holistic layout information, and 2) most of the methods ignore ordinal issues in image aesthetic assessment, i.e. image scored 5.3 is more likely to be in the high quality class than image scored 4.5. To address these challenges, we presents a novel convolutional networks with two branches to encode global and local features . The first branch not only captures the spatial layout information but also feedbacks the top-down neural attention. The second branch selects the important attended region to extract the fine details features. A sobel-based attention layer is integrated with the second branch to enhance fine details encoding. Regarding the second problem, we combine the strength of classification approach and regression approach by a multi-task learning framework. Extensive experiments on challenging Aesthetic and Visual Analysis (AVA) dataset and Photo.net dataset indicate the effectiveness of the proposed method.
Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He
Signal Process. Image Commun.3
2019 Blind Video Quality Assessment With Weakly Supervised Learning and Resampling Strategy
abstract
Due to the 3D spatiotemporal regularities of natural videos and small-scale video quality databases, effective objective video quality assessment (VQA) metrics are difficult to obtain but highly desirable. In this paper, we propose a general-purpose no-reference VQA framework that is based on weakly supervised learning with a convolutional neural network (CNN) and a resampling strategy. First, an eight-layer CNN is trained by weakly supervised learning to construct the relationship between the deformations of the 3D discrete cosine transform of video blocks and the corresponding weak labels judged by a full-reference (FR) VQA metric. Thus, the CNN obtains the quality assessment capacity converted from the FR-VQA metric, and the effective features of the distorted videos can be extracted through the trained network. Then, we map the frequency histogram calculated from the quality score vectors predicted by the trained network onto the perceptual quality. Especially, to improve the performance of the mapping function, we transfer the frequency histogram of the distorted images and videos to resample the training set. The experiments are carried out on several widely used VQA databases. The experimental results demonstrate that the proposed method is on a par with some state-of-the-art VQA metrics and has promising robustness.
Yu Zhang 0062, Xinbo Gao 0001, Lihuo He, Wen Lu 0004, Ran He 0002
IEEE Trans. Circuits Syst. Video Technol.4
2019 A Gated Peripheral-Foveal Convolutional Neural Network for Unified Image Aesthetic Prediction
abstract
Learning fine-grained details is a key issue in image aesthetic assessment. Most of the previous methods extract the fine-grained details via random cropping strategy, which may undermine the integrity of semantic information. Extensive studies show that humans perceive fine-grained details with a mixture of foveal vision and peripheral vision. Fovea has the highest possible visual acuity and is responsible for seeing the details. The peripheral vision is used for perceiving the broad spatial scene and selecting the attended regions for the fovea. Inspired by these observations, we propose a gated peripheral-foveal convolutional neural network. It is a dedicated double-subnet neural network (i.e., a peripheral subnet and a foveal subnet). The former aims to mimic the functions of peripheral vision to encode the holistic information and provide the attended regions. The latter aims to extract fine-grained features on these key regions. Considering that the peripheral vision and foveal vision play different roles in processing different visual stimuli, we further employ a gated information fusion network to weigh their contributions. The weights are determined through the fully connected layers followed by a sigmoid function. We conduct comprehensive experiments on the standard Aesthetic Visual Analysis (AVA) dataset and Photo.net dataset for unified aesthetic prediction tasks: 1) aesthetic quality classification; 2) aesthetic score regression; and 3) aesthetic score distribution prediction. The experimental results demonstrate the effectiveness of the proposed method.
Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He
IEEE Trans. Multim.3
2018 Blind Image Quality Assessment Based on Visuo-Spatial Series Statistics
abstract
Existing blind image quality assessment (BIQA) methods based on statistics attach limited attention to the relative position of pixels. Features in these BIQA methods are too flimsy to characterize quite a few distortions with strong locality or complexity. However, psychological studies have shown that according to the relative position within visual field, the cognitive system generates visuo-spatial serial memory used for cognitive tasks, e.g., subjective image quality assessment. Inspired by the visuo-spatial series generated by human visual system (HVS), we propose a BIQA method based on imitation Visuo-spatial Series Statistics (VSS). The proposed method simulates visual system to construct visuo-spatial series based on the relative position of pixels, and use statistical features of visuo-spatial series to predict image quality. Extensive experiments demonstrate the proposed method has a superior performance compared to the state-of-the-art BIQA methods.
Ziheng Zhou 0004, Wen Lu 0004, Lihuo He, Xinbo Gao 0001
ICASSP2
2018 Single Image Super Resolution Based on Deep Residual Network via Lateral Modules
abstract
Recently, convolutional neural networks have demonstrated high-quality reconstruction for single image super resolution (SISR). In this paper, we propose a Deep Residual Network via lateral modules (DRNLM), DRNLM is the structure with lateral modules, progressive and symmetric residual blocks (convolutional residual blocks and deconvolutional residual blocks). First, DRNLM introduces lateral modules, which are used to transmit low-level features (coarse residue) into high-level features (fine residue) effectively, thus finer residue can be obtained for better image reconstruction. Second, considering more channels can stack more details, progressive channels that vary from 64 to 256 are utilized in DRNLM through residual blocks. Third, symmetric residual blocks have same dimensions of input and output, which can ensure the gradient ranging within certain limits when the network goes deeper. Extensive experiments demonstrate that the proposed method outperforms the existing methods in accuracy and visual impression.
Rui Wang 0173, Wen Lu 0004, Yuanfei Huang, Xinbo Gao 0001, Lihuo He
ICIP2
2018 Spatiotemporal Masking for Objective Video Quality Assessment
Ran He 0002, Wen Lu 0004, Yu Zhang 0062, Xinbo Gao 0001, Lihuo He
PRCV (1)2
2018 Dominant vanishing point detection in the wild with application in composition analysis
Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He, Qi Liu 0054
Neurocomputing3
2018 Sparse representation based stereoscopic image quality assessment accounting for perceptual cognitive process
Bin Jiang 0003, Yafang Wang, Wen Lu 0004, Qinggang Meng
Inf. Sci.4
2018 No reference quality evaluation for screen content images considering texture feature based on sparse representation
Jiacheng Liu 0003, Bin Jiang 0003, Wen Lu 0004
Signal Process.4
2018 Single Image Super-Resolution via Multiple Mixture Prior Models
abstract
Example learning-based single image super-resolution (SR) is a promising method for reconstructing a high-resolution (HR) image from a single-input low-resolution (LR) image. Lots of popular SR approaches are more likely either time-or space-intensive, which limit their practical applications. Hence, some research has focused on a subspace view and delivered state-of-the-art results. In this paper, we utilize an effective way with mixture prior models to transform the large nonlinear feature space of LR images into a group of linear subspaces in the training phase. In particular, we first partition image patches into several groups by a novel selective patch processing method based on difference curvature of LR patches, and then learning the mixture prior models in each group. Moreover, different prior distributions have various effectiveness in SR, and in this case, we find that student-t prior shows stronger performance than the well-known Gaussian prior. In the testing phase, we adopt the learned multiple mixture prior models to map the input LR features into the appropriate subspace, and finally reconstruct the corresponding HR image in a novel mixed matching way. Experimental results indicate that the proposed approach is both quantitatively and qualitatively superior to some state-of-the-art SR methods.
Yuanfei Huang, Jie Li 0001, Xinbo Gao 0001, Lihuo He, Wen Lu 0004
IEEE Trans. Image Process.5
2018 Single Image Dehazing With Depth-Aware Non-Local Total Variation Regularization
abstract
Single image dehazing can benefit many computer vision applications hence has attracted much more attention in recent years. However, it still remains a challenging task due to its double uncertainty of scene transmission and scene radiance. The existing image dehazing methods usually impair edges in the estimated transmission which leads to halo effects in the dehazing results. Besides, most existing methods suffer from noise and artifacts amplification in dense haze region after dehazing. To address these challenges, we propose a transmission adaptive regularized image recovery method for high quality single image dehazing. An initial transmission map is first obtained by a boundary constraint on the haze model. Then it is refined by applying a non-local total variation (NLTV) regularization to keep depth structures while smoothing excessive details. Noticing that the artifacts amplification effect depends on scene transmission, a transmission adaptive regularized recovery method based on NLTV is proposed to simultaneously suppress visual artifacts and preserve image details in the final dehazing result. An efficient alternating optimization algorithm is also proposed to solve the regularization model. Thorough experimental results demonstrate that the proposed method can effectively suppress visual artifacts for degraded hazy images, and yields high-quality results comparative to the state-of-the-art dehazing methods both quantitatively and qualitatively.
Qi Liu 0054, Xinbo Gao 0001, Lihuo He, Wen Lu 0004
IEEE Trans. Image Process.4
2017 Video quality assessment by compact representation of energy in 3D-DCT domain
Lihuo He, Wen Lu 0004, Changcheng Jia
Neurocomputing2
2017 Single image super resolution based on sparse domain selection
Wen Lu 0004, Huxing Sun, Rui Wang 0173, Lihuo He, Ming-Jong Jou, Shensian Syu, JiShiang Li
Neurocomputing1
2017 A no-reference optical flow-based quality evaluator for stereoscopic videos in curvelet domain
Huanling Wang, Wen Lu 0004, Baihua Li, Atta Badii, Qinggang Meng
Inf. Sci.3
2017 Haze removal for a single visible remote sensing image
Qi Liu 0054, Xinbo Gao 0001, Lihuo He, Wen Lu 0004
Signal Process.4
2016 Fast image quality assessment via supervised iterative quantization method
Lihuo He, Di Wang 0011, Qi Liu 0054, Wen Lu 0004
Neurocomputing4
2016 On combining visual perception and color structure based image quality assessment
Wen Lu 0004, Tianjiao Xu, Yuling Ren, Lihuo He
Neurocomputing1
2016 Quality assessment metric of stereo images considering cyclopean integration and visual saliency
Yafang Wang, Baihua Li, Wen Lu 0004, Qinggang Meng, Zhihan Lyu, Dezong Zhao, Zhiqun Gao
Inf. Sci.4
2016 Statistical modeling in the shearlet domain for blind image quality assessment
Wen Lu 0004, Tianjiao Xu, Yuling Ren, Lihuo He
Multim. Tools Appl.1
2010 Spatio-temporal salience based video quality assessment
abstract
It is important to design an effective and efficient objective metric of the video quality in video processing areas. The most reliable way is subjective evaluation, thus the most reasonable objective metric should adequately consider characteristics of the human visual system (HVS). Visual attention (VA) is one of the essential visual phenomena of HVS, the realization of which relies on the saliency of visual field. Moreover, the saliency of visual field has a great influence on recognition, memorization and subjective evaluation of the image. This paper explores the saliency of visual field for objective quality assessment of videos. The proposed method first uses the VA model to obtain visual saliency map of the distorted video, including color, intensity and motion. Then the salient map is used to weight a structural similarity map between the original and distorted videos to get the final value of the video quality. Experimental results prove that the proposed method achieves a good correlation with subjective valuation.
Xinbo Gao 0001, Wen Lu 0004, Dacheng Tao, Xuelong Li 0001
SMC3
2010 A novel image quality metric based on morphological component analysis
abstract
Due to that human eye has different perceptual characteristics for different morphological components, so a novel image quality metric is proposed by incorporating morphological component analysis (MCA) and human visual system (HVS), which is capable of assessing the image with different types of distortion. Firstly, reference and distorted images are decomposed into texture and cartoon components by MCA respectively. Then these components are changed into perceptual features by just noticeable difference (JND) which integrates masking features, luminance adaptation and contrast sensitive function (CSF). Finally, the difference between reference and distorted images' perceptual features is quantified using a pooling strategy, and then the final result of the image quality is obtained. Experimental results demonstrate that the performance of the metric prevail over some existing methods on LIVE database II.
Xuelong Li 0001, Lihuo He, Wen Lu 0004, Xinbo Gao 0001, Dacheng Tao
SMC3
2010 An image quality assessment metric with no reference using hidden Markov tree model
abstract
No reference (NR) method is the most difficult issue of image quality assessment (IQA), which does not need the original image or its features as reference and only depends on the statistical law of the natural images. So, the NR-IQA is a high -level evaluation for image quality and simulates the complicated subjective process of human beings. This paper presents a NR-IQA metric based on Hidden Markov Tree (HMT) model. First, the HMT is utilized to model natural images, and the statistical properties of the model parameters are analyzed to mimic variation of image degradation. Then, by estimating the deviation degree of the parameters from the statistical law the distortion metric is constructed. Experimental results show that the proposed image quality assessment model is consistent well with the subjective evaluation results, and outperforms the existing models on difference distortions.
Fei Gao 0006, Xinbo Gao 0001, Wen Lu 0004, Dacheng Tao, Xuelong Li 0001
VCIP3
2010 Image quality assessment and human visual system
abstract
This paper summaries the state-of-the-art of image quality assessment (IQA) and human visual system (HVS). IQA provides an objective index or real value to measure the quality of the specified image. Since human beings are the ultimate receivers of visual information in practical applications, the most reliable IQA is to build a computational model to mimic the HVS. According to the properties and cognitive mechanism of the HVS, the available HVS-based IQA methods can be divided into two categories, i.e., bionics methods and engineering methods. This paper briefly introduces the basic theories and development histories of the above two kinds of HVS-based IQA methods. Finally, some promising research issues are pointed out in the end of the paper.
Xinbo Gao 0001, Wen Lu 0004, Dacheng Tao, Xuelong Li 0001
VCIP2
2010 A new quality metric for compressed images based on DDCT
abstract
As the performance-indicator of the image processing algorithms or systems, image quality assessment (IQA) has attracted the attention of many researchers. Aiming to the widely used compression standards, JPEG and JPEG2000, we propose a new no reference (NR) metric for compressed images to do IQA. This metric exploits the causes of distortion by JPEG and JPEG2000, employs the directional discrete cosine transform (DDCT) to obtain the detail and direction information of the images and incorporates with the visual perception to obtain the image quality index. Experimental results show that the proposed metric not only has outstanding performance on JPEG and JPEG2000 images, but also applicable to other types of artifacts.
Wen Lu 0004, Jing Li 0026, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001
VCIP1
2010 No-reference image quality assessment in contourlet domain
Wen Lu 0004, Dacheng Tao, Yuan Yuan 0001, Xinbo Gao 0001
Neurocomputing1
2009 A natural image quality evaluation metric
Xuelong Li 0001, Dacheng Tao, Xinbo Gao 0001, Wen Lu 0004
Signal Process.4
2009 Image Quality Assessment Based on Multiscale Geometric Analysis
abstract
Reduced-reference (RR) image quality assessment (IQA) has been recognized as an effective and efficient way to predict the visual quality of distorted images. The current standard is the wavelet-domain natural image statistics model (WNISM), which applies the Kullback-Leibler divergence between the marginal distributions of wavelet coefficients of the reference and distorted images to measure the image distortion. However, WNISM fails to consider the statistical correlations of wavelet coefficients in different subbands and the visual response characteristics of the mammalian cortical simple cells. In addition, wavelet transforms are optimal greedy approximations to extract singularity structures, so they fail to explicitly extract the image geometric information, e.g., lines and curves. Finally, wavelet coefficients are dense for smooth image edge contours. In this paper, to target the aforementioned problems in IQA, we develop a novel framework for IQA to mimic the human visual system (HVS) by incorporating the merits from multiscale geometric analysis (MGA), contrast sensitivity function (CSF), and the Weber's law of just noticeable difference (JND). In the proposed framework, MGA is utilized to decompose images and then extract features to mimic the multichannel structure of HVS. Additionally, MGA offers a series of transforms including wavelet, curvelet, bandelet, contourlet, wavelet-based contourlet transform (WBCT), and hybrid wavelets and directional filter banks (HWD), and different transforms capture different types of image geometric information. CSF is applied to weight coefficients obtained by MGA to simulate the appearance of images to observers by taking into account many of the nonlinearities inherent in HVS. JND is finally introduced to produce a noticeable variation in sensory experience. Thorough empirical studies are carried out upon the LIVE database against subjective mean opinion score (MOS) and demonstrate that 1) the proposed framework has good consistency with subjective perception values and the objective assessment results can well reflect the visual quality of images, 2) different transforms in MGA under the new framework perform better than the standard WNISM and some of them even perform better than the standard full-reference IQA model, i.e., the mean structural similarity index, and 3) HWD performs best among all transforms in MGA under the framework.
Xinbo Gao 0001, Wen Lu 0004, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Image Process.2
2009 Reduced-Reference IQA in Contourlet Domain
abstract
The human visual system (HVS) provides a suitable cue for image quality assessment (IQA). In this paper, we develop a novel reduced-reference (RR) IQA scheme by incorporating the merits from the contourlet transform, contrast sensitivity function (CSF), and Weber's law of just noticeable difference (JND). In this scheme, the contourlet transform is utilized to decompose images and then extract features to mimic the multichannel structure of HVS. CSF is applied to weight coefficients obtained by the contourlet transform to simulate the appearance of images to observers by taking into account many of the nonlinearities inherent in HVS. JND is finally introduced to produce a noticeable variation in sensory experience. Thorough empirical studies are carried out upon the Laboratory for Image and Video Engineering database against the subjective mean opinion score and demonstrate that the proposed framework has good consistency with subjective perception values and the objective assessment results can well reflect the visual quality of images.
Dacheng Tao, Xuelong Li 0001, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Syst. Man Cybern. Part B3
2008 An image quality assessment metric based contourlet
abstract
In reduced-reference (RR) image quality assessment (IQA), the visual quality of distorted images is evaluated with only partial information extracted from original images. In this paper, by considering the information of textures and directions during image distortion, we propose a new reduced- reference IQA metric to calculate the diversifications based on contourlet transform Experimental results illustrate that even with low data rate, the presented metric has still good consistency with the subjective perception.
Wen Lu 0004, Xinbo Gao 0001, Xuelong Li 0001, Dacheng Tao
ICIP1
2008 Frequency structure analysis for IQA
abstract
Over the past years, research and applications of image quality assessment have attracted increasing attention. Basically, naked human eyes are the final receivers of an image and human visual system is able to extract structural information from the viewing field with high adaptation. Different frequency components of an image play different roles on image semantic contents, and the extraction of the visual information is also crucial in computerized image quality assessment systems. In this paper, an image quality assessment metrics is proposed based on wavelet structure and human perception. With the proposed metric, structural similarity is well expanded from pixel-wise to frequency field. Experimental results illustrate that the proposed metric gives good consistency with subjective assessment results of naked human eyes, i.e., it fits well the perception of human visual system.
Xuelong Li 0001, Wen Lu 0004, Dacheng Tao, Xinbo Gao 0001
SMC2
2008 Wavelet-based contourlet in quality evaluation of digital images
Xinbo Gao 0001, Wen Lu 0004, Xuelong Li 0001, Dacheng Tao
Neurocomputing2